[Data Engineering] [Production Released]
Trench Group · Global Industrial Leader
Global Cloud Migration & Establishment of a Hub-and-Spoke Azure Infrastructure
Context
Greenfield development of a global 5-region Hub-and-Spoke Azure Data Platform with strict compliance requirements.
Our Contribution / Solution
- Multi-tenant 5-region Hub-and-Spoke architecture on Azure with regional Databricks workspaces and Unity Catalog
- Highly secure network topologies (Private Endpoints, Azure Firewall, Hub VNETs) for GDPR, USMCA, and China CSL/DSL
- Entire infrastructure deployed as standardized Terraform (IaC), orchestrated via Databricks Asset Bundles (DABs)
- AI agents for Jira management, repository screening during pull requests, and Gantt charts via secure MCP servers
Result
- Delivery of a fully standardized Terraform MVP in just 8 weeks
- Significant reduction in infrastructure costs through consolidation of isolated factory workspaces into regional shared hubs
- An absolutely compliance-secure foundation for all future global AI initiatives
- [Azure]
- [Databricks]
- [Terraform]
- [Unity Catalog]
- [dbt]
- [MCP]
- [Industrial & High-Tech]
Read Case Study → [Data Engineering] [Production Released]
Denner AG · Swiss Food Retail
Stabilization and Automation of a Microsoft Fabric Data Platform
Context
Stabilization and performance optimization of the central Microsoft Fabric data platform.
Our Contribution / Solution
- Full architectural responsibility for over 1,200 dbt models within a Data Mesh and Medallion architecture
- Design and development of a reusable Python library (lib_pegasus) for the complete automation of recurring data engineering tasks
- Seamless, context-rich logging mechanisms and dbt tests for improved failure diagnosis in production
Result
- Development of a bespoke Python library for the full automation of data engineering tasks
- Substantial increase in development velocity across the entire internal team through standard libraries
- Complete stabilization of mission-critical ETL processes and elimination of all runtime-critical performance bottlenecks under massive data volumes
- [Microsoft Fabric]
- [dbt]
- [Python]
- [Retail]
- [Retail]
Read Case Study → [Data Engineering] [Production Released]
Carl Zeiss Vision · Precision Optics
Modernization and Migration of Lens Production Data into an Azure Lakehouse
Context
Decomposition of a complex legacy reporting infrastructure into a scalable Azure Lakehouse.
Our Contribution / Solution
- Highly scalable Data Lakehouse based on Azure SQL, Databricks, and Delta Lake
- Delta Live Tables (DLT) and Apache Kafka for zero-latency processing of incoming sensor data (REST, MongoDB, Blob Storage)
- Refactoring of historical SSIS pipelines and complex T-SQL logic into scalable PySpark and Spark SQL workflows
- AI-supported search platform (Azure OpenAI + Azure AI Search) based on structured metadata
Result
- Reduction of ETL data processing times by over 80% and lowering of operational costs by 30%
- Decrease in data processing time by 80% (from 2 hours down to 15–20 minutes)
- Reduction of operational cloud operating costs by 30%
- Increase in overall system scalability by 400%
- [Azure]
- [Databricks]
- [Delta Lake]
- [Kafka]
- [PySpark]
- [Industrial & High-Tech]
Read Case Study → [Data Modeling] [Production Released]
dm-drogerie markt / dmTECH · Retail
Development of a High-Performance Snowflake ETL Solution for Store and Project Data
Context
Development of a GCP-based Snowflake solution for automated ERP and Planisware data streams.
Our Contribution / Solution
- Scalable Snowflake Data Warehouse on the Google Cloud Platform (GCP) following Medallion architecture principles
- Highly efficient interface to the Planisware data source utilizing Denodo virtualization layers
- Robust data pipelines utilizing Apache Airflow and Snowpark Python (SCD and CDC logic)
- CI/CD infrastructure based on GitLab and Terraform, automated deployment on Kubernetes clusters (AKS)
Result
- Reduction of data loading times from 160 down to 10 minutes
- Shortening of ETL load times from a previous 160 minutes to merely 10 minutes
- Near real-time data provision for MicroStrategy reporting across over 500 dm retail stores
- [Snowflake]
- [GCP]
- [Denodo]
- [Airflow]
- [Snowpark]
- [Retail]
Read Case Study → [Data Engineering] [Production Released]
E.ON · Energy Sector
Automation of Real-Time Data Pipelines for Dynamic Pricing Systems
Context
Data processing automation for dynamic pricing within the European energy market.
Our Contribution / Solution
- Highly scalable and resilient ELT pipelines based on Snowflake and Azure
- dbt development model combined with Dagster for consistent transformations
- Fully automated generation of Data Lineage to increase comprehensibility for internal business analysts
Result
- Establishment of fully automated dbt Data Lineage and reduction of the data error rate by 15%
- Decrease of erroneous data records within the production pipelines by 15%
- Reduction of development time for new data pipelines by 50%
- [Snowflake]
- [dbt]
- [Dagster]
- [Azure]
- [Grafana]
- [Energy]
Read Case Study → [Data Engineering] [Production Released]
Encavis AG · Renewable Energy
IoT Real-Time Pipelines for Renewable Energy Installations
Context
IoT real-time data processing and time series optimization utilizing Prefect and Snowflake.
Our Contribution / Solution
- State-of-the-art, hybrid Data Lakehouse architecture with Snowflake accommodating structured and unstructured time series data
- Continuous data streams via Snowpipe directly into the data platform
- Bespoke ingestion scripts in Python employing Prefect and dbt to orchestrate parallel data processes (multithreading)
Result
- Significant acceleration in the acquisition of solar installation sensor data through multithreading data pipelines
- Substantial performance enhancement in processing complex, time-critical sensor data via multithreaded execution
- Stable real-time data streams enabling faultless ad-hoc analytics and yield monitoring
- [Snowflake]
- [Snowpipe]
- [Prefect]
- [dbt]
- [Python]
- [Energy]
Read Case Study →