Data Engineer
Azure Data Factory
Azure Data Platform
Azure Platform
Big Data
Bigdata
Data
Data Analysis
Data Architecture
Data Engineer
Data Factory
Data Factory Azure
Data Integration
Data Pipeline
Data Platform
Data Processing
Data Warehouse
Database
Databases
ETL
Informatica
Information Technology (IT)
Integration
Microsoft Azure
Microsoft Fabric
Programming
Programming Language
Programming Languages
Pyspark
Rag Architectures
Spark
SQL
Job Description
Baker Group is hiring a Data Engineer to design, build, and maintain data pipelines and infrastructure in Microsoft Fabric. In this onsite role in Ankeny, IA, you will help establish a single source of truth by owning ingestion, transformation, orchestration, governance, and preparation of data for reporting and AI/ML workloads.
You will work across the full data lifecycle, from integrating enterprise systems into Fabric to ensuring curated, trustworthy datasets are ready for analysts, data scientists, and downstream applications. The role also includes infrastructure stewardship, monitoring, and ongoing improvements to support reliable and timely data availability.
What you’ll do
- Design, build, and maintain ETL/ELT pipelines that ingest data from enterprise systems into Microsoft Fabric.
- Architect and maintain the Fabric medallion Lakehouse structure (bronze, silver, gold) as Baker Group’s single source of truth.
- Develop infrastructure and environment best practices, including Development/Test/Production setups and Git-based version control.
- Own pipeline orchestration, scheduling, and monitoring to support reliable, timely, and accurate data availability.
- Curate and maintain core datasets across employee, finance, project, service, and manufacturing domains.
- Establish and enforce data quality controls, including validation and reconciliation across pipelines.
- Design and manage data models, schemas, and semantic layers that support Data Analyst reporting and Data Scientist modeling.
- Define and maintain data ontologies and canonical business definitions (for example, what constitutes a “project,” “employee,” or “cost code”).
- Prepare data for AI and machine learning use cases, including feature-ready datasets, RAG pipelines, and vector embedding storage.
- Manage Fabric capacity planning, workspace organization, and performance optimization.
- Implement data governance practices aligned with Baker Group’s data classification standards, including access controls, lineage tracking, and metadata management.
- Partner with business system owners (including ERP, HRIS, and MRP) to understand upstream structures and manage change impacts.
- Collaborate with Data Scientists and Data Analysts to ensure pipeline outputs support modeling, paginated reporting, dashboards, and self-service BI.
- Coordinate with Software Development and DevOps teams to support application development needs.
- Work with 3rd party consultants when needed to deliver data engineering projects and augment capacity.
- Develop and maintain documentation for pipelines, schemas, and integration logic.
- Troubleshoot and resolve pipeline failures, latency issues, and data quality incidents.
- Monitor and maintain data infrastructure, including Fabric capacity, orchestration tools, and monitoring/alerting systems.
- Evaluate and recommend new data engineering tools, patterns, and best practices, and stay current on emerging trends.
What you bring
- Bachelor’s degree in Computer Science, Data Engineering, Information Systems, or another relevant quantitative field.
- 3 to 5 years of experience in data engineering, ETL/ELT development, or a related field.
- Proficiency with SQL and database technologies for data extraction, transformation, and loading.
- Experience with Microsoft Fabric, Azure Data Factory, or similar cloud ETL/orchestration tools.
- Experience with medallion architecture and modern data warehousing patterns.
- Programming experience with Python, PySpark, or T-SQL for data transformation.
- Familiarity with data modeling techniques such as dimensional modeling and star schema.
- Understanding of data governance, data quality, and metadata management practices.
Preferred add-ons
- Experience preparing data for AI/ML consumption (for example, vector embeddings and RAG architectures).
- Business acumen with understanding of construction or related industries.
Tools and technologies
- Microsoft Fabric
- ETL, ELT
- Azure Data Factory
- Git
- SQL
- Python, PySpark, T-SQL
- Medallion architecture
- Dimensional modeling, star schema
- Retrieval-augmented generation (RAG)
- Vector embedding storage
Certifications
- No specific requirements; relevant certifications such as Microsoft Certified: Fabric Data Engineer Associate, Azure Data Engineer Associate, or similar are a plus.
Work environment
- Prolonged periods of sitting at a desk and working on a computer.
- Must be able to lift 10 pounds occasionally.
- May have occasional visits to a job site requiring periods of standing, walking, and/or climbing stairs.
- Use a computer for 8 hours a day.