This position is no longer accepting applications
Closed on September 9, 2026.
This role is filled — get an email when new Data Analytics roles open on DataJobs.io:
Sr Data Engineer, Data Analytics & Intelligence, NA
Senior
Artificial Intelligence
Azure Data Factory
Azure Data Lakehouse
Big Data
Data Analytics
Data Engineer
Data Integration
Data Pipeline
Data Platform
Data Processing
ETL
Microsoft Azure
Microsoft Fabric
SQL
View similar jobs
Get alerted when similar jobs are posted — set up a New Data Analytics jobs on DataJobs.io alert.
See other roles at Vantage Data Centers Management Company LLC.
Job Description
Build and scale governed Azure data foundations for enterprise reporting, operational intelligence, and AI-enabled consumption for Operations in North America.
Responsibilities
- Design, build, and maintain reliable, scalable data pipelines using Python and PySpark on the Microsoft Azure data platform
- Develop and operate batch and incremental pipelines using Azure Data Factory for orchestration and Azure Data Lake Storage Gen2 as the primary data store
- Create and maintain curated lakehouse / gold-layer datasets and semantic-model inputs for governed operational insights and AI-enabled consumption
- Independently implement SQL- and Spark-based transformations to produce curated datasets for enterprise reporting, analytics, and downstream use
- Own assigned pipelines and datasets, including monitoring, troubleshooting, performance optimization, documentation, and production support
- Work with Azure Synapse and Microsoft Fabric / Lakehouse patterns where applicable, plus related Azure analytics services
- Prepare structured operational data for AI-enabled use cases by documenting business rules, source lineage, data reliability constraints, known quality limitations, and data dictionary definitions
- Support source visibility and confidence context, including Data Reliability & Trust Indicator integration where applicable
- Contribute to ontology, taxonomy, semantic model, and data dictionary alignment across operational domains (KPIs, incidents, work orders, and more)
- Collaborate with business analysts, operations SMEs, data stewards, IT Global, and cross-functional stakeholders to translate requirements into working data solutions
- Apply data governance, security, access-control, data classification, and engineering standards for compliant, maintainable, scalable solutions
- Identify, document, and route data-quality issues to accountable owners to improve source correction rather than masking defects downstream
- Participate in code reviews, technical discussions, sprint planning, and platform improvement initiatives
- Proactively identify data quality issues, pipeline risks, platform dependencies, and improvement opportunities, and communicate them clearly
- Develop and maintain PySpark notebooks and jobs to ingest, transform, validate, and curate data within the enterprise data platform
- Build and modify Azure Data Factory pipelines for batch and incremental ingestion
- Implement Spark transformations to write curated datasets to Azure Data Lake Storage Gen2 and/or Fabric Lakehouse patterns using established folder structures, naming conventions, and governance standards
- Create and maintain SQL views, tables, lakehouse objects, and semantic-model inputs to support analytics, operational intelligence, and AI-enabled consumption patterns
- Prepare datasets for Fabric Data Agent / AI agent use cases with documentation including business rules, joins, grain, quality limitations, source lineage, and operational definitions
- Respond to pipeline failures, data validation issues, operational alerts, and data-quality escalations with root-cause analysis and remediation steps
- Perform performance tuning for Spark jobs and SQL workloads (partitioning, filtering, incremental logic, query optimization, and resource-aware design)
- Validate outputs with business partners, operations SMEs, and data stewards; address defects or discrepancies through documented correction paths
- Support observability, logging, and auditability practices for pipelines and AI-consumable datasets where applicable
- Commit code using Git, follow branching standards, support pull request reviews, and participate in CI/CD using GitHub, Azure DevOps, or similar tools
- Update documentation for pipelines, datasets, data contracts, data dictionaries, business rules, and operational runbooks as changes are made
- Execute assigned backlog items within sprint timelines; raise risks, dependencies, or blockers early
- Additional duties as assigned by management
Requirements
- Bachelor’s degree in Engineering, Computer Science, Data Analytics, or a related field, or equivalent experience
- Minimum 5–8 years of experience in data engineering, analytics engineering, or a closely related technical data role
- Proficiency in Python for data pipelines, automation, and data processing workflows, including PySpark-based transformations
- Proficiency in SQL for querying, transformation, analytical processing, model validation, and data quality checks
- Solid understanding of ETL/ELT pipelines, data transformation patterns, data integration concepts, incremental processing, and production support practices
- Experience analyzing enterprise data sources to identify relationships, transformations, business rules, grain, ownership, and quality constraints
- Experience building solutions on the Microsoft Azure platform, including exposure to Azure Data Factory, Azure Synapse, Azure Data Lake Storage Gen2, Microsoft Fabric / Lakehouse patterns, and related analytics services
- Working knowledge of data modeling fundamentals, including fact/dimension tables and semantic models
- Experience supporting governed data products, including metadata, lineage, issue documentation, access-control awareness, data-quality validation, and operational runbooks
- Experience with source control and CI/CD workflows using GitHub or Azure DevOps
- Strong communication and collaboration skills across IT Global, business SMEs, data governance partners, platform teams, and operations stakeholders
- Experience in Agile environments and using collaboration or tracking tools such as Jira or similar tools
- Travel required is expected to be up to 10% but may increase over time
Technologies
- Python, PySpark, Spark
- Microsoft Azure, Azure Data Factory, Azure Data Lake Storage Gen2, Azure Synapse
- Microsoft Fabric, Fabric Lakehouse, Microsoft Fabric / Lakehouse patterns
- SQL
- Git, GitHub, Azure DevOps
- Jira, CI/CD
Benefits
- Medical, dental, and vision coverage
- Life and AD&D
- Short and long-term disability coverage
- Paid time off
- Employee assistance
- 401k participation with company match
- Above market total compensation package
- Comprehensive suite of health and welfare, retirement, and paid leave benefits
Salary Range
- USD 130,000 - 155,000 per year
- Range is based on Colorado market data and may vary in other locations
- Compensation may depend on qualifications, skills, competencies, and experience and may fall outside the range shown
Additional Details
- On-site role in Denver, CO (3 days on site required, 2 days flexible)
- Eligible for company benefits including medical, dental, and vision coverage, life and AD&D, short and long-term disability coverage, paid time off, employee assistance, and 401k with company match
Physical Demands and Special Requirements
- Reasonable accommodations may be made to enable individuals with disabilities to perform the essential functions
- Occasionally required to stand, walk, sit; use hands to handle or feel objects; reach with hands and arms; climb stairs; balance; stoop or kneel; talk and hear
- Occasionally required to lift and/or move up to 25 pounds
Desired Qualifications
- Experience with distributed data processing frameworks, including Apache Spark
- Experience preparing governed data products for AI-enabled use cases, including Microsoft Fabric Lakehouse, semantic models, Fabric/Data Agent patterns, ontology or taxonomy alignment, and explainable AI outputs
- Familiarity with data observability, metadata management, lineage, data contracts, reliability indicators, and operational best practices in production environments
- Familiarity with additional Azure services such as Azure Functions or Logic Apps in support of data workflows
- Experience supporting data platform enhancement, refactoring, modernization, or reusable architecture initiatives
- Experience working with structured and unstructured operational sources, such as enterprise applications, operational workflows, documents, dashboards, and knowledge assets
- Experience working in a scaling or fast-paced organization where priorities evolve quickly and practical delivery discipline is required