Sr Databricks Data Engineer
Apache Airflow
Automation
Big Data
Bigdata
Data Architecture
Data Engineer
Data Engineering
Data Governance
Data Integration
Data Lake
Data Lakehouse
Data Pipeline
Data Platform
Data Security
Data Warehouse
Database
Databases
Databricks
Databricks Lakeflow
Databricks Workflows
Delta Lake
Delta Live Tables
ETL
Spark
SQL
Structured Streaming
Job Description
Senior Databricks Data Engineer within Deloitte AI and Data practice, accountable for designing, building, and optimizing cloud-based data engineering solutions on Databricks to modernize data platforms, enable analytics and AI, and drive measurable business outcomes. This role is onsite in Pittsburgh, PA, with a salary range of USD 116,200 to 229,100 per year.
Responsibilities
- Establish and advocate leading practices for data architecture, integration, and modeling, documenting standards and promoting adherence across teams.
- Own the end-to-end design, development, and maintenance of robust data pipelines and architectures to support enterprise-scale data needs.
- Lead initiatives to enhance data quality, streamline operations, and scale data processes.
- Assess, pilot, and integrate new big data and analytics technologies to keep the organization at the cutting edge; provide leadership and mentorship to data engineers and architects to support growth and project success.
- Advise on and implement governance, security, and compliance strategies tailored to cloud-based data ecosystems.
- Translate technical concepts and business value for stakeholders across the organization, including executives, business leads, and technology teams.
- Guide the adoption of CI/CD practices using tools such as Azure DevOps, AWS CodePipeline, Jenkins, TFS, and PowerShell to streamline deployments and operations.
Requirements
- Bachelor's degree in Computer Science, Engineering, or a related field
- More than five years of hands-on data engineering experience focusing on Databricks across AWS, Azure, or Google Cloud Platform (GCP)
- Proficiency with Lakehouse architectures, Apache Spark, Delta Lake, cloud-native databases and storage solutions, and distributed compute platforms
- Experience with data warehousing and 3NF, dimensional modeling, enterprise data lakes, incremental data loads, and metadata-driven ingestion and data quality frameworks using PySpark
- At least one year leading complex, cross-functional data projects and technical teams, including Delta Live Tables, Autoloader, Structured Streaming, Databricks Workflows, Apache Airflow, Unity Catalog, automated CI/CD pipelines, and performance optimization of data pipelines, code, and compute resources
- Ability to travel approximately 50 percent, depending on client engagements
- Limited immigration sponsorship may be available
- Master's degree in Computer Science, Engineering, or a related field
- Experience with one or more cloud ecosystems (AWS, Azure, GCP) and associated big data services
- Experience tuning and optimizing performance in Databricks and Apache Spark environments
- Experience with Databricks Lakeflow
- Experience with artificial intelligence and machine learning solutions
Technologies
- Databricks
- Amazon Web Services (AWS)
- Microsoft Azure
- Google Cloud Platform (GCP)
- Apache Spark
- Delta Lake
- Unity Catalog
- Delta Live Tables
- Autoloader
- Structured Streaming
- Databricks Workflows
- Apache Airflow
- Azure DevOps
- AWS CodePipeline
- Jenkins
- TFS
- PowerShell
- PySpark
- Databricks Lakeflow
Benefits
- Discretionary annual incentive program
- Benefits package aligned with Core Talent Model
Qualifications Required
- Bachelor's degree in Computer Science, Engineering, or a related field
- Five or more years of hands-on data engineering experience with a focus on Databricks on AWS, Azure, or GCP
- Experience with Lakehouse architecture, Apache Spark, Delta Lake, cloud-native databases, storage solutions, and distributed compute platforms
- Experience with data warehousing, 3NF, dimensional modeling, enterprise data lakes, incremental data loads, and metadata-driven ingestion and data quality frameworks using PySpark
- At least one year leading complex, cross-functional data projects and technical teams, including Delta Live Tables, Autoloader, Structured Streaming, Databricks Workflows, Apache Airflow, Unity Catalog, automated CI/CD pipelines, and performance optimization of data pipelines, code, and compute resources
- Ability to travel approximately 50 percent
- Limited immigration sponsorship may be available
Preferred
- Master's degree in Computer Science, Engineering, or a related field
- Experience with one or more cloud ecosystems (AWS, Azure, GCP) and their big data services
- Experience tuning and optimizing performance in Databricks and Apache Spark environments
- Experience with Databricks Lakeflow
- Experience with artificial intelligence and machine learning solutions