Senior Data Engineer
Senior
Azure
Azure Data Factory
Azure Databricks
Big Data
Bigdata
Cloud
Cloud Platforms
Data Architecture
Data Engineer
Data Governance
Data Integration
Data Lake
Data Lakehouse
Data Pipeline
Data Platform
Data Processing
Data Security
Database
Databricks
Databricks Workflows
Delta Lake
Engineering
ETL
Event Streaming
Microsoft Azure
Spark
SQL
Job Description
Deloitte seeks a Senior Data Engineer to design, build, and optimize end-to-end data pipelines onsite in Jersey City, NJ, with a salary range of USD 95,000 to 150,000 per year and responsibilities spanning design leadership, collaboration and mentoring within a Project Delivery Model.
Responsibilities
- Maintain regular communications with Engagement Managers (Directors), project teammates, and stakeholders from functional and technical teams, escalating issues that require engagement management attention.
- Design, develop, and optimize ETL and ELT pipelines using Azure Data Factory and Databricks.
- Write and tune PySpark and Spark SQL notebooks for large-scale data transformations.
- Architect end-to-end data solutions across dev, UAT, and prod environments using Unity Catalog.
- Lead design discussions with client architects and other counterparts.
- Collaborate with multiple teams to establish data contracts and schema agreements.
- Lead the design and optimization of high-volume data pipelines.
- Define and enforce data engineering standards, including naming conventions, partitioning strategies, cluster configurations, and Spark tuning.
- Drive performance improvements through AQE tuning, liquid clustering, broadcast joins, and shuffle partition management.
- Design Databricks cluster policies, autoscaling setups, and cost optimization strategies.
- Perform root cause analysis of production incidents and implement durable fixes.
- Mentor junior and mid-level engineers through code reviews and pair programming.
- Evaluate new technologies and recommend adoption, such as DABs, DLT, Auto Loader, Serverless Compute, and event hubs.
Requirements
- Proficiency in Python, PySpark, Spark SQL, and SQL Server.
- Experience with Azure components including Data Factory, ADLS Gen2, Key Vault, and Azure Monitor.
- Hands-on work with Databricks features like Delta Lake, Unity Catalog, and Workflows.
- Familiarity with Apache Airflow for orchestrating data pipelines.
- Version control and CI/CD experience using Git and Azure DevOps.
- In-depth knowledge of Spark internals, including DAG optimization, spill analysis, and skew handling.
- Delta Lake advanced features such as time travel, deletion vectors, and predictive I/O.
- Unity Catalog governance covering row and column security, external locations, and system tables.
- Infrastructure as Code experience with Terraform and Azure Resource Manager templates.
- Bachelor's degree in Computer Science, Information Technology, Computer Engineering, or a related IT discipline, or equivalent experience.
- Limited immigration sponsorship may be available.
- Ability to travel approximately 10 percent, depending on client needs and assignments.
Technologies
- Python
- PySpark
- Spark SQL
- SQL Server
- Azure Data Factory
- ADLS Gen2
- Key Vault
- Azure Monitor
- Databricks
- Delta Lake
- Unity Catalog
- Workflows
- Apache Airflow
- Git
- Azure DevOps
- Deep Spark internals
- DAG optimization
- spill analysis
- skew handling
- Delta Lake time travel
- deletion vectors
- predictive I/O
- Unity Catalog governance
- Terraform
- Azure ARM templates
- DABs
- DLT
- Auto Loader
- Serverless Compute
- event hubs
Additional Requirements
- Information for applicants who need accommodations: accommodations.