DataJobs.io
← Back to all jobs

Job Description

Deloitte seeks a Senior Data Engineer to design, build, and optimize end-to-end data pipelines onsite in Jersey City, NJ, with a salary range of USD 95,000 to 150,000 per year and responsibilities spanning design leadership, collaboration and mentoring within a Project Delivery Model.

Responsibilities

  • Maintain regular communications with Engagement Managers (Directors), project teammates, and stakeholders from functional and technical teams, escalating issues that require engagement management attention.
  • Design, develop, and optimize ETL and ELT pipelines using Azure Data Factory and Databricks.
  • Write and tune PySpark and Spark SQL notebooks for large-scale data transformations.
  • Architect end-to-end data solutions across dev, UAT, and prod environments using Unity Catalog.
  • Lead design discussions with client architects and other counterparts.
  • Collaborate with multiple teams to establish data contracts and schema agreements.
  • Lead the design and optimization of high-volume data pipelines.
  • Define and enforce data engineering standards, including naming conventions, partitioning strategies, cluster configurations, and Spark tuning.
  • Drive performance improvements through AQE tuning, liquid clustering, broadcast joins, and shuffle partition management.
  • Design Databricks cluster policies, autoscaling setups, and cost optimization strategies.
  • Perform root cause analysis of production incidents and implement durable fixes.
  • Mentor junior and mid-level engineers through code reviews and pair programming.
  • Evaluate new technologies and recommend adoption, such as DABs, DLT, Auto Loader, Serverless Compute, and event hubs.

Requirements

  • Proficiency in Python, PySpark, Spark SQL, and SQL Server.
  • Experience with Azure components including Data Factory, ADLS Gen2, Key Vault, and Azure Monitor.
  • Hands-on work with Databricks features like Delta Lake, Unity Catalog, and Workflows.
  • Familiarity with Apache Airflow for orchestrating data pipelines.
  • Version control and CI/CD experience using Git and Azure DevOps.
  • In-depth knowledge of Spark internals, including DAG optimization, spill analysis, and skew handling.
  • Delta Lake advanced features such as time travel, deletion vectors, and predictive I/O.
  • Unity Catalog governance covering row and column security, external locations, and system tables.
  • Infrastructure as Code experience with Terraform and Azure Resource Manager templates.
  • Bachelor's degree in Computer Science, Information Technology, Computer Engineering, or a related IT discipline, or equivalent experience.
  • Limited immigration sponsorship may be available.
  • Ability to travel approximately 10 percent, depending on client needs and assignments.

Technologies

  • Python
  • PySpark
  • Spark SQL
  • SQL Server
  • Azure Data Factory
  • ADLS Gen2
  • Key Vault
  • Azure Monitor
  • Databricks
  • Delta Lake
  • Unity Catalog
  • Workflows
  • Apache Airflow
  • Git
  • Azure DevOps
  • Deep Spark internals
  • DAG optimization
  • spill analysis
  • skew handling
  • Delta Lake time travel
  • deletion vectors
  • predictive I/O
  • Unity Catalog governance
  • Terraform
  • Azure ARM templates
  • DABs
  • DLT
  • Auto Loader
  • Serverless Compute
  • event hubs

Additional Requirements

Similar Jobs