DataJobs.io
← Back to all jobs

Job Description

The Senior Data Engineer will support enterprise data engineering initiatives by designing, building, and optimizing scalable data solutions. The role includes owning and operating the Core Data platform, delivering both batch and streaming Spark pipelines, and applying governance across AWS, Kubernetes, and Airflow.

Role Overview

This position focuses on Core Data platform delivery and ongoing operations. The engineer will manage Databricks platform governance, develop pipelines with PySpark and related tooling, and ensure platform reliability through monitoring, cost awareness, automation, and troubleshooting across deployed services.

Key Responsibilities

  • Manage Databricks platform governance, including Unity Catalog, ACLs, lineage, and data discovery and privacy tooling
  • Design, write, test, and deploy batch and streaming data pipelines using PySpark, Scala, SQL, Python, and Kotlin
  • Meet with stakeholders to gather requirements and translate them into scalable data platform solutions
  • Use knowledge of Databricks platform and developer tooling to diagnose errors, audit platform activity, and automate updates across pipelines, objects, and integrations
  • Explain Spark architecture and pipeline behavior to stakeholders to identify root causes and recommend solutions
  • Provide solution architecture across AWS, Databricks, Kubernetes, and Airflow (MWAA), including cross-platform integrations
  • Build and maintain Kubernetes containers and containerized utilities that support deployed data platform services
  • Apply networking knowledge to troubleshoot connectivity and integration errors across platform components
  • Perform platform administration such as provisioning and removing access, assessing resource utilization, monitoring platform health and cost, and evaluating stakeholder requests
  • Collaborate with engineers, architects, and product managers to drive Core Data platform success and participate in agile or scrum ceremonies
  • Maintain documentation for platform changes, standards, and pipeline configurations to support data quality and governance
  • Develop and maintain scalable data engineering solutions, including building and enhancing data pipelines and workflows
  • Support cloud-based data platform initiatives and collaborate with engineering teams on system architecture and design
  • Implement infrastructure automation and deployment standards and troubleshoot and optimize data processing environments

Required Qualifications

  • Strong experience with Databricks
  • Advanced Python development skills
  • Experience with Terraform and Infrastructure as Code
  • 5+ years of relevant data engineering experience
  • Ability to participate in architecture and system design discussions
  • Strong problem-solving and data platform implementation experience

Additional Qualifications

  • 5+ years of data engineering experience developing and operating large-scale data pipelines
  • Deep hands-on experience with Databricks and Apache Spark (batch and streaming), including pipeline development in PySpark and/or Scala
  • Strong understanding of Spark architecture (executors, stages, partitioning, shuffle) with ability to explain performance tuning tradeoffs to technical and non-technical stakeholders
  • Proficiency with Databricks platform tooling (API, SDK, CLI) for automation, auditing, governance, and operational troubleshooting
  • Proficient in SQL with advanced performance tuning capabilities
  • Hands-on production experience with Airflow (MWAA) for orchestrating data pipelines
  • Experience managing Databricks governance, including ACLs, Unity Catalog, lineage, and access provisioning
  • Proficiency in Python and at least one additional language (Scala, Kotlin, or SQL-driven pipeline tooling)
  • Experience designing and optimizing scalable ETL/ELT pipelines integrating diverse structured and unstructured data sources
  • AWS-primary experience (compute, storage, networking, IAM); experience with other cloud providers is transferable
  • Proficiency with Docker and Kubernetes for building and maintaining containerized data platform services
  • Working knowledge of networking concepts to diagnose cross-platform integration and connectivity issues
  • Familiarity with Snowflake and comparable tooling relative to the Databricks ecosystem
  • Experience implementing CI/CD and DevOps practices using Git-based workflows
  • Experience implementing data quality checks, monitoring, and logging for pipeline reliability
  • Self-starting problem solver with strong analytical and communication skills and willingness to learn new tooling and trends
  • Familiar with Scrum and Agile methodologies
  • Bachelor’s Degree in Computer Science, Information Systems, or related field, or equivalent work experience; Master’s Degree is a plus

Preferred Qualifications

  • Experience with AWS
  • Cloud-based data platform experience
  • Data pipeline optimization and automation expertise

Technology Stack

  • Databricks, Apache Spark
  • Python, PySpark, Scala, SQL, Kotlin
  • Terraform, Infrastructure as Code
  • AWS, Kubernetes, Docker
  • Airflow (MWAA)
  • Unity Catalog
  • Git-based workflows, CI/CD

Location and Schedule

  • Glendale, CA (hybrid)
  • Duration: 12+ months

Compensation

  • Pay rate: USD 90 - 93 per hourly (DOE)

Benefits

  • Medical, dental, and vision coverage
  • 401(k) with company match
  • Short-term disability
  • Life insurance with AD&D

Similar Jobs