DataJobs.io
← Back to all jobs

Job Description

Pyx Health Inc is hiring a Healthcare Data Engineer to help build and maintain Azure-based data infrastructure and dependable data pipelines. In this remote role, you will work with Databricks, Airflow, Python, and SQL to support healthcare data workflows, with an emphasis on reliability, governance, and data quality for downstream reporting.

What you’ll do

  • Build and maintain data pipelines on Azure to ingest, transform, and clean healthcare data using Databricks, PySpark, and Delta Lake
  • Implement pipeline logic based on defined specifications, applying architectural guidance from senior engineers
  • Create reusable testing frameworks that help ensure pipeline reliability
  • Monitor pipelines for failures and performance issues, escalating complex problems when needed
  • Develop and maintain Airflow DAGs with appropriate error handling and retry logic
  • Support deployment and configuration of pipelines using Astronomer on Azure via Astro CLI
  • Reduce manual intervention by improving pipeline reliability over time
  • Implement data models in Delta Lake for efficient storage and retrieval, including merge/upsert patterns
  • Design and optimize Delta Lake tables for analytic workloads
  • Use Unity Catalog to support data governance, organization, and security
  • Apply Change Data Capture (CDC) patterns under the direction of senior team members
  • Write and optimize T-SQL and Databricks SQL scripts, stored procedures, and ETL support
  • Work across the Azure ecosystem including ADLS Gen2, Key Vaults, Logic Apps, and Azure DevOps
  • Follow established security and scalability standards when building data infrastructure
  • Ensure scripts and datasets are well-documented to support enterprise data governance
  • Design and continuously improve automated data quality monitoring, reconciliation, and alerting to support reliable downstream reporting
  • Troubleshoot and resolve pipeline failures by documenting root causes and resolutions
  • Collaborate with data scientists, analysts, and business stakeholders to understand reporting requirements and troubleshoot data issues
  • Participate in code reviews and incorporate feedback to improve code quality
  • Document pipelines, processes, and implementation decisions clearly and consistently
  • Stay current with data engineering technologies, especially in the healthcare space

What you’ll bring

  • 2–4 years of experience as a Data Engineer or closely related data role
  • Hands-on experience with Azure cloud services such as ADLS, Databricks, or similar
  • Working knowledge of SQL for scripting and data modeling, including T-SQL, Spark, and Databricks SQL
  • Ability to contribute to technical projects with moderate oversight
  • Strong communication skills, including comfort asking questions and sharing updates with cross-functional partners
  • Familiarity with CI/CD, Git, and version control workflows
  • Solid problem-solving skills with a systematic approach to debugging
  • Healthcare data standards and regulations knowledge (such as HIPAA and HL7) is a plus
  • Bachelor’s degree in Computer Science, Information Technology, Engineering, or related field (or equivalent work experience)

Core must-haves

  • Databricks SQL (Mid): Delta Lake fundamentals, basic merge/upsert patterns, familiarity with CDC concepts
  • Python (Mid): Pipeline logic, data transformation, SQL scripting; some PySpark/Spark DataFrame experience
  • Databricks Spark Notebooks (Mid): Notebook-based development, basic cluster usage, Delta table operations
  • Airflow Python Development (Mid): Ability to write and maintain DAGs, including error handling and retry patterns
  • Airflow Astro Configuration (Foundational): Familiarity with Astronomer or willingness to learn; basic Astro CLI usage
  • Azure Ecosystem (Foundational): Working knowledge of ADLS Gen2, Key Vaults, and Azure DevOps
  • T-SQL (Foundational–Mid): Basic stored procedures, SQL Server querying, and data manipulation

Nice to have

  • Azure Data Factory (Foundational): Basic familiarity with pipeline authoring and triggers
  • Git + Azure DevOps CI/CD (Foundational): Experience with branching, pull requests, peer review processes, version control best practices, and deploying production data pipelines using CI/CD workflows

Location: Remote (remote).

Similar Jobs