DataJobs.io
← Back to all jobs

Job Description

Dhanu Global Enterprises, Inc. is building a governed data and AI platform for a pharmaceutical client, with a focus on connecting device, laboratory, partner, and document information into a unified foundation. In this onsite role in Indianapolis, you will design and deliver production-ready data engineering capabilities that support regulatory reporting, advanced analytics, and AI-driven scientific insights.

Working in a GxP/FDA-regulated environment, you will help integrate structured and unstructured data through an Azure-based architecture, combining governed ETL/ELT pipelines with API integrations and strong operational controls.

What you will do

  • Design, build, and maintain production ETL/ELT pipelines that integrate laboratory and operational systems (including Darwin, Teamcenter/PLM, LabVantage LIMS, Jama, TurboAC, and Qdocs/Veeva) into Microsoft Azure Fabric Lakehouse and PostgreSQL.
  • Develop and optimize a Bronze, Silver, Gold medallion architecture, covering schema mapping, data modeling, referential integrity, and performance optimization.
  • Build scalable API integrations and cross-cloud data pipelines across Azure and AWS to support enterprise data integration.
  • Implement automated data quality controls, controlled vocabulary normalization, schema validation, Q-gate/specification checks, data lineage, and audit trails to support GxP compliance.
  • Monitor, troubleshoot, and optimize pipeline performance, reliability, error handling, and operational monitoring for production environments.
  • Collaborate with business, engineering, and IT teams to connect data sources and establish a governed, scalable digital thread for analytics, AI, and regulatory reporting.

Key requirements

  • 8+ years of hands-on Data Engineering experience building production ETL/ELT pipelines and API integrations.
  • Expert-level Python and/or PySpark for ingestion, transformation, orchestration, testing, and CI/CD.
  • Hands-on experience with Microsoft Azure Fabric (Lakehouse, Data Factory, Fabric Pipelines, and Delta Lake) delivering end-to-end production solutions.
  • Hands-on AWS experience, including S3, Glue (or equivalent), and RDS/Aurora.
  • Strong SQL and PostgreSQL skills, including normalized schema design, query optimization, indexing, and performance tuning.
  • Experience with data modeling, medallion architecture, and Lakehouse design patterns.
  • Experience implementing data lineage, quality controls, schema validation, error handling, and monitoring in regulated environments.
  • Knowledge of GxP, GMP, GCP in pharmaceutical, biotechnology, or medical device environments.
  • Experience with SCM, WMS, PLM, MM, QA, and/or validated applications in FDA-regulated systems.

Technologies

  • ETL, ELT, Python, PySpark
  • Microsoft Azure Fabric: Lakehouse, Data Factory, Fabric Pipelines, Delta Lake
  • PostgreSQL
  • AWS: S3, Glue, RDS, Aurora
  • SQL, Bronze, Silver, Gold, Medallion architecture, Q-gate, Data lineage

Role logistics

  • Location: Indianapolis, IN (onsite, 5 days per week)
  • Salary: USD 60 - 70 per monthly
  • Duration: Through December 2026, with possible extension into 2027
  • Experience level: 8+ years

Similar Jobs