DataJobs.io
← Back to all jobs

Job Description

Xenon7 is hiring a Senior Data Engineer for a life sciences client on a contract basis. This hybrid role is built around owning the architectural direction while delivering hands-on data platform and pipeline engineering that connects scientific and clinical informatics with manufacturing process and OT data.

The work centers on scalable enterprise data integration across modalities such as genomics, proteomics, and LIMS, alongside systems like MES, SCADA, and OSIsoft PI. The setup includes a 3-day onsite cadence in the Indianapolis, IN area, with openness to regional/EST candidates who can travel onsite.

Responsibilities

  • Design, build, and maintain production-grade data pipelines and data architecture for scientific, clinical trial, and research informatics (including small/large molecule, genomics, proteomics, and LIMS).
  • Structure complex multi-modal clinical and scientific datasets to support advanced analytics, enterprise reporting, and downstream machine learning readiness.
  • Ingest, harmonize, and model operational technology (OT) and manufacturing process data including API manufacturing pipelines, batch processing data, MES, SCADA, and OSIsoft PI.
  • Unify laboratory and facility data pipelines into centralized, highly available enterprise data platforms.
  • Build robust ETL/ELT pipelines using modern cloud platforms such as Databricks, Snowflake, and AWS/Azure with PySpark and SQL.
  • Operate within enterprise governance requirements, including data residency and GxP regulatory standards in a heavily monitored environment.
  • Collaborate directly with process engineers, chemical engineering leads, and research informatics directors to convert operational friction into clear technical specifications.
  • Set engineering best practices, data modeling standards, and pipeline monitoring frameworks across the enterprise data stack.

Requirements

  • Senior-level experience (10–20+ years) in software development, data platform architecture, and complex ETL/ELT engineering.
  • Proven ability to engineer data pipelines across specialized, non-standard domains, including transitions between process/chemical engineering data and clinical/scientific research informatics.
  • Unrestricted US work authorization (no sponsorship available) and ability to work 3 days onsite per week in the Indianapolis, IN area.
  • Strong problem-solving and adaptability, with the ability to explain complex technical architecture to cross-functional engineering teams.
  • Advanced Python, PySpark, and expert-level SQL.
  • Hands-on experience with Databricks, Snowflake, or AWS/Azure enterprise data ecosystems.
  • Extensive experience with Airflow, dbt, Spark, and enterprise data orchestration engines.
  • Demonstrated streaming and batch architecture work using REST APIs, message brokers, and database integrations.
  • Deep exposure to either scientific/clinical informatics (CDISC/SDTM, LIMS, clinical trials, multi-omics) or chemical/process engineering data (API manufacturing, batch data, SCADA, MES, OSIsoft PI).

Technologies

  • Python, PySpark, SQL
  • Databricks, Snowflake, AWS, Azure
  • Airflow, dbt, Spark
  • REST APIs, message brokers
  • CDISC/SDTM, LIMS, MES, SCADA, OSIsoft PI

Location & Contract Details

  • Location: Indianapolis, IN Metro (Hybrid / 3-Day Onsite). Open to regional/EST candidates with onsite travel.
  • Contract Type: Contractor Full-Time / Enterprise Project Engagement (outsourced via Xenon7).

Nice-to-Haves & Certifications

  • Academic background in Chemical Engineering, Bio-process Engineering, Computer Science, or related STEM discipline.
  • Direct experience working in regulated GxP environments in Life Sciences or Specialty Chemicals.
  • Certifications such as Databricks Certified Data Engineer Senior/Professional, Snowflake SnowPro Core/Advanced, or AWS Data Engineer Associate/Professional.

Important context: This is a hands-on data engineering lead role focused on building and scaling data architecture, pipelines, and platform infrastructure. It is not a pure Data Scientist or ML Researcher position, and it is not fully remote.

Similar Jobs