DataJobs.io
← Back to all jobs

Job Description

The role of Associate Director, Data Engineer within DSCS Digital Data Strategy is based onsite in Boston, MA. It leads the design, construction, and governance of biologics data pipelines and defines end-to-end data strategy to enable modeling, optimization, and decision support across Digital Insights. A Ph.D. is required, and the compensation ranges from USD 129,000 to 203,100 per year.

Responsibilities

  • Act as the domain owner for biologics data engineering, maintaining a comprehensive view of all digital projects, data sources, systems, and data flows within the domain.
  • Contribute to shaping the DSCS Digital Data Strategy by informing relevant work and initiatives.
  • Design and build robust, scalable data pipelines that ingest experimental and process data from biologics source systems such as process historians, chromatography systems, electronic lab notebooks, and analytical instruments.
  • Provide analysis-ready datasets to support digital initiatives, including process characterization models, data lineage, multivariate analytics, and cross-site manufacturing connectivity.
  • Define and enforce data standards, metadata schemas, and ontology mappings to ensure biologics data is interoperable for modeling and optimization workflows.
  • Collaborate with automation colleagues, anticipating when new or modified automated workflows create new data streams that require pipeline development and ontology mapping.
  • Own and govern system of record standards for biologics, ensuring consistent configuration and data entry practices across experiments, molecules, and sites.
  • Catalog all processes, analytical methods, instruments, and digital systems within the biologics domain to create a comprehensive data landscape map.
  • Develop and maintain data visualizations, dashboards, and reports enabling scientists to explore process data across runs, molecules, scales, and manufacturing sites.
  • Influence the digital data strategy for biologics by identifying opportunities to improve data capture at the source and reduce friction between experimentation and modeling.
  • Mentor and guide supporting data engineers across modalities, ensuring alignment with domain strategy and ontology governance.
  • Maintain and version all pipeline code in GitHub, following team standards for code review, documentation, and deployment.
  • Build strong partnerships with process development scientists, analytical scientists, and manufacturing teams to gather requirements and shape the domain's digital data strategy.
  • Demonstrate strong interpersonal, communication, and collaboration skills.
  • Model and promote the core values of diversity and inclusion, fostering a supportive culture where all can thrive.
  • Collaborate effectively in a dynamic, integrated, and multidisciplinary team environment.

Requirements

  • Proficiency in Python and/or R, with comfort using development environments such as Jupyter, Posit/RStudio, or VS Code.
  • Solid SQL skills with hands-on experience writing and optimizing queries against relational databases and data warehouses.
  • Experience with ETL/ELT processes and building data pipelines in a scientific or pharmaceutical context.
  • Familiarity with cloud platforms (AWS, Azure, or GCP) for data storage, processing, and integration.
  • Working knowledge of how process models, multivariate analyses, and statistical tools consume experimental data to anticipate modeler needs and deliver structured datasets.
  • Experience defining or enforcing data standards, metadata schemas, or ontology mappings in a scientific or pharmaceutical context.
  • Familiarity with version control systems (Git/GitHub) and collaborative software development practices.
  • Demonstrated ability to lead technical initiatives, mentor junior engineers, and influence data strategy across multiple stakeholders.
  • Ability to deliver complex solutions under compressed timelines in a dynamic environment.

Technologies

  • Python
  • R
  • Jupyter
  • Posit/RStudio
  • VS Code
  • SQL
  • AWS
  • Azure
  • GCP
  • Databricks
  • Delta Lake
  • Git
  • GitHub
  • Streamlit
  • Shiny
  • PowerBI
  • Spotfire
  • Tableau
  • Allotrope Simple Model
  • ISA-88
  • OPC-UA

Benefits

  • Medical, dental, and vision coverage for employee and family
  • Retirement benefits including a 401(k) plan
  • Paid holidays
  • Vacation time
  • Compassionate and sick leave
  • Annual bonus and long-term incentive, if applicable

Preferred Experience and Skills

  • Hands-on biologics process development experience such as chromatography, filtration, purification, or formulation, transitioning into a data engineering, data science, or computational role.
  • Experience with Databricks, including notebook-based development, workflow orchestration, and Delta Lake.
  • Experience with data visualization tools (Streamlit, Shiny, PowerBI, Spotfire, Tableau) for scientist-facing dashboards and exploratory apps.
  • Familiarity with ontology frameworks or standardized data models (eg, Allotrope Simple Model, ISA-88, OPC-UA) and mapping instrument data to structured schemas.
  • Understanding of Design of Experiments and process characterization study designs to structure data for statistical analysis of CPPs and CQAs.
  • Experience with data lineage and building traceability across experimental systems, materials, and manufacturing steps.
  • Knowledge of regulatory expectations relevant to biologics process development, process characterization, and process validation (ICH Q8-Q12, validation lifecycle, comparability studies).
  • Experience collaborating with lab automation teams and integrating data from newly automated workflows.
  • Experience with cross-site data integration, harmonizing data from multiple facilities with different systems and conventions.
  • Evidence of cross-functional collaborations spanning laboratory, manufacturing, modeling, and digital teams.

Similar Jobs