Associate Director, Data Engineer
Manager
Big Data
Cloud Data Engineering
Data
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Engineering
Data Engineering Lead
Data Integration
Data Pipeline
Data Platform
Data Processing
Data Warehouse
Database
Databases
Director
ETL
Informatica
Information Technology (IT)
Integration
Programming
Programming Language
SQL
Job Description
Merck’s DSCS Digital Technologies team is building the data foundation that enables data-driven modeling for sterile drug product development. In this role, you will design, build, and maintain pipelines and datasets that capture, curate, and deliver SPD experimental and process data for machine learning, statistical, and hybrid modeling workflows. The position is based onsite in Rahway, NJ with hybrid flexibility, supporting a multidisciplinary environment across scientific, engineering, and digital disciplines.
What you’ll do
- Partner with SPD experimentalists, process engineers, and analytical scientists to define requirements for data solutions that feed modeling pipelines.
- Design and implement robust, scalable data pipelines that ingest experimental and process data, including SPD unit operations such as mixing, pooling, pumping, filling, filtration, and freeze-drying, plus analytical characterization data.
- Work alongside process modeling teams to convert model requirements into analysis-ready, feature-rich datasets.
- Define and enforce data standards, metadata schemas, and ontologies to improve interoperability and downstream usability for modeling workflows.
- Automate ingestion from laboratory instruments, electronic lab notebooks, PAT systems, and manufacturing systems, integrating outputs with cloud-based storage and compute environments.
- Create data analysis and visualization workflows to surface insight from SPD data.
- Design and build dashboards, reports, and data exports for scientific and cross-functional stakeholders.
- Curate data and define requirements to automate data ingestion at scale.
- Influence SPD digital data strategy by identifying opportunities to improve data capture at the source and reduce friction between experimentation and modeling.
- Embody core values of inclusion and help foster a supportive culture where all team members can thrive.
- Collaborate effectively in a dynamic, integrated, and multidisciplinary team environment to deliver trusted partnerships across broad stakeholders.
Qualifications
- Ph.D. in Computer Science, Data Science, Engineering, Chemistry, Physics, Biology, Pharmaceutical Sciences, or a closely related field, with at least 3 years of industrial/pharmaceutical or relevant experience; OR
- M.S. in Computer Science, Data Science, Engineering, Chemistry, Physics, Biology, Pharmaceutical Sciences, or a closely related field, with at least 5 years of industrial/pharmaceutical or relevant experience; OR
- B.S. in Computer Science, Data Science, Engineering, Chemistry, Physics, Biology, Pharmaceutical Sciences, or a closely related field, with at least 7 years of industrial/pharmaceutical or relevant experience.
- Hands-on experience in sterile drug product development, sterile DS and DP manufacturing processes, or closely related pharmaceutical development, with a demonstrated transition into a data engineering, data science, or computational role.
- Experience developing and deploying data pipelines, ETL/ELT workflows, and data integration solutions in a scientific or pharmaceutical context.
- Proficiency programming in Python and/or R, with Posit/RStudio/Jupyter.
- Working knowledge of how data-driven models consume experimental data and the ability to anticipate modeler needs for appropriately structured datasets.
- Excellent communication, creativity, and interpersonal skills.
- Proven ability to deliver complex solutions under compressed timelines in a dynamic environment.
- Motivation to learn new skills, take on new challenges, and maintain scientific curiosity.
Tools and technologies
- Languages & tooling: Python, R, Posit/RStudio, Jupyter, SQL
- Pipeline and integration: ETL/ELT
- Visualization: Shiny, Streamlit, Spotfire, Dash, Power BI, Tableau
- Data platforms and governance: Dataiku, Databricks, Graph databases
- AWS services: S3, Redshift, Glue, Athena, SageMaker
- Scientific and manufacturing systems: Electronic lab notebooks, ELN, LIMS, historian/SCADA systems, PAT
- Infrastructure context: cloud-based storage and compute environments
Benefits
- Annual bonus and long-term incentive, if applicable
- Medical, dental, vision healthcare, and other insurance benefits for employee and family
- Retirement benefits, including 401(k)
- Paid holidays, vacation, and compassionate and sick days
Location: Rahway, NJ (onsite). Flexible work arrangements: Hybrid. Employee status: Regular. Requisition ID: R418480. Job posting end date: 10/5/2026. Salary range: $129,000.00 - $203,100.00 per year.