Senior Specialist, Data Engineer at Merck in Boston, MA is responsible for designing and maintaining data pipelines to ingest Upstream Biologics Process Development data, enabling downstream analytics, process characterization, scale-up predictions, and multivariate analyses.
Responsibilities
- Design and sustain robust, scalable data pipelines that ingest experimental and process data from upstream biologics source systems
- Provide analysis-ready datasets to support upstream digital initiatives, including process characterization models, scale-up predictions, multivariate analytics, and high-throughput process development workflows
- Map instrument outputs and experimental results ensuring ontology alignment and interoperability across upstream data sources
- Develop and maintain data visualizations, dashboards, and reports that enable upstream scientists to explore process data across runs, molecules, and scales
- Support system of record standards by ensuring consistent data entry practices
- Identify and flag data quality issues, metadata gaps, and inconsistencies across source systems, contributing to continuous improvement of upstream data capture practices
- Collaborate with upstream process development scientists, analytical scientists, and engineers to understand evolving data needs and translate them into pipeline requirements
- Coordinate with adjacent domain engineers to ensure seamless data handoffs at domain boundaries
- Maintain and version all pipeline code in GitHub, following team standards for code review, documentation, and deployment
- Demonstrate excellent interpersonal, communication, and collaboration skills
- Embrace and model core values of inclusion, fostering a supportive culture where all can thrive
- Collaborate effectively in a dynamic, integrated, and multidisciplinary team environment
Qualifications
- PhD required
- Minimum 2 years of experience in a relevant role
Requirements
- Proficient in Python and/or R programming
- Comfortable working in development environments such as Jupyter, Posit/RStudio, or VS Code
- Solid SQL skills with hands-on experience writing and optimizing queries against relational databases and data warehouses
- Experience with ETL/ELT processes and building data pipelines in a scientific or pharmaceutical context
- Familiarity with version control systems (Git/GitHub) and collaborative software development practices
- Ability to work in a team environment with cross-functional interactions
- Motivated to learn new skills, willingness to take on new challenges, and scientific curiosity
Technologies
- Python, R
- Jupyter, Posit/RStudio, VS Code
- SQL
- Git, GitHub
- Databricks, Delta Lake
- Streamlit, Shiny
- PowerBI, Spotfire, Tableau
- Allotrope Simple Model, ISA-88, OPC-UA
Benefits
- Medical, dental, vision healthcare and other insurance benefits for employee and family
- Retirement benefits, including 401(k)
- Paid holidays
- Vacation
- Compassionate and sick days
Location
Boston, MA (onsite)
Flexible Work Arrangements
Hybrid
Compensation
USD 117,000 - 184,200 per year
Requisition ID
R404620
Job Posting End Date
07/09/2026
Employee Status
Regular
US and Puerto Rico Residents Only
Our company is committed to inclusion, ensuring that candidates can engage in a hiring process that exhibits their true capabilities. Please click here if you need an accommodation during the application or hiring process.
San Francisco Residents Only
We will consider qualified applicants with arrest and conviction records for employment in compliance with the San Francisco Fair Chance Ordinance
Los Angeles Residents Only
We will consider for employment all qualified applicants, including those with criminal histories, in a manner consistent with the requirements of applicable state and local laws, including the City of Los Angelesβ Fair Chance Initiative for Hiring Ordinance