Data Engineer
Job Description
BioAgilytix is hiring a hands-on Data Engineer technical lead to design, build, and support its Enterprise Data Platform. In this onsite Durham, NC role, you will develop scalable, validated, production-ready data pipelines, enterprise data models, and certified data products that support laboratory operations, analytics, sponsor reporting, regulatory compliance, and AI initiatives.
How you’ll make an impact
- Design, develop, and support scalable enterprise data platforms that deliver trusted, governed, high-quality data products for analytics, scientific operations, sponsor reporting, regulatory compliance, and AI.
- Build and maintain production-grade ELT pipelines integrating data from LIMS, ERP, CRM, REST APIs, sponsor systems, cloud applications, and other enterprise sources.
- Create modular, reusable transformation frameworks using modern ELT practices, including automated testing, documentation, lineage, version control, and deployment automation.
- Develop dimensional and semantic models, plus governed datasets aligned to enterprise business definitions across the organization.
- Deliver validated, traceable, and auditable pipelines for regulated laboratory operations and enterprise reporting, maintaining integrity, reproducibility, lineage, and compliance with GxP, GLP, HIPAA, 21 CFR Part 11, and enterprise data governance standards.
- Implement automated validation, reconciliation, data quality controls, audit logging, monitoring, observability, and operational alerting to keep enterprise data reliable and trusted.
- Optimize platform performance, scalability, security, governance, and operational efficiency.
- Provide production operational support including monitoring, incident resolution, root cause analysis, and continuous reliability improvements.
- Support sponsor-facing delivery by building automated harmonization, transformation, validation, lineage, and regulatory reporting processes.
- Develop certified, governed, AI-ready data products supporting enterprise analytics, machine learning, semantic search, and generative AI initiatives.
- Collaborate with Laboratory Operations, Quality teams, IT, and business stakeholders to deliver scalable, reusable, governed enterprise data solutions.
- Perform other duties as needed.
What you bring
- Bachelor’s degree in computer science, Information Systems, Engineering, Mathematics, Data Science, or related field (Master’s preferred).
- 5+ years of experience in data engineering, data management, software engineering, business intelligence, or related technical disciplines, preferably in life sciences, biotechnology, pharmaceuticals, CROs, healthcare, or other regulated industries.
- 3+ years hands-on experience designing, developing, and supporting enterprise-scale data engineering solutions in production environments.
- 3+ years hands-on experience architecting, developing, and administering enterprise solutions using Snowflake as a primary cloud data platform, including performance optimization, security, governance, workload management, and operational support.
- Strong hands-on experience with dbt Cloud or dbt Core for modular transformations, automated testing, documentation, lineage, and deployment.
- Demonstrated expertise in enterprise dimensional modeling (star schemas, conformed dimensions, slowly changing dimensions, snapshot fact tables, analytical data warehouse design).
- Experience designing semantic models, enterprise business vocabularies, ontology-driven data products, or knowledge graph concepts.
- Strong proficiency in SQL and Python.
- Experience developing enterprise data integration using Talend, Fivetran, or equivalent ETL/ELT platforms.
- Experience integrating enterprise applications using REST APIs, GraphQL APIs, file-based interfaces, CDC, and event-driven messaging platforms.
- Experience with AWS (S3, Lambda, ECS, Glue) and/or Azure.
- Experience maintaining certified enterprise data products with documented business definitions, transformation logic, lineage, ownership, and lifecycle management.
- Experience implementing least-privilege security, RBAC, data masking, row-level security, encryption, secrets management, and secure data sharing.
- Experience supporting enterprise production platforms including incident management, root cause analysis, operational monitoring, performance tuning, release management, and reliability engineering.
- Experience with Laboratory Information Management Systems and regulated laboratory environments is strongly preferred.
- Expert proficiency in SQL and Python, including Snowflake architecture (Snowpark, Dynamic Tables, Streams, Tasks, data sharing, security, governance, workload management, performance tuning).
- Git and GitHub Actions, CI/CD, Infrastructure-as-Code, and DataOps practices.
- Ability to translate complex scientific, laboratory, and business requirements into scalable enterprise data and certified data products based on established standards.
- Strong analytical, troubleshooting, and optimization skills, including investigation and root cause analysis for complex issues.
- Ability to develop validated, traceable, auditable data solutions aligned to compliance, validation, quality, and data-integrity requirements.
- Collaboration and communication skills across Scientific Operations, Quality Engineering, Quality Assurance, and IT.
- Ability to work independently on complex assignments, use judgment, and escalate decisions affecting architecture, governance, security, compliance, or platform direction.
- Excellent English communication (written and spoken), documentation, and interpersonal skills.
Technologies you’ll work with
Snowflake, dbt Cloud, dbt Core, ELT, dimensional and semantic modeling, CI/CD automation, DataOps, data governance standards, GxP, GLP, HIPAA, 21 CFR Part 11, SQL, Python, Talend, Fivetran, REST APIs, GraphQL APIs, CDC, event-driven messaging platforms, AWS (S3, Lambda, ECS, Glue), Azure, Snowpark, Dynamic Tables, Streams, Tasks, data sharing, RBAC, Power BI, Sigma, Tableau, Git, GitHub Actions, Infrastructure-as-Code, semantic reporting platforms.
Benefits
- Medical Insurance (HDHP with HSA and PPO)
- Dental Insurance
- Vision Insurance
- Flexible Spending Account (medical and dependent care)
- Short Term Disability and Long Term Disability
- Life Insurance
- Paid Time Off: 4 weeks per year
- Parental Leave
- Paid Holidays: 9 scheduled and 5 floating
- 401k with Employer Match
- Employee Referral Program
Position details: Full-time role, onsite in Durham, NC. Some flexibility in hours is allowed, with availability required during core work hours per the BioAgilytix Employee Handbook. Occasional weekend, holiday, and evening work may be needed.