Senior Bioinformatics Data Engineer
Job Description
This senior data and platform engineering role focuses on accelerating biomarker data lake modernization on AWS, building and operating production pipelines across oncology clinical studies.
Responsibilities
- Build and maintain Dagster‑driven ingestion pipelines for genomics vendors (Caris, Predicine, Tempus, Olink, CellCarta) with IO managers, Iceberg writers, and row‑level accounting.
- Hardene dbt transformations from Silver to Gold, including real‑data test coverage, store‑failures patterns, staging/intermediate/mart models, and macro consolidation.
- Implement clinical data ingestion paths for SDTM and ADaM, add reconciliation logic, and route subject‑dimension data.
- Deliver platform infrastructure such as FastAPI endpoints, CI/CD pipelines, containerized deployments, observability instrumentation, and Redshift performance tuning.
- Extract transformation rules from legacy R and PySpark code and align them with new platform implementations.
- Identify repetitive processes and convert them into automated workflows, guardrails, or reusable tooling.
- Participate in adversarial design and code reviews, identify edge cases, and push back on suboptimal patterns.
- Collaborate with the lead engineer on design decisions and sustain delivery velocity through paired sessions and PR reviews.
- Ensure reproducibility standards: CI on every PR, automated tests, and avoid ad hoc notebook‑based production processes.
Requirements
- AI native engineering practice: proven experience building systems and workflows around AI coding agents (Claude Code, Cursor, Codex, or equivalent) beyond prompting; know when to automate repetitive tasks, add guardrails for agent outputs, and build infrastructure to accelerate future work.
- Education: Bachelor's or master's degree in computer science, data engineering, bioinformatics, or a related field.
- Experience: 5+ years in data engineering with shipped production pipelines on AWS (S3, ECS/Fargate, Redshift or equivalent MPP).
- Programming: strong Python and SQL skills with working knowledge of modern data engineering libraries.
- Orchestration and dbt: advanced proficiency with dbt and a workflow tool (Dagster, Airflow, or Prefect).
- Data quality: track record of catching silent failures, questioning data correctness, and detecting lossy joins or incomplete deliveries.
- Architecture: solid understanding of lakehouse patterns, ETL processes, and schema design for complex multi‑modal datasets.
- Compliance: ability to handle PHI‑adjacent clinical data under contractor policy (background check, compliance training, VPN access).
- Legacy code: willingness to work with R and PySpark to extract business rules and validate new implementations.
- Communication: excellent ability to collaborate in an embedded pair model with tight feedback loops.
Technologies
- Dagster, dbt, Apache Iceberg, Amazon S3, Redshift
- FastAPI, Docker, Amazon ECS, AWS Fargate
- Python, SQL, R, PySpark
- Airflow, Prefect
- CloudFormation, AWS Glue Catalog, CI/CD
Preferred Qualifications
- Direct experience with Apache Iceberg, AWS Glue Catalog, or lakehouse table formats.
- Comfort reading genomic data (VAF, HGVS nomenclature, VCFs, CNV/fusion semantics) or ability to ramp on unfamiliar scientific domains quickly.
- Familiarity with clinical data standards such as SDTM, ADaM, and CDISC.
- Pharma, clinical research, or life sciences background.
- Experience with containerization (Docker/ECS) and infrastructure‑as‑code (CloudFormation).
- Proficiency in R for interoperability with bioinformatics teams.