HealthLeap is building reliable data systems that help turn messy clinical information into dependable inputs for ML pipelines, analytics, and user-facing APIs. This role focuses on owning the core data pipelines end to end, with a strong emphasis on correctness, monitoring, and recovery.
What you’ll build and operate
You will design, build, and run data pipelines that move data from hospital ingestion through transformation and delivery. The goal is trusted, production-ready datasets that downstream teams can rely on for machine learning, customer analytics, and API experiences.
- Build and operate pipelines from hospital ingestion through transformation and delivery.
- Produce trusted data for ML pipelines, customer analytics, and user-facing APIs.
- Define data contracts and implement checks to catch missing records, schema changes, and incorrect values.
- Manage backfills, late-arriving data, failures, and recovery workflows.
- Partner with integration, ML, and product engineers to ensure data lands reliably where it needs to go.
How you’ll ensure reliability
This position is centered on keeping data dependable as real-world inputs change. You’ll use monitoring, validation, and recovery practices to handle late data, schema changes, and system failures while maintaining continuity for model inputs and analytics.
- Strong judgment around correctness, monitoring, and recovery.
- Ability to trace issues across systems and own the fix through production.
Requirements
- 5+ years building production data systems, with strong Python and SQL.
- Experience owning pipelines that depend on messy, changing external data.
- Strong data modeling skills and judgment about correctness, monitoring, and recovery.
- Ability to trace a problem across systems and own the fix through production.
Tools
Python, SQL
Compensation and benefits
Base salary: $175,000–$275,000 per year, plus meaningful equity.
- 100% covered healthcare premiums
- Unlimited PTO with a 20-day minimum
- 4% 401(k) match
- Laptop and home office budget
What stands out
- Experience building data pipelines for ML products.
- Experience with clinical data, EHRs, HL7, or FHIR.
- Early-stage experience building and operating core data systems.
Interview process
- No LeetCode or puzzles; you’ll use the tools you’d use on the job, including AI.
- Intro call
- Data pipeline design
- Practical data exercise
- Onsite with the team in San Francisco
- Decision made the same week as the onsite
Location and work style
San Francisco, CA (onsite). The team works together in the office by default, with flexibility to work from home when needed. They judge output, not hours.
Note: This role isn’t a fit if you want to own only one piece of the data stack or if you require predictable 9-to-5 hours. A hospital go-live can mean a 60+ hour week, even though the company protects deep rest.