DataJobs.io
← Back to all jobs

Job Description

Responsibilities

  • Develop production ML systems that score and refine outputs where correctness is nuanced and labels are imperfect.
  • Architect evaluation frameworks for tasks where ground truth is partial, delayed, or disputed.
  • Create feedback loops that translate reviews, disagreements, corrections, and adjudication into measurable model and system improvements.
  • Own end-to-end production ML behavior, including precision and recall tradeoffs, regression detection, drift, latency, cost, and explainability.
  • Enhance model quality using appropriate approaches such as prompting, fine-tuning, retrieval, active learning, heuristics, and error analysis.
  • Collaborate with backend engineers to integrate inference into durable, long-running workflows while preserving debuggability and human oversight.

Requirements

  • Proven track record shipping ML systems that improved a real product, workflow, or business metric.
  • Strong instincts for model quality, evaluation design, error analysis, and production failure modes.
  • Comfort operating in ambiguous problem spaces with imperfect labels and evolving correctness.
  • Judicious decision making about when to apply prompting, fine-tuning, retrieval, human review, or simpler constraints.
  • Solid engineering fundamentals across the full ML stack, not limited to modeling.
  • Familiarity with LLM applications, model-assisted workflows, evaluation frameworks, or human-in-the-loop ML is a strong plus.
  • Preference for simple, inspectable ML systems that improve quickly and fail in understandable ways, rather than overly complex architectures.
  • Comfort deploying models only when there is a clear evaluation story.
  • Ability to navigate ambiguity without paralysis and make reasonable bets with incomplete information.
  • Focus on real-world impact and output of the system, not solely benchmark metrics.

Technologies

  • Python
  • Temporal
  • Postgres
  • AWS
  • LiteLLM

Benefits

  • Bi-annual performance bonus structure
  • Generous equity grant vested over 4 years
  • Up to $15k relocation bonus
  • $10k housing bonus if you live within 0.5 miles of the office
  • $1.5k monthly meal stipend
  • Free Equinox membership
  • $200 monthly laundry reimbursement
  • $200 monthly personal wellness reimbursement
  • Health, dental, and vision insurance

Day to Day

  • Move quickly within a young, high-ownership codebase where decisions have long-term architectural impact
  • Work across models, data, backend systems, and product surfaces, with frequent context switching
  • Debug production ML failures in live, long-running workflows where silent errors matter
  • Collaborate closely with backend engineers on a stack that includes Python, Temporal, Postgres, AWS, and LiteLLM
  • Balance automation confidence with human review, recognizing when to defer versus ship

What Makes This Role Different

  • The architecture is not fixed; early engineers define how quality is measured, how models and humans interact, where automation is trusted, and how the system compounds over time
  • The feedback loop is short, with visible impact on what customers receive when model behavior changes ship
  • You operate on a strategically central product area at Mercor during a moment when frontier AI challenges lack robust solutions

Similar Jobs