DataJobs.io
← Back to all jobs

Job Description

Distyl AI builds evaluation systems that help AI teams ship measurable improvements into customer environments. In this hybrid San Francisco role, you will help create the tooling and signal-quality foundations for Evaluation-Driven Development, so iteration is guided by outcomes you can verify. The position comes with meaningful equity, comprehensive benefits, and the opportunity to work on high-impact projects with access to state-of-the-art AI models and modern AI tooling.

Responsibilities

  • Design and implement evaluation frameworks that support Evaluation-Driven Development for AI systems deployed in customer environments.
  • Define how system quality is measured across domains, ensuring evaluation signals reflect user needs, domain constraints, and business objectives.
  • Build and maintain golden test cases and regression suites in Python, using human-authored and AI-assisted test generation to capture critical behaviors and edge cases.
  • Treat test suites as first-class system components that evolve alongside the AI system.
  • Develop and maintain evaluation pipelines that run offline and online, integrating results into system iteration loops.
  • Ensure evaluation results inform prompt design, agent logic, model selection, and release readiness using measurable improvements instead of intuition.
  • Define, calibrate, and operate LLM-based graders, aligning automated judgments with expert human assessments and refining grading approaches when signals diverge from real-world outcomes.
  • Collaborate with Forward Deployed AI Engineers, Architects, Product Engineers, AI Strategists, and domain experts to ensure evaluation frameworks guide development and deployment in production.

Requirements

  • 2+ years of software engineering experience.
  • Strong Python engineering skills, including the ability to build evaluation and experimentation pipelines that run in production with the same rigor as application code.
  • Experience with Evaluation-Driven Development or Experiment-Driven Development, including awareness of pitfalls like overfitting to metrics that do not reflect real outcomes.
  • Ability to translate human judgment into code, working with subject matter experts to encode judgments into test cases, scoring functions, and graders that scale.
  • Systems-oriented mindset for designing evaluation systems that connect prompts, agents, data, and deployment while supporting fast iteration and trust and safety in production.
  • AI-native working style, using AI tools to generate tests, analyze failures, explore edge cases, and accelerate debugging and iteration.
  • Travel between 10–50% of the time, depending on the project and role.

Benefits

  • Meaningful equity plus a comprehensive benefits package.
  • 100% coverage of medical, dental, and vision insurance for employees and dependents.
  • Flexible time off.
  • Retirement and financial planning benefits, including access to pre-tax HSA, FSA, and commuter accounts, 401(k), and financial coaching resources.
  • Comprehensive wellness benefits, including physical fitness, mental well-being, and fertility and family-building benefits through Carrot.
  • Complimentary in-office lunches and snacks provided.
  • Access to state-of-the-art AI models, generous usage of modern AI tools, and real-world business problems.
  • Ownership of high-impact projects across top enterprises.
  • A mission-driven, fast-moving culture that values curiosity, pragmatism, and excellence.

Work model: Hybrid collaboration with 3+ days per week (Tuesday–Thursday) in-office.

Compensation: Base salary range is $150,000 to $250,000 per year, depending on experience, location, and level.

Technology focus: Python, LLM-based graders, Evaluation-Driven Development, Experiment-Driven Development, HSA, FSA, 401(k), Carrot.

Similar Jobs