Distyl AI builds evaluation systems that help AI teams ship measurable improvements into customer environments. In this hybrid San Francisco role, you will help create the tooling and signal-quality foundations for Evaluation-Driven Development, so iteration is guided by outcomes you can verify. The position comes with meaningful equity, comprehensive benefits, and the opportunity to work on high-impact projects with access to state-of-the-art AI models and modern AI tooling.
Responsibilities
- Design and implement evaluation frameworks that support Evaluation-Driven Development for AI systems deployed in customer environments.
- Define how system quality is measured across domains, ensuring evaluation signals reflect user needs, domain constraints, and business objectives.
- Build and maintain golden test cases and regression suites in Python, using human-authored and AI-assisted test generation to capture critical behaviors and edge cases.
- Treat test suites as first-class system components that evolve alongside the AI system.
- Develop and maintain evaluation pipelines that run offline and online, integrating results into system iteration loops.
- Ensure evaluation results inform prompt design, agent logic, model selection, and release readiness using measurable improvements instead of intuition.
- Define, calibrate, and operate LLM-based graders, aligning automated judgments with expert human assessments and refining grading approaches when signals diverge from real-world outcomes.
- Collaborate with Forward Deployed AI Engineers, Architects, Product Engineers, AI Strategists, and domain experts to ensure evaluation frameworks guide development and deployment in production.
Requirements
- 2+ years of software engineering experience.
- Strong Python engineering skills, including the ability to build evaluation and experimentation pipelines that run in production with the same rigor as application code.
- Experience with Evaluation-Driven Development or Experiment-Driven Development, including awareness of pitfalls like overfitting to metrics that do not reflect real outcomes.
- Ability to translate human judgment into code, working with subject matter experts to encode judgments into test cases, scoring functions, and graders that scale.
- Systems-oriented mindset for designing evaluation systems that connect prompts, agents, data, and deployment while supporting fast iteration and trust and safety in production.
- AI-native working style, using AI tools to generate tests, analyze failures, explore edge cases, and accelerate debugging and iteration.
- Travel between 10–50% of the time, depending on the project and role.
Benefits
- Meaningful equity plus a comprehensive benefits package.
- 100% coverage of medical, dental, and vision insurance for employees and dependents.
- Flexible time off.
- Retirement and financial planning benefits, including access to pre-tax HSA, FSA, and commuter accounts, 401(k), and financial coaching resources.
- Comprehensive wellness benefits, including physical fitness, mental well-being, and fertility and family-building benefits through Carrot.
- Complimentary in-office lunches and snacks provided.
- Access to state-of-the-art AI models, generous usage of modern AI tools, and real-world business problems.
- Ownership of high-impact projects across top enterprises.
- A mission-driven, fast-moving culture that values curiosity, pragmatism, and excellence.
Work model: Hybrid collaboration with 3+ days per week (Tuesday–Thursday) in-office.
Compensation: Base salary range is $150,000 to $250,000 per year, depending on experience, location, and level.
Technology focus: Python, LLM-based graders, Evaluation-Driven Development, Experiment-Driven Development, HSA, FSA, 401(k), Carrot.