AI Engineer, Evals & Agent Quality
Job Description
Town.com is building an AI assistant, and this role will own evaluation and agent quality end to end, from measurement to routing and regression prevention.
Responsibilities
- Design and implement a generalized eval system to measure assistant quality across all user touchpoints
- Ensure evaluation coverage includes multi-step agent trajectories, not just single-turn outputs
- Create and maintain golden datasets along with a labeling loop to keep data current
- Run continuous validation checks to confirm improvements and avoid regressions
- Build model routing capabilities and supporting online evaluation tooling to learn which models perform best
- Make every prompt and system change measurable so the team can iterate quickly without breaking working behavior
- Collaborate with product engineers to instrument quality and close the loop from quality signals to implementation fixes
Requirements
- Experience building or owning LLM eval systems, or delivering offline/online quality measurement at scale
- Strong, rigorous approach to measurement design
- Hands-on familiarity with the eval landscape, including off-the-shelf tooling and eval frameworks, plus the ability to choose approaches based on judgment
- Comfort reasoning about model routing and tradeoffs between models
- Track record of shipping fixes in addition to producing dashboards and metrics
- Senior or staff-level capability operating in a greenfield environment where the evaluation system does not yet exist
Location
- San Francisco, CA (onsite)
- Five days a week in person at the Financial District office
Compensation
- USD 250,000 - 300,000 per yearly