AI Engineer
Job Description
Realtor.com is hiring an AI Engineer for its AI Integrations Team in Austin, TX (onsite). In this role, you will own the AI safety layer for RealAssist, a consumer-facing LLM real-estate assistant. You will build and evaluate guardrails, tune compliance and content-moderation classifiers, and help ensure safety stays reliable as the product and models evolve.
What you’ll do
- Build and tune classifiers that help keep the AI assistant compliant, safe, and useful, working at the intersection of AI and consumer product needs.
- Develop and own LLM-as-a-judge classifiers for Fair-Housing compliance and content moderation, targeting high recall on disallowed content while avoiding unnecessary blocking for legitimate users.
- Design and run guardrail evaluation pipelines, including curated labeled datasets, maintaining train/test/validate splits, and running offline prod-replay evaluations to measure overblocking.
- Use eval reporting to make defensible safety decisions by producing confusion matrices and tracking precision/recall/F1 per category.
- Integrate and operate cloud guardrail services as a runtime screening layer for prompt injection and jailbreak attempts, with fail-open behavior and alerting.
- Red-team the assistant and wire safety regression checks into CI to prevent guardrails from silently degrading.
- Manage guardrail infrastructure as code using Terraform, including templates, IAM, and project structure, and participate in architecture discussions and technical design reviews.
- Partner with ML, backend, product, and legal/compliance stakeholders to define risk tiering and route new AI capabilities through safety review before launch.
- Own the AI safety layer end-to-end, from design through deployment, with continuous, data-driven decision-making grounded in evaluation metrics.
Requirements
- AI Safety & Evaluation Expertise (Required)
- 4+ years of professional software/ML engineering experience, with hands-on LLM application work
- Bachelor’s degree or equivalent experience
- Strong Python proficiency
- Real classification-metrics literacy: precision/recall/F1 (macro vs weighted), confusion-matrix debugging, dataset curation, and calibration to production distributions
- Experience building or tuning LLM prompts/classifiers and evaluating them using frameworks such as DeepEval, RAGAS, Arize Phoenix, LangSmith, OpenAI Evals, or a homegrown harness
- Experience turning small seed sets into robust labeled eval datasets using dataset-synthesis tooling and prompt-optimization loops
- Familiarity with prompt-injection/jailbreak defense concepts including OWASP LLM Top 10, input/output filtering, least-privilege tool access, and adversarial testing
Bonus experience
- Hands-on with a cloud guardrail service: Google Cloud Model Armor, AWS Bedrock Guardrails, or Azure AI Content Safety
- Terraform / IaC for cloud infrastructure
- Experience with Google Cloud Vertex AI / Gemini
- Monitoring and observability exposure (for example, New Relic or similar)
- Experience in regulated or compliance-sensitive domains (fair housing, fair lending, healthcare, finance, trust & safety)
- Red-teaming or AI-security background
Tools you may use
Python, DeepEval, RAGAS, Arize Phoenix, LangSmith, OpenAI Evals, Google Cloud Model Armor, Terraform, Google Cloud Vertex AI / Gemini, New Relic, OWASP LLM Top 10, CI, Google Cloud, AWS Bedrock Guardrails, Azure AI Content Safety
Benefits
- Inclusive and competitive medical, Rx, dental, and vision coverage
- Family forming benefits
- 13 Paid Holidays
- Flexible Time Off
- 8 hours of paid Volunteer Time Off
- Immediate eligibility into the Company 401(k) plan with a 3.5% company match
- Tuition Reimbursement program for degreed and non-degreed programs
- 1:1 personalized Financial Planning Sessions
- Student Debt Retirement Savings Match program
- Free snacks and refreshments in each office location