DataJobs.io
← Back to all jobs

Job Description

Realtor.com is hiring an AI Engineer for its AI Integrations Team in Austin, TX (onsite). In this role, you will own the AI safety layer for RealAssist, a consumer-facing LLM real-estate assistant. You will build and evaluate guardrails, tune compliance and content-moderation classifiers, and help ensure safety stays reliable as the product and models evolve.

What you’ll do

  • Build and tune classifiers that help keep the AI assistant compliant, safe, and useful, working at the intersection of AI and consumer product needs.
  • Develop and own LLM-as-a-judge classifiers for Fair-Housing compliance and content moderation, targeting high recall on disallowed content while avoiding unnecessary blocking for legitimate users.
  • Design and run guardrail evaluation pipelines, including curated labeled datasets, maintaining train/test/validate splits, and running offline prod-replay evaluations to measure overblocking.
  • Use eval reporting to make defensible safety decisions by producing confusion matrices and tracking precision/recall/F1 per category.
  • Integrate and operate cloud guardrail services as a runtime screening layer for prompt injection and jailbreak attempts, with fail-open behavior and alerting.
  • Red-team the assistant and wire safety regression checks into CI to prevent guardrails from silently degrading.
  • Manage guardrail infrastructure as code using Terraform, including templates, IAM, and project structure, and participate in architecture discussions and technical design reviews.
  • Partner with ML, backend, product, and legal/compliance stakeholders to define risk tiering and route new AI capabilities through safety review before launch.
  • Own the AI safety layer end-to-end, from design through deployment, with continuous, data-driven decision-making grounded in evaluation metrics.

Requirements

  • AI Safety & Evaluation Expertise (Required)
  • 4+ years of professional software/ML engineering experience, with hands-on LLM application work
  • Bachelor’s degree or equivalent experience
  • Strong Python proficiency
  • Real classification-metrics literacy: precision/recall/F1 (macro vs weighted), confusion-matrix debugging, dataset curation, and calibration to production distributions
  • Experience building or tuning LLM prompts/classifiers and evaluating them using frameworks such as DeepEval, RAGAS, Arize Phoenix, LangSmith, OpenAI Evals, or a homegrown harness
  • Experience turning small seed sets into robust labeled eval datasets using dataset-synthesis tooling and prompt-optimization loops
  • Familiarity with prompt-injection/jailbreak defense concepts including OWASP LLM Top 10, input/output filtering, least-privilege tool access, and adversarial testing

Bonus experience

  • Hands-on with a cloud guardrail service: Google Cloud Model Armor, AWS Bedrock Guardrails, or Azure AI Content Safety
  • Terraform / IaC for cloud infrastructure
  • Experience with Google Cloud Vertex AI / Gemini
  • Monitoring and observability exposure (for example, New Relic or similar)
  • Experience in regulated or compliance-sensitive domains (fair housing, fair lending, healthcare, finance, trust & safety)
  • Red-teaming or AI-security background

Tools you may use

Python, DeepEval, RAGAS, Arize Phoenix, LangSmith, OpenAI Evals, Google Cloud Model Armor, Terraform, Google Cloud Vertex AI / Gemini, New Relic, OWASP LLM Top 10, CI, Google Cloud, AWS Bedrock Guardrails, Azure AI Content Safety

Benefits

  • Inclusive and competitive medical, Rx, dental, and vision coverage
  • Family forming benefits
  • 13 Paid Holidays
  • Flexible Time Off
  • 8 hours of paid Volunteer Time Off
  • Immediate eligibility into the Company 401(k) plan with a 3.5% company match
  • Tuition Reimbursement program for degreed and non-degreed programs
  • 1:1 personalized Financial Planning Sessions
  • Student Debt Retirement Savings Match program
  • Free snacks and refreshments in each office location

Similar Jobs