DataJobs.io
← Back to all jobs

Job Description

Apple’s Human-Centered AI team focuses on building trustworthy generative AI experiences. In this role, you will evaluate and improve foundation models by creating measurement frameworks, scalable MLOps pipelines, and automated insights that connect human perception to model performance.

Responsibilities

  • Architect and execute comprehensive evaluation suites for LLMs and multimodal models, identifying edge cases across multi-step reasoning, factuality, adversarial robustness, safety, and alignment.
  • Build deterministic, heuristic, and LLM-assisted evaluation frameworks (including LLM-as-a-judge and reward modeling) to quantify human-perceived quality metrics such as helpfulness and hallucination rates.
  • Convert qualitative failure modes into quantifiable loss patterns, programmatic guardrails, and actionable data-mixture adjustments for both training and inference.
  • Partner with engineering teams to refine model behavior using evaluation telemetry to guide prompt engineering, Retrieval-Augmented Generation (RAG) strategies, and model fine-tuning.
  • Apply advanced ML methods (such as embedding-based clustering, representation learning, and perturbation analysis) to map error taxonomies and latent failure manifolds in model outputs.
  • Develop robust MLOps workflows that codify evaluation metrics, automate regression testing across model checkpoints, and incorporate human-centric assessments into ML CI/CD pipelines.
  • Architect scalable, distributed inference and processing pipelines (for example, Ray and vLLM) to support high-throughput evaluation, automated annotation, and large-scale output analysis.
  • Define quantitative evaluation frameworks that capture nuanced human factors, including trust calibration, conversational state tracking, and interpretability.
  • Create automated evaluation pipelines using LLMs to assess outputs at scale, optimizing for high correlation with human baseline annotations.
  • Collaborate with ML researchers, software developers, and product managers across Apple to convert product requirements into scalable, reliable, and efficient model evaluation infrastructure.

Requirements

  • 5+ years of relevant industry experience in ML Engineering or Applied Research.
  • Knowledge of human factors, HCI, or cognitive science methodologies as applied to AI system design.
  • Advanced proficiency in Python and modern deep learning ecosystems, including PyTorch, JAX, and Hugging Face.
  • Experience building scalable ML inference pipelines, model-evaluation workflows, and structured rating frameworks for large-scale AI systems.
  • Strong ability to interpret unstructured model outputs (text, transcripts, and embedding spaces) and translate qualitative findings into actionable engineering guidance and training objectives.
  • Hands-on experience developing, fine-tuning, or evaluating LLMs, multimodal models, and NLP systems.
  • Deep familiarity with AI quality metrics, hallucination detection techniques (including SelfCheckGPT), model alignment methods (including RLHF and DPO), and LLM-as-a-judge frameworks such as G-Eval and DeepEval.
  • Experience building internal tools or automated pipelines for ML workflows using tools such as MLflow, Weights & Biases, or similar platforms.
  • Strong familiarity with advanced prompt engineering, RAG architectures (vector databases and semantic search), and fine-tuning.
  • Bachelor’s or Master’s degree in Computer Science, Machine Learning, Artificial Intelligence, Cognitive Science, or a related technical field.

Technologies

  • Python, PyTorch, JAX, Hugging Face
  • Ray, vLLM
  • MLflow, Weights & Biases
  • SelfCheckGPT, RLHF, DPO
  • G-Eval, DeepEval
  • LLM-as-a-judge
  • Retrieval-Augmented Generation (RAG), vector databases, semantic search

Benefits

  • Comprehensive medical and dental coverage.
  • Retirement benefits.
  • Range of discounted products and free services.
  • Reimbursement for certain educational expenses (including tuition) for formal education related to advancing your career.
  • Opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs.
  • Eligible for discretionary restricted stock unit awards.
  • Can purchase Apple stock at a discount if voluntarily participating in Apple’s Employee Stock Purchase Plan.
  • May be eligible for discretionary bonuses or commission payments and relocation.

Role Location and Salary

This position is based in Seattle, WA and is onsite. The salary range is USD 142,300 - 263,300 per year.

Additional Notes

The role will help bridge human perception and algorithmic performance to evaluate and optimize foundation models and generative AI systems. Responsibilities include architecting evaluation frameworks, designing scalable MLOps pipelines for model assessment, and working cross-functionally with Software Engineering, Product, Research, and Responsible AI teams to support reliable, safe, and human-aligned AI experiences.

Preferred Qualifications

  • Knowledge of human factors, HCI, or cognitive science methodologies as applied to AI system design.

Similar Jobs