This position is no longer accepting applications
Closed on September 23, 2026.
This role is filled — get an email when new Artificial Intelligence roles open on DataJobs.io:
Evaluation & Insights Machine Learning Engineer
Ai Evaluation
Ai Testing And Evaluation
Artificial Intelligence
Engineer
Generative Ai Engineer
Machine Learning Evaluation
Machine Learning Infrastructure
View similar jobs
Get alerted when similar jobs are posted — set up a New Artificial Intelligence jobs on DataJobs.io alert.
See other roles at Apple.
Job Description
As part of Apple’s Human-Centered AI team, you will evaluate and improve AI systems through data science, model behavior analysis, and qualitative insights.
Responsibilities
- Architect and run comprehensive evaluation suites for LLMs and multimodal models, surfacing edge cases across multi-step reasoning, factuality, adversarial robustness, safety, and alignment.
- Build deterministic, heuristic, and LLM-assisted evaluation frameworks (for example, LLM-as-a-judge and reward modeling) to measure human-perceived quality metrics such as helpfulness and hallucination rates.
- Convert qualitative failure modes into quantified loss patterns, programmatic guardrails, and actionable data-mixture adjustments for model training and inference.
- Work with engineering teams to refine model behavior using evaluation telemetry to guide prompt engineering, Retrieval-Augmented Generation (RAG) strategy, and model fine-tuning.
- Apply advanced ML methods (embedding-based clustering, representation learning, perturbation analysis) to map error taxonomies and latent failure patterns.
- Implement MLOps workflows to formalize evaluation metrics, automate regression testing across model checkpoints, and connect human-centric assessments to ML CI/CD pipelines.
- Design scalable, distributed inference and processing pipelines (for example, Ray and vLLM) to support high-throughput evaluation, automated annotation, and large-scale output analysis.
- Define quantitative evaluation frameworks that capture human factors such as trust calibration, conversational state tracking, and interpretability.
- Build automated evaluation pipelines using LLMs to score outputs at scale, optimizing for high correlation with human baseline annotations.
- Collaborate with ML researchers, software developers, and product managers to translate product requirements into reliable and efficient evaluation infrastructure.
Requirements
- Knowledge of human factors, HCI, or cognitive science methodologies as applied to AI system design.
- Bachelor’s or Master’s degree in Computer Science, Machine Learning, Artificial Intelligence, Cognitive Science, or a related technical field.
- 8+ years of relevant industry experience in ML Engineering or Applied Research.
- Advanced proficiency in Python and modern deep learning ecosystems (PyTorch, JAX, Hugging Face).
- Proven experience building scalable ML inference pipelines, model-evaluation workflows, and structured rating frameworks for large-scale AI systems.
- Strong ability to interpret unstructured model outputs (text, transcripts, embedding spaces) and synthesize qualitative findings into actionable engineering guidance and training objectives.
- Hands-on experience developing, fine-tuning, or evaluating LLMs, multimodal models, and NLP systems.
- Deep familiarity with AI quality metrics, hallucination detection techniques (including SelfCheckGPT), model alignment (including RLHF and DPO), and LLM-as-a-judge frameworks (including G-Eval and DeepEval).
- Experience building internal tools or automated pipelines for ML workflows using platforms such as MLflow and Weights & Biases or similar.
- Strong familiarity with advanced prompt engineering, RAG architectures (vector databases, semantic search), and fine-tuning.
Technologies
- Python, PyTorch, JAX, Hugging Face
- LLM-as-a-judge, reward modeling
- Retrieval-Augmented Generation (RAG), vector databases, semantic search
- Embedding-based clustering, representation learning, perturbation analysis
- MLOps, CI/CD pipelines
- Ray, vLLM
- Trust calibration, SelfCheckGPT
- RLHF, DPO
- G-Eval, DeepEval
- MLflow, Weights & Biases
- Fine-Tuning
Benefits
- Comprehensive medical and dental coverage
- Retirement benefits
- Range of discounted products and free services
- Reimbursement for certain educational expenses, including tuition
- Discretionary restricted stock unit awards
- Opportunity to purchase Apple stock at a discount through the Employee Stock Purchase Plan
- Base pay range between $184,700 and $324,800
- Comprehensive total compensation package may include discretionary bonuses or commission payments as well as relocation
Pay & Benefits
- Base pay range: $184,700 - $324,800 (depends on skills, qualifications, experience, and location).
- Employees may participate in discretionary employee stock programs and purchase stock at a discount via the Employee Stock Purchase Plan.
- Eligibility for benefits, compensation, and stock programs is subject to plan or program terms and requirements.
Location
- Cupertino, CA (onsite)
Experience Level
- 8+ years