AI Engineer - Algorithm Evaluation & Agentic Systems
Agentic Ai
Agentic Ai Orchestration
Ai Agent
Ai Agent Platform
Ai Evaluation
Ai Evaluation
Artificial Intelligence
Automation
Computer Vision + Llm
Data Pipeline
Data Processing
Engineer
Engineering
Generative AI
Industrial Automation
Llm As Judge
Machine Learning Evaluation
Machine Vision
Mechatronics
Multimodal Ai
Robotics
Vision Language Models
Job Description
Apple’s DAQ team builds evaluation to advance advanced visual technologies. This AI Engineer role focuses on algorithm evaluation and designing agentic systems for computer vision and video understanding, bridging rigorous experimentation with production-ready benchmarking and integration.
Role Focus
Within DAQ, you will lead the benchmarking and integration of state-of-the-art image and video understanding models. The position emphasizes building evaluation pipelines to uncover edge-case failure modes and architect autonomous multi-modal workflows that connect experimentation with production deployments.
Responsibilities
- Design, build, and scale comprehensive evaluation pipelines for holistic end-to-end system evaluation and granular component-level testing on complex image and video understanding tasks.
- Analyze model outputs to identify root causes of visual hallucinations, temporal inconsistencies in video, and edge-case failures.
- Build, deploy, and evaluate agentic workflows that use vision models to autonomously solve multi-step user problems such as video summarization and visual search.
- Use component-level evaluation to isolate and triage which parts of an agentic workflow (including tool selection, memory retrieval, and visual reasoning) are succeeding or failing.
- Lead dataset curation and ground-truth benchmark development, tailored to evaluating multi-modal capabilities.
- Partner with core model training teams to provide actionable, data-driven insights and metrics to inform model training and fine-tuning iterations.
- Lead benchmarking and integration efforts for state-of-the-art models for image and video understanding.
Requirements
- MS and a minimum of 3 years of relevant industry experience.
- 3+ years of applied experience in Machine Learning, Computer Vision, or AI System Evaluation.
- Deep understanding of core Machine Learning principles, including probability, statistics, data distributions, and model bias/variance.
- Deep theoretical and practical understanding of Computer Vision and Vision-Language Models, including Vision Transformers (ViTs), spatial-temporal modeling, and image/video processing.
- Proven track record defining robust metrics and KPIs and designing rigorous evaluation frameworks for generative AI or foundation models, including custom benchmark creation, automated regression testing, LLM/VLM-as-a-judge methodologies, and human-in-the-loop evaluation.
- Experience building and evaluating LLM/VLM-powered agents, including tool use, multi-step reasoning, planning, and memory management workflows.
- Strong intuition for probing ML models to discover edge cases, hallucinations, and performance bottlenecks in constrained environments, and translating findings into improvement recommendations.
- Proficiency in Python and experience with deep learning frameworks (PyTorch) for inference, embedding extraction, and scalable evaluation pipeline development.
Technologies
- Python
- PyTorch
- Vision Transformers (ViTs)
- Computer Vision (CV)
- Vision-Language Models (VLMs)
- LLM/VLM-as-a-judge methodologies
- LLM/VLM-powered agents
What We Value
- Production mindset, including correctness, observability, and maintainability.
- Ability to reason about system-level tradeoffs beyond model performance.
- Balance between experimentation speed and engineering rigor.
- Comfort working in ambiguous problem spaces and defining metrics from first principles.
- Clear communication of technical findings to technical and non-technical audiences.
Preferred Qualifications
- Ability to lead technical evaluation strategies end-to-end, drive architectural decisions for testing infrastructure, and mentor engineers.
- Strong foundation in statistics, including hypothesis testing, confidence intervals, and experimental design.
- Knowledge of reinforcement learning, planning, or decision-making systems.
- Experience evaluating multi-modal or multi-agent systems.
- Prior work on AI reliability, safety, or benchmarking.
Pay & Benefits
- Base pay range: USD 150,400 to USD 277,600 per year.
- Comprehensive medical and dental coverage.
- Retirement benefits.
- Range of discounted products and free services.
- Reimbursement for certain educational expenses, including tuition for formal education related to advancing your career.
- Opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs.
- Discretionary restricted stock unit awards.
- Ability to purchase Apple stock at a discount if voluntarily participating in the Apple Employee Stock Purchase Plan.
- Potential discretionary bonuses or commission payments.
- Relocation may be eligible.
- Learn more about Apple Benefits.
- Note: Apple benefit, compensation, and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program.
Location and Experience Level
Onsite role in Sunnyvale, CA. Minimum experience requirement is 3 years, with an MS degree required.