DataJobs.io
← Back to all jobs

Job Description

Apple is hiring a Machine Learning Engineer to build and optimize intelligent search and AI experiences. The role focuses on transformer-based and foundation models, semantic retrieval, and evaluation systems designed for on-device performance in Cupertino.

Location and Work Style

  • Location: Cupertino, CA
  • Work mode: Onsite

Compensation

  • Base salary range: USD 150,400 to 225,300 per year
  • Base pay determination: Depends on skills, qualifications, experience, and location

Role Summary

In this position, you will design, train, fine-tune, and deploy transformer-based language models and foundation models. You will also develop semantic retrieval and retrieval-augmented generation systems, with an emphasis on on-device deployment and measurable evaluation.

Key Responsibilities

  • Build semantic retrieval, embedding, reranking, and retrieval-augmented generation systems, along with models for query understanding, intent prediction, personalization, retrieval, and ranking.
  • Analyze search relevance and user behavior to define evaluation methodologies, offline benchmarks, and online metrics covering retrieval quality, ranking, personalization, and language model performance.
  • Build scalable experimentation and evaluation pipelines for LLMs and search models, measuring model quality, robustness, latency, efficiency, and end-to-end product metrics.
  • Design, train, fine-tune, distill, and optimize transformer-based language models and foundation models for efficient on-device deployment.
  • Develop LLM fine-tuning and post-training approaches, including supervised fine-tuning, instruction tuning, preference optimization, parameter-efficient fine-tuning, and task-specific adaptation.
  • Research and prototype on-device generative AI methods such as knowledge distillation, model compression, quantization, pruning, and low-latency inference.
  • Develop techniques for transferring capabilities from large foundation models into compact on-device models while balancing quality, latency, memory footprint, power consumption, and compute constraints.
  • Partner with engineers, researchers, product managers, and designers to move AI capabilities from research into production, contributing to technical strategy and exploring applications of foundation models, multimodal AI, agentic retrieval, and personalized intelligence.

Required Education

  • Master’s or Ph.D. in Computer Science, Machine Learning, Artificial Intelligence, or a related field.
  • Bachelor degree in Computer Science, Machine Learning, Artificial Intelligence, or a related field.

Minimum Requirements

  • Experience optimizing machine learning models for resource-constrained environments, including knowledge distillation, model compression, quantization, and pruning.
  • Experience with on-device machine learning or edge AI, or mobile inference frameworks, including optimizing latency, memory, compute, and power constraints.
  • Experience distilling capabilities from large foundation models into small language models or task-specific models for efficient inference.
  • Experience building retrieval-augmented generation, vector search, embedding retrieval, neural reranking, or semantic search systems.
  • Experience with query understanding, query rewriting, intent classification, personalized retrieval, learning-to-rank, or recommendation models.
  • Experience with transformer architectures and foundation model families such as BERT, T5, Llama, Gemma, Mistral, or related architectures.
  • Experience evaluating language models, designing AI quality metrics, and building automated and human-in-the-loop evaluation pipelines.
  • Experience building large-scale production search, recommendation, personalization, or generative AI systems.
  • Familiarity with multimodal foundation models, tool use, agentic AI, or agentic retrieval systems.
  • Strong understanding of tradeoffs among model quality, latency, memory, power consumption, privacy, and reliability for production on-device AI systems.
  • Ability to prototype new ideas, run rigorous experiments, address ambiguous technical problems, and translate research into production-quality ML solutions.
  • Ability to work onsite in Cupertino, California, in accordance with Apple’s applicable work policies.
  • Programming skills in Python and/or C/C++, with production software experience using modern ML frameworks such as PyTorch, JAX, or TensorFlow.
  • Background in machine learning, deep learning, natural language processing, information retrieval, search, recommender systems, or generative AI.
  • Experience training, fine-tuning, or deploying transformer-based models and large language models.
  • Experience with modern deep learning architectures and techniques including transformers, embeddings, representation learning, and neural ranking.

Technologies

  • Python
  • C/C++
  • PyTorch, JAX, TensorFlow
  • Transformer architectures
  • BERT, T5, Llama, Gemma, Mistral
  • Vector search, embeddings
  • Neural reranking
  • Retrieval-augmented generation
  • Mobile inference frameworks
  • On-device machine learning, edge AI

Benefits

  • Comprehensive medical and dental coverage
  • Retirement benefits
  • Range of discounted products and free services
  • Reimbursement for certain educational expenses, including tuition
  • Discretionary bonuses or commission payments, as well as relocation (might be eligible)
  • Opportunity to become an Apple shareholder through participation in Apple’s discretionary employee stock programs
  • Eligible for discretionary restricted stock unit awards
  • Apple stock purchase at a discount via voluntary participation in the Employee Stock Purchase Plan

Pay & Benefits Notes

At Apple, base pay is one part of the total compensation package and is determined within a range. Benefit, compensation, and employee stock programs are subject to eligibility requirements and other terms of the applicable plan or program. Learn more about Apple Benefits.

Similar Jobs