Machine Learning Engineer - Multimodal Intelligence
Job Description
Apple’s Multimodal Intelligence team is building the systems that help turn multimodal foundation models into shipping Apple Intelligence features. In this role, you will help deliver model pipelines and production infrastructure that support end-to-end training, evaluation, and deployment across on-device and hybrid environments.
Based in Seattle, WA (onsite), this Machine Learning Engineer position focuses on taking multimodal models from curated data through optimization and testing, then hardening promising approaches into reliable, maintainable production systems that operate under real constraints.
What you’ll do
- Build the pipelines, infrastructure, and production systems that turn multimodal foundation models into shipping Apple Intelligence features.
- Own end-to-end model delivery, including scalable data curation and training pipelines, fine-tuning and optimization of large multimodal models for on-device and hybrid execution, and reproducible evaluation and regression testing for text and visual understanding.
- Harden approaches into robust, maintainable production systems while accounting for latency, memory, power, and privacy constraints.
- Collaborate closely with modeling, platform, hardware, and product engineering teams across Apple, incorporating future hardware design and product needs into implementation decisions.
- Work broadly with teams to help deliver strong product outcomes.
Required qualifications
- BS with a minimum of 3 years of relevant industry experience.
- Master’s or PhD, or equivalent practical experience, in Computer Science, Computer Vision, Machine Learning, or a related technical field.
- Deep expertise in multimodal foundation models with emphasis on practical applications.
- Proven ability to translate research into practical applications through published work or industry experience.
- Applied research experience in at least one major area of model development, such as data curation, pre-training, fine-tuning, alignment, or evaluation, particularly for multimodal systems.
- Experience building large-scale training pipelines, including working with large datasets and scaling models across distributed systems.
- Experience bridging research ideas with production constraints.
- Experience with deep learning demonstrated in at least one area of multimodal systems (for example: vision, language, video, or audio).
- Proficiency in Python and a modern deep learning framework such as PyTorch or JAX.
- Experience with rapid prototyping, reproduction, and validation of research ideas.
- Ability to work in a collaborative environment.
- Ability to communicate analysis results clearly and effectively.
Technology
- Python
- PyTorch
- JAX
Compensation and benefits
The base pay range for this role is $142,300 to $263,300 per year. Your base pay will depend on skills, qualifications, experience, and location.
- Comprehensive medical and dental coverage
- Retirement benefits
- Discounted products and free services
- Reimbursement for certain educational expenses, including tuition
- Discretionary restricted stock unit awards
- Eligibility to purchase Apple stock at a discount through the Employee Stock Purchase Plan (voluntary participation)
- Opportunity to become an Apple shareholder through participation in discretionary employee stock programs
- Discretionary bonuses or commission payments and relocation (eligibility may apply)