DataJobs.io
← Back to all jobs

Job Description

Tesla AI is hiring an AI Engineer, Model Distillation to design and execute distillation pipelines that compress large teacher models into efficient students for autonomy and robotics.

Responsibilities

  • Design and execute state-of-the-art distillation pipelines that transfer capabilities from large teacher models to ultra-efficient student models for FSD (Full Self-Driving), Optimus, and Digital Optimus
  • Transfer reasoning, vision-language understanding, and action policies into compact students
  • Develop and iterate on distillation methodologies across logit-level, sequence-level, and trajectory-level objectives
  • Pioneer novel distillation recipes combining synthetic data generation, soft-label training, supervised fine-tuning (SFT), and Reinforcement Learning (RL)
  • Run experiments to identify distillation scaling behavior, including teacher size, student capacity, student architecture design, data mixture, and compute allocation
  • Measure how design choices impact capability retention, latency, and on-robot performance
  • Build and maintain infrastructure for efficient teacher inference, distillation data curation, and distributed student training
  • Resolve compute and memory bottlenecks end-to-end
  • Evaluate distilled models for teacher fidelity and product metrics such as task success, robustness, and real-world behaviors
  • Close performance gaps when students diverge from teachers
  • Collaborate cross-functionally to ship distilled models to production while meeting performance, safety, and reliability standards
  • Contribute tools and frameworks to keep distillation reproducible, measurable, and reusable across Tesla AI model families

Requirements

  • Deep, proven expertise in deep learning fundamentals with experience in training, compressing, or distilling large-scale language, vision, or multimodal models
  • Strong grasp of teacher-student interaction, synthetic data, and tradeoffs between capability, latency, and compute
  • In-depth knowledge of loss design (example: KL / soft-target objectives) and modern architectures such as mixture of experts and hybrid attention
  • Hands-on experience with post-training methods including distillation, supervised fine-tuning, and policy optimization / RL
  • Expertise in distributed computing and large-scale training or inference pipelines
  • Proficiency in Python and knowledge of software engineering best practices
  • Experience with deep learning frameworks: PyTorch, TensorFlow, or JAX
  • Demonstrated ability to work in a cross-functional team environment
  • Strong problem-solving skills to troubleshoot complex issues across data, training, and deployment

Technologies

  • Python
  • PyTorch
  • TensorFlow
  • JAX
  • Mixture of experts
  • Hybrid attention

Compensation and Benefits

  • Expected compensation: $124,000 - $558,000/annual salary + cash and stock awards + benefits
  • Pay offered may vary based on market location and job-related knowledge, skills, and experience; total compensation may include other elements depending on the position offered
  • Medical plans: plan options with $0 payroll deduction
  • Family-building, fertility, adoption and surrogacy benefits
  • Dental (including orthodontic coverage) and vision plans with options at $0 paycheck contribution
  • Company paid HSA contribution when enrolled in the High-Deductible medical plan with HSA
  • Healthcare and Dependent Care Flexible Spending Accounts (FSA)
  • 401(k) with employer match, Employee Stock Purchase Plans, and other financial benefits
  • Company paid Basic Life and AD&D
  • Short-term and long-term disability insurance (90 day waiting period)
  • Employee Assistance Program
  • Sick and Vacation time (Flex time for salary positions, accrued hours for hourly positions) and Paid Holidays
  • Back-up childcare and parenting support resources
  • Voluntary benefits: critical illness, hospital indemnity, accident insurance, theft & legal services, and pet insurance
  • Weight Loss and Tobacco Cessation Programs
  • Tesla Babies program
  • Commuter benefits
  • Employee discounts and perks program

What to Expect

  • Work on real-world autonomy at scale by training frontier-scale foundation models on multimodal telemetry, vision, language, and robotic action data
  • Collaborate in an environment using driving, robotics, and human-interaction datasets spanning vision, language, action, and on-robot telemetry
  • Study how capability transfers from large teacher models into compact students under strict latency, power, and reliability constraints
  • Access high GPU resources per engineer to train frontier-scale teachers, generate large distillation corpora, and iterate on student architectures
  • Run distillation experiments across language, multimodal, and policy models with high fidelity and scale, then deploy distilled models to FSD / Optimus in real-world settings

Similar Jobs