Staff Machine Learning Engineer
Job Description
Unity Technologies SF is hiring a Staff Machine Learning Engineer to build production-grade AI for game experiences with a focus on computer vision and multi-modal modeling.
Responsibilities
- Set technical vision and roadmap for computer vision and multi-modal AI models, covering transformers, diffusion models, vision-language models, and JEPA-style generative architectures
- Design and implement models for image and video understanding and generation, including segmentation, detection, and dense prediction
- Develop multi-modal reasoning over images, text, and 3D inputs
- Make architecture, training, data pipeline, and evaluation trade-offs balancing quality, capability, latency, and cost across cloud, server, and on-device targets
- Drive research-to-production delivery: training, fine-tuning, distillation, export, and serving for deployment scenarios from cloud GPUs to efficient on-device inference
- Partner with research scientists to translate novel CV and multi-modal architectures into deployable, well-engineered implementations
- Build scalable multi-modal inference systems that ingest diverse inputs (images, video, text, primitives, and metadata) and produce outputs ranging from semantic predictions to pixel-level generation
- Monitor and adopt field breakthroughs including vision-language pretraining and alignment, efficient diffusion approaches (consistency models, flow matching), efficient attention (FlashAttention, linear-attention variants), and vision tokenization/representation learning
- Where needed for latency or device constraints, apply compression and optimization such as compression, quantization, pruning, and knowledge distillation, and integrate runtimes like TensorRT, ONNX Runtime, CoreML, and TFLite
- Lead and mentor ML engineers, establishing engineering best practices, code review standards, and rigorous benchmarking and evaluation methodology
- Collaborate with research, platform engineering, product managers, and runtime teams to align ML capabilities to product roadmaps and target-platform constraints
- Define and enforce measurement practices using KPIs for model quality, accuracy, latency, memory, and cost
Requirements
- 6+ years of ML engineering with strong depth in computer vision and/or multi-modal modeling
- Production experience with transformer-based and diffusion-based vision models (examples: ViT, CLIP/SigLIP-style encoders, Stable Diffusion, DETR/SAM-style architectures)
- End-to-end model lifecycle experience including data curation, training and fine-tuning, evaluation, and serving at scale
- Familiarity with efficient attention, diffusion samplers, multi-modal fusion, and vision-language alignment methods
- Strong Python skills and modern deep-learning tooling such as PyTorch, plus solid software engineering fundamentals
- Proven technical leadership: setting direction, influencing cross-functional partners, and growing engineers
Technologies
- Python, PyTorch
- TensorRT, ONNX Runtime, CoreML, TFLite
- FlashAttention
- ViT, CLIP, SigLIP
- Stable Diffusion
- DETR, SAM
Benefits
- Comprehensive health, life, and disability insurance
- Commute subsidy
- Employee stock ownership
- Competitive retirement/pension plans
- Generous vacation and personal days
- Support for new parents through leave and family-care programs
- Office food snacks
- Mental Health and Wellbeing programs and support
- Employee Resource Groups
- Global Employee Assistance Program
- Training and development programs
- Volunteering and donation matching program
Additional information
- Location: Mountain View, CA (onsite)
- Salary (USD per year): USD 172,200 - 283,900
- Zone A: $218,400 - $283,900
- Zone B: $194,100 - $252,300
- Zone C: $172,200 - $223,900
- Beyond base salary, the role may be eligible for equity awards and participation in company incentive plans (including annual discretionary bonuses or sales commissions)
- Final offer depends on geographic location, relevant experience, professional background, and skill set
You might also have
- Experience with world-model, video-generation, or neural rendering pipelines (NeRF, 3DGS, or similar)
- Experience deploying models to constrained or on-device targets, including quantization (INT8/INT4/FP16), pruning, distillation, and runtimes such as CoreML, TFLite, ONNX
- Familiarity with mobile SoC accelerators (Apple Neural Engine, Qualcomm Hexagon/Adreno, ARM Mali) or compiler stacks such as MLIR, TVM, or XLA
- Contributions to open-source ML frameworks or peer-reviewed CV/ML research publications
- Background in real-time graphics or game engine pipelines (Metal, Vulkan, OpenGL ES)