DataJobs.io
← Back to all jobs

Job Description

In this Senior Machine Learning Engineer role, you will help design and evolve Unity Vector’s online model inference platform. The focus is on production infrastructure for low-latency, high-reliability inference, with strong emphasis on optimization, observability, and safe experimentation.

Responsibilities

  • Design and operate large-scale online inference infrastructure that serves production ML models with low latency and high reliability, including solutions such as PyTorch, Triton Inference Server, Kubernetes, GKE, Ray, or comparable distributed serving frameworks.
  • Build infrastructure that supports distributed training workflows using tools such as Pytorch, Ray Data, and Ray Train.
  • Integrate ML pipelines with workflow orchestration systems (such as Flyte, Airflow, or similar) to support dependable multi-stage training workflows.
  • Optimize model performance through techniques including model compilation, GPU and CPU utilization improvements, request scheduling, kernel fusion, and runtime-level tuning.
  • Enhance ML system observability by monitoring latency, throughput, error rates, cost, saturation, and model health.
  • Collaborate with ML engineers to accelerate model iteration while maintaining production safety, scalability, and cost efficiency.
  • Improve reliability and reproducibility of model serving workflows, including model packaging, artifact validation, compatibility testing, and deployment automation.
  • Lead architectural improvements to increase robustness, usability, scalability, and cost efficiency for the online ML platform.

Requirements

  • Experience building and operating production-grade online ML inference systems, such as NVIDIA Triton Inference Server, TorchServe, Ray Serve, TensorFlow Serving, or similar platforms.
  • Experience with model serving frameworks including NVIDIA Triton Inference Server, TorchServe, Ray Serve, TensorFlow Serving, or similar systems.
  • Experience optimizing inference workloads using dynamic batching, model compilation, quantization, GPU acceleration, GPU kernel optimization, caching, or runtime tuning.
  • Strong background in distributed systems, Kubernetes, autoscaling, service reliability, and production observability.
  • Strong programming skills in Python, including practical experience in production ML systems and high-scale services.
  • Experience with PyTorch and modern model deployment workflows, covering model packaging, validation, and end-to-end serving lifecycle management.
  • Experience designing infrastructure for safe rollout approaches such as canary testing, A/B experimentation, and automated rollback.
  • Strong systems thinking to balance latency, throughput, reliability, scalability, and cost tradeoffs in online environments.
  • Demonstrated ability to lead technical direction and influence architectural decisions across teams without formal authority.

Technologies

  • PyTorch
  • NVIDIA Triton Inference Server
  • Triton Inference Server
  • Kubernetes
  • GKE
  • Ray
  • Ray Data
  • Ray Train
  • Flyte
  • Airflow
  • TorchServe
  • Ray Serve
  • TensorFlow Serving
  • Python

Benefits

  • Comprehensive health, life, and disability insurance
  • Commute subsidy
  • Employee stock ownership
  • Competitive retirement/pension plans
  • Generous vacation and personal days
  • Support for new parents through leave and family-care programs
  • Office food snacks
  • Mental Health and Wellbeing programs and support
  • Employee Resource Groups
  • Global Employee Assistance Program
  • Training and development programs
  • Volunteering and donation matching program

Location

Olympia, WA (onsite)

Compensation

Salary range (yearly): USD 165,600 - 273,400

  • Zone A: $210,300 - $273,400
  • Zone B: $187,200 - $243,300
  • Zone C: $165,600 - $215,200

Similar Jobs