DataJobs.io
← Back to all jobs

Job Description

LILT is building a live translation product, and this role focuses on the real-time speech translation backend. You will own the end-to-end path from live audio input to translated output, combining streaming speech recognition and adaptive machine translation in a low-latency system deployed on GPU Kubernetes infrastructure.

What you’ll do

  • Build and run services that support high-throughput, real-time audio and text streaming.
  • Own signal processing, session lifecycle management, and concurrency to keep the system stable under load.
  • Integrate and serve streaming speech recognition and machine translation models, working with research teams to meet latency budgets.
  • Develop model-driven confidence scoring and routing logic for human intervention when needed.
  • Broadcast real-time updates and corrections to end users.
  • Architect and scale production ML infrastructure on GPU-accelerated Kubernetes, including Ray Serve deployments.
  • Implement batching, load balancing, and autoscaling strategies to maintain both performance and cost efficiency.
  • Set up instrumentation to measure real-time performance and establish production-grade observability.
  • Identify bottlenecks, optimize end-to-end throughput, and reduce end-to-end latency to meet production standards.
  • Define technical contracts and interfaces for audio ingestion and downstream service integrations.
  • Collaborate with frontend and platform engineering teams to maintain robust integration points.
  • Drive cross-team alignment through clear API interfaces and technical contracts across engineering and product teams.

Key requirements

  • BS or MS in Computer Science (or related) or equivalent practical experience.
  • 3+ years building production backend or ML serving systems in Python, with strong async skills (asyncio).
  • Hands-on experience with real-time streaming transport such as WebSocket or gRPC bidirectional streaming, including session state, backpressure, and connection lifecycle handling.
  • Experience serving ML models on GPUs in production using Ray Serve, Triton, vLLM, or similar, along with Docker and Kubernetes.
  • Experience integrating speech or NLP models into production systems, ideally streaming ASR with partial hypotheses, endpointing, and VAD.
  • A latency-engineering mindset: you have profiled, instrumented, and optimized real-time or low-latency systems and can reason using per-stage budgets.
  • Ability to use AI coding agents (Claude Code, Codex, or similar) with strong fundamentals: you can debug, review, and reason about every line and know when not to trust generated output.
  • US citizenship and residence in the United States (contract requirement).

Technologies

  • Python, asyncio, WebSocket, gRPC
  • Ray Serve, Triton, vLLM
  • Docker, Kubernetes
  • Claude Code, Codex
  • COMET, CometKiwi
  • RabbitMQ
  • Datadog, Prometheus
  • WebRTC, SFU
  • LiveKit Agents, Pipecat

Location and eligibility

  • Boston, MA (onsite)
  • Requires US citizenship and residence in the United States.
  • Preferred locations include Washington, D.C.; Boston, MA; and Indianapolis, IN (East Coast / ET timezone preferred).

Compensation

USD 120,000 - 161,434 per yearly.

Preferred qualifications

  • Ray Serve experience, including streaming responses and model multiplexing.
  • Familiarity with simultaneous or incremental MT concepts (retranslation, prefix stability, wait-k policies).
  • Machine translation quality estimation in the COMET/CometKiwi class, or other production confidence estimation.
  • Message brokers for real-time fan-out and state distribution (RabbitMQ or similar).
  • Streaming text-to-speech integration and time-to-first-audio optimization.
  • WebRTC and SFU concepts, or voice pipeline frameworks such as LiveKit Agents or Pipecat.
  • Handling CJK and other non-Latin text in NLP pipelines (Japanese, Korean, and English are first languages).
  • Observability tooling for production ML systems (Datadog, Prometheus).

Similar Jobs