DataJobs.io
← Back to all jobs

Job Description

Placement Force seeks a Senior Machine Learning Engineer to own the complete lifecycle of a production large language model, from initial training and fine-tuning to scalable deployment and rigorous evaluation.

Responsibilities

  • Design and run pretraining, continued pretraining, and fine-tuning workflows, including supervised fine-tuning, alignment methods such as DPO and RLHF, as well as LoRA/QLoRA and full-parameter training, all tied to well-defined, measurable goals.
  • Own data pipeline decisions that materially affect model quality, covering data curation, deduplication, mixture weighting, and contamination checks against evaluation datasets.
  • Execute and interpret distributed training across multiple GPUs and nodes using FSDP, DeepSpeed, or Megatron-style parallelism, diagnosing scale-specific issues such as loss spikes, stragglers, and checkpoint corruption.
  • Make and justify tradeoffs between model size, training cost, and downstream performance.
  • Deploy trained models to production with quantization, batching strategies, KV-cache management, and serving framework selection (for example vLLM, TensorRT-LLM, TGI), accompanied by explicit latency, throughput, and cost targets.
  • Design for LLM serving failure modes, including tail latency under load, graceful degradation, prompt injection risk, and safe fallback behavior.
  • Develop the operational capabilities around production models, including monitoring, alerting, and rollback paths, applying the same rigor as for other critical services.
  • Create and maintain evaluation harnesses that go beyond published benchmarks, incorporating task-specific evaluation sets aligned to real product use cases.
  • Establish human evaluation protocols where automated metrics fall short, and determine the appropriate contexts for each approach.
  • Own regression detection to catch quality drops caused by new checkpoints, prompt template changes, or serving optimizations before they impact users.
  • Contribute to safety and robustness evaluation, including measures of hallucination, adversarial testing, and behavior under distribution shift, as an integral part of the release process.

Requirements

  • More than four years of applied ML engineering experience, with at least two years directly working on large language models in production.
  • Proven production deployment experience shipping models that served live traffic, with demonstrated awareness of latency, cost, and quality tradeoffs.
  • Strong software engineering fundamentals.
  • Fluency with the modern LLM tooling landscape, including training frameworks, serving frameworks, and evaluation tooling.
  • Comfort with ambiguity and the ability to define what constitutes acceptable model behavior in the absence of established benchmarks.

Technologies

  • FSDP
  • DeepSpeed
  • Megatron-style parallelism
  • vLLM
  • TensorRT-LLM
  • TGI

About the Role

This role involves owning the full lifecycle of a production large language model, spanning training and fine-tuning through deployment at scale and rigorous evaluation of what actually ships. It requires fluency across research, engineering, and production, because the most challenging problems emerge at the intersections of these domains.

Similar Jobs