DataJobs.io
← Back to all jobs

This position is no longer accepting applications

Closed on September 13, 2026.

This role is filled — get an email when new Data Processing roles open on DataJobs.io:

Job Description

Hybrid in the New York City Metro area (3 days onsite). In this Principal Machine Learning Engineering role, you will own the ML infrastructure behind real-time compliance enforcement systems. Expect a position centered on building the pipelines, evaluation workflow, and production serving needed to train, measure, and deploy models reliably, with a focus on latency, reliability, and cost.

Pay range: USD $200,000 - $250,000 per year.

What you’ll do

  • Build and own training pipelines including data preparation, reproducible fine-tuning runs, experiment tracking, and release automation.
  • Develop evaluation infrastructure with automated eval runs, regression gates, dashboards, and dataset versioning.
  • Own production model serving for low-latency inference, including batching, optimization, autoscaling, and cost management.
  • Ship model updates safely using versioning, canarying, rollback, and drift monitoring.
  • Create repeatable workflows to adapt models to new domains and changing customer needs.
  • Convert expert labels and reviewer feedback into clean training and evaluation datasets.
  • Help raise the team’s engineering bar for ML infrastructure as the organization grows.

What you’ll bring

  • 8+ years of software engineering experience, including 4+ years building ML or LLM infrastructure for production.
  • Hands-on experience with the modern LLM stack: PyTorch, distributed training, and fine-tuning at scale (e.g., LoRA, SFT) using inference engines such as vLLM or TensorRT-LLM.
  • Experience building eval harnesses, regression gates, or dataset pipelines, with strong understanding of precision, recall, and calibration.
  • Proven ownership of production model serving with real latency, reliability, and cost constraints.
  • Strong fundamentals in Python, containers, CI/CD, cloud infrastructure, and observability.
  • Ability to scope work, ship frequently, and make pragmatic build-vs-buy decisions.
  • Experience collaborating tightly with research partners and defining clear interfaces.

Technologies you’ll work with

  • PyTorch, LoRA, SFT, vLLM, TensorRT-LLM
  • Python, containers, CI/CD, cloud infrastructure, observability

Additional qualifications that may help

  • Experience productionizing small or specialized language models.
  • Experience with structured-output serving or constrained decoding in production.
  • Prior work in regulated or high-stakes domains (fintech, healthcare, legal, trust and safety).
  • Experience deploying models into customer-controlled environments.

Work location: Hybrid remote in New York, NY 10001 (3 days onsite).

This role may fill quickly. Submit your resume to be considered.

Similar Jobs