DataJobs.io
← Back to all jobs

Job Description

Founding Principal Machine Learning Engineer at Arcade to design, train, evaluate, and productionize small, fast, on-prem, agentic tool recommendation and runtime ML models. This role owns the full model lifecycle for agent calls inside enterprise environments, from pipeline execution to model release and runtime readiness.

Role Overview

You will run Arcade’s ML pipeline end to end, ensuring model shipping is routine. You will fine-tune and train models spanning tool selection, routing, retrieval, and agent memory, while pushing the system into new territories such as search and recommendation over tools and agent context. You will also lead evaluation practices that combine offline assessments on real agent traces with online measurement in production, including head-to-head comparisons with Claude, GPT, and Gemini.

Key Responsibilities

  • Own the ML pipeline end to end, from data and training through evaluation and release.
  • Make model releases routine by improving reliability across training, evaluation, and deployment.
  • Fine-tune and train models for tool selection, routing, retrieval, and agent memory.
  • Apply methods such as distillation, embeddings, rerankers, and other approaches as appropriate.
  • Extend capabilities into new areas, beginning with search and recommendation over tools and agent context.
  • Build the evaluation system with offline evals on real agent traces, online measurement in production, and head-to-head comparisons with Claude, GPT, and Gemini.
  • Create a loop from production agent traces and tool-call data back into better models within enterprise data boundaries.
  • Quantize, optimize, and package models for customer VPCs and air-gapped environments, and collaborate with the Runtime team on serving.
  • Select the ML stack, define the strategy, and help shape hiring priorities for future ML hires.
  • Use depth in modern agent systems (harnesses, memory, skills) to identify where models improve agent outcomes, and call early whether approaches will work.
  • Drive speed improvements so projects that take a week today take a day next time.
  • Own the full model pipeline from data ingestion to model publishing for embeddings, reranking, classification, recognition, and other tasks.
  • Stand behind model performance using available able data, including reviewing outcomes row by row in a spreadsheet.
  • Ensure everything runs in customer environments including their cloud, hardware, and sometimes air-gapped networks.
  • Bring practical insights to implement in a stable but young model pipeline.
  • Lead development of agent capabilities beyond API calls, focusing on intelligent and efficient behaviors that outpace competitors.
  • Own models on the runtime path of every agent call served.
  • Productionize the first agent recommendation model and set forward-looking patterns for the ML stack.
  • Own build-buy decisions for the stack going forward with a healthy budget to spend.
  • Report directly to the Head of Engineering.

Minimum Requirements

  • 7+ years of software engineering experience, including 4+ years training and shipping production ML systems.
  • Trained or fine-tuned models that reached production and improved customer-relevant metrics, not only leaderboard results.
  • Production experience with agent systems including harnesses, memory, skills, tool use, and sub-agents.
  • Knowledge of fine-tuning: when it works and how to apply it.
  • Evaluation depth with statistics fluency to determine whether changes are meaningful or noise.
  • Experience deploying models under real constraints such as tight latency budgets, limited GPUs, or non-standard infrastructure (vLLM, ONNX, TensorRT, llama.cpp, or equivalent).
  • Strong Python for training and ML work, plus TypeScript or Go for production services that serve models.
  • Ability to select a stack, document decisions, and defend them over time.
  • Preference for shipping a working v0.5 over a polished v2.0 on a longer timeline.
  • Comfort with ambiguity in an early team environment where the charter expands and decisions are made with incomplete data.
  • A strong desire to ship.

Technologies

  • Python, TypeScript, Go
  • vLLM, ONNX, TensorRT, llama.cpp
  • Claude, GPT, Gemini
  • MCP

Compensation and Benefits

Compensation is aligned with the stated range and determined based on a candidate’s background, experience, and performance. Starting at $230,000 base salary, plus equity and competitive benefits, including in-person work at Arcade’s San Francisco office.

Traction and Market Context

  • Real deployments with Fortune-100 customers such as Morgan Stanley and Open Table.
  • Enterprises are racing to deploy agents in production, while most have not yet reached deployment.

Additional Founder and Team Context

The CEO previously founded Stormpath (acquired by Okta), where he created the first Authentication API for developers. The CTO led the vector database team at Redis, shipped 100+ LLM applications, and contributes to LangChain and LlamaIndex.

Arcade has assembled authentication, integrations, distributed systems, and AI experts from Okta, Redis, Microsoft, Splunk, Ngrok, Google, Airbyte, Disney, and HPE, with experience building and founding multiple successful developer platforms.

Backers

Arcade’s Series A is led by SYN Ventures, with strategic investment from Morgan Stanley and Wipro. Earlier investors have also backed Databricks, Clickhouse, MongoDB, Perplexity, Cohere, ScaleAI, Confluent, Elastic, and Firebase.

Similar Jobs