Lead Machine Learning Engineer
Manager
Agentic Ai
Ai Agent
Ai Agent Platform
Ai Ml
Artificial Intelligence
Automation
Azure Ai
Azure Openai
CI/CD
Cloud
Cloud Infrastructure
Cloud Native
Cloud Operations
Cloud Platform
Cloud Platforms
Cloud Technology
Data Analysis
Data Architecture
Data Platform
Database
Databases
DevOps
Devops Tools
DevSecOps
Engineer
Engineering
Generative AI
Generative Ai Platform
Google Cloud
Infrastructure
Infrastructure As Code
Kubernetes
Machine Learning Engineer
Machine Learning Infrastructure
Ml Ops
Platform Engineering
Programming
Programming Language
Programming Languages
Rag Architectures
Security Automation
Software Development
Technical Lead
Job Description
Motion Recruitment is hiring a Lead Machine Learning Engineer in Raleigh, NC to help architect scalable AI and agentic platform capabilities for a global AI platform powering LLM-powered research assistants, retrieval systems, and enterprise-grade agent workflows. In this role, you will shape AI platform strategy, technical standards, and governance so the systems operate reliably at scale while aligning with responsible AI principles.
What you’ll do
- Architect scalable AI platforms, including defining a reference architecture for LLM, ML, and agent-based systems across products
- Design high-availability, low-latency inference platforms for global scale
- Establish reusable platform components to support the model lifecycle, deployment, and monitoring
- Architect multi-step, reasoning-driven agent systems
- Design orchestration patterns for tool use, API invocation, and structured function calling
- Lead implementation and governance of Model Context Protocol (MCP) servers to standardize tool integration and context management
- Define guardrails, permissions, and audit mechanisms for enterprise-safe AI systems
- Set best practices for MLOps, CI/CD, observability, and system reliability
- Embed Responsible AI principles across platform architecture
- Mentor senior engineers and influence technical direction across teams
Required experience and education
- 10+ years of experience with a Master’s degree, or 12+ years of experience with a bachelor’s degree
- 10+ years building production-grade ML systems at scale
- Extensive experience with LLMs, generative AI, and RAG systems in real-world deployments
- Proven expertise designing distributed systems in cloud environments (AWS, Azure, or GCP)
- Hands-on experience with Kubernetes, containerization, and scalable inference systems
- Experience designing agentic systems and tool orchestration frameworks
- Experience implementing or governing MCP servers or structured tool-calling architectures
- Strong Python engineering background
- Experience with vector databases and search systems
- Deep understanding of model evaluation, reliability, and monitoring
- Strong architectural judgment and systems thinking
- Leadership experience influencing technical direction across teams
- Strong communication skills and executive presence
- Experience mentoring senior engineers or leading cross-functional initiatives
Technologies
- LLMs, RAG systems, Python
- AWS, Azure, GCP
- Kubernetes, containerization
- Model Context Protocol (MCP)
- Vector databases, search systems
- MLOps, CI/CD, observability