Staff Machine Learning Engineer
Job Description
Xometry is hiring a Staff Machine Learning Engineer (senior individual contributor) to lead end-to-end ML systems work for the DFM AI + IQE initiative and partner integrations.
Responsibilities
- Own the full ML lifecycle, from requirements gathering through release, delivering high-quality results on-time across complex, cross-functional initiatives
- Lead the partner integration AI/ML plane for the embedded DFM AI + IQE integration with Teamcenter and Designcenter
- Design real-time ML serving architecture and implement a low-latency signal path that delivers DFM and pricing feedback inside the designer environment
- Define data contracts for model inputs and outputs, and implement MLOps, governance, and observability for a mission-critical partner integration
- Develop cloud-based production systems for real-time endpoints and MLOps, integrated with Xometry’s broader systems and infrastructure
- Handle cross-domain technical complexity by evaluating variable factors and aligning solutions to both business and technical objectives
- Identify opportunities, drive new processes and solutions, and create multi-quarter roadmaps for key technical goals
- Apply best practices for automated testing, parallel and distributed computing, and secure software development across ML systems
- Collaborate with engineers, product managers, data scientists, and business stakeholders to translate requirements into robust technical solutions
- Conduct and support design reviews, code reviews, and technical mentorship to raise team capability
- Stay current with ML/AI advances and integrate relevant tools, frameworks, and approaches into production systems
Requirements
- Bachelor’s degree in a STEM field (or equivalent experience) plus 6-8 years of machine learning engineering experience
- Proven experience owning and delivering complex ML systems in production
- Deep expertise in ML and AI, including Gradient Boosting, Deep Learning, and/or Generative AI frameworks, with a focus on backend scalability and reusability
- Hands-on experience deploying real-time ML products at scale in cloud environments (AWS strongly preferred), including auto-scaling, monitoring, and alerting
- Strong proficiency in Python and advanced ML/AI frameworks such as TensorFlow or PyTorch
- Solid software engineering fundamentals, including data structures and algorithms
- Demonstrated MLOps experience: model monitoring, data and concept drift detection, and automated retraining and redeployment pipelines
- Experience with CI/CD (example: GitHub Actions), test driven development, and infrastructure as code (example: Terraform)
- Experience profiling and optimizing ML model deployments for latency and throughput
- Ability to operate independently on new and ambiguous assignments and communicate effectively across engineering, product, and business audiences
- Experience with state-of-the-art modeling techniques including transformers, self-supervised pre-training, large language models (LLMs), or generative AI
- Knowledge of containers, Kubernetes, and cloud-native distributed systems
- Background in manufacturing, supply chain, or marketplace environments is a plus
Location
- Denver, CO (hybrid)
Salary
- USD 200,000 - 220,000 per year
Technology Stack
- Python, TensorFlow, PyTorch
- Gradient Boosting, Deep Learning, Generative AI frameworks
- AWS, CI/CD pipelines, GitHub Actions, test driven development, Terraform (infrastructure as code)
- Transformers, self-supervised pre-training, large language models (LLMs)
- Containers, Kubernetes, cloud-native distributed systems
- Solid Edge, NX, Designcenter, Teamcenter
- MLOps, model monitoring, data drift detection, concept drift detection, automated retraining and redeployment pipelines
Benefits
- 401(k) match
- Medical, dental and vision insurance
- Life and disability insurance
- Generous paid time off including vacation, sick leave, floating and fixed holidays, maternity and bonding leave
- EAP and other wellbeing resources