Principal Machine Learning Engineer
Job Description
As a Principal Machine Learning Engineer on Oracle's Generative AI Services team, you will lead the architecture, design, and development of distributed, scalable infrastructure for training, fine-tuning, and inference. You will collaborate with partner teams and engineers to ensure robust production deployments and contribute to open source projects such as vLLM and SGLang.
Location
Seattle, WA (onsite)
Overview
This role focuses on building and refining high-performance AI infrastructure capable of supporting large scale generative workloads. You will provide architectural leadership and hands-on development across the model lifecycle, driving the adoption of advanced technologies and coordinating with internal partners to ship reliable, production-ready systems. Engagement with open source communities is also part of the scope, including contributions to vLLM and SGLang.
Responsibilities
- Lead the architecture, design, and development of distributed, scalable, high-performance systems for AI model training, fine-tuning, and inference.
- Build and optimize next-generation AI infrastructure that powers large-scale generative AI workloads.
- Direct analysis of model architectures to improve performance, efficiency, and scalability.
- Leverage cutting-edge technologies to develop state-of-the-art AI systems and onboard frontier models.
- Benchmark, diagnose, troubleshoot, and resolve issues across the AI model lifecycle, including training, fine-tuning, and serving, ensuring reliable production deployments.
- Contribute to open source frameworks such as vLLM and SGLang and strengthen Oracle Cloud Infrastructure's position in the ecosystem.
- Lead a team of senior and junior engineers, guiding delivery of the roadmap on time with high quality.
Technologies
- vLLM
- SGLang