J
Lead Machine Learning Engineer-MLOps
Job Description
Join JPMorgan Chase as a Lead Machine Learning Engineer MLOps on the Recommendation Engine team, working onsite in Palo Alto, CA. You will build and deploy ML models using a modern MLOps stack, spanning distributed training on GPU clusters, real-time and batch inference, and scalable hyperparameter tuning with production-grade validation. The role offers a competitive base salary in the range of USD 164,350 to 260,000 per year.
Responsibilities
- Design, deploy, and maintain robust pipelines for distributed training on GPU-enabled clusters to enable scalable ML workflows.
- Develop pipelines for high throughput real time as well as batch inference, prioritizing performance and reliability.
- Implement quantization techniques and deploy large language models to improve efficiency across specific GPU architectures.
- Oversee the management and optimization of vector databases to support advanced AI and ML applications.
- Establish and maintain comprehensive monitoring and observability pipelines to ensure system health, performance, and rapid issue resolution.
- Collaborate with product, architecture, and other engineering teams to integrate new technologies and continuously improve existing infrastructure.
- Partner with cross functional teams to define scalable and performant technical solutions.
Requirements
- BS in Computer Science or related Engineering field with 6+ years of experience.
- MS in Computer Science or related Engineering field with 4+ years of experience.
- Strong proficiency in Python and cloud computing, preferably AWS.
- Understanding of quantization techniques such as PTQ and AWQ used to accelerate LLM inference on specific GPU architectures.
- Experience in systems engineering fundamentals including caching, CUDA, autoscaling, high throughput, low latency, and cross-region resilience.
- Solid knowledge and passion for data science fundamentals, including training and deploying models.
- Experience with monitoring and observability tools to track model inputs/outputs and feature statistics.
- Operational experience with big data and ML tools such as Ray, DuckDB, Spark, and with training/inference systems like Ray and vllm/SGLang.
- Strong engineering fundamentals and an analytical mindset.
Technologies
- Python
- AWS
- CUDA
- PTQ
- AWQ
- Ray
- DuckDB
- Spark
- vllm
- SGLang
- Docker
- Kubernetes
- ECS
- Airflow
- Kubeflow
- Vector databases
Benefits
- Base salary determined by role, experience, skills, and location.
- Commission-based pay and discretionary incentive compensation, paid in cash and/or equity, recognizing individual contributions.
- Comprehensive health care coverage.
- On site health and wellness centers.
- Retirement savings plan.
- Backup childcare.
- Tuition reimbursement.
- Mental health support.
- Financial coaching.
Similar Jobs
J
J
J