Sr Lead Machine Learning Engineer
Ai Ml
Artificial Intelligence
Automation
Big Data
Bigdata
Cloud Infrastructure
Cloud Native
Cloud Platform
Cloud Platforms
Cloud Technology
Data & Ai
Data Analysis
Data Analytics
Data Engineer
Data Platform
Data Processing
Deep Learning
DevOps
DevSecOps
Engineer
Engineering
Kubernetes
Machine Learning
Machine Learning Engineer
Platform Engineering
Job Description
Capital One is hiring a Sr Lead Machine Learning Engineer for an onsite role in McLean, VA. You’ll work with an Agile team that focuses on productionizing machine learning applications and systems at scale, with a focus on real-time decisioning across customers’ credit journeys.
In this position, you’ll design, develop, and deploy machine learning models and infrastructure using modern Python and Kubernetes-based technology. The work includes building the software and data pipelines that support big data and ML capabilities, then keeping models healthy in production through retraining, monitoring, and governance.
What you’ll do
- Design, build, and deliver machine learning models and components to address real-world business needs in collaboration with Product and Data Science teams
- Make ML infrastructure decisions informed by modeling considerations such as model and data selection, feature selection, training, hyperparameter tuning, dimensionality, bias/variance, and validation
- Solve complex problems by writing and testing application code, developing and validating ML models, and automating tests and deployment
- Work in a cross-functional Agile environment to create and enhance software enabling state-of-the-art big data and ML applications
- Retrain, maintain, and monitor ML models once deployed
- Use cloud-based architectures, technologies, and platforms to deliver optimized ML models at scale
- Construct optimized data pipelines to supply ML models
- Apply CI/CD best practices, including test automation and monitoring, to support reliable deployments of ML models and application code
- Support secure, well-managed code to reduce vulnerabilities, ensure risk-aware governance of models, and follow Responsible and Explainable AI best practices
- Use programming languages including Python, Scala, or Java
Requirements
- Bachelor’s Degree
- 8+ years of experience designing and building data-intensive solutions using distributed computing (internship experience does not apply)
- 4+ years of experience programming with Python, Scala, or Java
- 3+ years of experience building, scaling, and optimizing ML systems
- 2+ years of experience leading teams developing ML solutions
Technologies
- Python, Kubernetes, Scala, Java
- AWS, Azure, Google Cloud Platform
- scikit-learn, PyTorch, Dask, Spark, TensorFlow
Benefits
- Eligible to earn performance-based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI)
- Comprehensive, competitive, and inclusive set of health, financial, and other benefits
Preferred qualifications
- Master’s or Doctoral Degree in computer science, electrical engineering, mathematics, or a similar field
- Experience developing and deploying ML solutions in a public cloud such as AWS, Azure, or Google Cloud Platform
- 4+ years of on-the-job experience with an industry recognized ML framework such as scikit-learn, PyTorch, Dask, Spark, or TensorFlow
- 3+ years of experience developing performant, resilient, and maintainable code
- 3+ years of experience with data gathering and preparation for ML models
- 3+ years of people management experience
- ML industry impact through conference presentations, papers, blog posts, open source contributions, or patents
- 3+ years of experience building production-ready data pipelines that feed ML models
- Ability to communicate complex technical concepts clearly to a variety of audiences
- Experience leveraging interactive AI tooling to accelerate productivity, using capabilities beyond basic code completion
Compensation: USD 229,900 - 262,400 per year.