Machine Learning Engineer 5 (Senior Manager, IC)
Job Description
Capital One is building AI-powered risk management capabilities, and this role supports that effort within Risk Tech by partnering with the GRC team and company partners. You will help design, build, deploy, and operate machine learning models and governed Responsible/Explainable AI solutions, with an emphasis on production monitoring and scalable multi-tenant platforms.
This Machine Learning Engineer 5 (Senior Manager, IC) position is based in Richmond, VA (onsite), with an annual salary range of USD 209,000 - 238,500. The role is eligible for performance-based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI).
What you’ll do
- Design, build, and/or deliver ML models and components that address real business problems in collaboration with Product and Data Science.
- Build and scale multi-tenant platforms to support large footprint ML model training and/or serving at scale.
- Apply ML expertise to guide infrastructure and modeling decisions, including model choice, data and feature selection, training, hyperparameter tuning, dimensionality, and bias/variance considerations and validation.
- Solve complex problems by writing and testing application code, developing and validating ML models, and automating tests and deployment.
- Work on an Agile cross-functional team to create and improve software for big data and ML applications.
- Retrain, maintain, and monitor models in production.
- Leverage or build cloud-based architectures and platforms to deliver optimized ML models at scale.
- Construct optimized data pipelines that feed ML models.
- Use CI/CD best practices, including test automation and monitoring, to support successful deployments of both ML models and application code.
- Ensure well-managed code to reduce vulnerabilities, and support risk-governed model governance aligned with Responsible and Explainable AI best practices.
- Use programming languages such as Python, Scala, or Java.
Minimum qualifications
- Bachelor’s Degree or higher in Computer Science, Machine Learning, or a related quantitative field (Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering).
- At least 6 years of experience programming with Python, Java, Golang, or C++.
- At least 6 years of ML experience using industry standard frameworks PyTorch or TensorFlow and libraries including Pandas, NumPy, Scikit-learn.
- At least 6 years using and operating large scale distributed systems (Spark, Ray) to prepare AI/ML data.
- At least 4 years deploying and operating ML solutions in production and operating production services in the cloud (AWS, GCP, Azure), including using Kubernetes to manage large scale containerized ML systems.
Preferred qualifications
- Master’s or doctoral degree in computer science, electrical engineering, mathematics, or related field.
- 5+ years of experience optimizing ML algorithms, configurations, and infrastructure.
- 5+ years of experience following software development best practices including source control, testing, code reviews, and CI/CD.
- 5+ years of experience building resilient software solutions with pre-production testing, advanced deployment techniques (one-box, blue/green, gradual dial-up), monitoring, alarms, and incident response plan preparation.
- 5+ years of experience with ML techniques (Supervised, semi-supervised, unsupervised, reinforcement learning, etc.), model types (Regression, Classification, Clustering, etc.), architectures (RNNs, CNNs, LSTMs, Transformers), training concepts (loss function, hyperparameters, regularization), and diagnosing common issues such as underfitting and overfitting.
- 5+ years of experience designing, implementing, and scaling production-ready data pipelines for training and evaluating ML models.
- ML industry impact through conference presentations, papers, blog posts, open source contributions, or patents.
- Ability to communicate complex technical and machine learning concepts clearly to a variety of audiences.
Technology you’ll work with
Python, Scala, Java, Golang, C++, PyTorch, Tensorflow, Pandas, NumPy, Scikit-learn, Spark, Ray, AWS, GCP, Azure, Kubernetes
Benefits
- Performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI).
- Comprehensive, competitive, and inclusive health, financial, and other benefits that support your total well-being.