Senior Machine Learning Engineer (AI Foundations)
Job Description
The Senior Machine Learning Engineer will join Capital One's AI Foundations team to productionize machine learning applications at scale, design robust ML architectures, and ensure the high availability and performance of ML systems. This role encompasses end-to-end ML engineering work, including data pipelines, cloud infrastructure, CI/CD, and responsible AI practices.
Responsibilities
- Design, build, and deploy ML models and components to address real world business needs, in collaboration with Product and Data Science teams.
- Guide ML infrastructure choices using knowledge of modeling techniques, including model selection, data and feature decisions, training, hyperparameters, dimensionality, bias/variance, and validation.
- Tackle complex problems by writing and testing application code, developing and validating ML models, and automating tests and deployment.
- Collaborate within a cross functional Agile team to create software that enables state-of-the-art big data and ML applications.
- Retrain, maintain, and monitor models in production.
- Leverage or build cloud-based architectures, technologies, and platforms to deliver optimized ML models at scale.
- Construct optimized data pipelines to feed ML models.
- Apply continuous integration and continuous deployment best practices, including test automation and monitoring, to ensure successful deployment of ML models and application code.
- Ensure code quality and governance, manage risk, and follow responsible and explainable AI practices.
- Utilize programming languages such as Python, Scala, or Java.
Requirements
- Bachelor’s Degree
- At least 4 years of experience programming with Python, Scala, or Java (internship experience does not apply)
- At least 3 years designing and building data‑intensive solutions using distributed computing
- At least 2 years of on‑the‑job experience with ML frameworks (scikit-learn, PyTorch, Dask, Spark, or TensorFlow)
- At least 1 year of experience productionizing, monitoring, and maintaining models
Technologies
- Python, Scala, Java
- scikit-learn, PyTorch, Dask, Spark, TensorFlow
- AWS, Azure, Google Cloud Platform
Benefits
- Health benefits
- Financial benefits
- Performance-based incentive compensation, including cash bonuses and long-term incentives
What you’ll do
- Design, build, and deliver ML models and components that solve real world business problems, in collaboration with Product and Data Science teams.
- Inform ML infrastructure choices with expertise in modeling techniques, including model selection, data and feature choices, training, hyperparameters, dimensionality, bias/variance, and validation.
- Address complex problems by authoring and testing application code, developing and validating ML models, and automating tests and deployment.
- Work within a cross-functional Agile team to create software that enables cutting‑edge big data and ML capabilities.
- Retrain, maintain, and monitor models in production.
- Utilize or build cloud-based architectures and platforms to deliver optimized ML models at scale.
- Develop optimized data pipelines to feed ML models.
- Apply CI/CD practices with test automation and monitoring to support successful deployments of ML models and code.
- Ensure code quality and governance, mitigate vulnerabilities, and adhere to responsible and explainable AI standards.
- Work with Python, Scala, or Java as primary programming languages.
Basic Qualifications
- Bachelor’s Degree
- At least 4 years programming with Python, Scala, or Java (internship experience does not apply)
- At least 3 years designing and building data‑intensive solutions using distributed computing
- At least 2 years on-the-job experience with ML frameworks (scikit-learn, PyTorch, Dask, Spark, or TensorFlow)
- At least 1 year of experience productionizing, monitoring, and maintaining models
Preferred Qualifications
- 1+ years building, scaling, and optimizing ML systems
- 1+ years gathering and preparing data for ML models
- 2+ years developing performant, resilient, and maintainable code
- Experience deploying ML solutions in a public cloud such as AWS, Azure, or Google Cloud Platform
- Master’s or doctoral degree in computer science, electrical engineering, mathematics, or a related field
- 3+ years working with distributed file systems or multi-node database paradigms
- Contributed to open source ML software
- Authored or co-authored a paper on a ML technique, model, or concept
- 3+ years building production-ready data pipelines that feed ML models
- Experience designing, implementing, and scaling complex data pipelines for ML models and evaluating their performance
- Experience leveraging interactive AI tooling to accelerate productivity beyond basic code completion