Machine Learning Engineer 5 (Senior Manager, IC)
Manager
Ai Ml
Application Security
Artificial Intelligence
Automation
Big Data
Cloud
Cloud Infrastructure
Cloud Native
Cloud Operations
Cloud Platform
Cloud Platforms
Cloud Platforms Cloud Platforms
Cloud Technology
Data Analysis
Data Platform
Data Science
DevOps
DevSecOps
Engineer
Engineering
Engineering Software
Facilities Management
Information Technology (IT)
Infrastructure
Infrastructure As Code
Kubernetes
Machine Learning Engineer
Machine Learning Engineering
Machine Learning Inference
Machine Learning Operations
Platform Engineering
Programming
Programming Language
Risk Management
Security Automation
Software Security
Job Description
Build and deploy AI-powered risk management solutions by designing ML models, scalable platforms, and production-ready pipelines.
Responsibilities
- Design, build, and/or deliver machine learning models and components to address real-world business problems with Product and Data Science
- Build and scale massive multi-tenant platforms for large-footprint ML model training and/or serving
- Drive ML infrastructure decisions using knowledge of modeling techniques and considerations, including model choice, data and feature selection, training, hyperparameter tuning, dimensionality, bias/variance, and validation
- Solve complex problems by writing and testing application code, developing and validating ML models, and automating tests and deployment
- Collaborate in cross-functional Agile teams to create and enhance software for big data and ML applications
- Retrain, maintain, and monitor models in production
- Leverage or build cloud-based architectures, technologies, and/or platforms to deliver optimized ML models at scale
- Construct optimized data pipelines to feed ML models
- Apply continuous integration and continuous deployment best practices, including test automation and monitoring, for successful deployment of ML models and application code
- Manage code to reduce vulnerabilities; ensure risk-governed models and adherence to Responsible and Explainable AI best practices
- Use programming languages including Python, Scala, or Java
Requirements
- Bachelor’s Degree or higher in Computer Science, Machine Learning, or a related quantitative field (Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering)
- At least 6 years of experience programming with Python, Java, Golang, or C++
- At least 6 years of machine learning experience using industry-standard frameworks PyTorch or Tensorflow and libraries Pandas, NumPy, Scikit-learn
- At least 6 years operating and using large scale distributed systems (Spark, Ray) to prepare AI/ML data
- At least 4 years deploying and operating machine learning solutions in production, including cloud production services (AWS, GCP, Azure) and using Kubernetes to manage large-scale containerized ML systems
Technologies
- Python, Scala, Java, Golang, C++
- PyTorch, Tensorflow
- Pandas, NumPy, Scikit-learn
- Spark, Ray
- AWS, GCP, Azure
- Kubernetes
Preferred Qualifications
- Master’s or doctoral degree in computer science, electrical engineering, mathematics, or a related field
- 5+ years optimizing ML algorithms, configurations, and infrastructure
- 5+ years following software development best practices including source control, testing, code reviews, and CI/CD
- 5+ years building resilient software with pre-production testing, advanced deployment techniques (one-box, blue/green, gradual dial-up), monitoring, alarms, and incident response plan preparation
- 5+ years working with machine learning techniques (Supervised, semi-supervised, unsupervised, reinforcement learning) and model types (Regression, Classification, Clustering)
- 5+ years working with model architectures (RNNs, CNNs, LSTMs, Transformers) and training concepts (loss function, hyperparameters, regularization), including evaluating accuracy and diagnosing underfitting and overfitting
- 5+ years designing, implementing, and scaling production-ready data pipelines for training and evaluating ML models
- ML industry impact through conference presentations, papers, blog posts, open source contributions, or patents
- Ability to communicate complex technical and machine learning concepts to a variety of audiences
Compensation & Location
- McLean, VA: $229,900 - $262,400 for Machine Learning Engineer 5
- Richmond, VA: $209,000 - $238,500 for Machine Learning Engineer 5
- Location listed: McLean, VA (onsite)
- Salary range provided: USD 209,000 - 262,400 per year
Incentives
- Eligible to earn performance-based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI)
- Incentives may be discretionary or non-discretionary depending on the plan
Additional Notes
- Applications expected to accept for a minimum of 5 business days
- No agencies please
- Equal opportunity employer (EOE, including disability/vet) committed to non-discrimination under applicable laws
- Drug-free workplace
- Considering qualified applicants with criminal history consistent with applicable laws
Sponsorship
- Capital One will consider sponsoring a new qualified applicant for employment authorization for this position
Full-Time
- Full-time type: Full-time
Technical Support / Accommodations
- Accommodation request contact: Capital One Recruiting at 1-800-304-9102 or [email protected]
- Technical support/questions: [email protected]