Machine Learning Engineer 4 (Manager, IC)
Backend Developer
Manager
Artificial Intelligence
Automation
Big Data
Cloud
Cloud Infrastructure
Cloud Native
Cloud Native Technologies
Cloud Platform
Cloud Platforms
Cloud Platforms Cloud Platforms
Cloud Technology
Data Analysis
Data Engineer
Data Integration
Data Pipeline
Data Platform
Data Processing
Data Science
DevOps
DevSecOps
Engineer
Engineering
Engineering Software
Facilities Management
Information Technology (IT)
Infrastructure
Infrastructure As Code
Kubernetes
Machine Learning
Machine Learning Engineer
Ml Ops
Platform Engineering
Programming
Programming Language
Programming Languages
Risk Management
Scaling Compute
Job Description
Design, build, deploy, and run machine learning solutions at scale in a production-focused role on an Agile team.
Responsibilities
- Design, build, and/or deliver ML models and components to solve real-world business problems in partnership with Product and Data Science
- Shape ML infrastructure decisions using expertise in modeling techniques and common issues, including model and feature selection, training, hyperparameter tuning, dimensionality, bias/variance, and validation
- Solve complex technical problems by writing and testing application code, developing and validating ML models, and automating tests and deployment
- Collaborate within a cross-functional Agile team to create and enhance software enabling big data and ML applications
- Retrain, maintain, and monitor models in production
- Build and/or leverage cloud-based architectures, technologies, and platforms to deliver optimized ML models at scale
- Construct optimized data pipelines that feed ML models
- Apply continuous integration and continuous deployment practices, including test automation and monitoring, to support successful releases of ML models and application code
- Ensure managed, secure code to reduce vulnerabilities, and support governance with an emphasis on risk and Responsible and Explainable AI
- Use programming languages including Python, Scala, or Java
Requirements
- Bachelor's Degree or higher in Computer Science, Machine Learning, or a related quantitative field: Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering
- 4+ years of experience programming with Python, Java, Golang, or C++
- 4+ years of Machine Learning experience with industry frameworks PyTorch or Tensorflow and libraries Pandas, NumPy, Scikit-learn
- 4+ years operating large scale distributed systems (e.g., Spark, Ray) to prepare AI or Machine Learning data
- 2+ years deploying and operating production ML solutions and production services in the cloud (AWS, GCP, Azure), using Kubernetes to manage large scale containerized ML systems
Preferred Qualifications
- Master's or Doctoral degree in Computer Science, Electrical Engineering, Mathematics, or related field
- 3+ years optimizing ML algorithms, configurations, and infrastructure
- 3+ years applying software development best practices (source control, testing, code reviews, CI/CD, etc.)
- 3+ years building resilient software with pre-production testing, advanced deployment techniques (one-box, blue/green, gradual dial-up), monitoring and alarms, and incident response plan preparation
- 3+ years experience with ML techniques (supervised, semi-supervised, unsupervised, reinforcement learning) and model types (regression, classification, clustering), including architectures (RNNs, CNNs, LSTMs, Transformers), plus training and evaluation (loss function, hyperparameters, regularization, underfitting/overfitting)
- 3+ years designing, implementing, and scaling production-ready data pipelines for training and evaluating ML models
- 1+ years experience as a technical lead developing ML solutions using industry best practices, patterns, and automation
- Authored or co-authored a paper on an ML technique, model, or proof of concept
Technologies
- Python, Scala, Java, Golang, C++, PyTorch, Tensorflow
- Pandas, NumPy, Scikit-learn
- Spark, Ray
- AWS, GCP, Azure
- Kubernetes
Benefits
- Eligible for performance-based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI); incentives may be discretionary or non-discretionary depending on the plan
- Comprehensive, competitive, and inclusive health, financial, and other benefits supporting total well-being
Compensation, Location, and Schedule
- Location: New York, NY (onsite)
- Salary: USD 215,200 - 245,600 per year (Machine Learning Engineer 4)
- Full-time
- This role is expected to accept applications for a minimum of 5 business days
- No agencies please
Work Authorization
- Capital One will not sponsor a new applicant for employment authorization or provide immigration-related support for this position