Machine Learning Engineer 4
Job Description
Transform real business outcomes with large-scale AI. Capital One is seeking a Machine Learning Engineer (ML Engineer 4) to help drive major AI transformations and scale production ML models. You will design, build, deploy, monitor, and govern machine learning solutions and the data and software pipelines that support them, using modern cloud and distributed systems. This is an onsite role in Plano, TX.
Compensation: $179,400 - $204,700 per year (Plano, TX).
What you’ll do
In this role, you’ll perform a wide range of ML engineering work, including:
- Design, build, and/or deliver ML models and components that solve real-world business problems in collaboration with Product and Data Science teams
- Shape ML infrastructure decisions using your understanding of modeling techniques and tradeoffs, including model choice, data and feature selection, training, hyperparameter tuning, dimensionality, bias/variance, and validation
- Address complex problems by writing and testing application code, developing and validating ML models, and automating tests and deployment
- Collaborate within a cross-functional Agile team to create and improve software supporting big data and ML applications
- Retrain, maintain, and monitor models in production
- Use and/or build cloud-based architectures, technologies, and platforms to deliver optimized ML models at scale
- Construct optimized data pipelines that feed ML models
- Apply continuous integration and continuous deployment best practices, including test automation and monitoring, to support successful deployments
- Ensure code is well-managed to reduce vulnerabilities, that models are well-governed from a risk perspective, and that ML follows best practices in Responsible and Explainable AI
- Use programming languages such as Python, Scala, or Java
Skills and experience needed
- Bachelor’s Degree or higher in Computer Science, Machine Learning, or a related quantitative field (including Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering)
- At least 4 years of experience programming with Python, Java, Golang, or C++
- At least 4 years of Machine Learning experience with industry frameworks and libraries, including PyTorch or TensorFlow and Pandas, NumPy, Scikit-learn
- At least 4 years operating large-scale distributed systems (such as Spark and Ray) to prepare AI or ML data
- At least 2 years deploying and operating ML solutions in production services in the cloud (AWS, GCP, or Azure) and using Kubernetes to manage large-scale containerized ML systems
Tools you’ll work with
Python, Scala, Java, Golang, C++, PyTorch, Tensorflow, Pandas, NumPy, Scikit-learn, Spark, Ray, AWS, GCP, Azure, Kubernetes, CI/CD
Incentives and benefits
- Eligible to earn performance-based incentive compensation, which may include cash bonus(es) and/or long-term incentives (LTI)
- Comprehensive, competitive, and inclusive health, financial, and other benefits supporting total well-being (eligibility varies by full or part-time status, exempt or non-exempt status, and management level)
Preferred qualifications
- Master’s or Doctoral Degree in Computer Science, Electrical Engineering, Mathematics, or a related field
- 3+ years of experience optimizing ML algorithms, configurations, and infrastructure
- 3+ years following software development best practices including source control, testing, code reviews, CI/CD, and related practices
- 3+ years building resilient software solutions with pre-production testing, advanced deployment techniques (one-box, blue/green, gradual dial-up), monitoring, alarms, and preparing incident response plans
- 3+ years working with ML techniques (supervised, semi-supervised, unsupervised, reinforcement learning, etc.), model types (regression, classification, clustering, etc.), model architectures (RNNs, CNNs, LSTMs, Transformers), training concepts (loss function, hyperparameters, regularization), and evaluating accuracy and diagnosing issues (underfitting, overfitting)
- 3+ years designing, implementing, and scaling production-ready data pipelines for training and evaluating ML models
- 1+ years of experience as a technical lead developing ML solutions using industry best practices, patterns, and automation
- Authored or co-authored a paper on a ML technique, model, or proof of concept
Additional notes: Capital One will not sponsor a new applicant for employment authorization or offer immigration-related support for this position. Applications are expected to be accepted for a minimum of 5 business days. No agencies please. Capital One is an equal opportunity employer (EOE) committed to non-discrimination. Capital One promotes a drug-free workplace.
Accommodation and recruiting contact: If you require an accommodation, contact Capital One Recruiting at 1-800-304-9102 or via email at [email protected]. For technical support or questions about the recruiting process, email [email protected].