Machine Learning Engineer 5 (IC)
Backend Developer
Ai Ml
Artificial Intelligence
Big Data
Data Analysis
Data Engineer
Data Pipeline
Data Platform
Data Processing
Data Science
Data Science Ml
DevOps
Engineer
Information Technology (IT)
Machine Learning Engineer
Machine Learning Engineering
Machine Learning Operations
Machine Learning Platform
Programming
Programming Language
Programming Languages
Job Description
The Machine Learning Engineer 5 role focuses on building and scaling production-grade AI and ML systems at scale, with an emphasis on responsible AI practices. The work involves designing models and platforms, creating data pipelines, and supporting deployment and ongoing model operations in collaboration with Product and Data Science teams.
What You’ll Do
- Design, build, and deliver machine learning models and components that address real-world business problems with Product and Data Science partners.
- Build and scale large multi-tenant platforms for running large-footprint ML training and/or serving.
- Guide ML infrastructure decisions using expertise across modeling and engineering topics, including model and data selection, feature selection, training, hyperparameter tuning, dimensionality, bias/variance, and validation.
- Solve complex problems by writing and testing application code, developing and validating ML models, and automating tests and deployment.
- Collaborate on a cross-functional Agile team to create and enhance software for state-of-the-art big data and ML applications.
- Retrain, maintain, and monitor production models.
- Use cloud-based architectures, technologies, and/or platforms to deliver optimized ML models at scale.
- Construct optimized data pipelines to support ML model inputs for training and evaluation.
- Apply continuous integration and continuous deployment best practices, including test automation and monitoring, to support successful deployment of ML models and application code.
- Ensure code is well-managed to reduce vulnerabilities, support risk-governed model practices, and follow best practices in Responsible and Explainable AI.
- Use programming languages including Python, Scala, or Java.
Core Requirements
- Bachelor’s Degree or higher in Computer Science, Machine Learning, or a related quantitative field (Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering).
- At least 6 years of programming experience with Python, Java, Golang, or C++.
- At least 6 years of machine learning experience using industry-standard frameworks and libraries, including PyTorch or TensorFlow, and Pandas, NumPy, Scikit-learn.
- At least 6 years operating large-scale distributed systems (including Spark and Ray) to prepare AI/ML data.
- At least 4 years deploying and operating machine learning solutions in production within cloud environments (AWS, GCP, Azure) and using Kubernetes to manage large-scale containerized ML software systems.
Technologies
Python, Scala, Java, Golang, C++, PyTorch, Tensorflow, Pandas, NumPy, Scikit-learn, Spark, Ray, AWS, GCP, Azure, Kubernetes, CI/CD
Preferred Qualifications
- Master’s or Doctoral Degree in Computer Science, Electrical Engineering, Mathematics, or a related field.
- 5+ years of experience optimizing ML algorithms, configurations, and infrastructure.
- 5+ years of experience following software development best practices, including source control, testing, code reviews, and CI/CD.
- 5+ years of experience building resilient software solutions with pre-production testing, advanced deployment techniques (one-box, blue/green, gradual dial-up), monitoring, alarms, and incident response plan preparation.
- 5+ years of experience working with ML techniques (Supervised, semi-supervised, and unsupervised, reinforcement learning, etc.), model types (Regression, Classification, Clustering, etc.), model architectures (RNNs, CNNs, LSTMs, Transformers), training concepts (loss function, hyperparameters, regularization), and diagnosing evaluation issues such as underfitting and overfitting.
- 5+ years of experience designing, implementing, and scaling production-ready data pipelines for training and evaluating ML models.
- ML industry impact through conference presentations, papers, blog posts, open source contributions, or patents.
- Ability to clearly communicate complex technical and machine learning concepts to a variety of audiences.
Compensation
Salary Range (New York, NY): USD 250,800 - 286,200 per year.
Benefits
- Capital One provides a comprehensive, competitive, and inclusive set of health, financial, and other benefits supporting total well-being.
- Performance-based incentive compensation, which may include cash bonus(es) and/or long-term incentives (LTI).
Additional Information
- Capital One will consider sponsoring a new qualified applicant for employment authorization for this position.
- This role is expected to accept applications for a minimum of 5 business days.
- No agencies please.
- Capital One is an equal opportunity employer (EOE, including disability/vet) committed to non-discrimination in compliance with applicable federal, state, and local laws.
- Capital One promotes a drug-free workplace.
- If you require an accommodation for employment information or to apply, contact Capital One Recruiting at 1-800-304-9102 or via email at [email protected].
- For technical support or questions about Capital One’s recruiting process, email [email protected].
- Capital One does not provide, endorse, or guarantee and is not liable for third-party products, services, educational tools, or other information available through this site.
- Capital One Financial is made up of several different entities; positions posted in Canada are for Capital One Canada, in the United Kingdom for Capital One Europe, and in the Philippines for Capital One Philippines Service Corp. (COPSSC).
Location and Work Model
New York, NY (onsite)