AI Engineer 5 (MLX, Agentic AI, Gen AI platform Services)
Agentic Ai
Ai Agent
Ai Agent Platform
Ai Security
Artificial Intelligence
Artificial Intelligence Engineer
Data Analysis
Data Platform
Data Science
Engineer
Facilities Management
Generative AI
Information Technology (IT)
Machine Learning Engineering
Programming
Programming Language
Programming Languages
Project Management
Risk Management
Job Description
The Intelligent Foundations and Experiences (IFX) team at Capital One is seeking an AI Engineer 5 to help build and deploy responsible, scalable AI systems and platform components, with emphasis on foundation model workflows, agentic AI, and production-grade GenAI capabilities.
Role Focus
In this role, you will design and deliver AI-powered software components that support how associates work and how customers interact with Capital One. The work spans foundation model training and inference, agentic and multi-agent workflows, orchestration and orchestration pipelines, optimization, governance, and observability.
Key Responsibilities
- Collaborate with a cross-functional group including engineers, research scientists, technical program managers, and product managers to deliver AI-powered products.
- Design, develop, test, deploy, and support AI software components such as foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability.
- Apply a broad stack of Open Source and SaaS AI technologies, including AWS Ultraclusters, Huggingface, VectorDBs, and PyTorch.
- Develop state-of-the-art foundation model optimization techniques to improve performance across scalability, cost, latency, and throughput.
- Help shape the technical vision and long-term roadmap for foundational AI systems at Capital One.
- Design, implement, and optimize multi-model orchestration pipelines that integrate LLMs, vector search, and domain-specific models into unified systems.
- Establish and lead cost-performance governance reviews, including tracking GPU utilization, model throughput, and inference cost efficiency.
- Lead team design councils or design review boards to maintain technical consistency and alignment with AI engineering standards.
- Mentor Principal and Manager-level AI engineers to support cross-domain learning and increase organizational technical maturity.
Required Qualifications
- Bachelor’s degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or a related field with at least 6 years of experience developing AI and ML algorithms or technologies, or a Master’s degree with at least 4 years of experience developing AI and ML algorithms or technologies.
- At least 6 years of programming experience with Python, Go, Scala, CUDA, or Java.
Preferred Qualifications
- Experience leading development of AI systems with tradeoffs across cost, latency, throughput, and accuracy.
- 7 years of experience deploying scalable and responsible AI solutions on cloud platforms (AWS, Google Cloud, Azure, or equivalent private cloud).
- Experience designing, developing, delivering, and supporting complex AI systems.
- Experience developing AI and ML algorithms or technologies (including LLM inference, similarity search and VectorDBs, guardrails, and memory) using Python, C++, C#, Java, CUDA, or Golang.
- Experience applying state-of-the-art techniques to optimize training and inference software for improved hardware utilization, latency, throughput, and cost.
- Experience building agentic AI systems and agentic workflows.
- Passion for staying current with AI research and applying new techniques appropriately in production.
- Excellent communication and presentation skills for articulating complex AI concepts to peers.
- Experience architecting and integrating heterogeneous AI systems, including rule-based, retrieval-augmented, and generative components, into unified production pipelines.
- Experience defining and enforcing standards for ethical AI deployment, including explainability, fairness, and human-in-the-loop review processes.
- Ability to balance model performance and operational cost using dynamic inference strategies and model compression.
- Experience right-sizing models, instance counts, and hardware types based on requirements such as context length and token inputs and outputs.
Technologies and Tools
- AWS Ultraclusters
- Huggingface
- VectorDBs
- PyTorch
- AWS, Google Cloud, Azure
- Python, Go, Scala, CUDA, Java
- LLMs, vector search, similarity search
- Guardrails, retrieval-augmented, rule-based
- Foundation model, multi-agent workflows, agents
- Observability
Location
New York, NY (onsite)
Compensation
USD 250,800 - 286,200 per year (salary information by location: New York, NY).
Benefits
- Performance-based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI)
- Comprehensive, competitive, and inclusive set of health, financial, and other benefits supporting total well-being