AI Engineer 5
Job Description
Capital One is hiring an AI Engineer to help build responsible, scalable foundation AI and LLM capabilities for banking.
Responsibilities
- Collaborate with a cross-functional group of engineers, research scientists, technical program managers, and product managers to deliver AI-powered products for associates and customers.
- Design, develop, test, deploy, and support AI software components, including foundation model training, LLM inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability.
- Apply open source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, and more.
- Develop state-of-the-art foundation model optimization techniques to improve scalability, cost, latency, and throughput in production AI systems.
- Help shape the technical vision and long-term roadmap for foundational AI systems at Capital One.
- Design, implement, and optimize multi-model orchestration pipelines that integrate LLMs, vector search, and domain-specific models into unified systems.
- Establish and lead cost-performance governance reviews across AI systems, including GPU utilization, model throughput, and inference cost efficiency.
- Lead team design councils or design review boards to ensure technical consistency and compliance with AI engineering standards.
- Mentor Principal and Manager-level AI engineers to support cross-domain learning and raise organizational technical maturity.
Requirements
- Bachelor’s degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or a related field plus at least 6 years of experience developing AI/ML algorithms or technologies, or a Master’s degree plus at least 4 years of experience developing AI/ML algorithms or technologies.
- At least 6 years of programming experience with Python, Go, Scala, CUDA, or Java.
Technologies
- Python, Go, Scala, CUDA, Java
- AWS Ultraclusters, Huggingface, VectorDBs, PyTorch, AWS, Google Cloud, Azure
- C++, C#, Golang
- LLM Inference, Similarity Search, Guardrails, Memory, retrieval-augmented
- Foundation model training, agents and multi-agent workflows, multi-model orchestration pipelines, vector search
Team
- The Intelligent Foundations and Experiences (IFX) team is central to bringing Capital One’s AI vision to life.
- Partners with teams across the company to advance science and AI engineering, and builds and deploys proprietary solutions used across business and delivered to millions of customers.
- Enables teams to enhance products with responsible and scalable AI for high-leverage impact.
Ideal Candidate
- Enjoys building systems, takes pride in quality, and focuses on doing the right thing.
- Stays current with the latest research and can apply new techniques carefully in production.
- Works through big, undefined problems by finding root causes and communicating findings clearly.
- Has strong engineering and mathematics fundamentals, with expertise across hardware, software, and AI to uncover optimization opportunities.
- Can pursue business goals even when the path is not fully known.
Preferred Qualifications
- Experience leading development of AI systems with tradeoffs across cost, latency, throughput, and accuracy.
- 7+ years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g., AWS, Google Cloud, Azure, or equivalent private cloud).
- Experience designing, developing, delivering, and supporting complex AI systems.
- Experience developing AI and ML algorithms or technologies (e.g., LLM inference, similarity search and VectorDBs, guardrails, memory) using Python, C++, C#, Java, CUDA, or Golang.
- Experience optimizing training and inference software to improve hardware utilization, latency, throughput, and cost.
- Experience building agentic AI systems and agentic workflows.
- Excellent communication and presentation skills for articulating complex AI concepts to peers.
- Experience architecting and integrating heterogeneous AI systems including rule-based, retrieval-augmented, and generative components into unified production pipelines.
- Experience defining and enforcing ethical AI deployment standards, including explainability, fairness, and human-in-the-loop review processes.
- Ability to balance model performance and operational cost through dynamic inference strategies and model compression.
- Experience right-sizing models, instance counts, and hardware types based on requirements such as context length and token inputs/outputs.
Location: McLean, VA (onsite)
Salary: USD 229,900 - 262,400 per year
Experience: 6+ years (minimum)