AI Engineer 4 (AI Foundations)
Job Description
The AI Engineer 4 role within Capital One’s Intelligent Foundations and Experiences (IFX) team focuses on building and deploying responsible, scalable foundation model systems. You will own end-to-end architecture for complex AI platforms while partnering across engineering, research, and product to deliver AI-powered capabilities for associates and customers.
Responsibilities
- Work with cross-functional partners including engineers, research scientists, technical program managers, and product managers to deliver AI-powered products that change how associates work and how customers interact with Capital One.
- Design, develop, test, deploy, and support AI software components spanning foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability.
- Apply a broad stack of open source and SaaS AI technologies, including AWS Ultraclusters, Huggingface, VectorDBs, and PyTorch.
- Develop and introduce state-of-the-art foundation model optimization techniques to improve scalability, cost, latency, and throughput in large-scale production AI systems.
- Contribute to the technical vision and long-term roadmap for foundational AI systems at Capital One.
- Own end-to-end architecture for complex AI systems, with focus on maintainability, observability, and ethical alignment.
- Define and maintain service-level objectives (SLOs) for AI reliability, including latency, uptime, and model performance drift.
- Collaborate with infrastructure engineering teams to optimize GPU and TPU utilization and accelerate model inference pipelines.
- Lead cross-functional technical reviews for new AI system deployments, ensuring security, data governance, and compliance standards are met.
- Mentor Principal and Senior Associates on scalable system design, performance tuning, and research-to-production translation.
Requirements
- Bachelor’s degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or a related field, plus at least 4 years of experience developing AI and ML algorithms or technologies, or a Master’s degree in a related field plus at least 2 years of such experience.
- At least 4 years of programming experience with Python, Go, Scala, CUDA, or Java.
Technologies
- AWS Ultraclusters, Huggingface, VectorDBs, PyTorch
- AWS, Google Cloud, Azure
- Python, Go, Scala, CUDA, Java, C++, C#, Golang
- GPU, TPU
Team Description
The Intelligent Foundations and Experiences (IFX) team is central to delivering Capital One’s AI vision. The team works with partners across the company to advance the state of the art in science and AI engineering, building and deploying proprietary solutions that support core business needs and deliver value to millions of customers. IFX builds AI models and platforms that help teams across Capital One enhance products with responsible, scalable AI.
What You’ll Bring
- Enjoy building high-quality systems while applying care to responsible outcomes.
- Maintain strong awareness of the latest research and apply novel techniques thoughtfully in production.
- Adapt quickly to clarify complex, undefined problems and communicate findings concisely.
- Share new ideas even when approaches have not yet been proven.
- Bring deep technical capability, including strong foundations in engineering and mathematics, with the ability to recognize and exploit optimization opportunities across hardware, software, and AI.
- Demonstrate resilience and the ability to forge paths toward business goals when the route is not predetermined.
Preferred Qualifications
- Experience leading development of AI systems with tradeoff decisions across cost, latency, throughput, and accuracy.
- 6 years of experience deploying scalable and responsible AI solutions on cloud platforms such as AWS, Google Cloud, Azure, or equivalent private cloud.
- Experience designing, developing, delivering, and supporting AI services.
- Experience developing AI and ML algorithms or technologies including LLM inference, similarity search and VectorDBs, guardrails, and memory using Python, C++, C#, Java, CUDA, or Golang.
- Experience applying state-of-the-art optimization techniques for training and inference to improve hardware utilization, latency, throughput, and cost.
- Experience building agentic AI systems and agentic workflows.
- Proficiency designing distributed systems for model training, evaluation, and online inference at petabyte scale.
- Experience defining AI model governance processes, including producibility, lineage tracking, and automated retention schedules.
- Proven ability to influence architectural decisions across multiple AI product lines or platforms.
Location and Salary
- Location: Cambridge, MA (onsite)
- Salary: USD 197,300 to 225,100 per year