DataJobs.io
← Back to all jobs

Job Description

The Intelligent Foundations and Experiences (IFX) team at Capital One is seeking an AI Engineer 5 to help build and deploy responsible, scalable AI systems and platform components, with emphasis on foundation model workflows, agentic AI, and production-grade GenAI capabilities.

Role Focus

In this role, you will design and deliver AI-powered software components that support how associates work and how customers interact with Capital One. The work spans foundation model training and inference, agentic and multi-agent workflows, orchestration and orchestration pipelines, optimization, governance, and observability.

Key Responsibilities

  • Collaborate with a cross-functional group including engineers, research scientists, technical program managers, and product managers to deliver AI-powered products.
  • Design, develop, test, deploy, and support AI software components such as foundation model training, large language model inference, agents and multi-agent workflows, similarity search, guardrails, model evaluation, experimentation, governance, and observability.
  • Apply a broad stack of Open Source and SaaS AI technologies, including AWS Ultraclusters, Huggingface, VectorDBs, and PyTorch.
  • Develop state-of-the-art foundation model optimization techniques to improve performance across scalability, cost, latency, and throughput.
  • Help shape the technical vision and long-term roadmap for foundational AI systems at Capital One.
  • Design, implement, and optimize multi-model orchestration pipelines that integrate LLMs, vector search, and domain-specific models into unified systems.
  • Establish and lead cost-performance governance reviews, including tracking GPU utilization, model throughput, and inference cost efficiency.
  • Lead team design councils or design review boards to maintain technical consistency and alignment with AI engineering standards.
  • Mentor Principal and Manager-level AI engineers to support cross-domain learning and increase organizational technical maturity.

Required Qualifications

  • Bachelor’s degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or a related field with at least 6 years of experience developing AI and ML algorithms or technologies, or a Master’s degree with at least 4 years of experience developing AI and ML algorithms or technologies.
  • At least 6 years of programming experience with Python, Go, Scala, CUDA, or Java.

Preferred Qualifications

  • Experience leading development of AI systems with tradeoffs across cost, latency, throughput, and accuracy.
  • 7 years of experience deploying scalable and responsible AI solutions on cloud platforms (AWS, Google Cloud, Azure, or equivalent private cloud).
  • Experience designing, developing, delivering, and supporting complex AI systems.
  • Experience developing AI and ML algorithms or technologies (including LLM inference, similarity search and VectorDBs, guardrails, and memory) using Python, C++, C#, Java, CUDA, or Golang.
  • Experience applying state-of-the-art techniques to optimize training and inference software for improved hardware utilization, latency, throughput, and cost.
  • Experience building agentic AI systems and agentic workflows.
  • Passion for staying current with AI research and applying new techniques appropriately in production.
  • Excellent communication and presentation skills for articulating complex AI concepts to peers.
  • Experience architecting and integrating heterogeneous AI systems, including rule-based, retrieval-augmented, and generative components, into unified production pipelines.
  • Experience defining and enforcing standards for ethical AI deployment, including explainability, fairness, and human-in-the-loop review processes.
  • Ability to balance model performance and operational cost using dynamic inference strategies and model compression.
  • Experience right-sizing models, instance counts, and hardware types based on requirements such as context length and token inputs and outputs.

Technologies and Tools

  • AWS Ultraclusters
  • Huggingface
  • VectorDBs
  • PyTorch
  • AWS, Google Cloud, Azure
  • Python, Go, Scala, CUDA, Java
  • LLMs, vector search, similarity search
  • Guardrails, retrieval-augmented, rule-based
  • Foundation model, multi-agent workflows, agents
  • Observability

Location

New York, NY (onsite)

Compensation

USD 250,800 - 286,200 per year (salary information by location: New York, NY).

Benefits

  • Performance-based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI)
  • Comprehensive, competitive, and inclusive set of health, financial, and other benefits supporting total well-being

Similar Jobs