Lead AI Engineer (FM Hosting, LLM Inference)
Job Description
Capital One offers an on-site Lead AI Engineer role in New York, NY with a competitive annual salary of USD 215,200 to 245,600. This position is part of the IFX team, focusing on foundation model training, large language model inference, and the design, development, and deployment of AI powered products that impact how our associates work and how customers interact with Capital One.
The role provides the opportunity to influence technical direction while collaborating with a cross-functional team to deliver measurable outcomes in a fast-paced, data-driven environment.
Responsibilities
- Collaborate with engineers, research scientists, technical program managers, and product managers to deliver AI driven products that transform internal workflows and customer experiences.
- Design, build, test, deploy, and support AI software components including foundation model training, LLM inference, similarity search, guardrails, model evaluation, experimentation, governance, and observability.
- Utilize a broad stack of Open Source and SaaS AI technologies such as AWS Ultraclusters, Huggingface, VectorDBs, Nemo Guardrails, PyTorch, and more.
- Develop and apply advanced LLM optimization techniques to improve scalability, cost efficiency, latency, and throughput for large scale production AI systems.
- Contribute to the technical vision and the long term roadmap of foundational AI systems at Capital One.
Requirements
- Bachelor's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 4 years of experience developing AI and ML algorithms or technologies.
- Master's degree in Computer Science, AI, Electrical Engineering, Computer Engineering, or related fields plus at least 2 years of experience developing AI and ML algorithms or technologies.
- At least 4 years of experience programming with Python, Go, Scala, or Java.
- 6 years of experience deploying scalable and responsible AI solutions on cloud platforms (e.g., AWS, Google Cloud, Azure, or an equivalent private cloud).
- Experience designing, developing, delivering, and supporting AI services.
- Experience developing AI and ML algorithms or technologies (for example LLM Inference, Similarity Search and VectorDBs, Guardrails, Memory) using Python, C++, C#, Java, or Golang.
- Experience developing and applying state-of-the-art techniques for optimizing training and inference software to improve hardware utilization, latency, throughput, and cost.
- A demonstrated passion for staying current with AI research and AI systems, and applying novel techniques in production.
Technologies
- Python
- Go
- Scala
- Java
- AWS Ultraclusters
- Huggingface
- VectorDBs
- Nemo Guardrails
- PyTorch
Benefits
- Health benefits
- Financial benefits
- Incentives including performance based incentive compensation (cash bonuses and/or long term incentives)