Sr AI Engineer I
Job Description
American Express is seeking a Senior AI Engineer to design and build foundational Agentic AI Platform capabilities that help AI agents operate safely, reliably, and in line with enterprise governance across the organization. This role is hands-on across AI systems, distributed platforms, cloud infrastructure, developer tooling, and control and oversight mechanisms.
Role Summary
In this position, you will contribute to the architecture and implementation of Agent Runtime & Execution, orchestration, unified control plane services, governance controls, developer self-service tooling, CI/CD for AI agents, and evaluation and telemetry to support continuous improvement.
Key Responsibilities
- Contribute to the architecture and implementation of capabilities across the Agentic AI Platform, including Agent Runtime & Execution.
- Design and build scalable agent runtime infrastructure using Kubernetes and cloud-native technologies.
- Develop runtime abstractions that let agents execute across internally managed sandboxed environments and third-party AI platforms.
- Build capabilities for orchestration, tool execution, state and context management, memory, asynchronous workloads, event-driven execution, and multi-agent workflows.
- Build a unified control plane for managing agents across heterogeneous execution environments and AI providers.
- Develop APIs and services for agent registration, configuration, deployment, versioning, lifecycle management, policy enforcement, and runtime management.
- Create abstractions that reduce provider-specific complexity while preserving differentiated capabilities across AI ecosystems.
- Engineer platform-level controls for enterprise AI governance, including agent identity, authentication and authorization, tool permissions, policy enforcement, data boundaries, auditability, and lifecycle controls.
- Build mechanisms for governing models, prompts, tools, MCP servers, knowledge sources, agent-to-agent interactions, and external integrations.
- Partner with security, risk, privacy, and governance teams to translate enterprise requirements into scalable technical controls.
- Build registries and catalogs so developers and AI systems can discover reusable agents, tools, skills, prompts, models, knowledge sources, and other platform capabilities.
- Develop metadata, ownership, versioning, dependency, certification, and discovery mechanisms to support a healthy enterprise agent ecosystem.
- Build end-to-end telemetry for agent execution, including traces, events, model interactions, tool calls, latency, token consumption, cost, failures, policy decisions, and quality signals.
- Develop capabilities that make complex and multi-agent workflows explainable and debuggable.
- Enable platform and application teams to understand agent behavior across multiple models, runtimes, tools, and external systems.
- Develop evaluation frameworks for measuring agent quality, reliability, safety, and task performance.
- Build automated evaluation pipelines with offline evaluations, production signals, human feedback, regression testing, and experimentation.
- Help create continuous learning loops that translate production telemetry and feedback into improvements to agents, prompts, tools, models, and platform capabilities.
- Build SDKs, APIs, CLIs, templates, local development environments, and self-service workflows that simplify and accelerate agent development.
- Create paved roads that incorporate enterprise security, governance, observability, and operational standards while enabling developer speed.
- Build CI/CD capabilities designed for AI agents, including automated evaluation, policy validation, security checks, artifact/version management, deployment, promotion, rollback, and release controls.
- Design GitOps and infrastructure-as-code patterns for deploying and managing agent workloads.
- Help establish engineering standards for moving agents from experimentation to production safely and repeatedly.
Required Qualifications
- Strong software engineering experience building production systems using languages such as Python, Java, Go, or TypeScript.
- Strong understanding of Generative AI, LLMs, agent architectures, tool/function calling, retrieval, context management, and agent orchestration.
- Experience building production applications or platforms using major model providers or AI platforms.
- Experience with Kubernetes, containers, microservices, distributed systems, and cloud-native architecture.
- Experience designing production APIs and event-driven or asynchronous systems.
- Strong understanding of modern cloud infrastructure and infrastructure-as-code practices.
- Experience with CI/CD, automated testing, production observability, and software delivery practices.
- Strong understanding of security fundamentals including identity, authentication, authorization, secrets, and least-privilege access.
- Ability to navigate ambiguous technical problems and turn emerging technologies into reliable production systems.
- Strong communication skills and ability to collaborate across engineering, architecture, product, security, and governance organizations.
Technologies
- Python, Java, Go, TypeScript
- Generative AI, LLMs, agent architectures
- Tool/function calling, retrieval, context management, agent orchestration
- Kubernetes, cloud-native technologies, microservices, distributed systems
- Event-driven systems, asynchronous systems
- CI/CD, infrastructure-as-code, GitOps
- OpenTelemetry
- LangGraph, Semantic Kernel, Google ADK, OpenAI Agents SDK
- Model Context Protocol (MCP), A2A
- Vector databases, embeddings
- LLM-as-judge techniques, policy-as-code
Preferred Qualifications
- Experience with agent frameworks and orchestration technologies such as LangGraph, Semantic Kernel, Google ADK, OpenAI Agents SDK, or similar frameworks.
- Experience with Model Context Protocol (MCP) and emerging agent interoperability protocols like A2A.
- Experience with multi-agent systems and agent-to-agent communication patterns.
- Experience with Kubernetes operators, controllers, service meshes, or sophisticated Kubernetes platform engineering.
- Experience with agent/model gateways and intelligent model routing.
- Experience with vector databases, retrieval systems, embeddings, and enterprise knowledge architectures.
- Experience with LLM and agent evaluation frameworks, LLM-as-judge techniques, human feedback systems, experimentation, and quality measurement.
- Experience with AI observability, distributed tracing, OpenTelemetry, Langfuse, and production monitoring.
- Experience with Evals.
- Experience with policy engines and policy-as-code.
- Experience with AI security topics such as prompt injection defenses, tool security, data-loss prevention, and AI-specific threat modeling.
- Experience with platform engineering, internal developer platforms, developer portals, and enterprise service catalogs.
- Experience operating large-scale distributed systems in highly regulated environments.
What Success Looks Like
- Help developers across American Express move from an agent idea to a governed production deployment through a consistent, self-service platform.
- Enable a secure and reliable default path so developers can build once and operate agents across multiple AI ecosystems while receiving governance, identity, observability, evaluation, deployment, and operational capabilities by default.
- Establish engineering foundations for an enterprise ecosystem where AI agents can be built, discovered, trusted, governed, observed, evaluated, deployed, and continuously improved at scale.
Location and Compensation
- Location: Sunrise, FL (onsite)
- Salary Range: USD 123,000 - 215,250 per year
Work Authorization
Depending on factors such as business unit requirements, the nature of the position, cost and applicable laws, American Express may provide visa sponsorship for certain positions.