Senior AI Engineer
AI
Ai Ml
API
APIs
Artificial Intelligence
Automation
Cloud
Cloud Infrastructure
Cloud Native
Cloud Platform
Cloud Platforms
Cloud Technology
Collaboration
DevOps
DevSecOps
Embeddings
Engineer
Engineering
Generative AI
Git
Google Cloud
Google Cloud Platform
Google Kubernetes Engine
Infrastructure
Infrastructure As Code
Kubernetes
Machine Learning
Platform Engineering
Software Development
Systems
Vertex Ai
Job Description
Senior AI Engineer role in Denver, CO (remote) at Home Depot / THD, focused on designing, building, and optimizing production-grade agentic AI systems to drive measurable business outcomes.
Responsibilities
- 70% Delivery and Execution: Collaborate with UX, engineering, and product management to create secure, reliable, and scalable ML solutions; document and review to meet quality and change control standards; ensure user stories are developer-ready, clear, and testable; write custom code or scripts to automate infrastructure, monitoring services, and test cases.
- 10% Learning: Engage in learning activities around modern software design, machine learning, and development best practices; proactively review articles, tutorials, and videos to stay current with technologies used in other organizations.
- 20% Support and Enablement: Field questions from product and support teams; monitor tools and foster cross-team collaboration; provide production application support; monitor production SLOs; review performance and capacity across code, infrastructure, data, messaging, and prediction quality.
Requirements
- Experience: 6+ years in AI, Machine Learning Engineering, or Software Engineering with strong Python development skills and modern software practices.
- AI Delivery: Proven track record building and deploying production-grade AI solutions using LLMs, SLMs, RAG frameworks, copilots, agents, and multi-agent systems.
- AI Foundations: Deep understanding of transformers, embeddings, deep learning, prompt engineering, agentic reasoning patterns, and vector databases.
- Orchestration & Integration: Experience developing orchestration layers (task execution, routing, planning, workflows) and integrating AI solutions with enterprise platforms, APIs, and business systems.
- Infrastructure & MLOps: Expertise in cloud-native architectures, Docker, Kubernetes/GKE, Terraform, and AI pipeline design; hands-on MLOps/LLMOps practices across CI/CD, automated testing, model versioning and registries, governance, security, and lifecycle management.
- AIOps & Deployment Reliability: Background building automated CI/CD for AI/agent systems; implement canary, blue-green, and shadow deployments with automated rollback; establish end-to-end observability (logging, metrics, tracing, alerting) across models, agents, and orchestration for reliability and cost efficiency at scale.
- Optimization & Debugging: Ability to optimize complex AI systems for performance, reliability, scalability, latency, cost, and token efficiency; diagnose and resolve operational failure modes.
- Execution & Collaboration: Excellent cross-functional communication and collaboration skills to bring AI solutions from concept to production in large enterprise environments.
Technologies
- Python, LLMs, SLMs, RAG frameworks
- Vertex AI, Gemini, Google ADK, LangGraph, CrewAI, AutoGen
- Node.js, React, REST
- Linux, Git, Docker, Kubernetes, GKE
- Terraform, Transformers, Embeddings, Vector databases
Benefits
- Health care benefits
- 401K
- ESPP
- Paid time off
- Success sharing bonus
Travel Requirements
- Typically requires overnight travel 5% to 20% of the time
Physical Requirements
- Most of the time spent sitting in a comfortable position; frequent movement possible; occasional lifting of light objects
Working Conditions
- Located in a comfortable indoor area; infrequent and unobjectionable exposure to any unpleasant conditions
Minimum Qualifications
- Must be eighteen years of age or older
- Must be legally permitted to work in the United States
Preferred Qualifications
- Tools & Frameworks: Hands-on with Vertex AI, Gemini, Google ADK, LangGraph, CrewAI, AutoGen, or similar orchestration tools
- AI Infrastructure & Platform Tooling: IaC (Terraform), Kubernetes/GKE, GPU provisioning and autoscaling, model registries, feature stores, vector databases at production scale
- Full-stack Skills: Node.js/React/REST, API design, performance optimization, Linux, Git, modern deployment toolchains
- Industry Context: Retail, supply chain, manufacturing, eCommerce, logistics, or finance with mature Applied ML
- Guardrails & Reliability: Responsible AI, evaluation frameworks, reliability engineering, AI governance guardrails
- Leadership & Innovation: Track record of driving innovation, delivering business impact, mentoring teams, establishing AI engineering standards
- Education: Master’s or bachelor’s in computer science, AI, ML, or related field
Minimum Education
- High school diploma or GED
Preferred Education
- No additional education
Minimum Years of Work Experience
- 2
Preferred Years of Work Experience
- No additional years of experience
Minimum Leadership Experience
- None
Preferred Leadership Experience
- None
Certifications
- None
Competencies
- Global Perspective
- Manages Ambiguity
- Nimble Learning
- Self-Development
- Collaborates
- Cultivates Innovation
- Situational Adaptability
- Communicates Effectively
- Drives Results
- Interpersonal Savvy