Location: onsite in Denver, Colorado. This role offers a path to influence healthcare through agentic AI, with a competitive compensation package and the backing of a large, well-supported platform. You will join Deloitteβs multidisciplinary team to design, build, and operate end-to-end agentic AI systems for healthcare decisioning, deployed across payers, providers, and life sciences, with a strong emphasis on reliability, safety, and auditable outcomes in high-stakes settings.
Responsibilities
- Design and implement agentic systems capable of multi-step reasoning, planning, tool use, and workflow execution within complex, regulated operational processes.
- Develop stateful workflows using LangGraph and LangChain, including branching, retries, self-correction, human-in-the-loop checkpoints, and reusable orchestration patterns.
- Engineer for long-horizon reliability with multi-step task completion, recovery from compounding errors, planning under uncertainty, and robust tool use when individual steps fail.
- Craft the reasoning behind regulated decisions with policy- and criteria-grounded outputs, structured proposer/critic/judge-style review, and auditable rationales for decisions across clinical review, prior authorization, claims integrity, and care management.
- Develop end-to-end Retrieval-Augmented Generation pipelines: ingestion, chunking, embeddings, vector and hybrid retrieval, reranking, contextual compression, and grounding strategies.
- Engineer memory and context management including conversational state, persistent memory, retrieval-aware context assembly, and token-efficient context selection.
- Apply modern context-delivery patterns so agents access the right information at the right time.
- Implement observability and tracing for prompts, tool calls, retrieval quality, agent traces, failures, drift, latency, and production behavior.
- Apply guardrails and safety controls to reduce hallucinations and unsafe actions.
- Evaluate agents at the trajectory and task level with multi-step task success analysis, failure-mode and regression analysis, sandboxed testing, and metrics for retrieval and generation quality, automated checks, and human review.
- Ensure healthcare-grade safety with deployment eval gates, human oversight and escalation models, auditability and traceability for regulated decisions, and PHI/HIPAA-conscious data handling.
- Build integrations with internal and external tools, APIs, enterprise systems, databases, and model providers so agents operate safely within real business workflows.
- Deliver production-quality code with rigorous testing, CI/CD, logging, versioning, and documentation; make architecture decisions that balance quality, safety, latency, cost, and model risk.
- Collaborate with modeling and post-training engineers to improve model behavior for tool use, grounding, and long-horizon reasoning through evaluation-driven feedback and, where helpful, fine-tuning or reasoning-optimized models.
- Translate ambiguous, high-complexity operational processes into robust system logic and reusable AI patterns; stay current with advances in agentic systems and translate research into practical engineering decisions.
Requirements
- Bachelor's degree in Computer Science, Engineering, Data Science, Computational Linguistics, or a related field.
- Proven track record shipping production agentic systems as a primary craft, with strong software and ML fundamentals and substantial recent hands-on agentic work.
- Hands-on experience building production agent systems with modern orchestration such as LangGraph/LangChain or equivalents, including custom orchestration.
- Experience designing and optimizing end-to-end RAG systems: indexing, retrieval, reranking, grounding, and evaluation.
- Strong memory and context management understanding, including context windows, retrieval-driven context assembly, persistent memory, and high-signal context selection.
- Deep practical understanding of LLM behavior, including strengths, limitations, hallucination risks, reasoning constraints, and latency/cost trade-offs; familiarity with evaluation methods.
- Experience evaluating and debugging agent behavior beyond output quality, focusing on task success and trajectory analysis.
- Strong Python engineering skills and modern software practices: testing, CI/CD, version control, API integration; experience implementing observability, tracing, and debugging for production LLM-based systems.
- Hands-on experience with at least one frontier model platform (e.g., Anthropic, Google, OpenAI) and/or open-weight/self-hosted models (e.g., Llama via vLLM), including production tool use and agent capabilities.
- Ability to travel 0-50 percent depending on client work.
- Limited immigration sponsorship may be available.
Technologies
- LangGraph
- LangChain
- vLLM
- Llama via vLLM
- Pinecone
- Weaviate
- Milvus
- Anthropic
- Google
- OpenAI
- Python
- FHIR
The Team
Deloitte brings together AI researchers, modeling and platform engineers, architects, clinical and domain specialists, and product leaders to build, deploy, and operate verticalized AI systems across software, data, models, and cloud infrastructure, engineered for one of the most complex operating environments in the world. The work spans the healthcare industry β payers, providers, and life sciences β and involves genuinely hard reasoning problems, nuanced operational workflows, and a high bar.
Compensation
Base salary is benchmarked to leading technology firms, with a substantial performance-based incentive designed to grow with the value you help create. The estimated base salary range is $110,700 to $372,900 (not geographic differential). Actual base pay depends on your skills, experience, and level.