An onsite role in Deloitte’s healthcare AI practice in New York, NY, focusing on architecting and delivering end-to-end agentic AI systems for healthcare decisioning. The position combines LLM- and SLM-powered reasoning, orchestration, retrieval, memory, and control layers to support payers, providers, and life sciences within regulated workflows.
Responsibilities
- Design and implement agentic systems capable of multi-step reasoning, planning, tool use, and workflow execution against complex, regulated operational processes.
- Build stateful workflows using frameworks such as LangGraph and LangChain, including branching, retries, self-correction, human-in-the-loop checkpoints, and reusable orchestration patterns.
- Engineer for long-horizon reliability, ensuring multi-step task completion, recovery from compounding errors, planning under uncertainty, and robust tool use when individual steps fail.
- Develop the reasoning behind regulated decisions with policy- and criteria-grounded outputs, structured proposer/critic/judge-style review, and auditable rationales for high-stakes decisions across clinical review, prior authorization, claims integrity, and care management.
- Design end-to-end Retrieval-Augmented Generation (RAG) pipelines covering ingestion, chunking, embeddings, vector and hybrid retrieval, reranking, contextual compression, and grounding strategies.
- Engineer memory and context management, including conversational state, persistent memory, retrieval-aware context assembly, and token-efficient context selection.
- Apply modern context-delivery patterns so agents access the right information at the right time.
- Implement observability and tracing for prompts, tool calls, retrieval quality, agent traces, failures, drift, latency, and production behavior.
- Incorporate guardrails, safety controls, and failure-handling to reduce hallucinations and unsafe actions.
- Evaluate agents at trajectory and task levels, including multi-step task success, failure modes, regression analysis, and sandboxed testing, alongside retrieval and generation quality metrics, automated checks, and human review.
- Engineer healthcare-grade safety, deploying eval gates, human oversight and escalation models, and auditability and traceability for regulated decisions with PHI/HIPAA-aware data handling.
- Build integrations with internal and external tools, APIs, enterprise systems, databases, and model providers to ensure agents operate safely within real business workflows.
- Deliver production-quality code with solid testing, CI/CD, logging, versioning, and documentation; make architecture decisions balancing quality, safety, latency, cost, and model risk.
- Collaborate with modeling and post-training engineers to improve model behavior for tool use, grounding, and long-horizon reasoning through evaluation-driven feedback and, where helpful, fine-tuned or reasoning-optimized models.
- Translate ambiguous, high-complexity operational processes into robust system logic and reusable AI patterns; stay current with advances in agentic systems and translate research into practical engineering decisions.
Requirements
- Bachelor’s degree in Computer Science, Engineering, Data Science, Computational Linguistics, or a related field.
- Demonstrated depth building and shipping production agentic systems; this is your primary craft, with a track record of shipped systems, research, model releases, or open source work over years.
- Strong hands-on experience building production agent systems with modern orchestration — LangGraph/LangChain or equivalent, including custom orchestration.
- Experience designing and optimizing end-to-end RAG systems: indexing, retrieval, reranking, grounding, and evaluation.
- Solid understanding of memory and context management, including context windows, retrieval-driven context assembly, persistent memory, and high-signal context selection.
- Deep, practical understanding of LLM behavior, including strengths, limitations, hallucination risks, reasoning constraints, and latency/cost trade-offs, plus evaluation methods to measure them.
- Experience evaluating and debugging agent behavior through task-success and trajectory analysis, not solely output quality.
- Strong Python engineering skills and modern software practices: testing, CI/CD, version control, and API integration; experience implementing observability, tracing, and debugging for production LLM-based systems.
- Hands-on experience with at least one frontier model platform (e.g., Anthropic, Google, OpenAI) and/or open-weight/self-hosted models (e.g., Llama via vLLM), including production tool use and agent capabilities.
- Ability to travel 0-50% on average, depending on client work and industry engagements.
- Limited immigration sponsorship may be available.
Technologies
- LangGraph, LangChain, Python, Llama via vLLM, Anthropic, Google, OpenAI, Pinecone, Weaviate, Milvus, FHIR
The team
Deloitte brings together AI researchers, modeling and platform engineers, architects, clinical and domain specialists, and product leaders to build, deploy, and operate vertical AI systems across software, data, models, and cloud infrastructure. The healthcare-focused work spans payers, providers, and life sciences, tackling genuinely hard reasoning problems, nuanced operational workflows, and a high bar for quality and safety.
Compensation
Base salary is benchmarked toward leading technology firms, with a substantial performance-based incentive opportunity designed to align with the value you help create. The estimated base salary range is $110,700 to $372,900 per year (not adjusted for geographic differential); actual base pay depends on your skills, experience, and level.