Agentic AI Engineer — Healthcare AI
Job Description
Deloitte in Raleigh, NC offers an onsite role focused on healthcare AI within an AI-first initiative. You will build end-to-end LLM- and SLM-powered agentic systems for healthcare decisioning and deploy them into live clinical and operational settings. This position blends advanced engineering with practical, regulated workflows and a collaborative, impact-driven culture. The role carries a competitive salary and the opportunity to contribute to cutting edge AI in healthcare.
Location and compensation
Onsite in Raleigh, North Carolina. Salary range of USD 134,500 to 265,100 per year. A bachelor’s degree is required.
Responsibilities
- Design and implement agentic systems capable of multi-step reasoning, planning, tool use, and workflow execution against complex, regulated operational processes.
- Build stateful workflows using frameworks such as LangGraph and LangChain, including branching, retries, self-correction, human-in-the-loop checkpoints, and reusable orchestration patterns.
- Engineer for long-horizon reliability, enabling multi-step task completion, recovery from compounding errors, planning under uncertainty, and robust tool use when steps fail.
- Build the reasoning behind regulated decisions with policy- and criteria-grounded outputs, structured proposer/critic/judge style review, and auditable rationales for high-stakes decisions across clinical review, prior authorization, claims integrity, and care management.
- Develop end-to-end Retrieval-Augmented Generation pipelines, covering ingestion, chunking, embeddings, vector and hybrid retrieval, reranking, contextual compression, and grounding strategies.
- Engineer memory and context management, including conversational state, persistent memory, retrieval-aware context assembly, and token-efficient context selection.
- Apply modern context-delivery patterns to ensure agents access the right information at the right time.
- Implement observability and tracing for prompts, tool calls, retrieval quality, agent traces, failures, drift, latency, and production behavior.
- Apply guardrails, safety controls, and failure-handling to reduce hallucinations and unsafe actions.
- Evaluate agents at the trajectory and task level with multi-step task success metrics, failure-mode analysis, sandboxed testing, and integrated quality checks alongside human review.
- Engineer healthcare-grade safety with deployment eval gates, human oversight and escalation models, auditability and traceability for regulated decisions, and PHI/HIPAA-aware data handling.
- Build integrations with internal and external tools, APIs, enterprise systems, databases, and model providers to operate safely within real business workflows.
- Deliver production-quality code with robust testing, CI/CD, logging, versioning, and documentation; balance quality, safety, latency, cost, and model risk in architectural decisions.
- Partner with modeling and post-training engineers to improve model behavior for tool use, grounding, and long-horizon reasoning through evaluation-driven feedback and fine-tuning where appropriate.
- Translate ambiguous, high-complexity operational processes into robust system logic and reusable AI patterns; stay current with agentic system advances and translate research into practical engineering decisions.
Requirements
- Bachelor's degree in Computer Science, Engineering, Data Science, Computational Linguistics, or a related field.
- Demonstrated depth building and shipping production agentic systems as a primary craft, with a track record of shipped systems, research, model releases, and open source work over years.
- Strong hands-on experience building production agent systems with modern orchestration, including LangGraph/LangChain or equivalent with custom orchestration.
- Experience designing and optimizing end-to-end RAG systems: indexing, retrieval, reranking, grounding, and evaluation.
- Solid understanding of memory and context management, including context windows, retrieval-driven context assembly, persistent memory, and high-signal context selection.
- Deep practical knowledge of LLM behavior, limitations, hallucination risks, reasoning constraints, and latency/cost trade-offs, plus evaluation methodologies.
- Experience evaluating and debugging agent behavior through task-success and trajectory analysis, not just output quality.
- Strong Python engineering skills and modern software practices: testing, CI/CD, version control, API integration; experience implementing observability, tracing, and debugging for production LLM-based systems.
- Hands-on experience with at least one frontier model platform (e.g., Anthropic, Google, OpenAI) and/or open-weight/self-hosted models (e.g., Llama via vLLM), including production tool use and agent capabilities.
- Ability to travel 0-50 percent, depending on client needs and engagement scope.
- Limited immigration sponsorship may be available.
Technologies
- LangGraph, LangChain, Python, vLLM, Llama
- Pinecone, Weaviate, Milvus
- LoRA, QLoRA
- RAG, Anthropic, Google, OpenAI
- FHIR
Preferred Qualifications
- Experience with multi-agent systems and agent collaboration patterns.
- Familiarity with vector databases and retrieval infrastructure such as Pinecone, Weaviate, or Milvus.
- Exposure to model adaptation and fine-tuning techniques such as LoRA or QLoRA.
- Understanding of traditional NLP concepts including tokenization, semantic similarity, entity extraction, summarization, and transformer fundamentals.
- Experience operating in regulated, high-stakes environments; healthcare exposure or standards such as FHIR is a plus, not required.
- Proven habit of staying current with AI research, benchmarks, and emerging engineering patterns.