DataJobs.io
← Back to all jobs

Job Description

Deloitte in Raleigh, NC offers an onsite role focused on healthcare AI within an AI-first initiative. You will build end-to-end LLM- and SLM-powered agentic systems for healthcare decisioning and deploy them into live clinical and operational settings. This position blends advanced engineering with practical, regulated workflows and a collaborative, impact-driven culture. The role carries a competitive salary and the opportunity to contribute to cutting edge AI in healthcare.

Location and compensation

Onsite in Raleigh, North Carolina. Salary range of USD 134,500 to 265,100 per year. A bachelor’s degree is required.

Responsibilities

  • Design and implement agentic systems capable of multi-step reasoning, planning, tool use, and workflow execution against complex, regulated operational processes.
  • Build stateful workflows using frameworks such as LangGraph and LangChain, including branching, retries, self-correction, human-in-the-loop checkpoints, and reusable orchestration patterns.
  • Engineer for long-horizon reliability, enabling multi-step task completion, recovery from compounding errors, planning under uncertainty, and robust tool use when steps fail.
  • Build the reasoning behind regulated decisions with policy- and criteria-grounded outputs, structured proposer/critic/judge style review, and auditable rationales for high-stakes decisions across clinical review, prior authorization, claims integrity, and care management.
  • Develop end-to-end Retrieval-Augmented Generation pipelines, covering ingestion, chunking, embeddings, vector and hybrid retrieval, reranking, contextual compression, and grounding strategies.
  • Engineer memory and context management, including conversational state, persistent memory, retrieval-aware context assembly, and token-efficient context selection.
  • Apply modern context-delivery patterns to ensure agents access the right information at the right time.
  • Implement observability and tracing for prompts, tool calls, retrieval quality, agent traces, failures, drift, latency, and production behavior.
  • Apply guardrails, safety controls, and failure-handling to reduce hallucinations and unsafe actions.
  • Evaluate agents at the trajectory and task level with multi-step task success metrics, failure-mode analysis, sandboxed testing, and integrated quality checks alongside human review.
  • Engineer healthcare-grade safety with deployment eval gates, human oversight and escalation models, auditability and traceability for regulated decisions, and PHI/HIPAA-aware data handling.
  • Build integrations with internal and external tools, APIs, enterprise systems, databases, and model providers to operate safely within real business workflows.
  • Deliver production-quality code with robust testing, CI/CD, logging, versioning, and documentation; balance quality, safety, latency, cost, and model risk in architectural decisions.
  • Partner with modeling and post-training engineers to improve model behavior for tool use, grounding, and long-horizon reasoning through evaluation-driven feedback and fine-tuning where appropriate.
  • Translate ambiguous, high-complexity operational processes into robust system logic and reusable AI patterns; stay current with agentic system advances and translate research into practical engineering decisions.

Requirements

  • Bachelor's degree in Computer Science, Engineering, Data Science, Computational Linguistics, or a related field.
  • Demonstrated depth building and shipping production agentic systems as a primary craft, with a track record of shipped systems, research, model releases, and open source work over years.
  • Strong hands-on experience building production agent systems with modern orchestration, including LangGraph/LangChain or equivalent with custom orchestration.
  • Experience designing and optimizing end-to-end RAG systems: indexing, retrieval, reranking, grounding, and evaluation.
  • Solid understanding of memory and context management, including context windows, retrieval-driven context assembly, persistent memory, and high-signal context selection.
  • Deep practical knowledge of LLM behavior, limitations, hallucination risks, reasoning constraints, and latency/cost trade-offs, plus evaluation methodologies.
  • Experience evaluating and debugging agent behavior through task-success and trajectory analysis, not just output quality.
  • Strong Python engineering skills and modern software practices: testing, CI/CD, version control, API integration; experience implementing observability, tracing, and debugging for production LLM-based systems.
  • Hands-on experience with at least one frontier model platform (e.g., Anthropic, Google, OpenAI) and/or open-weight/self-hosted models (e.g., Llama via vLLM), including production tool use and agent capabilities.
  • Ability to travel 0-50 percent, depending on client needs and engagement scope.
  • Limited immigration sponsorship may be available.

Technologies

  • LangGraph, LangChain, Python, vLLM, Llama
  • Pinecone, Weaviate, Milvus
  • LoRA, QLoRA
  • RAG, Anthropic, Google, OpenAI
  • FHIR

Preferred Qualifications

  • Experience with multi-agent systems and agent collaboration patterns.
  • Familiarity with vector databases and retrieval infrastructure such as Pinecone, Weaviate, or Milvus.
  • Exposure to model adaptation and fine-tuning techniques such as LoRA or QLoRA.
  • Understanding of traditional NLP concepts including tokenization, semantic similarity, entity extraction, summarization, and transformer fundamentals.
  • Experience operating in regulated, high-stakes environments; healthcare exposure or standards such as FHIR is a plus, not required.
  • Proven habit of staying current with AI research, benchmarks, and emerging engineering patterns.

Similar Jobs