DataJobs.io
← Back to all jobs

Job Description

Build and scale governed AI systems for Violet and the broader ThirdBase analytics platform using LLMs, retrieval, and healthcare data.

Responsibilities

  • Develop and improve LLM-powered analytics including agentic workflows, natural-language querying, RAG, and tool-use over structured data (SQL/Snowflake) and unstructured content (documents)
  • Contribute to text-to-SQL and semantic-layer systems that help non-technical users query complex healthcare datasets with accuracy and safety
  • Build retrieval pipelines for document stores using vector search and hybrid keyword + semantic retrieval
  • Implement evaluation, guardrails, and hallucination checks with a focus on accuracy for healthcare domain data
  • Support data-governance and compliance in the AI layer, including derived-insights-only design and avoiding PHI/PII exposure with appropriate access controls
  • Integrate multiple tools and data sources (warehouse, campaign platforms, web, document repositories) into agent workflows
  • Deploy and maintain services on AWS, monitoring model performance, latency, and cost
  • Partner with product, analytics, and senior engineering to translate requirements into shipped features

Requirements

  • 2–4 years building production software, with hands-on experience in LLM/GenAI applications such as RAG, agents, or text-to-SQL (work, internships, or substantial personal projects)
  • Strong Python and SQL; comfort querying a cloud data warehouse (Snowflake is a plus)
  • Working knowledge of AWS and deploying services in the cloud
  • Substantive experience with agentic frameworks such as LangGraph, CrewAI, n8n, or similar
  • Experience with AI-assisted coding and CI/CD processes
  • Familiarity working with Node, React, or other JavaScript frameworks
  • Exposure to retrieval systems, embeddings, and vector search
  • Experiments-driven design using evaluation harnesses for change management
  • Understanding of prompt engineering and LLM guardrails focused on reliability of outputs
  • Awareness of data privacy and compliance basics and willingness to build with them
  • Ability to own a feature or component end-to-end and collaborate across a team

Nice to Have

  • Experience with healthcare or life-sciences data (claims, referrals, HCP, ICD-10) or another regulated domain
  • Familiarity with HIPAA and PHI/PII constraints, including de-identified or derived-insights data models
  • Exposure to agent orchestration frameworks and multi-tool workflows
  • Basic MLOps/LLMOps: monitoring and cost/latency optimization
  • Background in analytics, BI, or data engineering

Technology Stack

  • Cloud: AWS (Lambda, S3, ECS/Fargate, API Gateway)
  • Data warehouse: Snowflake (SQL, semantic layers)
  • Languages: Python, SQL
  • AI/ML: LLMs (OpenAI/Anthropic-class models, Amazon Bedrock), RAG, embeddings, vector search, prompt engineering, function-calling/tool-use
  • Data & document tooling: vector databases, hybrid search, Egnyte, Google Drive
  • Analytics libraries: pandas and modern data-analysis/plotting tooling
  • Governance: compliance-aware data access, PII/PHI handling, source citation
  • Agent frameworks & orchestration: LangGraph, CrewAI, n8n

Location

  • Nashville, TN (onsite)

Department / Reporting

  • Department: Data Solutions & Analytics
  • Reports to: VP, Data Products & Technology

Company Context

BPD is a strategic growth partner delivering technology-enabled, AI-infused solutions for healthcare’s leading brands, supported by a proprietary data platform.

Similar Jobs