DataJobs.io
← Back to all jobs

Job Description

Xenoss is seeking a Staff AI Engineer / AI Solution Architect to lead the applied AI architecture for a long-term In-Call Assistant initiative in New York, NY (onsite). This role shapes a real-time conversational AI system that helps front-office employees during live customer interactions, with compliance guardrails, confidence management, and a clear path from PoC-first development to production-grade conversation intelligence.

What you’ll build

You’ll define the end-to-end AI approach for conversation intelligence, including the pipeline for low-latency signal detection, context preparation, specialist recommendation generation, and RAG over approved product and policy knowledge. The system is designed to identify customer needs, objections, and buying signals, then recommend required process steps using grounded, policy-compliant behavior.

Responsibilities

  • Lead the applied AI architecture across the In-Call Assistant lifecycle, from data and taxonomy design to model training, evaluation, and production readiness.
  • Design the end-to-end AI architecture for the In-Call Assistant.
  • Define signal and trigger taxonomies for live conversations.
  • Design training strategies for signal detection and specialist recommendation models.
  • Shape data preparation, annotation, and SME validation workflows.
  • Evaluate fine-tuning, post-training, RAG, and hybrid approaches.
  • Design low-latency signal detection, routing, context preparation, and confidence management.
  • Design evaluation frameworks, golden datasets, and model improvement cycles.
  • Define grounding, guardrails, abstention, and policy-compliance behavior.
  • Make trade-offs across model quality, latency, cost, explainability, and governance.
  • Partner with AI engineering, data engineering, MLOps, and client SMEs.
  • Act as an escalation point for AI architecture, evaluation, and data strategy decisions.
  • Translate ambiguous business use cases into testable AI hypotheses and validation plans.

PoC-first scope and delivery context

  • Define the AI approach for the conversation intelligence PoC.
  • Establish the event / intent / insight taxonomy.
  • Define the golden dataset strategy and annotation workflow.
  • Establish evaluation frameworks and acceptance criteria.
  • Drive trade-offs between accuracy, explainability, latency, cost, and governance.
  • Decide which modeling approaches fit each use case.
  • Work within a cross-functional team spanning AI engineering, data engineering, MLOps, solution architecture, and client stakeholders.

What you bring

  • Hands-on experience with applied AI / ML systems in production-oriented environments.
  • Experience with NLP, conversational AI, or transcript-based intelligence systems.
  • Ability to design evaluation frameworks, not only run experiments.
  • Experience building or validating structured datasets from unstructured text.
  • Strong understanding of LLM-based extraction, classification, RAG, and fine-tuning trade-offs.
  • Practical knowledge of classical ML or predictive modeling.
  • Understanding of probability-based prediction, calibration, and outcome evaluation.
  • Comfort working with messy enterprise data and incomplete labels.
  • Ability to communicate with both technical teams and business stakeholders.
  • Strong ownership of ambiguity, scope control, and PoC validation strategy.
  • Financial services domain exposure.
  • Experience with sales, call center, or customer conversation analytics.
  • Speech / ASR pipeline familiarity.
  • Model governance and auditability experience.
  • Experience with real-time AI systems or low-latency inference.
  • Experience combining unstructured conversation signals with structured CRM, transaction, or customer profile data.
  • Experience designing golden datasets and SME review workflows.

Technologies you’ll use

  • LLM, SFT, DPO, preference optimization, LoRA, QLoRA, PEFT, PyTorch, Hugging Face
  • RAG, embeddings, retrieval
  • MLOps, monitoring, feedback loops, and model governance

Infrastructure and data residency: work is executed within the client perimeter using the client environment only (no external training or data processing environments). Delivery is PoC-first, with evolution toward live production conversation intelligence and prediction systems. Engagement is FTE-equivalent via long-term B2B contract.

Similar Jobs