AI Engineer
Job Description
LawPro.ai is hiring an AI Engineer to own how large language models power its data insights and analytics platform. This Virginia onsite role brings together AI research and production engineering, with end-to-end responsibility for model evaluation, selection, orchestration pipeline optimization, monitoring, and documentation. You will help keep AI systems accurate, cost-effective, and resilient as the LLM ecosystem evolves.
What you’ll do
- Run a systematic, ongoing evaluation process for new and emerging LLMs, benchmarking them against the specific tasks in the orchestration pipeline while continuously optimizing accuracy, relevance, speed, and cost outcomes.
- Build and maintain evaluation frameworks (Evals) and help establish an EvalOps culture to measure LLM output quality, including accuracy, relevance, faithfulness, and speed. A key focus area is reducing hallucinations in medical record summarization and legal document analysis.
- Monitor the LLM landscape across providers, identify deprecation timelines, and execute full model transitions by integrating replacements into the production pipeline and adjusting for model behavior. Take initiative to decommission stale, costly, or lower-performing legacy prompts and endpoints.
- Directly implement orchestration pipeline optimizations for document understanding, medical record summarization, case chronology generation, and drafting support. Own code changes, deployments, and production validation, with a preference for precise execution over large refactors.
- Communicate model evaluation findings with product and GTM stakeholders, then lead the technical implementation yourself, ensuring smooth handoffs across discovery, staging, and live production.
- Ship model changes into production end-to-end by writing integration code, managing deployments, running validation tests, and ensuring a clean rollout.
- Implement monitoring and observability for production performance, benchmarking outputs and cost, detecting drift, and providing continuous reporting to management. Use micro-benchmarking to track token-level latency, output drift, and cost efficiency across pipeline components.
- Maintain thorough documentation covering evaluation methodologies, model comparison results, transition decisions, and runbooks for the systems you own.
What you bring
- 5+ years of AI/ML engineering experience evaluating, fine-tuning, and deploying large language models in production, including building and deploying models to AWS or GCP infrastructure at scale.
- Hands-on development and implementation of multiple RAG solutions.
- Hands-on experience with embedding models and vector databases.
- Hands-on experience building agentic workflows and implementing EvalOps or Evals-as-a-Service architecture.
- Deep familiarity with the LLM ecosystem and the ability to assess model fit and tradeoffs, including heuristic-gated model routing across cost, quality, speed, and capabilities.
- Proven experience designing and operating evaluation frameworks to measure LLM quality, including accuracy, relevancy, and hallucination detection in high-stakes domains such as legal or medical.
- Strong software engineering foundation with production-deployed solutions, including LLM orchestration frameworks and multi-model pipelines.
- Ability to work in a fast-paced, high-ambiguity environment with strong ownership and tight feedback loops, prioritizing systematic process-building.
- Excellent communication skills to translate complex model evaluation findings into clear recommendations for engineering, product, and non-technical stakeholders.
- Bonus: experience with unstructured medical or legal document processing, or background in classical ML (statistics, embeddings, retrieval-augmented generation).
Tech stack
- AWS
- GCP
Role fit: You will own evaluation, selection, and continuous optimization of the large language models and AI processes powering LawPro.ai’s data insights and analytics platform, proactively managing transitions to new models and technologies while handling both recommendation and implementation with end-to-end technical rigor.