Principal Machine Learning Engineer
Job Description
Lead the technical direction of AI Canvas on the AI layer powering next-generation cybersecurity data exploration.
Responsibilities
- Provide technical leadership for the architecture and development of AI capabilities for AI Canvas, from early design through production deployment and continuous improvement
- Design and build scalable, production-grade systems using LLMs, RAG, and AI agents, including assistants, tools, skills, and agent harnesses
- Architect intelligent, natural-language workflows for exploring complex security data, including natural-language-to-query generation, follow-up interactions, clarification, investigation, and troubleshooting
- Define and evolve the architecture for agentic AI systems covering orchestration, context management, tool invocation, memory, reasoning workflows, and multi-step task execution
- Build robust evaluation frameworks and harnesses to measure AI quality, including correctness, relevance, reliability, regression detection, and end-to-end product behavior
- Establish evaluation methodologies using automated metrics, LLM-based evaluators, human evaluation, and representative production datasets
- Design and implement guardrails and safety mechanisms to improve reliability, reduce hallucinations, enforce system constraints, and support responsible AI behavior
- Drive improvements in model and system quality through prompt engineering, retrieval strategies, model selection, fine-tuning where appropriate, and systematic experimentation
- Build scalable RAG and knowledge-retrieval systems that ground responses in large, complex, and evolving security datasets
- Partner with backend and platform engineers to build reliable APIs, services, and infrastructure for AI workloads at production scale
- Establish best practices for observability, debugging, tracing, versioning, reproducibility, testing, and monitoring across LLM and agent-based systems
- Evaluate emerging models, frameworks, and AI infrastructure and determine when and how to incorporate them into the product
- Balance rapid experimentation with engineering rigor needed for mission-critical production AI systems
- Mentor engineers, influence technical direction across teams, and raise the engineering bar for applied AI development
Requirements
- Bachelor’s degree in Computer Science, Machine Learning, Engineering, or a related technical field, or equivalent practical experience
- 7+ years of software engineering, machine learning engineering, or related industry experience, including significant experience building production systems
- Strong experience designing and building ML or AI-powered applications at scale
- Hands-on experience building applications using large language models and generative AI technologies
- Experience with one or more: RAG, AI agents, AI assistants, tool-calling systems, agent orchestration, or LLM-based workflows
- Strong proficiency in Python and experience developing production-grade backend or ML services
- Experience designing evaluation systems for AI/ML applications, including offline evaluation, regression testing, quality measurement, and production monitoring
- Strong understanding of modern ML and AI concepts, including embeddings, retrieval, ranking, prompt engineering, model inference, and experimentation
- Experience building scalable systems on public cloud platforms such as GCP, AWS, or Azure
- Strong software engineering fundamentals: system design, distributed systems, APIs, testing, CI/CD, and observability
- Ability to lead complex technical initiatives across multiple teams while staying deeply hands-on
- Strong communication skills, translating ambiguous product problems into clear technical architectures and execution plans
Technologies
- LLMs, retrieval-augmented generation (RAG), AI agents, AI assistants
- Natural-language-to-query generation, natural-language-to-SQL
- Python
- GCP, AWS, Azure
- Embeddings, prompt engineering
- CI/CD, Docker, Kubernetes
- Vector databases
- LLM-based evaluators
Preferred Experience
- Master’s or PhD in Computer Science, Machine Learning, Artificial Intelligence, or a related technical field
- Deep experience building and operating LLM-powered products or agentic AI platforms in production
- Experience with AI development frameworks and tooling for model orchestration, agent systems, evaluation, tracing, or observability
- Experience building natural-language-to-SQL or natural-language-to-query systems, semantic layers, or AI-powered data exploration products
- Experience designing LLM evaluation harnesses, synthetic test generation, LLM-as-a-judge approaches, or human-in-the-loop evaluation workflows
- Experience with vector databases, search and retrieval infrastructure, embeddings, and large-scale knowledge systems
- Experience with containerization and orchestration technologies such as Docker and Kubernetes
- Experience designing highly available, low-latency AI inference and backend services
- Familiarity with AI security, prompt injection defenses, data privacy, access control, and guardrails for enterprise AI applications
- Experience in cybersecurity, security analytics, observability, or large-scale data platforms
- Contributions to open-source AI, ML, agent, or infrastructure projects
Compensation
- $163,200.00 - $264,000.00/yr
- Compensation depends on qualifications, experience, and work location
- For candidates who receive an offer at the posted level, starting base salary (non-sales) or base salary plus commission target (sales/commissioned) is expected to fall within the annual range listed
- Offered compensation may include restricted stock units and a bonus
Immigration Sponsorship
- Yes