AI Engineer
Job Description
Entertainment Partners is hiring a Senior Software Engineer to build, train, evaluate, and deploy production AI and agentic systems for its intelligent product suite.
Responsibilities
- Design, develop, train, fine-tune, and evaluate machine learning models using PyTorch and related ecosystem libraries (torchvision, torchaudio, torch.nn, torch.optim).
- Build and maintain ML training pipelines, experiment tracking workflows, and model evaluation frameworks.
- Implement transformer-based models and LLM integrations for production use cases including NLP, information extraction, classification, and generation.
- Use parameter-efficient fine-tuning approaches (LoRA, QLoRA, PEFT) to adapt foundation models for EP domains including payroll, residuals, and production management.
- Design and implement RAG (Retrieval-Augmented Generation) using vector databases (pgvector, Pinecone, Weaviate) and semantic search pipelines.
- Optimize inference for latency and throughput with quantization, batching, and caching for production serving.
- Develop and maintain AI evaluation frameworks with automated evals as unit tests to support reliable, safe, production-grade behavior.
- Implement LLM-powered agentic workflows using LangChain, LangGraph, and EP’s internal MCP (Model Context Protocol) server architecture.
- Build multi-step reasoning pipelines, tool-calling agents, and autonomous task execution systems integrating with EP’s enterprise data and product APIs.
- Apply prompt engineering strategies including few-shot templates, chain-of-thought scaffolding, and structured output validation.
- Maintain EP AI quality engineering (QE) standards such as failure taxonomy, runtime guardrails, and evidence-driven release gates.
- Contribute to EP’s Enterprise Context Engine, a governed zero-data-retention AI context layer exposed via MCP to Tabnine Agent and Claude Code.
- Build and maintain MLOps infrastructure for training, experiment tracking (MLflow, Weights & Biases), versioning, and deployment.
- Containerize and deploy ML services with Docker and Kubernetes; integrate with CI/CD using GitHub Actions and Azure DevOps.
- Monitor model performance in production using drift detection, feedback loops, and automated retraining triggers.
- Ensure AI systems meet security, privacy, and compliance needs including data minimization and access control for sensitive payroll data.
- Collaborate with data engineering on feature stores, data pipelines, and training data infrastructure.
- Partner with the Chief Architect AI & Data and CAIO to define AI architecture patterns and best practices.
- Work with product managers, UX designers, and full stack engineers to translate AI capabilities into product features.
- Conduct AI/ML code reviews focused on reproducibility, correctness, and production readiness.
- Mentor engineers on AI engineering fundamentals, LLM integration patterns, and responsible AI practices.
- Stay current with the AI/ML landscape, evaluate new models/frameworks, and assess applicability.
- Advance EP’s AI capability maturity by contributing to the PE AI Maturity Scorecard (S1–S3).
- Represent EP’s AI engineering practices in Architecture Review Board discussions.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Machine Learning, Statistics, Mathematics, or a related quantitative field.
- 6–10+ years of professional software engineering experience, including 3+ years focused on ML/AI engineering in production environments.
- Expert-level Python proficiency and deep familiarity with the Python ML/AI ecosystem.
- Hands-on production experience with PyTorch: nn.Module, custom training loops, autograd, GPU acceleration (CUDA), and model serialization (TorchScript, ONNX).
- Experience with Hugging Face libraries: Transformers, Datasets, and PEFT for fine-tuning/adaptation.
- Demonstrated ability building RAG pipelines including chunking, embedding models, vector store selection, and retrieval evaluation.
- Production experience integrating LLM APIs (OpenAI, Anthropic, open-source via vLLM/Ollama) and building reliable prompt engineering systems.
- Experience with LangChain or LangGraph for multi-step agents and tool-calling workflows.
- Strong ML fundamentals: supervised/unsupervised learning, loss functions, regularization, evaluation metrics, and statistical validation.
- Experience with experiment tracking and reproducible workflows using MLflow, Weights & Biases, or Comet.
- Working knowledge of Docker and cloud ML services (AWS SageMaker, Azure ML, or OCI Data Science).
- Experience with SQL and NoSQL databases to design pipelines for ML training and inference.
Technologies
- Python, PyTorch, torchvision, torchaudio, torch.nn, torch.optim
- LoRA, QLoRA, PEFT
- RAG, pgvector, Pinecone, Weaviate
- quantization, MLflow, Weights & Biases, Docker, Kubernetes
- GitHub Actions, Azure DevOps
- LangChain, LangGraph, MCP (Model Context Protocol)
- Tabnine Agent, Claude Code
- TorchScript, ONNX
- Hugging Face Transformers, Hugging Face Datasets, CUDA
- OpenAI, Anthropic, vLLM, Ollama
- Comet, AWS SageMaker, Azure ML, OCI Data Science
- SQL, NoSQL, vector databases
- Triton Inference Server, TorchServe, Ray Serve
- TensorFlow, JAX, OpenCV, Kubeflow, KFServing
- Enterprise Context Engine, MLOps, CI/CD
Benefits
- Health, Dental, and Vision options
- 401(k) retirement savings plan and company match
- Paid holidays, vacation time, and sick time
- Participation in company equity plans
- Employee Assistance Program and mental health and wellness programs
- Training and development
- Annual bonus and merit reviews
Location & Compensation
- Tempe, AZ (hybrid)
- $140,000 - $180,000 per year
Preferred Qualifications
- Experience with additional deep learning frameworks (TensorFlow, JAX) or framework interoperability (ONNX).
- Familiarity with computer vision (torchvision, OpenCV) or speech/audio processing (torchaudio) domains.
- Experience with model compression techniques: quantization (INT8, FP16, BF16), pruning, distillation.
- Experience serving ML models at scale using Triton Inference Server, TorchServe, Ray Serve, or similar.
- Contributions to open-source ML projects or published research (papers, patents, or technical blog posts).
- Experience with responsible AI frameworks, bias evaluation, and AI governance practices.
- Familiarity with MCP (Model Context Protocol) server development for exposing tools to AI agents.
- Prior domain experience in payroll, fintech, media, or enterprise SaaS environments.
- Experience with Kubernetes-based ML workload orchestration (Kubeflow, KFServing, or similar).
- Hybrid work environment: Burbank, CA headquarters with flexible remote schedule.
- On-call availability as needed for production AI system incidents and model deployment events.
- Access to GPU-accelerated compute environments (cloud-based) for model training workloads.
- Ability to sit for extended periods of time at a computer workstation.
- Dexterity of hands and fingers to operate a computer keyboard and mouse.
- Occasional participation in early-morning or evening sessions to coordinate with distributed teams or international partners.