Applied AI Engineer
Agentic Ai
Agentic Automation
Ai Agent
Ai Agent Platform
Analytics
Artificial Intelligence
Artificial Intelligence Engineer
Azure
Azure Ai
Azure Ai Search
Azure Ai Services
Azure Openai
Business Analytics
Business Intelligence
Cloud
Cloud Data Engineering
Cloud Data Platform
Cloud Platform
Cloud Platforms
Cloud Platforms Cloud Platforms
Data Analysis
Data Analytics
Data Analytics Tools
Data Architecture
Data Engineering
Data Integration
Data Pipeline
Data Platform
Data Processing
Data Science
Engineer
Generative AI
Generative Ai Applications
Generative Ai Engineer
Generative Ai Platform
Hr Technology
Information Technology (IT)
Microsoft
Microsoft Agent Framework
Office Tools
Power BI
Power Platform
Programming
Programming Language
Programming Languages
Rag Architectures
Reporting and Analytics
SQL
Job Description
ERCOT, the Electric Reliability Council of Texas, is seeking an Applied AI Engineer to build and deploy production generative AI capabilities in a regulated environment. The role focuses on architecting reliable systems, including agentic workflows and RAG pipelines, with governance, evaluation, and dependable operations.
What you’ll do
- Translate ambiguous business problems into scoped technical roadmaps, identifying constraints such as data access, compliance, latency, and cost before development starts.
- Design and build production agentic systems covering planning, tool-calling, multi-step reasoning, memory, and error recovery using orchestration frameworks such as LangGraph or Microsoft Agent Framework.
- Implement production RAG pipelines, including chunking, embeddings, hybrid search, reranking, retrieval-quality evaluation, and content freshness.
- Build and extend connectors that provide agents secure, standardized access to enterprise tools and data.
- Deploy applications to managed cloud platforms and integrate them with enterprise systems and collaboration tools.
- Create evaluation suites, tracing, and rollback paths so agent behavior is reliable in production, not limited to demonstrations.
- Monitor, debug, and continuously improve deployed applications using evaluation metrics.
- Apply system design fundamentals by defining architecture, data flows, and integration boundaries with attention to scalability, reliability, latency, and cost.
- Codify repeatable patterns by turning successful builds into reusable components and reference architecture.
- Work with non-technical business owners to understand workflows, maintain awareness of evolving LLM capabilities, and apply current implementation patterns and AI development stacks.
Requirements
- Proven experience building and deploying production-grade autonomous agents, not prototypes.
- Experience with agent orchestration frameworks such as LangGraph, Microsoft Agent Framework, or comparable tools.
- Production RAG experience using vector search and vector databases such as pgvector, Azure AI Search, or Databricks Vector Search.
- Strong Python skills and hands-on integration with LLM APIs.
- System design fundamentals including scalable, reliable, maintainable services, API and integration-boundary design, and trade-offs across latency, throughput, and cost.
- Ability to build or extend tool and data connectors for LLM applications.
- Experience deploying and operating applications on a managed cloud platform.
- Knowledge of AI governance, model lifecycle practices, and evaluation methodology.
- Stakeholder and discovery skills to scope ambiguity, work with non-technical business owners, and operate autonomously.
Technologies you’ll work with
- Agent & LLM frameworks: LangGraph, Microsoft Agent Framework, LangChain, LlamaIndex
- LLM platforms & APIs: Claude API, Azure OpenAI, OpenAI API, model routing and evaluation frameworks
- AI coding assistants: Claude Code, OpenAI Codex, GitHub Copilot, Microsoft Copilot Studio
- Retrieval & vector search: Azure AI Search, Databricks Vector Search, pgvector
- Data & analytics: Databricks, Power BI, SQL, Oracle DB, PostgreSQL
- Connectors & integration: MCP (Model Context Protocol), REST APIs, enterprise system connectors, Teams integration
- Cloud & deployment: Azure, OpenShift, Docker, Kubernetes, Helm
- CI/CD & source control: GitHub, GitHub Actions, Git pull-request workflows
- Observability & evaluation: Tracing, evaluation harnesses, LLM observability, logging and monitoring
- ITSM & Agile tools: ServiceNow, Jira
- Scripting: Python, PowerShell
Preferred qualifications
- Solution and system architecture across multiple applications, including security-by-design and reference architecture.
- Experience with large-scale data platforms such as Databricks for retrieval, feature work, or pipeline development.
- Experience in regulated industries (energy, finance, healthcare) or audit-driven environments.
- Background in multi-agent orchestration and context engineering.
Education, location, and compensation
- Minimum experience: 5 years
- Education: Bachelor’s degree in Computer Science, Data Science, Information Systems, Engineering, or a related field (or equivalent knowledge gained through a combination of education and experience)
- Certification (preferred): Cloud or AI/ML certification such as Azure AI Engineer, AWS Machine Learning, or Databricks
- Location: Taylor, TX (hybrid, 2 days per week)
- Salary: USD 145,000 - 200,000 per year