DataJobs.io
← Back to all jobs

Job Description

Lead technical ownership of OCI AI platform capabilities and enterprise, production-grade agentic AI systems on Oracle Cloud Infrastructure (OCI).

Responsibilities

  • Act as the senior technical owner of OCI AI platform capabilities, covering agent execution, inference systems, model serving, AI workflow orchestration, evaluation, and observability.
  • Design and deliver scalable agentic AI systems that can reason, plan, use tools, execute workflows, orchestrate multi-step tasks, and support safe human-in-the-loop escalation.
  • Develop production-grade services for tool calling, agent memory, context management, MCP integration, vector retrieval, multi-agent coordination, policy enforcement, and evaluation.
  • Steer architecture across distributed services optimized for low latency, high throughput, GPU efficiency, reliability, cost, operability, and secure multi-tenant operation.
  • Define service boundaries, APIs, data models, state management, consistency tradeoffs, failure modes, SLIs/SLOs, rollout strategies, and readiness criteria for AI platform services.
  • Guide technical strategy across infrastructure, platform, security, data, and application engineering teams, translating broad goals into multi-quarter plans with measurable milestones.
  • Integrate AI agents securely and reliably with enterprise APIs, cloud services, databases, identity systems, secrets management, and external systems.
  • Establish AgentOps and LLMOps practices for tracing, monitoring, evaluation suites, regression testing, experimentation, safety guardrails, prompt and tool versioning, and production reliability.
  • Assess and operationalize emerging technologies in generative AI, agentic workflows, inference optimization, long-context systems, reasoning models, AI developer tooling, and agentic-first development.
  • Promote engineering excellence through code and design reviews, test strategy, deployment automation, incident analysis, documentation, and AI-assisted development using tools like Codex, Claude Code, Cursor, Copilot, or equivalents.
  • Mentor staff and senior engineers, raise architectural standards, and influence OCI engineering practices without requiring direct management authority.
  • Own critical production outcomes including reliability, performance, security posture, cost efficiency, and supportability for delivered systems.

Requirements

  • Advanced degree (PhD, MS, or BS with equivalent practical experience) in computer science, AI/ML, engineering, or related field.
  • 12+ years of professional software engineering experience with ownership of production systems, or equivalent impact at senior staff / principal level.
  • Proven track record as a staff, senior staff, principal, or equivalent technical leader influencing architecture and execution across multiple teams.
  • Deep experience designing, building, and operating high-scale distributed systems, cloud services, infrastructure platforms, or AI/ML platform services.
  • Hands-on experience with production AI systems, agentic AI applications, autonomous workflows, tool-using agents, multi-step orchestration, or multi-agent systems.
  • Practical experience with orchestration frameworks such as LangGraph, LangChain, CrewAI, AutoGen, LlamaIndex, or similar ecosystems.
  • Strong understanding of LLM application patterns including prompt design, structured outputs, function/tool calling, context management, RAG, memory, tool safety, and evaluation.
  • Proficiency in Python with ability to contribute production code, reviews, tests, and debugging in complex distributed environments.
  • Expertise with Kubernetes, Docker, cloud-native infrastructure, service-to-service communication, scalability, fault tolerance, observability, and performance analysis.
  • Experience defining SLIs/SLOs, production readiness criteria, incident response practices, monitoring, tracing, experiments, and reliability programs for AI or distributed systems.
  • Strong understanding of AI safety, governance, security, and operational risks for autonomous or semi-autonomous systems, including data handling, access control, auditability, and human accountability.
  • Excellent written and verbal communication, with proven ability to lead technical direction, resolve ambiguity, and influence senior stakeholders.

Technologies

  • Python
  • Kubernetes
  • Docker
  • LangGraph
  • LangChain
  • CrewAI
  • AutoGen
  • LlamaIndex
  • Model Context Protocol (MCP)
  • Oracle Cloud Infrastructure (OCI)
  • Codex
  • Claude Code
  • Cursor
  • Copilot

Preferred qualifications

  • Experience optimizing large-scale GPU inference or training workloads for latency, throughput, utilization, availability, and cost.
  • Experience building or operating model serving, inference gateways, agent runtimes, workflow engines, developer platforms, or internal AI productivity platforms.
  • Experience integrating AI systems with enterprise APIs, databases, cloud services, vector databases, embeddings, retrieval systems, identity systems, and policy enforcement layers.
  • Experience with LLM fine-tuning, long-context systems, reasoning models, model routing, caching, batching, quantization, or related generative AI research.
  • Experience building evaluation frameworks for agentic systems, including offline evals, online experiments, golden tasks, adversarial testing, regression gates, and observability dashboards.
  • Experience using AI-assisted software development tools such as Codex, Claude Code, Cursor, Copilot, or similar in large-scale engineering environments.
  • Track record of defining architectural standards, platform capabilities, or engineering practices adopted across multiple teams or organizations.
  • Experience in enterprise, cloud infrastructure, regulated, security-sensitive, or mission-critical environments.

Similar Jobs