DataJobs.io
← Back to all jobs

Job Description

This role focuses on designing and implementing production-ready generative AI services that connect Java-based systems with modern LLM capabilities. You will build scalable microservices, deliver AI-enabled interfaces via REST and event-driven patterns, and support evaluation, monitoring, testing, and documentation within an agile delivery environment.

Responsibilities

  • Design, develop, and deploy production-grade Java microservices (for example using Spring Boot) that integrate with large language models and AI/ML platforms.
  • Architect and implement RESTful APIs and event-driven communication to expose generative AI capabilities to internal and external clients.
  • Build prompt engineering strategies, including prompt templating, chaining, and dynamic context management, to optimize outputs for business use cases.
  • Develop Retrieval-Augmented Generation (RAG) capabilities, integrating vector databases to produce grounded, context-aware responses from proprietary knowledge sources.
  • Create resilient integration layers for model inference, including retry logic, timeouts, rate limiting, and fallback mechanisms to support high availability.
  • Implement asynchronous processing and streaming responses for efficient handling of long-running inference workflows.
  • Deliver monitoring, logging, and observability to track token usage, latency, and model performance.
  • Partner with data scientists and ML engineers to translate experimental models into production-ready code, with a focus on architecture and performance optimization.
  • Develop and execute unit, integration, and load tests to ensure code quality and system reliability under stress.
  • Evaluate new generative AI models, libraries, and tools by producing technical assessments and recommendations.
  • Write technical documentation and maintain internal AI service frameworks and shared libraries.
  • Stay current on the GenAI landscape and propose improvements to architecture and development practices.

Requirements

  • Bachelor’s degree in Computer Science, Software Engineering, or a related technical field (or equivalent practical experience).
  • 5+ years of professional software development experience, emphasizing object-oriented programming and enterprise application development.
  • Expert-level Java skills, including concurrency, functional programming patterns, and strong error handling.
  • Hands-on experience with microservices architecture, including service discovery, API gateways, and distributed tracing.
  • Strong knowledge of Spring Boot or a comparable Java microservices framework, including testing and configuration.
  • Demonstrated experience integrating third-party AI/LLM APIs (such as OpenAI, Anthropic, or open-source models) in a server-side environment.
  • Solid understanding of cloud platforms (AWS, GCP, or Azure) and experience with Docker and Kubernetes is highly desirable.
  • Proficiency with SQL and NoSQL databases, plus practical experience using vector databases (for example Pinecone, Weaviate, pgvector) for semantic search.
  • Working knowledge of observability practices, including logging frameworks, metrics collection, and monitoring dashboards.
  • Strong problem-solving, debugging, and analytical abilities, with a proactive and self-directed work style.
  • Effective communication skills for collaboration and the ability to explain complex technical concepts to non-technical stakeholders.
  • Interest in learning new technologies, especially in generative AI and machine learning engineering.

Technologies

  • Java, Spring Boot
  • RESTful APIs, microservices architecture
  • OpenAI, Anthropic
  • Service discovery, API gateways, distributed tracing
  • AWS, GCP, Azure
  • Docker, Kubernetes
  • SQL, NoSQL
  • Vector databases: Pinecone, Weaviate, pgvector
  • Retrieval-Augmented Generation (RAG)

Location

Atlanta, GA (onsite)

Minimum Experience

5 years

Similar Jobs