Gen AI Engineer with Java & Microservices
Job Description
This role focuses on designing and implementing production-ready generative AI services that connect Java-based systems with modern LLM capabilities. You will build scalable microservices, deliver AI-enabled interfaces via REST and event-driven patterns, and support evaluation, monitoring, testing, and documentation within an agile delivery environment.
Responsibilities
- Design, develop, and deploy production-grade Java microservices (for example using Spring Boot) that integrate with large language models and AI/ML platforms.
- Architect and implement RESTful APIs and event-driven communication to expose generative AI capabilities to internal and external clients.
- Build prompt engineering strategies, including prompt templating, chaining, and dynamic context management, to optimize outputs for business use cases.
- Develop Retrieval-Augmented Generation (RAG) capabilities, integrating vector databases to produce grounded, context-aware responses from proprietary knowledge sources.
- Create resilient integration layers for model inference, including retry logic, timeouts, rate limiting, and fallback mechanisms to support high availability.
- Implement asynchronous processing and streaming responses for efficient handling of long-running inference workflows.
- Deliver monitoring, logging, and observability to track token usage, latency, and model performance.
- Partner with data scientists and ML engineers to translate experimental models into production-ready code, with a focus on architecture and performance optimization.
- Develop and execute unit, integration, and load tests to ensure code quality and system reliability under stress.
- Evaluate new generative AI models, libraries, and tools by producing technical assessments and recommendations.
- Write technical documentation and maintain internal AI service frameworks and shared libraries.
- Stay current on the GenAI landscape and propose improvements to architecture and development practices.
Requirements
- Bachelor’s degree in Computer Science, Software Engineering, or a related technical field (or equivalent practical experience).
- 5+ years of professional software development experience, emphasizing object-oriented programming and enterprise application development.
- Expert-level Java skills, including concurrency, functional programming patterns, and strong error handling.
- Hands-on experience with microservices architecture, including service discovery, API gateways, and distributed tracing.
- Strong knowledge of Spring Boot or a comparable Java microservices framework, including testing and configuration.
- Demonstrated experience integrating third-party AI/LLM APIs (such as OpenAI, Anthropic, or open-source models) in a server-side environment.
- Solid understanding of cloud platforms (AWS, GCP, or Azure) and experience with Docker and Kubernetes is highly desirable.
- Proficiency with SQL and NoSQL databases, plus practical experience using vector databases (for example Pinecone, Weaviate, pgvector) for semantic search.
- Working knowledge of observability practices, including logging frameworks, metrics collection, and monitoring dashboards.
- Strong problem-solving, debugging, and analytical abilities, with a proactive and self-directed work style.
- Effective communication skills for collaboration and the ability to explain complex technical concepts to non-technical stakeholders.
- Interest in learning new technologies, especially in generative AI and machine learning engineering.
Technologies
- Java, Spring Boot
- RESTful APIs, microservices architecture
- OpenAI, Anthropic
- Service discovery, API gateways, distributed tracing
- AWS, GCP, Azure
- Docker, Kubernetes
- SQL, NoSQL
- Vector databases: Pinecone, Weaviate, pgvector
- Retrieval-Augmented Generation (RAG)
Location
Atlanta, GA (onsite)
Minimum Experience
5 years