Data Scientist II - ML Engineering
Artificial Intelligence
CI/CD
Data Pipeline
Data Platform
Data Science
Data Science Ml
Databricks
Devops Tools
Engineer
Engineering
Generative AI
Generative Ai Engineer
Kubernetes
Large Language Models
Machine Learning
Machine Learning Engineer
Machine Learning Infrastructure
Machine Learning Models
Machine Learning Pipelines
MLOps
Model Serving
Platform Engineering
Rag Architectures
SQL
Job Description
This role is focused on ML engineering, including the build and maintenance of an ML platform and the production model lifecycle. You will also contribute to Generative AI engineering, with emphasis on RAG, search, knowledge graphs, and multi-agent workflows.
What You’ll Do
- Design, create, and maintain an ML platform and related environments.
- Manage Docker containers and Kubernetes clusters, including dependencies and configurations.
- Implement CI/CD pipelines to automate building, testing, and deployment of machine learning models.
- Monitor and optimize model training performance and resource usage.
- Deploy ML models to production, with model versioning and rollback mechanisms.
- Enable scalable, reliable model serving using tools such as Vertex, Databricks, TensorFlow Serving, Flask, or FastAPI, to support increasing model complexity and volume.
- Collaborate with data scientists, data engineers, and stakeholders to understand and meet infrastructure needs.
- Stay current with the latest technologies and best practices in ML infrastructure.
- Architect and develop Generative AI solutions using Machine Learning and GenAI techniques.
- Specialize in engineering and deploying Generative AI models, with focus on Retrieval-Augmented Generation (RAG) systems, search, knowledge graphs, and multi-agent workflows.
- Handle both unstructured and structured data to prepare context for Language Model Learning (LLM), including embedding large text corpora, developing generative SQL queries, and building connectors to structured databases.
- Train models on prepared data and optimize fine-tuning hyperparameters for strong performance.
- Build a framework to stitch cross-domain learning and optimize toward mission-specific and multi-mission tasks.
- Serve as an expert in AI interpretation and causality by identifying model causality relationships and building a framework to measure bias, underspecification, and latent drivers with their connections.
- Create an enterprise domain-specific reasoning system to improve actionable insights and optimize resources used across the machine learning process.
- Orchestrate a reusable storytelling methodology to support AI translation, applying inquisitive, transparency-focused approaches to business stakeholders.
- Apply AI reasoning into business action recommendations and use AI research to accelerate business innovation.
Required Qualifications
- A related degree or comparable formal training, certification, or work experience.
- 5+ years of experience in a retail or retail-related decision science role.
- Expertise in ML visualization flow.
- Expertise in optimizing distributed machine learning in a heterogeneous domain environment.
- Technical knowledge in programming languages: SQL, R, Python, Scala, Java, C/C++.
- Technical knowledge in big data and ML optimization: GPU code optimization, Horovod, Spark MLlib optimization, Cython, JNI, Numba.
- Technical knowledge in mainstream ML/AI: manifold learning, distributed clustering, graph network, hierarchical model, Bayesian network, deep learning, computer vision, NLP/NLU, reinforcement learning, meta-Learning, federated learning.
- Technical skills to apply causal reasoning representation and learning, including human-centric, explainable, responsible AI.
- Ability to work comfortably with imperfect or incomplete data.
- Ability to apply AI reasoning into business action recommendation.
- Ability to work in a fast-paced retail environment with frequently shifting priorities.
- Work extended hours; sit for long periods.
Core Technologies
- Docker, Kubernetes, CI/CD
- Vertex, Databricks, TensorFlow Serving, Flask, FastAPI
- Machine Learning, GenAI, Retrieval-Augmented Generation (RAG), Language Model Learning (LLM)
- SQL, R, Python, Scala, Java, C/C++, GPU code optimization
- Horovod, Spark MLlib, Cython, JNI, Numba
- Manifold learning, distributed clustering, graph network, hierarchical model, Bayesian network, deep learning, computer vision, NLP/NLU, reinforcement learning, meta-Learning, federated learning
- Generative AI
Location
San Antonio, TX (onsite)