AI Engineer
Job Description
Build and scale AI platform infrastructure for AI/ML and generative AI workloads on Google Cloud at Capgemini. This contract role is based in Charlotte, NC (onsite) and focuses on cloud platform engineering, CI/CD automation, and MLOps practices to support reliable, secure, and scalable delivery. If you enjoy turning experimentation into production-ready systems, you will work closely with AI/ML engineers and data scientists to streamline the full path from training to monitoring.
What you’ll do
- Design, build, and maintain scalable GCP cloud infrastructure for AI/ML and application workloads
- Architect and manage CI/CD pipelines for model training, deployment, and application release workflows (Cloud Build, GitHub Actions, Jenkins, GitLab CI, or similar)
- Create and run MLOps pipelines for model versioning, training, deployment, and monitoring using tools such as Vertex AI Pipelines, Kubeflow, and MLflow
- Use Infrastructure as Code (Terraform, Deployment Manager) for repeatable and auditable environment provisioning
- Containerize and orchestrate services with Docker and Kubernetes (GKE)
- Set up observability including logging and alerting for AI/ML services using Cloud Monitoring, Cloud Logging, and Prometheus/Grafana
- Partner with AI/ML engineers and data scientists to productionize models and reduce friction between experimentation and deployment
- Implement platform security with IAM, secrets management, and compliance across environments
- Drive automation to reduce manual toil across build, test, deployment, and rollback workflows
- Troubleshoot platform and infrastructure issues and lead root-cause resolution for production incidents
What you bring
- 7+ years of experience in platform engineering, DevOps, or infrastructure engineering
- Hands-on expertise in Google Cloud Platform (GCP), including Vertex AI, GKE, Cloud Run, IAM, VPC/networking, and BigQuery
- Deep experience designing and managing CI/CD pipelines, including options such as Cloud Build, Jenkins, GitHub Actions, GitLab CI, and ArgoCD
- Proficiency with Infrastructure as Code (Terraform preferred)
- Strong scripting and programming skills in Python, Bash, or Go
- Experience with containerization and orchestration (Docker, Kubernetes)
- Working knowledge of MLOps concepts across the ML lifecycle: training, versioning, deployment, and monitoring
- Experience with observability tooling such as Cloud Monitoring, Prometheus, Grafana, and ELK/EFK
- Strong understanding of cloud security, IAM, and secrets management best practices
Tech stack you may work with
GCP, Vertex AI, GKE, Cloud Run, IAM, VPC/networking, BigQuery, Cloud Build, GitHub Actions, Jenkins, GitLab CI, Vertex AI Pipelines, Kubeflow, MLflow, Terraform, Deployment Manager, Docker, Kubernetes, Prometheus, Grafana, Cloud Monitoring, Cloud Logging, Python, Bash, Go, ELK/EFK, ArgoCD.
Benefits
- Medical, dental, vision, and retirement benefits