Lead Data Engineer – AI/Machine Learning
Manager
Apache Airflow
Artificial Intelligence
Big Data
Bigdata
Cloud Data Engineering
Cloud Data Platform
Cloud Platform
Cloud Platforms
Data
Data Analysis
Data Analytics
Data Architecture
Data Build Tool
Data Engineer
Data Engineering
Data Engineering Lead
Data Integration
Data Management
Data Modeling
Data Operations
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Data Science
Data Warehouse
Database
Databases
Databricks
DevOps
ETL
Informatica
Llm Operations
Machine Learning
Ml Ops
Programming Language
Programming Languages
Spark
Streaming Data
Vector Databases
Job Description
Lead AI/Machine Learning Data Engineering role partnering with the VP, Head of Data to drive AI/ML enablement and readiness across the organization.
Responsibilities
- Design, build, and optimize data pipelines, ingestion frameworks, and platform components for analytics, reporting, and AI/ML use cases.
- Operate with autonomous ownership of complex initiatives, owning technical design through implementation and rollout with minimal oversight.
- Identify and resolve performance, scalability, and reliability issues across the existing data platform.
- Propose improvements proactively by recognizing gaps in data engineering and platform capabilities.
- Produce clean, well-tested, well-documented code and infrastructure-as-code while maintaining strong engineering hygiene.
- Help define AI/ML frameworks by evaluating and recommending tools, platforms, and standards for building and deploying AI/ML solutions.
- Build working AI/ML prototypes that deliver immediate value to engineering teams.
- Shape and support implementation of MLOps strategy, including model deployment, monitoring, versioning, and lifecycle management.
- Collaborate with Data Governance to ensure AI/ML frameworks align with data governance, security, and compliance requirements.
- Design and advocate for scalable AI/ML infrastructure patterns such as feature stores, curated/governed datasets, and streaming access for training and inference.
- Partner with Data Science, Data Engineering, and business stakeholders to assess AI/ML readiness gaps and build a roadmap to close them.
- Serve as a subject-matter expert and thought partner to the VP, Head of Data on emerging AI/ML technologies, practices, and industry trends.
- Document AI/ML standards, frameworks, and decisions to support consistent adoption as practices mature.
- Act as a senior technical resource for architecture guidance, design patterns, and best practices for AI readiness and ML Ops frameworks.
- Partner with Enterprise Architecture on establishing architectural blueprints for AI readiness.
- Other Duties as Assigned.
Requirements
- Data engineering fundamentals: deep expertise in data pipeline design, optimization, and distributed data processing (Spark, dbt, Airflow, Kafka, or equivalent).
- Data platforms: hands-on experience with Snowflake, Databricks, and/or Azure Synapse Analytics, with the ability to architect and optimize workloads.
- Cloud: strong knowledge of AWS, Azure, or GCP, plus modern data warehouse/lakehouse architectures.
- Python: strong Python skills and production-grade software engineering practices (testing, version control, code review) for shipping systems beyond notebooks.
- API and integration: experience building systems around models via orchestration, tool-calling, and retrieval systems.
- LLM experience: practical experience with LLM APIs (e.g., OpenAI) and open-weight models.
- Prompt engineering: prompt engineering and evaluation as an established discipline.
- Model fundamentals: understanding of context windows, tokenization, embeddings, and limitations such as hallucination, latency, and cost tradeoffs.
- Vector search: experience with vector databases (Pinecone, Weaviate, pgvector, etc.) and embedding models.
- Search techniques: chunking strategies, hybrid search, and reranking.
- Agent and orchestration: frameworks such as LangChain, LangGraph, LlamaIndex, or custom orchestration.
- Tool use / agents: function-calling design, multi-step reasoning chains, and agent memory/state management.
- Model adaptation: practical understanding of when to fine-tune vs. prompt vs. RAG.
- Parameter-efficient methods: familiarity with LoRA and related approaches.
- MLOps / LLMOps: experience with MLOps/LLMOps practices.
- Evaluation and observability: model evaluation frameworks, A/B testing for model outputs, and observability (tracing, logging model calls).
- Deployment patterns: latency/cost optimization, caching, streaming responses, and fallback handling.
- Versioning: versioning prompts and models, not only code.
- Safety and governance: safety, evaluation, governance awareness; bias/safety evaluation and appropriate handling of PII.
Technologies
- Spark, dbt, Airflow, Kafka
- Snowflake, Databricks, Azure Synapse Analytics
- AWS, Azure, GCP
- Python, OpenAI
- Pinecone, Weaviate, pgvector
- LangChain, LangGraph, LlamaIndex
- LoRA
- MLOps, LLMOps
- Data Vault 2.0, Ensemble data modeling techniques
Experience
- Minimum 7+ years in data engineering with experience on large-scale, mature data platforms.
- 3+ years developing ML or AI deliverables, including deployment to production.
- Bachelor's degree in a related field or demonstrated equivalent experience in a related field required.
- Working knowledge of agentic workflows for engineering and architecture.
- Demonstrated autonomous technical ownership of complex projects from design through delivery with minimal oversight.
- Experience shaping AI/ML enablement (defining frameworks, evaluating MLOps tooling, or building infrastructure for training and deployment).
- Experience partnering with Data Governance, Data Science, or Compliance to align technical practices with governance and regulatory requirements.
- Track record of proposing and driving innovative technical solutions rather than only executing predefined plans.
- Experience designing or implementing agentic workflows for data engineering preferred.
- Experience working with Property & Casualty insurance carriers preferred.
- Experience with Data Vault 2.0 or Ensemble data modeling techniques preferred.
Location
- Cincinnati, OH (hybrid)
Benefits
- Medical, dental, vision, and life insurances
- Short and long-term disability
- Company-match of 100% of a 6% contribution 401(k) plan
- Employee Assistance Plan
- Health Savings Account
- Flexible Spending Account
- Health Reimbursement Account
- Wellness program
- Opportunities for professional development and advancement
Similar Jobs
S