Senior Data AI Engineer
Senior
Artificial Intelligence
Big Data
Bigdata
Cloud
Cloud Native
Cloud Operations
Cloud Platforms
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Governance
Data Integration
Data Lake
Data Pipeline
Data Platform
Data Processing
Data Security
Data Warehouse
Database
Databases
ETL
Google Cloud
Google Cloud Platform
Machine Learning
Machine Learning Engineer
Programming Language
Programming Languages
Spark
SQL
Vertex Ai
Job Description
CNA Insurance is seeking a Senior Data AI Engineer to lead the design, construction, and operationalization of end-to-end AI and ML solutions that transition complex data assets to a modern cloud data lakehouse. This senior individual contributor will tackle both structured and unstructured data—ranging from documents and media to transactional records—building scalable pipelines, retrieval architectures, and knowledge graphs, with the potential to mentor peers as needed.
Location
Chicago, IL (onsite)
Salary
USD 72,000 - 141,000 per year
Experience
Minimum 2 years of experience
Education
Bachelor's Degree
Responsibilities
- Design and implement AI solutions that accelerate migration from legacy systems to the cloud while ensuring scalability, reliability, and governance compliance.
- Build scalable ingestion and transformation pipelines for both structured (SQL, relational) and unstructured (documents, images, audio, email, call transcripts) data, applying OCR, NLP preprocessing, and document chunking optimized for large language model consumption.
- Deploy modern lakehouse patterns on Google Cloud Platform, including data governance, cataloging, and lineage tracking to support AI/ML workloads at scale.
- Create vector databases, embedding pipelines, and knowledge graph structures that serve as the retrieval layer for RAG and related AI applications.
- Productionize AI solutions and advanced analytics within a DevOps/MLOps framework, implementing automated testing, monitoring, and rollback capabilities.
- Foster innovation by proposing new ideas and selecting appropriate tools and frameworks to translate business problems into analytics solutions.
- Investigate and implement process improvements to close technology gaps and deepen knowledge of enabling technologies.
Requirements
- Extensive experience building scalable ingestion and transformation pipelines for structured and unstructured data, with a track record of migrating workloads to modern cloud platforms.
- Proficiency in parsing and normalizing diverse content types (PDFs, emails, images, call transcripts) using OCR and NLP preprocessing (tokenization, entity extraction, summarization) and document chunking for LLM use.
- Hands-on experience designing and implementing vector databases (e.g., Vertex AI Vector Search, Pinecone, pgvector), embedding pipelines, and knowledge graphs that support RAG and semantic search.
- Strong SQL and data analytics skills; experience building data marts and feature datasets for data science and ML applications.
- Solid Python coding skills; practical experience with BigQuery, Claude Code, RAG architectures, LLMs, ADK, and prompt engineering techniques.
- Expertise in building ML platforms and data pipelines at scale; familiarity with ML algorithms, deep learning, NLP, information retrieval, and data mining.
- Experience with Google Cloud Platform services (Vertex AI, Dataflow, BigQuery, Cloud Run, Pub/Sub) and comfort with distributed computing frameworks (Apache Spark, Dataproc) for large-scale processing.
- Ability to manage diverse data sources, including preprocessing, cleansing, and validating data integrity for data science and ML needs.
- Demonstrated experience in machine learning, deep learning, NLP, information retrieval, or data mining, especially with unstructured or semi-structured data.
- Hands-on experience with vector databases, embedding models (e.g., text-embedding-gecko, OpenAI Ada, Cohere), and end-to-end RAG pipeline design.
- Experience with Agile methodologies is preferred.
- Strong communication and collaboration skills to work effectively in a matrixed environment.
- Preferred exposure to the insurance industry and its products and services.
- Experience implementing big data processing technologies; Apache Spark experience is a plus.
Technologies
- Python, SQL, Java
- BigQuery, Claude Code, Vertex AI, Vertex AI Vector Search
- Pinecone, pgvector, OpenAI Ada, Cohere
- text-embedding-gecko, Dataproc, Apache Spark
- Dataflow, Cloud Run, Pub/Sub
- Google Cloud Platform, ADK
Reporting relationship
Typically reports to a Director or above