DataJobs.io
← Back to all jobs

Job Description

CNA Insurance is seeking a Senior Data AI Engineer to lead the design, construction, and operationalization of end-to-end AI and ML solutions that transition complex data assets to a modern cloud data lakehouse. This senior individual contributor will tackle both structured and unstructured data—ranging from documents and media to transactional records—building scalable pipelines, retrieval architectures, and knowledge graphs, with the potential to mentor peers as needed.

Location

Chicago, IL (onsite)

Salary

USD 72,000 - 141,000 per year

Experience

Minimum 2 years of experience

Education

Bachelor's Degree

Responsibilities

  • Design and implement AI solutions that accelerate migration from legacy systems to the cloud while ensuring scalability, reliability, and governance compliance.
  • Build scalable ingestion and transformation pipelines for both structured (SQL, relational) and unstructured (documents, images, audio, email, call transcripts) data, applying OCR, NLP preprocessing, and document chunking optimized for large language model consumption.
  • Deploy modern lakehouse patterns on Google Cloud Platform, including data governance, cataloging, and lineage tracking to support AI/ML workloads at scale.
  • Create vector databases, embedding pipelines, and knowledge graph structures that serve as the retrieval layer for RAG and related AI applications.
  • Productionize AI solutions and advanced analytics within a DevOps/MLOps framework, implementing automated testing, monitoring, and rollback capabilities.
  • Foster innovation by proposing new ideas and selecting appropriate tools and frameworks to translate business problems into analytics solutions.
  • Investigate and implement process improvements to close technology gaps and deepen knowledge of enabling technologies.

Requirements

  • Extensive experience building scalable ingestion and transformation pipelines for structured and unstructured data, with a track record of migrating workloads to modern cloud platforms.
  • Proficiency in parsing and normalizing diverse content types (PDFs, emails, images, call transcripts) using OCR and NLP preprocessing (tokenization, entity extraction, summarization) and document chunking for LLM use.
  • Hands-on experience designing and implementing vector databases (e.g., Vertex AI Vector Search, Pinecone, pgvector), embedding pipelines, and knowledge graphs that support RAG and semantic search.
  • Strong SQL and data analytics skills; experience building data marts and feature datasets for data science and ML applications.
  • Solid Python coding skills; practical experience with BigQuery, Claude Code, RAG architectures, LLMs, ADK, and prompt engineering techniques.
  • Expertise in building ML platforms and data pipelines at scale; familiarity with ML algorithms, deep learning, NLP, information retrieval, and data mining.
  • Experience with Google Cloud Platform services (Vertex AI, Dataflow, BigQuery, Cloud Run, Pub/Sub) and comfort with distributed computing frameworks (Apache Spark, Dataproc) for large-scale processing.
  • Ability to manage diverse data sources, including preprocessing, cleansing, and validating data integrity for data science and ML needs.
  • Demonstrated experience in machine learning, deep learning, NLP, information retrieval, or data mining, especially with unstructured or semi-structured data.
  • Hands-on experience with vector databases, embedding models (e.g., text-embedding-gecko, OpenAI Ada, Cohere), and end-to-end RAG pipeline design.
  • Experience with Agile methodologies is preferred.
  • Strong communication and collaboration skills to work effectively in a matrixed environment.
  • Preferred exposure to the insurance industry and its products and services.
  • Experience implementing big data processing technologies; Apache Spark experience is a plus.

Technologies

  • Python, SQL, Java
  • BigQuery, Claude Code, Vertex AI, Vertex AI Vector Search
  • Pinecone, pgvector, OpenAI Ada, Cohere
  • text-embedding-gecko, Dataproc, Apache Spark
  • Dataflow, Cloud Run, Pub/Sub
  • Google Cloud Platform, ADK

Reporting relationship

Typically reports to a Director or above

Similar Jobs