Principal AI / Machine Learning Data Engineer
Job Description
The Principal AI / Machine Learning Data Engineer role at UnitedHealth Group focuses on designing and building end-to-end AI pipelines for large-scale unstructured data to enable advanced analytics and Generative AI. Located in Eden Prairie, MN with remote work options, this position sits at the crossroads of data engineering and AI, delivering production-grade data platforms and scalable AI capabilities.
Responsibilities
- Design, implement, and sustain scalable data pipelines and platforms that support analytics, machine learning, and AI initiatives
- Develop and optimize ingestion frameworks for both structured and unstructured data, including streaming and event-driven sources
- Collaborate with cross-functional partners to understand evolving data and AI needs and define long-term technical roadmaps
- Support and enable ML and AI workflows, including feature engineering, data preparation, and model deployment assistance
- Lead strategic programs around Generative AI, data quality, observability, lineage, and governance
- Create and maintain frameworks that enable rapid experimentation and deployment of AI/ML solutions
- Promote and evolve best practices in data modeling, orchestration, testing, and monitoring
- Identify opportunities to improve platform scalability, performance, and cost efficiency
- Partner with product, analytics, and infrastructure teams to deliver impactful data and AI solutions
- Develop reusable parsing, enrichment, analytics, and service libraries to accelerate delivery across teams
- Operate effectively under time-sensitive conditions while ensuring thoroughness and accuracy
- Maintain high ethical standards, objectivity, and confidentiality in all work
- Build and operate production data platforms and pipelines across batch and streaming workloads
- Hands-on engineering in Python and SQL, with experience in JVM languages (Java/Scala) within Spark ecosystems
- Work with distributed processing and lakehouse/warehouse patterns (for example Spark/PySpark, Databricks, Snowflake)
- Develop pipelines for OCR, document parsing, and text extraction from image-based or scanned data
- Enable production Generative AI solutions, including retrieval patterns and evaluation/monitoring practices
- Adopt knowledge-centric data approaches (metadata-driven systems, entity resolution, graph concepts) to improve discoverability
- Maintain a data quality, observability, and monitoring mindset with profiling, validation, alerting, and reliability improvements
- Orchestrate, implement CI/CD, containerization, and infrastructure-as-code with tools such as Airflow, GitHub Actions, Docker, Terraform, and Kubernetes
- Work in cloud environments (AWS, Azure, or GCP), ensuring secure handling of sensitive data (PII/PHI) and collaboration with compliance teams
- Provide leadership through influence, mentor engineers, and translate ambiguous problems into scalable technical roadmaps
Requirements
- Bachelor's degree or equivalent experience
- 5+ years designing, building, and operating scalable data pipelines and platforms (batch and streaming)
- 2+ years deploying Generative AI solutions to production (for example RAG, LLM-powered pipelines, semantic search)
- Strong hands-on development experience in Python and SQL, with Spark/PySpark and Databricks or similar distributed platforms
- Experience building ingestion and processing frameworks for unstructured data, including OCR, documents, and images, with parsing and enrichment
- Proficiency with cloud platforms (AWS, Azure, GCP), DevOps/CI/CD, and infrastructure-as-code, including secure handling of sensitive data
- Proven ability to design scalable solutions, implement data quality and observability practices, and collaborate with diverse stakeholders
Technologies
- Python
- SQL
- Spark
- PySpark
- Java
- Scala
- Databricks
- Snowflake
- AWS
- Azure
- Google Cloud Platform (GCP)
- Kafka
- Kinesis
- Event Hubs
- Airflow
- GitHub Actions
- Docker
- Terraform
- Kubernetes
- Delta Lake
- Plotly
- Seaborn
- Chartjs
- Great Expectations
- Deequ
Benefits
- Comprehensive benefits package
- Incentive and recognition programs
- Equity stock purchase opportunities
- 401k contribution
Preferred Qualifications
- Experience with cloud platforms such as AWS, Azure, or Google Cloud, including managed data services
- Experience with streaming and event-driven architectures (Kafka, Kinesis, Event Hubs)
- Experience with data quality and validation frameworks and data observability tooling
- Experience enabling MLOps practices, including feature stores, model registries, experiment tracking, and deployment automation
- Experience with lakehouse architectures, Delta Lake, and Spark performance tuning
- Experience with data visualization libraries such as Plotly, Seaborn, and Chartjs
- Experience with machine learning and predictive analytics
- Familiarity with security and privacy concepts for data platforms and collaboration with compliance partners
- Strong hands-on Python and SQL skills; familiarity with Java/Scala in Spark ecosystems
Application Deadline
The posting will remain active for at least two business days or until a sufficient candidate pool is reached. The listing may be withdrawn earlier if applicant volume is high.