AI Data Engineer β Platform & Analytics
Analytics
Apache Airflow
Artificial Intelligence
Automation
Big Data
Bigdata
Cloud
Cloud Platform
Data
Data Analysis
Data Analytics
Data Architecture
Data Build Tool
Data Engineer
Data Integration
Data Lake
Data Lakehouse
Data Management
Data Modeling
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Database
Databases
Databricks
Engineer
ETL
Iceberg
Large Language Models
Machine Learning
Snowflake
Spark
SQL
Job Description
Responsibilities
- Design and sustain a scalable medallion data lakehouse architecture organized into Bronze, Silver, and Gold layers, implemented with Apache Iceberg tables.
- Develop and optimize Gold-layer data products with semantic richness for fast, cross-engine access across Snowflake, Databricks (Spark), Trino/Starburst, AWS Athena, and Presto.
- Design, build, and operate robust data workflows and DAGs using Apache Airflow.
- Develop, deploy, and monitor end-to-end ETL and ELT pipelines to ingest diverse semiconductor data streams into the data lake.
- Create and implement high-performance consumption data models that yield clean, transformed, production-ready datasets.
- Author and optimize advanced SQL queries for transformation, analysis, and performance benchmarking.
- Leverage generative AI coding assistants and automation tools to accelerate pipeline development, documentation, and testing.
- Implement data quality checks, schema evolution rules, and governance practices within Iceberg and Snowflake environments.
Requirements
- Solid experience in software development, data engineering, and data management.
- Hands-on experience scheduling and monitoring production-grade pipelines with Apache Airflow.
- Proven track record designing medallion architectures and extensive work with Apache Iceberg.
- Advanced Snowflake proficiency and practical experience building transformation models in DBT Core.
- Expert-level SQL and Python skills with thorough mastery of contemporary ETL/ELT patterns and design principles.
- Ability to produce high-quality code with meticulous attention to detail.
- Experience with concurrent programming and threading APIs.
- Experience with software development tools and processes such as debuggers, version control (GitHub), and profilers is a plus.
- Experience using AI tooling such as GitHub Copilot, Snowflake Cortex, LLM APIs, Claude Code to speed up coding and problem-solving.
- Experience delivering multiple enterprise-grade, production-level end-to-end data pipeline solutions from scratch.
Technologies
- Apache Iceberg
- Snowflake
- Databricks (Spark)
- Trino/Starburst
- AWS Athena
- Presto
- Apache Airflow
- DBT Core
- SQL
- Python
- GitHub Copilot
- Snowflake Cortex
- LLM APIs
- Claude Code
Academic Credentials
- BS Degree in Engineering or related field
Location
- Santa Clara, CA (onsite)
- Austin
- Seattle
- Secaucus