Data Engineer
Python
Amazon Emr
Amazon Web Services
Analytics
AWS
Aws Glue
Big Data
Bigdata
Cloud
Cloud Operations
Cloud Platforms
Data Analysis
Data Analytics
Data Engineer
Data Integration
Data Pipeline
Data Platform
Data Processing
Data Warehouse
Database
EMR
ETL
Firehose
Java Language
Node Js
Scala
Software Engineering
Spark
SQL
Job Description
Onsite Data Engineer role in Bellevue, WA, joining Amazon's PXT Central Science team to build scalable data pipelines, feature extraction, and production-ready analytics.
Responsibilities
- Data Pipeline Development: design and maintain scalable data pipelines using native AWS services (Glue, EMR, Lambda); implement monitoring and error handling for data workflows; optimize for performance, reliability, and cost efficiency
- Model Productionization and API Development: build and maintain APIs and data serving layers that productionize science models for downstream use; create batch and real-time inference pipelines
- Data Integration and Quality: develop scalable feature extraction and processing frameworks for diverse data types; implement robust data quality and validation checks; design flexible schemas to support evolving requirements
- Cross-team Collaboration: partner with economics, data science, and software engineering teams to translate analytical requirements into production-ready solutions; participate in technical design reviews and architecture discussions
- Analytics and Infrastructure: maintain layered data systems used by economists and scientists; build automated reporting solutions; work across multiple interconnected AWS accounts with security best practices
Requirements
- Proficiency in professional software engineering practices across the full software development lifecycle, including coding standards, architectures, code reviews, source control, CI/CD, testing, and operational excellence
- 3+ years of data engineering experience
- Experience with at least one modern programming language such as Python, Java, Scala, or NodeJS
- Experience with data modeling, data warehousing, and building ETL pipelines
- Experience with AWS technologies including Redshift, S3, Glue, EMR, Kinesis, Firehose, Lambda, and IAM roles and permissions
- Experience with non-relational databases or data stores (object storage, document or key-value stores, graph databases, column-family databases)
- Bachelor's degree or foreign equivalent in computer science, engineering, mathematics, or related field
Technologies
- Python
- Java
- Scala
- NodeJS
- AWS Glue
- AWS EMR
- AWS Lambda
- Amazon Redshift
- Amazon S3
- Amazon Kinesis
- Amazon Firehose
- IAM
- Hadoop
- Hive
- Spark
Benefits
- Sign-on payments
- Restricted stock units (RSUs)
- Health insurance (medical, dental, vision, prescription; Basic Life and AD&D; option for supplemental life plans; EAP; mental health support; medical advice line; FSAs; adoption and surrogacy reimbursement)
- 401(k) matching
- Paid time off
- Parental leave
About the Team
The Central Science Team within Amazon’s People Experience and Technology organization (PXTCS) uses economics, behavioral science, statistics, machine learning, and Generative AI to proactively identify mechanisms and process improvements that enhance Amazon and the lives, well-being, and value of work for Amazonians. This interdisciplinary group blends science, engineering, and UX to deliver solutions with measurable impact.