The AWS Specialist Technology Team (STT) is building a centralized analytics data platform to bring product telemetry, usage metrics, and business outcomes together in one place. In this Data Engineering role, you will design, build, and operate scalable ETL/ELT pipelines and data models that support agent-driven analytics experiences and executive decision-making.
You will join a high-growth engineering organization applying generative AI and agentic technologies to transform how AWS field teams operate, with the chance to help shape foundational architecture from the start.
What you’ll do
- Design, build, and run scalable ETL/ELT pipelines to ingest product telemetry, usage events, and business outcome data from multiple heterogeneous sources across the STT product portfolio
- Architect and implement a centralized data platform using AWS-native services, including Redshift, S3, Glue, Lake Formation, Lambda, and Athena, to serve as a single source of truth for organizational analytics
- Build and maintain data models that connect product usage signals to business outcomes, including pipelines related to content effectiveness, field engagement, pipeline progression, and revenue impact
- Develop data infrastructure for AI/ML pipelines and agentic systems, including MCP tools and natural-language data access layers
- Implement data quality frameworks with automated monitoring, alerting, and validation to keep data accurate and reliable as the platform scales
- Create self-service data products with defined SLAs, documentation, and governance to reduce ad-hoc requests and enable stakeholders to answer their own questions
- Partner with Applied Scientists and SDE teams to deliver clean, well-modeled data for agent evaluation frameworks, retrieval quality measurement, and content effectiveness scoring
- Establish data contracts, lineage tracking, and catalog metadata to improve discoverability and trust across the organization
- Maintain operational excellence by owning on-call responsibilities, monitoring pipeline health, and resolving data freshness or quality issues before they affect consumers
- Help evolve from static dashboards toward agentic data systems by building foundational data layers that AI agents can query and reason over
Requirements
- 3+ years of data engineering experience
- 3+ years developing and operating large-scale data structures for business intelligence analytics using ETL/ELT processes
- 3+ years developing and operating large-scale data structures for business intelligence analytics using SQL
- 3+ years developing and operating large-scale data structures for business intelligence analytics using data modeling
- 3+ years of experience in the job offered or a related occupation
Technologies
- Redshift, S3, Glue, Lake Formation, Lambda, Athena
- ETL, ELT
- SQL
- EMR, Kinesis, FireHose
- IAM roles and permissions
- Natural-language data access layers, MCP tools
Preferred qualifications
- Experience with AWS technologies such as Redshift, S3, AWS Glue, EMR, Kinesis, FireHose, Lambda, and IAM roles and permissions
- Experience with non-relational databases and data stores, including object storage, document or key-value stores, graph databases, or column-family databases
Benefits
- Health insurance (medical, dental, vision, prescription, Basic Life & AD&D, with option for Supplemental life plans, EAP, Mental Health Support, Medical Advice Line, Flexible Spending Accounts, Adoption and Surrogacy Reimbursement coverage)
- 401(k) matching
- Paid time off
- Parental leave
- Sign-on payments
- Restricted stock units (RSUs)
Location: New York, NY (onsite) • Compensation: USD 132,100 - 196,600 per year
About the team: You will be one of the first two Data Engineers on a centralized analytics team built from the ground up, working alongside Business Intelligence Engineers, a Senior BD, an Applied Scientist, and a TPM. The pace of innovation is high, problems are ambiguous, and the impact spans thousands of field team members and the customers they serve, with your work shaping how the organization consumes and acts on data.