DataJobs.io
← Back to all jobs

Job Description

GE Aerospace’s Commercial Engine Services BI team is hiring a Staff Data Engineer to design and support production data pipelines that convert raw operational inputs into trusted, analytics-ready datasets. In this hands-on role, you will build scalable transformations for both real-time and batch analytics, while enabling AI/ML-ready data through strong data quality, monitoring, and cross-team collaboration.

This position is based in Evendale, OH and supports fully remote arrangements across the United States. In-person attendance is required for New Hire Orientation on Day 1.

What you’ll do

  • Build production-grade data pipelines that transform raw operational data into analytics-ready datasets for applications, reports, and AI/ML models
  • Implement multi-layer transformation logic using medallion architecture, including efficient data cleaning, enrichment, and aggregation
  • Develop incremental loading patterns with schema evolution handling and data versioning to maintain reliability and backward compatibility
  • Schedule and orchestrate automated data refreshes for real-time reporting
  • Optimize pipeline performance for large datasets using partitioning, caching, indexing, and aggregation strategies aligned to dashboard performance needs
  • Troubleshoot pipeline failures, data quality problems, and performance bottlenecks, then implement fixes and preventive measures
  • Create automated data quality checks covering null validation, range checks, referential integrity, business rule enforcement, and schema drift detection
  • Implement validation frameworks to identify data issues early before they impact downstream dashboards or models
  • Monitor data quality metrics and alerts, investigate anomalies, communicate with stakeholders, and coordinate remediation with source system owners
  • Build data quality monitoring systems and reporting that provide visibility into pipeline health, data freshness, record counts, and quality trends
  • Document known data quality issues, workarounds, and resolution plans in a knowledge base for the BI team
  • Implement pipeline monitoring and alerting for failures, freshness, quality issues, compute costs, and execution times
  • Perform root cause analysis for data incidents, document findings, and implement preventive actions
  • Partner with BI analysts to translate dashboard and reporting requirements into transformation code
  • Collaborate with software engineers on training dataset preparation, feature pipelines, and data quality for forecasting and machine learning models
  • Work with the Data Platform Architect to follow architectural patterns, coding standards, and adopt platform capabilities such as data cataloging, monitoring frameworks, and CI/CD pipelines
  • Support BI with data questions, query optimization, and dataset troubleshooting, including guidance on efficient querying
  • Coordinate with the CDAIO team on source system integrations, data contracts, and ingestion layer requirements
  • Write documentation for pipelines including business logic, transformation steps, data lineage, dependencies, refresh schedules, and SLAs
  • Create dataset data dictionaries with column definitions, data types, expected values, refresh frequency, and usage examples
  • Maintain runbooks for common troubleshooting scenarios and document validation logic, data quality rules, and known issues
  • Follow software engineering best practices including Git-based version control, code review, automated testing, and CI/CD integration
  • Contribute reusable SQL and Python utilities, templates, and patterns to accelerate pipeline development

Required qualifications

  • Bachelor’s Degree in Computer Science, Information Systems, or related field from an accredited college/university (or a high school diploma/GED with a minimum of 4 years of relevant data engineering experience)
  • Minimum of 5 years of hands-on experience building production data pipelines and ETL/ELT processes
  • Expert-level SQL skills including complex joins, window functions, CTEs, aggregations, and query optimization for large datasets
  • Strong Python skills and familiarity with the PySpark DataFrame API, including transformations, actions, and optimization
  • Proven experience building ETL/ELT pipelines on cloud data platforms such as Databricks, Snowflake, AWS Glue, or similar
  • Understanding of dimensional modeling, slowly-changing dimensions, aggregate tables, and analytics-optimized data structures
  • Experience implementing automated data validation, schema checks, and data quality frameworks
  • Familiarity with cloud data services, including compute optimization and cost management
  • Experience with Git workflows, code review practices, and automated testing for data pipelines

Technologies

  • SQL, Python, PySpark
  • Databricks, Snowflake, AWS Glue
  • Git, CI/CD
  • Medallion architecture, ETL, ELT

Compensation and posting details

  • Base pay range: $112,000 to $150,000 per year
  • Eligible for an annual discretionary bonus based on a percentage of base salary; commission eligibility may apply based on the plan
  • Expected posting close date: Friday, October 2nd, 2026

Benefits

  • Healthcare benefits: medical, dental, vision, and prescription drug coverage
  • Health Coach access from GE Aerospace
  • Employee Assistance Program (24/7 confidential assessment, counseling, and referral services)
  • GE Aerospace Retirement Savings Plan including a 401(k) with company matching contributions and company retirement contributions
  • Access to Fidelity resources and planning consultants
  • Tuition assistance
  • Adoption assistance
  • Paid parental leave
  • Disability insurance
  • Life insurance
  • Paid time-off for vacation or illness

Location and eligibility

  • Onsite at the Evendale, OH campus is indicated; however, the role is eligible for fully remote arrangements across the United States
  • In-person requirement: New Hire Orientation on Day 1

Relocation

  • Relocation assistance: Not provided

Similar Jobs