DataJobs.io
← Back to all jobs

Job Description

The Senior Data Engineer role at Neptune Technology Group Inc. focuses on owning and modernizing the data pipeline architecture, migrating ETL workloads from Redshift stored procedures and legacy SSIS to scalable, maintainable pipelines built with AWS Glue and S3, and guiding the shift toward near real-time analytics using stream processing technologies such as Apache Flink and ClickHouse.

Responsibilities

  • Design and implement modern ETL and ELT pipelines leveraging AWS Glue, S3, and related services to replace legacy stored procedures and SSIS jobs.
  • Architect data transformation workflows that are testable, version-controlled, and observable.
  • Optimize and maintain the Redshift data warehouse, focusing on materialized views, query performance, and cost efficiency.
  • Advance the platform from batch ETL to near real-time stream processing by evaluating and deploying technologies such as Apache Flink, ClickHouse, Kafka, Kinesis, or equivalent.
  • Design pipelines capable of supporting both near real-time and batch workloads during the transition.
  • Collaborate with product and analytics teams to ensure data models support reporting, AI/ML initiatives, and customer-facing features.
  • Establish patterns and best practices for pipeline development adopted by the broader team.
  • Participate in production support and incident response for data infrastructure.

Requirements

  • Bachelor’s or Master’s degree in Computer Science, Engineering, or a related field, with a minimum of five years in data engineering roles.
  • Extensive experience with AWS Glue (PySpark/Python), S3, and Redshift.
  • Proven track record migrating ETL workloads from legacy tools such as SSIS or stored procedures to modern cloud-native pipelines.
  • Strong SQL skills with best practices and SQL linting, particularly in Redshift or other columnar/MPP databases.
  • Experience with or strong interest in stream processing frameworks (Flink, Spark Streaming, Kafka Streams, or similar).
  • Familiarity with data pipeline orchestration, monitoring, and error handling patterns.
  • Experience with infrastructure-as-code and CI/CD for data pipelines.

Technologies

  • Python
  • PySpark
  • SQL
  • AWS Glue
  • S3
  • Redshift
  • SSIS
  • Apache Flink
  • ClickHouse
  • Kafka
  • Kinesis
  • Apache Spark Streaming
  • Kafka Streams
  • Aurora MySQL
  • DynamoDB
  • Apache Druid
  • dbt
  • Airflow
  • Step Functions
  • MQTT

Location

Onsite in either Tallassee, Alabama or Duluth, Georgia. Travel to manufacturing or customer locations may be required up to 20% of the time when necessary.

Similar Jobs