DataJobs.io
← Back to all jobs

Job Description

Build and optimize scalable AWS data solutions using PySpark and modern data lake patterns.

Responsibilities

  • Design, develop, and deploy scalable, cost-effective data solutions on AWS using services including S3, EC2, EMR, Glue, Athena, Lambda, Redshift, and Kinesis
  • Build and maintain robust ETL and ELT data pipelines with PySpark for data ingestion, transformation, and loading into data stores (including open table formats such as Iceberg)
  • Develop and optimize big data processing jobs using PySpark on AWS EMR or AWS Glue to handle large datasets efficiently with integration for Iceberg
  • Implement and manage data warehousing solutions with schema design, data modeling, and query optimization, with a focus on Hive and modern data lake table formats like Iceberg for historical data and analytics
  • Set up secure, robust cloud infrastructure components including VPCs, subnets, routing, and security groups to ensure connectivity and isolation
  • Design, deploy, and manage containerized data processing applications on Amazon Elastic Kubernetes Service (EKS)
  • Optimize performance and efficiency by tuning AWS resources and big data applications across Spark, Hive, and Iceberg
  • Apply data governance and security best practices in AWS, including access control and compliance via IAM policies, S3 bucket policies, and encryption
  • Set up monitoring, logging, and ingestion/processing visibility; troubleshoot and resolve issues for data pipelines and AWS infrastructure
  • Develop automation using Python and shell scripting for provisioning, deployment, and operational tasks
  • Collaborate with data scientists, analysts, and other engineering teams to understand data needs and deliver reliable data solutions

Requirements

  • At least one AWS certification, such as:
    • AWS Certified Solutions Architect - Associate
    • AWS Certified Data Analytics - Specialty
    • AWS Certified Developer - Associate
  • Hands-on experience with key AWS services for data processing and storage, including:
    • Amazon S3 (data lakes) and EC2
    • EMR, Glue, Athena, Lambda
    • VPC, subnets, routing, security groups
    • EKS
  • Strong PySpark proficiency for complex data transformations and analytics
  • Practical experience with Apache Iceberg for managing and querying data lakes
  • In-depth, practical experience with Apache Hive for data storage, querying, and schema management
  • Expert-level Python for data manipulation and AWS automation using Boto3
  • Proficient in shell scripting for automation and operational tasks
  • Strong SQL skills for querying and data manipulation
  • Solid understanding of ETL/ELT, data modeling, distributed computing, and data governance

Technologies

  • AWS, PySpark, Amazon S3, EC2, EMR, AWS Glue, AWS Athena, AWS Lambda, Amazon Redshift, Amazon Kinesis
  • Apache Iceberg, Apache Hive, VPC, subnets, routing, security groups
  • Amazon Elastic Kubernetes Service (EKS), Kubernetes
  • IAM policies, encryption
  • Python, shell scripting, SQL, Boto3, Apache Spark, SparkSQL
  • Data lakes, ETL/ELT, Iceberg, Spark Hive Iceberg
  • Apache Airflow, Git, Apache Kafka, Flink, Presto, AWS CodePipeline, GitHub Actions, GitLab CI

Benefits

  • Comprehensive Medical Plan covering Medical, Dental, Vision
  • Short Term and Long-Term Disability Coverage
  • 401(k) Plan with Company match
  • Life Insurance
  • Vacation Time, Sick Leave, Paid Holidays
  • Paid Paternity and Maternity Leave

Good to Have Skills

  • Workflow orchestration experience with tools like Apache Airflow
  • CI/CD experience with tools and practices such as AWS CodePipeline, GitHub Actions, and GitLab CI
  • Version control experience with Git
  • Exposure to other big data technologies such as Apache Kafka, Flink, or Presto
  • Containerization and orchestration experience with Kubernetes

Certifications

  • AWS Certified Solutions Architect - Associate
  • AWS Certified Solutions Architect - Professional
  • AWS Certified Data Analytics - Specialty
  • AWS Certified Developer - Associate

Skills

  • Mandatory Skills: Apache Spark, Big Data Hadoop Ecosystem, Python, Python for DATA, SparkSQL

Location: Tampa, FL (onsite)

Salary: USD 83,912 - 113,900 per yearly

Mandatory: Karat Interview

Similar Jobs