DataJobs.io
← Back to all jobs

Job Description

This role focuses on designing, building, and optimizing scalable PySpark data pipelines for Tata Consultancy Services in Irving, TX.

Responsibilities

  • Design, develop, and maintain ETL/ELT pipelines using PySpark
  • Craft efficient PySpark transformations with DataFrames and Spark SQL
  • Create modular, reusable Python-based data processing components
  • Uphold data quality, integrity, and performance across pipelines
  • Debug, tune performance, and optimize PySpark jobs
  • Collaborate with cross functional teams including Data Analysts, Architects, and DevOps
  • Contribute to CI/CD pipelines and deployment workflows for data applications
  • Monitor and troubleshoot production data workloads

Requirements

  • Hands-on PySpark experience with Spark SQL and DataFrame API
  • Advanced Python skills for data processing, performance tuning, and modular coding
  • Solid understanding of ETL design patterns and data pipeline architecture
  • Working knowledge of SQL for data transformation and analysis
  • Experience with data processing in distributed environments
  • 3 to 8 years of experience in Data Engineering or PySpark development
  • Proven hands-on project experience using PySpark and Python
  • Bachelor of Computer Science

Must Have Technical / Functional Skills

  • Strong hands-on experience in PySpark and Python for designing scalable data processing pipelines
  • Practical knowledge of distributed data processing and production-grade Python coding

Technologies

  • PySpark, Spark SQL, DataFrame API
  • Python, SQL
  • Amazon S3, AWS Glue, AWS EMR
  • Airflow, Snowflake, Git

Experience

  • 3–8 years of experience in Data Engineering or PySpark development
  • Proven hands-on project experience in PySpark and Python

Job Details

  • Salary: USD 100,000 – 120,000 per year
  • Location: Irving, TX (onsite)
  • Job Function: TECHNOLOGY
  • Role: Engineer
  • Job ID: 416300
  • Education: Bachelor of Computer Science

Similar Jobs