Sr Hadoop+Spark(scala) Data Engineer
Job Description
Virtues is offering a senior Big Data engineering role focused on Hadoop, Spark, and Scala, with an onsite presence in Irving, TX. The position provides a competitive annual salary ranging from USD 100,696 to 110,268, and an opportunity to design, develop, and support scalable batch and real-time data pipelines across multiple data platforms.
Responsibilities
- Design, develop, implement, and maintain scalable, high-performance data ingestion and processing pipelines using Hadoop ecosystem technologies.
- Develop and manage data pipelines supporting batch processing, real-time/event-driven processing, and streaming data workflows.
- Read data from structured, semi-structured, and unstructured sources.
- Ingest batch data and real-time event streams, including Kafka events.
- Perform complex data validation, cleansing, enrichment, and transformation.
- Deliver processed data to target data stores, curated layers, publishing zones, and downstream endpoints.
- Develop and optimize Apache Spark applications using Scala for large-scale distributed processing.
- Design and implement Kafka-centric event processing and real-time data pipelines.
- Develop streaming data transformation logic using Apache Spark Streaming and/or Spark Structured Streaming.
- Build and maintain scalable batch processing solutions using Apache Spark.
- Develop data processing and analytical solutions using HiveQL, Pig Latin, HBase, and custom MapReduce programs.
- Develop data transformation and integration processes to move data from raw zones to curated and published data warehouse layers.
- Collaborate with data architects, application teams, business stakeholders, and platform teams to translate requirements into scalable technical solutions.
- Work extensively with Hadoop ecosystem technologies including HDFS, MapReduce, Hive, Pig, Sqoop, HBase, ZooKeeper, Oozie, Apache Spark, Scala, Flume/Flume NG, Kafka, and Hue.
- Apply strong knowledge of Hadoop architecture and core components such as NameNode, DataNode, HDFS, JobTracker, TaskTracker, and MapReduce programming/execution.
- Install, configure, integrate, and support Hadoop ecosystem components within Cloudera-based environments.
- Work with distributed storage and processing frameworks to ensure scalability, reliability, fault tolerance, and high performance.
- Monitor and optimize data pipeline performance, resource utilization, throughput, and processing efficiency.
Requirements
- Bachelor’s or Master’s degree in Computer Science, Information Technology, Engineering, Data Science, or a related technical discipline.
- 7+ years of hands-on experience in the Hadoop framework and the broader Hadoop ecosystem.
- 6+ years of hands-on experience developing data ingestion and integration solutions across multiple data platforms.
- 5+ years of strong hands-on experience with Apache Spark using Scala for distributed data processing.
- 5+ years of experience in data modeling, data transformation, detailed technical design, and data integration.
- Proven experience designing and developing large-scale batch and real-time data pipelines.
- Strong experience with HiveQL, Pig Latin, HBase, and custom MapReduce programming.
- Experience developing and managing Kafka-centric event-driven data pipelines.
- Solid understanding of batch processing, stream processing, and event-driven architectures.
- Hands-on experience with Spark Streaming and/or Spark Structured Streaming.
- Experience installing and configuring Cloudera Hadoop ecosystem components, including Hive, HBase, ZooKeeper, Oozie, Spark, Sqoop, Flume, Pig, and Hue.
- Strong understanding of Hadoop architecture, HDFS, distributed storage, and MapReduce concepts.
- Strong analytical, problem-solving, debugging, and performance-tuning skills.
- Excellent communication and collaboration skills.
Technologies
- Hadoop, HDFS, MapReduce, Hive, Pig, Sqoop, HBase, ZooKeeper, Oozie
- Apache Spark, Spark Streaming, Spark Structured Streaming
- Scala, Flume/Flume NG, Kafka, Hue
- HiveQL, Pig Latin
- BigQuery, Cloudera, MapR, Hortonworks
Key Competencies
- Strong expertise in distributed data processing and big data architecture
- Deep understanding of batch, real-time, streaming, and event-driven processing
- Proficiency in Scala and distributed data engineering frameworks
- Ability to design scalable, fault-tolerant, high-performance data solutions
- Analytical problem-solving, debugging, and root-cause analysis capabilities
- Ability to work independently and collaborate with cross-functional teams
- Ownership, attention to detail, and commitment to data quality and operational excellence
Desirable and Nice-to-have Skills
- End-to-end Hadoop administration and production support
- Hadoop infrastructure setup, installation, configuration, upgrades, patching, monitoring, troubleshooting, and maintenance
- Administration of Hadoop distributions such as Cloudera, MapR, Hortonworks
- Installing, configuring, and managing ecosystem components: Hive, Pig, HBase, ZooKeeper, Oozie, Spark, Sqoop, Flume, Hue
- Managing and monitoring HDFS, distributed file systems, and Hadoop clusters
- Managing, monitoring, scheduling, and troubleshooting MapReduce and distributed processing jobs
- Cluster capacity planning, resource management, health monitoring, and operational support
- Automation of operational activities via scripting (backup/restore, monitoring, health checks, maintenance, reporting)
- Experience with version control, change management, release management, incident management, problem management, and root-cause analysis
- Nice-to-have: Hadoop Platform Administration, GCP BigQuery