Senior Data Engineer
Senior
Apache Airflow
AWS
Big Data
Bigdata
Cloud
Cloud Data Engineering
Cloud Platform
Cloud Platforms
Data
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Engineering
Data Governance
Data Integration
Data Lake
Data Lakehouse
Data Management
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Data Security
Database
Databases
Databricks
Databricks Pyspark
Databricks Workflows
Dataops
Delta Lake
ETL
Informatica
Information Technology (IT)
Programming Languages
Pyspark
Spark
SQL
Step Functions
Job Description
Steampunk is hiring a Senior Data Engineer (onsite) to design and deliver enterprise-grade data platforms, services, and pipelines using Databricks.
Responsibilities
- Lead and architect data migrations using Databricks, with emphasis on performance, reliability, and scalability
- Assess and understand existing ETL jobs, workflows, data marts, BI tools, and reports
- Respond to technical inquiries related to customization, integration, enterprise architecture, and feature/functionality of data products
- Support an Agile software development lifecycle
- Contribute to the growth of the AI & Data Exploitation Practice
- Inspect existing data pipelines, identify their purpose and functionality, and re-implement them efficiently in Databricks
Requirements
- Current ICE US Public Trust Clearance (Full Background Investigation)
- Ability to hold a position of public trust with the US government
- 5–7 years industry experience coding commercial software and a passion for solving complex problems
- 5–7 years direct experience in Data Engineering, including experience with:
- Big data tools: Databricks, Apache Spark, Delta Lake, etc.
- Relational SQL (preferably T-SQL; alternatively pgSQL, MySQL)
- Data pipeline and workflow management tools: Databricks Workflows, Airflow, Step Functions, etc.
- AWS and Azure cloud services (for example: Databricks on AWS, S3, EC2, RDS)
- Object-oriented/object function scripting languages: PySpark/Python, Java, C++, Scala, etc.
- Data lakehouse architecture and Delta Lake/Apache Iceberg
- Advanced working knowledge of SQL, including query authoring and optimization, plus working familiarity with a variety of databases
- Experience manipulating, processing, and extracting value from large, disconnected datasets
- Experience manipulating structured and unstructured data
- Experience architecting data systems (transactional and warehouses)
- Experience with SDLC, CI/CD, and operating in dev/test/prod environments
- Commitment to data governance
- Experience working in an Agile environment
- Experience supporting project teams of developers and data scientists building web-based interfaces, dashboards, reports, and analytics/machine learning models
- Plus: Experience with data cataloging tools such as Informatica EDC, Unity Catalog, Collibra, Alation, Purview, or DataZone
Technologies
- Databricks
- SQL, T-SQL, pgSQL, MySQL
- PySpark, Python
- Apache Spark
- Delta Lake
- Databricks Workflows, Airflow, Step Functions
- AWS, Azure
- S3, EC2, RDS
- Java, C++, Scala
- Data Lakehouse, Apache Iceberg
- Informatica EDC, Unity Catalog, Collibra, Alation, Purview, DataZone
- CI/CD, SDLC
Additional Information
- Location: McLean, VA (onsite)
- Salary: USD 140,000 - 180,000 per yearly
- Minimum experience: 5 years
Identity Verification
- During the application process, you are expected to be on camera during interviews and assessments
- Steampunk reserves the right to take your picture to verify your identity and prevent fraud
Company Focus
- Steampunk is a change agent in federal contracting, bringing new thinking to clients in Homeland and Federal Civil
- Human-Centered delivery methodology emphasizes shared accountability for solving mission challenges