Senior Data Engineer
Senior
Analytics
Automation
Azure Event Hubs
Azure Stream Analytics
Big Data
Cloud Operations
Data
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Engineering
Data Integration
Data Lake
Data Management
Data Modeling
Data Pipeline
Data Platform
Data Processing
Data Warehouse
Database
Databases
Databricks
Digital Marketing
ETL
Informatica
Information Technology (IT)
Integration
SQL
Systems
Job Description
FirstPROPhiladelphia is seeking a Senior Data Engineer to design and operate real-time streaming pipelines for connected medical devices in a hybrid contract-to-hire engagement in Philadelphia. The role centers on building scalable data flows with Azure Databricks and related Azure services, enabling ML and analytics workflows within a Medallion architecture. The successful candidate will collaborate with data scientists, software engineers, and business stakeholders to translate telemetry into reliable, actionable insights.
Engagement and compensation
- Location: Philadelphia, PA (hybrid)
- Job type: Contract (contract-to-hire)
- Hourly rate: USD 60 - 70
- Education and experience requirements: 6+ years of data engineering/analytics/warehousing with a Bachelor's degree in Computer Science, Mathematics, Engineering, or a related field; OR 3+ years with a Master's degree in a related technical field
Responsibilities
- Design and implement batch and streaming data pipelines using Azure and Databricks, focusing on real-time IoT telemetry within a Medallion architecture
- Develop ETL/ELT workflows to ingest, transform, and validate large volumes of structured and unstructured data
- Build and maintain data services, APIs, and microservices for application, analytics, and ML/AI teams
- Implement real-time streaming solutions via Azure Event Hubs, Azure Stream Analytics, and related Azure integration patterns with cost-effective throughput, partitioning, and downstream delivery to Databricks
- Optimize production Databricks pipelines using PySpark, Spark SQL, and Delta Lake, including Spark tuning for performance, reliability, and cost
- Troubleshoot and resolve complex pipeline issues across Databricks, Azure, and on‑premises systems with root-cause analysis and corrective actions
- Collaborate with data analysts, software engineers, ML engineers, and business stakeholders to translate requirements into technical designs and delivery priorities
- Apply data quality, validation, and privacy-first practices, delivering reliable pipelines through engineering standards, documentation, testing, and CI/CD
Requirements
- Bachelor's degree in Computer Science, Mathematics, Engineering, or a related field and 6+ years of professional experience in data engineering, analytics, or warehousing; OR Master's degree in a related technical field and 3+ years of professional experience in data engineering, analytics, or warehousing
- 5+ years designing, building, and operating big data and real-time streaming pipelines across cloud and on‑premises environments
- 5+ years applying DevOps and CI/CD practices to data and analytics workloads
- Production experience delivering data services, APIs, or microservices for downstream data consumption
- Strong, hands-on experience designing and building production-grade data pipelines on Databricks
- Demonstrated Spark optimization and tuning in Databricks, including performance analysis, partitioning strategies, caching, shuffle optimization, and cost‑aware pipeline design
- Strong experience with Azure cloud services for data engineering and streaming workloads
- Experience delivering data services, APIs, or microservices for data consumption
- Data quality, validation, and privacy‑aware handling for regulated or sensitive data
Technologies
- Databricks
- Spark, PySpark, Spark SQL
- Delta Lake
- Azure, including Azure Event Hubs and Azure Stream Analytics
- Azure Databricks
- Medallion architecture
Outcomes
- Onboard to the Azure Databricks environment and contribute to troubleshooting, stabilization, and optimization of existing batch and streaming pipelines
- Stand up Azure streaming ingestion for telemetry data and deliver production-ready pipelines integrated with Databricks to support API, ML, and downstream analytics within the Medallion framework
- Design and deliver data services or consumption patterns that enable business and ML teams to access near‑real‑time telemetry data reliably, securely, and at scale
Work hours and travel
- Willingness to assist in troubleshooting and analysis during off-hour production problems as needed
- The hybrid setup requires a minimum of two days per week in the downtown Philadelphia office