This position is no longer accepting applications
Closed on August 11, 2026.
This role is filled — get an email when new Data Analysis roles open on DataJobs.io:
Senior Data Engineer
Get alerted when similar jobs are posted — set up a New Data Analysis jobs on DataJobs.io alert.
See other roles at firstPROPhiladelphia.
Job Description
FirstPROPhiladelphia is seeking a Senior Data Engineer to design and operate real-time streaming pipelines for connected medical devices in a hybrid contract-to-hire engagement in Philadelphia. The role centers on building scalable data flows with Azure Databricks and related Azure services, enabling ML and analytics workflows within a Medallion architecture. The successful candidate will collaborate with data scientists, software engineers, and business stakeholders to translate telemetry into reliable, actionable insights.
Engagement and compensation
- Location: Philadelphia, PA (hybrid)
- Job type: Contract (contract-to-hire)
- Hourly rate: USD 60 - 70
- Education and experience requirements: 6+ years of data engineering/analytics/warehousing with a Bachelor's degree in Computer Science, Mathematics, Engineering, or a related field; OR 3+ years with a Master's degree in a related technical field
Responsibilities
- Design and implement batch and streaming data pipelines using Azure and Databricks, focusing on real-time IoT telemetry within a Medallion architecture
- Develop ETL/ELT workflows to ingest, transform, and validate large volumes of structured and unstructured data
- Build and maintain data services, APIs, and microservices for application, analytics, and ML/AI teams
- Implement real-time streaming solutions via Azure Event Hubs, Azure Stream Analytics, and related Azure integration patterns with cost-effective throughput, partitioning, and downstream delivery to Databricks
- Optimize production Databricks pipelines using PySpark, Spark SQL, and Delta Lake, including Spark tuning for performance, reliability, and cost
- Troubleshoot and resolve complex pipeline issues across Databricks, Azure, and on‑premises systems with root-cause analysis and corrective actions
- Collaborate with data analysts, software engineers, ML engineers, and business stakeholders to translate requirements into technical designs and delivery priorities
- Apply data quality, validation, and privacy-first practices, delivering reliable pipelines through engineering standards, documentation, testing, and CI/CD
Requirements
- Bachelor's degree in Computer Science, Mathematics, Engineering, or a related field and 6+ years of professional experience in data engineering, analytics, or warehousing; OR Master's degree in a related technical field and 3+ years of professional experience in data engineering, analytics, or warehousing
- 5+ years designing, building, and operating big data and real-time streaming pipelines across cloud and on‑premises environments
- 5+ years applying DevOps and CI/CD practices to data and analytics workloads
- Production experience delivering data services, APIs, or microservices for downstream data consumption
- Strong, hands-on experience designing and building production-grade data pipelines on Databricks
- Demonstrated Spark optimization and tuning in Databricks, including performance analysis, partitioning strategies, caching, shuffle optimization, and cost‑aware pipeline design
- Strong experience with Azure cloud services for data engineering and streaming workloads
- Experience delivering data services, APIs, or microservices for data consumption
- Data quality, validation, and privacy‑aware handling for regulated or sensitive data
Technologies
- Databricks
- Spark, PySpark, Spark SQL
- Delta Lake
- Azure, including Azure Event Hubs and Azure Stream Analytics
- Azure Databricks
- Medallion architecture
Outcomes
- Onboard to the Azure Databricks environment and contribute to troubleshooting, stabilization, and optimization of existing batch and streaming pipelines
- Stand up Azure streaming ingestion for telemetry data and deliver production-ready pipelines integrated with Databricks to support API, ML, and downstream analytics within the Medallion framework
- Design and deliver data services or consumption patterns that enable business and ML teams to access near‑real‑time telemetry data reliably, securely, and at scale
Work hours and travel
- Willingness to assist in troubleshooting and analysis during off-hour production problems as needed
- The hybrid setup requires a minimum of two days per week in the downtown Philadelphia office