DataJobs.io
← Back to all jobs

Job Description

Arrivia, Inc. is hiring a Fabric Data Engineer to help build and evolve enterprise data pipeline architectures in the Microsoft Fabric ecosystem. This hybrid role in Scottsdale, AZ supports cloud modernization efforts such as migrating relational data to OneLake, decommissioning legacy warehouse systems, and partnering with analysts and business stakeholders to deliver reliable, governed reporting.

What you’ll do

  • Architect and optimize end-to-end pipelines using Microsoft Fabric Data Factory, Dataflows Gen2, and PySpark and Spark SQL notebooks for scale and performance.
  • Lead Lakehouse and Warehouse design using the medallion approach (Bronze, Silver, Gold) and establish the best practices the team follows.
  • Support the move from on-premises relational databases into OneLake, retiring legacy data-warehouse systems as the platform transitions to the cloud.
  • Build low-latency streaming pipelines with Fabric Eventstream using sources such as Azure Event Hubs, IoT Hub, and custom applications.
  • Write and optimize T-SQL, Spark SQL, and PySpark, owning incremental loads, refresh scheduling, and SLA monitoring for dependable pipelines.
  • Design pipelines that support Retrieval-Augmented Generation through chunking, embeddings, and vector search, and use LLMs and Model Context Protocol (MCP) servers to improve team workflows.
  • Promote governance using sensitivity labels, role-based access controls, and data cataloging in Microsoft Purview.
  • Drive CI/CD with Fabric deployment pipelines, Git branching strategies, and automated testing for data assets.
  • Maintain Power BI semantic models where needed to keep enterprise reporting consistent and accurate.
  • Coach Fabric Data Engineer I team members through code reviews, pair programming, and knowledge sharing, and help lead architectural reviews and continuous improvement.

Qualifications

  • 3 to 5 years of data engineering, ETL and ELT development, or a related analytics engineering role.
  • Strong SQL experience across T-SQL and Spark SQL, plus strong Python with PySpark.
  • Working knowledge of Scala (a plus).
  • Several years of Apache Spark and lakehouse data work at scale, including performance tuning and building PySpark and Spark SQL notebooks against large datasets.
  • Solid grounding in lakehouse architecture, data warehousing, dimensional modeling, and data vault methodology.
  • Hands-on experience with Microsoft Fabric or a comparable platform such as Azure Synapse or Databricks.
  • Experience with real-time and streaming data at scale, including event-driven architectures and tools like Azure Event Hubs or Kafka.
  • Familiarity with vector search, embeddings, and Retrieval-Augmented Generation patterns, plus hands-on use of LLMs and AI-assisted development tools.
  • Strong CI/CD practices across Fabric deployment pipelines, Git, and automated testing.
  • Demonstrated experience mentoring junior engineers and leading technical initiatives.
  • Microsoft Certified: Fabric Data Engineer Associate (DP-700) is highly preferred.
  • A bachelor’s degree in a related field, or equivalent practical experience.

Technologies

  • Microsoft Fabric, Microsoft Fabric Data Factory, Dataflows Gen2, PySpark, Spark SQL, OneLake, Fabric Eventstream
  • Azure Event Hubs, IoT Hub, T-SQL
  • LLMs, Model Context Protocol servers, Retrieval-Augmented Generation, chunking, embeddings, vector search
  • Microsoft Purview, CI/CD, Fabric deployment pipelines, Git
  • Power BI, Power BI semantic models
  • Azure Synapse, Databricks, Apache Spark, Kafka
  • Sensitivity labels, role-based access controls

Benefits

  • Unlimited PTO
  • Exclusive employee travel rates
  • Travel discounts through arrivia programs
  • Medical, dental, and vision insurance
  • 401(k) with company participation

Similar Jobs