Fabric Data Engineer
Job Description
Hybrid role in Scottsdale, AZ supporting AI-driven data engineering at arrivia. In this position, you’ll take ownership of complex pipeline architecture inside the Microsoft Fabric ecosystem, modernizing data platforms and building governed, high-performance solutions. You’ll work with the latest AI technologies, including LLMs and Model Context Protocol (MCP) servers, while partnering with analysts, data scientists, and business stakeholders to turn complex needs into reliable data products.
Responsibilities
- Architect and optimize end-to-end pipelines using Microsoft Fabric Data Factory, Dataflows Gen2, and PySpark and Spark SQL notebooks for scale and performance.
- Lead Lakehouse and Warehouse design using the medallion pattern (Bronze, Silver, Gold), and establish the best practices the team follows.
- Support migration of on-premise relational data into OneLake and help retire legacy data-warehouse systems to accelerate the move to cloud.
- Build low-latency streaming pipelines with Fabric Eventstream from sources such as Azure Event Hubs, IoT Hub, and custom applications.
- Write highly optimized T-SQL, Spark SQL, and PySpark, owning incremental loads, refresh scheduling, and SLA monitoring.
- Design pipelines that support Retrieval-Augmented Generation, including chunking, embeddings, and vector search, using LLMs and MCP servers to improve team delivery.
- Champion governance with sensitivity labels, role-based access controls, and cataloging in Microsoft Purview.
- Drive CI/CD with Fabric deployment pipelines, Git branching strategies, and automated testing for data assets.
- Maintain Power BI semantic models where needed to keep enterprise reporting consistent and accurate.
- Coach Fabric Data Engineer I team members through code reviews, pair programming, and knowledge sharing, and help lead architectural reviews and continuous improvement.
- Design and optimize enterprise-scale data solutions, set technical standards, and mentor junior engineers.
- Partner with analysts, data scientists, and business stakeholders to translate complex requirements into well-governed, performant data platforms.
Requirements
- 3 to 5 years in data engineering, ETL and ELT development, or a related analytics engineering role.
- Strong SQL across T-SQL and Spark SQL, plus strong Python with PySpark.
- Working knowledge of Scala is a plus.
- Several years with Apache Spark and lakehouse data at scale, including performance tuning with PySpark and Spark SQL notebooks on large datasets.
- Solid grounding in lakehouse architecture, data warehousing, dimensional modeling, and data vault methodology.
- Hands-on experience with Microsoft Fabric or a comparable platform such as Azure Synapse or Databricks.
- Experience with real-time and streaming data at scale using event-driven architectures and tools like Azure Event Hubs or Kafka.
- Familiarity with vector search, embeddings, and Retrieval-Augmented Generation, including hands-on work with LLMs and AI-assisted development tools.
- Strong CI/CD habits across Fabric deployment pipelines, Git, and automated testing.
- Demonstrated experience mentoring junior engineers and leading technical initiatives.
- Microsoft Certified: Fabric Data Engineer Associate (DP-700) is highly preferred.
- Bachelor’s degree in a related field, or equivalent practical experience.
Benefits
- Unlimited PTO
- Exclusive employee travel rates
- Travel discounts through arrivia programs
- Medical, dental, and vision insurance
- 401(k) with company participation
About arrivia
arrivia powers travel loyalty and rewards programs for leading brands. Its teams help millions of travelers book memorable experiences while delivering innovative technology and travel solutions for partners. With a global workforce and a culture built on curiosity, ownership, authenticity, and collaboration, arrivia is creating the future of travel.