Lead Data Engineer
Job Description
Protective is hiring a Lead Data Engineer (Remote) to set the technical direction for a delivery pod building data products on Voyager, Protective’s Databricks lakehouse on Azure. This role leads end-to-end design and production delivery for datasets across Bronze/Raw, Silver/Prep, and Gold/Prod medallion layers, with hands-on ownership of production code, data contract standards, and consumer compatibility.
What you’ll be responsible for
- Design data products end to end, including what gets ingested, how data is cleansed and conformed, how it is modeled, and what the Gold layer looks like to querying teams.
- Own dimensional design, covering grain, natural and surrogate keys, Type 2 history, facts, bridges, and conformed dimensions shared across pod products.
- Define where logic belongs, separating cleansing in Silver, business logic in Gold, and concerns that should stay with consumers.
- Keep models aligned to the questions they answer, pushing back on designs that will not hold over time.
- Partner with ML engineering when Gold is a training or feature source, ensuring datasets are contracted, versioned, and reproducible as consumer-facing products.
- Own ODCS data contracts as real interfaces: named owners and consumers, enforceable quality rules, freshness and update expectations, and an explicit breaking-change policy.
- Make the compatibility call for contract changes and drive consumer notifications when changes are genuinely breaking.
- Represent pod contracts in cross-domain conversations, especially when one pod’s Gold layer becomes another team’s dependency.
- Set and enforce engineering standards across Python, SQL, dbt, testing, model structure, naming, and repository conventions, aligned with paved paths and Azure DevOps CI gates.
- Ensure quality rules are enforced via tests and asset checks (not just documentation), and that pipeline health is observable through freshness, volume, latency, and cost instrumentation with alerting against contracted SLAs and SLOs.
- Own operational posture for pod pipelines: failure diagnosis, data-issue triage, backfills, performance and cost tuning, on-call coverage and escalation, root-cause analysis, and runbooks others can execute.
- Operate within regulated carrier expectations, including change management through pull requests and pipelines, segregation of duties between authoring and deploying, least-privilege access, and CI/CD-generated audit evidence.
- Develop reusable frameworks, templates, and patterns to improve consistency and delivery speed.
- Work with the Product Owner and Scrum Master on decomposition and refinement, turning use cases into estimable, testable stories with identified target layer and repository.
- Hold Definition of Ready and Definition of Done standards, including merged and approved code, passing CI and coverage gates, and evidence that outcomes are real.
- Identify unknowns that require a spike instead of an estimate, and surface them during planning rather than mid-sprint.
- Grow engineers through design review, pairing, and code review, reducing single points of knowledge and onboarding new engineers to platform conventions and tooling.
- Work with the platform team to raise capability gaps as demand signals rather than building private workarounds.
- Partner with the DataOps/MLOps Lead on shared CI/CD, orchestration, and observability standards, strengthening the paved path and articulating what’s still missing.
- Collaborate with data architecture and governance on solution shape, Unity Catalog placement, and access requirements.
What you bring
- Bachelor’s degree in Computer Science, Information Systems, Engineering, or a related field (equivalent practical experience considered).
- 6+ years building and operating production data pipelines and consumer-facing data models, from source ingestion through published data products.
- Strong hands-on Python and SQL, with credibility to make design calls and willingness to still write and review code.
- Hands-on experience with Databricks or a comparable Spark-based lakehouse, including Delta Lake, MERGE, incremental processing, and performance tuning.
- Deep dimensional modeling experience: grain, keys, slowly changing dimensions, facts and dimensions, and conformed dimensions.
- Demonstrated technical leadership through setting standards and leading design and reviews, regardless of whether the title was “lead.”
- Experience owning data that other teams depend on, including handling breaking changes and production data incidents.
- Experience with orchestration (Dagster, Databricks Workflows, Airflow, or similar), Git-based collaborative development, code review, and CI/CD (Azure DevOps or comparable).
- Ability to set observability and SLA/SLO expectations for dependent data and run the incident and communication path when expectations are missed.
- Clear communication of trade-offs to engineers, product owners, and business stakeholders, including the ability to say no to designs that will not hold.
Tools you’ll work with
Databricks, Azure, Spark-based lakehouse, Delta Lake, MERGE, incremental processing, Python, SQL, dbt, Dagster, Databricks Workflows, Airflow, Git, CI/CD, Azure DevOps, Unity Catalog, ODCS, MLOps, MLflow, model registries, model serving, dlt (dltHub), Great Expectations, Monte Carlo.
Benefits
- Comprehensive health, dental and vision insurance
- Mental health benefits and an employee assistance program
- Paid time away benefits (e.g., paid time off, paid parental leave, short-term disability, and a cultural observance day)
- Contributions to healthcare accounts
- Pension plan
- 401(k) plan with Company matching
- ProHealth Rewards to improve wellbeing while earning cash rewards
Preferred qualifications
- Databricks certification (Data Engineer Professional or equivalent demonstrated depth)
- Unity Catalog at multi-team scale (catalogs, schemas, external locations, permissions, and lineage)
- dbt at scale on Databricks, plus Python-based modeling frameworks over Delta Lake
- Dagster and Dagster Cloud (assets, asset checks, and branch deployments)
- Practical experience with data contracts, ODCS, or data-mesh style data product ownership
- Declarative Python ingestion framework experience such as dlt (dltHub) or comparable
- Data quality and observability tooling such as Great Expectations, Monte Carlo, or similar
- Familiarity with MLOps practice (MLflow, model registries, model serving) sufficient to design data products that ML systems can depend on
- Azure and Azure DevOps
- Financial services, insurance, or another regulated industry experience, including data access, lineage, and audit expectations
- Experience introducing AI-assisted development into a team’s normal workflow in a disciplined way
If you require an accommodation to complete the application and recruitment process due to a disability, email [email protected].