DataJobs.io
← Back to all jobs

Job Description

Navanta, LLC is building a governed data backbone that powers public Call Report data and secure bank-core ingestion as client environments come online. In this role, you will lead a lakehouse foundation that is clean, well-modeled, and reconcilable, enabling trustworthy data products for Navanta’s AI platform through close collaboration with AI/ML, security, and platform teams.

Responsibilities

  • Design and evolve the lakehouse using Apache Iceberg (or similar) on object storage, including a catalog for table management and per-bank isolation, plus dbt models and a query engine.
  • Build secure, least-privilege ingestion from bank systems using log-based CDC where permitted, with query-based and batch/SFTP fallbacks, and an in-bank collector pattern.
  • Own data modeling for the semantic and metric layer, covering deposits, concentration, uninsured exposure, asset quality, and peer groups.
  • Manage schema drift, data quality, and reconciliation, while making ingestion observable and recoverable.
  • Partner with AI/ML on the structured-query path, and with Security on PII classification at landing, aligned with regulatory data-handling requirements.
  • Document data lineage, transformation logic, and access controls to support audit and exam readiness.
  • Define and enforce data contracts, quality thresholds, and alerting for pipeline failures.

Requirements

  • 8–12+ years in data engineering with end-to-end ownership of ingestion through serving; 2+ years in a lead or senior role.
  • Strong Python and expert SQL, with rigorous data modeling for analytics.
  • Hands-on lakehouse experience using Iceberg/Delta/Hudi or equivalent and modern transformation tooling.
  • Proven ability to build reliable pipelines from messy operational and transactional source systems.
  • Comfort with CDC mechanics and the realities of pulling from databases you do not control.
  • Bachelor’s degree in computer science, mathematics, information systems, or a related field, or equivalent hands-on experience.
  • Experience in financial services or a regulated data environment is strongly preferred.

KPIs

  • Data freshness and pipeline reliability with SLAs met for data ingestion and bank-core feeds.
  • Data quality score across key metrics versus source reconciliation.
  • Time to onboard a new bank’s environment, from kickoff to queryable lakehouse.
  • PII classification coverage at landing and zero unauthorized data-access incidents.
  • Semantic layer adoption: percentage of assistant queries resolved via governed metrics versus ad hoc SQL.

Technologies

Python, SQL, Apache Iceberg, Polaris, Nessie, Lakekeeper, dbt, Trino/Presto, DuckDB, Debezium, Kafka, Redpanda, Dagster, Airflow, S3, MinIO, SQL Server, PostgreSQL, pgvector, Delta, Hudi.

Work Structure & Expectations

This is a full-time role combining ongoing pipeline operations with initiative-based lakehouse build-out and new bank onboarding. You will collaborate closely with AI/ML, platform engineering, and Security, and participate in an on-call rotation focused on data pipeline reliability.

Location

Alpharetta, GA (onsite); up to 20% travel time may be required.

Physical Demands & Work Environment

Typical office environment. Duties may require sitting and using hands to finger, handle, or touch objects and controls; frequently talking or hearing; occasionally standing, walking, stooping, kneeling, crouching, or crawling; and occasionally lifting or moving up to 10 pounds (usually waist high, up to 50 feet away). Specific vision abilities include close vision and the ability to adjust focus.

Similar Jobs