Lead Data Engineer
Job Description
Navanta, LLC is building a governed data backbone that powers public Call Report data and secure bank-core ingestion as client environments come online. In this role, you will lead a lakehouse foundation that is clean, well-modeled, and reconcilable, enabling trustworthy data products for Navanta’s AI platform through close collaboration with AI/ML, security, and platform teams.
Responsibilities
- Design and evolve the lakehouse using Apache Iceberg (or similar) on object storage, including a catalog for table management and per-bank isolation, plus dbt models and a query engine.
- Build secure, least-privilege ingestion from bank systems using log-based CDC where permitted, with query-based and batch/SFTP fallbacks, and an in-bank collector pattern.
- Own data modeling for the semantic and metric layer, covering deposits, concentration, uninsured exposure, asset quality, and peer groups.
- Manage schema drift, data quality, and reconciliation, while making ingestion observable and recoverable.
- Partner with AI/ML on the structured-query path, and with Security on PII classification at landing, aligned with regulatory data-handling requirements.
- Document data lineage, transformation logic, and access controls to support audit and exam readiness.
- Define and enforce data contracts, quality thresholds, and alerting for pipeline failures.
Requirements
- 8–12+ years in data engineering with end-to-end ownership of ingestion through serving; 2+ years in a lead or senior role.
- Strong Python and expert SQL, with rigorous data modeling for analytics.
- Hands-on lakehouse experience using Iceberg/Delta/Hudi or equivalent and modern transformation tooling.
- Proven ability to build reliable pipelines from messy operational and transactional source systems.
- Comfort with CDC mechanics and the realities of pulling from databases you do not control.
- Bachelor’s degree in computer science, mathematics, information systems, or a related field, or equivalent hands-on experience.
- Experience in financial services or a regulated data environment is strongly preferred.
KPIs
- Data freshness and pipeline reliability with SLAs met for data ingestion and bank-core feeds.
- Data quality score across key metrics versus source reconciliation.
- Time to onboard a new bank’s environment, from kickoff to queryable lakehouse.
- PII classification coverage at landing and zero unauthorized data-access incidents.
- Semantic layer adoption: percentage of assistant queries resolved via governed metrics versus ad hoc SQL.
Technologies
Python, SQL, Apache Iceberg, Polaris, Nessie, Lakekeeper, dbt, Trino/Presto, DuckDB, Debezium, Kafka, Redpanda, Dagster, Airflow, S3, MinIO, SQL Server, PostgreSQL, pgvector, Delta, Hudi.
Work Structure & Expectations
This is a full-time role combining ongoing pipeline operations with initiative-based lakehouse build-out and new bank onboarding. You will collaborate closely with AI/ML, platform engineering, and Security, and participate in an on-call rotation focused on data pipeline reliability.
Location
Alpharetta, GA (onsite); up to 20% travel time may be required.
Physical Demands & Work Environment
Typical office environment. Duties may require sitting and using hands to finger, handle, or touch objects and controls; frequently talking or hearing; occasionally standing, walking, stooping, kneeling, crouching, or crawling; and occasionally lifting or moving up to 10 pounds (usually waist high, up to 50 feet away). Specific vision abilities include close vision and the ability to adjust focus.