Senior Director, Data Engineer
Job Description
Capital One is hiring a Senior Director, Data Engineer to lead engineering teams focused on the future of data platforms and banking in the cloud. Based in McLean, VA (onsite), the role partners across product, data science, architects, and executive leaders to deliver scalable, data-driven solutions spanning identity analytics, access decisioning, and AI/ML security.
Role summary
This senior leadership position drives enterprise strategy and execution for identity analytics and intelligence, with emphasis on low-latency, real-time decisioning, ML and security platforms, and governance for AI/ML security and reliability. The work includes architecting event-driven systems, advancing MLOps practices, building intelligence for agentic and non-human identities, and leading enterprise readiness for emerging cyber risk posed by frontier AI models.
Responsibilities
- Own the enterprise strategy and multi-year roadmap for identity analytics, access decisioning, and AI/ML security platforms, aligning AI/ML investment with business, audit, and regulatory priorities and partnering with C-suite and CISO-org stakeholders on funding next-generation capabilities.
- Lead and grow a 50+ person organization across Data Engineering, ML, Generative AI, and Security Products, including managers and senior individual contributors.
- Hire, develop, and retain top technical talent; set engineering standards and expectations for well-managed delivery.
- Set technical direction hands-on, including reviewing architectures, model designs, and system trade-offs.
- Stay close to the code and data to make strong decisions under uncertainty and to build credibility with engineers.
- Architect low-latency, event-driven systems for real-time identity decisioning and threat detection using streaming telemetry, behavioral signals, and contextual graph data.
- Design, build, and maintain ML infrastructure and pipelines for end-to-end workflows including feature extraction, training, testing, guardrails, evaluation, deployment, and both real-time and batch inference with high performance, scalability, and reliability.
- Drive the evolution of MLOps through automated, metrics-backed deployment workflows, integration validation and testing systems, and scalable monitoring and observability for models in production.
- Build the intelligence layer for agentic and non-human identity using behavioral analytics, ML-based intent determination, and context graphs to secure AI agents and workload identities.
- Apply production-grade LLM and GenAI techniques for threat-intelligence summarization, automated incident triage, and adaptive access-policy recommendation, and introduce optimization to improve scalability, cost, latency, and throughput.
- Operate an automated governance function with pipelines from enterprise and IGA platforms, operationalizing NIST controls through continuous control validation and audit-ready evidence.
- Lead enterprise readiness for emerging cyber risk posed by frontier AI models and agentic attack patterns.
Required qualifications
- Bachelor’s Degree in Computer Science or a related quantitative field (Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering).
- At least 9 years of experience in data engineering.
- At least 7 years of people management experience.
- At least 7 years of programming experience with at least one of: Python, Java, or Scala.
- At least 5 years of experience driving technical delivery of roadmap features.
- At least 6 years of experience designing and developing data pipelines.
- At least 4 years of experience in data modeling and designing end-to-end data solutions using both relational and non-relational database systems.
Technologies
- Python, Java, Scala
- NIST
- AWS, Microsoft Azure, Google Cloud
- EMR, Spark, Glue, Databricks, Airflow, Dagster
- Monte Carlo, Splunk
- MongoDB, Cassandra, DynamoDB
- Snowflake, Redshift
- SQL, NoSQL
- Generative AI, LLMs, GenAI
Compensation
USD 314,800 - 359,300 per yearly for McLean, VA.
Additional location ranges are available for New York, NY, Plano, TX, and Richmond, VA. Candidates hired to work in other locations will be subject to the pay range associated with that location, and the actual annualized salary amount offered at the time of hire will be reflected in the candidate’s offer letter.
Benefits
- Comprehensive, competitive, and inclusive set of health, financial and other benefits that support your total well-being.
- Performance based incentive compensation, which may include cash bonus(es) and/or long term incentives (LTI).
Preferred qualifications
- Master’s Degree in Computer Science or a related field.
- 12+ years of experience in application development with demonstrated proficiency in Python, SQL, Scala, or Java.
- 8+ years of hands-on experience designing, deploying and operating data workloads in at least one public cloud environment (AWS, Microsoft Azure, or Google Cloud).
- 8+ years of experience building or supporting distributed data or compute workloads using tools such as EMR, Spark, Glue, or Databricks.
- 8+ years of experience designing, implementing, and operating real-time or streaming data pipelines.
- 6+ years of experience working on data observability (e.g., Monte Carlo, Splunk) or data orchestration tools (e.g., Airflow, Dagster).
- 8+ years of experience working with unstructured or semistructured data using NoSQL databases (e.g., MongoDB, Cassandra, DynamoDB).
- 8+ years of experience designing and supporting data warehousing solutions (e.g., Snowflake, Redshift).
- 6+ years of experience working in an Agile development environment.
- 6+ years of experience developing user-centric reusable data products.
- 5+ years of experience in Data Governance, Data Governance Platforms, Data Standardization, and Data Modeling.