Data Engineer 4
Apache Airflow
Artificial Intelligence
AWS
Azure
Big Data
Bigdata
Cassandra
Cloud
Cloud Data Engineering
Cloud Data Platform
Cloud Data Warehouse
Cloud Data Warehouse
Cloud Native
Cloud Operations
Cloud Platform
Cloud Platforms
Cloud Platforms Cloud Platforms
Data
Data Analysis
Data Analytics
Data Architecture
Data Engineer
Data Engineering
Data Integration
Data Lakehouse
Data Pipeline
Data Pipelines
Data Platform
Data Processing
Data Warehouse
Data Warehousing
Database
Databases
Databricks
Dataops
Dynamodb
Engineer
ETL
Google Cloud
Informatica
Information Technology (IT)
MongoDB
Programming Language
Programming Languages
Snowflake
SQL
Job Description
Capital One is hiring Data Engineers to build cloud-first data platforms, pipelines, and operational capabilities using full-stack and distributed data technologies.
Responsibilities
- Work with Agile teams to design, develop, test, implement, and support technical solutions
- Guide developers, data analysts, and data scientists with experience across machine learning, distributed microservices, lakehouse architecture, and full-stack systems
- Build with Python and Spark and integrate open-source relational and NoSQL databases
- Develop and operate cloud data warehousing solutions, including Databricks and Snowflake
- Stay current with data and engineering trends by experimenting with new technologies and participating in internal and external tech communities
- Mentor members of the data community
- Partner with product managers and software engineers to deliver robust cloud-first data solutions
- Independently design, build, and deliver cloud data solutions and applications with little or no supervision
- Architect and enforce reusable data engineering design patterns to improve code quality and maintainability
- Design and build data pipelines and platforms focused on scalability, resilience, and operational efficiency
- Implement data security standards, including encryption at rest/transit and fine-grained access control, to support data privacy compliance
- Communicate technical concepts and data outcomes clearly to internal and external stakeholders to drive alignment
Requirements
- Bachelor’s Degree or higher in Computer Science or a related quantitative field (Statistics, Economics, Operations Research, Analytics, Mathematics, Engineering)
- At least 4 years of experience in application development (internship experience does not apply)
- At least 2 years of experience in distributed data
- At least 2 years of experience with SQL
- At least 2 years of experience with one programming language: Python, Java, or Scala
- At least 2 years of experience in data pipeline design and development
- At least 1 year of experience in data modeling and designing end-to-end data solutions using both relational and non-relational databases
Technologies
- Python, Spark
- Databricks, Snowflake
- SQL
- NoSQL databases; open-source relational databases
- Machine learning
- Distributed microservices; lakehouse architecture
- Encryption at rest; encryption in transit; fine-grained access control
- AWS, Microsoft Azure, Google Cloud
- EMR, Glue, Airflow, Dagster
- Monte Carlo, Splunk
- MongoDB, Cassandra, DynamoDB, Redshift
Benefits
- Comprehensive, competitive, and inclusive health, financial, and other benefits supporting total well-being
Preferred Qualifications
- 7+ years of application development with demonstrated proficiency in Python, SQL, Scala, or Java
- 4+ years of hands-on experience designing, deploying, and operating data workloads in at least one public cloud (AWS, Microsoft Azure, or Google Cloud)
- 4+ years of experience building or supporting distributed data or compute workloads using EMR, Spark, Glue, or Databricks
- 4+ years of experience designing, implementing, and operating real-time or streaming data pipelines
- 2+ years of experience in data observability (Monte Carlo, Splunk) or data orchestration (Airflow, Dagster)
- 4+ years of experience with unstructured or semistructured data using NoSQL databases (MongoDB, Cassandra, DynamoDB)
- 4+ years of experience designing and supporting data warehousing solutions (Snowflake, Redshift)
- 2+ years of experience working in an Agile development environment
- 2+ years of experience developing user-centric reusable data products
Compensation / Incentives
- Plano, TX: $179,400 - $204,700 per year
- Eligible for performance-based incentive compensation, which may include cash bonus(es) and/or long-term incentives (LTI)
Employment Authorization / Immigration Support
- Capital One will not sponsor a new applicant for employment authorization and will not offer immigration-related support for this position
Application / Other Statements
- Expected minimum application window: 5 business days
- No agencies
- Capital One is an equal opportunity employer (EOE, including disability/veter), committed to non-discrimination under applicable federal, state, and local laws
- Capital One promotes a drug-free workplace
- Qualified applicants with criminal history will be considered consistent with applicable laws regarding criminal background inquiries
- Accommodations: contact Capital One Recruiting at 1-800-304-9102 or [email protected]
- Technical support or recruiting process questions: [email protected]
- Capital One does not provide, endorse, or guarantee third-party products, services, educational tools, or other information available through this site
- Entity note: positions posted in Canada, United Kingdom, and Philippines correspond to their respective Capital One entities