Lead Data Engineer
Job Description
Capital Technology Group seeks a Lead Data Engineer to build scalable AWS-native data platforms for mission-critical analytics in federal environments.
Responsibilities
- Design, build, and maintain scalable data pipelines, ETL/ELT workflows, and data models using Python, Apache Spark (PySpark), SQL (PostgreSQL), and AWS Glue
- Develop and optimize AWS-native data platforms using AWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), Amazon S3, RDS, and CloudWatch
- Build high-performance ingestion, transformation, and orchestration workflows for structured and semi-structured data with Apache Iceberg and Parquet
- Design and optimize analytical platforms using Amazon Athena, Trino, Hive, OpenSearch, and enterprise data catalog technologies
- Build AI-enabled data solutions using Amazon Bedrock, RAG pipelines, and vector search technologies including Amazon S3 Vectors and OpenSearch vector indexes
- Develop cloud infrastructure via CloudFormation (Infrastructure as Code) and implement changes through GitHub with enterprise CI/CD pipelines
- Improve reliability, scalability, performance, and maintainability through monitoring, troubleshooting, automation, and continuous optimization
- Support mission-critical analytics and reporting in large-scale AWS-based federal data environments, implementing solutions aligned to FedRAMP and NIST SP 800-53 security controls
- Mentor junior engineers through technical guidance, architecture discussions, and code reviews, promoting engineering best practices
- Collaborate with cross-functional teams in an Agile environment to define requirements, deliver high-quality data solutions, and communicate technical concepts to technical and non-technical stakeholders
Requirements
- Bachelor’s degree in Computer Science, Engineering, or a related technical field
- 15+ years of professional experience in data engineering, data architecture, or related fields
- Strong hands-on experience with Apache Spark (PySpark) required, plus Python, SQL (PostgreSQL), and dbt for large-scale data engineering, ETL/ELT development, data transformation, and data modeling
- Experience with AWS Glue, Amazon EMR, Amazon MWAA (Apache Airflow), AWS Lambda, Amazon S3, Amazon RDS
- Experience developing scalable data pipelines, workflow orchestration, and data integration solutions across enterprise environments
- Experience with modern data lake technologies and formats such as Parquet and Iceberg
- Experience designing and optimizing solutions using relational and NoSQL databases
- Ability to build reliable, high-performance data platforms using performance tuning and enterprise-scale ETL/ELT architectures
- Strong analytical and problem-solving skills
- Experience working in Agile, iterative software development environments
- Ability to quickly learn and apply new technologies and domain knowledge
- Excellent written and verbal communication skills, including the ability to explain complex topics to diverse audiences
Technologies
- Python; Apache Spark (PySpark); SQL (PostgreSQL); AWS Glue; Amazon EMR; Amazon MWAA (Apache Airflow); Amazon S3; RDS; CloudWatch
- Apache Iceberg; Parquet; Amazon Athena; Trino; Hive; OpenSearch; enterprise data catalog technologies
- Amazon Bedrock; RAG pipelines; vector search technologies; Amazon S3 Vectors; OpenSearch vector indexes
- CloudFormation (Infrastructure as Code); GitHub; CI/CD pipelines
- FedRAMP; NIST SP 800-53; AWS Lambda; dbt; NoSQL databases
Benefits
- Remote Work (Hybrid roles will be specified in the job post)
- Competitive Compensation Package
- Medical, Dental, and Vision
- Life Insurance, Short/Long Term Disability
- Employee Assistance Program
- 401(k) with 4% matching
- Liberal PTO vacation policy
- Generous Annual Continuing Education
- Annual Wellness Budget
- Bonus Incentive Programs (Employee referrals and performance-based rewards)
Client Requirements
- Applicants MUST be US Citizens and be able to obtain Public Trust clearance
Location & Salary
- Location: Washington, DC (hybrid)
- Salary: USD 150,000 - 200,000 per year (estimated range; final offer may vary based on experience, skills, and other factors; range is not a guarantee and subject to change)