Lead Data Engineer – Physical AI Platform
Job Description
Caterpillar is hiring a Lead Data Engineer to help build scalable, cloud-based data capabilities that power reliable, high-quality data for business and engineering teams. This full-time onsite role is based in Irving, TX (Dallas), with domestic relocation assistance available and visa sponsorship offered. The salary range for this position is $128,470 - $208,770 per year.
What you’ll build and improve
You will design, develop, and maintain data pipelines, microservices, and cloud data platforms that support both real-time and batch processing across enterprise systems. The work centers on cloud-native ingestion and streaming, data integration frameworks, operational monitoring, and data quality testing in an agile environment.
Responsibilities
- Collaborate with Principal Software Engineers and Data Architects to define solution architecture
- Lead solution design and optimization of scalable data pipelines and microservices using Python to support real-time and batch processing
- Drive development of cloud-native data ingestion and streaming solutions using AWS services such as Kinesis, S3, DynamoDB, EventBridge, and related technologies
- Own the design, implementation, and operational excellence of data integration frameworks and source data pipelines supporting CI Autonomy initiatives
- Partner with business, product, and engineering stakeholders to translate complex requirements into scalable data architectures, workflows, mappings, and system designs
- Establish automated testing, data quality controls, and validation frameworks to protect integrity, reliability, and compliance across distributed data ecosystems
- Lead operational monitoring, performance tuning, and root-cause analysis of production data platforms using observability tooling such as CloudWatch to maintain availability and service reliability
Requirements
- Bachelor’s degree in Computer Science, Computer Engineering, or a related field
- 8+ years of experience in data engineering or related disciplines with increasing responsibility
- Extensive experience on modern, large-scale, complex data platforms
- Strong foundation developing and deploying Python solutions in production environments
- Experience leading teams to build high-throughput, scalable data pipelines
- Hands-on experience with AWS data services including Kinesis, S3, DynamoDB, EventBridge at scale
- Strong skills in SQL, including data quality and validation practices
- Experience deploying software using CI/CD tools such as Azure DevOps, Jira, Jenkins, etc.
- Experience developing microservices that support real-time data ingestion
- Experience developing software applications using relational and NoSQL databases
- Ability to ensure data integrity across distributed and streaming systems
- Experience with monitoring, testing, and automation in large-scale data environments
Tools you’ll work with
Python, Java, AWS (Kinesis, S3, DynamoDB, EventBridge, CloudWatch), CI Autonomy, Azure DevOps, Jira, Jenkins, SQL, APIs, microservices, and real-time and batch data processing.
Benefits
- Medical, dental, and vision benefits
- Paid time off plan (Vacation, Holidays, Volunteer, etc.)
- 401(k) savings plans
- Health Savings Account (HSA)
- Flexible Spending Accounts (FSAs)
- Health Lifestyle Programs
- Employee Assistance Program
- Voluntary Benefits and Employee Discounts
- Career Development
- Incentive bonus
- Disability benefits
- Life Insurance
- Parental leave
- Adoption benefits
- Tuition Reimbursement
Additional details: Any offer of employment is conditioned upon the successful completion of a drug screen. This role requires full-time work at the Irving, TX office (Dallas).