Lead Data Engineer – Physical AI Platform, Data Engineering
Job Description
Caterpillar is seeking a Lead Data Engineer to design, build, and maintain scalable data pipelines, microservices, and cloud-based data platforms for business and engineering teams. This onsite role in Chicago focuses on data architecture, performance, reliability, and continuous improvement using AWS and Python.
Key Responsibilities
- Collaborate with Principal Software Engineers and Data Architects to define solution architecture.
- Lead solution design and optimization of scalable data pipelines and microservices in Python for real-time and batch processing across enterprise platforms.
- Drive development of cloud-native data ingestion and streaming solutions using AWS services including Kinesis, S3, DynamoDB, EventBridge, and related technologies.
- Own the design, implementation, and operational excellence of data integration frameworks and source data pipelines supporting CI Autonomy initiatives.
- Partner with business, product, and engineering stakeholders to translate complex requirements into scalable data architectures, workflows, mappings, and system designs.
- Establish automated testing, data quality controls, and validation frameworks to ensure integrity, reliability, and compliance across distributed data ecosystems.
- Lead operational monitoring, performance tuning, and root-cause analysis of production data platforms using observability tools such as CloudWatch to maintain high availability and service reliability.
Required and Demonstrated Experience
- Ability to lead analysis and resolution of complex issues within distributed data platforms, designing scalable and resilient solutions.
- Effective communication across teams, including constructive feedback, active listening, and documentation to make data systems and processes easy to understand and support.
- Experience leading design and development of backend systems and data pipelines using Python, Java, and modern frameworks, including providing technical direction.
- Experience delivering data engineering solutions in an Agile environment by guiding work across the development lifecycle and ensuring quality, reliability, and business value.
- Capability to lead design and integration of APIs, data pipelines, streaming platforms, and databases for reliable data exchange across enterprise systems.
- Expertise leading design of scalable, event-driven data systems and architectures, guiding technical decisions to support reliability and maintainability aligned with business needs.
- Strong knowledge of AWS services and data engineering tools to define requirements, support testing and deployment, and troubleshoot issues across environments.
- Ability to define and implement testing strategies covering functional, performance, and data quality testing across the development lifecycle.
Technologies and Tools
- Python, Java
- AWS, Kinesis, S3, DynamoDB, EventBridge
- CloudWatch, CI/CD
- Azure DevOps, Jira, Jenkins
- SQL, Relational databases, NoSQL databases
- Microservices, APIs
- CI Autonomy, Helios Data Platform
Top Candidate Profile
- Bachelor’s degree in Computer Science, Computer Engineering, or related field.
- 8+ years of experience in data engineering or related disciplines with increasing responsibility.
- Extensive experience on modern, large scale, complex Caterpillar data platforms such as Helios Data Platform.
- Strong foundation developing and deploying Python solutions to production.
- Experience leading teams to build high-throughput, scalable data pipelines.
- Strong hands-on experience with AWS data services (Kinesis, S3, DynamoDB, EventBridge, etc.) at scale.
- Strong SQL experience, including data quality and validation practices.
- Experience deploying software using CI/CD tools such as Azure DevOps, Jira, Jenkins, etc.
- Experience developing microservices that support real-time data ingestion.
- Experience developing software applications using relational and noSQL databases.
- Ability to ensure data integrity across distributed and streaming systems.
- Experience with monitoring, testing, and automation in large-scale data environments.
Location and Work Arrangement
This full-time position requires working onsite five days a week at the Chicago, IL office. Domestic relocation assistance is available. Visa sponsorship is available for eligible applicants.
Salary Range
USD 128,470 - 208,770 per year.
Benefits
- Medical, dental, and vision benefits.
- Paid time off plan (Vacation, Holidays, Volunteer, etc.).
- 401(k) savings plans.
- Health Savings Account (HSA).
- Flexible Spending Accounts (FSAs).
- Health Lifestyle Programs.
- Employee Assistance Program.
- Voluntary Benefits and Employee Discounts.
- Career Development.
- Incentive bonus.
- Disability benefits.
- Life Insurance.
- Parental leave.
- Adoption benefits.
- Tuition Reimbursement.
- These benefits also apply to part-time employees.