AI and Machine Learning Engineer
Job Description
Hewlett Packard Enterprise is hiring an AI and Machine Learning Engineer for a hybrid role based in Spring, TX, with an average of 2 days per week from an HPE office. This position supports work on high-performance computing and AI and Large Language Model performance on HPE GPU servers, including benchmarking, analysis, and technical documentation.
What you’ll get at HPE
- A comprehensive suite of benefits designed to support physical, financial, and emotional wellbeing, including support for your loved ones.
- Career programs to help you reach your goals, whether you want to become a knowledge expert or apply your skills in another division.
- Flexibility to manage work and personal needs in a semi-remote style.
- An unconditionally inclusive culture that celebrates individual uniqueness in the way teams work.
Responsibilities
- Install and configure complex IT infrastructure, including servers, storage, and network components.
- Develop software scripts and configuration to automate deployment and operations.
- Study, measure, and improve the performance of Large Language Models running on HPE GPU servers.
- Perform system-level analysis of server workloads across HPE platforms running DL and ML code, including accelerated hardware and high-speed networks such as InfiniBand.
- Run AI and HPC benchmarks, capture performance data, and interpret logs and traces to understand workload behavior.
- Build scripts and software tools to analyze AI workload performance data.
- Write white papers and guidance documents covering AI workload and model selection.
- Communicate progress and concerns to management in a timely manner; provide clear summaries to non-technical colleagues.
- Partner with software and hardware partners to optimize systems and resolve performance issues.
- Document issues found during testing and evaluation.
- Provide guidance to less-experienced staff members.
Requirements
- Master’s degree or PhD in Computer Science, Engineering, Information Technology or Systems, or a relevant field.
- 5+ years of experience.
- 5+ years in Machine Learning/Artificial Intelligence and 5+ years in HPC.
- Experience running NCCL and HPL alongside AI benchmarks.
- Experience with containers and distributed deep learning and neural networks, including transformers used in generative AI projects.
- Experience with high-performance computer servers, high-performance networking, and associated software including Slurm resource managers.
- Experience with Weka I/O, NFTS, and Lustre file systems.
- Programming experience in Python, C, and C++.
- Strong analytical and critical thinking skills.
- Scripting, process automation, and CI/CD are strongly desired.
- Must be a self-starter who can work with minimum supervision in a semi-remote setting.
Technologies you’ll work with
- Large Language Models, HPE GPU servers, and InfiniBand
- NCCL, HPL, and AI benchmarks
- Containers, distributed deep learning, neural networks, and transformers
- High Performance Computer Servers, High Performance Networking, Slurm
- Weka I/O, NFTS, and Lustre File Systems
- Python, C, C++, and CI/CD
Additional skills
- Artificial Intelligence Technologies
- Cross Domain Knowledge
- Data Engineering
- Data Science
- Design Thinking
- Development Fundamentals
- Full Stack Development
- IT Performance
- Machine Learning Operations
- Scalability Testing
- Security-First Mindset
Recruitment fraud alert
Scams may occur through false websites, emails, social media, or chat-based applications and often aim to obtain personal information or money. Hewlett Packard Enterprise (HPE), its direct and indirect subsidiaries and affiliated companies, and its authorized recruitment agencies/vendors will never charge a candidate a registration fee, hiring fee, or any other fee in connection with recruitment and hiring. HPE also never requests personal information such as back account details, Social Security numbers, or national IDs via social media or chat applications.