Principal Machine Learning Engineer, Accelerated Apache Spark
Job Description
NVIDIA seeks a Principal Machine Learning Engineer to guide ML engineering on the GPU-accelerated Apache Spark program in Santa Clara, focusing on ML-driven performance prediction and optimization for Spark workloads. The role also involves building AI-based deployment tools and providing mentorship across environments. This onsite position offers a salary range of USD 272,000 to 431,250 per year.
Role overview
In this capacity, you will lead the ML engineering efforts for a GPU-accelerated Spark initiative, developing ML-based methods to forecast performance and optimize Spark workloads. You will additionally create AI-enabled tools and provide technical mentorship to enable deployment across diverse environments.
Responsibilities
- Design and implement ML solutions for forecasting performance and optimizing GPU-accelerated enterprise Apache Spark workloads.
- Develop advanced algorithms and adaptive systems to continually enhance Spark performance on GPUs.
- Create AI-driven agents and tools to assist with diagnosing system issues and optimizing applications.
- Collaborate with partners and customers to deploy complex ML solutions across varied environments.
- Maintain deep domain expertise by staying current with the latest advances in ML systems and algorithms.
- Provide technical leadership and mentorship in data science and ML for a team of engineers.
Requirements
- Education: BS, MS, or PhD or equivalent experience in Machine Learning, Data Science, Computer Science, or a closely related field.
- 12+ years of professional experience designing, implementing, and productionizing high-quality ML/DL solutions.
- 5+ years of experience as a technical lead in ML model development.
- Proven hands-on experience (2+ years) with large-scale data processing platforms such as Apache Spark.
- Demonstrated ability to apply modern tooling and robust practices across all stages of building, deploying, and maintaining ML models.
- Excellent programming skills in Python and related data science libraries (NumPy, pandas, scikit-learn, SciPy, PyTorch, TensorFlow).
- Deep expertise in advanced ML methodologies, including LLMs/GenAI, reinforcement learning, and adaptive online ML systems.
- Strong capability in feature engineering, feature importance assessment, and developing boosted tree models (e.g., XGBoost).
Technologies
- Python
- NumPy
- pandas
- scikit-learn
- SciPy
- PyTorch
- TensorFlow
- Apache Spark
- XGBoost
- Scala
- Java
- C++
- CUDA
Benefits
- Equity and benefits
Ways to stand out from the crowd
- Understanding of the internal workings and architecture of Apache Spark
- Familiarity with NVIDIA GPUs and CUDA
- Experience coding in Scala, Java, and/or C++