DataJobs.io
← Back to all jobs

Job Description

Palo Alto Networks is seeking a Principal Machine Learning Engineer to help design and deliver robust, next-generation cloud security solutions. Based in Santa Clara, CA (onsite), the role provides technical leadership across the full machine learning lifecycle, taking models from development and training through production deployment and real-time inference.

In this position, you will shape scalable architecture and delivery practices for backend and machine learning systems, while collaborating closely with Product, SRE, QA, and Support to align execution with business goals. You will also drive MLOps evaluation and continuous integration, delivery, and monitoring so security-focused ML capabilities run reliably at scale.

What you will do

  • Provide technical leadership for end-to-end solution delivery, partnering with cross-functional teams including Product, SRE, QA, and Support.
  • Drive scalable cloud security architecture through a mix of strategic planning and hands-on engineering.
  • Establish and promote best practices for model versioning, reproducibility, auditing, and compliance to support code quality and data privacy.
  • Architect and lead the end-to-end ML lifecycle, from initial development and training to production deployment and real-time inference.
  • Build and maintain automated, resilient systems for CI/CD and monitoring across backend and ML components.
  • Continuously evaluate and integrate cutting-edge MLOps tools and frameworks to improve scalability, reliability, and efficiency.
  • Design and implement next-generation cloud security solutions addressing complex backend infrastructure and ML model challenges.
  • Strategically manage and optimize ML infrastructure and pipelines to improve performance, support smooth production integration, and reduce operational costs.

Skills and qualifications

  • Strong background in machine learning and ML frameworks (e.g., TensorFlow, PyTorch).
  • Experience with Infrastructure-as-Code (IaC) tools such as Terraform or CloudFormation.
  • Bachelor’s degree in Computer Science, Engineering, or a related field (or equivalent practical experience).
  • 10+ years of software development experience with a focus on cloud-native and SaaS applications.
  • Proven experience designing and building large-scale, distributed systems on public cloud platforms: AWS, GCP, or Azure.
  • Strong proficiency in at least one modern programming language such as Python, Go, or Java.
  • Demonstrated experience across the full ML lifecycle, including model deployment and MLOps.

Technologies

  • TensorFlow, PyTorch
  • Terraform, CloudFormation
  • Python, Go, Java
  • AWS, GCP, Azure
  • CI/CD, MLOps
  • Docker, Kubernetes
  • Kafka, Flink

Team and work approach

  • Engineering teams are directly connected to the mission of preventing cyberattacks, with a focus on innovation in cybersecurity.
  • Collaboration is primarily in person, with most teams working from the office full time, with flexibility when needed.

Preferred qualifications

  • Master’s or PhD in Computer Science or a related technical field.
  • Experience in the cybersecurity domain or with network security products.
  • Expertise with containerization and orchestration, particularly Docker and Kubernetes.
  • Experience with real-time data processing and streaming technologies such as Kafka and Flink.
  • Contributions to open-source projects in the cloud-native or MLOps space.

Compensation: USD 163,200 - 264,000 per year.

Similar Jobs