Sr. Machine Learning Engineer
Artificial Intelligence
Automation
Big Data
Bigdata
Cloud
Cloud Infrastructure
Cloud Native
Cloud Platform
Cloud Platforms
Cloud Technology
Data Analysis
Data Analytics
Data Engineer
Data Integration
Data Pipeline
Data Platform
Data Processing
DevOps
DevSecOps
Engineer
Engineering
Infrastructure As Code
Kubernetes
Machine Learning Engineer
Ml Ops
Platform Engineering
Job Description
CrowdStrike is hiring a Sr. Machine Learning Engineer for the AIDR Engineering Team to support high-throughput, low-latency LLM inferencing and related post-training and data engineering.
Responsibilities
- Innovate with state-of-the-art machine learning technology to accelerate data science
- Provide pragmatic engineering support for fast-moving R&D teams
- Build and maintain high-quality, scalable solutions for customer-facing applications
- Conduct in-depth analysis to identify potential vulnerabilities or gaps
- Construct and maintain data pipelines, and contribute to training and implementation of custom models
- Collaborate across teams to brainstorm, define, and devise solutions
- Embrace ongoing learning and self-improvement
- Stay aligned with customer challenges and improve support to address them
- Maintain top-tier coding quality with best practices, rigorous testing, and thorough logging and metrics
- Work in a collaborative, agile team environment
- Mentor fellow engineers across a spectrum of technologies, while learning from them
- Explore improvements to product architecture, knowledge models, user experience, performance, and reliability
- Own work end to end with autonomy: develop, test, deploy, and monitor changes
- Thrive in an environment that highly values trust
Requirements
- Prior experience in data engineering and architecture supporting advanced data science use cases
- Deep understanding of LLM post-training methods and computational architectures
- Understanding of scalability and distributed systems (including sharding, partitioning, and concurrency)
- Ability to work as a team player
- Strong engineering best practices, including appropriate testing, peer code reviews, and resilient architecture
- Ability to thrive in a test-driven, collaborative, iterative development environment
- Track record of delivering time-bound, high-quality software with unit testing, code review, and regular checks-in for continuous integration
- Proven experience applying AI technologies to enhance decision-making, streamline workflows and processes, improve efficiency, and drive business outcomes
Technologies
- Python
- Docker
- Kubernetes
- AWS
- GCP
- MaaS
- Kafka
- Cassandra
- Spark
- ElasticSearch
- Terraform
- Chef
- Ansible
Tech Stack (Learning Capacity Required)
- High-level coding language such as JVM technologies or Python
- Docker
- Kubernetes
- AWS, GCP, or MaaS
- Kafka, Cassandra, and Spark
- ElasticSearch
- Terraform, Chef, or Ansible
- Experience scaling inference across GPUs or GPU clusters
Bonus Points
- Existing, demonstrable applied work in scalable architectures for LLM post-training or fine-tuning
- Prior experience in cybersecurity or intelligence fields
Benefits
- Market leader in compensation and equity awards
- Comprehensive physical and mental wellness programs
- Competitive vacation and holidays for recharge
- Paid parental and adoption leaves
- Professional development opportunities for all employees regardless of level or role
- Employee Networks, geographic neighborhood groups, and volunteer opportunities to build connections
- Vibrant office culture with world class amenities
- Great Place to Work Certified™ across the globe
- Health insurance
- 401k
- Paid time off
Location: Remote
Salary: USD 140,000 - 215,000 per year