Staff Machine Learning Engineer
Job description
About the role
Zscaler is seeking a technical leader to join the Digital Experience department. You will contribute to the Zero Trust Exchange platform, focusing on core intelligence and data to support over 15 million users. In this capacity, you will own the design and execution of machine learning initiatives that directly influence the reliability and performance of the platform. You will act as a hands-on architect, translating complex operational challenges into scalable intelligent solutions. The role requires you to manage an agentic troubleshooting framework by creating workflows, playbooks, and processes that standardize problem resolution. You will integrate GenAI advancements into production, including LLM fine-tuning, inference optimization, and robust data processing strategies. Success in this position will depend on your ability to utilize cloud platforms and data lakes for feature generation and deep exploratory analysis. You will handle high-volume data through real-time processing and aggregation pipelines to ensure timely insights. Furthermore, you will design and operate scalable microservices, data pipelines, orchestration systems, and caching layers that form the backbone of the service.
Key facts
What you'll do
- Implement advanced monitoring strategies to establish baselines and detect deviations within automated troubleshooting workflows.
- Evaluate and pilot emerging GenAI technologies, integrating them into production environments through optimized inference and data handling.
- Architect resilient data platforms on cloud infrastructure, leveraging data lakes to drive sophisticated feature engineering and exploration.
- Construct high-throughput streaming pipelines capable of processing massive data volumes with low-latency aggregation.
- Orchestrate complex microservice ecosystems, ensuring reliable deployment, scalability, and efficient caching strategies.
- Develop and maintain containerized applications using Docker and Kubernetes to ensure portability and operational consistency.
- Engineer data pipelines that transform raw telemetry into structured formats suitable for machine learning and real-time analytics.
- Collaborate with cross-functional partners to define clear requirements and translate business needs into technical specifications.
- Conduct rigorous performance testing and optimization of LLM and ML models to meet strict latency and accuracy targets.
- Document system architectures and operational procedures to ensure maintainability and knowledge transfer across the team.
Requirements
- Hold a BS in Computer Science with 6+ years of professional experience in the field.
- Alternatively, possess an MS or PhD with 5+ years of experience specifically in AI/ML and distributed systems.
- Demonstrate exceptional proficiency in core programming concepts, algorithms, and data structures.
- Apply deep knowledge of machine learning principles throughout the entire model lifecycle.
- Exhibit hands-on experience with feature generation, prompt engineering, deployment, monitoring, and ongoing optimization.
- Show expertise in designing and building distributed microservices using Python, Go, or Java.
- Have practical experience working with containerization and orchestration tools such as Docker and Kubernetes.
- Possess a strong understanding of distributed systems architecture and the challenges inherent in scaling such systems.
- Bring a proven track record of developing data pipelines that are robust, scalable, and efficient.
Nice to have
- Demonstrated experience scaling autonomous AI agents and orchestration workflows to automate incident response or cloud infrastructure troubleshooting.
- Solid background in fine-tuning and deploying proprietary SLMs or LLMs with a focus on cost efficiency, safety, latency optimization, and rigorous evaluation metrics.
- Experience building production-ready AI systems that incorporate event correlation, anomaly detection, and sophisticated incident investigation techniques.
Practical notes
This role operates on a hybrid schedule requiring presence in the San Jose office for three days each week. Candidates must be eligible to work in the United States without sponsorship for this position. Zscaler is committed to fostering an inclusive environment and provides reasonable accommodations for neurodivergent, differently abled, or other candidates requiring support during the recruiting process. All team members are expected to adhere strictly to company security and privacy policies.