Staff Site Reliability Engineer
Job description
About the role
Zscaler is seeking a Staff Site Reliability Engineer to join the Cloud Infrastructure and Operations department. You will report to the Chief Architect of REIS and work from the San Jose office three days per week. In this role, you will act as a senior technical force multiplier guiding the reliability and efficiency of critical cloud platforms. You will own the design and execution of resilient infrastructure strategies that align with enterprise scale demands. The position requires deep collaboration with architecture and security teams to evolve operational best practices. You will influence platform direction through hands-on leadership in complex production environments. Your work will directly impact the stability and performance of services that power customer and business operations. This is a hands-on leadership position where technical depth meets strategic infrastructure ownership.
Key facts
What you'll do
- Construct and enhance scalable systems leveraging KVM Linux, Kubernetes clusters, and multiple public cloud providers.
- Diagnose and remediate performance bottlenecks across distributed applications and underlying operating systems.
- Oversee platform security and observability implementations through advanced monitoring frameworks and nftables configurations.
- Orchestrate robust software deployments across Openstack infrastructures and Kubernetes environments with precision.
- Engineer and sustain custom operational tools using BASH scripting, Golang, and Python for automation excellence.
- Lead incident response efforts by applying systematic troubleshooting methodologies to resolve critical platform issues.
- Partner with development teams to embed reliability patterns and operational readiness into application design lifecycles.
- Mentor junior engineers on infrastructure as code paradigms and contribute to technical documentation standards.
- Evaluate emerging technologies and translate them into scalable platform enhancements or integrations.
- Drive continuous improvement initiatives that optimize cost, performance, and resilience of production cloud services.
- Define and evolve platform standards for networking, compute, and storage components in line with security compliance.
- Act as a technical escalation owner for cross-functional issues that impact service reliability and availability.
Requirements
- Bring a minimum of 5 years of hands-on experience in Linux and UNIX System Administration with a proven track record.
- Demonstrate foundational knowledge of AI and ML technologies along with practical experience applying AI-driven solutions in operational contexts.
- Exhibit technical security proficiency with public key infrastructure (PKI), SSH, PGP encryption, and multi-factor authentication mechanisms.
- Master networking fundamentals including dynamic routing, subnetting, DHCP services, ARP operations, NAT configurations, IPv4 and IPv6 dual-stack environments, and enterprise firewalls.
- Show practical experience with container orchestration using Kubernetes and Docker alongside infrastructure automation using Ansible playbooks.
- Hold the ability to work legally in the United States for the position duration without requiring company sponsorship.
- Commit to adherence of Zscaler internal policies, security protocols, and regulatory compliance standards at all times.
- Maintain availability and reliability expectations for on-call rotations and critical production support windows.
Nice to have
- Hands-on experience with AIOps platforms and intelligent log parsing to enable AI-driven anomaly detection for predictive auto-scaling and root-cause analysis.
- Advanced expertise in Kubernetes ecosystems combined with deep experience in CEPH storage and Openstack infrastructure deployments.
- Background in dedicated Linux System Administration at scale with enterprise grade performance tuning and optimization.
- Familiarity with Secrets Management tools and operational workflows, particularly Hashicorp Vault for secure credential lifecycle management.
Practical notes
This position operates on a hybrid schedule with three mandatory days in the San Jose office each week. Employment is full-time and requires adherence to standard working hours and on-call responsibilities as scheduled. Candidates must be authorized to work in the United States without requiring sponsorship for employment visa or work authorization. Zscaler maintains an equal employment opportunity policy and provides reasonable accommodations for individuals with disabilities throughout the recruitment process. Reasonable accommodation requests related to the application or interview process should be directed in advance to the hiring team. Benefits coverage details, including health plans, retirement options, parental leave, vacation, sick time, and education reimbursement, are administered according to company policies and eligibility criteria.