Staff Engineer - Platform
Job description
About the role
You will architect and deliver the core platform capabilities that power the Illumio AI Security Graph across global hybrid multi-cloud deployments. You will own the design and execution of scalable control planes that enforce Zero Trust policies dynamically and at scale. You will lead the integration of distributed systems components to ensure cohesive operation across heterogeneous environments. You will drive the adoption of cloud native patterns that balance architectural integrity with pragmatic business delivery. You will mentor engineers on best practices for resilient and observable infrastructure. You will translate complex security requirements into robust platform abstractions. You will ensure that every deployment reflects the highest standards of performance, security, and operational excellence.
Key facts
What you'll do
Architect and own the platform roadmap that underpins the Illumio AI Security Graph and its breach containment capabilities.
Design and implement scalable control plane services that operate consistently across on-premises, AWS, Azure, and GCP environments.
Collaborate closely with product managers and developers to translate business requirements into resilient cloud architectures.
Define and enforce standards for Kubernetes orchestration, service mesh, and containerized workload execution.
Optimize distributed systems for fault tolerance, high availability, and rapid disaster recovery at global scale.
Implement comprehensive monitoring, logging, and alerting frameworks to provide deep visibility into platform health.
Diagnose and resolve performance bottlenecks, security misconfigurations, and operational anomalies in real time.
Evaluate emerging cloud technologies and open source tools to drive continuous innovation in platform capabilities.
Automate provisioning, deployment, and lifecycle management of platform components using infrastructure as code.
Partner with UI and visualization teams to ensure that platform telemetry and control surfaces deliver actionable insights.
Champion Zero Trust principles through platform features that enable dynamic segmentation and adaptive policy enforcement.
Engage with the broader engineering community to share knowledge and elevate platform thinking across the organization.
Ensure all platform deliverables comply with export control regulations and security compliance frameworks.
Lead incident response efforts for platform-wide issues, driving postmortems and preventative improvements.
Act as a technical leader in shaping architectural governance across the engineering organization.
Requirements
Candidates must possess a Bachelor's degree in Computer Science, Engineering, or a related technical field or equivalent practical experience.
You must have 8+ years of professional experience in software engineering, cloud architecture, or platform engineering roles.
You must have proven experience designing and operating large-scale distributed systems in production environments.
You must have extensive hands-on experience with AWS, Azure, and Google Cloud Platform APIs and services.
You must have in-depth knowledge of networking, security controls, and identity and access management in cloud environments.
You must have substantial experience with Kubernetes, Docker containers, and orchestration at scale.
You must be proficient in at least one high-level programming language such as Java or Python.
You must have demonstrated ability to build highly scalable and resilient cloud-native services from the ground up.
You must have strong experience with Infrastructure as Code tools such as Terraform or CloudFormation.
You must have excellent troubleshooting, debugging, and performance optimization skills across distributed stacks.
You must have strong written and verbal communication skills to collaborate effectively with cross-functional teams.
You must be comfortable working autonomously in an agile environment with minimal supervision.
You must be able to obtain U.S. export control compliance clearance due to access to controlled technology.
You must be willing to relocate to or maintain work authorization in the Sunnyvale, HQ location.
Nice to have
Experience with service mesh technologies such as Istio or Linkerd.
Contributions to open source projects related to cloud infrastructure, security, or networking.
Familiarity with CI/CD pipelines and GitOps workflows.
Knowledge of security information and event management (SIEM) platforms and log analytics.
Experience with chaos engineering and resilience testing practices.
Practical notes
This role is based in Sunnyvale and requires onsite presence at the headquarters office.
Employment is full-time and contingent upon successful completion of background checks and export control compliance verification.
All offers are issued through the official recruitment team and require electronic signature documentation.