Staff Reliability Engineer
ServiceNowUSAFull-time2d ago
EngineeringReliabilityremotecurated-jd
Job description
Staff Reliability Engineer at ServiceNow.
About the role
You will develop cloud-native platforms focused on reliability, release management, and testing to support ServiceNow operations. This role involves building AI-driven tools and automation to improve developer productivity and ensure high-confidence software deployments.
Key facts
What you'll do
- Architect and manage cloud-native engineering platforms for software validation and production readiness.
- Create automated test pipelines and integrate observability, reliability signals, and quality gates into CI/CD workflows.
- Build reusable frameworks for test data management, mock services, and developer self-service environments.
- Develop Kubernetes-based infrastructure to support scalable test workloads and release automation.
- Implement automated checks for security, policy enforcement, resilience, and operational health.
- Mentor engineering staff through technical guidance and code reviews.
- Resolve complex infrastructure and networking issues using software engineering and systems design.
Requirements
- 8+ years of experience in SRE, DevOps, Platform, or Software Engineering with a Bachelor degree; 6 years with a Master degree; or 3 years with a PhD.
- Hands-on experience with Kubernetes, including cluster operations, networking, storage, and autoscaling.
- Proficiency in Python, Go, Java, or Ruby.
- Experience with cloud-native platform operations and CI/CD integration.
- Knowledge of progressive delivery, such as canary deployments, feature flags, and automated rollbacks.
- Understanding of observability, SLI/SLOs, and incident management for distributed systems.
- Experience integrating AI into engineering workflows for decision-making or diagnostics.
Nice to have
- Experience with GitLab CI/CD, Argo CD, or Flux.
- Proficiency in test automation frameworks like Playwright, Selenium, Cypress, REST Assured, PyTest, or JUnit/TestNG.
- Familiarity with Infrastructure as Code tools like Terraform or Ansible.
- Experience with the Kubernetes ecosystem, including Helm, Istio, Linkerd, Prometheus, and OpenTelemetry.
- Background in operating Kubernetes across AWS (EKS), Azure (AKS), or Google Cloud (GKE).
Skills & tools
- Kubernetes, Python, Go, Java, Ruby, AWS, Azure, GCP, Terraform, Ansible, Prometheus, OpenTelemetry, Istio, Linkerd, GitLab CI/CD, Argo CD, Flux, Playwright, Selenium, Cypress, REST Assured, PyTest, JUnit/TestNG.
Practical notes
- Base pay is $166,500 - $291,400, plus equity, variable compensation, and benefits including 401(k) match and health plans.
- Employment is subject to export control regulations where required.
- Reasonable accommodations for the application process are available by contacting globaltalentss@servicenow.com.