
Staff Site Reliability Engineer
Job description
About the role
You will own the design and operation of large-scale infrastructure that powers a national digital identity platform serving over 152 million users across federal, state, and healthcare ecosystems. You will drive the transformation from monolithic architectures to scalable, resilient microservices and platform APIs that define our product foundation. You will lead high-availability initiatives that directly define what "nine nines" of reliability means for our services and our customers. You will architect systems and deliver code that enables application teams to adopt robust practices by default rather than by manual instruction. You will integrate and unify a complex, heterogeneous infrastructure stack into a cohesive platform that supports rapid, secure growth. You will provide cross-functional technical leadership, aligning infrastructure strategy with product and business objectives across engineering teams. You will mentor a team of engineers, shaping the direction of infrastructure engineering and elevating the standard of what good operations look like.
Key facts
What you'll do
Execute on the transformation from monolith to scalable microservices with an API and platform engineering focus.
Drive initiatives to continually improve reliability, with a deep understanding of the operational and business implications of each "9" in availability.
Architect systems and write code that guides application teams toward best practices by default, not by exception.
Integrate and unify diverse infrastructure components into a cohesive, scalable platform within a massive, complex tech stack.
Design observability, reliability, and CI/CD frameworks that support growth and operational excellence at scale.
Collaborate cross-functionally with product, application, and integration teams to align infrastructure direction with business goals.
Provide technical leadership to shift the team from reactive support to a proactive, strategic function that anticipates needs.
Mentor and guide a team of 6 engineers while defining and shaping the future direction of infrastructure engineering.
Champion the adoption of platform thinking and infrastructure as code to enable consistent, repeatable delivery.
Evaluate and integrate new technologies and patterns that improve efficiency, resilience, and developer experience.
Own critical production incidents, conducting thorough reviews and driving remediation to prevent recurrence at scale.
Partner with security and compliance to ensure infrastructure meets federal and industry standards for identity and data protection.
Influence cost optimization strategies through efficient resource utilization and workload consolidation where appropriate.
Act as a subject matter expert, communicating complex infrastructure concepts to both technical and non-technical stakeholders.
Requirements
Bachelor's degree in Computer Science or a related field of study is required.
Possess at least 10 years of hands-on coding experience building internal platforms and tools that support developer experience and operational best practices.
Bring at least 5 years of experience working with cloud platforms, with a preferred background in GCP and acceptable experience in AWS, and a cloud engineering background is mandatory.
Demonstrate a proven track record of operating and scaling infrastructure in production environments serving high user loads.
Show experience leading infrastructure transformation initiatives, including migration strategies from monolithic to microservices architectures.
Have a strong history of building and maintaining observability, monitoring, and alerting systems to ensure system reliability.
Exhibit deep knowledge of CI/CD principles and the ability to design frameworks that enable safe and rapid deployments at scale.
Commit to following security and compliance standards relevant to identity verification, data privacy, and federal technology procurement.
Nice to have
Only items explicitly indicated as preferred in the source are included; no additional preferences are added.
Practical notes
Work is full-time and based in the Mountain View, California office five days per week.
This role is located in Mountain View, California, and requires presence at that office.
No specific visa, travel, or deadline information is provided in the source material.