Senior Staff Software Engineer-SRE
Job description
About the role
This role defines and delivers the platform foundations that enable reliable, scalable, and performant products. The Senior Staff Software Engineer-SRE designs strategic approaches to system stability and partners across the engineering organization to embed reliability into every phase of delivery. You will influence technical direction, lead incident response, and mentor teams to achieve operational excellence.
Software engineers turn product ideas into working code. Engineers work in small teams, review each other's work, and ship in small batches. Most teams follow agile practices such as sprints and daily standups. Engineers also write tests, fix bugs, and improve performance. The field values clear communication as much as technical skill. Engineers spend part of every week on planning, code review, and debugging, not just writing new code. The ability to explain a technical decision in plain words separates strong engineers from the rest.
Key facts
What you'll do
Incident response efforts are led, root causes are analyzed, and long-term fixes are implemented to prevent recurrence.
System behavior is analyzed, bottlenecks and saturation points are identified, and solutions are implemented to improve resilience.
Reliability is embedded into the software development lifecycle through partnership with engineering teams.
Emerging technologies are evaluated and tools are recommended to enhance productivity, observability, and system robustness.
Technical leadership and mentorship are provided across the engineering organization to elevate overall capability.
Requirements
A degree in Computer Science, Engineering, or related technical field is required.
You bring 10+ years of experience in software engineering with a strong focus on backend systems and distributed architecture.
You have extensive experience building and operating Java-based systems using RESTful APIs, Spring Boot, and Microservices architecture.
You understand distributed systems concepts deeply, including fault tolerance, eventual consistency, and scalability.
You have proven experience with cloud platforms such as AWS, Azure, and GCP, and with cloud-native architectures.
You are expert in observability tools for monitoring, logging, and tracing, such as Prometheus, Grafana, ELK, or similar.
You define and manage SLIs, SLOs, and error budgets as part of reliability engineering.
You have hands-on experience with CI/CD pipelines, automation, and infrastructure as code.
You have practical experience with incident management, root cause analysis, and postmortems.
You possess strong analytical, debugging, and problem-solving skills.
You communicate, collaborate, and lead effectively across diverse teams.
Nice to have
Experience with emerging technologies and the ability to recommend tools that enhance observability and robustness is valued.
Skills & tools
The role relies on platforms built with Java, RESTful APIs, and Spring Boot.
Practical notes
The role follows an office-first culture, with an expectation of three days per week in the office for most positions. Specific flexibility may vary based on scope and will be confirmed with your recruiter during the interview process.
Typical interview steps
Hiring for engineering roles usually starts with a recruiter screen, followed by one or two technical rounds. Candidates often solve a coding problem, discuss past projects, and answer system design questions. Some loops include a take-home task. Final rounds typically cover team fit and give candidates a chance to ask questions. Interviewers look for how you break down an unfamiliar problem, not just whether you reach the answer. Practicing a few problems aloud and reviewing your own past projects are the best preparation.
Good to know
Reliability engineering focuses on ensuring systems perform as expected under real-world conditions. Observability tools help teams understand system behavior and accelerate troubleshooting. Cloud-native practices support scalable and resilient architectures. Incident response and postmortems drive continuous improvement. Mentorship and cross-team collaboration strengthen the overall engineering organization.