Senior Site Reliability Engineer
Job description
About the role
You will proactively and reactively improve the reliability of Block's platform and critical infrastructure as a core member of the SRE team. You are metrics-driven, systems-oriented, and focused on building distributed platforms that enable safe, scalable product development. You will leverage and continuously improve AI-driven tooling and automation to enhance observability, accelerate incident detection and response, and reduce operational toil across the organization. This includes applying AI to incident analysis, alert tuning, and operational workflows to streamline processes. You will participate in primary platform oncall, supporting Block's most critical Tier 0 services, and lead incident command, coordinate mitigation, and drive effective escalation during high-severity events.
Key facts
What you'll do
- Build and extend platforms to improve system reliability and resiliency across distributed environments.
- Work on team goals that encompass reliability for the entire company, aligning initiatives with cross-functional priorities.
- Standardize reliability tools across multiple platforms and organizations to create consistency and efficiency.
- Triage, coordinate, and lead stabilization of sev 0-1 incidents, ensuring minimal customer impact.
- Serve as primary oncall, maintaining structured escalation paths and exercising leadership during complex scenarios.
- Drive platform-wide reliability improvements, shared operational tooling, and deploy-safety patterns for sustainable growth.
- Use AI-driven systems to improve signal detection, reduce noise, and accelerate root cause analysis with data-backed insights.
- Design and implement safe deployment patterns, including progressive delivery, automated rollback, and robust guardrails.
- Create and maintain evidence-based maturity assessments using trailing 90-day data windows to inform strategic decisions.
- Maintain vendor and dependency management practices, ensuring validated escalation contacts reachable within ≤ 5 minutes.
- Collaborate with engineering teams to integrate reliability best practices into the full development lifecycle.
- Champion automation opportunities that reduce operational toil and improve system observability at scale.
- Partner with product and infrastructure teams to define and evolve reliability standards and service objectives.
- Mentor and enable other engineers through knowledge sharing and hands-on support during critical incidents.
Requirements
- Drive to root cause systems with many moving parts and take the necessary steps to fix them effectively.
- Demonstrated technical initiative and leadership on previous projects, especially those with a backend or platform focus.
- Familiarity with AI-driven tooling for observability, incident analysis, or automation to enhance workflows.
- A mindset that naturally reaches for AI to accelerate problem-solving and reduce operational toil across teams.
- Experience running production oncall for high-availability systems in demanding environments.
- Strong incident management skills, including structured triage, mitigation under pressure, and blameless postmortems.
- Fluency with CI/CD pipelines, progressive rollout strategies, and rollback automation to ensure safe deployments.
- Monitoring and observability expertise, including building and tuning alerts for uptime, error rates, latency regression, and resource exhaustion.
- Ability to create and maintain evidence-based maturity assessments using trailing 90-day data windows for continuous improvement.
- Comfort with vendor and dependency management, maintaining validated escalation contacts reachable within ≤ 5 minutes.
- Boundless curiosity, autonomy, and a strong sense of accountability in dynamic and ambiguous situations.
- A strong desire to perform, learn, and grow as an engineer while contributing to high-impact outcomes.
- 5+ years of software development experience, with a focus on building and operating reliable systems.
Practical notes
- Hours: Please refer to the source material for detailed working hours and scheduling expectations.
- Travel: No additional travel requirements specified in SOURCE beyond standard operational needs.
- Visa: SOURCE does not provide specific visa sponsorship details; interested candidates should refer to official application instructions.
- Deadlines: SOURCE does not specify application deadlines; candidates are encouraged to apply as soon as possible through official channels.
We are working to build a more inclusive economy where our customers have equal access to opportunity, and we strive to live by these same values in building our workplace.