Staff Site Reliability Engineer
Job description
About the role
You will lead the development of Domino's internal AI-assisted reliability tooling, owning the systems that analyze tickets, logs, traces, and documentation to resolve outages faster with less recurring toil. You will improve observability coverage and signal quality for our most critical customer-facing systems, ensuring engineers have rich, reliable data throughout the development and support lifecycle. You will own incident response end-to-end, from detection to remediation, and leave each problem space better documented, better understood, and less likely to recur. You will guide the development of customer and user-facing observability tools embedded in our products, aligning implementation with real operational needs. You will define and mature SLO and SLI frameworks for priority services, turning abstract reliability goals into measurable, actionable standards. You will scale cloud operations practices for Domino's single-tenant SaaS offering, working closely with engineering teams to improve the reliability and repeatability of customer deployments and upgrades. You will mentor other engineers and shape how SRE is practiced at Domino, including incident response workflows, operational readiness expectations, and a post-incident learning culture. You will use sound judgment about AI and LLM tooling, understanding where it genuinely helps operational workflows and where it adds noise instead of signal.
Key facts
What you'll do
Lead the development of Domino's internal AI-assisted reliability tooling, including systems that analyze tickets, logs, traces, and documentation to help teams resolve outages faster with less recurring toil.
Improve the observability coverage and signal quality for our most critical customer-facing systems, so engineers have more to work with throughout the development and support lifecycle.
Own incident response end-to-end, from detection to remediation, and leave each problem space better documented, better understood, and less likely to recur.
Guide the development of customer and user-facing observability tools within our products to align implementation with operational realities and user needs.
Define and mature SLO and SLI frameworks for priority services, turning abstract reliability goals into measurable, actionable standards that teams can track and enforce.
Scale cloud operations practices for Domino's single-tenant SaaS offering, and collaborate with engineering teams to improve the reliability and repeatability of customer deployments and upgrades.
Mentor other engineers and shape how SRE is practiced at Domino, including incident response workflows, operational readiness expectations, and post-incident learning culture.
Assess and implement AI and LLM-based tooling where they reduce manual toil and improve signal quality, while avoiding solutions that add noise or complexity without proportional value.
Partner closely with product, support, and infrastructure teams to ensure observability, reliability, and operational practices evolve as customer needs and product capabilities grow.
Contribute to on-call rotations and runbooks, ensuring that operational procedures are clear, automated where possible, and continuously refined based on incident learnings.
Drive automation of repetitive operational tasks, using software engineering practices to build tools that make day-to-day reliability work more efficient and scalable.
Champion a culture of learning and transparency around incidents, using them as opportunities to improve systems, processes, and team collaboration.
Requirements
You have deep experience in Site Reliability Engineering, platform engineering, or a software engineering role with genuine, hands-on operational ownership in production environments.
You are fluent with Kubernetes, Linux, cloud platforms, and observability tooling, and you can use these technologies to investigate and resolve complex, real-world production problems.
You have a strong ability to perceive and close reliability gaps in technical products, tools, and processes, turning qualitative concerns into concrete improvements.
You possess strong software engineering skills in Python or Go, with a track record of building internal tools or services that people in your organization actually rely on.
You are comfortable leading technically ambiguous work and influencing direction across teams without needing direct authority to get things done effectively.
You have a history of improving reliability through engineering and automation, rather than only reacting to fires manually or through ad hoc fixes.
You have strong communication skills and real experience mentoring engineers or shaping technical decision-making within your team or organization.
You demonstrate sound judgment about AI and LLM tooling, understanding where these tools genuinely help operational workflows and where they introduce noise or distraction.
You have experience working in SaaS environments, with exposure to multi-tenant considerations, customer impact analysis, and operational runbooks.
You are comfortable working asynchronously in a distributed team and can maintain clarity and alignment without constant synchronous communication.
You care deeply about building reliable systems and are willing to take ownership of complex, cross-functional operational challenges end-to-end.
You are based in or eligible to work remotely from Argentina and can commit to the expected working hours for this role.
Practical notes
LENGTH: 700-900 words. No HTML, no markdown, no em dashes.