Software Engineer II, Reliability
Job description
About the role
At Klaviyo, we value the unique backgrounds, experiences and perspectives each Klaviyo (we call ourselves Klaviyos) brings to our workplace each and every day. We believe everyone deserves a fair shot at success and appreciate the experiences each person brings beyond the traditional job requirements. If you're a close but not exact match with the description, we hope you'll still consider applying. Want to learn more about life at Klaviyo? Visit klaviyo.com/careers to see how we empower creators to own their own destiny. As a Software Engineer II, Reliability, you will help ensure Klaviyo's critical platforms are reliable, scalable, and sustainable while enabling rapid product development. We treat reliability as a core product feature and use software engineering to solve complex systems and operational challenges. Our work spans infrastructure, security, and software engineering, and focuses on building and operating systems that are reliable, secure, and performant at scale. The SRE team's charter is to build and operate foundational services and infrastructure, reduce operational toil through automation, and continuously improve systems based on real production learnings. Your work will directly impact how Klaviyo engineers build software and how customers experience our platform every day.
Key facts
What you'll do
Define, deliver, and iterate on well-scoped reliability projects as an early-to-mid career SRE with support from senior engineers. Own services and operational tasks that directly uphold the reliability, scalability, and performance of Klaviyo's critical platforms. Apply software engineering practices to automate operational work, aiming to minimize manual toil and improve system consistency. Contribute to the design and implementation of systems using established SRE practices and patterns. Define, measure, and refine SLIs and SLOs for services you support to reflect real user needs and service expectations. Improve observability across the stack by enhancing metrics, dashboards, logging, and tracing to support faster detection and diagnosis. Participate in on-call rotations and respond to production incidents with guidance and support from more senior team members. Assist with incident investigation and contribute to post-incident reviews and follow-up actions to prevent recurrence. Perform basic analysis of system behavior, capacity usage, and scaling characteristics to inform operational decisions. Identify reliability issues or operational pain points and collaborate with teammates to implement practical fixes. Partner with product, platform, and security engineers to ship reliable systems that meet customer and business needs. Write and maintain clear operational runbooks and system documentation to make operations transparent and repeatable.
Requirements
Have experience operating cloud-native production systems and services in a professional setting. Write production-quality code (e.g. Python, Go, or similar) to automate operations and improve reliability of services. Understand common failure modes in distributed systems, such as dependency failures, resource exhaustion, and partial outages. Have hands-on experience with containerized workloads and platforms (e.g. Kubernetes) in production environments. Be comfortable participating in on-call rotations and diagnosing straightforward production issues using available tools. Have experience using observability tools and responding to alerts in a production environment. Be familiar with SRE concepts such as SLIs, SLOs, and error budgets, and be in the process of learning how to apply them in practice. Have hands-on experience with infrastructure as code or declarative configuration (e.g. Terraform, Kubernetes manifests) to manage environments safely and consistently. Be able to follow incident response processes and contribute meaningfully during outages to restore service and limit impact. Be comfortable receiving feedback and collaborating closely with cross-functional teams to iterate on solutions.
Nice to have
Only if the source specifies preferred items. The source text does not include explicit Nice to have qualifications, so this section remains empty.
Practical notes
Employment type: Full-time.
Location: Dublin, Ireland. This role requires participation in on-call rotations, including availability during weekends and holidays as needed. Travel is not specified in the source, and no visa sponsorship details are provided. Candidates must meet the stated experience and technical requirements. The compensation details are not provided in the source.