Senior Site Reliability Engineer
Job description
About the role
PlayOn is looking for an experienced Senior Site Reliability Engineer to help us strengthen the reliability, performance, and scalability of our systems. This role sits at the intersection of software engineering and operations - focused on building the tools, automation, and visibility that enable our teams to deliver resilient software at scale. You will own the design and evolution of the infrastructure that powers our platform, ensuring that services remain performant and available under real-world conditions. You will translate complex operational challenges into automated, maintainable solutions that raise the reliability bar for the entire engineering organization. This is a hands-on role that requires deep technical judgment, strong collaboration, and a commitment to continuous improvement. You will shape industry-standard practices for monitoring, incident response, and delivery while mentoring other engineers on reliability principles.
Key facts
What you'll do
Assess and improve visibility across our platforms by reviewing current dashboards, metrics, and logs, identifying the biggest gaps, and implementing targeted improvements that clarify system health.
Refine monitoring and alerting for the most critical services so issues are detected earlier and teams can coordinate faster responses with clearer ownership.
Automate routine operational tasks and build tooling that reduces manual effort, enabling engineers to focus on higher-value work and long-term reliability improvements.
Define and implement observability practices by adding instrumentation and telemetry into existing build and deploy processes so reliability checks become part of the normal release workflow.
Establish and evolve service-level indicators and service-level objectives for core user flows, aligning teams on shared definitions of performance and availability.
Partner with application and quality engineering teams to implement best practices in reliability, release automation, testing, and change management.
Lead incident response improvements by streamlining communication, coordination, and follow-up processes during outages and service disruptions.
Drive operational excellence through proactive incident prevention, blameless postmortems, and capacity planning that informs future infrastructure decisions.
Participate in on-call rotations to support critical services and ensure rapid response to incidents while contributing to a healthy on-call culture.
Mentor and set standards for how engineering teams measure, monitor, and improve reliability across all services using data-driven decision-making.
Collaborate cross-functionally to translate complex operational requirements into clear technical specifications and implementation plans.
Evaluate and adopt observability tools and practices, including metrics, tracing, and logging solutions, to maintain a high level of insight into system behavior.
Continuously analyze performance trends and reliability metrics to guide infrastructure investments and prioritize automation efforts.
Act as a technical leader for reliability initiatives, balancing short-term firefighting with long-term platform health and scalability.
Requirements
Solid experience in Python, especially for automation, tooling, and data-driven operational tasks.
Proficiency in at least one statically typed language from the following set Java, C++, or Go.
Strong understanding of Linux systems, cloud infrastructure on major platforms such as AWS, GCP, or Azure, and modern deployment practices including Docker, Kubernetes, and Terraform.
Experience with CI/CD pipelines, version control systems, and automated testing frameworks is essential for effective collaboration with engineering teams.
Experience with observability tools such as Prometheus, Grafana, ELK, or Datadog, and the ability to use log and metric data to diagnose and resolve issues.
Proven experience facilitating and documenting Critical User Journeys and translating them into actionable service-level agreements and objectives.
Demonstrated ability to work effectively with cross-functional teams and communicate clearly in high-stakes, high-impact situations.
A problem-solver who approaches reliability as a shared responsibility across engineering and who is committed to improving practices across the organization.
Familiarity with incident management processes, postmortem culture, and blameless communication norms is strongly preferred.
Experience contributing to open source or building internal tools that improve developer productivity and operational visibility is valued.
Comfort working in fast-paced environments where priorities shift and ambiguity is common, with a track record of delivering results under pressure.
Strong written and verbal communication skills for documenting systems, procedures, and decisions for both technical and non-technical audiences.
A commitment to learning and applying industry-standard reliability engineering practices and collaborating with security and platform teams.