Platform Reliability Engineer
Job description
About the role
The role centers on owning the reliability and operational health of the Apify platform through a developer-centric lens. You will own the design and evolution of how we observe production systems and turn raw data into meaningful signals. This position focuses on sustainable practices and long-term improvements rather than reactive firefighting. You will define what healthy system behavior looks like and ensure incidents become opportunities for meaningful learning. Collaboration with product and platform engineers will shape how reliability practices integrate into everyday workflows. You will own concrete operational artifacts like runbooks, status page updates, and post-incident reviews. Ultimately, this role enables engineering teams to move faster with confidence by reducing noise and increasing signal quality.
Key facts
What you'll do
Operate and evolve the monitoring stack built on Prometheus, Grafana, and OpenTelemetry while defining which metrics truly represent customer experience.
Instrument services across the stack to expose the right signals, balancing cost, cardinality, and actionability for the platform and its consumers.
Define and refine alerting policies so that engineering teams receive timely, clear, and actionable signals without being overwhelmed by noise.
Shape incident management practices including communication protocols, timeline documentation, and structured learning sessions after every event.
Own status page communication during incidents, ensuring external and internal stakeholders receive consistent and timely updates.
Build and maintain runbooks that are practical, discoverable, and followed, turning tribal knowledge into repeatable procedures.
Partner closely with platform and product engineers to turn reliability requirements into practical standards that scale across teams.
Identify repetitive operational tasks and automate them, reducing manual toil and freeing engineers to focus on product work.
Write documentation that teams actually use, focusing on clarity, examples, and the reasoning behind reliability decisions.
Champion a blameless post-incident culture where learning is prioritized over punishment and improvements are tracked to closure.
Collaborate on the selection and rollout of observability tools, ensuring they meet the needs of both platform and application teams.
Evaluate new signals and dashboards continuously, pruning low-value metrics and adding high-value indicators as the platform evolves.
Participate in on-call rotations only when necessary, focusing on design and process work that reduces the need for after-hours intervention.
Guide deployment and release practices to improve reliability, drawing from CI/CD and release engineering best practices where relevant.
Contribute to the technical design of infrastructure components so that reliability considerations are built in from the start.
Requirements
You have hands-on experience choosing and tuning production metrics that reflect real user behavior and system health.
You are comfortable leading incident responses from detection through resolution and ensuring follow-up actions are completed.
You have hands-on experience with Prometheus, Grafana, and OpenTelemetry, using them to build meaningful observability pipelines.
You have used alert-routing tools such as PagerDuty to design escalation policies and reduce alert fatigue effectively.
You can read and write code to understand services, pipelines, and dependencies across the technology stack.
You have practical experience with post-incident reviews, tracking action items, and measuring improvements over time.
You know what a healthy post-incident culture looks like in practice and actively promote learning and psychological safety.
You can write clear, concise guidance that teams adopt, using examples and rationale to drive sound decisions.
You are self-driven to automate manual tasks and improve developer workflows through tooling and process changes.
Nice to have
Meaningful hands-on experience as an application or backend developer who has built services that run in production.
Experience building and maintaining infrastructure on AWS using services such as EC2, EKS, S3, CloudFormation, and related tooling.
Hands-on familiarity with container technologies and orchestration, including container lifecycle and networking concepts.
Some familiarity with CI/CD pipelines or release practices and an informed opinion on what makes deployments reliable and safe.
Practical notes
The role is full-time and based in Prague.
No travel requirements are specified.
No visa sponsorship information is provided in the source material.
No application deadline is mentioned in the source material.
Our tech stack
Infra: AWS Compute (Kubernetes (EKS), EC2, Lambda), Helm, ArgoCD, MongoDB, Redis, DynamoDB, S3, GitHub Actions
Monitoring: Grafana, Prometheus, OpenTelemetry, Mezmo, PagerDuty
Frontend: React.js, styled-components, Storybook, Chromatic, Cypress, Playwright
Backend: TypeScript/Node.js, Nest.js, Next.js, Express.js, Docusaurus, Vitest
Tools: GitHub, Notion, Google Workspace
Editor and AI assistant of your choice (GH Copilot, Cursor, Claude, Gemini, or JetBrains AI)
Process: two-week sprints, code reviews, tests, automating whatever we can, and deploying multiple times per day.
Have completed the general onboarding process and settled into the team rhythm.
Have built working relationships with platform engineers, engineering leads, and others involved in production response.
Understand, in principle, how the Apify platform works, and be able to handle smaller problems, incidents, or bugs on the infrastructure you work with most.
Have mapped how we handle monitoring, incidents, and alerts today, identifying friction points and opportunities for focused improvement.
Have published initial monitoring, observability, and alerting guidelines covering signals, naming, key dashboards, and alerting principles such as severity, routing, and noise reduction.
Be participating in incident reviews and translating patterns into improved playbooks and practices that reduce future risk.