Software Engineer - Site Reliability Engineering
Neo4jUK1d ago
EngineeringReliabilityremotecurated-jd
Job description
Software Engineer - Site Reliability Engineering at Neo4j.
About the role
Join the team maintaining Neo4j Aura, our managed database service that supports thousands of production instances across major cloud providers. You will move beyond reactive firefighting to build the infrastructure, automation, and culture that ensure long-term system resilience.
Key facts
What you'll do
- Develop automation to improve troubleshooting and scaling across our global Kubernetes fleet.
- Replace manual scripts with software-defined operations to ensure predictable and repeatable outcomes.
- Refine incident response processes, including blameless reviews and alert management.
- Partner with product teams to establish and track SLIs and SLOs.
- Build observability systems that provide actionable insights and early issue detection.
Requirements
- Proficiency in Go for backend tool development, with a focus on architecture and testing.
- Experience applying SRE methodologies in production environments.
- Ability to troubleshoot large-scale distributed systems.
- Experience with observability stacks such as Prometheus, Grafana, OTel Collector, or Google Cloud operations suite.
- Proficiency in managing applications on Kubernetes.
- Experience with infrastructure management tools like Terraform and Kustomize.
- Familiarity with CI/CD workflows, specifically GitHub Actions.
- Willingness to participate in on-call rotations and contribute to postmortems.
Skills & tools
- Go, Python
- Kubernetes, Kustomize, Terraform
- GitHub Actions
- Prometheus, Grafana, OTel Collector, Google Cloud operations suite
Practical notes
This is a hybrid role. Neo4j is a global company with over $200M in ARR and a valuation exceeding $2B. We encourage applicants from underrepresented groups to apply even if they do not meet every listed qualification. Please review our recruitment privacy notice on the Neo4j website before applying.