
Site Reliability Engineer
Job description
About the role
At Careers, the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely. This is a remote position, so you'll be working remotely from your home. You may occasionally visit a Careers office to meet with your team for events or meetings.
Join Careers's Global Storage Engineering team, which operates one of the largest Ceph environments in the industry, powering the object, block, and file storage platforms that underpin hosting, applications, internal infrastructure, and next-generation AI/HPC workloads. If you are passionate about distributed systems, large-scale storage architecture, and solving complex reliability challenges, you will work on infrastructure that few engineers ever experience. At Careers, Ceph is not a side project - it is a critical platform. Our environment spans 80+ production clusters, 20,000+ OSDs, and approximately 300 PB of raw storage capacity, supporting tens of billions of objects across multiple continents. The scale demands deep technical expertise in storage architecture, automation, observability, and performance engineering.
As a Senior Site Reliability Engineer, you will be a key technical owner of the platform, responsible for maintaining reliability, driving operational excellence, and influencing the future evolution of our storage ecosystem. You will tackle challenging production problems, develop automation that operates at massive scale, contribute to architectural decisions, and collaborate with some of the industry's most experienced Ceph engineers. This is an opportunity to have direct impact on a storage platform that serves millions of customers worldwide.
What you'll do
- Own the reliability, performance, scalability, and capacity of large-scale production Ceph environments supporting object, block, and file storage workloads (RADOS, RGW, RBD, CephFS).
- Diagnose and resolve complex distributed systems issues including recovery/backfill events, OSD instability, PG imbalance, storage latency, and RGW performance bottlenecks.
- Design and build automation using Python, Shell, SaltStack, and Ansible to reduce operational toil and improve platform resilience at scale.
- Define and improve observability for the storage platform through SLIs, SLOs, PromQL, LogQL, Grafana dashboards, and proactive alerting strategies.
- Lead storage lifecycle initiatives including cluster expansions, hardware refreshes, software upgrades, OpenStack integrations, and large-scale migration projects.
- Partner with development and infrastructure teams to integrate Ceph with OpenStack, Kubernetes, and other platform services.
- Implement and maintain disaster recovery and business continuity strategies for critical storage services.
- Analyze capacity trends and drive data lifecycle management strategies including compression, erasure coding, and tiered storage.
- Collaborate with security and compliance teams to ensure storage controls, encryption, and auditability meet organizational standards.
- Mentor junior engineers and contribute to technical documentation and runbooks for operational procedures.
Requirements
- 5+ years operating large-scale Linux infrastructure with significant experience supporting distributed storage systems in production environments.
- 2+ years hands-on Ceph administration including OSD, MON, MDS, and RGW operations, CRUSH map management, pool design, placement groups, and performance troubleshooting.
- Strong understanding of Linux internals, storage architecture, networking, filesystems, block devices, and performance analysis under production workloads.
- Experience developing operational automation and tooling with Python, Shell, and configuration management platforms such as Ansible, SaltStack, Puppet, or Chef.
- Proven incident response expertise, including production troubleshooting, root-cause analysis, postmortem creation, and implementation of durable corrective actions.
- Demonstrated ability to work effectively in a remote-first environment with strong written communication skills.
- Comfort with monitoring and observability tools such as Prometheus, Grafana, Loki, and alerting pipelines.
- Familiarity with storage networking concepts and hardware constraints in high-density environments.
Practical notes
- Location Details: At Careers the future of work looks different for each team. Some teams work in the office full-time; others have a hybrid arrangement (they work remotely some days and in the office some days) and some work entirely remotely.
- This is a remote position, so you'll be working remotely from your home. You may occasionally visit a Careers office to meet with your team for events or meetings.
-
Engagement: Remote