Senior Software Engineer, Storage Infrastructure
Job description
About the role
The Emerging Technologies and Incubation team develops and launches new products by utilizing Cloudflare's global network scale. This role focuses on the core storage layer, managing everything from physical hardware to distributed databases that support our stateful services. You will join a team dedicated to ensuring our storage systems remain performant, reliable, and scalable as we build the next generation of infrastructure. In this capacity, you will drive the implementation of storage solutions that directly impact the reliability of our edge network. You will be responsible for translating high-level infrastructure goals into concrete storage architectures and interfaces. The position demands a hands-on approach to solving complex durability and performance challenges across distributed environments. You will work closely with cross-functional partners to define and deliver storage primitives that enable future product capabilities. Your contributions will ensure that storage infrastructure keeps pace with the demands of emerging technologies incubated within the team.
Key facts
What you'll do
- Build and maintain globally distributed storage systems from initial design to production release.
- Author design documentation for new provisioning systems and model failure domain dependencies.
- Benchmark storage hardware and create standardized observability tools and runbooks for database clusters.
- Automate operational tasks and reduce manual toil through custom tooling.
- Collaborate across the full stack to ensure durability and performance across our edge network.
- Implement storage interfaces that abstract complex underlying hardware for higher-level services.
- Analyze production telemetry to identify storage bottlenecks and drive remediation efforts.
- Partner with hardware teams to validate new devices and integrate them into the storage stack.
- Develop testing frameworks that validate data integrity and system resilience under adverse conditions.
- Support the evolution of our storage architecture to align with long-term product roadmaps.
- Investigate and prototype novel storage techniques to address scalability and efficiency challenges.
- Document operational procedures and ensure runbooks reflect the current state of deployed systems.
- Work with SRE practices to define service level objectives for storage components.
- Mentor junior engineers on best practices for building robust storage infrastructure.
Requirements
- Proficiency in programming languages such as Rust, Go, or Python.
- Understanding of distributed systems, including consensus, consistency, data replication, fault tolerance, and partition tolerance.
- Experience working with distributed databases and storage systems.
- Familiarity with infrastructure as code and configuration management tools.
- Knowledge of storage fundamentals like filesystems, block devices, and SSD characteristics.
- Experience developing high-throughput, low-latency systems.
- Understanding of network fundamentals, including bandwidth constraints, latency, and cross-datacenter replication.
- Ability to communicate complex technical decisions clearly in written and verbal formats.
- Capability to work in a fast-paced environment with shifting priorities and deadlines.
- Willingness to participate in on-call rotations to support production storage incidents.
- Commitment to adhering to security and compliance standards for storage systems.
- Demonstrated ability to write and maintain code that operates at global scale.
- Experience with monitoring and debugging complex distributed systems in production.
- Willingness to collaborate with hardware and firmware teams to resolve issues.
Nice to have
- Experience with contributing to open source storage projects.
- Familiarity with hardware-level storage interfaces and protocols.
- Knowledge of performance tuning for database storage engines.
Practical notes
-
Compensation: Estimated annual salary of $185,000 - $254,000 for hires in New York City, New Jersey, Washington, Washington DC, and California (excluding Bay Area).
- Equity: This role is eligible for the Cloudflare equity plan.
- Benefits: Includes medical, dental, vision, 401(k), disability insurance, life insurance, flexible PTO, and various family/parental leave programs.
- Export Control: This position may require access to information protected by U.S. export control laws. Offers may be conditioned on authorization to receive controlled technology without sponsorship for an export license.
- Interview process: Candidates reaching the offer stage may be required to attend an in-person interview at a Cloudflare office or hub.