Distributed Systems Engineer
Job description
About the role
Become a key member of the Data Localization team, where you will be instrumental in developing the foundational infrastructure that governs the location and processing of customer data across Cloudflare's vast global network. This position emphasizes the creation of distributed systems that adhere to strict geographic regulations, ensuring the high level of scale and reliability that our clients expect.
Key facts
What you'll do
- Architect, implement, and manage systems that uphold data residency regulations, which includes policy enforcement and cryptographic key management at the network's edge.
- Create user-friendly APIs and dashboards that allow customers to control their data localization preferences effectively.
- Engage with various layers of the technology stack, from low-level policy execution to user interface development, primarily using Go and Rust.
- Take full ownership of features, guiding them from the initial design phase through to production deployment and subsequent maintenance.
- Tackle intricate challenges where regulatory compliance is paramount, and system failures could significantly affect customers.
- Collaborate with cross-functional teams to ensure that the systems you build are scalable and maintainable.
- Conduct thorough testing and validation of systems to ensure they meet performance and reliability standards.
- Continuously monitor system performance and implement improvements based on observed metrics and user feedback.
Requirements
- At least 5 years of experience in the design, construction, and maintenance of large-scale distributed systems in a production environment.
- Proficient in at least one backend programming language, such as Go, Rust, or C/C++, with a willingness to learn new languages as required.
- Strong grasp of distributed systems concepts, including various consistency models (strong, sequential, eventual) and consensus algorithms (such as Paxos or Raft).
- Familiarity with data replication, sharding, and partitioning strategies, and their effects on system availability and latency.
- Understanding of common failure scenarios, including partial failures, network partitions, split-brain issues, and clock skew.
- Experience in designing systems that are idempotent and implementing retry strategies, backpressure, timeouts, circuit breakers, and rate limiting.
- Knowledge of health checking, failure detection, leader election, and graceful degradation techniques.
- Ability to integrate observability features (metrics, logs, tracing) into system architecture from the beginning.
- Proven experience in API design (REST or gRPC), relational databases, and asynchronous messaging systems, with a solid understanding of transactional boundaries.
- Familiarity with AI-assisted development tools, demonstrating the ability to enhance productivity while maintaining correctness, security, and design integrity.
- A track record of owning production systems, including participating in on-call rotations, managing incident responses, conducting post-mortem analyses, and driving continuous reliability and performance enhancements.
- Strong written and verbal communication skills, essential for producing clear design documentation and collaborating effectively across different time zones.
Nice to have
- Experience in developing systems that are compliance-driven, manage sensitive security data, or function across multiple geographic regions.
- Basic understanding of cryptographic principles, including envelope encryption, key management, hardware security modules (HSMs), or public key infrastructure (PKI).
- Familiarity with edge computing, content delivery networks (CDN), L4/L7 proxies, or large-scale globally distributed storage and key-value solutions.
- Experience in leading or contributing to engineering projects that require collaboration across multiple teams, including platform, security, and product groups.
Skills & tools
- Go
- Rust
- PostgreSQL
- Kubernetes
- ClickHouse
- REST APIs
- gRPC
- AI-assisted development tools
Practical notes
This position may require access to information that is subject to U.S. export control laws. Employment offers may depend on the ability to receive controlled software or technology without needing an export license. Cloudflare is dedicated to being an equal opportunity employer, promoting diversity and inclusion. All qualified candidates will be considered without regard to race, color, religion, sex, gender identity, gender expression, sexual orientation, national origin, ancestry, citizenship, age, disability, medical condition, or any other protected status. Reasonable accommodations are available for individuals with disabilities during the application process. For assistance, please reach out to hr@cloudflare.com or visit us at 101 Townsend St. San Francisco, CA 94107.