Senior Distributed Systems Engineer
Job description
About the role
We are seeking a skilled engineer to join our team dedicated to data residency, where you will play a crucial role in developing the foundational infrastructure that dictates the geographical location of customer data processing within Cloudflare's vast global network. This position requires the creation of highly resilient and available distributed systems that adhere to strict compliance standards regarding data location.
Key facts
What you'll do
- Design, implement, and sustain backend services that enforce data residency policies across our extensive edge network, ensuring compliance with geographic regulations.
- Take full ownership of features, guiding them from the initial design phase to production deployment and ongoing maintenance, including the creation of technical documentation and specifications.
- Collaborate closely with teams across various engineering disciplines, such as cryptography, networking, and data management, to ensure residency guarantees are integrated seamlessly without compromising system performance or reliability.
- Conduct thorough analyses of potential failure scenarios, consistency implications, and overall system impact, focusing on compliance-first design principles that allow for graceful degradation.
- Engage in an on-call rotation, leading incident resolution efforts and contributing to the planning of system reliability and capacity.
- Foster improved engineering practices within the team through code reviews, mentorship, and effective communication.
- Develop and implement observability practices, including metrics, logging, and tracing, to enhance system monitoring from the design phase onward.
- AI-assisted development tools to enhance productivity while maintaining high standards of correctness and quality in your work.
Requirements
- At least 5 years of professional experience in designing, building, and maintaining large-scale distributed systems in a production environment.
- Proficiency in at least one backend programming language, such as Go or Rust, with a willingness to learn and adapt to others as necessary.
- In-depth knowledge of distributed systems principles, including various consistency models, consensus protocols, and strategies for data replication, sharding, and partitioning, along with their effects on availability and latency.
- Familiarity with common failure modes, including network partitions, split-brain scenarios, and clock synchronization challenges, as well as design patterns for managing failures, such as idempotency, retries, and circuit breakers.
- Experience with API design (REST or gRPC), relational databases, and asynchronous messaging systems, with a solid understanding of transactional boundaries.
- Demonstrated ability to take ownership of production systems, including responsibilities for on-call duties, incident management, and continuous enhancement of system reliability and performance.
- Strong written and verbal communication skills, enabling effective collaboration and the creation of comprehensive design documents across different teams and locations.
Nice to have
- Experience in developing systems that require strict compliance, security measures, or multi-region capabilities.
- Knowledge of fundamental cryptographic principles, including envelope encryption, key management, or public key infrastructure (PKI).
- Familiarity with edge computing, CDN platforms, L4/L7 proxies, or large-scale distributed storage solutions.
- Background in leading or contributing to complex engineering projects that involve multiple teams and cross-functional dependencies.
Skills & tools
- Go
- Rust
- PostgreSQL
- Kubernetes
- ClickHouse
- Workers
- Durable Objects
- REST
- gRPC
Practical notes
- Candidates who advance to the offer stage may need to participate in an in-person interview at a Cloudflare office.
- Job offers may be contingent upon the ability to receive U.S. export-controlled technology without the need for sponsorship.
This role offers a unique opportunity to work at the forefront of data residency and compliance within a leading technology company. If you are about building resilient distributed systems and are eager to collaborate with a diverse team of engineers, we encourage you to apply and contribute to our mission of enhancing internet security and performance for our customers worldwide.
About the company
Cloudflare operates one of the largest networks in the world, providing security, performance, and reliability services to websites and internet applications. Founded by Matthew Prince, Lee Holloway, and Michelle Zatlyn in 2009, Cloudflare went public on the NYSE in September 2019.