Distributed Systems Engineer
Job description
About the role
You will design and implement core distributed systems components that power Ritual's heterogeneous blockchain infrastructure, focusing on node specialization and consensus mechanisms. You will develop and optimize peer-to-peer networking protocols to ensure reliable AI model execution and blockchain consensus across geographically dispersed nodes. Your work will involve architecting solutions for efficient state management and synchronization to maintain network integrity under varying loads. You will collaborate closely with the research team to translate novel distributed systems algorithms into production-ready implementations. You will build and maintain infrastructure for node monitoring, metrics collection, and system observability to provide real-time insights into network health. You will create robust testing frameworks that validate the behavior and performance of distributed system components under failure conditions. You will ensure the reliability and performance of the Ritual network by identifying bottlenecks and implementing targeted improvements. You will take end-to-end ownership of complex problems, using excellent communication skills to align with cross-functional partners.
Key facts
What you'll do
Develop and deploy distributed systems components that form the backbone of Ritual's heterogeneous compute network.
Optimize peer-to-peer networking protocols to reduce latency and increase throughput for AI model execution and consensus operations.
Construct and maintain infrastructure for node monitoring, metrics collection, and system observability to ensure comprehensive visibility.
Architect strategies for efficient state management and synchronization across specialized nodes handling diverse workloads.
Work with the research team to implement and prototype novel distributed systems algorithms that address specific blockchain and AI challenges.
Establish and refine testing frameworks for distributed system components to verify correctness and resilience under stress.
Implement solutions that balance performance, reliability, and scalability for high-throughput blockchain operations.
Utilize advanced Linux expertise and command line proficiency to manage system administration and configuration at scale.
Leverage infrastructure orchestration tools like Kubernetes to automate deployment, scaling, and management of services.
Employ mastery of monitoring and observability tools such as Grafana and Prometheus to detect anomalies and drive improvements.
Design fault-tolerant, high-performance distributed applications that meet the rigorous demands of AI-native blockchain environments.
Demonstrate excellent problem-solving skills to debug intricate distributed systems issues that span multiple layers of the stack.
Act quickly and context switch at a high frequency to respond to evolving priorities and urgent production incidents.
Maintain a high level of end-to-end ownership and self-direction, ensuring timely delivery of critical milestones.
Requirements
You possess deep expertise in Go and/or Rust, with a track record of building production-grade distributed systems that operate at scale.
You have a strong understanding of peer-to-peer systems, networking, and messaging protocols that underpin modern distributed architectures.
You have proven experience operating and developing blockchain nodes, with deep knowledge of node architecture and consensus mechanisms.
You have experience optimizing high-throughput systems and implementing performance improvements that scale to meet growing demands.
You have advanced Linux expertise, with strong command line proficiency and system administration skills to manage complex infrastructure.
You are proficient with infrastructure orchestration tools like Kubernetes to ensure reliable and automated service delivery.
You have mastery of monitoring and observability tools such as Grafana and Prometheus to maintain system health and performance.
You have a strong background in system design, with the ability to make architectural decisions that balance performance, reliability, and scalability.
You have excellent problem-solving skills and the ability to debug complex distributed systems issues that require deep investigation.
You can act quickly under high pressure scenarios and context switch at a high frequency without losing focus or quality.
You bring a high level of end-to-end ownership and self-direction, with excellent communication skills to collaborate effectively with teams.
You are hungry, high-energy, and able to work within and meet deadlines in a fast-paced environment that demands rapid iteration.
Nice to have
Familiarity with AI/ML systems and their distributed computing requirements to better align infrastructure with model needs.
Knowledge of cryptography and secure systems design to enhance the robustness and trustworthiness of the platform.
Contributions to open-source distributed systems projects that demonstrate a commitment to the broader engineering community.
Practical notes
Participation in virtual and in-person events is expected as part of team engagement and professional development.
The role offers fully remote and/or hybrid arrangements, allowing flexibility in work location based on individual preference.
High-energy, fast-paced environment requires the ability to meet aggressive deadlines and adapt to shifting priorities.