Validator Engineer
Job description
About the role
You own the uptime, performance, and reliability of Ritual's validators end to end, ensuring the chain remains secure and available at all times. You will manage upgrades, snapshots, and sync performance while serving as the primary technical contact for validator operators across testnet and mainnet. This role owns deployment tooling, documentation, and support to enable seamless staking participation and healthy ecosystem growth. You will build and automate bare-metal and cloud deployments using our existing full-fledged simulation framework to validate critical paths. You will orchestrate client upgrades, parameter changes, and hard-forks with zero downtime across both staging and mainnet environments. You will profile and patch execution and consensus clients to eliminate bottlenecks and optimize resource utilization. You will produce run-books, maintain clear documentation, and actively support validator operators via Discord to keep staking participation robust and transparent.
Key facts
What you'll do
- Build and automate bare-metal and cloud deployments for our validators using our existing full-fledged simulation framework to ensure repeatable and reliable infrastructure.
- Orchestrate client upgrades, parameter changes, and hard-forks with zero downtime across staging and mainnet while maintaining strict safety and consistency checks.
- Profile and patch critical paths for optimal performance of the execution and consensus clients to reduce latency and maximize throughput.
- Produce run-books, maintain clear documentation, and actively support validator operators via Discord to keep staking participation healthy and accessible.
- Monitor system health and performance metrics to detect anomalies early and drive rapid incident response with minimal disruption.
- Coordinate with infrastructure partners and stakeholders to align deployments, maintenance windows, and operational best practices.
- Implement observability pipelines that emphasize logs, metrics, and traces as the core input for every operational decision.
- Automate routine operational tasks to reduce manual overhead and improve scalability across testnet and mainnet validator sets.
- Conduct post-incident reviews and contribute to continuous improvement of deployment, recovery, and monitoring processes.
- Collaborate with cross-functional teams to ensure that validator operations meet the highest standards of reliability and security.
Requirements
- 4+ years operating high-availability distributed systems with direct experience running Proof-of-Stake validators on EVM-based chains.
- Proficient in Go and Rust, with the ability to read, debug, and extend both Execution Layer and Consensus Layer client code effectively.
- Strong grasp of protocol-level concepts such as fork-choice, evidence handling, slashing conditions, and state sync to ensure correct chain behavior.
- Expertise in Kubernetes, Helm, Terraform, and CI/CD workflows to automate everything from container builds to canary deployments.
- Advanced Linux user with deep networking knowledge and a strong focus on system hardening and security to protect critical infrastructure.
- Passionate about observability, treating logs, metrics, and traces as the foundation for monitoring and troubleshooting.
- Experience managing production-grade infrastructure with an emphasis on uptime, reliability, and rapid recovery from failure modes.
- Comfortable working in a fast-paced environment where protocol changes, upgrades, and operational challenges require quick adaptation.
Nice to have
- Direct experience supporting external validator operators, including writing documentation, maintaining snapshots, and answering integration questions during testnet or mainnet launches.
- Upstream contributions to both execution and consensus clients to improve protocol reliability and performance.
- Experience orchestrating GPU workloads or other compute-intensive pipelines that align with AI infrastructure demands.
- Familiarity with deploying, maintaining, and tuning open-source AI models in production environments.
- Experience interfacing with infrastructure partners, custodians, and auditors in high-uptime environments to ensure compliance and operational integrity.
Practical notes
The role is fully remote, allowing flexible location arrangements without travel requirements. There are no visa or travel constraints associated with this position. The engagement is full-time, and candidates should be available to respond promptly to operational incidents and support needs. No specific deadlines are imposed beyond standard project timelines defined in collaboration with the team.