Site Reliability Engineer
Job description
About the role
The owns the design and execution of infrastructure that delivers uncompromised uptime for our global staking operations. You will architect resilient systems that ensure our validators and RPC nodes run with maximum efficiency and security. This role focuses on extracting peak performance from our stack while embedding proactive risk mitigation into every deployment. You will own the automation that transforms complex blockchain infrastructure into repeatable, reliable workflows. Collaboration with product and engineering teams will shape the future of our yield and restaking products. Your work will directly influence the trust our clients place in us to safeguard and grow their digital assets. You will be a key driver in maintaining the standards that define us as the largest institutional staking provider in the market.
Key facts
What you'll do
Automate the provisioning and configuration of bare-metal servers using Ansible to maintain our critical infrastructure stack.
Refactor and enhance deployment pipelines to reduce manual intervention and eliminate single points of failure across data centers.
Develop custom monitoring and alerting solutions for blockchain nodes, including validators, RPC endpoints, and operator nodes, to ensure constant performance.
Conduct research and development on emerging blockchain networks such as TON, Avail, Monad, Babylon, Story, and Berachain to prepare for rapid deployment.
Optimize resource allocation and network configurations to achieve a Net Reward Rate that surpasses market benchmarks across ETH, SOL, and DOT.
Collaborate with product teams to integrate infrastructure changes that support new yield products and stablecoin aggregators.
Implement infrastructure as code practices to standardize environments from development through production using Terraform and Kubernetes.
Coordinate with on-call responsibilities to provide rapid response and resolution for infrastructure incidents affecting client services.
Utilize configuration management and container orchestration to ensure high availability and scalability of our platform.
Design and maintain secure network architectures with HAProxy and load balancing strategies to optimize client connectivity.
Leverage GitHub Actions to streamline CI/CD workflows, ensuring that every change meets rigorous security and performance criteria.
Analyze system performance metrics to identify bottlenecks and drive continuous improvement in node reliability and API response times.
Build and maintain integrations with major exchanges and custodians like BitGo, Copper, Crypto.com, and ByBit to support unified API and widget deployments.
Contribute to the expansion of our service offerings, including RWA initiatives and data products, by providing robust infrastructure foundations.
Document operational procedures and automation frameworks to ensure knowledge sharing and continuity across the distributed team.
Requirements
Expert level proficiency in Ansible is mandatory because the majority of our automation relies on this tool for configuration and orchestration.
Demonstrate solid programming capabilities in high-level languages such as Python or Go, with Go experience being a significant advantage for solving complex infrastructure challenges.
Exhibit strong ownership mentality, requiring self-organization and the ability to drive tasks forward with minimal guidance toward team objectives.
Possess a deep understanding of blockchain infrastructure components, including validators, RPC nodes, and operator nodes, and their operational demands.
Have experience managing infrastructure on bare-metal environments to ensure optimal performance and security for mission-critical services.
Show competence in using Terraform, Docker, and Kubernetes (K8s) to manage containerized applications and infrastructure as code.
Be skilled in configuring and maintaining HAProxy and load balancing solutions to handle high-traffic decentralized applications.
Commit to adhering to strict SLIs, SLAs, and SLAs to guarantee service reliability for our global client base.
Practical notes
The role operates on a fully remote basis from Europe, allowing flexible work arrangements without location constraints.
Employment is structured as a full-time contractor under an Indefinite-term Consultancy Agreement, providing long-term professional engagement.
The compensation package is competitive and denominated in US dollars, with the option to receive payment in cryptocurrency for flexibility.
You are entitled to paid vacation and sick leave, ensuring a healthy balance between professional duties and personal well-being.
Access to a well-being program and mental health care program supports your holistic health throughout your tenure.
Reimbursement is available for equipment and co-working expenses to create an optimal remote working environment.
The company invests in professional development by funding education, including foreign language courses and other professional growth opportunities.
You may be required to attend overseas conferences and engage with community immersion programs to represent P2P.org and expand our network.
Candidates must be able to work within the European time zone to ensure seamless collaboration with distributed team members.
All employment terms are aligned with the policies of P2P.org, which champions equal opportunity and non-discriminatory practices for all applicants.