Staff Software Engineer, Infrastructure
Job description
.
About the role
At Ripple, we are building a world where value moves like information does today. It is big, it is bold, and we are already doing it through crypto solutions for financial institutions, businesses, governments, and developers. This role improves the global financial system and creates greater economic fairness and opportunity for more people in more places around the world. You will see your impact while unlocking incredible career growth opportunities and growing your skills with colleagues who have your back. You will own the design, operation, and continuous improvement of the platform that powers our custody business.
Key facts
What you'll do
Design, build, and operate scalable, resilient infrastructure across cloud providers including Azure, AWS, GCP, and IBM Cloud.
Lead infrastructure architecture and reliability improvements for critical custody services to meet stringent security and uptime requirements.
Implement and improve monitoring, alerting, logging, and observability across distributed systems to provide clear insight into platform health.
Own and evolve blockchain node infrastructure, including high availability, failover strategies, and provider management to ensure continuity.
Build automation for infrastructure provisioning, deployments, testing, failover, incident response, and operational maintenance to reduce manual effort and risk.
Drive deployment operations and release reliability for platform and product services to enable frequent, safe, and predictable releases.
Proactively identify performance bottlenecks, reliability risks, and operational gaps before they impact customers and degrade experience.
Participate in on-call rotations, support production incidents, and help improve incident response processes and post-incident learning.
Contribute to internal platform tools, services, and developer workflows to increase productivity and consistency across engineering teams.
Create clear documentation, runbooks, and operational procedures for critical systems to ensure clarity, consistency, and compliance.
Mentor engineers and provide technical guidance on infrastructure, reliability, and platform engineering decisions to elevate the entire team.
Champion cost optimization, security best practices, and infrastructure scalability as foundational aspects of every solution you design.
Collaborate closely with product, security, and operations teams to align infrastructure roadmaps with business objectives and regulatory requirements.
Act as a technical leader in shaping engineering practices and fostering a culture of reliability, ownership, and continuous improvement.
Requirements
10+ years of experience in software engineering, platform engineering, infrastructure engineering, or systems operations for highly available production systems.
Experience managing blockchain nodes, including high availability, failover, and operational resilience in production environments.
Experience operating infrastructure in high-traffic, customer-critical, or security-sensitive environments with a strong focus on uptime and reliability.
Strong production experience with PostgreSQL for designing, tuning, and troubleshooting database workloads at scale.
Proven experience provisioning and managing infrastructure with Terraform or similar infrastructure-as-code tools to enable repeatable and auditable deployments.
Deep experience with containerized infrastructure and Kubernetes in highly available environments, including cluster lifecycle management and networking.
Proficiency with programming and scripting languages such as .NET, Go, Bash, or TypeScript to implement tools and automation.
Experience with messaging and queueing technologies such as RabbitMQ or AMQP for building reliable distributed systems.
Experience with observability tools such as OpenTelemetry, Grafana, Loki, Prometheus, or similar platforms to drive data-driven improvements.
Familiarity with GitOps practices using platforms like Argo CD or Flux to automate and govern deployments across environments.
Experience working with confidential computing solutions like AWS Nitro Enclaves, IBM Hyper Protect Virtual Servers, or GCP Confidential Computing to protect sensitive workloads.
Excellent ability in solving problems and the ability to diagnose complex distributed system issues using structured investigation methods.
Strong written and verbal communication skills to collaborate effectively with cross-functional teams and stakeholders in a global organization.
A commitment to following security and compliance frameworks relevant to financial services and infrastructure operations.
Practical notes
LENGTH: 700-900 words. No HTML, no markdown, no em dashes.
Output the page only.