Site Reliability Engineer, Observability
Job description
.
About the role
At Ripple, we are building a world where value moves like information does today. It is big, it is bold, and we are already doing it through crypto solutions for financial institutions, businesses, governments, and developers. In this SRE role focused on observability, you will own the design and implementation of monitoring, alerting, and dashboards that keep our payment infrastructure reliable and performant. You will work hands-on with New Relic, Terraform, and incident management tools while collaborating closely with stream-aligned product teams. The position sits within Technical Operations, primarily on Azure and some AWS, supporting a predominantly Windows environment that handles significant enterprise treasury transaction volumes. You will be instrumental in establishing practices for incident management and operational maturity, since these programs are early-stage and need thoughtful ownership.
Key facts
What you'll do
Design, implement, and maintain monitoring, alerting, and dashboards in New Relic (APM, Infrastructure, Logs, Synthetics) across Azure and AWS environments.
Write efficient NRQL queries to support troubleshooting, analysis, reporting, and service-level objectives for payment and treasury workloads.
Define and implement SLOs, SLIs, and error budgets, and coach teams on using these metrics to balance velocity with reliability.
Lead alert noise reduction and signal quality engineering by tuning thresholds, eliminating false positives, and ensuring alerts are actionable.
Partner with engineering teams to advance observability maturity through structured logging, metrics instrumentation using RED and USE methods, distributed tracing, and effective dashboard patterns.
Develop and maintain Terraform infrastructure as code for provisioning and managing monitoring resources, alert configurations, and observability infrastructure as a core engineering responsibility.
Establish and enforce governance standards for IaC related to observability infrastructure, creating repeatable and auditable models for monitoring resource management.
Author and troubleshoot Azure DevOps pipelines to improve deployment visibility, change tracking, and release hygiene as they relate to production reliability.
Administer and configure Incident.IO, including alert routing, notification workflows, Slack and OpsGenie integration, and runbook management to operationalize incident response capabilities.
Build out incident management foundations and practices that can scale as the organization grows, defining processes and playbooks where none currently exist.
Optimize observability costs by managing log ingestion, pipeline rules, and New Relic configuration to balance insight with efficiency.
Collaborate cross-functionally with product, security, and infrastructure teams to ensure observability solutions meet compliance, performance, and business requirements.
Contribute to on-call rotations and incident response activities, helping to stabilize production systems and reduce time to resolution.
Communicate system health and reliability trends to both technical and business stakeholders, translating complex data into clear, actionable insights.
Continuously evaluate new observability tools, patterns, and practices, and drive adoption where they can provide measurable improvements.
Requirements
Demonstrate experience with observability and reliability engineering in production environments, with a strong track record of managing monitoring and alerting at scale.
Bring proven skills in writing queries and creating dashboards in New Relic, including NRQL, APM, infrastructure, and log management.
Showcase experience designing and implementing Terraform configurations for cloud infrastructure and monitoring resources in Azure and AWS.
Have a solid understanding of SLOs, SLIs, error budgets, and alerting best practices, including strategies for alert noise reduction and signal quality.
Exhibit hands-on experience with incident management tools such as Incident.IO, OpsGenie, Slack, and runbook automation.
Display familiarity with Windows-based server environments and infrastructure, given the predominant Windows footprint in this role.
Provide evidence of experience working in Azure and some AWS environments, supporting cloud-hosted services and integrations.
Bring strong scripting or programming abilities, preferably in Python, PowerShell, or similar languages used for automation and observability tooling.
Nice to have
Experience contributing to open source observability projects or publishing reusable dashboards and Terraform modules.
Background in financial services, treasury, or high-volume transaction systems where reliability and compliance are critical.
Familiarity with compliance frameworks and audit-friendly documentation practices for observability and incident management.
Practical notes
This is a full-time position based in Chicago, Illinois, United States.
No specific visa sponsorship details or travel requirements are outlined in the source.
No application deadline is specified in the source material.