Sr. Site Reliability Engineer
Job description
About the role
MeridianLink provides financial software as a service to credit unions, banks, and lenders. We are looking for a Senior Site Reliability Engineer to join our cloud engineering team and take ownership of how our critical financial platforms behave in production. In this role you will work across cloud environments to keep services reliable, secure, and fast for every customer who depends on them. This position is remote in the United States and reports into the infrastructure and platform organization. You will spend your time designing systems that hold up under load, catching problems before customers notice them, and building the tooling that lets the whole engineering team move quickly without breaking things. You will be expected to think about the system end to end, from the moment a request enters the network to the moment it is written to storage.
Key facts
- Full time position, located remotely in the United States.
- Team focus is reliability, scalability, observability, and automation.
- Primary cloud platforms are AWS and Azure.
- Daily tools include Kubernetes, Terraform, Ansible, Python, and CI/CD pipelines.
- Work touches security, compliance, and SOC requirements because we handle financial data.
What you'll do
- Define, track, and review Service Level Objectives and Service Level Indicators for all critical services, and make sure the team knows when targets are at risk.
- Lead the observability strategy by designing monitoring, logging, and distributed tracing architectures that give clear visibility into system health.
- Select and deploy the right observability tooling, then teach other engineers how to use it well.
- Write and maintain runbooks so that any on-call engineer can respond to an incident with confidence.
- Run incident response and coordinate post-incident reviews that focus on learning rather than blame.
- Build infrastructure as code with Terraform and Ansible so that cloud resources are reproducible and auditable.
- Manage Kubernetes clusters and the containerized workloads that run on them.
- Design and improve CI/CD pipelines to make releases safer and faster.
- Automate repetitive operational work so the team can spend time on problems that need human judgment.
- Work with security and compliance teams to keep our financial SaaS platform aligned with SOC and regulatory expectations.
- Mentor less experienced engineers on reliability practices, capacity planning, and production operations.
- Participate in on-call rotation and use that time to find and remove the causes of alerts.
Requirements
- Several years of experience in a site reliability, platform, or production engineering role.
- Strong working knowledge of at least one major public cloud, AWS or Azure.
- Hands-on experience running Kubernetes in production.
- Proficiency with Terraform and Ansible for infrastructure automation.
- Solid programming ability in Python or a similar language.
- Experience building and maintaining CI/CD pipelines.
- A clear understanding of SLOs, SLIs, error budgets, and incident response.
- Experience designing monitoring, logging, and tracing for distributed systems.
- Ability to communicate clearly with engineers, managers, and business stakeholders.
- A track record of working in or alongside regulated environments such as finance.
Nice to have
- Experience with financial services or fintech SaaS products.
- Familiarity with SOC 2, ISO 27001, or other compliance frameworks.
- Experience with Azure services in addition to AWS.
- Background in capacity planning and performance testing under load.
- Certification such as AWS Solutions Architect or Kubernetes Administrator.
Skills & tools
- AWS and Azure cloud services.
- Kubernetes and container orchestration.
- Terraform and Ansible.
- Python for scripting and automation.
- CI/CD tools such as Jenkins, GitHub Actions, or GitLab CI.
- Monitoring and logging platforms like Prometheus, Grafana, Datadog, or CloudWatch.
- Incident management processes and on-call tooling.
Practical notes
- This is a fully remote role based in the United States; you should be comfortable working with a distributed team across time zones.
- On-call rotation is part of the position, and the team actively works to reduce alert fatigue.
- You will work directly with security, compliance, and engineering teams on a regular basis.
- Interviews focus on real production scenarios, runbooks, and post-incident reviews rather than abstract puzzles.
- The company provides a stipend or reimbursement for home office and equipment needs.
Project highlights
Your first major projects will include raising reliability standards on the core loan origination platform, building a unified observability dashboard that spans AWS and Azure, and automating infrastructure changes so that releases can happen without manual steps.
Why you should apply
Reliability work at MeridianLink directly affects whether thousands of financial institutions can serve their customers every day. You get the autonomy to set technical direction on reliability, the support of a team that cares about doing the work properly, and the chance to work on systems with real scale and real consequences. If you want ownership, clear responsibility, and the freedom to improve systems from the inside out, this role gives you all three.