Senior Site Reliability Engineer
BranchRemote (USA)3d ago
EngineeringSREReliabilityremotecurated-jd
Job description
Senior Site Reliability Engineer at Branch
About the role
Branch is seeking a Senior Site Reliability Engineer to enhance the performance, scalability, and dependability of our financial platform. You will be instrumental in driving improvements to our development and deployment workflows, establishing best practices, and ensuring the operational excellence of our systems.
Key facts
What you'll do
- Collaborate with development teams to ensure services are high-performing and reliable through thorough testing and release processes.
- Architect infrastructure, monitoring solutions, and operational procedures for our systems and applications.
- Provide support for services throughout their design, development, load testing, and launch phases.
- Measure and track key performance indicators, including availability and latency, to assess overall system health.
- Work with service owners to define and implement Service Level Indicators (SLIs) and Service Level Objectives (SLOs), and promote their adoption across teams.
- Analyze and optimize platform performance, resilience, and efficiency, including capacity planning under various load conditions.
- Participate in incident response activities and conduct root cause analyses.
- Address operational tasks and develop automated solutions to meet defined SLAs and SLOs.
- Manage the monitoring services used by our applications.
Requirements
- Bachelor's degree in a relevant engineering field or equivalent practical experience.
- A minimum of 3 years of experience in site reliability engineering.
- Proven experience building and operating Java / Spring Boot services in production environments.
- Hands-on experience with Terraform, Go, Java, Gradle, Docker, OpenTelemetry, and Kubernetes.
Nice to have
- Experience with scripting languages such as Python and Bash.
- Familiarity with developing and maintaining Kubernetes Operators.
- Exposure to Google Pub/Sub, Redis, Prometheus, Grafana, Google Spanner, and MySQL.
- Experience with JVM performance tuning and profiling.
- Prior work with Google Cloud Platform (GCP).
Skills & tools
- Java / Spring Boot
- Terraform
- Go
- Gradle
- Docker
- OpenTelemetry
- Kubernetes
- Python (nice to have)
- Bash (nice to have)
- Google Pub/Sub (nice to have)
- Redis (nice to have)
- Prometheus (nice to have)
- Grafana (nice to have)
- Google Spanner (nice to have)
- MySQL (nice to have)
Practical notes
- Candidates must be currently authorized to work in the USA without requiring sponsorship or transfer.
- This role is for remote work within the United States only.
- Benefits include comprehensive medical, dental, and vision insurance, stock options, a free premium Origin Financial Wellness subscription, a monthly home-office stipend, 401k, 12 weeks of paid parental leave, flexible time off, and 11 paid company holidays.
- Branch offers a same-day pay option.