Site Reliability Engineer
EarnInUSA2w ago
EngineeringReliabilityremotecurated-jd
Job description
Site Reliability Engineer at EarnIn.
About the role
EarnIn provides real-time financial flexibility to individuals living paycheck to paycheck. This role focuses on building resilient production systems and refining operational standards to ensure our community members have reliable access to their earnings.
Key facts
What you'll do
- Design systems with a focus on capacity planning, failure modes, and graceful degradation.
- Define and track SLIs and SLOs to guide reliability tradeoffs.
- Utilize observability tools including Datadog, CloudWatch, logs, metrics, traces, and APM.
- Manage incident response, including detection, triage, postmortems, and remediation.
- Improve alerting workflows to ensure pages are actionable and relevant.
- Build automation and infrastructure tooling to eliminate operational toil.
- Integrate AI-assisted tools to accelerate root-cause analysis, documentation, and infrastructure-as-code workflows.
- Partner with engineering teams to improve deployment safety and service ownership.
- Document operational knowledge to minimize information silos.
Requirements
- Bachelor's or master's degree in Engineering, Computer Science, or equivalent professional experience.
- 3+ years of experience in SRE, Software Engineering, or Infrastructure Engineering.
- Coding proficiency in Go, Python, or similar production-oriented languages.
- Experience with distributed systems, including timeouts, retries, backoff, and failure isolation.
- Hands-on experience with production operations, observability, and incident management.
- Familiarity with SLIs, SLOs, error budgets, and MTTR metrics.
- Experience using AI-assisted development tools like GitHub Copilot, Cursor, Claude, or ChatGPT.
- Strong communication skills with the ability to explain technical reliability concepts clearly.
Skills & tools
- Languages: Python, Go
- Observability: Datadog, CloudWatch, APM, logs, metrics, traces
- AI Tools: GitHub Copilot, Cursor, ChatGPT, Claude
- Concepts: SLIs, SLOs, MTTR, error budgets, distributed systems, infrastructure-as-code
Practical notes
This is a hybrid role requiring 2 days per week in the Mountain View office. EarnIn is an E-Verify participant. The company may use AI tools during the hiring process for resume review, scheduling, or interview summarization; final hiring decisions are made by human staff. Candidates may decline to have interviews recorded without impact on their evaluation. Unsolicited resumes from third-party recruiters are not accepted.