Site Reliability Engineer
Job description
About the role
EarnIn provides real-time financial flexibility to individuals living paycheck to paycheck. This role focuses on building resilient production systems and refining operational standards to ensure our community members have reliable access to their earnings. You will own the reliability and operational integrity of critical production services that directly impact how users access their earned wages. The position requires a proactive mindset to identify potential failure points before they affect the user experience and community trust. You will be responsible for designing systems that maintain high availability even under unpredictable load patterns. Collaboration with engineering teams will ensure that reliability practices are embedded into the software development lifecycle. Your work will help define how the platform scales, degrades, and recovers in complex production environments.
Key facts
What you'll do
- Design systems with a focus on capacity planning, failure modes, and graceful degradation to support fluctuating user demand.
- Define and track SLIs and SLOs to guide reliability tradeoffs and ensure alignment with business objectives.
- Utilize observability tools including Datadog, CloudWatch, logs, metrics, traces, and APM to maintain visibility into system health.
- Manage incident response, including detection, triage, postmortems, and remediation to reduce future risk.
- Improve alerting workflows to ensure pages are actionable and relevant, reducing noise and fatigue for on-call staff.
- Build automation and infrastructure tooling to eliminate operational toil and standardize common operational tasks.
- Integrate AI-assisted tools to accelerate root-cause analysis, documentation, and infrastructure-as-code workflows for faster iterations.
- Partner with engineering teams to improve deployment safety and service ownership through clear interfaces and documentation.
- Document operational knowledge to minimize information silos and create reusable playbooks for future scenarios.
- Analyze trends in production incidents to drive long-term improvements in system resilience and architecture.
- Collaborate with cross-functional teams to balance reliability requirements with delivery speed and product innovation.
- Implement monitoring strategies that provide early warnings for anomalies in critical user journeys and financial transactions.
- Evaluate new technologies and patterns that can enhance the stability and performance of distributed services.
- Support the rollout of new features by ensuring that operational readiness and rollback strategies are in place.
- Contribute to the evolution of best practices for reliability engineering across the organization.
Requirements
- Bachelor's or master's degree in Engineering, Computer Science, or equivalent professional experience in a related technical field.
- 3+ years of experience in SRE, Software Engineering, or Infrastructure Engineering focused on production systems.
- Coding proficiency in Go, Python, or similar production-oriented languages to develop and debug automation scripts.
- Experience with distributed systems, including timeouts, retries, backoff, and failure isolation to maintain service integrity.
- Hands-on experience with production operations, observability, and incident management to respond effectively to disruptions.
- Familiarity with SLIs, SLOs, error budgets, and MTTR metrics to make data-driven reliability decisions.
- Experience using AI-assisted development tools like GitHub Copilot, Cursor, ChatGPT, or Claude to enhance productivity and code quality.
- Strong communication skills with the ability to explain technical reliability concepts clearly to both technical and non-technical stakeholders.
- Demonstrated ability to work independently and collaboratively in a fast-paced, high-ownership environment.
- Willingness to engage in on-call rotations and respond to incidents outside of standard business hours as needed.
- Understanding of financial services operations and compliance considerations is beneficial but not required.
- Commitment to maintaining a secure and reliable environment for handling sensitive user data and transactions.
Practical notes
This is a hybrid role requiring 2 days per week in the Mountain View office. EarnIn is an E-Verify participant. The company may use AI tools during the hiring process for resume review, scheduling, or interview summarization; final hiring decisions are made by human staff. Candidates may decline to have interviews recorded without impact on their evaluation. Unsolicited resumes from third-party recruiters are not accepted.