Lead Site Reliability Engineer
Job description
Lead Site Reliability Engineer at Stuut Ai.
About the role
Stuut is transforming the B2B accounts receivable landscape by replacing outdated manual collection methods with an innovative automated platform. We are on the lookout for a Lead Site Reliability Engineer who will be responsible for designing our infrastructure and setting the operational benchmarks essential for our expansion in the industrials, chemicals, and manufacturing sectors. This role is pivotal in ensuring that our systems are robust, scalable, and capable of supporting our ambitious growth plans.
Key facts
What you'll do
- Define and implement a comprehensive reliability strategy that includes establishing Service Level Indicators (SLIs), Service Level Objectives (SLOs), and error budgets to guide our operational excellence.
- Oversee the management and scaling of our cloud infrastructure utilizing AWS and Kubernetes, ensuring that our systems are secure and resilient against failures.
- Create and maintain observability frameworks for logging, metrics, and tracing, enabling the team to derive actionable insights from system performance data.
- Lead incident response efforts and facilitate blameless postmortems to identify areas for improvement and enhance system reliability.
- Strengthen system resilience through effective capacity planning, redundancy measures, and failover strategies to mitigate risks.
- Optimize Continuous Integration/Continuous Deployment (CI/CD) pipelines to ensure that deployment processes are safe, observable, and easily reversible when necessary.
- Work closely with product and engineering teams to influence system architecture and scalability, ensuring that our solutions meet both current and future demands.
- Automate routine operational tasks to reduce manual effort and enhance developer productivity, allowing the team to focus on more strategic initiatives.
- Conduct thorough root cause analyses on complex reliability challenges to prevent recurrence and improve system performance.
- Provide mentorship to engineers on best practices in operational hygiene and reliability principles, fostering a culture of continuous improvement.
Requirements
- A minimum of 7 years of experience in site reliability engineering, infrastructure management, or backend development.
- Demonstrated expertise in designing and maintaining high-availability systems within dynamic and fast-paced environments.
- Strong programming skills in Python or TypeScript, particularly for the development of automation tools and scripts.
- In-depth technical knowledge of AWS, Kubernetes (specifically EKS), Docker, and cloud-native architectures is essential.
- Experience in implementing observability solutions and developing high-quality alerting mechanisms to enhance system monitoring.
- Familiarity with defining and enforcing SLOs and error budgets to ensure operational accountability.
- Knowledge of contemporary technology stacks, including FastAPI, Vue.js, PostgreSQL (RDS), and event-driven architectures.
- Proven experience in managing CI/CD processes, infrastructure as code, and modern deployment methodologies.
- Ability to strike a balance between maintaining system stability, accelerating development velocity, and optimizing costs.
Nice to have
- Experience with configuration management tools such as Terraform or Ansible.
- Familiarity with security best practices in cloud environments and experience in implementing security measures.
- Knowledge of machine learning or data analytics frameworks that could enhance our platform's capabilities.
Skills & tools
- Proficient in cloud services, particularly AWS, and container orchestration with Kubernetes.
- Strong understanding of monitoring and observability tools like Prometheus, Grafana, or similar.
- Experience with CI/CD tools such as Jenkins, GitLab CI, or CircleCI.
- Familiarity with programming languages and frameworks relevant to our tech stack, including Python, TypeScript, FastAPI, and Vue.js.
Practical notes
- The compensation package includes a competitive base salary ranging from $200K to $275K, along with equity options.
- Benefits for employees based in the U.S. encompass comprehensive medical, dental, and vision insurance plans.
- Additional perks include a 401(k) plan with matching contributions, flexible paid time off (PTO), and parental leave policies to support work-life balance.
META
Company: Stuut Ai
Title: Lead Site Reliability Engineer
Listed
location: San Francisco
Job type: full_time
About the company
Stuut is transforming accounts receivable for B2B companies - making collections smarter and faster for companies that have historically relied on manual processes that are labor intensive and costly. Our platform is gaining traction with finance teams across industrials, chemicals, and manufacturing sectors from Fortune 10 brands to scaling midmarkets.