Principal Site Reliability Engineer
SezzleLatin America6d ago
EngineeringReliabilityremotecurated-jd
Job description
Principal Site Reliability Engineer at Sezzle.
About the role
Sezzle is looking for a Principal Site Reliability Engineer to take ownership of our infrastructure and systems. You will operate with high autonomy to identify bottlenecks, improve reliability, and scale our platform while mentoring the engineering team.
Key facts
What you'll do
- Design, build, and upgrade scalable infrastructure using AWS, Kubernetes, and RDS.
- Lead the infrastructure roadmap to improve system recoverability and performance.
- Conduct capacity planning, benchmarking, and stress testing.
- Establish and enforce SLAs, alerts, and anomaly detection protocols.
- Spearhead AI enablement efforts to automate infrastructure tasks and improve developer productivity.
- Maintain consistency across a distributed microservices architecture.
- Set engineering standards for observability, security, and CI/CD.
- Mentor staff and translate business objectives into technical roadmaps.
Requirements
- 12+ years of professional software or infrastructure engineering experience.
- Must have deployed significant changes to production infrastructure or applications within the last 30 days.
- Strong proficiency in Golang and building RESTful APIs.
- Expert knowledge of SQL-based RDBMS (MySQL, PostgreSQL) with experience in schema and query optimization.
- Experience with observability tools like Prometheus, Grafana, Datadog, or New Relic.
- Solid understanding of distributed system patterns such as event-driven architecture, queues, and transactional outbox.
- Bachelor degree in Computer Science or equivalent practical experience.
Nice to have
- Experience with AWS Aurora RDS (MySQL and Postgres).
- Background in data engineering, warehousing, and pipelines.
- Proficiency in CI/CD pipelines and containerized microservices on Kubernetes.
- Familiarity with AI developer tools like Claude Code, Gemini CLI, Codex, or Cursor.
- History of shipping commercial APIs in high-growth environments.
Skills & tools
- Languages: Golang, Typescript, Python
- Frontend: React, React Native
- Database: MySQL, Postgres, Elasticsearch
- Cloud & DevOps: AWS, Kubernetes, Gitlab
- Version Control: Git
Practical notes
This is a remote position based in Latin America. We prioritize candidates who demonstrate high standards, a bias for action, and the ability to challenge decisions constructively. We favor open-source solutions and emphasize automated testing across our development lifecycle.