VP Engineering - Infrastructure & SRE
SezzleRemote (United States,)4d ago
VPEngineeringInfrastructureSREremotecurated-jd
Job description
VP Engineering - Infrastructure & SRE at Sezzle
About the role
Sezzle is seeking a senior engineering leader to oversee the strategy, reliability, and security of its core platform. This is a hands-on role that requires deep technical expertise and the ability to lead teams through a period of significant growth and increased operational rigor.
Key facts
What you'll do
- Define and execute the long-term vision and roadmap for Sezzle's infrastructure, focusing on scalability, resilience, and compliance.
- Lead, mentor, and grow a team of engineers responsible for cloud infrastructure, Kubernetes, databases, networking, and site reliability.
- Ensure the availability, disaster recovery, and business continuity of a critical payments platform.
- Drive the adoption of AI-powered tools and workflows within SRE practices for incident response, capacity planning, and automation.
- Collaborate with Security and Compliance teams to meet stringent financial industry standards, including PCI-DSS and SOC 2.
- Manage and participate in the on-call rotation, acting as incident commander during high-severity events.
- Oversee the AWS cloud strategy, Kubernetes platform, and database tier (Aurora RDS).
- Champion infrastructure-as-code principles and drive automation across all infrastructure operations.
- Manage significant cloud budgets and focus on cost efficiency.
Requirements
- A minimum of 15 years of combined experience in infrastructure, platform engineering, SRE, or software development, with at least 5 years in a leadership capacity.
- Extensive, hands-on experience with AWS, including designing and operating production environments at scale (compute, networking, IAM, multi-account structures).
- Deep, hands-on expertise with Kubernetes in production environments, including cluster management and running critical services (EKS experience is highly preferred).
- Profound knowledge of relational databases at scale, specifically Aurora RDS (MySQL and/or Postgres), covering high availability, performance, and recovery.
- Demonstrated experience owning and implementing disaster recovery and business continuity plans for production systems, including testing.
- Proven experience leading teams in adopting AI tools for engineering or operations.
- A track record of operating 24/7 high-availability platforms with direct revenue or customer impact.
- Willingness to participate in an on-call rotation and lead incident response for major outages.
- A strong technical foundation, remaining hands-on with systems and designs.
- Experience managing substantial cloud budgets and driving cost optimization.
- Proficiency with infrastructure-as-code tools like Terraform.
- A history of successfully hiring, developing, and retaining engineering talent.
- Bachelor's degree in Computer Science or a related technical field.
Nice to have
- Direct experience supporting PCI-DSS and SOC 2 compliance programs from an infrastructure perspective.
- Experience in the fintech, payments, or banking sectors, particularly in highly regulated environments.
- Experience deploying AI/ML tooling for production operations, such as AI-assisted incident response or automated runbooks.
- Experience with multi-region architectures, chaos engineering, and formal resilience programs.
- Proficiency with modern observability stacks (e.g., Prometheus, Grafana, Loki).
- Familiarity with service mesh, zero-trust networking, and secrets management.
- Experience with advanced CI/CD practices, progressive delivery, and platform engineering.
- Experience presenting to executive boards, auditors, or examiners.
Skills & tools
- AWS
- Kubernetes (EKS preferred)
- Aurora RDS (MySQL, Postgres)
- Terraform
- Observability stacks (Prometheus, Grafana, Loki, Tempo)
Practical notes
- Total compensation range: $400,000 - $600,000 per year, negotiable.
- Visa sponsorship is not available for this role.