Principal Software Engineer, Enterprise Scalability
Job description
About the role
You will serve as Klaviyo's senior individual contributor for enterprise scalability, operating without direct reports but with significant ownership over platform integrity. You will own the definition and enforcement of enterprise scalability fitness functions, including latency, throughput, and error rate targets across all major workloads. You will hunt for and eliminate systemic bottlenecks by designing and implementing sharding, partitioning, caching, back-pressure, and multi-region readiness strategies. You will partner closely with cross-functional teams to productionize improvements and ensure architectural changes are scalable and maintainable. You will lead incident deep dives and translate findings into systemic fixes that prevent recurrence at scale. You will act as the primary technical voice for scalability with executives, translating complex tradeoffs into clear business impact. You will integrate AI into scale and resiliency workflows to enable proactive anomaly detection and guided runbooks that reduce repeat incidents.
Key facts
What you'll do
Define enterprise scalability fitness functions covering latency, throughput, and error rates and maintain a scorecard that aligns teams to shared SLOs and budgets.
Design and implement sharding and partitioning strategies that enable safe multi-tenant growth while preserving performance and correctness.
Architect and operate caching and back-pressure mechanisms that protect downstream services during traffic spikes and failure scenarios.
Drive multi-region readiness and high-volume migration paths to support global customers and reduce regional risk.
Build lightweight enablement artifacts such as benchmarks, profiling harnesses, and reproducible testbeds that teams can reuse.
Pair directly with engineering teams to land fixes and validate changes against realistic load patterns before production rollout.
Lead scalability reviews and readiness gates that accelerate delivery by providing clear criteria and fast feedback.
Communicate tradeoffs and technical outcomes clearly to both executives and engineers, tying work to customer outcomes and business impact.
Integrate AI into scale and resiliency workflows, using it for profiling assistance, workload modeling, synthetic load generation, and guided runbooks.
Champion a culture of performance and reliability by evangelizing observability, hypothesis-driven testing, and data-driven decision making.
Requirements
12+ years of experience scaling multi-tenant SaaS platforms with a demonstrated record of removing major bottlenecks.
Strong performance engineering background including capacity planning, profiling, and latency and throughput optimization at scale.
Hands-on experience with sharding, partitioning, caching, and back-pressure patterns in high-volume systems.
Deep understanding of multi-region architectures, data residency considerations, and strategies for safe migrations at scale.
Proven ability to design and execute high-volume data migrations with minimal disruption and rigorous validation.
Experience building and maintaining scalable infrastructure that supports rapid product experimentation and feature iteration.
Comfortable using AI tools to augment scale work, with strict attention to observability, guardrails, and reproducibility.
Strong cross-organizational influence skills, including the ability to align teams through fitness functions, scorecards, and readiness gates.
Excellent written and verbal communication skills for presenting complex technical tradeoffs to executives and engineers.
Curiosity and a proactive mindset toward learning and applying AI in responsible ways to improve scalability outcomes.
Nice to have
Company-wide adoption of scale scorecards with clearly defined fitness functions for latency, throughput, and error rates.
Documented high-impact wins where 2-3 major bottlenecks were removed using reproducible testbeds and improved pXX latencies and error rates.
Evidence of AI-assisted scale engineering, including anomaly detection that reduces alert noise and generative load testing used in release readiness checks.
Demonstrated reductions in time-to-isolate regressions, ideally in the range of 20-30% through tooling and process improvements.
Practical notes
This is a full-time position based in Boston, MA.
Employment is at-will.
We use Covey as part of our hiring and promotional process. For jobs or candidates in NYC, certain features may qualify it as an AEDT. As part of the evaluation process we provide Covey with job requirements and candidate submitted applications. We began using Covey Scout for Inbound on April 3, 2025.