Principal Engineer, Core Infrastructure
Job description
About the role
At Klaviyo, we value the unique backgrounds, experiences and perspectives each Klaviyo brings to our workplace each and every day. We believe everyone deserves a fair shot at success and appreciate the experiences each person brings beyond the traditional job requirements. If you're a close but not exact match with the description, we hope you'll still consider applying. Want to learn more about life at Klaviyo? Visit klaviyo.com/careers to see how we empower creators to own their own destiny. As a hands-on principal for compute, networking, storage, runtimes (e.g., Kubernetes), CI/CD, and observability, you will architect the service platform that lets teams ship fast and safely. You will be an IC role with no direct reports, where you lead via design, code, and incident excellence, setting technical standards and SLOs for platform services. You will partner closely with product and engineering teams to ensure the platform scales reliably while enabling rapid experimentation. You will act as a force multiplier, removing roadblocks and enabling other engineers to deliver with confidence. Your work will directly influence the reliability, performance, and developer experience across the entire organization.
Key facts
What you'll do
Architect and evolve the Kubernetes platform, service mesh, networking, storage, and CI/CD pipelines; ship golden paths and IaC modules that standardize how teams build and deploy.
Define platform SLOs; use error budgets to guide reliability versus velocity trade-offs; drive incident learning and readiness reviews to improve future outcomes.
Improve developer velocity through measurable gains in build and deploy times, reduction of flaky tests, and enhancements to local development ergonomics.
Lead capacity planning and commitments; build guardrails for cost, security, and compliance in collaboration with Security and FinOps partners.
Write high-impact code, automation, and tooling; mentor across teams and raise the bar on operational excellence through clear standards and practices.
Embed AI into the developer experience - from code generation to observability and incident response - so teams ship faster and safer by default.
Establish platform design standards and review processes to ensure consistency, scalability, and resilience across services.
Own the reliability and performance of critical production systems, including conducting post-incident reviews and defining remediation plans.
Collaborate with product and infrastructure teams to enable new features with minimal friction and maximum stability for end users.
Participate in on-call rotations and respond to critical incidents, leading diagnosis and remediation efforts with clear communication.
Drive the adoption of best practices for security, compliance, and cost optimization across platform tooling and workflows.
Contribute to the broader technical community by documenting patterns, presenting learnings, and influencing engineering culture.
Evaluate emerging technologies and open-source projects to determine their applicability and value to the platform roadmap.
Partner with SRE, platform, and application teams to align on goals, metrics, and operational runbooks.
Requirements
10+ years of experience building and operating cloud platforms, with deep expertise in compute, networking, storage, and runtimes such as Kubernetes.
Demonstrated history of operating large-scale, multi-region, highly available services with rigorous SLO definition and management.
Strong proficiency in infrastructure-as-code tools such as Terraform, along with CI/CD pipelines and automation frameworks.
Solid understanding of databases and storage systems, including SQL and NoSQL databases, as well as object, block, or file storage platforms.
Hands-on experience with Kubernetes, service meshes, and production-grade networking and load balancing solutions.
Proven ability to write, maintain, and debug code in languages commonly used in infrastructure and platform engineering.
Experience with observability tools and practices, including metrics, tracing, and logging at scale.
Comfortable working in fast-paced, ambiguous environments and making decisions with incomplete information while communicating trade-offs clearly.
Strong written and verbal communication skills, with the ability to translate technical concepts into business terms for diverse audiences.
A collaborative mindset and a commitment to mentorship, enabling peers and junior engineers to grow through feedback and shared knowledge.
Nice to have
Core SLOs and velocity: Demonstrated success maintaining ≥99.95% SLOs for core services, with 25-50% faster build/deploy times and trending reductions in developer-reported friction.
AI-enabled platform: Experience integrating approved AI tooling into IDE and CI/CD with repository policies and auditability; evidence of ≥70% adoption among eligible engineers, MTTR reduced by 20-30% via AI-assisted triage, and lowered flaky-test rates through AI-suggested fixes.
Guardrails and governance: Established cost, security, and compliance controls codified as IaC modules and enforced through paved-road patterns.
Experience with enterprise governance, including compliance frameworks and audit readiness.
Familiarity with GDPR and data privacy considerations in large-scale, production environments.
Practical notes
This is a full-time position based in Boston, Massachusetts.
Employment eligibility requirements must be met, and sponsorship may be considered for qualified candidates.
The compensation range provided reflects local legal requirements and is subject to change based on business needs.
Klaviyo is an equal opportunity employer and encourages applicants of all backgrounds to apply.