Senior Platform Engineer
Job description
About the role
You will own the design and operation of WHOOP's Kubernetes clusters on AWS, shaping the core infrastructure that powers member insights. You will drive architectural decisions that enhance scalability, resiliency, performance, and security for the build and deployment platform. You will build systems and tooling that increase deployment safety and accelerate release velocity to Kubernetes across the organization. You will advance CI/CD capabilities to support frequent, reliable production deployments while optimizing the stability of the entire software delivery pipeline. You will partner with application, security, and data teams to embed secure-by-default infrastructure practices into everyday workflows. You will lead incident response, root cause analysis, and postmortems to continuously improve platform reliability and set the direction for the AppInfra team. You will mentor engineers and define the long-term roadmap for infrastructure and Kubernetes management across all WHOOP technology stacks.
Key facts
What you'll do
- Design, develop, and operate WHOOP's Kubernetes clusters running on AWS infrastructure while ensuring optimal performance and stability.
- Drive architectural decisions to improve scalability, resiliency, performance, and security across the build and deployment platform for large-scale systems.
- Build systems and tooling that increase deployment safety and accelerate release velocity to Kubernetes through automation and platform integrations.
- Advance CI/CD capabilities to support frequent, reliable production deployments with robust testing, monitoring, and rollback strategies.
- Lead developer productivity improvements through tooling, automation, and platform integrations that remove friction for engineering teams.
- Partner with application, security, and data teams to embed secure-by-default infrastructure practices and compliance into platform design.
- Participate in incident response, root cause analysis, and postmortems to continuously improve platform reliability and prevent future issues.
- Mentor and provide technical leadership to engineers on the Application Infrastructure team, fostering growth and knowledge sharing.
- Help define and execute the long-term roadmap for infrastructure and Kubernetes management across all WHOOP technology stacks.
- Collaborate cross-functionally to align infrastructure initiatives with product, security, and data objectives.
- Implement and operate control planes and supporting services that enable reliable and auditable software delivery.
- Evaluate emerging technologies and patterns to modernize infrastructure while maintaining operational excellence and security.
- Instrument systems to collect telemetry and metrics that drive data-informed infrastructure improvements.
- Champion best practices for infrastructure as code, versioning, and change management across the platform.
- Work closely with SRE and platform teams to standardize operational runbooks and improve observability.
- Contribute to open source and internal tools that enhance the developer experience and reduce operational burden.
- Ensure infrastructure components meet stringent security, audit, and availability requirements for a health-focused organization.
- Track and communicate platform metrics to stakeholders to guide investment and prioritization decisions.
Requirements
- 5+ years of experience in DevOps, Platform, Site Reliability, Cloud Engineering, or Backend Software Engineering roles with a strong track record of delivery.
- Deep understanding of Kubernetes architecture and core components including API server, etcd, controllers, and scheduler.
- Strong knowledge of container networking concepts, including overlay networking, service meshes, and network policies in production environments.
- Hands-on experience with multi-cluster Kubernetes environments and inter-cluster communication patterns for distributed systems.
- Hands-on experience operating cloud infrastructure, preferably in AWS (e.g., IAM, VPC, EC2, S3, RDS, CloudTrail, Organizations) for secure and scalable deployments.
- Hands-on experience with Infrastructure as Code tools (e.g. Terraform) to manage cloud resources and ensure reproducible environments.
- Experience developing backend or infrastructure-adjacent services using Java, C#, or Python to build reliable and performant platforms.
- Proven ability to evaluate system performance, identify bottlenecks, and use data to drive improvements in complex distributed systems.
- Experience collaborating with multiple stakeholders and prioritizing work for maximum business impact across engineering organizations.
- Strong understanding of security principles and practices as they apply to infrastructure, identity, and data protection.
- Demonstrated ability to work in fast-paced, high-scale production environments with minimal downtime and high reliability expectations.
- Excellent written and verbal communication skills for coordinating with cross-functional teams and documenting platform decisions.
- Comfort with command-line tools, scripting, and automation to streamline operations and reduce manual effort.
- Willingness to participate in on-call rotations and respond to incidents promptly with clear communication and remediation steps.
- Commitment to following established processes while advocating for improvements where they enhance safety and velocity.
- Openness to feedback and collaboration with engineers, security, and product teams to refine platform capabilities over time.
- Willingness to learn new technologies and adapt to evolving infrastructure landscapes in a growing health-focused company.
- Alignment with WHOOP's mission to unlock human performance and healthspan through secure and reliable systems.
Nice to have
- Experience operating Kafka or other large-scale distributed systems that require high throughput and low latency.
- Experience with Kubernetes security best practices, including RBAC, secrets management, and pod security standards for hardened clusters.
- Exposure to service reliability practices such as SLOs, SLIs, and error budgets to guide reliability investments.
- Prior experience supporting compliance or security-focused infrastructure initiatives in regulated environments.
- Familiarity with monitoring and observability stacks used to maintain platform health at scale.
- Background in supporting multi-tenant environments or platforms with diverse user needs.
- Experience contributing to or maintaining internal developer platforms or CI/CD frameworks.
Practical notes
This role is based in the WHOOP office located in Boston, MA. The successful candidate must be prepared to relocate if necessary to work out of the Boston, MA office.