Senior Platform Engineer
Job description
About the role
In this role you will own the design and operation of the AI Devex platform that underpins developer productivity across WHOOP. You will architect and run Kubernetes clusters on AWS that serve as the foundation for our model-agnostic AI runtime environment. You will collaborate closely with other Platform teams to define standards that enhance scalability, resiliency, and security for agentic workflows. You will build tooling that enables long running adhoc tasks to execute efficiently and cost effectively. Your work will directly accelerate engineering throughput by automating critical portions of the development lifecycle. You will set the technical direction for the team while mentoring engineers across the Platform organization. Finally you will partner with application, security, and data teams to embed secure-by-default practices into every layer of our infrastructure.
Key facts
What you'll do
Architect and deploy Kubernetes clusters on AWS infrastructure to support AI driven development tools at scale.
Drive architectural designs that improve scalability, resiliency, performance, and security of our agent runtime environment.
Build systems and tooling that increase our capability to run long running adhoc tasks on cost efficient infrastructure.
Advance WHOOP's ability to automate portions of the development lifecycle through platform automation and integration.
Lead developer productivity improvements by creating tooling, automation, and platform integrations that remove friction for engineers.
Partner with application, security, and data teams to embed secure-by-default infrastructure practices across the software development lifecycle.
Participate in incident response, root cause analysis, and postmortems to continuously improve platform reliability and uptime.
Mentor and provide technical leadership to engineers across the Platform organization to elevate overall capability.
Help define and execute the long-term roadmap for the AI Devex team in alignment with business and technology goals.
Champion best practices for cloud infrastructure using Infrastructure as Code tools to ensure repeatability and compliance.
Optimize resource utilization and cost efficiency for AI workloads running in our cloud environment.
Collaborate with cross functional stakeholders to translate business requirements into robust platform solutions.
Contribute to open source and internal projects that enhance the developer experience and accelerate innovation.
Act as a subject matter expert for Kubernetes, AI runtimes, and cloud infrastructure within the organization.
Requirements
5+ years of experience in DevOps, Platform, Site Reliability, Cloud Engineering, or Backend Software Engineering roles.
Deep understanding of Kubernetes architecture and core components including pods, services, deployments, and networking.
Good knowledge of AI runtimes and their environmental requirements including resource constraints and scheduling needs.
Hands on experience operating cloud infrastructure preferably in AWS and managing highly available systems.
Hands on experience with Infrastructure as Code tools such as Terraform for provisioning and managing cloud resources.
Proven ability to evaluate system performance identify bottlenecks and use data to drive improvements and inform decisions.
Experience collaborating with multiple stakeholders and prioritizing work to maximize business impact and delivery.
Ability to work independently and as part of a cross functional team in a fast paced environment.
Strong written and verbal communication skills for documenting designs and sharing knowledge.
Commitment to following security and compliance guidelines to protect sensitive data and systems.
Willingness to participate in on call rotations and respond to platform incidents as needed.
Dedication to continuous learning and improvement in emerging technologies related to AI and cloud platforms.
Capability to mentor junior engineers and provide technical guidance across the Platform organization.
Willingness to relocate if necessary to work out of the Boston MA office.
Nice to have
Experience with agentic workflows and building tools that support AI assisted development.
Familiarity with model serving platforms and GPU accelerated infrastructure for AI workloads.
Contributions to the open source community demonstrating real world use of Kubernetes and cloud native patterns.
Background in developer productivity tooling or internal platform team experience.
Practical notes
This role is based in the WHOOP office located in Boston, MA. The successful candidate must be prepared to relocate if necessary to work out of the Boston, MA office.
Candidates must be authorized to work in the United States and not require work authorization sponsorship at the time of hire.
Employment is contingent upon successful completion of a background check.
WHOOP participates in E-Verify https://www.e-verify.gov/ to determine employment eligibility. It is unlawful in Massachusetts to require or administer a lie detector test as a condition of employment or continued employment. An employer who violates this law shall be subject to criminal penalties and civil liability.