
Staff Infrastructure Engineer
Job description
About the role
This role owns the cloud platform that every engineer at Headway deploys on, making deployments boring, scaling automatic, infrastructure self-serve, and cost attributable. You will architect and lead the redesign of deployment topology to isolate blast radius and prevent failures from propagating across services. You will serve as the technical anchor for compute, networking, and the Python runtime, bringing Staff-level influence to a dependency for every engineer. You will build a self-serve infrastructure platform with guardrails so teams own their standard changes while Reliability Engineering focuses on exceptions. You will own capacity planning for spiky workloads, using data to derive scaling floors ahead of demand rather than chasing them reactively. You will stand up cloud cost attribution across AWS, Datadog, and LLM spend to make infrastructure spending visible and intentional. You will lead Python runtime and dependency health, resolving event loop contention and garbage collection issues that limit performance under load.
Key facts
What you'll do
- Redesign deployment architecture to contain blast radius and ensure mistakes in one service cannot block or take down others.
- Drive per-service deploy isolation and functional-area slices that contain failures rather than allowing them to propagate across the platform.
- Own ECS and EKS footprint, evaluating broader EKS adoption for AI workloads and designing the next iteration of inter-service network connectivity.
- Own capacity and scaling for spiky workloads, building floors computed ahead of demand and enabling self-deriving signals from data.
- Build a Terraform-based self-serve infrastructure platform with guardrails so engineering teams can manage standard infrastructure changes independently.
- Stand up per-team cost attribution across AWS, Datadog, and LLM spend to make infrastructure costs visible and attributable for informed tradeoffs.
- Own the Python runtime and dependency health of the monolith, managing garbage collection, event loop contention, and runtime limits under load.
- Lead framework and package upgrades that most teams defer, raising the baseline for reliability and performance across all services.
- Make infrastructure changes safe from the start as AI increases code velocity, ensuring safety keeps pace with speed.
- Serve as the technical anchor for Headway's compute, networking, and deployment platform, influencing Staff-level decisions across engineering.
- Partner with Reliability Engineering to reduce routine work and focus on strategic, unusual problems that protect the platform.
- Enable other engineering teams through paved-road tooling, architecture reviews, and runbooks that raise the standard for shipping on the platform.
- Ensure observability and control across the production infrastructure so performance, cost, and risk remain transparent and manageable.
- Guide adoption of container orchestration patterns in ECS and EKS to support scale, resilience, and efficient resource use.
Requirements
- 8 or more years in platform, infrastructure, or SRE roles at companies running significant production traffic.
- Deep AWS expertise with production ownership of compute and networking at scale, including ECS, EKS, RDS, networking, and IAM.
- Strong infrastructure-as-code experience, particularly Terraform, including designing self-serve platforms for other engineering teams.
- Hands-on experience with autoscaling and capacity engineering, and container orchestration with ECS and/or EKS.
- A track record of making deploys safe and self-serve for other teams, not only for your own services.
- Staff-level technical leadership experience influencing architecture and alignment across multiple teams.
- Comfort operating in ambiguity while driving technical direction and building systems that get safer as they get faster.
- A commitment to raising the baseline for engineering productivity through tooling, runbooks, and architecture reviews.
Nice to have
- Experience building and operating platforms that serve as the foundation for mental healthcare at scale.
- Familiarity with regulated environments and compliance considerations relevant to health data and software.
- Background working with high-traffic SaaS products where infrastructure directly supports patient care.
Practical notes
- This is a full-time remote role.
- No travel is required.
- No visa sponsorship is available at this time.
- The role is open until filled, and applications will be reviewed on a rolling basis.