Senior Infrastructure Engineer
Job description
About the Role
We are seeking a Senior Infrastructure Engineer to take ownership of the critical systems that power our global web platform. This role centers on driving reliability and performance for the infrastructure that serves millions of page views every hour. You will be responsible for the end-to-end health of our production environment, ensuring that tens of thousands of customer projects launch smoothly without interruption. The position requires a hands-on leader who builds robust, scalable foundations using cloud technologies and automation. You will partner closely with product and engineering teams to turn operational challenges into long-term strategic advantages. Success in this role means you keep the platform stable, secure, and observable so the business can move fast without compromising quality.
Key facts
- Application deadline: applications accepted on an ongoing basis until position is closed and filled
- This posting is for a new position.
What you'll do
- Own and evolve the cloud platform that Webflow's product and engineering teams depend on, including compute layer, EKS fleet, serverless infrastructure, networking, and cloud operations across AWS and GCP.
- Design, implement, and maintain infrastructure-as-code patterns and shared components that enable other teams to build quickly and safely.
- Architect and operate the networking layer that connects Webflow's services, ensuring reliability, security, and scalability in multi-cloud setups.
- Drive observability as the core operating principle, improving dashboards, alerts, and SLOs so issues are caught early and on-call signals remain actionable.
- Build and manage AI-powered automation for infrastructure tasks, covering policy-as-code, drift detection, and LLM-assisted runbook generation.
- Collaborate with security and compliance stakeholders to embed controls into infrastructure pipelines and runtime guardrails.
- Define and lead best practices for multi-region and multi-cloud architectures, including data residency, failover strategies, and latency optimization.
- Mentor engineers across the organization on cloud fundamentals, debugging techniques, and reliability principles.
- Participate in on-call rotations to respond to incidents and improve runbooks based on real-world events.
- Contribute to the growth and culture of an expanding international engineering organization.
Requirements
- Have a background as an infrastructure, site reliability engineer or cloud engineer with enthusiasm for automation and code, or a software engineer background with deep enthusiasm for cloud infrastructure and distributed systems.
- Bring 5+ years of experience owning and operating customer-facing cloud infrastructure where downtime is not acceptable.
- Demonstrate deep hands-on experience with AWS services and a strong, well-formed opinion on effective cloud operations.
- Show experience managing Kubernetes clusters at scale, including upgrades, node group management, autoscaling, and add-on lifecycle.
- Have fluency with infrastructure-as-code tools such as Pulumi or Terraform, preferring changes driven by code over console-driven workflows.
- Have navigated multi-region or multi-cloud environments on AWS or GCP, understanding tradeoffs in latency, resilience, and cost.
- Stay curious and open to growth, actively embracing AI tools and building fluency with emerging technologies to improve how work is done.
- Work effectively in a fast-paced environment, balancing delivery speed with operational rigor and attention to detail.
- Communicate clearly and constructively with both technical and non-technical stakeholders across distributed teams.
- Take ownership of your work, following through on commitments and driving outcomes end to end.
Nice to have
- Experience with Karpenter, cluster autoscaler, or other Kubernetes-native scaling tooling.
- Hands-on experience with OpenTelemetry, Datadog, Prometheus, or Grafana for observability.
- Background building AI-assisted infrastructure tooling, including cost optimization loops, anomaly detection, or policy-as-code enhanced by LLMs.
- Contributions to multi-region architectures involving data residency, regional failover, or latency-based routing.
Practical notes
Applications are accepted on an ongoing basis until the position is closed and filled. This is a remote role based in Argentina. You will need valid right to work authorization depending on the country of employment. If extended an offer, you may be required to pass a background check.