Senior Infrastructure Engineer
WebflowRemote (Argentina)3w ago
EngineeringInfrastructureremotecurated-jd
Job description
Senior Infrastructure Engineer at Webflow
About the role
Webflow seeks a Senior Infrastructure Engineer to enhance the stability and performance of our customer-facing production systems, which handle millions of page views hourly. You will ensure our platform remains secure and scalable for our global user base.
Key facts
What you'll do
- Manage and advance the cloud platform supporting Webflow's products and engineering teams, including compute, EKS, serverless functions, networking, and operational aspects across AWS and GCP.
- Develop and maintain our infrastructure-as-code foundation, creating reusable patterns and components for other teams.
- Design and oversee the networking infrastructure connecting our services, ensuring reliability, security, and scalability across cloud environments.
- Enhance observability by improving dashboards, alerts, and SLOs for owned infrastructure, enabling proactive issue detection and actionable on-call responses.
- Build and maintain AI-driven automation for cloud infrastructure management, such as policy-as-code, drift detection, and LLM-assisted runbook generation.
- Contribute to shaping the culture of a growing, international team.
Requirements
- Minimum of 5 years of experience managing and operating cloud infrastructure in a customer-facing, low-downtime environment.
- Extensive hands-on experience with AWS and a strong understanding of effective cloud operations.
- Proven experience managing Kubernetes clusters at scale, including upgrades, node management, autoscaling, and add-on lifecycles.
- Experience with infrastructure-as-code tools like Pulumi or Terraform, with a preference for code-driven changes over manual console operations.
- Experience working in multi-region or multi-cloud setups on AWS or GCP.
- A curious and growth-oriented mindset, with a proactive approach to AI and emerging technologies.
Nice to have
- Experience with Karpenter, cluster autoscaler, or other Kubernetes-native scaling tools.
- Familiarity with OpenTelemetry, Datadog, Prometheus, or Grafana.
- Experience building AI-assisted infrastructure tools, including cost optimization, anomaly detection, or LLM-enhanced policy-as-code.
- Experience with multi-region architecture, including data residency, regional failover, or latency-based routing.
Skills & tools
- AWS
- GCP
- Kubernetes (EKS)
- Pulumi
- Terraform
- Serverless technologies
- Networking
- Observability tools (e.g., Datadog, Prometheus, Grafana, OpenTelemetry)
- AI/ML for infrastructure automation
Practical notes
- Applications are accepted on an ongoing basis until the position is filled.
- Authorization to work in the country of employment is required.