SRE/Infrastructure Engineer
Job description
About the role
You will own the Terraform, Kubernetes, and cloud plumbing that lets E2B run millions of sandboxes.
You will be the reliability and deployment expert who ensures our infrastructure can scale to millions of sandboxes while remaining observable and recoverable.
You will partner closely with enterprise customers during pre-sales and deployments, translating their requirements into robust, repeatable infrastructure patterns.
You will extend our BYOC offering to public cloud marketplaces, making the deployment experience smooth and consistent across AWS, Google Cloud, and Azure.
You will build and maintain reusable Terraform components for networking, IAM, and secrets, establishing standards that reduce toil and prevent configuration drift.
You will wire up observability and monitoring, tightening the feedback loop between infrastructure changes and production behavior to drive quick improvements.
You will be the primary on-call owner for infrastructure incidents, investigating root causes and implementing safeguards to prevent recurrence.
You will work hands-on with modern AI tooling to accelerate development, testing, and debugging of infrastructure pipelines.
Key facts
What you'll do
- Design and maintain reusable Terraform modules for GCP, AWS, and Azure, focusing on VPCs, subnets, security groups, and IAM.
- Drive the migration from Nomad to Kubernetes, including cluster lifecycle, upgrades, and multi-cluster management.
- Implement Kubernetes operators and controllers to automate the lifecycle of E2B sandbox workloads.
- Build and refine observability dashboards, alerts, and runbooks to surface infrastructure health and performance in near real time.
- Own the end-to-end reliability of BYOC and self-hosted deployments, ensuring fast, repeatable provisioning for large customers.
- Conduct technical pre-sales discussions, translating customer requirements into infrastructure designs and deployment plans.
- Partner with enterprise customers during rollout and ongoing operations, serving as the main point of contact for infrastructure-related issues.
- Collaborate with the Platform team to align infrastructure standards with product needs, avoiding custom per-customer forks.
- Optimize deployment pipelines to reduce lead time and increase confidence in infrastructure changes across multiple clouds.
- Contribute to open-source infrastructure projects when aligned with product goals, such as Terraform providers and Kubernetes integrations.
Requirements
- 5+ years operating production cloud infrastructure, with demonstrated experience running services at scale.
- You've owned Terraform and at least one orchestrator such as Kubernetes, Nomad, or ECS in production environments.
- You have built Terraform modules, managed remote state, debugged configuration drift, and resolved production incidents.
- Hands-on Kubernetes experience at meaningful scale, including real workloads, real traffic, and real on-call responsibilities.
- You can speak knowledgeably about ingress, RBAC, and autoscaling in Kubernetes deployments.
- Strong Linux fundamentals, with deep comfort in networking, namespaces, cgroups, systemd, and filesystems.
- You are able to follow low-level debugging scenarios where the root cause lies below the orchestrator.
- Comfortable working directly with customers, including pre-sales technical discussions, deployment, and operational support.
- Multi-cloud comfort with expert depth in at least one major cloud provider, such as Google Cloud, AWS, or Azure.
- You have configured VPCs, managed IAM, set up private networking, and debugged routing and DNS issues in production.
- Terraform is a first-class skill for you, and you have opinions on state management, workspaces, and module boundaries.
- You read and write code outside of YAML, including Go and Terraform, and are not blocked when debugging across configuration and source.
- You have startup experience where ambiguity is common and you operate effectively with a "figure it out" mindset.
- You are comfortable proposing and driving technical decisions without needing a formal ticket to begin work.
- You have worked in high-scale environments, either with high-traffic online products or infrastructure-product companies where infrastructure is the core product.
- You are excited to work in person from San Francisco on a DevTool product and thrive in a close-knit, collaborative team setting.
Nice to have
- Deep cloud expertise in AWS, GCP, or Azure, with production proof of running critical workloads.
- Experience building self-hosted or BYOC deployments for enterprise customers, including marketplace integrations.
- Contributions to open-source infrastructure projects such as Terraform providers, Kubernetes operators, or related tools.
- Experience with Cloudflare Workers or similar edge compute platforms in an infrastructure or observability context.
- Familiarity with marketplace deployment patterns for AWS, Google Cloud, or Azure.
- Contributions to open-source infrastructure projects (Terraform providers, Kubernetes operators, or similar).
Practical notes
Work is full-time and based in San Francisco.
The role expects in-person presence in San Francisco to maintain fast, collaborative execution.
There are no specific visa sponsorship details, hours, or travel requirements stated in the source material.
Apply only if you meet the stated requirements and are prepared to own infrastructure end-to-end for enterprise-facing deployments.