DevOps Engineer
Job description
About the role
Flux is taking the hard out of hardware. To do this, we're building the world's first AI hardware engineer. Founded in 2019, our platform enables anyone to go from idea to manufacturable board using nothing more than a natural language prompt. This democratizes a process that has historically required years of specialized expertise. We're going after a $15B+ electronic design automation market, backed by 8VC, Bain Capital Ventures, Liquid 2 Ventures, Outsiders Fund, Figma board member John Lilly, and GitHub founder Tom Preston-Werner. In February 2026, we closed a total of $37M in funding to accelerate that mission. What happens when we unlock the ability for anyone to make hardware? We fundamentally reshape the world. That's our mission, and we're just getting started.
As a DevOps Engineer, you will own the reliability and scalability of the infrastructure that powers Flux's complete platform experience. You are responsible for the systems that sit behind the editor, handling billing, authentication, onboarding flows, and critical integrations. You will ensure these production services remain highly available, performant, and secure for every user. You will define the operational standards that allow Flux to move fast without breaking things. Ultimately, your work directly enables makers and engineers to turn their ideas into real boards.
Key facts
What you'll do
- Improve the reliability, availability, and operational health of production systems across the entire Flux stack.
- Set and enforce observability standards for metrics, logs, and errors to provide clear insight into system behavior.
- Define SLOs and SLIs, design alerting policies, and maintain on-call readiness with an emphasis on signal quality and actionable responses.
- Partner closely with application engineers to design resilient architectures and reduce operational risk during the design phase.
- Build and maintain internal tooling that enhances system safety, simplifies debugging, and accelerates developer velocity.
- Manage infrastructure deployments and configurations using Pulumi across GCP, AWS, and Firebase environments.
- Investigate and lead incident response efforts, performing deep debugging across interconnected systems to resolve issues rapidly.
- Collaborate with product and platform teams to ensure new features are delivered with production-readiness as a core requirement.
- Optimize cloud costs and resource utilization while maintaining high standards for performance and reliability.
- Mentor junior engineers on best practices for operating distributed systems in cloud-native environments.
Requirements
- Bring 5 or more years of experience in SRE, DevOps, or production operations roles managing real-world systems.
- Demonstrate a proven track record of operating and scaling production systems while meeting strict uptime and latency goals.
- Show strong hands-on competence with observability tools and platforms such as Datadog, Sentry, or similar solutions.
- Have experience designing and implementing SLOs, SLIs, and effective alerting strategies that drive responsible on-call practices.
- Be proficient with modern CI/CD pipelines and infrastructure-as-code methodologies and tools.
- Possess deep familiarity with cloud-native architectures and serverless platforms, specifically GCP, AWS, and related managed services.
- Exhibit robust cross-system debugging capabilities and a disciplined approach to incident response and postmortems.
- Hold the ability to read, understand, and modify infrastructure and application code to collaborate effectively across teams.
- Maintain a strong commitment to security, compliance, and operational best practices in all deployed services.
Nice to have
- Hands-on experience with multi-product distributed tracing that stitches together data from Datadog, Sentry, and GCP tracing.
- Comfortable working with containerized and serverless workloads in environments such as Cloud Run, Firebase, and AWS Lambda.
- Experience building frontend interfaces using TypeScript, Node.js, and React, or the capacity and eagerness to learn these technologies quickly.
- Background supporting AI-powered systems or workloads that exhibit high variance and complex dependencies.
- Prior startup experience and a history of owning systems through periods of rapid growth and evolving requirements.
Practical notes
Work is full_time. The position is based in the San Francisco Office.