Platform Engineer
Job description
About the role
You will architect and operate the cloud infrastructure that hosts millions of concurrent AI sandboxes, owning the core platform that enables agents to build software autonomously. Your work will directly underpin the execution environments for industry-leading AI agents from Groq, Manus, and Black Forest Labs, translating abstract agent workflows into deterministic compute resources. You will design the low-level orchestration logic that decides where every sandbox lives, ensuring optimal hardware utilization and isolation. This role requires you to own the entire lifecycle of performance-critical infrastructure, from kernel tuning through distributed systems coordination. You will implement the mechanisms that allow live migration of active sandboxes without disrupting running agent workflows. Your contributions will be visible in every interaction that starts a sandbox in under 200ms. You will help define the future of infrastructure for AI by building the control plane that manages billions of ephemeral compute instances.
Key facts
What you'll do
Design and implement a distributed systems platform capable of scheduling and managing millions of concurrent AI sandboxes with strict latency requirements.
Architect an intelligent orchestrator responsible for bin-packing sandbox workloads across heterogeneous physical nodes while respecting affinity and isolation constraints.
Implement live migration capabilities for running sandboxes, allowing workload movement across nodes with zero disruption to active agent sessions.
Champion the developer experience for self-hosting our platform, creating tools and documentation that lower the barrier for teams deploying their own private instances.
Enforce a hard startup latency budget of 200 milliseconds from user input to fully initialized sandbox execution.
Scale the platform architecture to support current demands and future growth toward billions of concurrent sandboxes running globally.
Build a comprehensive observability stack that exposes metrics and traces from the kernel and hypervisor layers upward to application-level signals.
Optimize critical paths using systems programming techniques to reduce overhead and improve throughput of network and block I/O operations.
Collaborate closely with product and research teams to prototype new virtualization features that unlock novel agent capabilities and performance gains.
Ensure the reliability and stability of the control plane through rigorous testing, chaos engineering, and deep analysis of failure modes in distributed scenarios.
Maintain and evolve the open source components of the platform, fostering a healthy community of users and contributors.
Partner with networking specialists to design secure multi-tenant network topologies that provide strong isolation between customer workloads.
Continuously profile production systems to identify bottlenecks in CPU, memory, and I/O subsystems under extreme load conditions.
Define and track service level objectives that reflect the real-world performance of AI agents running inside the sandboxes.
Requirements
Bring a minimum of 5 years of professional experience building distributed systems that operate at significant scale.
Have operated infrastructure handling over 100,000 requests per second across multiple regions while managing petabyte-scale data volumes.
Possess deep expertise in Linux internals, including the ability to debug performance issues using eBPF tools and understand CPU scheduling policies.
Understand memory management strategies at the operating system level and be able to explain the practical differences between cgroups version 1 and version 2 without referencing documentation.
Have hands-on experience with VM hypervisors such as Firecracker, QEMU, or KVM and comprehend the implications of nested virtualization.
Know what a hypercall is and can discuss the performance trade-offs between different virtualization architectures.
Demonstrate strong systems programming abilities in at least one systems-level language such as Go, Rust, C, or C++.
Have written performance-critical code where decisions about lock-free data structures, memory-mapped files, or io_uring directly impacted throughput or latency.
Have built or operated production-grade orchestration systems similar to Kubernetes or Nomad, including implementing custom schedulers.
Understand bin-packing algorithms and resource scheduling strategies used to solve the noisy neighbor problem in large-scale environments.
Have shaved milliseconds off hot code paths and used profiling to understand CPU cache behavior and memory locality.
Know what "p99 latency" means in practice and have actively worked to reduce tail latency in distributed services.
Possess a strong grasp of networking fundamentals including L4 and L7 load balancing techniques.
Understand how network namespaces, iptables, and nftables can be used to construct secure and isolated environments for multi-tenant scenarios.
Are physically located in San Francisco or are willing to relocate to the area to work in person with the team.
Are passionate about open source principles and are comfortable with the public nature of our code and infrastructure.
Contribute to technical discussions, write clear documentation, and help users who self-host the platform succeed.
Nice to have
Hands-on experience with userfaultfd (UFFD) and copy-on-write mechanisms for advanced memory optimization.
Knowledge of GPU passthrough or PCIe device virtualization techniques.
Background building or maintaining infrastructure tailored for AI or machine learning workloads.
Contributions to Firecracker, Cloud Hypervisor, or other related open source virtualization projects.
Experience designing observability solutions for systems at kernel-level scale using distributed tracing.
Practical notes
We operate as a hybrid team with four days per week in the office and one day remote. Our offices are located in San Francisco, USA, and Prague, Czech Republic. The role requires in-person collaboration a few days each week to solve complex problems together. We offer competitive benefits including healthcare, vision, and dental insurance, unlimited paid time off, a 401k plan, and additional perks for employees who work from our offices.