General Application
Job description
About the role
You own the design and delivery of the platform that powers the fastest and most cost efficient CI infrastructure in the industry. You will translate demanding product requirements into robust, observable services that operate at the edge of thousands of bare metal machines. You will collaborate closely with infrastructure and product teams to iterate quickly while maintaining high standards for reliability and performance. You are responsible for writing production grade code that is tested, documented, and maintainable over the long term. You own the full lifecycle of your work from initial design through deployment, monitoring, and iterative improvement. You proactively identify risks and constraints, communicate tradeoffs, and drive alignment across cross functional partners. You contribute to a culture of ownership where your work directly impacts the experience of thousands of engineers and the reliability of critical CI pipelines.
Key facts
What you'll do
Architect and implement services that orchestrate the lifecycle of tens of millions of ephemeral Firecracker microVMs each month across multiple global regions.
Design and extend the control plane that schedules, monitors, and scales our bare metal fleet to consistently support 100k+ vCPUs concurrently while meeting strict SLAs.
Build and refine the storage layer that manages petabyte scale data on our self managed Ceph cluster, ensuring durability, performance, and efficient capacity planning.
Partner with product and GTM teams to translate CI workload patterns into reliable primitives that enable faster and cheaper GitHub Actions execution for 3000+ companies.
Instrument systems end to end to surface actionable metrics and traces, driving decisions that improve utilization, latency, and cost efficiency at scale.
Implement automation for deployment, upgrades, and recovery of infrastructure components across heterogeneous bare metal environments with minimal manual intervention.
Collaborate with security and reliability teams to harden the platform, respond to incidents, and evolve runbooks for complex distributed systems.
Evaluate and integrate new technologies to evolve the platform, balancing innovation against operational burden and long term maintainability.
Own debugging and postmortem processes for complex failures, turning observations into durable improvements in code, tests, and documentation.
Mentor engineers by clearly articulating design decisions, code patterns, and operational practices that raise the standard of engineering excellence across the team.
Contribute to roadmap prioritization by analyzing capacity, tradeoffs, and dependencies to deliver meaningful progress against strategic objectives.
Work closely with cross functional stakeholders to align technical constraints with business goals while maintaining a high bar for quality and user experience.
Translate ambiguous problems into concrete technical tasks, define interfaces, and deliver incremental value that unblocks downstream teams.
Champion best practices in testing, observability, and maintainability to ensure that the platform remains robust as it grows in scale and complexity.
Requirements
You have a Bachelor's degree or equivalent practical experience in Computer Science, Distributed Systems, or a related technical field.
You are experienced with building or operating distributed systems that manage large numbers of processes or virtualized workloads.
You are comfortable managing infrastructure at scale, including provisioning, networking, and operating bare metal or virtualized environments.
You have hands on experience with container or VM orchestration, scheduling, and resource management concepts and tools.
You have strong programming skills in at least one systems level language such as Go, Rust, or C++, and you write safe, testable, and performant code.
You understand storage systems and have worked with technologies such as Ceph or similar distributed storage platforms.
You are fluent with Linux environments, including debugging, performance analysis, and tuning at the system and network level.
You are comfortable using and contributing to open source projects that power infrastructure and CI/CD tooling.
Nice to have
Experience with Firecracker or similar microVM technologies is a clear advantage for this role.
Deep experience operating Ceph or other distributed storage systems in production.
Contributions to open source projects related to virtualization, orchestration, or CI/CD tooling.
Background in security or compliance sensitive environments where auditability and reliability are critical.
Practical notes
Hours are full time, with flexibility for deep work blocks and asynchronous collaboration across regions.
Travel is not required, though occasional in person collaboration may be considered for team offsites.
We sponsor visas for qualified candidates and support remote work where policy and role eligibility allow.
New roles are reviewed on a rolling basis, so early application is encouraged.