Senior Software Engineer, Compute Platform
Job description
About the role
You will own the primary compute orchestration infrastructure that powers every service running at Coinbase, shaping the platform that underpins global economic freedom. This role centers on designing and shipping the core tooling, automation, and net-new capabilities that keep the platform reliable, scalable, and efficient at every scale. You will partner deeply with Security, Reliability, and Observability teams to elevate standards across the entire stack while integrating AI-driven processes into developer workflows. Expect an intense environment where you are pushed to solve hard infrastructure problems and remove operational toil. You will be responsible for making Kubernetes and related CNCF technologies sing across hundreds of engineers and countless workloads. Coinbase is a remote-first, but not remote-only company, requiring quarterly in-person "surges" for deep collaboration. If you thrive in high-bar settings and refuse to settle for good enough, this is your opportunity to build the future of finance and infrastructure.
Key facts
What you'll do
- Own the design, build, and operation of Kubernetes cluster management tooling and automation that keeps our compute platform reliable and self-healing at scale.
- Build developer-facing tooling and workflows that improve how engineers across Coinbase interact with Kubernetes, with a heavy emphasis on integrating AI-driven processes and support.
- Deliver net-new compute capabilities for service owners, such as one-off jobs, cron scheduling, deployment strategies, EFS support, and automated right-sizing.
- Drive operational excellence by automating toil, reducing on-call burden, and continuously improving platform observability and incident response.
- Partner with Security, Reliability, and Observability teams to ensure the compute platform meets Coinbase's standards for security, uptime, and performance.
- Utilize generative AI responsibly, maintaining human oversight to deliver business-ready outputs and drive measurable improvements in workflow efficiency, cost, and quality.
- Apply AI tooling to infrastructure workflows, improving automation, developer productivity, or operational efficiency in production-critical systems.
- Leverage the CNCF ecosystem, including Helm, Prometheus, ArgoCD, and Envoy, to solve real infrastructure problems and advance platform capabilities.
- Diagnose complex distributed system failures and drive them to root-cause resolution across compute, networking, and storage layers.
- Implement and maintain deployment strategies, monitoring dashboards, and reliability practices that raise the bar for service ownership.
- Collaborate closely with product and infrastructure teams to translate business requirements into robust, scalable platform solutions.
- Contribute to architectural decisions that balance trade-offs between scalability, resilience, developer experience, and operational cost.
Requirements
- 5+ years of software engineering experience, including 3+ years building and operating Kubernetes or similar compute orchestration systems (e.g., Mesos, Nomad, ECS).
- Hands-on experience with AWS and/or GCP infrastructure services (e.g., EC2, EKS, IAM, VPC, networking) in a production environment at scale.
- Demonstrated ability to design, implement, and operate distributed infrastructure systems, including diagnosing complex failures and driving them to root-cause resolution.
- Hands-on experience with the CNCF ecosystem (e.g., Helm, Prometheus, ArgoCD, Envoy) and a track record of applying these tools to solve real infrastructure problems.
- Proven ability to apply AI tooling to infrastructure workflows, improving automation, developer productivity, or operational efficiency.
- Utilizes generative AI responsibly, maintaining human oversight to deliver business-ready outputs and drive measurable improvements in workflow efficiency, cost, and quality.
- Strong proficiency in Linux, networking, and distributed systems concepts.
- Experience with infrastructure-as-code tools and version-controlled deployment pipelines.
Nice to have
- Experience contributing to open source projects related to Kubernetes, observability, or infrastructure tooling.
- Familiarity with financial services environments and their unique security, compliance, and reliability requirements.
- Background in performance optimization, capacity planning, or large-scale distributed tracing.
Practical notes
- Location is Remote
- Canada, with no mandatory relocation required.
- Engagement is Full-time.
- Coinbase is a remote-first, but not remote-only company, requiring quarterly in-person working sessions called "surges."
- Candidates may submit a maximum of 3 applications within a 6-month period.