Senior/Staff Platform Engineer
coderabbitUSAFull Time2w ago
TypeScriptGoNode.jsAWSGCPDockerKubernetesTerraformAILinuxSecuritySQL
Job description
Senior/Staff Platform Engineer at coderabbit.
About the role
Join coderabbit as a founding Platform Engineer, taking charge of the core infrastructure that powers our AI engine and product suite. This is a foundational role where you will establish the technical direction and architectural patterns for all future development. You will build critical systems from scratch, ensuring they are robust and scalable across a multi-region cloud environment.
Key facts
What you'll do
- Own the compute, orchestration, and infrastructure layers for coderabbit's AI engine and products.
- Define the technical strategy and architectural decisions for the platform.
- Build and operate large-scale distributed systems in a multi-region cloud environment.
- Establish best practices and patterns for other engineering teams to follow.
- Debug and resolve complex distributed failures in production systems.
Requirements
- At least 7 years of experience in Platform, Infrastructure, or Site Reliability Engineering, with a focus on building platforms.
- Extensive hands-on experience with Kubernetes, including control plane configuration, operator development, scheduler tuning, and runtime debugging.
- Proven ability to build and run large-scale distributed systems, understanding trade-offs and failure modes.
- Strong background in cloud compute on GCP or AWS, understanding underlying primitives beyond managed services.
- Deep proficiency in Docker and container runtime internals, including image layering, networking, and security.
Skills & tools
- Container & Orchestration: CRDs, operators, admission webhooks, RBAC, network policies, autoscaling, Docker/OCI toolchain
- Distributed Systems: Queuing, eventual consistency, backpressure, graceful degradation
- Infrastructure as Code: Advanced Terraform for module design, state management, programmatic provisioning
- Cloud Platforms: GCP (GKE, Cloud Run, VPC, IAM, Cloud SQL, Cloud Storage, Load Balancing)
- Programming: Node.js/TypeScript or Go for tooling and automation
- Observability: Datadog, Prometheus/Grafana, custom instrumentation, distributed tracing, SLO-based alerting
- Systems: Linux internals, networking fundamentals (TCP/IP, DNS, load balancing, eBPF), storage systems