Senior Staff Infrastructure Engineer - Kubernetes Platform
tensorwaveRemoteFull Time3w ago
KubernetesCI/CDAILinuxLegalOperationsSupportRecruitingGrowthStrategyEngineeringInfrastructure
Job description
Senior Staff Infrastructure Engineer - Kubernetes Platform at tensorwave.
About the role
Join our team as a Staff Infrastructure Engineer focused on our Kubernetes Platform. This role involves leading the design, development, and operational stability of our Kubernetes control plane architecture. You will work with various teams to meet business goals while maintaining high standards for quality and collaboration.
Key facts
What you'll do
- Design and evolve the Kubernetes control plane architecture across different regions.
- Implement multi-tenant cluster models, including shared control planes and virtual clusters.
- Transition from standalone clusters to a regionally managed platform.
- Define standards for isolation, resource segmentation, and policy enforcement.
- Ensure the reliability and behavior of Kubernetes platforms in production environments.
- Participate in on-call rotations and lead incident response efforts.
- Diagnose and resolve issues like control plane instability, API server saturation, and resource contention.
- Manage consistent lifecycle for clusters, including provisioning, upgrades, and scaling.
- Design and implement strategies for regional scaling and multi-data center cluster deployments.
- Ensure consistent behavior and reliability across all environments.
- Define cluster topology and failure domain strategies.
- Design ingress and egress architectures at both cluster and regional levels.
- Troubleshoot and optimize pod-to-pod networking, north-south traffic, and CNI behavior.
- Collaborate with network engineering on high-performance networking integration.
- Improve observability for control plane components, cluster health, and performance.
- Define and implement resilience strategies aligned with platform goals.
- Lead root cause analysis for production incidents.
- Work with DevOps engineers on automation and CI/CD, and with infrastructure teams on compute, storage, and networking.
- Align Kubernetes platform design with underlying infrastructure capabilities.
Requirements
- 7 or more years of experience in infrastructure, platform engineering, or distributed systems.
- Extensive experience operating Kubernetes at scale in production settings.
- Experience in cloud service provider, hyperscale, or similar large-scale environments is highly preferred.
- Proven experience scaling Kubernetes across multiple clusters and multiple regions or data centers.
- Strong understanding of Kubernetes internals, including the API server, scheduler, controller manager, and etcd.
- Experience designing or evolving control plane architectures and multi-tenant cluster models.
- Strong Linux systems expertise.
- Deep troubleshooting ability across Kubernetes, container runtime, and the networking stack.
- Experience with CNI plugins, with a preference for Cilium.
- Strong understanding of networking and traffic patterns, and resource isolation and scheduling.
Nice to have
- Experience with virtual cluster technologies like vcluster or Kamaji.
- Experience supporting GPU workloads in Kubernetes.
- Familiarity with NUMA-aware scheduling and topology-aware workloads.
- Awareness of RDMA and high-throughput networking environments.
- Experience with observability platforms such as Prometheus and Grafana.
Skills & tools
- Kubernetes
- Linux
- Cilium (preferred)
- vcluster (nice to have)
- Kamaji (nice to have)
- Prometheus (nice to have)
- Grafana (nice to have)
Practical notes
Applicants must be authorized to work in the United States. Background checks may be required. TensorWave provides reasonable accommodations for the hiring process; contact accomodations@tensorwave.com for assistance.