
Staff/Principal DevOps Engineer, AI Inference
Job description
About the role
Platform for AI Inference at Scale
The role centers on owning the platform that enables fast, reliable model serving while protecting cost control. You will design paths from intake through build, review, and ship to partners, with decisions shaping how teams move models through production. The work spans platform engineering, site reliability, and ML infrastructure, focusing on systems that deliver low-latency, high-throughput inference across GPU clusters and cloud accelerators. You will partner closely with ML engineers, research scientists, and software engineers to build platforms that serve models reliably to production users while maximizing compute efficiency.
What you will build includes GPU and accelerator infrastructure on Kubernetes, with work on scheduling, resource isolation, multi-tenant GPU sharing, device plugins, and topology-aware placement for inference workloads. You will design model serving platforms using frameworks such as vLLM, Triton Inference Server, TGI, or custom serving stacks, optimizing batching, caching, and request routing. Intelligent request routing and load balancing across heterogeneous accelerator fleets, including NVIDIA GPUs and AWS Inferentia/Trainium, will be central to maximizing utilization and minimizing latency. You will define autoscaling systems that dynamically match inference compute supply with demand across production, research, and experimental workloads.
Production-grade deployment pipelines for ML models will be another focus, enabling canary rollouts, A/B testing, model versioning, and safe rollback across multi-region deployments. Infrastructure-as-code for GPU-accelerated EKS clusters will be implemented using Terraform and Helm, covering node pools, spot and on-demand strategies, and accelerator-specific networking. Observability and performance optimization will include GPU utilization monitoring, inference latency profiling, token throughput dashboards, and SLO and SLI tracking for model endpoints. You will own CI/CD pipelines for model artifacts, managing container image builds with CUDA and driver dependencies, model registry integration, and automated inference benchmarking in CI. AWS cloud infrastructure for ML will be orchestrated by you, including EKS with GPU node groups, EC2 accelerated instances such as P4, Inf2, and Trn1, S3 model storage, EFA and high-bandwidth networking, and IAM with least privilege. Cost optimization and capacity planning will involve right-sizing accelerator instances, spot instance strategies for inference, and fleet-wide efficiency reporting.
To succeed, you must have operated GPU or accelerator infrastructure at scale in production environments. Deep experience with Kubernetes for ML, including device scheduling, resource quotas, node affinity, and accelerator device management, is required. You must deploy to AWS using infrastructure-as-code, with hands-on experience managing GPU-based compute on EKS and GPU-heavy EC2 using proven workflows. Experience with model serving infrastructure, including inference servers, request batching, KV-cache optimization, or LLM serving frameworks, is essential. You must also have a strong understanding of networking for distributed inference, including high-bandwidth interconnects, NCCL, VPC PrivateLink, and load balancing at L4 and L7. Strong proficiency in Python for automation, tooling, and integration with ML frameworks is required.
Bonus points are awarded for experience with LLM inference optimization tactics such as continuous batching, speculative decoding, and quantization methods like GPTQ and AWQ, along with tensor parallelism and pipeline parallelism. Hands-on experience with multiple accelerator families, including NVIDIA A100/H100, AWS Inferentia2, Trainium, and AMD MI300X, and maintaining hardware-agnostic serving infrastructure is valued. Multi-region deployment experience with geographic routing and failover for latency-sensitive inference endpoints will be an advantage. Fluency in Rust or Go for performance-critical infrastructure components is a plus, as is SRE practice for ML systems, including chaos engineering on GPU workloads, incident management, and capacity modeling for bursty inference traffic. Experience with model registries, artifact versioning, and ML supply chain security will strengthen your profile, along with observability work focused on token-level throughput, time-to-first-token, and per-request GPU memory profiling. Background in fast-growth settings balancing velocity with reliability in rapidly scaling AI systems is also recognized as valuable.
The position is based in Cambridge, MA USA and is full-time. Compensation includes a base of $180,000 and a $60,000 bonus, along with equity. Key tools and skills include Kubernetes, AWS, Terraform, Helm, vLLM, Triton Inference Server, TGI, NVIDIA GPUs, AWS Inferentia, Trainium, EKS, EC2, S3, EFA, and CI/CD. Please What you'll do
- Meet the bar Practical notes