Distributed Systems Engineer, Data & Inference Platform
Job description
About the role
You will design and operate the distributed systems that convert raw compute into actionable intelligence for real-world deployment. One week you are investigating a tail-latency regression in a production inference service handling millions of requests; the next you are re-architecting a Ray Data pipeline to prevent it from failing at petabyte scale. Your daily work spans architecture, implementation, and the on-call pager that holds you accountable for reliability and performance. Researchers and ML engineers will arrive with workloads that barely run on a single node, and you will refactor them into systems that run reliably, efficiently, and cheaply enough to matter at production scale. You will own the full lifecycle of these services, from initial design through sustained operation and eventual deprecation. The role demands comfort with ambiguity and the ability to translate ambiguous requirements into robust, observable infrastructure. You will be a systems partner embedded in the product and research workflow, ensuring that ambitious experiments can graduate to reliable services without losing speed or quality.
Key facts
What you'll do
- Serve Models at Scale by designing and operating distributed inference systems for LLMs, optimizing throughput, latency, and cost across heterogeneous GPU fleets while managing batching, scheduling, KV cache management, autoscaling, and the levers that make inference economical.
- Move the Data by building large-scale data pipelines with Ray Data, Spark, or equivalent frameworks to ingest, transform, and curate the datasets that power training and evaluation, identifying and eliminating bottlenecks that are not obvious on paper.
- Debug the Undebuggable by chasing failure modes that only appear under real production traffic, such as stragglers, head-of-line blocking, silent data corruption, and GPU memory fragmentation, and by writing postmortems that prevent the next ten incidents.
- Define SLOs, construct the observability stack needed to measure them, and own the on-call rotation that defends those service levels against regressions and outages.
- Partner Across the Stack by working directly with researchers and ML engineers to take experimental workloads that run on one node and evolve them into production-grade services that run reliably and efficiently.
- Own the data and model infrastructure that turns raw datasets and training runs into curated, high-quality assets, ensuring pipelines are maintainable, observable, and cost-effective.
- Evaluate and integrate emerging inference engines and framework features, such as LLM inference engines, modern lakehouse formats, and related open-source projects, to keep the platform competitive and efficient.
- Balance trade-offs between performance, reliability, and cost by making pragmatic engineering decisions that deliver durable solutions rather than temporary fixes.
- Drive operational excellence by automating manual steps, improving deployment pipelines, and reducing toil so the team can focus on high-impact work.
- Mentor and enable other engineers by sharing knowledge, documenting systems, and promoting best practices across the data and inference platforms.
Requirements
- 5+ years building and operating distributed systems in production environments, with a proven ability to handle incidents that span multiple systems and teams.
- Deep experience with at least one large-scale data or compute framework such as Ray, Spark, Flink, Beam, or Dask, including tuning, debugging, and capacity planning at scale.
- Strong fluency in Python and at least one systems programming language such as Go, Rust, or C++, allowing you to build performant services and understand low-level behavior.
- Working knowledge of the GPU and accelerator stack, including CUDA fundamentals, NCCL, mixed precision training and inference, and memory layout, enough to diagnose why a workload is bound by compute, memory, or communication.
- Experience operating Kubernetes-based infrastructure in production, including custom operators, controllers, or schedulers to manage specialized workloads.
- A track record of owning hard production incidents end-to-end, from initial diagnosis through mitigation to the durable fix and follow-up prevention.
- Comfort writing SQL and building queries against large datasets, as well as experience with data warehouse concepts and query performance optimization.
- Strong judgment around trade-offs between correctness, latency, throughput, and cost when designing systems that will run at scale.
- Excellent written and verbal communication, with the ability to translate technical constraints into clear recommendations for both technical and non-technical stakeholders.
- A collaborative mindset that seeks to unblock partners, share context, and build solutions that are maintainable as teams and requirements evolve.
Nice to have
- Hands-on experience with LLM inference engines such as vLLM, SGLang, TensorRT-LLM, or TGI.
- Practical experience with modern lakehouse formats like Iceberg, Delta, or Hudi, including performance tuning and schema evolution strategies.
- Contributions to open-source projects relevant to data platforms, distributed systems, or LLM serving stacks that demonstrate depth and community engagement.
Practical notes
This role is full-time and based in San Francisco. We offer an annual travel stipend through the Adaption Passport program, a weekly lunch stipend, comprehensive medical benefits, and generous paid time off.