
Member of Technical Staff
Job description
About the role
Perplexity is looking for a dedicated platform engineer to take ownership of its GPU infrastructure supporting AI inference workloads. This role involves managing a large GPU fleet spread across multiple cloud providers, building and maintaining a self-serve compute platform that enables inference engineers and researchers to launch training and inference jobs efficiently. You will be responsible for designing systems that abstract away the complexity of GPU provisioning, cluster configuration, and provider-specific infrastructure, allowing teams to focus on model development and deployment. Your work will ensure reliable operation of both long-running distributed training jobs and high-availability inference services, supporting the company's mission to serve hundreds of millions of queries a month with real-time AI inference.
Key facts
What you'll do
- Design, develop, and maintain a self-serve compute platform that simplifies launching training and inference workloads, removing the need for manual GPU provisioning and cluster configuration.
- Manage the GPU fleet across multiple cloud providers, including provisioning, lifecycle management, capacity planning, and ensuring reliability.
- Build scheduling and placement algorithms that efficiently find available GPU capacity across providers, optimize hardware utilization, and handle real-world constraints such as capacity limits and workload priorities.
- Support diverse workloads by maintaining the health of long-running distributed training jobs while guaranteeing the availability and low latency of production inference services on the same infrastructure.
- Develop and manage Kubernetes operators and Custom Resource Definitions (CRDs) to orchestrate GPU resources across multiple clusters and cloud providers, ensuring consistent behavior everywhere.
- Implement fault-tolerance mechanisms, autoscaling policies, and observability features that keep the GPU fleet utilized efficiently and allow workloads to survive node failures, provider hiccups, or capacity shifts without manual intervention.
- Collaborate with inference and cloud infrastructure teams to define technical strategies, turn operational constraints into scalable platform architecture, and develop a roadmap for future enhancements.
- Own end-to-end operational issues, troubleshooting problems related to GPU provisioning, network performance, workload scheduling, and infrastructure reliability.
- Ensure the platform handles provider-specific differences seamlessly, maintaining a unified experience for inference engineers regardless of where compute resources are located.
- Contribute to the development of infrastructure that supports hundreds of millions of queries per month, focusing on efficiency, scalability, and robustness.
- Write systems-level code in languages such as Go, Rust, or C++ to implement core platform components, operators, and automation tools.
- Support the integration of high-performance networking technologies like InfiniBand or RoCE to ensure fast GPU communication and data transfer.
- Participate in defining best practices for GPU cluster management, workload scheduling, and infrastructure automation to improve reliability and developer productivity.
- Stay current with advancements in distributed systems, Kubernetes, and GPU hardware to continuously improve platform capabilities.
- Document platform architecture, operational procedures, and best practices to facilitate team onboarding and knowledge sharing.
- Engage in cross-team communication to align platform development with research and inference engineering needs, ensuring the infrastructure evolves to meet future demands.
Requirements
- Extensive experience with Kubernetes, including developing custom operators, CRDs, and managing multi-cluster federation.
- Proven track record managing large-scale GPU clusters with NVIDIA hardware, CUDA, and associated networking technologies such as InfiniBand or RoCE.
- Experience orchestrating compute workloads across multiple cloud providers such as AWS, GCP, or similar, understanding the differences and challenges involved.
- Strong understanding of distributed systems fundamentals, including scheduling algorithms, resource allocation, fault tolerance, and high-availability design.
- Proficiency in systems programming languages such as Go, Rust, or C++, with experience developing infrastructure components.
- Experience supporting both long-running distributed training jobs and high-availability inference services, understanding their differing infrastructure requirements.
- Ability to own complex problems from end to end, including troubleshooting, designing solutions, and implementing improvements without predefined solutions.
- Familiarity with GPU hardware management, including NVIDIA hardware, CUDA programming, and GPU networking.
- Knowledge of container orchestration, cloud infrastructure automation, and infrastructure-as-code practices.
- Strong communication skills to collaborate effectively with cross-functional teams and document technical solutions clearly.
- Experience with high-performance networking technologies and their integration into compute clusters.
- Ability to work in a fast-paced environment, balancing reliability, scalability, and operational efficiency.
Nice to have
- Experience with inference serving stacks such as vLLM, SGLang, or TensorRT-LLM.
- Familiarity with HPC schedulers like Slurm or similar workload managers.
- GPU kernel development experience in CUDA or Triton, although not strictly required.
- Knowledge of high-speed interconnects such as InfiniBand, RoCE, or RDMA in production environments.
- Experience with observability tools like Prometheus, Grafana, or Weights & Biases to monitor and troubleshoot ML workloads.
- Understanding of ITAR compliance and handling sensitive infrastructure environments, if applicable.
Skills & tools
- Kubernetes, including custom operators, CRDs, and multi-cluster federation
- GPU hardware management (NVIDIA, CUDA)
- Cloud platforms: AWS, GCP, and others
- Distributed systems design and implementation
- Programming in Go, Rust, or C++
- Infrastructure automation and orchestration
- High-performance networking technologies (InfiniBand, RoCE)
- Monitoring and observability tools (Prometheus, Grafana, Weights & Biases)
- Containerization and cloud infrastructure management
Practical notes
an on-site role based in San Francisco. Candidates should be prepared to work within a hardware and cloud infrastructure environment that demands high reliability and operational efficiency. The position involves deep technical ownership of GPU infrastructure supporting large-scale AI inference workloads, requiring strong systems programming skills and experience managing complex distributed systems across multiple cloud providers. The role offers an opportunity to influence the core infrastructure that enables real-time AI services serving hundreds of millions of queries per month. You will be working closely with research and inference teams to develop scalable, reliable, and efficient GPU management solutions, ensuring the platform can meet the growing demands of AI inference at scale.