HPC Support Engineer
lambdaRemote (USA)Full Time2d ago
GoDockerKubernetesTerraformAnsibleCI/CDMachine LearningAIMLLinuxSupportSolutions
Job description
HPC Support Engineer at Lambda
About the role
Lambda is building the world's best AI cloud, and we need a skilled engineer to ensure our high-performance computing infrastructure runs smoothly. This role involves tackling complex technical challenges and improving our support systems.
Key facts
What you'll do
- Act as a senior resource for resolving difficult infrastructure and platform issues, investigating down to the hardware, driver, or kernel level.
- Quickly differentiate between hardware malfunctions, driver problems, kernel issues, and customer configuration errors to ensure accurate and efficient problem resolution.
- Proactively identify and address shortcomings in processes, tools, and documentation, taking initiative to implement improvements.
- Utilize AI tools to develop scripts, automations, or small internal utilities that solve operational gaps.
- Conduct root-cause analysis across distributed systems, clusters, and GPU hardware.
- Create clear documentation for solutions and contribute to the evolution of support procedures.
- Partner with engineering teams to translate common customer difficulties into lasting solutions.
- Handle escalated issues from colleagues, providing training and mentorship.
- Participate in a rotating on-call schedule, taking ownership of significant incidents and customer problems.
- Be prepared to assist wherever needed, particularly during rapid, high-volume deployments.
Requirements
- A minimum of 3 years of practical experience in HPC administration, support, or engineering.
- Extensive experience and a strong understanding of supporting Linux in a system administration capacity.
- Demonstrated experience in HPC environments, specifically with Linux cluster administration.
- Proficiency in coding and CI/CD practices, with a history of using AI-assisted tools for faster development.
- Familiarity with monitoring and logging tools such as Prometheus, Grafana, or Datadog.
- Strong abilities in analyzing logs, debugging kernel-level issues, and performance profiling.
- Experience with CUDA, NCCL, NVLink, and GPUDirect RDMA.
- Experience with high-throughput networking technologies, including Infiniband (IB) and RoCE.
- Knowledge of distributed AI/ML or HPC workloads.
- Understanding of TCP/IP, VPNs, and firewalls in cloud environments.
- The capacity to work independently and guide less experienced support engineers.
Nice to have
- Experience with containerization technologies like Docker and Kubernetes.
- Experience with GPU cloud providers.
- Experience with high-performance storage systems.
- Familiarity with infrastructure-as-code tools such as Terraform or Ansible.
- Experience with Nvidia GPUs and Infiniband.
Skills & tools
- Linux
- HPC Cluster Administration
- Kubernetes
- Slurm
- Python (or other scripting languages)
- CI/CD
- Prometheus
- Grafana
- Datadog
- Log Analysis
- Kernel Debugging
- Performance Profiling
- CUDA
- NCCL
- NVLink
- GPUDirect RDMA
- Infiniband (IB)
- RoCE
- TCP/IP
- VPN
- Firewalls
Practical notes
- This is a full-time, remote position based in the USA.
- Lambda offers competitive cash and equity compensation, health, dental, and vision insurance, a 401k plan with a company match, and flexible paid time off.
- Lambda is an Equal Opportunity Employer.