HPC Support Engineer
Job description
About the role
Lambda is building the world's best AI cloud platform, and we are looking for a highly skilled HPC Support Engineer to join our team. In this role, you will be responsible for maintaining and troubleshooting our high-performance computing infrastructure, ensuring optimal performance and reliability. You will work closely with engineering and support teams to resolve complex technical issues, improve support processes, and develop automation tools. This position offers an exciting opportunity to contribute to cutting-edge AI infrastructure and support a rapidly growing technology platform. The ideal candidate will have extensive experience in HPC environments, strong Linux support skills, and a proactive approach to problem-solving.
Key facts
What you'll do
- Act as a senior resource for resolving difficult infrastructure and platform issues, investigating problems down to the hardware, driver, or kernel level. You will be expected to troubleshoot complex issues that may involve GPU hardware, network components, or system software.
- Quickly differentiate between hardware malfunctions, driver problems, kernel issues, and customer configuration errors to ensure accurate and efficient problem resolution. Your expertise will help reduce downtime and improve customer satisfaction.
- Proactively identify and address shortcomings in existing processes, tools, and documentation. You will take initiative to implement improvements that streamline support workflows and enhance operational efficiency.
- Utilize AI tools to develop scripts, automations, or small internal utilities that address operational gaps. Your automation solutions will help scale support efforts and improve response times.
- Conduct root-cause analysis across distributed systems, clusters, and GPU hardware to identify underlying issues and prevent future incidents. Your analysis will contribute to the stability and performance of Lambda's infrastructure.
- Create clear, comprehensive documentation for solutions and contribute to the evolution of support procedures. Your documentation will serve as a valuable resource for support engineers and engineering teams.
- Partner with engineering teams to translate common customer difficulties into durable, scalable solutions. Your collaboration will help improve product robustness and reduce recurring support issues.
- Handle escalated issues from colleagues, providing guidance, training, and mentorship to support engineers. Your leadership will help build a strong, knowledgeable support team.
- Participate in a rotating on-call schedule, taking ownership of significant incidents and customer problems outside of regular hours. Your availability will be critical during high-volume deployment periods or major outages.
- Be prepared to assist wherever needed, particularly during rapid, high-volume deployments or critical incident responses. Flexibility and teamwork are essential in this fast-paced environment.
Requirements
- A minimum of 3 years of practical experience in HPC administration, support, or engineering roles. You should have a proven track record of managing high-performance computing systems.
- Extensive experience and a strong understanding of supporting Linux in a system administration capacity. You must be comfortable troubleshooting Linux-based clusters and related infrastructure.
- Demonstrated experience working in HPC environments, specifically with Linux cluster administration. Familiarity with cluster management and job scheduling systems is essential.
- Proficiency in coding and CI/CD practices, with a history of using AI-assisted tools for faster development. Your coding skills will support automation and tooling efforts.
- Familiarity with monitoring and logging tools such as Prometheus, Grafana, or Datadog. You should be able to analyze logs, set up dashboards, and interpret system metrics.
- Strong abilities in analyzing logs, debugging kernel-level issues, and performance profiling. Your expertise will help identify bottlenecks and hardware issues.
- Experience with CUDA, NCCL, NVLink, and GPUDirect RDMA. Knowledge of GPU programming and interconnect technologies is crucial for troubleshooting GPU-related problems.
- Experience with high-throughput networking technologies, including Infiniband (IB) and RoCE. You should understand network configurations and performance tuning for HPC clusters.
- Knowledge of distributed AI/ML or HPC workloads. Familiarity with the demands of large-scale machine learning workloads will be beneficial.
- Understanding of TCP/IP, VPNs, and firewalls in cloud environments. You should be able to configure and troubleshoot network security and connectivity issues.
- The capacity to work independently and guide less experienced support engineers. Leadership and mentorship skills are valued in this role.
Nice to have
- Experience with containerization technologies like Docker and Kubernetes. Container orchestration knowledge will help in deploying and managing services efficiently.
- Experience with GPU cloud providers. Familiarity with cloud-based GPU offerings can enhance support capabilities.
- Experience with high-performance storage systems. Knowledge of storage solutions used in HPC environments will be advantageous.
- Familiarity with infrastructure-as-code tools such as Terraform or Ansible. Your ability to automate infrastructure provisioning and configuration will be a plus.
- Experience with Nvidia GPUs and Infiniband networks. Hands-on experience with these technologies will help in troubleshooting and optimizing performance.
Skills & tools
- Linux
- HPC Cluster Administration
- Kubernetes
- Slurm
- Python or other scripting languages
- CI/CD pipelines
- Prometheus
- Grafana
- Datadog
- Log analysis tools
- Kernel debugging techniques
- Performance profiling methods
- CUDA
- NCCL
- NVLink
- GPUDirect RDMA
- Infiniband (IB)
- RoCE
- TCP/IP
- VPN
- Firewalls
Practical notes
- This is a full-time, on-site position based in the USA. The role requires working at our office location to support our infrastructure effectively.
- Lambda offers a competitive compensation package including salary and equity options. We also provide comprehensive health, dental, and vision insurance plans, a 401k plan with company matching, and flexible paid time off to promote work-life balance.
- We are committed to diversity and equal opportunity employment. Lambda welcomes applicants from all backgrounds and is dedicated to fostering an inclusive environment.
- The role involves participation in a rotating on-call schedule, which may require availability outside of standard working hours during critical incidents or high-volume deployment periods.
- Candidates should be prepared for ongoing learning and collaboration, as the environment involves cutting-edge AI infrastructure and continuous technological advancements.
- Prior experience working in fast-paced, high-stakes support environments will be beneficial for success in this role.