Senior HPC Infrastructure Engineer
Job description
About the role
Firmus Technologies is seeking a highly skilled and driven Kubernetes HPC Engineer to join our Software Defined Infrastructure team. In this role, you will build high-performance, fault-tolerant, and reliable infrastructure to support bare-metal provisioning, performance benchmarking, and platform validation. You will be instrumental in ensuring the stability, performance, and continuous improvement of our complex and mission-critical bare-metal HPC GPU clusters. This position requires a hands-on expert who thrives on solving deep infrastructure challenges and turning demanding AI and HPC requirements into robust, scalable platform capabilities. You will own the full lifecycle of critical infrastructure components from design through production validation. Your work will directly impact the performance and reliability of large-scale AI training and scientific computing workloads.
Location: Australia
What you'll do
Design and implement bare-metal provisioning workflows using Ironic and Kubernetes CRDs to enable automated and repeatable cluster deployment.
Deploy and manage GPU-enabled AI compute nodes with RDMA, InfiniBand, and RoCE networking to ensure low-latency, high-bandwidth communication for distributed training.
Optimise Kubernetes and Slurm platforms for multi-node AI training performance, including NCCL, UCX, GPUDirect, and fabric tuning to extract maximum throughput and minimize latency.
Implement Kubernetes primitives for GPU scheduling, isolation, and resource management models that enforce strong multi-tenancy and quality of service.
Design, deploy, and fine-tune Slurm GPU clusters with topology-aware configurations that align with hardware capabilities and workload patterns.
Develop and execute performance benchmarking workloads, including MLPerf, NCCL tests, microbenchmarks, and throughput/latency validation to quantify platform capabilities and identify bottlenecks.
Establish observability across GPU, InfiniBand fabric, storage, and provisioning components to enable rapid detection and remediation of issues.
Document architecture designs, operational procedures, and performance results to create durable knowledge and support cross-team collaboration.
Collaborate with L2 SRE engineers, site operations, and networking teams to ensure platform reliability, reproducibility, and performance across all deployed workloads.
Support hardware bring-up activities, including BIOS tuning, GPU topology verification, NUMA alignment, and PCIe/NVLink checks to ensure optimal hardware behavior.
Contribute to continuous improvement in cluster validation, CI/CD automation, and provisioning and testing frameworks to increase speed and reliability over time.
Contribute to the development of custom Kubernetes operators and intelligent orchestration frameworks that optimise AI workload performance for large-scale GPU cluster commissioning.
Requirements
Bachelor's or Master's degree in Computer Science, Engineering, or a related field.
Experience with bare-metal cluster provisioning using tools such as Metal3, OpenStack Ironic, MaaS, xCAT, or similar.
Deep knowledge of Kubernetes internals, including CRDs, controllers, operators, and cluster lifecycle management.
Strong understanding of Slurm configuration and compiling AI and HPC applications for performance and scalability.
Strong understanding of GPU systems (NVIDIA H100/H200 SXM platforms), CUDA/NCCL, and GPU topology (NVLink, NVSwitch, PCIe).
Familiarity with container runtimes for compute workloads, including Docker, Enroot, Singularity, and Podman.
Practical Linux systems engineering experience, including kernel, cgroups, system services, networking, and drivers.
Experience participating in an on-call rotation supporting production services and responding to infrastructure incidents.
Key Competencies
Systems Architecture: Ability to design and integrate bare-metal, GPU, RDMA, and Kubernetes/Slurm platforms.
Infrastructure Automation: Skilled in automated provisioning and lifecycle management of hardware and clusters.
GPU and HPC Performance: Understanding of GPU systems, RDMA fabrics, and distributed AI workload performance.
Technical Communication: Ability to communicate technical concepts effectively across diverse engineering and operations teams.
Continuous Improvement: Demonstrates curiosity, proactive learning, and innovation in AI and HPC infrastructure.
Success Metrics
Reliable provisioning of Kubernetes-based bare-metal HPC GPU clusters with high stability and performance.
Effective deployment and optimization of GPU-enabled AI compute nodes with RDMA, InfiniBand, and RoCE networking.
Improved multi-node AI training performance through Kubernetes and Slurm optimization, including NCCL, UCX, GPUDirect, and fabric tuning.
Successful implementation and scheduling of GPU workloads using Kubernetes primitives for isolation and resource management.
Well-tuned Slurm GPU clusters with topology-aware configurations aligned to hardware and workload demands.
Comprehensive performance benchmarking and validation using MLPerf, NCCL tests, microbenchmarks, and throughput/latency tests.
Enhanced observability across GPU, InfiniBand fabric, storage, and provisioning components for rapid issue detection and remediation.
Clear documentation of architecture designs, operational procedures, and performance results supporting cross-team collaboration.
Reliable platform operation through active participation in on-call rotation and incident response.
Demonstrable continuous improvement in cluster validation, CI/CD automation, and provisioning and testing frameworks.
Progress in developing custom Kubernetes operators and intelligent orchestration frameworks that optimize AI workload performance at scale.