Head of Infrastructure Support APAC
Job description
About the role
You own the end-to-end success of Infrastructure Support for the APAC region at Nscale, directly shaping service performance, escalation quality, and customer experience for GPU estates. You report to the VP of Support and operate alongside counterparts in EMEA and the US to ensure cohesive global support standards and true follow-the-sun coverage. You remain technically credible with GPU infrastructure, high-performance fabrics, and Linux operations so you can lead complex incident response while spending most of your time leading and developing your team. You will grow your regional Infrastructure Support team during rapid scaling, embed consistent operating models, and act as the organizational accountability layer ensuring strategic and tactical work lands with clear owners and gets driven to completion.
Key facts
What you'll do
- Assume regional ownership and accountability for Infrastructure Support success in APAC, owning service outcomes, customer impact, and team performance.
- Drive regional service performance against defined KPIs, including SLA adherence, MTTR, first-response time, backlog health, and CSAT with accurate reporting to the VP of Support and senior leadership.
- Identify APAC regional risks across capacity, capability, coverage, or customer early, and either resolve them or escalate them with a clear recommendation.
- Own APAC regional capacity modelling and headcount planning, forecasting support demand against fleet growth and customer onboarding, and making the business case for investment to the VP of Support.
- Act as the APAC regional accountability layer during rapid growth, ensuring that cross-functional work spanning Support, DC Operations, deployment, firmware, and Engineering lacks a clear owner gets one and gets driven to completion.
- Partner with counterparts across EMEA and US to run a single global function, maintaining consistent standards, processes, and quality with true follow-the-sun handover between regions.
- Own a global capability area on behalf of all regions, such as escalation management standards, the knowledge and runbook system, or the tooling and automation roadmap, working with other Heads of Infrastructure Support to define the standard every regional Support team operates to.
- Manage the day-to-day people management for your regional Infrastructure Support team, including regular 1:1s, performance reviews, development planning, and documented performance management from underperformance through to outcomes.
- Hire and grow the team by defining role requirements, running structured interviews, and building a bench of engineers who meet Nscale's technical and communication bar.
- Design your team's structure as the region scales, appointing and developing team leads and building second-line management capability as headcount grows.
Requirements
- Bring 5+ years of direct line management of engineers in an operational support environment, with end-to-end ownership of performance management including reviews, development plans, and documented underperformance processes through to outcomes.
- Demonstrate experience owning team workload, prioritisation, and service delivery against SLAs, with accountability for the numbers and the ability to explain those numbers to senior leadership.
- Show experience hiring, scaling, or standing up support/operations capability in a fast-moving environment, including capacity modelling and headcount planning against demand, and comfort operating where processes are still evolving while helping define them without slowing delivery.
- Exhibit excellent written and verbal communication with clear, specific, and concise messaging at every level, from ticket notes to executive updates to difficult customer conversations, with communication quality treated as a core leadership skill.
- Display a bias for decisive action and calculated risk in ambiguous situations, taking ownership of outcomes, speaking candidly, disagreeing when appropriate, and committing fully once decisions are made.
- Maintain hands-on technical foundation across Linux systems engineering in production, with proven troubleshooting across compute, storage, and network layers at scale.
- Apply working knowledge of GPU infrastructure, including NVIDIA or AMD platforms, driver/firmware stacks, hardware diagnostics (nvidia-smi, DCGM), fault isolation, and RMA workflows on AI or HPC estates.
- Demonstrate understanding of high-performance east-west fabrics, including RDMA over InfiniBand and/or RoCE, link-level diagnostics, and how fabric health drives cluster performance, with the ability to lead and challenge fabric-related incident response.
- Show hands-on experience with HPC scheduling and orchestration, such as Slurm for multi-GPU workloads, including containerised execution via Pyxis/Enroot, MPI-based communication, and deep diagnostics of queue health, network topology, and job-level failures.
- Possess networking fundamentals across L2/L3, routing, VLANs, load balancing, and how east-west cluster traffic differs from north-south traffic.
- Provide solid understanding of data centre operations, including servers, networking, storage, power, and virtualisation in an operational support context, including working with onsite DC Operations and smart-hands teams.
- Exhibit competence in observability and incident response, interpreting metrics and alerts, driving incidents to resolution, and leading post-incident improvement.
- Apply scripting and automation skills using Bash, Python, or similar, along with familiarity with Infrastructure as Code tools such as Ansible, Terraform, or similar.
- Demonstrate strong process literacy with ITIL-aligned incident, problem, and change management, and familiarity with SRE practices including runbooks, toil reduction, and continuous improvement.
- Maintain adaptability and comfort with out-of-hours escalations, regional on-call participation, and travel for onsite leadership when required.
Nice to have
- Deeper GPU/HPC exposure, including NCCL-based performance troubleshooting, NVLink/NVSwitch, Slurm-scheduled multi-GPU workloads, or rack-scale systems.
- High-performance storage exposure, such as VAST or comparable AI-optimised storage platforms, Ceph, or NFS at scale.
- OpenStack and fleet operations tooling experience, including OpenStack operations or fleet-scale provisioning and health tooling like MAAS, NetBox, Redfish-driven automation, or similar.
- Experience operating or supporting clusters and GPU operator stacks, which provides helpful context for the platform though is not the core of this role.
- Background in multi-region or follow-the-sun operations, including experience running or coordinating sup