Head of Infrastructure Support
Job description
About the role
You own the end-to-end success of Infrastructure Support for your region, managing the team, function, and direct impact on customers. You report directly to the VP of Support and collaborate with counterpart Heads across EMEA, the US, and APAC to ensure cohesive global support operations. You are accountable for regional service performance, escalation quality, customer experience, and the operational health of the GPU estates your team supports. You will grow your regional Infrastructure Support team during rapid company scaling, embedding a consistent operating model and acting as the organizational accountability layer to ensure strategic and tactical work gets clear ownership and completion. You remain technically credible with GPU infrastructure, high-performance fabrics, and Linux operations so you can lead complex incident response while spending the majority of your time leading people and processes.
Key facts
What you'll do
Own the success of Infrastructure Support for your region, including service outcomes, customer impact, and team performance.
Own regional service performance against defined KPIs such as SLA adherence, MTTR, first-response time, backlog health, and CSAT, with accurate reporting to the VP of Support and senior leadership.
Identify regional risks related to capacity, capability, coverage, or customers early, and either resolve them or escalate them with a clear recommendation.
Own regional capacity modelling and headcount planning by forecasting support demand against fleet growth and customer onboarding, and making the business case for investment to the VP of Support.
Act as the regional accountability layer during rapid growth when cross-functional work spanning Support, DC Operations, deployment, firmware, and Engineering lacks a clear owner, ensuring that work gets owned and completed.
Partner with the Heads of Infrastructure Support in other regions across EMEA, the US, and APAC to run a single global function with consistent standards, processes, and quality, enabling true follow-the-sun handover between regions.
Own a global capability area on behalf of all regions, such as escalation management standards, the knowledge and runbook system, or the tooling and automation roadmap, working with other Heads of Infrastructure Support to define standards for regional teams.
Own day-to-day people management for your regional Infrastructure Support team through regular 1:1s, performance reviews, development planning, and documented performance management, including underperformance, through to outcome.
Hire and grow the team by defining role requirements, running structured interviews, and building a bench of engineers who meet Nscale's technical and communication standards.
Design your team's structure as the region scales, appointing and developing team leads and building second-line management capability as headcount grows.
Set and monitor individual and team objectives, driving accountability and continuous improvement.
Design and own shift planning, rota coverage, and on-call scheduling for the region, ensuring sustainable 24/7 support in coordination with the global coverage model.
Identify skills gaps and drive upskilling through training, mentoring, and knowledge sharing across teams.
Ensure roles, responsibilities, and expectations are clearly understood and consistently applied.
Own ticket queue health for the region, including accurate prioritization, timely resolution, and clean escalation flow from frontline triage into L2/L3.
Monitor team productivity and workload trends, addressing bottlenecks before they become service risks.
Ensure adherence to ITIL-aligned processes across incident, request, change, and problem management.
Improve dashboards, alerting, and runbooks to reduce repeat incidents and drive right-first-time resolution.
Maintain consistent standards, processes, and documentation across regional teams and ensure compliance with audit, security, and operational requirements.
Act as the senior regional escalation point for complex or high-impact incidents, including customer-facing escalations, participating in regional on-call as required.
Lead post-incident reviews, identify recurring patterns, and ensure follow-up actions are tracked and delivered, converting incidents into problem records and durable fixes.
Represent Infrastructure Support to regional customers and internal senior stakeholders, communicating clearly, candidly, and concisely at every level from engineer to executive.
Contribute to readiness and support planning for new services, data centre deployments, and customer onboarding in the region.
Work alongside Senior Engineers on complex incidents, technical improvements, and operational tooling, remaining close enough to the work to lead it credibly.
Maintain hands-on fluency across GPU infrastructure, including drivers, firmware, hardware fault isolation, and RMA workflows, as well as Linux at scale and east-west high-performance fabrics such as InfiniBand and RoCE.
Understand east-west fabrics including RDMA over InfiniBand and/or RoCE, link-level diagnostics, and how fabric health drives cluster performance, and lead and challenge fabric-related incident response.
Demonstrate hands-on experience with HPC scheduling and orchestration, including operational Slurm for multi-GPU workloads, containerised execution via Pyxis/Enroot, MPI-based communication, and deep diagnostics of queue health, network topology, and job-level failures.
Understand networking fundamentals such as L2/L3, routing, VLANs, load balancing, and how east-west cluster traffic differs from north-south traffic.
Maintain data centre operations knowledge, including servers, networking, storage, power, and virtualization in an operational support context, including working with onsite DC Operations and smart-hands teams.
Demonstrate strong observability and incident response skills, interpreting metrics and alerts, driving incidents to resolution, and leading post-incident improvement.
Contribute to scripting and automation direction using Bash, Python, or similar languages, and demonstrate familiarity with Infrastructure as Code tools such as Ansible, Terraform, or similar.
Travel to Nscale or customer sites when needed to lead onsite support activity.
Requirements
5+ years of direct line management of engineers in an operational support environment, with end-to-end ownership of performance management including reviews, development plans, and documented underperformance processes through to outcome.
Experience owning team workload, prioritization, and service delivery against SLAs, with accountability for the numbers and the ability to explain those numbers to senior leadership.
Experience hiring, scaling, or standing up support/operations capability in a fast-moving environment, including capacity modelling and headcount planning against demand; comfortable operating where processes are still evolving and helping define them without slowing delivery.
Excellent written and verbal communication, clear, specific, and concise at every level from ticket notes to executive updates and difficult customer conversations, treating communication quality as a core leadership skill.
A bias for decisive action and calculated risk in ambiguous situations, taking ownership of outcomes, speaking candidly, disagreeing when appropriate, and committing fully once decisions are made.
8+ years of infrastructure, operations, or support engineering in production environments, including 4+ years of direct line management of engineers in an operational support function, with demonstrable ownership of performance management and significant exposure to GPU, HPC, or large-scale data centre estates.
Linux systems engineering in production, with proven troubleshooting across compute, storage, and network layers.
GPU infrastructure working knowledge including NVIDIA or AMD platforms, driver/firmware stacks, hardware diagnostics (nvidia-smi, DCGM), fault isolation, and RMA workflows on AI or HPC estates.
High-performance east-west fabrics understanding including RDMA over InfiniBand and/or RoCE, link-level diagnostics, and how fabric health drives cluster performance.
HPC scheduling and orchestration with hands-on Slurm operation for multi-GPU workloads, including containerised execution via Pyxis/Enroot, MPI-based communication, and deep diagnostics of queue health, network topology, and job-level failures.
Networking fundamentals including L2/L3, routing, VLANs, load balancing, and how east-west cluster traffic differs from north-south traffic.
Data centre operations including servers, networking, storage, power, and virtualization in an operational support context, including working with onsite DC Operations and smart-hands teams.
Observability and incident response including interpreting metrics and alerts, driving incidents to resolution, and leading post-incident improvement.
Automation scripting in Bash, Python, or similar, plus familiarity with Infrastructure as Code tools such as Ansible, Terraform, or similar.
Strong understanding of ITIL-aligned incident, problem, and change management, and of SRE practices including runbooks, toil reduction, and continuous improvement.
Practical notes
LENGTH: 700-900 words. No HTML, no markdown, no em dashes. Output the page only.