Infrastructure Support Engineer
Job description
About the role
You own the end-to-end health of Nscale's GPU fleets and the ticket lifecycle across nodes, networks, and data centre operations. You escalate early and appropriately, turning complex incidents into clean handovers and documented resolutions. You communicate technical detail with precision to customers, colleagues, and vendors, treating clarity as a core engineering skill. You absorb new hardware and platform concepts quickly, ask the right questions, and build competence under pressure. You maintain discipline in record-keeping, structured troubleshooting, and reliable follow-through. You seek feedback and use it to progress deliberately toward Senior ownership. You contribute to a culture of transparency, ownership, and continuous improvement.
Key facts
What you'll do
- Join the Support duty rotation and handle day-to-day tickets and alerts, escalating early and appropriately while collaborating with Engineering.
- Perform GPU node triage and hardware troubleshooting by interpreting nvidia-smi/DCGM output and system logs, isolating faults across GPU, NIC, and server hardware.
- Carry out physical remediation such as reseats, swap testing, and component checks, and prepare clean evidence for vendor RMA.
- Run fabric and link diagnostics following established runbooks (mlxlink or equivalent), capturing evidence accurately and enabling smooth escalations.
- Assist with storage and data-path investigations on high-performance platforms, gathering evidence for Senior or Engineering-led diagnosis.
- Follow established runbooks to resolve common issues and propose incremental improvements that undergo review.
- Accurately record, update, manage, and resolve tickets while keeping all parties informed with clear notes and next steps.
- Participate in monitoring, troubleshooting, and triage, capturing logs and facts to support efficient handovers.
- Participate in changes under peer review, learning risk assessment and backout practices in live customer environments.
- Help maintain source-of-truth accuracy across DCIM, inventory, and asset records (NetBox or similar patterns).
- Identify opportunities for automation and contribute simple scripts and tooling improvements to optimise processes.
- Act as the escalation point for onsite DC Operations staff and coordinate smart-hands tasks within your scope.
- Learn the Platform fundamentals so you can help customers get value from our services and ask for support when deeper expertise is required.
- Share knowledge by documenting steps you have validated and contributing to training materials, shadowing Seniors during complex work.
- Take part in incident reviews as a contributor and help track preventative follow-ups within your scope.
- Deliver assigned tasks and project work to agreed quality and timelines, flagging blockers early and seeking help when needed.
- Participate in on-call and out-of-hours work when scheduled and after onboarding, including travel to Nscale or customer locations to assist with deployments and operational tasks.
Requirements
- 3+ years in infrastructure support or support engineering, including deep support/service desk experience in structured, customer-facing environments.
- Clear written communication skills with a track record of concise updates and reliable follow-through.
- Working knowledge of GPU infrastructure with hands-on experience using nvidia-smi or similar diagnostics.
- Comfortable physically troubleshooting servers, including reseating components, swap testing, and working via BMC/out-of-band management.
- Solid working knowledge of Linux, including systemd, filesystems, permissions, and standard networking tools.
- Solid grasp of networking concepts such as IP addressing, subnets, VLANs, routing, DNS, and firewalls.
- Experience with ticketing and ITSM discipline, including prioritisation, escalation, SLA awareness, and accurate documentation.
- Able to use observability tools like dashboards and alerts to identify symptoms, gather evidence, and follow runbooks.
- Comfortable reading and writing simple Bash or Python and using Git for version control.
- Understanding of servers, networks, storage, and virtualisation concepts from a support or operations background.
- Curiosity and a growth mindset, with a desire to seek feedback and progress toward Senior.
- Adaptability to work in a fast-moving environment, participate in on-call after onboarding, and travel when necessary.
Nice to have
- Exposure to high-performance fabrics and GPU-HPC concepts such as RDMA/InfiniBand, link-level diagnostics, NCCL-based troubleshooting, or NVLink.
- Experience with high-performance storage platforms like VAST, Ceph, or large-scale NFS, including storage-network troubleshooting.
- Familiarity with OpenStack and fleet operations tooling such as MAAS, NetBox, Redfish, or similar platforms.
- Understanding of Kubernetes core concepts and basic troubleshooting via runbooks.
- Experience with automation and access tooling such as Ansible, Terraform, GitHub Actions, Teleport, or Vault.
- Progress toward relevant certifications in Linux, networking, Kubernetes, cloud, or security.
Practical notes
- Travel to Nscale or customer locations may be required.
- Participation in on-call and out-of-hours work is expected after onboarding.
- Compensation details are not provided in the source material.
Equal Opportunities Statement
We strongly encourage applications from people of colour, the LGBTQ+ community, people with disabilities, neurodivergent people, parents, carers, and people from lower socio-economic backgrounds. If there's anything we can do to accommodate your specific situation, please let us know. The responsibilities outlined in this job description are not exhaustive and are intended to provide a general overview of the position. The employee may be required to perform additional duties, tasks, and responsibilities as assigned by management, consistent with the skills and qualifications required for the role.