
Technical Support Engineer (GPU Clusters)
Job description
About the role
The role delivers first-line customer support for AI training, fine tuning, and inference workloads on Together AI platforms. The hire partners with product and sales to drive continuous improvement of offerings in a fast-paced, innovative environment. This position owns the direct resolution of complex technical challenges related to AI workloads on Kubernetes GPU clusters. The customer-facing SRE role is responsible for maintaining stable and healthy Kubernetes clusters to protect uptime and performance for clients. Patterns identified in support cases are transformed into product and roadmap inputs, guiding future models and capabilities alongside engineering and go-to-market teams. The hire will work cross-functionally with sales, engineering, support, product, and research to drive customer success. A strong sense of ownership and willingness to learn new skills supports both team and customer success. Clear communication and interpersonal skills are used to explain complex technical concepts to non-technical stakeholders.
Key facts
What you'll do
Customers receive swift and effective solutions through direct engagement with complex technical challenges on Kubernetes GPU clusters.
The customer-facing SRE role maintains stable and healthy Kubernetes clusters to protect uptime and performance for clients.
Patterns in support cases are transformed into product and roadmap inputs, guiding future models and capabilities alongside engineering and go-to-market teams.
Requirements for the position include at least 3 years of experience in a customer-facing technical role with a minimum of 1 year in support for AI services or mission-critical SaaS APIs.
SRE or DevOps experience with Kubernetes operations is required for this position.
A strong technical background covers AI, ML, GPU technologies, and their integration into high-performance computing environments.
Mastery of infrastructure services such as Kubernetes and SLURM, infrastructure as code tools like Ansible, high-performance network fabrics, NFS-based storage management, container infrastructure, and scripting languages is necessary.
Hands-on experience with HPC/Slurm cluster environments including node draining, job scheduling, and maintenance workflows is required.
Familiarity with high-speed networking concepts such as InfiniBand, RDMA, and network interface diagnostics is required.
Experience with distributed storage systems like Weka and NFS and the ability to troubleshoot I/O and bandwidth issues is required.
A foundational understanding of installation, configuration, administration, troubleshooting, and securing compute clusters is mandatory.
Complex technical problem solving and proactive issue resolution are core expectations for this role.
The ability to work cross-functionally with sales, engineering, support, product, and research drives customer success.
A strong sense of ownership and willingness to learn new skills supports both team and customer success.
Clear communication and interpersonal skills explain complex technical concepts to non-technical stakeholders.
Adaptability in dynamic environments manages multiple projects and handles frequent context switching and prioritization.
Practical notes
The role begins as a Monday to Friday position for ramping up and learning, then transitions to a weekend shift after becoming fully ramped. Remote work is supported under company policies.
Typical interview steps
Hiring for engineering roles usually starts with a recruiter screen, followed by one or two technical rounds. Candidates often solve a coding problem, discuss past projects, and answer system design questions. Some loops include a take-home task. Final rounds typically cover team fit and give candidates a chance to ask questions. Interviewers look for how you break down an unfamiliar problem, not just whether you reach the answer. Practicing a few problems aloud and reviewing your own past projects are the best preparation.
Good to know
Technical support in AI infrastructure combines deep system knowledge with customer communication. Common tools include Kubernetes, Slurm, and distributed storage platforms. Success depends on proactive monitoring, clear documentation, and cross-team collaboration.