Senior Manager, Technical Support Engineering
Job description
About the role
CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. You will own, scale, and continuously improve CoreWeave's infrastructure support function, a large, globally distributed 24/7/365 organization of skilled engineers who resolve customers' most complex technical challenges with deep expertise, efficiency, and empathy. You will set the technical and operational bar for the entire function and design the operating systems, including coverage models, staffing, Direct-to-Expert partnerships, quality programs, and metrics that let support scale reliably. You will lead with empathy and invest in each person's growth, while building the structure and processes that let the team scale with CoreWeave's hyper-growth without compromising the culture we care deeply about. This role provides prominent visibility and influence across the company as the connective tissue between customers and CoreWeave's engineering organization.
Key facts
What you'll do
Own the strategy, health, and performance of the entire 24/7/365 infrastructure support function, helping scale coverage and capability across regions and domains as CoreWeave grows.
Own talent acquisition and retention by hiring, onboarding, and developing engineers through diligent performance management and coaching tailored to each individual's needs.
Stay hands-on: dig into complex, customer-impacting issues alongside your team and serve as a senior technical escalation point, ensuring the highest quality of support.
Build and facilitate the enablement and career-development frameworks for onboarding, technical training, and progression paths that raise capability across the entire function.
Partner directly with Product Engineering, Specialist Field Engineers, and domain specialists to triage, resolve, and close the most challenging customer issues end-to-end.
Design and operate the Direct-to-Expert model so that the right expert is on the problem fast, with clear ownership and fewer handoffs for a high-touch customer experience.
Establish and maintain the operational systems, including coverage and staffing models, that keep the 24/7/365 support organization resilient and responsive.
Define and drive quality programs and metrics that measure support effectiveness, customer impact, and team health at scale.
Coordinate cross-functional responses to large-scale, mission-critical training workloads, GPU compute issues, high-performance networking, and Slurm/HPC clusters.
Feed insights from the field directly into the product roadmap to ensure infrastructure improvements are driven by real customer needs.
Champion a culture of empathy, ownership, and continuous improvement within the support engineering organization.
Evaluate and implement tools, processes, and platforms that improve issue resolution speed and customer satisfaction.
Maintain current technical depth in infrastructure technologies relevant to GPU compute, high-performance networking, storage, and HPC scheduling.
Act as a trusted advisor to both internal stakeholders and external customers, aligning technical solutions with business outcomes.
Drive incident management practices to ensure fast mitigation, clear communication, and thorough postmortems that prevent recurrence.
Requirements
8+ years of experience in technical support, infrastructure support, or a related field.
Demonstrated ability to operate at scale with complex, distributed, 24/7/365 organizations.
Strong experience hiring, managing, and developing high-performing engineering teams.
Exceptional problem-solving skills with a hands-on mindset for resolving deep technical issues.
Excellent communication and coaching abilities for working with customers, executives, and cross-functional partners.
Proficiency in operating Kubernetes, GPU compute, high-performance networking, and storage systems.
Experience supporting large-scale, mission-critical training workloads and HPC environments.
Strong understanding of metrics, quality programs, and operational systems for scaling support functions.
Nice to have
Experience with Direct-to-Expert support models and building specialized support organizations.
Background in AI, machine learning infrastructure, and high-performance computing domains.
Practical notes
This is a full-time position.
The role may involve travel as needed.
Applicants must be eligible to work in the locations specified in the key facts.
No specific visa or deadline information is provided in the source.