Account Solution Architect
Job description
About the role
CoreWeave is seeking a technical expert to bridge the gap between our high-performance AI infrastructure and client requirements. You will design custom cloud architectures that enable organizations to train and deploy large-scale AI models efficiently. This role demands a deep understanding of how computational infrastructure directly impacts the speed and success of AI initiatives. You will act as a trusted advisor, guiding clients through the complexities of building AI workloads in the cloud. A significant portion of your day will be spent analyzing intricate customer problems and matching them with the most effective technical solutions. You will leverage your expertise to ensure that every architecture is scalable, reliable, and aligned with business objectives. This position is critical for translating the immense power of CoreWeave's platform into tangible client outcomes. You will be the technical cornerstone of the sales and delivery teams, ensuring that client expectations are not just met but exceeded.
Key facts
What you'll do
- Analyze client AI workload patterns to architect bespoke GPU cloud environments that optimize performance and cost.
- Engineer resilient and scalable infrastructure using Managed Kubernetes, ensuring high availability for critical AI training jobs.
- Evaluate and implement specialized storage solutions to handle the massive input/output demands of large model training.
- Diagnose and resolve complex performance bottlenecks within distributed compute clusters to minimize client downtime.
- Translate ambiguous business objectives into precise technical configurations that maximize the efficiency of GPU compute resources.
- Guide clients through the deployment lifecycle, from initial setup through ongoing optimization of their cloud infrastructure.
- Collaborate with internal engineering teams to gather insights on platform capabilities for crafting innovative solutions.
- Serve as the primary technical consultant during the sales process, de-risking complex client proposals.
- Demonstrate the tangible benefits of CoreWeave's hardware and software stack through proof-of-concept implementations.
- Advise on best practices for securing and maintaining robust AI model development environments in the cloud.
- Synthesize feedback from client deployments to influence future product enhancements and feature development.
- Foster strong relationships with technical stakeholders to ensure alignment between infrastructure strategy and business goals.
Requirements
- You possess proven experience in cloud architecture, specifically within AI, machine learning, or high-performance computing environments where scale is a critical factor.
- You have technical proficiency with GPU compute resources, including their allocation, management, and optimization in production settings.
- You demonstrate strong ability to communicate complex technical concepts clearly to both engineering teams and non-technical stakeholders.
- You bring experience managing large-scale infrastructure deployments, including the troubleshooting of intricate performance bottlenecks and failures.
- You show a history of designing solutions that balance performance requirements with strict cost constraints for cloud-based operations.
- You have a solid understanding of modern container orchestration platforms and how they apply to AI and HPC workloads.
- You exhibit the capability to rapidly learn and adapt to new cloud technologies and hardware generations specific to the AI space.
- You maintain a meticulous attention to detail to ensure architectural designs are accurate, repeatable, and secure.
Nice to have
- Familiarity with NVIDIA hardware ecosystems and the specific requirements of AI model development lifecycles.
- Background in VFX rendering or other compute-intensive industry workloads that involve massive parallel processing.
- Experience working in a fast-paced, high-growth cloud infrastructure company where agility and ownership are paramount.
- Understanding of the nuances involved in scaling AI workloads from prototype to production scale.
- Knowledge of the competitive landscape and how CoreWeave's offerings compare regarding performance and flexibility.
Skills & tools
- Managed Kubernetes
- GPU Compute
- AI Model Training and Inference
- Cloud Infrastructure Design
- Distributed Systems
- Performance Tuning
- Technical Consulting
- Solution Design
- Stakeholder Communication
- Problem Solving
Practical notes
CoreWeave is an AI-native cloud provider recognized for its specialized infrastructure. Applicants should be prepared to discuss their experience with large-scale compute clusters and their approach to solving infrastructure challenges for AI pioneers. The Toronto role is a full-time position deeply embedded within the client-facing technical team. Success in this role requires a proactive mindset and a willingness to travel to client sites as necessary to ensure deployment and optimization goals are met.