Staff Product Manager, Insights
Job description
About the role
You will define the roadmap and strategy for observability and data-driven insights across our specialized AI cloud infrastructure. This role focuses on building tools that provide transparency into large-scale GPU clusters and complex AI workloads. You will own the end-to-end product lifecycle for observability features embedded within the Mission Control suite. The position requires close collaboration with infrastructure and platform engineering teams to ensure insights are accurate, timely, and relevant. You will be responsible for translating intricate infrastructure telemetry into clear and actionable information for AI engineering teams. The role demands a strategic mindset to align product vision with the evolving needs of AI workload customers. You will establish metrics and reporting standards that directly help users optimize AI model training and inference workflows. Continuous feedback collection from both internal and external stakeholders will guide your prioritization of the product backlog.
Key facts
What you'll do
- Lead the product lifecycle for observability features within the Mission Control suite from discovery through launch.
- Translate complex infrastructure performance data into actionable insights for AI engineering teams to drive faster decision-making.
- Partner with engineering teams to enhance visibility into GPU utilization, node health, and overall cluster efficiency metrics.
- Define and maintain core metrics and reporting standards that help users optimize AI model training and inference workflows.
- Gather qualitative and quantitative feedback from internal and external stakeholders to continuously prioritize the product backlog.
- Collaborate with cross-functional partners to identify opportunities for new insights that improve the user experience of AI cloud operations.
- Analyze user behavior and product usage data to inform iterative improvements and validate the success of new features.
- Work closely with data platform teams to ensure the underlying data pipelines support the requirements for high-fidelity observability.
- Champion the voice of the customer by articulating user needs, pain points, and desired outcomes to the broader product organization.
- Drive the development of dashboards and reporting tools that provide clarity into the performance and cost of AI infrastructure.
- Establish a clear product vision and long-term strategy for the insights portfolio in alignment with CoreWeave's business objectives.
- Facilitate discussions with stakeholders to resolve conflicting requirements and ensure alignment on product direction.
- Monitor competitive landscape and industry trends to identify best practices and opportunities for differentiation.
- Ensure all product initiatives adhere to CoreWeave's standards for quality, reliability, and user-centric design.
Requirements
- Proven experience in product management for technical infrastructure, cloud platforms, or observability tools in a prior role.
- Ability to communicate technical concepts effectively to both engineering teams and business stakeholders without relying on jargon.
- Experience working with distributed systems or large-scale compute environments that operate at the infrastructure level.
- Strong analytical skills demonstrated through a focus on data-driven decision making and evidence-based product choices.
- Demonstrated ability to manage multiple priorities in a fast-paced and dynamic environment.
- Experience leading cross-functional initiatives that require collaboration with engineering, design, and product teams.
- Understanding of the challenges and constraints inherent in operating AI and machine learning workloads in the cloud.
- Willingness to travel occasionally for meetings, customer visits, or industry events as required by business needs.
Skills & tools
- Observability and monitoring platforms such as Prometheus, Grafana, or similar systems.
- Deep familiarity with cloud infrastructure and GPU compute environments used for AI workloads.
- Hands-on experience with Kubernetes and container orchestration tools.
- Proficiency with data visualization and reporting tools to build dashboards and user-facing insights.
- Knowledge of logging and tracing systems that provide end-to-end visibility into distributed applications.
- Experience with scripting or programming languages that enable interaction with cloud APIs.
- Understanding of cost management and optimization principles in cloud environments.
Practical notes
CoreWeave is an AI-native cloud provider specializing in high-performance GPU infrastructure. Applicants must be authorized to work in the United States. Travel requirements are currently estimated at 0-10% of time. This role is based in Livingston, NJ, with New York, NY, Sunnyvale, CA, and Bellevue, WA as alternative locations. The engagement for this position is Full-time.