Staff Product Manager
Job description
About the role
You will lead the observability strategy for our AI cloud infrastructure, ensuring customers can effectively monitor and debug large-scale training runs. This role involves building the telemetry layer that powers both internal fleet reliability and the customer-facing console and API. You will report to the Head of Platform Product Management and collaborate with SRE and engineering teams to define the future of our diagnostic tools.
Key facts
What you'll do
- Define and execute the product roadmap for observability across single GPU instances and clusters ranging from 64 to 1,024+ GPUs.
- Determine which GPU health, cluster utilization, and InfiniBand fabric metrics are surfaced to customers via the API and console.
- Develop diagnostic experiences that allow users to isolate issues between their code, configuration, or platform infrastructure.
- Establish SLIs, SLOs, and SLAs to provide transparency into platform performance and reliability.
- Collaborate with SRE and fleet engineering to create a unified telemetry foundation for both internal operations and customer visibility.
- Engage directly with AI research labs and ML teams to gather feedback on debugging workflows and translate those insights into product requirements.
- Drive cross-functional alignment with engineering and design teams to ensure successful product launches.
- Instrument feature adoption and utilize performance metrics to inform iterative development cycles.
Requirements
- 7+ years of experience in product management.
- 3+ years of experience working on developer-facing, platform, or technical infrastructure products.
- Proven ability to influence stakeholders and drive cross-team adoption without formal authority.
- Experience managing products through the full lifecycle from initial concept to launch and post-launch iteration.
- Ability to engage in technical discussions regarding distributed systems, failure modes, and hardware health.
- Strong written communication skills with the ability to produce concise documentation for alignment.
- Capacity to thrive in ambiguous environments and establish structure for new product domains.
Nice to have
- Experience shipping monitoring or observability tools similar to Datadog, Grafana, or Prometheus.
- Background in distributed systems telemetry or high performance computing fabric-level metrics.
- Experience building API-first products or developer tools.
- Hands-on knowledge of distributed training stacks like PyTorch with NCCL or GPU fleet tooling such as DCGM.
- Experience working within usage-based cloud infrastructure business models.
Skills & tools
- Distributed systems
- Telemetry and observability pipelines
- API product management
- GPU and cluster infrastructure
- SLI/SLO/SLA definition
- InfiniBand fabric metrics
- DCGM
- PyTorch / NCCL
Practical notes
- This role requires 4 days per week of in-office presence at either the Bellevue or San Francisco location.
- Tuesday is the designated company-wide work from home day.
- Benefits include health, dental, and vision coverage for employees and dependents, a 401k plan with a 2% company match, wellness and commuter stipends, and flexible paid time off.
About the company
Lambda is a company named after the eleventh letter of the Greek alphabet. This letter, represented as Λ or λ, has a rich and influential history. It represents the voiced alveolar lateral approximant sound. Its origins trace back to the ancient Phoenician letter Lamed. From this foundational root, Lambda significantly contributed to the development of other alphabets. It gave rise to the Latin letter L and the Cyrillic letter El. In the system of Greek numerals, Lambda holds a value of