Staff Software Engineer, Observability
Job description
About the role
CoreWeave is seeking a technical leader to architect and scale our observability systems for high-performance AI infrastructure. You will design solutions that provide deep visibility into our massive GPU compute clusters and cloud services. This role requires a strategic thinker who can bridge the gap between raw infrastructure performance and actionable operational insights. The ideal candidate will own the end-to-end observability lifecycle from data collection to alert refinement. You will partner closely with platform and AI infrastructure teams to ensure telemetry aligns with evolving product needs. This position demands a balance between deep technical execution and high-level system design. You will be responsible for driving standardization and best practices across the observability stack. Success in this role will be measured by the reliability and clarity of the insights provided to our engineering teams.
Key facts
What you'll do
- Architect and implement large-scale observability platforms to monitor AI training and inference workloads across global data centers.
- Engineer high-throughput telemetry pipelines capable of processing petabyte-scale data streams from distributed GPU clusters without loss of fidelity.
- Partner with cross-functional engineering teams to diagnose complex performance regressions and reliability issues in production environments.
- Establish and enforce enterprise-wide standards for structured logging, granular metrics, and distributed tracing across hybrid cloud infrastructure.
- Optimize resource utilization of observability tooling to ensure collection and processing introduce minimal latency and overhead on critical AI workloads.
- Design and evolve dashboards and visualization frameworks that provide real-time operational intelligence to technical and executive stakeholders.
- Lead incident response efforts by providing deep forensic analysis through correlated logs, metrics, and traces during critical system events.
- Evaluate and integrate emerging open-source and commercial observability technologies to maintain a competitive edge in monitoring capabilities.
- Collaborate with security and compliance teams to ensure telemetry data governance meets regulatory and internal policy requirements.
- Mentor junior engineers and contribute to the technical roadmap of the observability platform within the Mission Control organization.
- Facilitate knowledge transfer sessions to improve team-wide understanding of monitoring best practices and debugging methodologies.
- Drive the automation of alerting rules to reduce noise and improve signal quality for on-call engineering staff.
- Work closely with product managers to translate business objectives into measurable service level objectives and key results.
- Champion the adoption of site reliability engineering principles to build a culture of ownership and reliability across the engineering organization.
Requirements
- Demonstrate extensive professional experience as a software engineer with a primary focus on observability, monitoring, or distributed systems over a significant portion of your career.
- Show a proven track record of designing and implementing monitoring and observability solutions for complex, cloud-native, and microservices-based environments.
- Exhibit strong proficiency in collecting, processing, and analyzing time-series data at scale, including the management of high-cardinality metrics.
- Have a deep background in building and operating tools that enable infrastructure automation, fleet management, and self-service platform capabilities.
- Possess the ability to work effectively in a hybrid work model, collaborating seamlessly across multiple physical office locations including Livingston, New York, Sunnyvale, and Bellevue.
- Bring a history of writing clean, maintainable, and well-documented code that can be understood and extended by other senior engineers.
- Show evidence of strong problem-solving skills, particularly when diagnosing root causes of elusive or intermittent system failures.
- Must be comfortable working with abstract concepts and translating ambiguous business requirements into concrete technical specifications.
Skills & tools
- Kubernetes
- Distributed systems architecture
- Telemetry and monitoring frameworks
- Large-scale data processing
- Cloud infrastructure management
Practical notes
CoreWeave is an AI-native cloud provider focused on high-performance GPU compute. This role sits within the Mission Control organization, which manages the lifecycle and operational transparency of our global data centers.