Staff Engineer, Network Observability
CoreWeaveUSA6d ago
Engineeringremotecurated-jd
Job description
Staff Engineer, Network Observability at CoreWeave.
About the role
CoreWeave is seeking a technical leader to architect and maintain visibility into our high-performance AI cloud network. You will design systems that monitor complex traffic patterns and infrastructure health to ensure our GPU clusters operate at peak efficiency.
Key facts
What you'll do
- Architect and implement observability solutions for large-scale AI cloud networking environments.
- Build tools to monitor and analyze network performance, latency, and packet flow across distributed data centers.
- Develop automated systems for detecting and diagnosing network anomalies in real time.
- Collaborate with engineering teams to improve the reliability and transparency of our infrastructure.
- Establish standards for data collection and visualization to support fleet-wide operations.
Requirements
- Extensive experience in network engineering or observability within a large-scale cloud environment.
- Proficiency in designing and managing monitoring stacks for high-throughput distributed systems.
- Deep understanding of networking protocols and traffic analysis.
- Proven ability to lead technical initiatives and mentor engineering staff.
Nice to have
- Experience with Kubernetes-based networking and observability tools.
- Background in building custom telemetry pipelines for massive GPU-compute clusters.
Skills & tools
- Network monitoring and diagnostic frameworks
- Distributed systems architecture
- Cloud infrastructure observability
- Data visualization and alerting systems
Practical notes
CoreWeave provides purpose-built infrastructure for AI and high-performance computing. Applicants should be prepared to work in a fast-paced environment focused on scaling AI workloads. Please submit your application via the CoreWeave careers portal.