Senior Software Engineer, Observability
Job description
About the role
You will design and build monitoring systems that provide visibility into our large-scale AI infrastructure. This role focuses on developing tools that track performance, reliability, and health across our GPU-accelerated cloud environment. You will own the end-to-end lifecycle of critical observability platforms, from initial design through production deployment and ongoing optimization. The work you do will directly impact the reliability and performance of AI compute services used by some of the largest enterprises and research institutions. You will partner closely with cross-functional teams to ensure monitoring coverage aligns with evolving product and infrastructure needs. Your contributions will help create a single pane of glass for complex, distributed AI workloads running in our data centers. You will play a key role in driving standardization and best practices for observability across the entire engineering organization.
Key facts
What you'll do
- Architect and implement observability solutions tailored for our distributed GPU cloud platform.
- Engineer data pipelines that reliably collect, process, and store infrastructure metrics and event streams.
- Construct dashboards and alerting rules that surface critical issues before they impact customers or workloads.
- Integrate monitoring instrumentation into node lifecycle controllers and fleet management components.
- Analyze trends in infrastructure health data to drive improvements in reliability and efficiency.
- Collaborate with SRE and platform teams to standardize metrics, logging, and tracing formats across services.
- Optimize the performance and scalability of the monitoring stack as our data center footprint expands globally.
- Work with security and compliance stakeholders to ensure observability data handling meets regulatory requirements.
- Support on-call rotations to investigate incidents and improve detection logic based on real-world failure patterns.
- Mentor junior engineers on instrumentation strategies and troubleshooting methodologies for complex systems.
Requirements
- You possess professional experience in software engineering with a demonstrated focus on observability or distributed systems.
- You have a proven track record of building and scaling monitoring infrastructure for cloud-native environments.
- You are able to work effectively within a team dedicated to infrastructure reliability and continuous improvement.
- You bring experience with high-performance computing or large-scale AI infrastructure and its unique monitoring challenges.
- You have hands-on experience writing production-grade code in languages commonly used for infrastructure tooling.
- You understand the fundamentals of metrics, logs, and traces and how they relate to system health and performance.
- You have experience designing systems that operate at scale with high availability and low latency requirements.
- You are comfortable collaborating with technical stakeholders across engineering, operations, and product teams.
Nice to have
- Experience with Kubernetes and fleet management systems and their associated observability needs.
- Background in developing internal tooling for cloud service providers or infrastructure platforms.
- Familiarity with metrics, tracing, and logging stacks at scale, including their deployment and operation.
- Prior work in AI or machine learning infrastructure environments where observability is critical.
- Contributions to open source projects related to monitoring, logging, or distributed tracing.
Skills & tools
- Observability and monitoring frameworks such as Prometheus, Grafana, or similar platforms.
- Distributed systems architecture and the challenges of debugging complex, networked environments.
- Cloud infrastructure management tools and provider APIs, including networking and compute resources.
- GPU-accelerated compute environments and the operational nuances of managing them at scale.
- Scripting and programming skills for automating data collection and analysis tasks.
- Experience with structured logging, log aggregation platforms, and search systems.
- Understanding of alerting strategies, incident response processes, and post-incident review practices.
Practical notes
- CoreWeave is an AI-native cloud provider focused on large-scale GPU infrastructure.
- Please apply through the official careers portal to be considered for this position.
- This is a full-time role based in either New York, NY or Sunnyvale, CA, with standard business hours expectations.
- No current information on compensation, bonuses, or benefits is provided in this listing.
- Candidates must be authorized to work in the United States without sponsorship for this role.
- Relocation details, if applicable, will be discussed during the hiring process with the hiring team.
- Travel requirements are not expected as part of the standard responsibilities for this position.
- The posting will remain open until the role is filled, and early application is encouraged to be considered.