Operations Enablement Analyst, Data Center Operations
Job description
About the role
This position is dedicated to enhancing the operational efficiency of our data center fleet through rigorous process improvement and advanced data analysis. You will serve as the critical bridge between technical operations teams and business stakeholders to ensure our infrastructure can scale effectively to meet demand. The role involves owning key metrics that drive reliability and transparency across our physical infrastructure. You will be responsible for analyzing complex operational data to uncover bottlenecks and streamline fleet management workflows. Additionally, you will develop and maintain essential documentation for standard operating procedures within the data center environment. Collaboration with cross-functional partners will be central to refining node and fleet lifecycle management processes for optimal performance. You will also monitor performance metrics proactively to identify issues before they impact operations. This position requires a focus on supporting our specialized GPU compute and storage infrastructure as an AI-native cloud provider.
Key facts
What you'll do
- Analyze operational data to identify bottlenecks and improve fleet management workflows using advanced analytical methods.
- Develop and maintain comprehensive documentation for standard operating procedures within the data center environment to ensure consistency.
- Collaborate with cross-functional teams to refine node and fleet lifecycle management processes for maximum efficiency.
- Monitor performance metrics to ensure infrastructure reliability and operational transparency through continuous observation.
- Assist in the implementation of tools that support large-scale AI infrastructure operations to drive automation.
- Translate complex technical data into actionable operational strategies that align with business objectives and growth targets.
- Design process improvements that optimize the utilization of hardware resources across distributed data center locations.
- Coordinate with technical teams and stakeholders to resolve operational issues and enhance communication protocols.
- Evaluate the effectiveness of existing workflows and recommend adjustments based on empirical data analysis and trends.
- Support the development of dashboards and reporting mechanisms that provide real-time visibility into operational status.
- Facilitate knowledge transfer sessions to ensure best practices are understood and adopted across the operations team.
- Drive initiatives that reduce downtime and improve the overall efficiency of AI infrastructure operations.
- Partner with engineering teams to ensure operational requirements are integrated into infrastructure deployment strategies.
- Maintain detailed records of operational changes and their impact on performance metrics for future reference.
Requirements
- Proven experience in data center operations or a related technical infrastructure field that demonstrates a history of successful execution.
- Strong analytical skills with the ability to translate complex data into actionable operational strategies that deliver measurable results.
- Proficiency in managing large-scale hardware lifecycles and infrastructure monitoring to ensure optimal performance and uptime.
- Ability to work effectively in a fast-paced, high-growth environment where priorities evolve rapidly and demands increase.
- Excellent communication skills for coordinating between technical teams and stakeholders to ensure alignment and understanding.
- Demonstrated capability to manage multiple projects simultaneously while maintaining attention to detail and accuracy.
- Experience with infrastructure monitoring tools and methodologies to track performance and identify issues proactively.
- A commitment to maintaining the highest standards of operational excellence and continuous improvement in all workflows.
Nice to have
- Experience with Kubernetes-based infrastructure or cloud-native environments that can enhance deployment flexibility.
- Familiarity with AI/ML workload requirements and GPU-based compute clusters to better support specialized infrastructure needs.
- Background in process automation or workflow optimization within physical data centers to drive efficiency gains.
Practical notes
CoreWeave is an AI-native cloud provider. This position requires a focus on supporting our specialized GPU compute and storage infrastructure. Candidates should be prepared to work within a distributed team structure across the listed US locations. The role is full-time and requires availability during standard business hours as needed for operational support. No specific travel requirements are outlined for this position. Employment is contingent upon the successful completion of any required background checks and authorization to work in the United States. Candidates must have the right to work in the United States without sponsorship for this role.