Data Center Site Operations Manager
Job description
About the role
Lightning AI is the company behind PyTorch Lightning. Founded in 2019, we build an end-to-end platform for developing, training, and deploying AI systems - designed to take ideas from research to production with less friction. Through our merger with Voltage Park, a neocloud and AI Factory, Lightning AI combines developer-first software with cost-efficient, large-scale compute. Teams get the tools they need for experimentation, training, and production inference, with security, observability, and control built in. We serve solo researchers, startups, and large enterprises. Lightning AI operates globally with offices in New York City, San Francisco, Seattle, and London, and is backed by Coatue, Index Ventures, Bain Capital Ventures, and Firstminute.
We are hiring a Data Center Site Operations Manager to lead operations at our Allen and Fort Worth facilities. You will own the end-to-end operational readiness and daily execution for our data center facilities in Texas, ensuring safe, scalable, and efficient deployment of our compute infrastructure. You will lead cross-functional coordination across engineering, construction, and vendor teams to align on timelines, compliance, and performance standards. You will build and mentor a high-performing operations team that delivers reliable uptime and rigorous process adherence. You will drive continuous improvement in workflows, safety protocols, and incident response mechanisms across the site. You will oversee the integration of power, cooling, and network systems specific to our deployment within existing colocation footprints. You will maintain strict inventory and vendor management for all equipment and services critical to operations. You will implement monitoring and reporting frameworks that provide clear visibility into site health and capacity. You will partner with leadership to scale processes as we grow from initial rollout to expanded production capacity.
What you'll do
Coordinate the tenant fit-out and operational readiness of our data center space in Allen and Fort Worth. Oversee the installation and commissioning of power, cooling, and network systems critical to 10MW deployment. Manage vendor relationships, scheduling, and performance for rack and roll delivery, cabling, and power-up procedures. Lead the setup of monitoring, alerting, and inventory systems to track site health and capacity in real time. Implement operational procedures and runbooks that ensure consistent, repeatable execution across shifts. Drive safety compliance and enforce protocols to maintain a secure and incident-free environment for staff and equipment. Collaborate with engineering and deployment teams to align infrastructure with workload requirements and growth plans. Track and manage all inventory, service tickets, and vendor SLAs to ensure accountability and timely resolution. Support the onboarding of new team members and provide training to uphold operational standards. Coordinate travel to other data center sites across the US as required for cross-site alignment and support. Facilitate communication between technical and non-technical stakeholders to maintain clarity on timelines and risks. Own the day-to-day management of the facility to ensure uptime, scalability, and adherence to company standards.
Requirements
Must be authorized to work in the United States without sponsorship for this position. Must reside in or be willing to relocate to Allen, Texas, and/or Fort Worth, Texas. Must commit to an on-site work model in Allen, Texas and Fort Worth, Texas with travel to other US sites. Must be available to work full-time during standard business hours with flexibility for on-call responsibilities as needed. Must have experience leading operations in a data center or high-tech facility environment. Must demonstrate strong leadership skills with a proven track record of managing and developing teams. Must possess excellent written and verbal communication skills for interacting with internal and external partners. Must be comfortable working with technical stakeholders and interpreting infrastructure requirements.
Nice to have
Experience in AI or GPU-intensive data center environments. Background in facilities or colocation management. Familiarity with deployment tools for large-scale compute infrastructure.
Practical notes
This is an on-site role based in Allen, Texas and Fort Worth, Texas. Travel to other US data center sites will be required. No sponsorship is available for this position. The compensation range is $140,000.00 to $180,000.00 per year, commensurate with experience.