Data Center Site Operations Manager/Sr. Manager
Job description
About the role
This position manages the complete operational lifecycle of a dedicated AI data center facility, starting from the commissioning phase and progressing through mature, full-scale functional operations. The hired individual is responsible for translating strategic infrastructure plans into daily execution ensuring that critical systems perform reliably and efficiently. They will establish rigorous operational frameworks that govern how the facility runs on a day to day basis. This role requires balancing technical depth with people management to solve complex physical and logistical problems. The manager will act as the primary operational leader ensuring that uptime, safety, and compliance are never compromised. They will oversee the coordination between technical teams and external partners to keep the data center aligned with business objectives. Success in this role is measured by the stability and performance of the infrastructure supporting high density computing workloads. The position demands a proactive mindset focused on preventing issues before they impact the business.
Key facts
What you'll do
Establish and lead a cross functional team of site managers, technicians, and contractors to maintain seamless data center operations around the clock.
Define and implement Standard Operating Procedures and Maintenance Methodologies that align with enterprise standards and industry best practices.
Monitor the performance and resiliency of critical infrastructure to ensure continuous uptime and service reliability for AI workloads.
Drive strict compliance with SOC2, ISO 27001, and OSHA safety standards through regular audits, thorough documentation, and timely corrective actions.
Champion the enforcement of security policies and data center best practices across all operational teams.
Partner with internal engineering groups and external technical vendors to align operational strategy and execute reconfigurable infrastructure changes.
Enable rapid deployment and integration processes specifically for AI and high performance computing environments.
Optimize facility level power, cooling, and network systems to support demanding compute and storage requirements.
Lead incident response efforts and coordinate root cause analysis to prevent future disruptions in the data center environment.
Develop and manage budgets, schedules, and resource plans to ensure operational efficiency and cost control.
Provide technical guidance and mentorship to senior level staff and junior team members to build internal capability.
Evaluate and integrate new technologies related to power, cooling, and IT equipment to future proof the data center operations.
Ensure accurate and up to date documentation of all operational procedures, configurations, and changes for audit and knowledge transfer purposes.
Serve as the key liaison between senior leadership and on site teams to communicate priorities, risks, and performance metrics.
Requirements
Hold a Bachelor's degree in computer science, electrical engineering, mechanical engineering, or physics from an accredited institution.
Possess 10 to 15 years of total engineering experience with at least 6 years dedicated to data center operations and management.
Demonstrate proven leadership in senior level roles that involve data center strategy, performance management, and oversight of large technical teams.
Bring hands on experience designing or managing cloud or on premises IT clusters including hardware, network, and power configurations.
Show a strong understanding of data center infrastructure components, reliability engineering principles, and operational risk management.
Have the ability to obtain and maintain necessary security clearances and work authorization in the United States.
Exhibit strong analytical, problem solving, and decision making skills under pressure in fast moving environments.
Display excellent written and verbal communication skills to effectively collaborate with technical and non technical stakeholders.
Nice to have
Show technical exposure to emerging power sources, advanced cooling technologies, and next generation IT equipment such as renewable energy, high power density racks, high voltage power supplies, and 2 phase cooling solutions.
Have working experience designing or managing CSP or crypto mining data centers, including interactions with OEMs and system integrators.
Possess in depth knowledge of GPU based systems, storage servers, state of the art liquid cooling technologies, and GPU cluster network topology.
Understand Generative AI fundamentals to support infrastructure decisions aligned with AI workloads.
Practical notes
This is an onsite role based in San Jose, California. The position requires authorization to work in the United States. The engagement is Full-Time. Typical interview steps for these data roles commonly include a SQL or coding exercise, a statistics question, and a case study. Candidates may be asked to design a metric, interpret an experiment, or build a small model. Some companies administer a take home analysis. Expect questions regarding past projects and the business impact of your work. Interviewers often assess how you communicate uncertainty and business impact, not only the technical correctness of the math. Bringing a clean write up of a past analysis to the interview is well received.
Good to know
Roles in semiconductor and data center operations often intersect engineering, facilities, and IT teams. Memory technologies such as DRAM and NAND flash underpin the infrastructure that powers cloud and AI workloads. Modern data centers increasingly adopt liquid cooling and high density architectures to support compute intensive workloads. Understanding of power, cooling, and network resilience is critical for large scale operations.