AI Infrastructure Data Center Deployment Lead
Job description
About the role
Lambda is seeking a leader to manage the physical deployment of AI infrastructure within our data centers. You will be instrumental in building out our high-performance computing environments, ensuring they meet stringent operational standards. This role involves direct oversight of site operations and collaboration with various engineering and supply chain teams. You will own the end-to-end execution of complex builds within highly regulated environments. The position requires a hands-on leader who is comfortable managing critical timelines and resource allocation. Success in this role depends on your ability to drive standards and consistency across distributed sites. You will play a key part in enabling the compute capacity that powers our AI and machine learning initiatives. This is a leadership position where your decisions directly impact the reliability and scalability of our infrastructure.
About the role
Lambda is seeking a leader to manage the physical deployment of AI infrastructure within our data centers. You will be instrumental in building out our high-performance computing environments, ensuring they meet stringent operational standards. This role involves direct oversight of site operations and collaboration with various engineering and supply chain teams. You will own the end-to-end execution of complex builds within highly regulated environments. The position requires a hands-on leader who is comfortable managing critical timelines and resource allocation. Success in this role depends on your ability to drive standards and consistency across distributed sites. You will play a key part in enabling the compute capacity that powers our AI and machine learning initiatives. This is a leadership position where your decisions directly impact the reliability and scalability of our infrastructure.
What you'll do
Oversee the physical installation and setup of major AI and machine learning clusters in data cages. Guide the construction and expansion of Lambda's data center facilities, ensuring alignment with power, cooling, and cabling requirements. Collaborate with Hardware Engineering, Supply Chain, and Operations teams to define and implement deployment standards. Work with cross-functional groups to address challenges and optimize regional data center operations. Interface with external vendors to ensure data center builds meet precise specifications. Understand and contribute to the configuration and integration of compute, storage, and networking components for AI/ML systems. Manage detailed deployment schedules and coordinate tasks across multiple stakeholders to maintain momentum. Lead site readiness assessments and validate that infrastructure meets technical and safety requirements. Monitor installation quality and perform reviews to ensure compliance with internal and external standards. Drive process improvements for tracking equipment delivery, installation, and commissioning. Provide technical guidance on data center layout, airflow management, and cable routing strategies. Support the integration of new hardware into existing environments with minimal disruption. Document procedures and configurations to create clear references for operations and support teams. Act as a primary point of contact for deployment-related inquiries and issue resolution.
What you'll do
Oversee the physical installation and setup of major AI and machine learning clusters in data cages. Guide the construction and expansion of Lambda's data center facilities, ensuring alignment with power, cooling, and cabling requirements. Collaborate with Hardware Engineering, Supply Chain, and Operations teams to define and implement deployment standards. Work with cross-functional groups to address challenges and optimize regional data center operations. Interface with external vendors to ensure data center builds meet precise specifications. Understand and contribute to the configuration and integration of compute, storage, and networking components for AI/ML systems. Manage detailed deployment schedules and coordinate tasks across multiple stakeholders to maintain momentum. Lead site readiness assessments and validate that infrastructure meets technical and safety requirements. Monitor installation quality and perform reviews to ensure compliance with internal and external standards. Drive process improvements for tracking equipment delivery, installation, and commissioning. Provide technical guidance on data center layout, airflow management, and cable routing strategies. Support the integration of new hardware into existing environments with minimal disruption. Document procedures and configurations to create clear references for operations and support teams. Act as a primary point of contact for deployment-related inquiries and issue resolution.
Requirements
A minimum of 5 years of experience in data center operations and infrastructure integration, with prior leadership responsibilities. Demonstrated experience managing large-scale data center and infrastructure deployments. Solid understanding of data center design principles, deployment processes, and architectural considerations. Proficiency in deployment planning and execution. Familiarity with industry standards and regulations governing data center operations. Strong analytical and problem-solving abilities to resolve data center-related issues. Effective communication and interpersonal skills for collaboration with diverse teams. The ability to work independently and make sound decisions in a fast-paced environment. A commitment to maintaining the highest levels of safety and quality in all deployment activities. Willingness to adapt to evolving priorities and shifting project requirements as business needs change. Strong organizational skills with the ability to manage multiple workstreams simultaneously. A disciplined approach to tracking progress, risks, and dependencies. Readiness to travel as required for project support or vendor engagements. Compliance with company policies and data center security protocols.
Requirements
A minimum of 5 years of experience in data center operations and infrastructure integration, with prior leadership responsibilities. Demonstrated experience managing large-scale data center and infrastructure deployments. Solid understanding of data center design principles, deployment processes, and architectural considerations. Proficiency in deployment planning and execution. Familiarity with industry standards and regulations governing data center operations. Strong analytical and problem-solving abilities to resolve data center-related issues. Effective communication and interpersonal skills for collaboration with diverse teams. The ability to work independently and make sound decisions in a fast-paced environment. A commitment to maintaining the highest levels of safety and quality in all deployment activities. Willingness to adapt to evolving priorities and shifting project requirements as business needs change. Strong organizational skills with the ability to manage multiple workstreams simultaneously. A disciplined approach to tracking progress, risks, and dependencies. Readiness to travel as required for project support or vendor engagements. Compliance with company policies and data center security protocols.
Nice to have
Experience with hardware engineering, supply chain management, and inventory systems like NetBox or Jira. Knowledge of high-performance computing technologies and their integration into data center environments. Expertise in the architecture of DGX and HGX based systems.
Skills & tools
Data center operations, infrastructure deployment, AI/ML infrastructure, hardware integration, network configuration, power and cooling systems, cabling infrastructure, project management, vendor management.
Practical notes
This role is for full-time employment. Lambda offers competitive cash and equity compensation, comprehensive health, dental, and vision coverage, a 401k plan with a 2% company match for US employees, and a flexible paid time off policy.