Industrial Compute
openaiRemote (USA)Full Time2d ago
OpenAIGPTAIMLSecurityComplianceOperationsSupportSolutionsEngineeringInfrastructureReliability
Job description
Industrial Compute at OpenAI
About the role
Join a team focused on building the foundational compute infrastructure that powers OpenAI's most advanced AI models. You will tackle complex challenges across hardware, software, and physical systems to deliver immense, reliable, and efficient computing power. This is an opportunity to contribute to the next generation of AI infrastructure at an unprecedented scale.
Key facts
What you'll do
- Contribute to the construction, expansion, and ongoing management of OpenAI's worldwide compute systems.
- Address intricate issues spanning software, hardware, manufacturing supply chains, and data center operations.
- Enhance the dependability, throughput, efficiency, and scalability of essential infrastructure.
- Collaborate with various departments to rapidly and dependably bring new computing resources online.
- Identify and resolve constraints within technical, operational, and physical frameworks.
- Develop tools, procedures, systems, or infrastructure to boost large-scale execution.
- Inform the long-term architectural direction and operational maturity of OpenAI's compute footprint.
Requirements
- Possess experience in building, scaling, or managing complex technical systems.
- Thrive on tackling ambiguous, high-impact problems where solutions are not always predetermined.
- Are adept at collaborating across diverse disciplines, including software, hardware, operations, and physical infrastructure.
- Demonstrate strong technical judgment and a proactive approach to execution.
- Prioritize reliability, speed, safety, and operational excellence.
- Are motivated by the challenge of constructing infrastructure at a massive scale.
- Desire to directly support the advancement and deployment of frontier AI.
Nice to have
- Experience with AI infrastructure, high-performance computing, distributed systems, GPU clusters, or cloud-scale platforms.
- Background in hardware systems, manufacturing, supply chain, data center development, or large capital infrastructure projects.
- Expertise in civil, controls, mechanical, hardware, electrical, thermal, power, networking, or facilities engineering.
- Experience bringing new technical platforms, data centers, factories, or large-scale systems from inception to production.
- Experience operating in rapidly evolving environments requiring both technical depth and swift execution.
Skills & tools
- Distributed Systems
- ML Infrastructure
- GPU Fleets
- Power and Cooling
- Networking
- Manufacturing and Supply Chain
- Data Center Delivery
- Software Engineering
- Hardware Engineering
- Operations
Practical notes
- Visa sponsorship may be available for qualified candidates.
- Occasional travel may be required.
- Benefits include equity and comprehensive health coverage.
- Applications will be considered in accordance with applicable laws, including fair chance ordinances in California and Los Angeles County.