Data Center Compute Infrastructure
Job description
About the role
Join OpenAI's Scaling team to build and operate the compute infrastructure that powers advanced AI models. This role is about delivering reliable, efficient compute capacity at a scale that has few precedents in the industry. You will tackle complex challenges across hardware, software, and physical systems, from the chips in the rack to the power delivering them and the cooling keeping them steady. Your work directly contributes to the foundation for next-generation AI systems, including models like GPT-5.6. It is hard infrastructure engineering in the broadest sense, and the problems you face will regularly have no obvious precedent.
Because the company's growth depends on compute, this role sits close to the center of the mission. When you get a new cluster online faster, or make an existing one more stable, your work compounds across model training and product features. You will not be confined to a single discipline; the role deliberately spans software, hardware, operations, manufacturing, supply chain, and data center work, and the people who thrive here are comfortable moving between these domains.
Key facts
What you'll do
You will design, construct, and manage OpenAI's global compute infrastructure, addressing intricate issues across software, hardware, manufacturing, supply chain, and data center operations. You will work to enhance the dependability, speed, efficiency, and reliability of essential infrastructure, and you will collaborate with many teams to rapidly bring new compute resources online. Identifying and resolving bottlenecks in the technical, operational, and physical systems is central to the job.
You will create tools, processes, systems, and infrastructure that improve execution at scale, making it easier for the whole organization to move faster. You will think in terms of the long-term architectural and operational maturity of OpenAI's compute footprint, anticipating needs rather than just reacting to current demand. A strong theme throughout the work is a balance between urgent execution and durable engineering: you need to ship things that work today while building the foundation that works for years.
Requirements
OpenAI is looking for people with experience building, scaling, or managing complex technical systems. You should enjoy solving ambiguous, high-impact problems with a path that is not predefined, and be comfortable collaborating across diverse disciplines, including software, hardware, operations, and physical infrastructure. You need strong technical judgment and focus on execution, and you should prioritize reliability, speed, safety, and operational excellence in everything you build. The role is ideal for someone who is excited by infrastructure at an unprecedented scale and whose desire is for their work to directly support the development and deployment of frontier AI. This is a hands-on and outcome-oriented role, and it rewards people who take full ownership of large, messy problems.
Nice to have
- Experience with AI infrastructure, high-performance computing, distributed systems, GPU clusters, or cloud-scale platforms.
- Background in hardware systems, manufacturing, supply chain, data center development, or large capital infrastructure projects.
- Expertise in civil, controls, or mechanical, hardware, electrical, thermal, power, networking, or facilities engineering.
- Experience bringing new technical platforms, data centers, factories, or large-scale systems from concept to production.
- Experience operating in fast-paced environments where both technical depth and execution speed matter greatly.
Skills & tools
- Distributed systems
- ML infrastructure
- GPU fleets
- Power systems
- Cooling systems
- Networking
- Manufacturing
- Supply chain
- Data center delivery
Practical notes
This is a hybrid role based in San Francisco. OpenAI is an equal opportunity employer and administers background checks in accordance with applicable law. Qualified applicants with arrest or conviction records will be considered consistent with laws including the San Francisco Fair Chance Ordinance, the Los Angeles County Fair Chance Ordinance for Employers, and the California Fair Chance Act for US-based candidates. Reasonable accommodations for applicants with disabilities are available.
If you are an infrastructure operator or engineer who loves solving big, messy problems that cut across software and physical systems, this role is unusual in its breadth and in the scale of what you will be responsible for. You will help shape the compute foundation that enables frontier AI development for years to come, and you will see your decisions reflected in the performance of the systems the whole company depends on.