Technical Program Manager, Compute Qualification
Job description
About the role
Together AI is on a mission to enhance its compute capabilities and is seeking a Technical Program Manager to play a pivotal role in this initiative. This position is crucial for ensuring that all new compute resources align with our high technical standards. You will be responsible for overseeing the thorough assessment of potential compute providers, ensuring they are equipped to handle customer workloads effectively. Your role will involve managing a structured evaluation process that encompasses various aspects, including compute, networking, storage, power, cooling, and operational readiness.
Key facts
What you'll do
- Lead the comprehensive process of evaluating new compute resources, starting from initial outreach to providers and culminating in the final acceptance decision.
- Handle multiple evaluations of providers simultaneously, setting timelines, tracking progress, and ensuring all stakeholders are kept up-to-date on requirements and deadlines.
- Collaborate closely with engineering teams specializing in infrastructure, networking, data centers, and site reliability to facilitate technical validation and translate findings into actionable recommendations for leadership.
- Review technical documentation and survey responses from providers for completeness and accuracy, pinpointing any gaps, inconsistencies, or potential risks that may need further exploration.
- Conduct preliminary analyses of provider data, comparing specifications, validating performance claims, and identifying issues before they are escalated to engineering experts.
- Develop and sustain documentation, templates, and standards that outline our requirements for compute, networking, storage, power, cooling, and operational support.
- Establish a systematic and auditable record of evaluation outcomes to support sourcing decisions and enable the qualification process to scale efficiently.
- Ensure that all evaluations are conducted in accordance with industry best practices and internal guidelines, maintaining a high level of quality throughout the process.
- Facilitate regular meetings with stakeholders to discuss progress, challenges, and next steps in the evaluation process.
- Provide training and support to junior team members involved in the evaluation process, fostering a collaborative team environment.
- Stay updated on industry trends and advancements in compute technology to continuously improve the evaluation process and criteria.
Requirements
- At least 5 years of experience in technical program management, infrastructure program management, or a similar technical operations role, ideally with experience in hardware, data centers, or large-scale compute environments.
- Proven track record of managing multiple complex, cross-functional projects to successful completion within established timelines, showcasing strong organizational skills and the ability to engage diverse stakeholders.
- Strong technical knowledge of data center infrastructure, including server and GPU hardware, high-performance networking (such as InfiniBand or Ethernet), storage systems, and the basics of power and cooling. You should be adept at reviewing detailed technical specifications and identifying areas that require further investigation.
- Experience working with data, including the ability to write scripts or queries (e.g., in Python or SQL) to independently compare, validate, and analyze provider specifications and test results.
- Exceptional written and verbal communication skills, with the ability to simplify complex technical information into clear, actionable recommendations for both technical teams and executive leadership.
- Willingness to travel to provider sites and data centers as necessary.
Nice to have
- Experience in qualifying, commissioning, or accepting GPU clusters or high-performance computing infrastructure against established performance and reliability benchmarks.
- Familiarity with the infrastructure utilized for AI training and inference, including interconnect architectures, cluster setup, and acceptance testing protocols.
- Background in designing AI or high-performance computing clusters.
- Previous experience collaborating directly with hardware vendors, colocation providers, or cloud capacity providers.
Skills & tools
- Proficiency in Python
- Experience with SQL
Practical notes
- US base salary range: $200,000 - $250,000
- Startup equity options
- Comprehensive health insurance coverage
- Additional benefits available
- Flexibility in remote work arrangements
Join Together AI and contribute to our mission of advancing compute capabilities while ensuring the highest standards of quality and performance. If you are a driven individual with a passion for technology and a knack for program management, we would love to hear from you!