Technical Program Manager, Infrastructure
Job description
Technical Program Manager, Infrastructure at Fal Ai.
About the role
You will own the operational and commercial performance of the hyperscalers, neoclouds, and GPU infrastructure providers that power fal's compute fleet. You will translate contract commitments into measurable operating standards, resolve performance issues, recover credits and remedies, and ensure our capacity is ready when customers need it. This is a hands-on role for someone who understands both infrastructure operations and commercial vendor management. You will build the systems, reporting, and operating cadence for a vendor portfolio that is growing quickly in scale and complexity. You partner closely with Infrastructure, Capacity Planning, Finance, Legal, and executive stakeholders to align capacity with demand forecasts and surface supply risks early. You drive post-signature relationships to ensure vendors meet stringent requirements for generative media workloads. You establish the governance backbone that allows fal to scale its infrastructure footprint without compromising reliability or performance.
Key facts
What you'll do
- Orchestrate post-signature relationships with GPU infrastructure providers, including hyperscalers, neoclouds, colocation partners, and other strategic compute vendors.
- Convert contractual commitments into measurable SLAs, scorecards, escalation procedures, and internal operating requirements.
- Monitor availability, provisioning timelines, capacity delivery, support responsiveness, and other vendor-performance indicators.
- Lead cross-functional escalations when vendors miss commitments, coordinating with Infrastructure, Capacity Planning, Finance, Legal, and executive stakeholders.
- Drive pursuit of service credits, remedies, and corrective-action plans, ensuring issues remain open until they are fully resolved.
- Execute vendor governance, including regular operating reviews, executive business reviews, performance reporting, renewal readiness, and contract-compliance tracking.
- Collaborate with Capacity Planning and Strategic Sourcing to align committed capacity with demand forecasts and identify supply risks early.
- Reconcile invoices against contracted pricing, delivered capacity, usage, credits, and other commercial terms with precision.
- Maintain accurate records of commitments, renewals, open claims, invoices, and vendor obligations in reliable systems.
- Construct vendor-management processes, dashboards, and operating cadence tailored to the needs of a scaling infrastructure footprint.
- Define and operationalize performance metrics that reflect the realities of generative media workloads on hyperscale and neocloud platforms.
- Partner with engineering teams to translate operational impact of vendor failures into clear remediation steps and communication.
- Establish reporting cadences that provide timely insight into vendor health, risk, and compliance across the portfolio.
- Implement tracking mechanisms for credits, remedies, and corrective actions to ensure closure and accountability.
- Proactively identify trends in vendor performance that could affect future capacity strategies or product roadmaps.
Requirements
- Bring 5+ years in infrastructure vendor management, cloud partnerships, strategic sourcing, procurement operations, or a related function.
- Demonstrate direct experience managing commercial relationships with hyperscalers, neoclouds, GPU providers, data-center operators, or other large-scale infrastructure vendors.
- Show history of operationalizing complex infrastructure agreements involving committed capacity, SLAs, service credits, provisioning requirements, renewals, and commercial remedies.
- Prove a track record of resolving vendor-performance issues while preserving long-term strategic relationships.
- Exhibit strong analytical and operational skills, including performance reporting, invoice reconciliation, credit calculations, and contract-compliance tracking.
- Possess enough technical fluency to work effectively with infrastructure engineering and capacity-planning teams and understand the operational impact of vendor failures.
- Have substantial experience running QBRs, vendor scorecards, executive escalations, and corrective-action plans in enterprise settings.
- Display comfort building processes and systems from scratch in a fast-moving, high-growth environment.
- Communicate clearly and directly, with excellent attention to detail in complex operational and commercial contexts.
- Thrive in situations where precise ownership of vendor performance translates into material business outcomes for a generative media platform.
- Navigate ambiguity and drive forward initiatives that span multiple stakeholders and time zones.
- Balance proactive problem-solving with disciplined follow-through on contractual and commercial obligations.
- Apply judgment to prioritize vendor issues based on impact to customers, capacity, and financial performance.
- Maintain integrity and transparency in all interactions with internal and external stakeholders.
Nice to have
- Experience managing GPU capacity across multiple providers in concurrent engagements.
- Experience allocating constrained capacity during shortages or periods of rapid demand growth for generative media workloads.
- Familiarity with GPU availability, provisioning, utilization, networking, and cloud infrastructure economics at scale.
- Experience recovering material service credits or negotiating remedies following major performance failures with hyperscalers or neoclouds.
- Experience establishing a vendor-management function at a hyperscaler, neocloud, AI infrastructure company, or high-growth technology company.
Practical notes
The role is full-time based in San Francisco.