
Senior Technical Program Manager, Machine Learning Infrastructure
Job description
About the role
You will manage the program portfolio covering inference, efficiency, serving, and endpoints to guarantee that Cohere's infrastructure continues to scale for a rapidly expanding user base of internal and external users. You will lead the end-to-end coordination and execution across the program, with a particular focus on cross-functional collaboration with Modeling and customer-facing teams. You will identify pain points and establish processes to ensure that the team can focus on development, while meeting needs of internal and external users of the models, and improving engineering best practices. You will build a strong culture around continuous improvement, such as liaising with incident management leads, ensuring that problems are accurately root caused, and that required fixes are provided in a timely manner. You will manage various overlapping projects and programs, ruthlessly prioritizing asks, to ensure that the company's top priorities are met. You will collaborate with stakeholders across the company to set, track, and manage timelines, deliverables, budgets, and scope to ensure successful program execution. You will deliver clear, timely, and consistent updates across engineering, leadership, and non-technical teams on the progress, plans, and incidents. You will act as a strong tactical and strategic partner to the senior leads across your program, whether that be providing low-level tactical support or help answer high-level strategic problems that drives impact at the company-level.
Key facts
What you'll do
Manage the program portfolio covering inference, efficiency, serving, and endpoints to guarantee that Cohere's infrastructure continues to scale for a rapidly expanding user base of internal and external users.
Lead the end-to-end coordination and execution across the program, with a particular focus on cross-functional collaboration with Modeling and customer-facing teams.
Identify pain points and establish processes to ensure that the team can focus on development, while meeting needs of internal and external users of the models, and improving engineering best practices.
Build a strong culture around continuous improvement, such as liaising with incident management leads, ensuring that problems are accurately root caused, and that required fixes are provided in a timely manner.
Manage various overlapping projects and programs, ruthlessly prioritizing asks, to ensure that the company's top priorities are met.
Collaborate with stakeholders across the company to set, track, and manage timelines, deliverables, budgets, and scope to ensure successful program execution.
Deliver clear, timely, and consistent updates across engineering, leadership, and non-technical teams on the progress, plans, and incidents.
Act as a strong tactical and strategic partner to the senior leads across your program, whether that be providing low-level tactical support or help answer high-level strategic problems that drives impact at the company-level.
Work closely with security and compliance teams to ensure that all infrastructure programs adhere to enterprise standards and regulatory requirements.
Partner with data platform teams to optimize data flows, storage strategies, and monitoring for ML workloads across deployment environments.
Champion the use of observability and metrics to drive decision-making and highlight trade-offs between speed, reliability, and cost.
Facilitate alignment on success metrics and key results across engineering, product, and operations teams for ML infrastructure initiatives.
Support scenario planning and capacity forecasting to guide infrastructure investment and resource allocation for model training and serving.
Drive post-incident reviews and translate learnings into actionable process and tooling improvements.
Requirements
5+ years of Technical Program Management experience focusing on Machine Learning Infrastructure, specifically covering areas such as model inference, serving, efficiency, and endpoints design & implementation.
In-depth technical knowledge around ML infrastructure design and implementation, as well as engineering best practices, supporting a balance of velocity and structure in rapidly growing organizations.
Experience working in a chaotic, fast paced, low structure environment. We need an EPM who is pragmatic, can roll up their sleeves when needed, and can function at the tactical and strategic level to do whatever it takes to make Cohere's models succeed.
(optional, but strongly preferred) Have technical experience in a hands-on capacity. We don't necessarily need someone who has been working hands-on in a technical role, but the programs are inherently extremely technical. You need to dive into the depths of internal systems, and develop a reputation as a subject matter expert who brings in engineering best practices and the right amount of structure without slowing the team down.
Exceptional written and verbal communication skills, with the ability to translate complex technical concepts to diverse audiences.
Strong ability to operate independently and make sound decisions with incomplete information while maintaining transparency and stakeholder trust.
Proven track record of managing priorities and delivering results in ambiguous and fast-moving environments.
Deep commitment to process improvement, documentation, and knowledge sharing to elevate the entire organization.
Nice to have
Preferred experience with cloud infrastructure platforms and ML deployment tooling.
Familiarity with monitoring and observability tools for ML systems.
Understanding of security and compliance considerations for enterprise AI deployments.
Practical notes
Please note: While we are remote-first and accept candidates from all over the world, we are strongly prioritizing hiring for candidates in the Eastern Time (ET) timezone.