Engineering Manager, Model Flywheel
Job description
About the role
The ChatGPT Model Flywheel group sits at the intersection of research breakthroughs and the product experiences delivered to millions of users. This team owns the critical infrastructure that transforms raw model advancements into reliable, safe, and continuously improving ChatGPT and Codex capabilities. You will lead engineers responsible for the full lifecycle of model integration, spanning automated experimentation frameworks, staged rollout tooling, capacity orchestration, and comprehensive measurement systems. Success in this position requires bridging deep technical execution in distributed systems with close collaboration across Research, Data Science, Inference, and Product organizations. The scope involves defining the technical strategy for how models are validated, deployed, and measured at a global scale.
Key facts
What you'll do
- Architect and unify the core harness, context management pipelines, and system prompt frameworks that underpin every ChatGPT model interaction, ensuring consistency and reliability across product surfaces.
- Drive the expansion of multi-tier model experiences, enabling the product to route requests intelligently across different model capabilities and sizes based on latency, cost, and quality requirements.
- Build and scale self-serve experimentation platforms that allow researchers and engineers to validate model changes rapidly, incorporating automated guardrails that prevent regressions in safety or quality before changes reach production.
- Own the end-to-end model rollout automation strategy, designing progressive delivery mechanisms, feature flagging systems, and instant rollback capabilities to minimize blast radius during deployments.
- Develop sophisticated capacity management tooling that dynamically allocates inference compute across model variants, integrating with the Fleet and Inference teams to optimize utilization and cost efficiency.
- Implement platform-wide health monitoring and observability stacks that provide real-time visibility into model serving performance, error rates, and latency distributions across global regions.
- Shape comprehensive measurement and evaluation systems, combining offline benchmarks, human grader signals, live user feedback loops, and A/B test scorecards to create a single source of truth for model quality.
- Establish continuous improvement flywheels where production insights automatically inform retraining priorities, evaluation set curation, and deployment criteria for subsequent model iterations.
- Partner directly with the Model Measurement Data Science team to define statistically rigorous evaluation methodologies and launch criteria that balance velocity with safety.
- Collaborate with the Codex and API teams to ensure deployment tooling and measurement infrastructure are reusable and consistent across the entire OpenAI product portfolio.
- Mentor and grow a high-performing engineering team, fostering a culture of operational excellence, psychological safety, and technical rigor in a fast-paced research-adjacent environment.
- Represent the team in cross-functional planning cycles, translating high-level product and research roadmaps into concrete infrastructure milestones and resource requests.
Requirements
- Proven track record leading engineering teams (typically 5+ years of management experience) in complex, cross-functional environments where software intersects with research or ML workflows.
- Demonstrated success shipping and operating large-scale production backend systems, preferably involving high-throughput inference serving, distributed computing, or real-time data processing.
- Deep practical understanding of the ML model deployment lifecycle, including versioning, canary analysis, shadow traffic validation, and model registry management.
- Strong fluency in building measurement and experimentation infrastructure, such as A/B testing frameworks, evaluation pipelines, or observability platforms for ML systems.
- Exceptional communication skills with a history of aligning diverse stakeholders - including Researchers, Product Managers, Data Scientists, and Infrastructure engineers - on technical strategy and execution priorities.
- Ability to navigate ambiguity and define technical roadmaps in a rapidly evolving domain where requirements shift as model capabilities advance.
- Experience managing on-call rotations and incident response processes for critical user-facing services with strict reliability targets (SLOs/SLAs).
- Proficiency in at least one major systems programming language (e.g., Python, Go, Rust, C++) and comfort with modern cloud-native tooling (Kubernetes, Terraform, CI/CD pipelines).
Nice to have
- Direct prior experience working with Large Language Models (LLMs) in production, including prompt engineering pipelines, context window optimization, or token-level streaming architectures.
- Background in building experimentation platforms specifically designed for ML model comparison, such as interleaving, side-by-side evaluation tooling, or reward model training loops.
- Familiarity with the unique challenges of GPU fleet management, inference optimization (e.g., quantization, speculative decoding, KV cache management), or hardware accelerator orchestration.
- Previous tenure at an AI research lab or a company deploying frontier models at consumer scale (hundreds of millions of users).
- Contributions to open-source projects in the MLOps, LLMOps, or model serving ecosystem (e.g., vLLM, Triton, MLflow, Kubeflow).
- Experience collaborating closely with Safety or Policy teams to implement automated content filtering, refusal analysis, or red-teaming infrastructure into the deployment pipeline.
Skills & tools
Python, Go, Kubernetes, Terraform, Prometheus/Grafana, Distributed Systems, LLM Serving (vLLM/Triton), A/B Testing Frameworks, MLflow, Model Registry, Feature Flagging, Capacity Planning, Incident Management, Cross-functional Leadership, Technical Strategy.
Practical notes
This is a hybrid role based in San Francisco, requiring regular in-office collaboration (typically 3 days per week) with the broader Applied AI, Research, and Infrastructure organizations. The compensation package includes a base salary range of $293,000 to $385,000 annually, supplemented by a significant equity grant (PPUs) and comprehensive benefits. Visa sponsorship is available for qualified candidates, including H-1B transfers and O-1 petitions; the legal team manages Green Card processing (PERM) for eligible employees after onboarding. The interview process generally consists of a recruiter screen, a hiring manager conversation, a system design session focused on ML infrastructure, a coding or debugging exercise, a behavioral leadership assessment, and a final cross-functional panel. OpenAI provides relocation assistance for candidates moving to the Bay Area. Background checks are conducted in compliance with the San Francisco Fair Chance Ordinance and California Fair Chance Act.
META
Company: OpenAI
Title: Engineering Manager, Model Flywheel
Listed
location: San Francisco
Job type: full_time
Department: Applied AI
Employment: FullTime
Workplace: Hybrid
Compensation JSON: {"compensationTierSummary":"$293K - $385K - Offers Equity","scrapeableCompensationSalarySummary":"$293K - $385K","compensationTiers":[{"id":"b21db7f6-baed-4973-84e7-fb2db6ec73f4","tierSummary":"$293K - $385K - Offers Equity","title":null,"additionalInformation":null,"components":[{"id":"4b9309ae-9a3d-4a86-ab7c-d6a5ce3e2b76","summary":"$293K - $385K","compensationType":"Salary","interval":"1 YEAR",
About the company
OpenAI is an AI research and deployment company dedicated to ensuring that general-purpose artificial intelligence benefits all of humanity. We push the boundaries of the capabilities of AI systems and seek to safely deploy them to the world through our products.