Senior Engineering Manager, Machine Learning Platform
Job description
About the role
Affirm is reinventing credit to make it more honest and friendly, giving consumers the flexibility to buy now and pay later without any hidden fees or compounding interest. We are seeking a Senior Engineering Manager to lead our ML Training & Serving team. This is a senior technical leadership role with company-wide impact. You will directly lead a team of platform engineers, including senior technical leaders, and combine strong people leadership with deep ML infrastructure judgment. In close partnership with senior ICs, you will shape the platform's strategy and technical direction. You will own the end to end lifecycle of our machine learning infrastructure, balancing ambitious innovation with the reliability our business demands. Your decisions will directly influence how quickly Affirm can deploy new AI capabilities while maintaining the highest standards of performance and security. You are expected to build a cohesive, high-performing engineering culture that empowers individuals to deliver complex systems at scale. This role is critical to ensuring our ML platforms remain robust, scalable, and aligned with evolving product and business needs.
Key facts
What you'll do
Own the technical strategy and roadmap for ML training and serving, covering large-scale model training, model deployment workflows, GPU infrastructure, and low-latency model serving at scale.
Lead and grow a team of platform engineers while remaining deeply engaged in technical strategy and architectural trade-offs in close partnership with senior ICs.
Continuously evolve ML training and serving infrastructure to stay ahead of the frontier - anticipating where AI and ML are heading and building the infrastructure that makes those capabilities possible at Affirm before they become urgent needs.
Partner with ML modeling, product, and infrastructure leadership to ensure ML training and serving infrastructure accelerates Affirm's most critical ML initiatives.
Establish engineering excellence across the organization: reliability, observability, developer experience, and operational rigor.
Recruit, develop, and retain world-class platform engineers.
Evaluate and adopt emerging technologies in machine learning, including large-scale training and serving of transformer-based models, advanced GPU compute strategies, and reinforcement learning.
Define and enforce platform standards that enable scalable, secure, and maintainable ML systems across the organization.
Collaborate closely with data engineering and SRE teams to ensure robust data pipelines, monitoring, and incident response for ML platforms.
Champion best practices in security, compliance, and cost optimization for ML workloads in production.
Drive data-driven decision making by establishing clear metrics for platform performance, developer productivity, and model serving quality.
Act as a thought leader both within Affirm and externally, contributing to the broader ML infrastructure community through documentation and knowledge sharing.
Mentor senior ICs and engineers, providing coaching on technical design, career growth, and effective execution in complex environments.
Balance long term strategic bets with short term execution, ensuring the platform delivers immediate value while building a durable foundation for future innovation.
Requirements
10+ years of industry experience in software and or machine learning engineering, with significant hands on software engineering experience, including 4+ years managing engineering teams.
Deep expertise in building and operating large scale ML infrastructure, with substantial experience in one or more of model training systems, model serving, deployment workflows, or GPU infrastructure.
Strong understanding of ML data needs - training datasets, data quality, reproducibility, evaluation data, and how data shapes model behavior and platform design.
Fluency with modern ML: deep neural networks, transformer architectures, reinforcement learning, large scale GPU training and serving.
Strong systems thinking - comfortable reasoning from low level infrastructure decisions to broad architectural trade offs.
Track record of building platforms that meaningfully accelerate the productivity and impact of ML teams.
Experience on the applied ML modeling side is a plus - understanding how models are built makes you a better platform builder.
Proven track record of recruiting, developing, and retaining high performing engineers, including coaching senior ICs and creating opportunities for them to expand their technical leadership.
Experience navigating ambiguity and leading through organizational complexity.
Bachelor's degree in a technical field or equivalent practical experience.
Practical notes
Hours, travel, visa, or deadlines not mentioned in SOURCE. Omit the section entirely.