Engineering Manager, Model Inference
Job description
About the role
You will lead the end-to-end technical strategy and execution for the Model Inference team, owning the systems that power real-time AI interactions in clinical workflows. You will define and drive the technical vision for how large language models are served, optimized, and monitored in production environments. You will partner deeply with ML Research and the GenAI Platform to translate research breakthroughs into reliable, scalable inference services. You will architect low-latency, high-throughput infrastructure that meets the stringent demands of healthcare data and clinical users. You will lead incident response and reliability efforts, ensuring that every clinician interaction is supported by performant and auditable systems. You will build and mentor a high-performing team of inference engineers, fostering an environment of ownership, rigor, and continuous improvement. You will establish engineering standards and operational practices that scale as the organization grows. You will work cross-functionally with Product, Data, and Clinical teams to align inference capabilities with evolving user needs and regulatory expectations.
Key facts
What you'll do
- Lead and grow a high-performing team of AI inference engineers focused on building and scaling infrastructure for Abridge's products and APIs.
- Own the technical direction of our inference systems - making key decisions around batching, throughput, latency, and GPU utilization.
- Architect and scale inference infrastructure for reliability, efficiency, and observability; lead incident response.
- Benchmark and eliminate bottlenecks throughout the inference stack.
- Partner with ML Research teams on model optimization, quantization, and deployment.
- Develop APIs for AI inference used by both internal teams and external customers.
- Recruit, mentor, and develop engineering talent; establish team processes, engineering standards, and operational excellence.
- Work closely with the GenAI Platform, Data, and Product teams to plan and execute projects that directly impact clinicians and patients.
- Evaluate and integrate emerging inference techniques to maintain performance and efficiency at scale.
- Define and track reliability, latency, and throughput metrics to guide infrastructure investments.
- Collaborate with security and compliance to ensure inference infrastructure meets enterprise standards.
- Drive adoption of best practices in distributed systems and performance engineering across the team.
- Conduct code reviews and technical design sessions to ensure high-quality, maintainable systems.
- Prioritize roadmap items based on impact to clinical workflows, customer needs, and technical debt reduction.
- Foster a culture of experimentation, data-driven decision making, and continuous learning.
Requirements
- 5+ years of engineering experience with 1+ years in a technical leadership or management role.
- Deep, hands-on experience with ML systems and inference frameworks (e.g., PyTorch, TensorRT, vLLM, TensorFlow).
- Strong understanding of LLM architecture (eg. Multi-Head Attention, Multi/Grouped-Query Attention, and common transformer components).
- Experience with inference optimizations (eg. batching, quantization, kernel fusion, FlashAttention).
- Familiarity with GPU characteristics, roofline models, and performance analysis.
- Experience deploying reliable, distributed, real-time systems at scale.
- Experience with parallelism strategies: tensor parallelism, pipeline parallelism, expert parallelism.
- Skilled at hiring and mentorship, with a demonstrated track record of helping engineers grow their skills and careers.
- Strong technical communication and cross-functional collaboration skills.
- Comfortable giving constructive feedback on technical designs and code reviews.
- Has thrived in a fast-growing startup and knows how to operate with urgency and focus.
- Proven ability to work in a fast-paced, ambiguous environment while maintaining high standards.
- Commitment to writing clean, maintainable, and well-documented code.
- Willingness to dive into details while keeping an eye on long-term system architecture.
Nice to have
- Background in training infrastructure and RL workloads.
- Skilled in building secure, compliant systems on major cloud platforms (GCP preferred, AWS experience welcome).
- Experience with Kubernetes and container orchestration at scale.
- Published work or contributions to inference optimization research.
Practical notes
- Full-time position.
- Office located in San Francisco.
- Work is expected in the Mission District office.