
Tech Lead Manager- MLRE, ML Systems
Job description
About the role
You will architect and own the end-to-end design of the LLM post-training and evaluation platform at Scale, balancing scalability, reliability, and performance for production workloads. You will partner with ML researchers and data scientists to translate their requirements into robust platform capabilities that accelerate experimentation and model development cycles. You will drive the technical vision for distributed training and inference frameworks, ensuring they meet evolving product and research needs. You will lead cross-functional collaboration to integrate state-of-the-art optimization techniques and algorithms into the core platform. You will own critical design decisions that shape how data curation, fine-tuning, and evaluation pipelines operate at scale. You will mentor engineers and influence coding standards to maintain high-quality, maintainable systems. You will act as a key technical liaison with product and operations to align platform roadmaps with business objectives.
Key facts
What you'll do
Architect and deliver the core infrastructure for large language model training and inference on a distributed scale-out platform.
Collaborate with machine learning and research teams to design workflows that accelerate development and data curation for next-generation LLMs.
Profile and optimize platform performance, throughput, and resource utilization across training and inference workloads.
Research, evaluate, and integrate state-of-the-art systems technologies to improve scalability, efficiency, and reliability of the ML stack.
Enable and support post-training methods such as RLHF and related algorithms including PPO and GRPO within the platform.
Build and maintain strong interfaces with CUDA, PyTorch, and transformer libraries to leverage cutting-edge deep learning primitives.
Partner with data science and product teams to ensure the platform meets evolving requirements for model development and evaluation.
Lead technical design reviews and contribute to standards that ensure platform robustness, observability, and maintainability.
Mentor engineers and foster a culture of high-quality software engineering and best practices across the team.
Drive initiatives to improve developer experience, automation, and reliability of the end-to-end ML lifecycle.
Contribute to the data quality evaluation pipeline by enhancing the underlying training framework that powers it.
Engage with internal and external partners to align platform capabilities with enterprise and research needs.
Champion best practices in system optimization, debugging, and performance tuning for large-scale ML workloads.
Represent Scale's technical leadership in discussions with partners and stakeholders regarding platform direction and adoption.
Requirements
Demonstrate a strong passion for system optimization and large-scale distributed computing challenges.
Have hands-on experience with multi-node LLM training and inference in production or research environments.
Possess deep experience developing and operating large-scale distributed ML systems, including orchestration and scheduling.
Show proven ability to work with post-training methods such as RLHF and RLVR, including algorithms like PPO and GRPO.
Exhibit strong software engineering proficiency with frameworks and tools such as CUDA, PyTorch, and transformers, including flash attention.
Maintain excellent written and verbal communication skills to operate effectively in a cross-functional team environment.
Have a track record of delivering high-performance, reliable, and maintainable systems under tight deadlines.
Demonstrate comfort with low-level debugging, performance profiling, and capacity planning for complex ML workloads.
Nice to have
Demonstrated expertise in post-training methods and next-generation use cases for large language models, including instruction tuning, RLHF, tool use, reasoning, agents, and multimodal applications.
Practical notes
This role is eligible for compensation details as outlined in the job description, including base salary, equity, and benefits. Employment is full-time. Candidates must be authorized to work in the country where the role is based. The company maintains a 90-day waiting period before reconsiderating candidates for the same role.
About Scale AI
Scale AI is hiring for Tech Lead Manager- MLRE, ML Systems. The listing location is San Francisco, CA; New York, NY.
This Tech Lead Manager- MLRE, ML Systems opening is posted for San Francisco, CA; New York, NY.