Principal Machine Learning Engineer
Job description
About the role
Join Grab's AI Platform team as a Principal Machine Learning Engineer, serving as a key technical leader and solution architect. This role focuses on enhancing Grab's core ML and AI infrastructure, supporting hundreds of data scientists and ML engineers across the company. You will drive the platform's strategic direction and ensure its effective adoption by internal teams. You will act as the primary technical architect for AI Platform users, guiding data scientists and ML engineers in designing comprehensive solutions. This position requires you to improve large-scale training capabilities on AIP, focusing on throughput, reliability, cost efficiency, and developer experience for various workloads including foundation models and LLM fine-tuning. You will accelerate the entire model development cycle, from data processing to deployment and monitoring, to reduce iteration times and boost model quality for critical use cases. Success in this role means integrating various AIP components like model serving, ML pipelines, and data serving into a unified platform experience.
Key facts
What you'll do
- Assume ownership of the end-to-end ML infrastructure stack, diagnosing bottlenecks and architecting resilient solutions for production-scale challenges.
- Champion the enhancement of large-scale training frameworks on AIP to optimize throughput, reliability, cost efficiency, and developer experience for diverse workloads such as foundation models and LLM fine-tuning.
- Drive initiatives to shorten the model development lifecycle, encompassing data processing, model training, deployment, and monitoring, to accelerate iteration speed and improve outcomes for high-impact use cases.
- Lead the integration of disparate AIP components, including model serving, ML pipelines, and data serving, into a cohesive and intuitive platform experience.
- Synthesize user feedback and translate complex challenges into clear actionable requirements for AIP teams, proposing cross-platform improvements that benefit data scientists and ML engineers.
- Create and maintain reference architectures and best practices that empower teams to adopt AIP independently for future ML projects, ensuring scalability and consistency.
- Monitor the evolving landscape of ML infrastructure and LLM technologies, collaborating with the Head of Engineering to influence Grab's technology roadmap and mentor senior engineers on emerging trends.
- Act as a strategic partner to product and engineering teams, aligning platform capabilities with business objectives and enabling innovation through robust technical leadership.
- Spearhead the development of tools for observability and debugging, ensuring platform reliability and rapid resolution of issues for critical ML workflows.
- Foster a culture of continuous improvement by identifying inefficiencies in existing workflows and implementing automations that enhance productivity across the AI Platform.
- Evaluate and pilot new infrastructure paradigms, ensuring that Grab remains at the forefront of practical ML deployment and training methodologies.
- Provide technical leadership during design reviews and post-mortems, ensuring that decisions are well-informed and aligned with long-term platform goals.
- Build strong partnerships with product teams to understand their unique requirements and tailor AIP features that deliver tangible value.
- Document architectural decisions and processes meticulously to ensure knowledge transfer and maintainability of the platform.
Requirements
- Demonstrate expert command of ML lifecycle platforms such as Kubeflow, MLflow, Triton, and TorchServe, alongside distributed training frameworks including PyTorch, Ray, and Horovod.
- Bring a minimum of 8 years of hands-on experience with distributed systems, specifically involving Kubernetes, containerization, and high-performance computing clusters utilizing GPUs and TPUs, with a focus on optimizing large-scale data and model pipelines.
- Show a proven ability in designing highly available, scalable, and secure multi-tenant platform services that meet the demands of a global user base.
- Highlight practical experience with LLM orchestration, fine-tuning infrastructure, or serving optimization techniques, including work with tools like vLLM or TensorRT-LLM.
- Exhibit a continuous learning approach to consistently evaluate and implement state-of-the-art infrastructure paradigms that solve real-world engineering problems.
- Display the capacity to work independently in ambiguous situations, taking full ownership of complex system integrations and platform consolidation efforts.
- Show a strong commitment to elevating engineering standards through active mentorship and fostering a collaborative technical culture within the AI Platform team.
- Possess excellent communication skills to bridge the gap between technical implementation and business requirements, ensuring clarity and alignment with stakeholders.
Nice to have
Preferred items are not specified in the SOURCE documentation for this role.
Practical notes
Grab provides comprehensive benefits including Term Life and Medical Insurance, a flexible GrabFlex package, parental and birthday leave, volunteering leave, and a confidential Grabber Assistance Programme. FlexWork arrangements, such as differentiated hours, support work-life balance. Grab is an equal opportunity employer committed to diversity and inclusion.
Deadline information and specific working hours are not provided in the SOURCE material. Travel and visa requirements are not detailed in the SOURCE documentation.