Software Engineer - ML Platform
Job description
Software Engineer - ML Platform at Avride.
About the role
In this position, you will play a key role in developing the ML infrastructure that supports extensive machine learning training and data processing for autonomous driving. You will collaborate closely with both the Cloud Platform and ML engineering teams to create a user-friendly ML platform that emphasizes scalable orchestration, distributed computing, and effective tools for managing the entire model lifecycle.
Key facts
What you'll do
- Develop and enhance our machine learning compute platform utilizing Kubernetes, with Argo Workflows for orchestrating training, evaluation, and data processing tasks.
- Create and implement essential platform features, such as a Ray-based internal SDK for distributed execution and multi-tenant resource management, which includes scheduling, priorities, quotas, and policy enforcement across GPU, CPU, memory, and I/O.
- Optimize overall training throughput and platform performance by refining data access patterns, improving caching strategies, and eliminating bottlenecks in storage, networking, and resource contention.
- Collaborate with ML teams to troubleshoot complex workload challenges, conduct root-cause analyses, and transform recurring issues into platform-wide solutions.
- Assess, integrate, and enhance open-source tools (including Argo Workflows, Ray, and components of the Kubernetes ecosystem) to adapt to the platform's evolving requirements.
Requirements
- Proficient in Python or Go; familiarity with C++ is advantageous.
- Proven experience in designing and building scalable and maintainable systems and services.
- Background in managing production services from start to finish, including APIs, reliability practices, and observability.
- In-depth understanding of Kubernetes, particularly regarding scheduling, resource management, controllers, and pod lifecycle under stress.
- Strong Linux and systems debugging abilities, including performance analysis and networking/storage/I/O issues.
- Competence in diagnosing complex production challenges using logs, metrics, and traces, and effectively resolving them.
Nice to have
- Familiarity with Argo Workflows, Ray, MLflow, or similar distributed ML tools.
- Practical experience in constructing or managing large-scale ML training systems, including GPU scheduling, distributed training, and training data pipelines.
- History of optimizing resource utilization and performance in distributed settings.
Practical notes
Candidates must be authorized to work in the U.S. The company does not provide relocation sponsorship, and remote work is not an option.
Avride is an equal opportunity employer and is dedicated to offering reasonable accommodations to qualified applicants and employees with disabilities to ensure equal access to employment opportunities. Avride adheres to the Americans with Disabilities Act (ADA). If you require a reasonable accommodation to assist with the application or hiring process, or to perform essential job functions, please contact jobs@avride.ai.