
Software Engineer, ML Dev Enablement
Job description
About the role
This position focuses on building the tools and infrastructure that empower our machine learning developers. You will play a crucial role in enhancing the efficiency and effectiveness of our ML workflows. Your contributions will directly impact the speed at which we can develop and deploy advanced autonomous driving systems. You will be responsible for designing the foundational components that allow data scientists and engineers to iterate rapidly. A significant part of your work will involve abstracting complex underlying infrastructure into intuitive interfaces. You will act as a bridge between cutting edge research and production grade implementation. Success in this role will be measured by the adoption and reliability of the tools you create.
Key facts
What you'll do
- Architect and maintain the core software systems that underpin the entire machine learning development lifecycle from experimentation to deployment.
- Design and implement scalable tools for data versioning, dataset management, and feature store integration to ensure consistency across training and inference.
- Partner closely with machine learning engineers and researchers to gather requirements, diagnose workflow bottlenecks, and deliver tailored technical solutions that unblock their work.
- Lead the development and optimization of end to end automation for machine learning pipelines, focusing on orchestration, scheduling, and resource management.
- Build and iterate on model evaluation frameworks that enable systematic comparison of model performance, data drift, and behavioral metrics against established benchmarks.
- Implement robust infrastructure for experiment tracking, artifact storage, and metadata management to provide full visibility into the model development journey.
- Create comprehensive developer documentation and internal SDKs that simplify the adoption of complex infrastructure components by non expert users.
- Collaborate with infrastructure and platform teams to ensure that the tools you build are aligned with the broader cloud strategy and security policies.
- Develop monitoring and debugging toolkits specifically for machine learning workloads to help data scientists diagnose issues in model behavior and data quality.
- Contribute to open source projects and internal standards that establish best practices for machine learning engineering and deployment hygiene.
Requirements
- Hold a Bachelor's degree in Computer Science, Engineering, or a related technical field that provides a strong foundation in algorithms and system design.
- Possess 2 years of professional software development experience building and shipping software products in a production environment.
- Demonstrate hands on experience with Python as a primary language and a deep understanding of common ML frameworks such as TensorFlow or PyTorch.
- Show familiarity with major cloud platforms like AWS, GCP, or Azure, including core services for compute, storage, and networking.
- Exhibit strong problem solving skills and the ability to translate ambiguous requirements into concrete technical specifications.
- Bring a proven track record of writing clean, maintainable, and well tested code that adheres to industry standard best practices.
- Display effective communication skills necessary to collaborate with diverse stakeholders including researchers, product managers, and operations teams.
- Commit to working within the Motional values of safety, integrity, and collaboration to ensure successful project outcomes.
Nice to have
- Hold a Master's degree or PhD in a relevant technical field with a focus on machine learning, distributed systems, or computer science.
- Bring direct experience with containerization technologies such as Docker and orchestration platforms like Kubernetes for deploying scalable services.
- Demonstrate knowledge of CI/CD principles and tooling, including the implementation of automated testing, linting, and deployment pipelines for machine learning projects.
- Have a background in distributed computing and parallel processing concepts that apply to training large scale models.
- Show familiarity with infrastructure as code tools that allow for the programmatic provisioning and management of cloud resources.
Practical notes
Motional is an equal opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all employees.