Machine Learning Engineer II
Job description
About the role
May Mobility is actively building the next generation of urban mobility through autonomous systems. Headquartered in Ann Arbor, Michigan, the company develops and deploys self-driving vehicles designed to improve safety, support sustainability, and enhance public transit. The role focuses on scaling the core decision-making technology that directs these vehicles. This position is based remotely within the United States and offers full-time engagement. The expected compensation range is $132,500 to $179,999 on an annual basis. You will own the design, implementation, and operational excellence of machine learning models that power autonomous vehicle decision-making. This role requires you to bridge the gap between cutting-edge research and reliable, city-scale deployment. You will be responsible for ensuring that data pipelines, training workflows, and inference systems meet the stringent demands of real-world driving. Your work will directly influence how autonomous vehicles perceive, reason, and act in complex urban environments.
Key facts
What you'll do
- Architect and maintain data and model processing pipelines that operate reliably at large scale across cloud and on-premise cluster environments.
- Design and implement containerized training workflows that manage GPU resource allocation, experiment tracking, and resilient data loading for high-throughput workloads.
- Build and curate data repositories and metadata systems to ensure that training and evaluation processes remain traceable, auditable, and reliable over time.
- Develop and refine continuous integration and delivery pipelines that enable safe, rapid updates to driving software while enforcing rigorous validation and testing gates.
- Collaborate closely with perception and planning teams to align machine learning objectives with practical driving constraints, safety requirements, and real-world operational conditions.
- Analyze experiment results and model training behaviors to identify performance degradation, data quality anomalies, and opportunities for systematic refinement of algorithms and data flows.
- Define standardized data formats, storage strategies, and query patterns that facilitate efficient access for large-scale model training, benchmarking, and offline analysis.
- Guide the movement of models from early research prototypes into production-grade services that can operate reliably across a global fleet of autonomous vehicles in diverse conditions.
- Partner with infrastructure and operations teams to monitor deployed models, diagnose issues in live systems, and implement improvements that enhance robustness and scalability.
- Contribute to the establishment of best practices, documentation, and knowledge-sharing initiatives that elevate the maturity of the machine learning organization.
- Support the execution of end-to-end experiments that validate new ideas, measure impact, and inform data-driven decisions for product and policy development.
- Ensure that all machine learning activities adhere to the company's standards for security, reliability, and performance in safety-critical applications.
- Work closely with cross-functional stakeholders to translate high-level product goals into concrete technical requirements and measurable success criteria.
- Continuously explore emerging methods and tools to improve the efficiency, accuracy, and scalability of the autonomous driving software stack.
Requirements
- Hold a Bachelor's or Master's degree in a technical discipline with a focus on mathematics, engineering, or computer science principles.
- Possess a minimum of two years of experience developing machine learning infrastructure, platforms, or distributed systems that operate in production environments.
- Demonstrate proficiency in writing code using C++, Python, and PyTorch, along with strong comfort working in Linux-based development and deployment environments.
- Show a solid understanding of fundamental machine learning concepts such as training loops, common operators, loss functions, and standard architectures.
- Exhibit the ability to write clean, maintainable, and well-documented code that can be reviewed and maintained by other engineers.
- Display strong problem-solving skills when diagnosing issues in data pipelines, model training jobs, and inference services.
- Communicate effectively with both technical and non-technical stakeholders to align on priorities, trade-offs, and implementation details.
- Have experience with version control systems, particularly Git, and understand software engineering best practices such as testing and code review.
- Be comfortable working with complex, multi-team environments where dependencies and priorities must be managed proactively.
Nice to have
- Familiarity with the concepts and methods used in autonomous driving perception and planning systems.
- Experience using programming languages such as Go or Rust in addition to primary application code.
- Working knowledge of ML orchestration and experiment management tools, including Ray, Kubeflow, Airflow, MLflow, or similar platforms.
- Experience with distributed training techniques, such as PyTorch Distributed Data Parallel or Fully Sharded Data Parallel, as well as frameworks like DeepSpeed.
- Experience with data pipeline technologies and storage systems, including Apache Spark, Parquet file formats, object storage, and feature or metadata stores.
About the company
May Mobility is transforming cities through autonomous technology to create a safer, greener, more accessible world. Based in Ann Arbor, Michigan, May develops and deploys autonomous vehicles (AVs) powered by our innovative Multi-Policy Decision Making (MPDM) technology that literally reimagines the way AVs think.