Senior ML Ops Engineer
Job description
About the role
The owns the design, implementation, and reliability of the machine learning infrastructure that powers autonomous rail operations. This role owns the end-to-end MLOps stack, encompassing data ingestion, experiment tracking, model training, deployment, and monitoring for safety-critical perception systems. The hire will architect scalable and resilient systems that bridge cutting edge research in autonomy with robust production workflows. They will collaborate daily with world class engineers in robotics, software, and perception to ensure platform capabilities keep pace with model innovation. This position requires ownership of infrastructure decisions that directly impact model performance, system latency, and operational reliability in real world environments. The role is responsible for driving standardization and automation across ML workflows to accelerate delivery and improve reproducibility.
Key facts
What you'll do
Architect and implement scalable MLOps pipelines that automate data management, model training, validation, deployment, and monitoring across research and production environments.
Design and operate cloud native infrastructure on platforms such as AWS and GCP, optimizing compute, storage, and networking for intensive ML workloads in both development and real world scenarios.
Collaborate closely with autonomy and perception engineers to define requirements, shape data strategies, and align model development with operational needs.
Build and maintain robust experiment tracking, versioning, and lineage systems to enable reproducibility and informed decision making across ML teams.
Implement deployment frameworks and CI CD workflows that support safe, auditable, and efficient model releases to edge and cloud environments.
Develop monitoring and alerting solutions for model performance, data drift, system health, and resource utilization to ensure reliable operation at scale.
Partner with software engineers to standardize tooling, improve developer experience, and create reusable components that accelerate ML workflows.
Evaluate and integrate emerging technologies and infrastructure patterns to enhance scalability, security, and efficiency of the ML platform.
Lead incident response and troubleshooting for ML infrastructure, coordinating with cross functional teams to resolve complex issues.
Champion best practices in security, compliance, and operational excellence to ensure infrastructure supports high integrity, safety critical applications.
Drive automation of repetitive tasks, reducing manual overhead and enabling engineers to focus on high impact model and system improvements.
Establish clear documentation and knowledge sharing practices to ensure continuity and enable cross team collaboration on infrastructure initiatives.
Contribute to hiring and mentorship by guiding junior engineers on MLOps principles, tooling, and operational responsibilities.
Requirements
Demonstrate extensive experience designing and operating ML infrastructure at scale, with a proven track record in production MLOps environments.
Showcase strong proficiency with distributed training, inference optimization, and deployment patterns for machine learning models in cloud and edge contexts.
Bring deep expertise in cloud platforms, focusing on AWS or GCP services relevant to compute, storage, networking, and security for ML workloads.
Exhibit mastery of containerization and orchestration technologies such as Docker and Kubernetes, along with infrastructure as code practices.
Possess strong scripting and programming skills in languages commonly used in data science and engineering, including Python and associated data science stacks.
Have a solid understanding of monitoring, logging, and observability techniques as they apply to ML systems and real time data pipelines.
Demonstrate experience with version control, CI CD pipelines, and collaborative software development practices essential for reliable infrastructure delivery.
Show commitment to security, compliance, and operational reliability, with experience implementing safeguards for sensitive data and safety critical systems.
Practical notes
The role can be hybrid, requiring a minimum of one week per month onsite in Los Angeles for a senior engineer with zero to one builds of perception systems. Work engagement details are to be confirmed from the source information. This position is integral to the development and deployment of autonomous battery electric rail vehicles, shifting freight from US trucking onto rail. Applicants should align with the mission of creating cleaner, safer, and more efficient logistics solutions.