ML Infrastructure Engineer, Training
Job description
ML Infrastructure Engineer, Training at Dyna Robotics.
About the role
This position is responsible for architecting and operating end-to-end research infrastructure that accelerates multimodal model experimentation and enables rapid iteration. The role focuses on optimizing production inference paths to achieve low-latency robot control across distributed cloud environments. You will deeply profile multi-cloud GPU hardware to maximize utilization and reliability for demanding workloads. The work requires applying distributed systems fundamentals to eliminate race conditions, manage memory, and optimize NCCL and inter-node communication. You will design, build, and operate infrastructure end-to-end to remove bottlenecks for fast-moving research teams. This position emphasizes the reliable operation of systems that unblock high-velocity research initiatives. You will work closely with teams building data-driven decisions, models, and pipelines to ensure infrastructure supports analytical and experimental goals. The role sits at the intersection of data engineering and ML systems, requiring a blend of statistics, coding, and communication to support business needs.
Key facts
What you'll do
Architect and operate end-to-end research infrastructure to accelerate multimodal model experimentation and streamline development workflows.
Optimize production inference paths for low-latency robot control to ensure precise and reliable execution in real-time environments.
Profile multi-cloud GPU fleets on GCP and AWS to drive higher utilization, performance, and system reliability.
Apply distributed systems fundamentals to eliminate race conditions, manage memory effectively, and optimize NCCL and inter-node communication.
Design, build, and operate infrastructure components to remove bottlenecks and support fast-moving research teams.
Tune infrastructure for high reliability and performance to support demanding robotic control and inference workloads.
Collaborate closely with data analysts, data scientists, and data engineers to ensure infrastructure aligns with analytical and model development needs.
Evaluate and implement optimizations for data pipelines, storage, and compute resources to improve efficiency and scalability.
Contribute to the development and maintenance of robust monitoring and operational practices for production-grade systems.
Engage with cutting-edge technologies and practices to continuously improve the capabilities and resilience of research infrastructure.
Requirements
The posting states a bachelor's degree requirement. 7+ years of engineering experience is required, with a track record of leading technical projects in high-performance computing or ML infrastructure.
Candidates must demonstrate deep experience with PyTorch and distributed training frameworks such as DeepSpeed and Accelerate, including nuances of mixed precision and gradient accumulation.
Hands-on experience managing cloud GPU environments on GCP or AWS and container orchestration with Kubernetes is required.
A fundamental understanding of distributed systems is required, including race conditions, memory management, and NCCL/inter-node communication.
Systems must be architected, built, and operated to unblock fast-moving research through reliable infrastructure and efficient workflows.
Experience with version control, code review, and collaborative engineering practices is expected as part of standard software development processes.
Strong problem-solving skills are required to diagnose complex issues and implement effective solutions in production environments.
Nice to have
Experience with robotics data formats such as MCAP or Protobuf, or multimodal models such as VLAs, is beneficial.
Deep ML systems experience with custom kernels in Triton, compilers, or runtime optimization is beneficial.
Experience as a founding or early-stage infrastructure hire is beneficial.
Practical notes
Typical interview steps
Data interviews commonly include a SQL or coding exercise, a statistics question, and a case study. Candidates may be asked to design a metric, interpret an experiment, or build a small model. Some companies give a take-home analysis. Expect questions about past projects and the business impact of your work. Interviewers often evaluate how you communicate uncertainty and business impact, not only the math. Bringing a clean write-up of a past analysis to the interview is well received.
Good to know
Questions to ask
Worth asking in any interview: how the team measures success, who the role works with daily, what the onboarding looks like, and what the company is trying to achieve this year. Asking what past hires did well is a strong final question. Keep the list short and pick the questions that matter most to you.
Career growth
Data careers grow toward senior analyst, staff data scientist, or data engineering lead. Many professionals specialize in machine learning, analytics, or infrastructure. Cross-functional work with product and engineering teams becomes more important at senior levels. The field changes quickly, so continuous learning is part of the job. Professionals who can translate numbers into decisions tend to advance fastest.