Fleet Reliability Engineer
Job description
Fleet Reliability Engineer at Dyna Robotics.
About the role
This position is defined by hands-on ownership of the entire lifecycle of our humanoid robot fleet in the field. You will be responsible for root-causing individual hardware failures at the bench level and simultaneously building the fleet-wide telemetry and analysis tooling that catches degradation before it escalates into downtime. The role requires equal parts physical repair skill and data-driven investigation, where you will pull robots apart, solder fixes, and then design the processes that prevent the same class of failure from recurring. You will operate at the critical intersection of hardware, firmware, controls, and data infrastructure, becoming the person the team calls when a robot behaves in an inexplicable manner. This is a genuinely hands-on role where writing degradation analysis is paired with the ability to re-commission a repaired electromechanical assembly. You will translate physical evidence directly into actionable design and process changes that improve long-term fleet health.
Key facts
What you'll do
- Root-cause complex failures across the fleet, isolating actuator dropouts, encoder and calibration drift, thermal events, CAN bus faults, and control-stack bugs to separate genuine hardware degradation from firmware, calibration, and model-side causes.
- Work hands-on with the hardware, bench-testing suspect actuators and electronics, performing rework and repair including soldering, harness and connector work, swapping and re-commissioning components, and building boot and power-cycle rigs for physical validation.
- Design and build mechanical solutions such as test fixtures, jigs, brackets, mounts, and small modifications to fielded robots, taking concepts through fabrication whether via 3D printing, shop resources, or vendor partnerships.
- Turn mechanical evidence into design changes, reasoning from data regarding backlash, windup, gear wear, preload loss, and hard-stop behavior to feed concrete feedback to the mechanical team who owns major structural design and simulation.
- Build and roll out fleet health metrics, including maintenance indices, hold-current precursors, backlash and deadband tracking, and bus-error monitoring, with alert rules carefully tuned to fire on real faults rather than noise floors.
- Develop calibration and test tooling, creating calibration tools with embedded safety gates, conducting characterization sweeps for new hardware revisions, and building pre-merge validation harnesses for firmware and software changes.
- Wrangle robot telemetry at scale, extracting and decoding logs in formats such as MCAP, npz, and episode archives from bandwidth-constrained robots, closing gaps in the telemetry pipeline, and building monitoring that should have existed from the start.
- Write clear documentation and hand them off, producing runbooks, standard operating procedures, incident reports, and handoff packages clean enough that a teammate can pick the work up cold and continue without interruption.
- Review pull requests and propose fixes across firmware defaults, control code, and deployment configuration, catching interaction bugs between calibration tools, boot checks, and safety gates before they reach production.
- Collaborate closely with mechanical, firmware, and controls engineers to ensure that physical repairs, data insights, and design changes are aligned with overall system performance and reliability goals.
Requirements
- 3 to 5 years in robotics, mechatronics, hardware test, or reliability engineering, with significant time spent on fielded hardware rather than lab prototypes.
- Genuinely hands-on instincts, comfort at the bench with a soldering iron, multimeter, and calipers, and the ability to tear down, repair, and re-commission electromechanical assemblies without depending on another team.
- An understanding of how actuators fail mechanically, including backlash, wear, preload loss, and end-stop damage, and the ability to connect physical evidence to design causes without performing large structural design or detailed FEA.
- Working knowledge of robot control including impedance and torque-controlled actuators, force-control parameters, and controller state lifecycles such as boot, e-stop, and reconnect, enough to tune deployed behavior and distinguish hardware faults from controls bugs.
- Strong Python skills for data analysis and tooling, with solid time-series and signal-statistics instincts around noise floors, drift, persistence rules, and the ability to recognize when an alert is driven by a tiny denominator.
- Fluency with actuators and motor control concepts, including quasi-direct-drive systems, encoders, torque and current telemetry, CAN bus, and calibration and zeroing procedures essential for fielded systems.
- Experience with observability stacks such as Datadog or similar platforms, and proficiency with version control using git and pull request workflows to maintain traceable changes.
- A methodical approach to documentation and handoff, ensuring that runbooks, incident reports, and operational procedures are complete, clear, and actionable for follow-on engineers.
Nice to have
Only items explicitly indicated as preferred in the source are included in this section; no additional preferences have been added.
Practical notes
The role is full-time based in Redwood City, California. Relocation support and visa sponsorship may be available for qualified candidates, and the position requires the ability to work in a fast-paced environment where priorities can shift rapidly based on fleet performance and customer needs.