Fleet Reliability Engineer
Job description
About the role
Quartermaster is building a live, real-time picture of the world's oceans. Our distributed sensor network turns ordinary civil and commercial vessels into intelligent nodes, delivering HD video, signal intelligence, and AI-powered insights that help coast guards, navies, insurers, energy operators, and researchers monitor, detect, and respond across global waters. At the center of the network is SmartMast™ - a ruggedized, vessel-mounted sensor package that installs on any vessel over 20 GT. Each unit combines a 360°/31x optical and infrared camera array (IP68-rated), dual software-defined radios, an onboard AI compute module, a configurable powerpack, and resilient global SATCOM that pushes data to users anywhere in under five seconds. We have deployed 600+ sensors across more than 25 countries, and our fleet has traveled over 10 million nautical miles. As the fleet scales, keeping every deployed SmartMast healthy, connected, and producing reliable data at sea is mission-critical. We are hiring our first dedicated Fleet Reliability Engineer to own that outcome. The hire will define and drive a first-of-its-kind reliability discipline for a rapidly scaling hardware fleet at sea. They will translate complex field performance data into decisive engineering actions that keep sensors operational in the world's most demanding maritime environments. This role demands deep accountability for uptime, availability, and long-term durability of every unit deployed on the ocean. The successful candidate will build the foundational processes, metrics, and tooling that allow the fleet to scale from hundreds to thousands of sensors without sacrificing robustness.
Key facts
FLEET HEALTH & MONITORING
- Own fleet-wide reliability metrics - uptime, availability, MTBF, MTTR, data-yield, and failure rates by component and by deployment environment - and report them to engineering and leadership.
- Build and refine dashboards, alerting, and telemetry pipelines that surface degrading units (power, thermal, connectivity, camera, radio, compute) before they go offline.
- Define what "healthy" means for each subsystem and set the thresholds that trigger proactive intervention.
- Translate raw telemetry into clear operational status for stakeholders, maintaining a living view of the global fleet at any moment.
- Correlate environmental conditions and operational patterns with unit performance to identify systemic risks.
FAILURE ANALYSIS & CONTINUOUS IMPROVEMENT
- Lead root-cause analysis (RCA) on field failures, from telemetry forensics through physical teardown of returned units.
- Maintain the fleet failure database and drive FMEA, reliability growth tracking, and corrective/preventive action (CAPA) to closure.
- Close the loop with hardware, firmware, and manufacturing teams - translating field failures into design-for-reliability, component-selection, and firmware changes.
- Own structured problem-solving using formal methods (8D, 5-Whys, fishbone) to ensure durable fixes.
- Track reliability trends across units, batches, and geographies to prioritize high-impact improvements.
FIELD SERVICE, MAINTENANCE & LOGISTICS
- Define preventive maintenance schedules, spares strategy, and the RMA / repair-and-return process for a globally distributed fleet.
- Update installation, diagnostic, and field-repair procedures and troubleshooting guides used by internal technicians and partner crews.
- Support field deployments and complex repairs directly, including periodic travel to vessels, ports, and installation sites.
- Coordinate with logistics partners to minimize downtime and optimize the flow of spares and repaired units.
- Drive standard work and safety protocols for all field activities to ensure consistent and reliable execution.
RELIABILITY ENGINEERING & SCALE
- Work with the hardware team to establish environmental and life-test protocols (vibration, salt-fog/corrosion, thermal, ingress, power) to qualify hardware and predict field life before deployment.
- Feed reliability requirements and acceptance criteria into new hardware revisions and supplier qualification.
- Design the reliability processes and tooling so they scale as the fleet grows into the thousands of units.
- Instrument the reliability function itself to measure effectiveness, identify bottlenecks, and drive automation.
- Partner with data science and product teams to integrate reliability signals into roadmap decisions and long-term system architecture.
Requirements
- Bachelor's degree in Electrical, Mechanical, Systems, Reliability, or a related engineering discipline - or equivalent hands-on experience.
- 5+ years of engineering experience with deployed electro-mechanical hardware, at least 2 of which are in reliability, sustaining/field engineering, or hardware operations for a fielded product.
- Demonstrated ownership of hardware reliability outcomes for a fleet or installed base - you have been directly responsible for uptime, failure rates, or MTBF/MTTR of real hardware in the field.
- Hands-on proficiency with root-cause analysis methods (8D, 5-Whys, fishbone) and reliability tools such as FMEA, fault-tree analysis, and CAPA.
- Practical experience diagnosing electro-mechanical systems using telemetry/logs, bench instruments, and physical teardown.
- Data fluency: able to query, analyze, and visualize fleet telemetry using SQL and Python (or equivalent) to find trends and drive decisions.
- Working knowledge of electronics, power systems, and embedded computing relevant to sensor platforms and edge compute hardware.
- Comfort operating in maritime and outdoor environments, with an understanding of the constraints of shipboard power, connectivity, and maintenance.
Nice to have
- Experience with global cellular and satellite communications architectures and the operational constraints of maritime networks.
- Familiarity with camera systems, image processing pipelines, and the challenges of edge AI inference in harsh conditions.
- Background in supply-chain, warranty, or RMA management for hardware products.
- Any experience with ruggedized, power-constrained, or mission-critical field deployments.
Practical notes
The role is full-time and based in Arlington, VA. Some travel is required to support field deployments, vessel inspections, and partner facilities. Candidates must be eligible to obtain a U.S. government security clearance as part of the onboarding process. This role reports to the Reliability and Hardware leadership and is expected to influence product roadmaps for multiple years as the fleet scales.