Software Engineer, II
Job description
About the role
Opportunity Overview
Torc Robotics presents a specialized position for an Autonomy Data Engineer at the II level. This role is situated within a fully remote framework, with the work location designated in Blacksburg, Virginia, United States. The engagement is structured as a full-time commitment. The compensation package includes a base salary of $137,000.00, with a total estimated value of $170,000.00. Candidates are encouraged to review the official application portal for the most current and detailed information regarding the position.
Company Background
Torc Robotics operates as a distinct entity focused exclusively on the development of software for automated trucks. The organization has maintained a leadership position in autonomous driving technology since 2007. For over a decade, the company has dedicated its efforts to commercializing advanced autonomy solutions in collaboration with experienced industry partners. Torc was the first autonomous vehicle software company to establish a direct partnership with a truck manufacturer, a milestone achieved through its work with Daimler. This historical focus provides the foundation for its current mission of transforming global freight movement through software. The organization is now a part of the Daimler family, reinforcing its commitment to this specialized domain.
Core Responsibilities
The successful candidate will own the entire data pipeline, from initial vehicle intake to the delivery of actionable insights. This involves defining robust data schemas and constructing curation tools that translate complex sensor information into structured, reliable datasets. These datasets directly inform the work of perception and planning engineers, shaping the autonomous truck's understanding of motion and environment.
A primary duty is to design scalable ingestion pipelines capable of handling high-bandwidth sensor logs collected from vehicles operating in demanding real-world conditions. This requires building systems that are resilient to intermittent and unreliable connectivity, ensuring data integrity from the vehicle to cloud storage. The engineer will implement rigorous data validation processes to identify corrupted information, missing sensors, or calibration inconsistencies before such data influences downstream model training.
The role also encompasses the development of dataset curation and labeling infrastructure. This includes building tools to query raw logs for the creation of curated training and evaluation datasets. The engineer will create automation for cost-effective pseudo-labeling workflows that can scale alongside data ingestion volumes. also, establishing data quality and model performance metrics is essential to direct labeling efforts toward the highest-value examples.
Autonomy data visualization represents another critical area. The engineer will deploy and maintain tools that support log review, annotation quality assurance, and debugging workflows. This involves creating integrations that allow teams to navigate directly from a model failure or dataset entry to the corresponding origin log data. Building dashboards that provide visibility into data coverage across different terrain types, operating environments, and geographic regions is also a key function.
Cross-functional collaboration is integral to this position. The engineer will establish and document data contracts between data services and model training consumers. This partnership extends throughout the data lifecycle, from helping to shape logging schemas and collection triggers to defining the dataset interfaces used for model training and evaluation. The role requires a commitment to evolving data engineering standards and best practices in conjunction with the wider technical teams.
Qualifications and Preferred Experience
Applicants must possess a minimum of three years of experience in building data platforms or pipelines within complex system environments. A strong understanding of how to define and enforce data contracts is necessary to ensure that intake pipelines deliver reliable and trusted outputs. The ability to design storage layouts that effectively balance query performance against cost at a large scale is a required skill.
Candidates must demonstrate the capability to write code that remains reliable even in the face of connectivity disruptions common to moving vehicles. The ability to transform unstructured operational scenes into quantifiable metrics that guide perception teams is essential.
Experience with autonomous trucks, simulation tools, or visualization systems is considered a valuable asset for this role. Proficiency with specific tools and technologies is expected, including Python, SQL, cloud storage solutions, sensor log formats, and data versioning methodologies.
What you'll do
- Meet the bar Practical notes