Software Engineer, Data Infra
Job description
Software Engineer, Data Infra at Dyna Robotics.
About the role
This role is the central nervous system for Dyna's learning loop, owning the critical infrastructure that turns chaotic robot interactions into structured intelligence. You will architect the data platforms that capture, refine, and expose the signals necessary for our models to generalize and self-improve. The position demands a rare blend of backend rigor and product intuition, as you build systems that are both performant and approachable for researchers. You will be responsible for designing the interactive tools that allow humans to guide robotic behavior through teleoperation and skill capture. Your work will directly influence the quality and speed of our model training cycles. If you thrive on solving hard geometric problems and enabling others, you will find this role deeply impactful. You will own the end-to-end data lifecycle from raw bytes to labeled datasets ready for ML.
Key facts
What you'll do
- Interactive Data Systems: Architect the engines and interfaces that unify raw robot logs, video, and 3D sensor data. You will enable seamless "human-in-the-loop" workflows, from episode annotation to analyzing manual interventions.
- Signal Extraction & Geometry: Design and implement algorithms to extract structured signals (trajectories, events, 3D poses) from raw captures. Strong mathematical intuition is required to turn pixels and point clouds into ground-truth insights.
- Evaluation & Benchmarking: Build high-performance tools to compare model-driven motion against human-captured data, helping the team quantify model progress across diverse tasks.
- Scalable ML Pipelines: Build and operate distributed data pipelines (using Python, GCP/AWS, and Kubernetes) for the ingestion, transformation, and validation of terabytes of multimodal data.
- Observability & Debugging: Develop visualization tools that make complex model behaviors and sensor data easy to interpret, reducing the time from "data collected" to "model trained."
- Startup Fluidity: Collaborate across ML, Robotics, and Product teams. As an early member of the data team, you will help define the roadmap where no blueprint yet exists.
- Data Curation & Quality: Implement robust validation frameworks to ensure the integrity and consistency of datasets used for training. You will create automated checks that catch anomalies before they poison the model.
- Infrastructure as Code: Treat data pipelines and visualization dashboards as production software, applying rigorous testing, monitoring, and versioning practices.
- Cross-functional Translation: Partner with robotics engineers to understand their data needs and translate them into scalable storage and retrieval strategies. You will bridge the gap between raw sensor output and actionable metrics.
Requirements
- The Experience: 5+ years of professional software experience, ideally with a focus on data-intensive or "human-in-the-loop" platforms.
- Technical Stack: Proficiency in Python (NumPy, Pandas) and a solid understanding of modern backend architectures. Experience with React/TypeScript is a major plus for building internal observability tools.
- Mathematical Core: Strong skills in algorithms and geometric calculations (e.g., coordinate transformations, 3D spatial reasoning).
- Data Mastery: Hands-on experience with relational and NoSQL databases (PostgreSQL, Redis) and cloud-native infrastructure (GCP/AWS).
- Problem-Solving: The ability to debug real-world data issues - from sensor drift to pipeline bottlenecks - independently.
- Reliability Focus: A mindset centered on building fault-tolerant systems that ensure data availability and consistency across distributed components.
- Collaboration Aptitude: Comfort working alongside researchers and engineers in a fast-moving environment where priorities evolve rapidly.
- Learning Agility: The capacity to quickly master new tools and domains, whether it is a new database or a robotic sensor modality.
Nice to have
- Experience with multimodal data (video, LiDAR, time-series) in robotics or autonomous systems.
- Familiarity with Airflow, Kubeflow, or similar distributed batch processing systems.
- Experience with experiment tracking frameworks like Weights & Biases or MLFlow.
- Experience as an early hire in a fast-paced startup environment.
Practical notes
- Compensation details are not provided in the source material.
- The role is full-time.
- The location is specified as Redwood City, California.
- No specific hours, travel requirements, visa sponsorship details, or application deadlines are mentioned in the source.
At Dyna Robotics, we build technology for the real world, which requires a team as diverse as the environments our robots inhabit. We are an equal opportunity employer committed to technical rigor and mutual respect.
Don't let a checklist stop you. Data shows that underrepresented groups often only apply if they meet 100% of the criteria. We value problem-solving and grit over keyword matching. If you're passionate about closing the loop between deployed robots and better models, we want to hear from you, even if you don't check every box.