Helix AI Engineer, Generative AI
Job description
About the role
Figure is pioneering the development of general-purpose humanoid robots intended to operate reliably within complex and unstructured real-world environments. As a member of the Helix team, you will own the design and implementation of generative systems that empower our robots to simulate and interpret physical surroundings through advanced multimodal models. You will be responsible for creating the foundational models that allow machines to perceive and reason about dynamic scenes using raw sensory inputs. This role requires you to build robust pipelines that translate high-dimensional observations into actionable representations for autonomous decision-making. You will work at the intersection of large-scale machine learning and robotics to push the boundaries of what humanoid systems can understand and predict. Your contributions will directly influence how Figure robots learn from and adapt to new environments over time. You will partner closely with researchers and engineers to turn novel generative techniques into production-grade capabilities that run on real hardware.
Key facts
What you'll do
- Architect, train, and launch large-scale generative models, specifically diffusion-based systems for video, vision, and multimodal data.
- Create models that enhance how robots perceive, predict, and model the world using raw sensory input.
- Develop generative pipelines to produce synthetic data and scale datasets for robot learning.
- Research and apply current methods in generative modeling and foundation models.
- Refine training workflows for large-scale models on distributed infrastructure.
- Collaborate with the agent, data, and training infrastructure teams to embed generative models into the autonomy stack.
- Assess model performance and generalization across real-world environments.
- Build frameworks for testing and experimenting with generative model development.
- Implement end-to-end experiments that validate the usefulness of generated content for robotic decision-making.
- Optimize data throughput and storage strategies to support continuous model training cycles.
- Instrument systems to capture failure modes and generate insights for future model improvements.
- Partner with product teams to align generative capabilities with real robotic use cases and customer needs.
- Explore novel evaluation metrics that capture perceptual quality and physical relevance of generated outputs.
- Maintain detailed documentation of model behaviors, data sources, and training configurations for reproducibility.
Requirements
- Proven track record of training and deploying generative models, such as autoregressive or diffusion approaches, at scale.
- Deep knowledge of modern deep learning methods for multimodal or vision systems.
- High proficiency in Python and PyTorch.
- Experience managing distributed training systems and large datasets.
- Ability to iterate on model performance with experimental rigor.
- Strong software engineering skills focused on building maintainable and reliable systems.
- Capacity to own technical problems independently in ambiguous environments.
- Demonstrated ability to debug complex model pipelines and resolve issues in production settings.
- Familiarity with version control, testing frameworks, and code review practices for machine learning projects.
- Willingness to follow security, privacy, and compliance guidelines relevant to robotics and data usage.
- Commitment to writing clean, modular, and well-tested code that can be maintained by cross-functional teams.
- Openness to receiving feedback from peers and incorporating it into iterative improvements of model performance.
- Readiness to work with incomplete data and make informed assumptions to drive projects forward.
- Understanding of hardware constraints and their impact on model design and deployment.
Nice to have
- Experience with diffusion models applied to video or image generation.
- Background in vision-language or vision-language-action foundation models.
- Expertise in synthetic data generation or simulation for embodied AI and robotics.
- Experience optimizing training on GPU clusters or multi-node systems.
- Familiarity with world models, 3D, or video prediction.
- Previous work in real-world ML systems or robotics.
- A history of publications in generative modeling, computer vision, or machine learning.
Practical notes
Base salary for this position ranges from $200,000 to $400,000. Final compensation depends on individual skills, experience, and job-related knowledge. Additional benefits and compensation components may be provided and will be detailed if an offer is extended. This role requires full-time, in-office presence in San Jose.