Staff AI Inference and Acceleration Engineer
Job description
About the role
Figure is at the forefront of building autonomous humanoid robots designed for both commercial and residential applications. This role involves taking ownership of the on-board inference architecture, ensuring that the robots can operate with human-level intelligence while maintaining strict power efficiency and low latency. The ideal candidate will work closely with hardware and AI/ML teams to optimize inference performance across heterogeneous compute platforms, including NPUs, GPUs, DSPs, and CPUs. The position requires a deep understanding of machine learning deployment, hardware acceleration, and embedded systems, with a focus on creating scalable and efficient inference solutions that meet real-world operational constraints.
Key facts
What you'll do
- Design and develop the on-board inference architecture by mapping models efficiently to various hardware accelerators such as NPUs, GPUs, DSPs, and CPUs to maximize performance and power efficiency.
- Balance inference workloads across heterogeneous hardware components to meet strict thermal, latency, and power consumption targets, ensuring optimal utilization of all available compute resources.
- Establish and manage a comprehensive compute budget for all inference tasks on the robot, coordinating resource allocation to meet real-time operational requirements.
- Evaluate emerging acceleration hardware options, including new NPUs, GPUs, and other accelerators, and contribute to defining the specifications for the robot's compute platform to ensure future scalability and performance.
- Refine and optimize inference toolchains, from initial model export through to final runtime execution, ensuring models run efficiently on embedded hardware with minimal latency and power consumption.
- Implement advanced model compression techniques such as pruning, operator fusion, and quantization methods including INT8 and INT4 to reduce model size and improve inference speed without sacrificing accuracy.
- Analyze inference pipelines to identify bottlenecks related to power usage, memory bandwidth, and latency, then develop solutions to mitigate these issues and improve overall inference throughput.
- Improve data movement strategies, kernel scheduling, and memory layout within the compute stack to enhance efficiency and reduce latency during inference operations.
- Collaborate closely with AI/ML teams to ensure that model architectures are designed with hardware constraints in mind, facilitating seamless deployment and optimal performance on target hardware.
- Work with the Platform Software team to integrate power management, scheduling, and runtime features that support inference workloads, ensuring system stability and efficiency.
- Maintain ongoing communication with silicon vendors and hardware suppliers to stay informed about developments in the accelerator market, influencing hardware design and feature sets to meet the company's needs.
- Contribute to the development of best practices and standards for inference deployment across the robot platform, ensuring consistency and reliability in inference performance.
- Participate in cross-disciplinary meetings to align hardware capabilities with AI/ML requirements, providing technical guidance on hardware limitations and opportunities for optimization.
- Document inference architecture designs, optimization strategies, and performance benchmarks to support knowledge sharing and future development efforts.
- Assist in troubleshooting inference-related issues during development and deployment phases, providing expert guidance to resolve hardware and software integration challenges.
- Support the evaluation and integration of new hardware accelerators, including testing and benchmarking to validate performance gains and compatibility.
- Engage in continuous learning about advancements in hardware acceleration, inference algorithms, and embedded systems to keep the team at the cutting edge of technology.
Requirements
- M.S. or Ph.D. in Computer Science, Electrical Engineering, Computer Engineering, or a related technical field, or equivalent industry experience demonstrating mastery in relevant areas.
- Minimum of 8 years of professional experience working with machine learning systems, compute architecture, or hardware acceleration, preferably in embedded or edge environments.
- Extensive hands-on experience with AI/ML inference deployment pipelines, including working with model formats such as ONNX and TFLite, and inference runtime environments.
- Practical experience tuning models for embedded or edge hardware, utilizing techniques such as pruning, quantization, and operator fusion to optimize inference speed and power consumption.
- Strong understanding of computer architecture principles, especially regarding heterogeneous compute systems, memory hierarchies, and power management strategies.
- Proven ability to benchmark and profile inference workloads on diverse hardware platforms, including NPUs, GPUs, DSPs, and CPUs, to identify performance bottlenecks and optimize accordingly.
- Familiarity with compilation frameworks and low-level toolchains such as MLIR, TVM, TensorRT, Torch, JAX, SNPE/QNN, ROCm, and CUDA, enabling efficient deployment and optimization of models.
- Proficiency in programming languages including Python and C++, with experience developing and debugging inference software and tools.
- Excellent communication skills, capable of translating complex technical requirements into actionable development tasks and collaborating effectively across hardware, software, and AI/ML teams.
- Ability to work in a fast-paced environment, managing multiple priorities while maintaining attention to detail and delivering high-quality results.
Nice to have
- Knowledge of how real-time operating system constraints influence inference scheduling and resource management, enabling more precise optimization strategies.
- Experience co-designing model architectures alongside ML teams to ensure compatibility with hardware limitations, leading to more efficient deployment and inference performance.
- Familiarity with hardware security considerations, especially related to inference on embedded systems, to ensure robustness and compliance with security standards.
- Understanding of ITAR regulations and ability to work within compliance frameworks when dealing with sensitive hardware or software components.
Skills & tools
- C++, Python, ONNX, TFLite, TVM, MLIR, TensorRT, Torch, SNPE/QNN, JAX, CUDA, ROCm, NPU, GPU, DSP, CPU.
Practical notes
Compensation for this role is determined by individual skills, experience, and relevant job knowledge. Additional benefits and compensation components may be provided upon an employment offer. The role involves working on-site at the San Jose, CA office, with no remote or hybrid options specified. The position is fully on-site, requiring physical presence at the company's location. The company emphasizes strict adherence to factual details, and all provided information reflects the current scope of the role without additional assumptions or inventions.