Principal AI SoC Runtime Software Architect
Job description
About the role
The is responsible for defining and executing the technical vision for the AI inference runtime across the entire System on Chip. This role owns the end-to-end architecture that spans user space applications, kernel drivers, and heterogeneous hardware engines to ensure a cohesive and efficient system. The position requires deep collaboration with hardware, firmware, and compiler teams to establish robust execution models and interfaces. You will set the standard for performance, reliability, and compatibility across all AI and media workloads. The role demands a hands-on leader capable of translating complex system requirements into clear, implementable architectures. You will own the definition of memory, buffer, and execution models that underpin the entire sensor-to-application pipeline. Success in this role will be measured by the ability to deliver a runtime platform that enables predictable, scalable, and optimized AI inference across diverse scenarios.
Key facts
What you'll do
- Define the SoC-wide execution model that orchestrates heterogeneous compute and media engines, incorporating dependency management, resource arbitration, priority, QoS, and concurrent pipeline behavior.
- Set the architectural and technical direction for the AI inference runtime, covering compiled-model execution, integration with industry-standard AI execution frameworks, and stable interfaces for applications and the Velaura SDK.
- Own the end-to-end dataflow architecture for sensor-to-application pipelines, ensuring seamless operation across camera inputs, media ingestion, preprocessing, inference execution, and postprocessing stages.
- Define the SoC-wide memory and buffer-sharing architecture spanning user space, kernel space, and heterogeneous hardware engines, establishing clear semantics for ownership, coherency, isolation, synchronization, and lifecycle management.
- Define the division of responsibility and interface contracts among runtime, kernel drivers, firmware, and hardware engines, covering execution flows, completion signaling, telemetry collection, fault management, and recovery procedures.
- Partner closely with the compiler team to define a compiler-runtime contract that ensures compiled artifacts contain necessary metadata and execution information for runtime loading, validation, execution, profiling, and compatibility maintenance across releases.
- Establish system-wide observability and performance analysis architectures that correlate behavior across software and hardware layers to enable optimization against latency, throughput, bandwidth, power consumption, utilization, and predictability targets.
- Define runtime resilience and validation frameworks, including fault-containment and recovery policies, architecture-level acceptance criteria, and qualification methodologies across correctness, concurrency, compatibility, performance, and sustained workload scenarios.
- Develop and maintain rigorous integration test strategies that validate dataflow integrity, resource management, and timing constraints across the entire stack under varying load and environmental conditions.
- Drive the creation of reference implementations and proof-of-concept pipelines to validate architectural decisions and demonstrate capabilities to internal and external stakeholders.
- Collaborate with product teams to ensure the runtime architecture aligns with evolving product requirements, balancing flexibility, performance, and time-to-market considerations.
- Champion best practices for software engineering, including code reviews, documentation, modularity, and testability, to ensure the runtime stack is maintainable and extensible over the long term.
- Act as the primary technical authority for runtime behavior, providing deep expertise during debugging, performance tuning, and platform certification activities.
- Mentor engineers across software layers on runtime principles, interfaces, and constraints to elevate the overall technical capability of the organization.
- Contribute to the broader technical community through internal design reviews, technical proposals, and knowledge transfer sessions that clarify trade-offs and rationale.
Requirements
- Demonstrate extensive experience designing and building production runtime systems, embedded middleware, multimedia frameworks, or other performance-critical systems software in demanding environments.
- Show strong C/C++ programming skills with a proven ability to architect and contribute hands-on code to production runtime software that spans application-facing APIs, user-space libraries, and low-level driver, firmware, and hardware interfaces.
- Exhibit a strong understanding of heterogeneous and asynchronous execution models, including command submission, queues, events, dependencies, synchronization, concurrency, scheduling, and resource management mechanisms.
- Possess a deep understanding of device memory, DMA, IOMMU/SMMU, cache coherency, memory mapping, shared buffers, buffer lifetimes, and kernel/user-space memory interfaces and their implications for system design.
- Have experience optimizing end-to-end data movement and execution across multiple hardware engines rather than focusing exclusively on the performance of individual kernels or accelerators.
- Demonstrate a track record of designing stable runtime APIs with well-defined compatibility guarantees, versioning strategies, error handling, diagnostics, and recovery behaviors.
- Show proven ability to debug complex cross-layer correctness and performance problems using disciplined, data-driven methods, profiling, and tracing methodologies.
- Display demonstrated technical leadership across component and organizational boundaries, including the capacity to translate system requirements into clear architectures, interfaces, implementation guidance, and validation strategies.
- Bring a strong sense of ownership for system-level quality, including performance, power, reliability, and security, and the ability to influence design decisions across the development lifecycle.
- Communicate effectively with diverse engineering teams, articulating technical trade-offs and rationale clearly to both specialized and generalist audiences.
Nice to have
- Direct experience developing or extending AI inference runtimes or execution providers, such as ONNX Runtime, TensorRT-like runtimes, OpenVINO, TensorFlow Lite delegates, Qualcomm QNN/SNPE, TVM runtimes, or comparable systems for NPUs, GPUs, DSPs, or other accelerators.
- Background in Linux kernel development or upstream contribution experience involving device, accelerator, media, or shared-memory subsystems.
- Hands-on experience with robotics, autonomous systems, edge AI, ROS 2, camera pipelines, ISP integration, V4L2/media frameworks, GStreamer, or other sensor-driven workloads.
- Familiarity with compiled-model artifacts, quantized execution, tensor layouts, graph partitioning, memory planning, and compiler/runtime integration techniques.
- Experience with embedded Linux SDKs, production deployment challenges, long-term runtime/API compatibility maintenance, or security and isolation requirements of multi-process accelerator systems.
Practical notes
Note: This role requires frequent travel within the Santa Clara area for collaboration and validation activities.