System Software Engineer - GPU & Accelerated Compute
Job description
System Software Engineer - GPU & Accelerated Compute at Sunday.
About the role
Work centers on GPU and accelerated compute to support diverse workloads. You will own the end to end stack that turns raw GPU hardware into reliable, low latency compute for robotics and ML pipelines. This includes scheduling, memory management, and synchronization so concurrent users get predictable performance for tasks such as model inference and SLAM. You will integrate tightly with hardware features like encode and decode through NVDEC and NVENC to move camera frames into GPU memory at low latency. Your code will keep inference pipelines full in real-time robotic systems by designing robust synchronization primitives and sharing patterns across users. You will explain complex GPU architecture and workload tradeoffs in plain language so product and partner teams can make informed decisions. You will work in small agile teams that review each other's work, ship in small batches, and spend part of every week on planning, code review, and debugging. The role values clear communication as much as technical skill, and you will be expected to defend design choices with data and reproducible results.
Key facts
What you'll do
GPU access is scheduled and time-sliced to serve concurrent users like model inference and SLAM with predictable latency.
Camera frames move into GPU memory at low latency by integrating with hardware encode and decode such as NVDEC and NVENC.
Inference pipelines stay full in real-time robotic systems through carefully designed synchronization primitives and patterns.
Requirements
Bring 2+ years of experience developing GPU systems software in demanding environments.
Write strong CUDA code in a systems language such as C++, C, or Rust at production scale with attention to correctness and performance.
Explain GPU architecture, workloads, and the tradeoffs of time-slicing and sharing compute across multiple users to both technical and non-technical audiences.
Implement robust accelerated compute solutions using the CUDA runtime API, CUDA Graphs, and CUDA IPC for efficient data and task sharing.
Support concurrent access patterns on robotic platforms with mechanisms like MPS and MIG to maximize utilization and isolation.
Profile accelerated workloads with tools such as Nsight Systems and Nsight Compute to meet strict latency and throughput targets.
Manage Linux fundamentals including scheduling, IPC, memory management, and performance tuning for robotic compute stacks that rely on GPU acceleration.
Collaborate with ML, perception, and hardware teams to align system software with product requirements and hardware constraints.
Own observability and debugging workflows for GPU-accelerated services to reduce mean time to resolution and improve reliability.
Contribute to making the stack extensible so that new accelerators and workloads can be integrated with minimal friction.
Nice to have
Contributions to CUDA libraries or other GPU programming libraries broaden accelerated compute options and simplify integration.
Camera pipeline integration and use of NVDEC/NVENC accelerate vision workloads and reduce CPU overhead on robotic hardware.
Optimize model inference on embedded GPU platforms such as Jetson for edge deployment scenarios where power and thermal constraints are critical.
Implement observability and tracing for GPU-accelerated workloads to simplify debugging, monitoring, and capacity planning.
Practical notes
US citizenship or permanent residency is required for this position. The role is based in Redwood City with potential for limited travel within the team. Typical interview steps include a recruiter screen, one or two technical rounds, and final team fit conversations. Candidates often solve a coding problem, discuss past projects, and answer system design questions, with some loops including a take-home task. Interviewers focus on how you break down unfamiliar problems, not just whether you reach the answer, so practicing a few problems aloud and reviewing your own projects is recommended.
Good to know
This role operates in a domain where real-time performance and efficient shared compute determine system behavior. General-purpose robotics software runs on GPU and accelerator hardware to meet strict latency constraints, requiring careful tuning of kernels, memory transfers, and scheduling policies. Tooling centers on CUDA, Linux performance tuning, and profiling systems for GPU workloads to ensure predictable throughput and latency. Collaboration across ML, perception, and hardware teams is central to success, as system software must balance the needs of researchers, engineers, and embedded platforms. You will work with production services that demand robustness, observability, and clear documentation so that multiple teams can depend on shared compute infrastructure.
Career growth
Engineering careers usually progress from individual contributor to senior, staff, and principal levels, with increasing ownership of cross team initiatives and architectural decisions. Some engineers move into management and lead teams of five to twenty people, while others stay on the technical track and deepen expertise in GPU systems and accelerated compute. Growth follows demonstrated impact, not tenure alone, and is evaluated through consistent delivery of reliable performance, clear communication, and mentorship of peers. A typical engineering ladder has clear levels with defined expectations for scope, quality, and mentorship, and moving up generally requires owning outcomes end to end rather than completing assigned tickets. You will have opportunities to influence platform direction, contribute to open source style internal libraries, and shape how observability and profiling tools evolve across the organization.