Senior Software Engineer
Job description
About the role
The Senior Software Engineer, ML Infrastructure is tasked with designing and building the production-grade inference infrastructure that powers SambaNova's serving stack on the Reconfigurable Dataflow Unit (RDU) architecture. This role is central to SambaNova's inference-first mission, translating cutting-edge inference techniques into reliable, high-throughput, low-latency services delivered via SambaStack and SambaCloud. The engineer will own the full lifecycle of complex systems, from request scheduling and advanced decoding algorithms to caching layers and accuracy infrastructure, ensuring the stack remains trustworthy and performant. Close collaboration with ML, compiler, runtime, and product teams is essential to ship inference features from initial prototype through to stable production deployment. This position represents a critical contribution to the core platform that enables enterprise and government organizations to unlock hidden value from their data using generative AI.
Key facts
What you'll do
Design and productionize advanced inference techniques on RDU to optimize for performance and cost, focusing on areas such as speculative decoding, constrained decoding, function and tool calling, prompt caching, and long-context inference.
Own SambaNova's integration with vLLM and other adjacent serving frameworks, adapting their capabilities to align with and leverage the unique characteristics of RDU's architecture.
Own the public inference API surface and its associated behaviors exposed through SambaStack and SambaCloud, defining contract stability and operational semantics.
Build and maintain the comprehensive accuracy verification and regression infrastructure that acts as a gate, ensuring every inference feature meets strict quality standards before reaching customers.
Partner proactively with ML, compiler, runtime, and product teams to shepherd inference features from early prototype stages through design reviews and into robust production services.
Contribute authoritative technical design inputs, lead code reviews, and drive architectural decisions as a senior individual contributor for the inference platform.
Develop and maintain deep observability into the inference stack, creating metrics, traces, and logs that illuminate performance characteristics and potential bottlenecks across the system.
Champion reliability and availability engineering practices, designing infrastructure that is resilient, fault-tolerant, and capable of sustaining enterprise-grade service levels.
Evaluate and prototype emerging inference methodologies, assessing their feasibility, performance, and integration complexity within the SambaNova hardware and software ecosystem.
Optimize the end-to-end dataflow and resource utilization on RDU, ensuring the serving stack efficiently handles diverse workloads while maximizing throughput and minimizing latency.
Collaborate with security and compliance teams to ensure that inference features and APIs adhere to organizational standards and regulatory requirements.
Document system behaviors, interfaces, and operational procedures to enable other engineers to effectively use, extend, and troubleshoot the serving infrastructure.
Requirements
Bachelor's degree in Computer Science, Electrical Engineering, or a related technical field is required.
Possess 5 or more years of industry experience focused on building and operating large-scale distributed systems, with demonstrated work in ML serving environments.
Demonstrate strong software engineering fundamentals, including mastery of algorithms, data structures, concurrency models, and robust systems design principles.
Have a proven track record of designing and maintaining production services that must meet strict latency, throughput, and availability requirements under varying loads.
Show working knowledge of modern Large Language Model inference techniques and hands-on experience with open-source serving stacks such as vLLM, TensorRT-LLM, or SGLang.
Be highly proficient in Python, with the ability to write clean, maintainable, and performant code for complex backend systems.
Have substantial experience collaborating across multiple disciplines, including ML researchers, compiler engineers, runtime specialists, and product managers, to deliver cohesive system-level solutions.
Commit to adhering to the company's contribution guidelines and code review standards to ensure the quality and consistency of the shared codebase.
Nice to have
Prior experience contributing to high-performance serving infrastructures or open-source inference projects is strongly preferred.
Familiarity with hardware acceleration concepts and co-design principles for AI accelerators provides a distinct advantage in optimizing the stack for RDU.
Experience with cloud-native deployment patterns, container orchestration, and infrastructure-as-code practices is beneficial for managing production services at scale.
Practical notes
This is a full-time remote position based in the United States.
Submission of an application form for this specific role is mandatory for consideration.