Senior Staff LLM Inference Engineer
Job description
Senior Staff LLM Inference Engineer at D Matrix.
About the role
You will own the full lifecycle of LLM inference features on D-Matrix's computational fabric, translating ambiguous research ideas into production-ready systems with measurable performance and power outcomes. You will collaborate directly with hardware architects and firmware teams to ensure software workloads fully leverage the unique capabilities of our silicon. This role requires you to design and validate novel deployment patterns that exploit heterogeneity across CPUs, GPUs, and custom accelerators. You will be responsible for building the tools and runtimes that allow frontier models to run efficiently and cost-effectively in real customer environments. You will partner with product and business development to turn internal prototypes into compelling customer demonstrations and proof points. You are expected to drive technical publications and open-source contributions that enhance D-Matrix visibility in the inference ecosystem.
Key facts
What you'll do
- Identify and prototype emerging LLM inference use cases that justify heterogeneous hardware deployments and unique D-Matrix capabilities.
- Build compelling proof-of-concept systems that demonstrate D-Matrix end-to-end stack advantages to customers, partners, and internal stakeholders.
- Develop and tune custom kernels and operator-level optimizations to maximize throughput and minimize latency across diverse workloads.
- Drive quantization, sparsity, and advanced batching strategies specifically tailored to the D-Matrix computational model and memory hierarchy.
- Build and maintain high-performance inference runtimes, serving frameworks, and evaluation tooling that integrate with production pipelines.
- Architect and optimize distributed inference systems, including tensor and pipeline parallelism, disaggregated prefill and decode paths, and adaptive KV-cache management.
- Work closely with hardware architects to provide firmware and compiler teams with detailed inference workload insights and actionable feedback.
- Partner with product and business development to translate novel POCs into customer-facing demonstrations and market-ready solutions.
- Implement benchmarking and evaluation pipelines that validate performance, power, and latency characteristics across heterogeneous deployments.
- Contribute to technical publications, whitepapers, and open-source projects that advance D-Matrix visibility and technical leadership.
- Explore and prototype next-generation serving techniques such as speculative decoding, mixture-of-experts routing, and long-context serving on real hardware.
- Ensure reliability, scalability, and maintainability of deployed inference systems operating in production-like environments.
Requirements
- Bachelor's degree in Computer Science, Electrical Engineering, or a related field, and 10+ years of relevant engineering experience; or equivalent demonstrated experience.
- Master's or PhD in Computer Science, Electrical Engineering, or a related field preferred, with 6+ years of relevant industry experience.
- Strong proficiency in Python and C/C++ for systems-level development and performance-critical implementations.
- Hands-on experience optimizing LLM inference, including attention kernels, KV-cache management, batching strategies, and quantization methods such as INT8, FP8, and INT4.
- Experience with at least one major inference framework (vLLM, SGLang, TensorRT-LLM, ONNX Runtime, or similar) at a contributor level for customization and extension.
- Familiarity with GPU kernel programming using CUDA or Triton and expertise in performance profiling and bottleneck analysis on real hardware.
- Demonstrated ability to design and debug complex software systems that interact closely with hardware and firmware components.
- Comfort with agile development practices, version control, and rigorous code review in fast-paced engineering environments.
Nice to have
- Experience with heterogeneous compute deployments, scheduling inference workloads across dissimilar hardware such as accelerators, CPUs, and GPUs.
- Familiarity with custom silicon or ASIC-based inference platforms that extend beyond conventional GPU-only environments.
- Contributions to open-source inference or ML systems projects that are widely used in production settings.
- Experience with production inference serving at scale, including strict latency SLOs, continuous batching mechanisms, and multi-model serving architectures.
- Working familiarity with speculative decoding, mixture-of-experts routing, and advanced long-context serving techniques.
- Prior exposure to distributed training-inference co-design or compiler-level optimizations targeting novel hardware fabrics.
Practical notes
The role is based in Santa Clara and requires full-time on-site availability. Travel may be required to support customer deployments or hardware validation activities. Candidates must be authorized to work in the United States without sponsorship for this position.