Principal AI/ML System Software Engineer
Job description
Principal AI/ML System Software Engineer at D Matrix.
About the role
D Matrix is pioneering high-performance computing solutions designed specifically to accelerate generative artificial intelligence workloads. The organization seeks a seasoned technical leader to drive the productization of its novel AI compute engine and to mature the surrounding deployment infrastructure. This position sits at the critical intersection of hardware innovation and software execution, requiring a practitioner who can navigate complex co-design trade-offs. The successful candidate will own significant portions of the software lifecycle, ensuring the stack delivers maximum throughput and efficiency on proprietary silicon. It is a high-impact opportunity to define the software foundation for a next-generation AI hardware platform.
Key facts
What you'll do
- Architect, develop, and sustain the comprehensive software stacks required to deploy next-generation AI models on D Matrix compute engines, ensuring robustness and performance.
- Manage and evolve the full-stack toolchain, making critical decisions that balance hardware capabilities against software usability and compiler efficiency.
- Deliver scalable, production-grade software solutions while operating within aggressive development schedules typical of a fast-moving silicon startup.
- Partner deeply with compiler engineers, machine learning researchers, and hardware architects to co-design and build the deployment infrastructure bridging models and silicon.
- Lead the optimization of inference pipelines, focusing on latency reduction, throughput maximization, and efficient memory utilization across distributed topologies.
- Define and enforce software engineering best practices, including code review standards, continuous integration pipelines, and automated testing frameworks for system-level validation.
- Troubleshoot and resolve complex cross-layer issues spanning kernel drivers, runtime libraries, compiler backends, and model serving layers.
- Mentor senior and staff engineers, providing technical guidance on system design, performance profiling, and debugging methodologies.
- Collaborate with product management to translate customer deployment requirements into concrete software features and roadmap priorities.
- Evaluate and integrate emerging model architectures (such as Mixture of Experts or novel attention mechanisms) into the supported software stack rapidly.
- Drive the bring-up and validation of new hardware revisions, working closely with RTL and physical design teams to debug silicon features.
- Contribute to the strategic technical direction of the software organization, identifying technical debt and proposing architectural refactors.
Requirements
- Bachelor of Science in Computer Science, Computer Engineering, Mathematics, Physics, or a closely related technical discipline with a minimum of 12 years of relevant industry experience, OR a Master of Science in a related field with 6+ years of experience.
- Deep, practical understanding of computer architecture principles, including memory hierarchies, cache coherence, interconnect topologies, and instruction set architectures.
- Expert-level proficiency in C and C++ for systems programming, alongside strong Python skills for automation, scripting, and ML framework integration, all within Linux environments.
- Proven track record designing, implementing, and shipping high-performance, distributed software systems that operate at scale.
- Demonstrated ability to act as a technical lead and owner for complex, multi-engineer projects from conception through production release.
- Solid grasp of machine learning fundamentals, specifically regarding model execution graphs, operator fusion, quantization techniques, and tensor manipulation.
- Experience working across the full stack, from low-level firmware or driver interactions up to high-level application APIs.
- Strong communication skills necessary for effective collaboration with hardware teams, compiler teams, and external partners.
Nice to have
- Advanced degree (Master of Science or Doctor of Philosophy) in Computer Science, Electrical Engineering, or a directly relevant field.
- Hands-on experience with modern Large Language Model serving frameworks such as vLLM, TensorRT-LLM, or SGLang.
- Deep familiarity with deep learning training and inference frameworks, particularly PyTorch and TensorFlow internals.
- Prior work with high-performance inference runtimes and execution engines, including ONNX Runtime and NVIDIA TensorRT.
- Knowledge of distributed communication libraries and collective operations, specifically OpenMPI and NCCL, for multi-node scaling.
- Background in rigorous software testing methodologies, including unit, integration, fuzz, and hardware-in-the-loop testing.
- Practical experience deploying Large Language Models (LLMs), Vision Language Models (VLMs), or complex NLP workloads on distributed GPU or accelerator clusters.
- Proficiency with MLOps orchestration platforms such as Ray and container orchestration via Kubernetes.
- Previous tenure at a major cloud provider, an AI accelerator startup, or a high-performance computing organization.
Skills & tools
- C, C++, Python, Linux, Distributed Systems, AI Inference, Hardware-Software Co-design, Compilers, Runtimes, PyTorch, TensorFlow, vLLM, TensorRT-LLM, ONNX Runtime, TensorRT, NCCL, MPI, Kubernetes, Ray, Git, CI/CD, Profiling Tools (perf, VTune, Nsight), Silicon Bring-up.
Practical notes
- D Matrix operates as an equal opportunity and affirmative action employer committed to workplace diversity.
- The company does not accept unsolicited resumes or candidate submissions from external recruitment agencies; all applications must be submitted directly through official company channels.
- The hybrid work policy mandates physical presence at the Santa Clara headquarters three days per week.
- Total compensation package includes a competitive base salary range of $195,000 to $285,000, supplemented by equity grants and performance-based bonuses.
- Role requires eligibility to work in the United States; specific visa sponsorship details should be confirmed during the recruitment process.