Principal System Software Engineer, AI Inference Execution
Job description
Principal System Software Engineer, AI Inference Execution at D Matrix
About the role
We are looking for a highly experienced Principal System Software Engineer to join our team at D Matrix, focusing on AI inference execution. In this role, you will be responsible for developing, optimizing, and maintaining the software stack that powers our AI compute engine. Your work will be central to productizing our AI deployment solutions, ensuring high performance, scalability, and reliability. You will collaborate closely with hardware and software teams to build infrastructure that supports AI workloads, including large language models and other advanced machine learning models. This position is based on-site at our Santa Clara headquarters, where you will work with a talented team dedicated to pushing the boundaries of AI technology.
Key facts
What you'll do
- Lead the development, enhancement, and maintenance of the next-generation AI deployment software stack for our AI compute engine.
- Collaborate with hardware, ML, and compiler teams to optimize hardware-software co-design trade-offs, ensuring maximum efficiency and performance.
- Build scalable, high-performance deployment infrastructure that can handle demanding AI workloads, including large language models and vision-language models.
- Work across the full software stack, including system software, data structures, and computer architecture, to deliver solutions.
- Develop and implement distributed, high-performance software systems tailored for AI inference workloads, ensuring low latency and high throughput.
- Optimize software components for deployment in production environments, including testing, debugging, and performance tuning.
- Contribute to the integration of inference frameworks such as TensorRT-LLM, vLLM, or SGLang, ensuring compatibility and performance.
- Work closely with ML engineers and hardware specialists to troubleshoot and resolve complex deployment issues.
- Lead efforts to improve automation, testing, and deployment pipelines to streamline software updates and releases.
- Support the deployment of large language models, vision-language models, and NLP workloads on distributed systems, ensuring scalability and robustness.
- Participate in defining best practices for deploying AI models at scale, including containerization and orchestration strategies.
- Contribute to documentation, training, and knowledge sharing within the team to foster continuous improvement.
- Stay current with emerging trends in AI hardware and software, applying new techniques to improve our deployment solutions.
- Mentor junior team members and promote a culture of technical excellence and innovation.
Requirements
- BS in Computer Science, Engineering, Math, Physics, or related field with 12+ years of industry software development experience; MS preferred with 6+ years.
- Strong understanding of system software, data structures, computer architecture, and machine learning fundamentals.
- Proficiency in C/C++ and Python development within a Linux environment, using standard development tools.
- Experience designing and implementing distributed, high-performance software systems for AI workloads.
- Demonstrated ability to build scalable software solutions within tight development timelines.
- Self-motivated team player with a strong sense of ownership, leadership, and collaboration skills.
- Excellent debugging, testing, and performance optimization skills.
- Familiarity with hardware acceleration techniques and optimization strategies for AI inference.
- Ability to work effectively across teams, communicating complex technical concepts clearly.
- Experience with deploying AI workloads, including large language models and vision-language models, on distributed systems.
- Knowledge of containerization and orchestration tools such as Docker and Kubernetes is a plus.
- Prior experience working in fast-paced environments, such as startups or incubation teams, is desirable.
- Experience working at a cloud provider or AI compute/subsystem company is advantageous.
Nice to have
- MS or PhD in Computer Science, Electrical Engineering, or related fields.
- Experience with inference servers and model serving frameworks such as TensorRT-LLM, vLLM, or SGLang.
- Deep understanding of deep learning frameworks like PyTorch and TensorFlow.
- Familiarity with deep learning runtimes such as ONNX Runtime or TensorRT.
- Experience with distributed systems collectives such as NCCL and OpenMPI.
- Knowledge of deploying ML workloads (LLMs, VLMs, NLP) on distributed systems.
- Experience with MLOps tools like Kubernetes, Ray, or similar platforms used from definition to deployment.
- Prior startup or small team experience, especially in incubation environments.
- Work experience at a cloud provider or AI hardware and software company.
Skills & tools
- Python, Kubernetes, Machine Learning, Deep Learning, NLP, TensorFlow, PyTorch, LLM, AI, ML, Linux, Talent, Solutions, Engineering, Infrastructure, Testing, Full Stack
Practical notes
This role is based at our Santa Clara headquarters and requires onsite presence three days per week. We operate in a hybrid work environment, emphasizing collaboration and innovation. Our company values humility, respect, and open communication, fostering an inclusive culture where diverse perspectives are welcomed. We are committed to equal opportunity employment and encourage applicants from all backgrounds to apply. We do not accept resumes or candidate submissions from external agencies; interested candidates should apply directly through our official channels. Our team is about solving challenging problems in AI and building impactful solutions that shape the future of technology.