Staff Software Engineer, SIMD Kernels
Job description
Staff Software Engineer, SIMD Kernels at D Matrix
About the role
We are looking for a highly skilled Staff Software Engineer to join the SIMD Kernels team at D Matrix, a company dedicated to advancing the capabilities of generative AI through innovative hardware and software solutions. In this role, you will be responsible for developing, optimizing, and maintaining software kernels that power machine learning operators such as softmax, layer normalization, and activation functions, which are critical components of our next-generation AI hardware platform. Your work will directly impact the performance, efficiency, and usability of our AI compute engine, enabling developers to build more powerful AI applications. You will collaborate closely with hardware engineers, software developers, and AI researchers to ensure integration of algorithms onto hardware architectures, and to improve the developer experience through intuitive SDK solutions. This position offers a unique opportunity to influence both the software stack and hardware design, contributing to AI technology in a fast-paced, innovative environment.
Key facts
What you'll do
- Develop, optimize, and maintain software kernels for core machine learning operators including softmax, layer normalization, activation functions, GEMMs, convolutions, and pooling operations.
- Collaborate with hardware teams to understand hardware capabilities and constraints, ensuring efficient mapping of algorithms onto AI hardware architectures.
- Build and improve SDK tools that make it easier for developers to our software and analyze performance metrics effectively.
- Contribute to the full software stack, from algorithm implementation to hardware integration, ensuring high performance and scalability.
- Conduct performance tuning and profiling of kernels across various hardware platforms to maximize throughput and minimize latency.
- Support deployment and integration of kernels into the broader AI compute ecosystem, ensuring stability and robustness.
- Engage in hardware-software co-design efforts to optimize the performance of AI workloads on specialized hardware accelerators.
- Document technical designs, implementation details, and performance results to facilitate knowledge sharing and future development.
- Participate in code reviews, testing, and continuous integration processes to maintain high-quality software standards.
- Stay informed about the latest advancements in hardware architectures, AI frameworks, and optimization techniques to incorporate best practices.
- Work closely with cross-functional teams to troubleshoot issues, improve existing kernels, and develop new features aligned with product goals.
- Contribute to the development of internal tools and frameworks that support kernel development and performance analysis.
- Assist in mentoring junior team members and sharing expertise to foster a collaborative and innovative team environment.
- Represent the company at technical discussions, conferences, or industry events related to AI hardware and software development.
Requirements
- MS or PhD in computer engineering, computer science, mathematics, physics, or a related technical field with at least 5 years of relevant industry experience.
- Strong understanding of computer architecture, data structures, system software, and machine learning fundamentals.
- Proven proficiency in C and C++ programming within Linux environments, with experience using standard development tools such as compilers, debuggers, and version control systems.
- Extensive experience implementing algorithms in C/C++ and Python, especially for high-performance computing.
- Hands-on experience developing software kernels for specialized hardware such as FPGAs, DSPs, GPUs, or AI accelerators, utilizing libraries like CUDA or similar.
- Demonstrated ability to develop and optimize operators used in ML workloads, including GEMMs, convolutions, softmax, layer normalization, pooling, and activation functions.
- Strong problem-solving skills and the ability to analyze performance bottlenecks and implement effective solutions.
- Self-motivated with a strong sense of ownership, leadership qualities, and the ability to work independently and collaboratively.
- Excellent communication skills, with the ability to document technical work clearly and effectively.
- Experience working in a fast-paced environment, managing multiple priorities, and delivering high-quality software on tight deadlines.
Nice to have
- Previous experience working in startup environments, small teams, or incubation projects.
- Familiarity with machine learning frameworks such as TensorFlow and PyTorch.
- Knowledge of ML compiler frameworks and algorithms like MLIR, LLVM, TVM, or Glow.
- Experience developing for embedded SIMD vector processors such as Tensilica.
- Background working with cloud providers or companies specializing in AI compute and subsystems.
- Experience designing or optimizing deep learning models for computer vision, natural language processing, or recommendation systems.
- Familiarity with hardware design and hardware-software co-design methodologies.
- Knowledge of performance profiling tools and techniques specific to AI hardware.
Skills & tools
- C and C++
- Python
- Linux operating system
- CUDA or similar GPU programming libraries
- Hardware architecture knowledge, especially AI accelerators
- ML frameworks such as TensorFlow and PyTorch
- Compiler frameworks like MLIR, LLVM, TVM, Glow
- Performance profiling and debugging tools
- Version control systems (e.g., Git)
Practical notes
This position is based at our headquarters in Santa Clara, California, with the role being on-site at the office. While the company considers candidates from regional offices, the primary location is Santa Clara. We value humility, collaboration, and diversity, fostering an inclusive environment where different perspectives are welcomed and appreciated. Our culture emphasizes respect, direct communication, and a passion for tackling challenging problems in AI hardware and software development. We are an equal opportunity employer and committed to creating a workplace free from discrimination. We do not accept resumes or candidate submissions from external agencies; interested applicants should apply directly through our official channels to ensure a fair and streamlined hiring process.