Principal Software Engineer, Kernels
Job description
Principal Software Engineer, Kernels at D Matrix.
About the role
Join our software team to productize the stack for our AI compute engine. You will develop and maintain high-performance kernels for our next-generation hardware while collaborating across disciplines to optimize hardware-software co-design. In this role, you will own the implementation of critical software kernels that bridge algorithmic innovation with silicon efficiency. You will work at the intersection of compilers, computer architecture, and machine learning to solve demanding performance challenges. Your contributions will directly influence how AI models execute on our specialized compute fabrics in production environments. You will partner with cross-functional experts to define, prototype, and deliver robust software solutions under tight constraints. This position requires a deep commitment to low-level optimization and a passion for building the foundational layers of an AI compute stack. You will be expected to translate abstract requirements into concrete, high-quality, and maintainable kernel implementations.
Key facts
What you'll do
- Develop, enhance, and maintain software kernels for new AI hardware across the entire compute pipeline.
- Map complex computational graphs from mainstream AI frameworks onto our unique underlying architecture with precision.
- Partner with compiler experts to design, prototype, and iteratively improve our compiler infrastructure and toolchains.
- Liaise closely with teams spanning machine learning, systems, mixed-signal, digital signal processing, and CPU engineering.
- Analyze and manage intricate full-stack toolchain development while navigating difficult hardware-software trade-offs.
- Implement optimized versions of fundamental machine learning operators including GEMMs, Convolutions, and BLAS routines.
- Leverage advanced SIMD techniques and instruction-level optimizations to maximize throughput and minimize latency.
- Integrate and utilize specialized compiler and MLIR-based toolchains such as LLVM, TVM, and Glow for our target platforms.
- Build and debug kernels for demanding AI workloads, ensuring correctness, stability, and peak performance on embedded targets.
- Create comprehensive tests and profiling methodologies to validate kernel behavior and drive continuous performance improvements.
- Adapt and optimize kernels for specialized processing elements like embedded SIMD vector processors including those from Tensilica.
- Collaborate on the definition and refinement of low-level interfaces between firmware, drivers, and application-level ML models.
- Document kernel implementations, optimization strategies, and architectural assumptions for internal knowledge sharing.
- Participate in on-call rotations to triage and resolve critical performance regressions or functional issues in deployed kernels.
- Contribute to open-source compiler projects and internal frameworks when strategically aligned with product goals.
Requirements
- Hold an MS degree in computer engineering, math, physics, or a closely related field with a minimum of 10 years of relevant industry experience.
- Alternatively, possess a PhD degree with at least 5 years of industry experience in relevant fields.
- Demonstrate exceptional proficiency in the C and C++ programming languages within complex Linux-based development environments.
- Apply strong Python scripting skills for automation, testing, and integration tasks throughout the development lifecycle.
- Have hands-on experience implementing algorithms specifically designed for specialized hardware such as GPUs, FPGAs, DSPs, or AI accelerators.
- Utilize industry-standard tools like CUDA where applicable to accelerate development and verification of hardware-specific code.
- Maintain deep knowledge of modern computer architecture principles, system software internals, and advanced data structures.
- Apply a solid understanding of machine learning fundamentals, including model architectures, training workflows, and inference patterns.
- Implement a broad range of ML operators from scratch, including GEMMs, Convolutions, BLAS, SIMD operations, softmax, layer normalization, and pooling.
- Bring experience developing software for embedded SIMD vector processors, with a background on platforms such as Tensilica.
- Show a proven track record of writing efficient, low-latency, and memory-conscious code for performance-critical systems.
- Exhibit strong problem-solving abilities when dealing with obscure bugs, timing issues, or hardware-specific limitations.
- Communicate effectively with multidisciplinary engineering teams to align on requirements and resolve technical challenges.
- Take ownership of tasks end-to-end, driving issues from initial investigation through design, implementation, and final validation.
- Adhere to strict coding standards, version control practices, and software quality metrics within a fast-paced engineering organization.
Nice to have
- Bring experience working within startups or small, fast-paced engineering teams where impact is immediate and direct.
- Apply familiarity with major ML frameworks such as PyTorch or TensorFlow to prototype and validate kernel functionality.
- Demonstrate hands-on experience with ML compiler stacks and associated projects like MLIR, LLVM, TVM, or Glow.
- Have a background in optimizing ML models for natural language processing, computer vision, or recommendation systems.
- Show a professional history within an AI compute company or a large cloud provider's infrastructure team.
Practical notes
- We are an equal opportunity employer committed to inclusive hiring.
- We do not accept submissions from external recruitment agencies; please apply directly.
- This role requires a hybrid work arrangement with a minimum of 3 days onsite in Santa Clara, California.
- Candidates must be authorized to work in the United States without sponsorship for this position.
- The listed compensation range reflects minimum and maximum target values inclusive of base salary, equity, and bonus components.