Senior Software Engineer - Kernels
Job description
Senior Software Engineer - Kernels at D Matrix.
About the role
You will own the design and delivery of software kernels that unlock the performance of next-generation AI hardware. You will collaborate closely with compiler experts and hardware engineers to translate high-level computational graphs into highly optimized low-level implementations. The role demands a strong partnership with machine learning and systems teams to ensure kernels meet accuracy, latency, and throughput targets. You will be responsible for navigating the full-stack toolchain to resolve complex performance and correctness challenges. Your work will directly influence how efficiently AI models execute on d-Matrix compute engines. You will contribute to building scalable software components under aggressive development timelines. You will champion direct communication and humility to drive solutions in a fast-paced, inclusive environment.
Key facts
What you'll do
Architect and develop high-performance software kernels tailored for AI accelerators and specialized hardware platforms.
Analyze computational graphs from AI frameworks and devise strategies to map them efficiently onto underlying hardware architectures.
Implement optimized versions of machine learning operators such as GEMMs, convolutions, and SIMD-based functions for activations and normalization.
Leverage expertise in C/C++ and Python to build, debug, and profile low-level kernel modules in Linux-based development environments.
Work alongside compiler specialists to integrate kernel implementations into broader compiler infrastructure such as MLIR and LLVM ecosystems.
Collaborate with mixed-signal, DSP, and CPU engineers to align software behavior with hardware specifications and constraints.
Utilize frameworks like CUDA and related libraries to accelerate workloads on GPUs and AI accelerators where applicable.
Adapt and optimize algorithms for embedded SIMD processors, including architectures such as Tensilica, to meet real-time processing needs.
Conduct performance analysis and trade-off evaluations to balance hardware utilization, memory bandwidth, and compute intensity.
Deliver reliable software increments within tight development windows while maintaining code quality and documentation.
Partner with machine learning experts to validate kernel behavior against model outputs and ensure numerical correctness.
Contribute to the productization of the software stack by turning experimental capabilities into robust, deployable features.
Support debugging and optimization across the full-stack pipeline from model ingestion to kernel execution.
Requirements
Must hold a Master of Science in computer engineering, math, physics, or a related discipline with a minimum of five years of industry experience.
Alternatively, hold a PhD in computer engineering, math, physics, or a related discipline with at least one year of industry experience.
Demonstrate a strong grasp of computer architecture, data structures, system software, and machine learning fundamentals.
Be proficient in C/C++ and Python development within Linux environments and comfortable using standard development and debugging tools.
Have hands-on experience implementing algorithms in high-level languages such as C/C++ and Python for performance-critical systems.
Possess experience implementing algorithms for specialized hardware such as FPGAs, DSPs, GPUs, and AI accelerators using libraries such as CUDA.
Have experience implementing operators commonly used in ML workloads, including GEMMs, convolutions, BLAS, and SIMD operators for operations like softmax, layer normalization, and pooling.
Have experience developing for embedded SIMD vector processors such as Tensilica.
Be a self-motivated team player with a strong sense of ownership and leadership in cross-functional settings.
Nice to have
Prior startup, small team, or incubation experience that demonstrates comfort in fast-paced, resource-constrained environments.
Experience with ML frameworks such as TensorFlow and/or PyTorch to understand model behavior and integration requirements.
Experience working with ML compilers and algorithms, such as MLIR, LLVM, TVM, Glow, or similar infrastructure.
Experience with deep learning frameworks for computer vision, natural language processing, or recommendation systems to evaluate model-specific kernel needs.
Work experience at a cloud provider or AI compute or AI subsystem company to understand large-scale deployment and performance constraints.
Practical notes
Hybrid schedule with expectation to work onsite at our Bangalore, India, offices 3-5 days per week.
Equal Opportunity Employment Policy
d-Matrix is proud to be an equal opportunity workplace and affirmative action employer. We're committed to fostering an inclusive environment where everyone feels welcomed and empowered to do their best work. We hire the best talent for our teams, regardless of race, religion, color, age, disability, sex, gender identity, sexual orientation, ancestry, genetic information, marital status, national origin, political affiliation, or veteran status. Our focus is on hiring teammates with humble expertise, kindness, dedication and a willingness to embrace challenges and learn together every day.
d-Matrix does not accept resumes or candidate submissions from external agencies. We appreciate the interest and effort of recruitment firms, but we kindly request that individual interested in opportunities with d-Matrix apply directly through our official channels. This approach allows us to streamline our hiring processes and maintain a consistent and fair evaluation of al applicants. Thank you for your understanding and cooperation.