Senior Staff Software Engineer
Job description
About the role
Senior Staff Software Engineer
Kernels
d-Matrix seeks a Senior Staff Software Engineer to advance generative AI capabilities. This role operates at the intersection of software and hardware innovation. Our culture values respect and direct communication. Collaboration among diverse perspectives drives our solutions.
The position is located in our Belgrade, Serbia, office. This is a hybrid role requiring 3-5 days onsite per week.
Responsibilities
You will join the team responsible for productizing the software stack for our AI compute engine. Your work will involve designing low-level kernels that extract performance from novel AI accelerator architectures. You will guide compiler teams in constructing resilient infrastructure for mapping graph operations to silicon. A core part of this role involves analyzing computational graphs from frameworks and aligning them with underlying hardware constraints. Implementation focuses on operators for specialized silicon, including GPUs and AI accelerators. You will ship software deliverables within strict development windows. Coordination with hardware experts in mixed signal, DSP, and CPU is essential. You will refine execution paths for neural network workloads. Partnering with systems experts validates behavior on embedded processors like Tensilica.
Qualifications
The minimum requirements include a Master of Science in computer engineering, math, or physics. Candidates must bring 10+ years of industry experience to this role. A PhD in computer engineering, math, or physics is an alternative, requiring 1+ years of industry experience.
A strong grasp of computer architecture, data structures, system software, and machine learning fundamentals is mandatory. Proficiency in C/C++ and Python development within Linux environments is required. You must be adept at using standard development tools. Implementation skills in high-level languages like C++ and Python are essential for performance-critical contexts.
Experience implementing kernels for specialized hardware is required. This includes work with libraries such as CUDA for FPGAs, DSPs, and GPUs. You will build operators common in ML workloads. These include GEMMs, Convolutions, and BLAS. Proficiency in SIMD functions for operations like softmax, layer normalization, and pooling is expected. Development for embedded SIMD vector processors such as Tensilica is a key part of the role.
You must be a self-motivated team player. Ownership and leadership are expected, balanced with humility.
Preferred Qualifications
Prior startup, small team, or incubation experience is valued. Hands-on work with ML frameworks such as TensorFlow and PyTorch in production is preferred. Experience with ML compilers and algorithms, such as MLIR and LLVM, is advantageous. Deep learning framework expertise for computer vision, language, or recommendation models is a plus. Work history at cloud providers or AI compute companies informs our infrastructure decisions.
Equal Opportunity
d-Matrix is an equal opportunity workplace and affirmative action employer. We are committed to fostering an inclusive environment where everyone feels welcomed and empowered. We hire the best talent for our teams, regardless of race, religion, color, age, disability, sex, gender identity, sexual orientation, ancestry, genetic information, marital status, national origin, political affiliation, or veteran status. We focus on hiring teammates with humble expertise, kindness, dedication, and a willingness to embrace challenges.
d-Matrix does not accept resumes or candidate submissions from external agencies. We appreciate the interest of recruitment firms but request that interested individuals apply directly through our official channels. This approach streamlines our hiring process and ensures a consistent evaluation for all applicants.
What you'll do
Senior Staff Software Engineer
Kernels
You will join the team responsible for productizing the software stack for our AI compute engine. Your work will involve designing low-level kernels that extract performance from novel AI accelerator architectures. You will guide compiler teams in constructing resilient infrastructure for mapping graph operations to silicon. A core part of this role involves analyzing computational graphs from frameworks and aligning them with underlying hardware constraints. Implementation focuses on operators for specialized silicon, including GPUs and AI accelerators. You will ship software deliverables within strict development windows. Coordination with hardware experts in mixed signal, DSP, and CPU is essential. You will refine execution paths for neural network workloads. Partnering with systems experts validates behavior on embedded processors like Tensilica.
Requirements
The minimum requirements include a Master of Science in computer engineering, math, or physics. Candidates must bring 10+ years of industry experience to this role. A PhD in computer engineering, math, or physics is an alternative, requiring 1+ years of industry experience.
A strong grasp of computer architecture, data structures, system software, and machine learning fundamentals is mandatory. Proficiency in C/C++ and Python development within Linux environments is required. You must be adept at using standard development tools. Implementation skills in high-level languages like C++ and Python are essential for performance-critical contexts.
Experience implementing kernels for specialized hardware is required. This includes work with libraries such as CUDA for FPGAs, DSPs, and GPUs. You will build operators common in ML workloads. These include GEMMs, Convolutions, and BLAS. Proficiency in SIMD functions for operations like softmax, layer normalization, and pooling is expected. Development for embedded SIMD vector processors such as Tensilica is a key part of the role.
You must be a self-motivated team player. Ownership and leadership are expected, balanced with humility.