Senior Software Engineer
Job description
Senior Software Engineer at D Matrix.
About the role
You will own the design and implementation of software kernels that power the next generation of AI compute hardware inside the d-Matrix ecosystem. This role places you at the intersection of hardware and software where your code directly dictates silicon behavior and performance characteristics. You will partner daily with compiler experts and hardware architects to translate high level computational graphs into finely tuned executable kernels. The position demands a high level of ownership where you diagnose intricate issues and deliver robust software under aggressive development timelines. You will leverage your deep understanding of diverse hardware architectures to map complex algorithms efficiently across heterogeneous compute units. Your work will influence the core infrastructure that drives inference workloads for AI applications across varied domains. You are expected to act as a technical leader within the team, proposing improvements and mentoring peers on best practices for kernel development. Together, your contributions will help realize the full potential of d-Matrix technology by ensuring software extracts maximum value from the underlying silicon.
Key facts
What you'll do
Architect and develop high performance software kernels specifically tailored for AI inference workloads on custom silicon.
Analyze computational graphs produced by machine learning frameworks and devise strategies to map them optimally onto the target hardware architecture.
Implement foundational operators such as GEMMs, convolutions, and SIMD functions for activations including softmax, layer normalization, and pooling with strict attention to cycle accuracy.
Build and maintain low level software components in C and C++ while integrating seamlessly with Python based orchestration and control systems.
Collaborate closely with hardware engineers, DSP specialists, and mixed signal teams to validate functionality and drive performance optimizations across the full stack.
Utilize development tools and workflows in Linux environments to compile, debug, and profile kernel code throughout the entire product lifecycle.
Adapt algorithms for specialized processors such as embedded SIMD vector units, ensuring efficient use of available compute and memory bandwidth.
Work within compressed development schedules to define, implement, and scale software deliverables that meet stringent quality and reliability standards.
Contribute to the evolution of compiler infrastructure alongside experts in ML compilers, potentially interacting with frameworks such as MLIR, LLVM, or TVM.
Serve as a hands on contributor and technical owner, balancing implementation work with design reviews, code quality, and long term maintainability.
Requirements
Candidates must hold a Master of Science degree in computer engineering, mathematics, physics, or a closely related discipline with a minimum of five years of industry experience in a relevant role.
Alternatively, candidates may hold a PhD in computer engineering, mathematics, physics, or a closely related discipline with at least one year of industry experience.
A strong foundational grasp of computer architecture is mandatory, including topics such as pipelining, caching, memory hierarchies, and instruction level optimization.
Deep knowledge of system software is required, encompassing operating systems concepts, concurrency, synchronization primitives, and low level programming techniques.
Proficiency in both C and C++ is non negotiable, along with demonstrated competence in developing and debugging software in Linux based environments.
Experience implementing algorithms using high level languages like C and C++ must be substantial, with a portfolio of complex projects that highlight performance sensitive coding.
Hands on experience implementing algorithms for specialized hardware such as field programmable gate arrays, digital signal processors, graphics processing units, and AI accelerators is essential.
Familiarity with graphics or compute libraries such as CUDA or similar parallel programming frameworks is a strict prerequisite for this position.
Demonstrated experience implementing core machine learning operators including GEMMs, convolutions, basic linear algebra subprograms, and SIMD friendly activations is required.
Background developing for embedded SIMD vector processors, such as those from Tensilica or comparable architectures, must be evidenced through concrete project history.
You must be a self motivated team player who exhibits strong ownership mentality and the ability to lead technical discussions without needing formal authority.
Nice to have
Prior startup, small team, or incubation experience that involves wearing multiple hats and moving quickly may be advantageous in this fast paced environment.
Hands on experience with machine learning frameworks such as TensorFlow and/or PyTorch is preferred for understanding model workflows and integration challenges.
Exposure to ML compilers and associated toolchains, including but not limited to MLIR, LLVM, TVM, or Glow, will deepen your ability to contribute effectively to compiler related tasks.
Experience working with deep learning models for computer vision, natural language processing, or recommendation systems provides context for the problems you will help solve.
Work history at a major cloud provider or a company focused on AI compute or subsystems may offer insight into scale and reliability expectations.
Practical notes
Hybrid schedule with mandatory onsite presence at our Belgrade, Serbia, office for 3 to 5 days per week.
This is a full time position aligned with standard employment regulations in Serbia.
No compensation details are specified in the source material.
No travel requirements, visa sponsorship details, application deadlines, or relocation information are provided in the source material.