Performance Research Engineer
Job description
About the role
The owns the end to end performance validation of the company's energy-efficient processor across the entire software stack. This role owns the design and execution of performance research experiments that directly inform hardware microarchitecture decisions and compiler optimization strategies. The hire will discover novel optimization techniques for the Fabric dataflow architecture and translate academic ideas into production quality contributions. They will design and maintain AI assisted tools that automate performance analysis and guide optimization decisions for future releases. Collaboration with the architecture team is central, as the role requires proposing and evaluating design changes that improve efficiency and throughput. The engineer will also partner with the world class compiler team to assess code generation quality and influence language extensions and new intrinsics. Finally, this position owns the integration of all new techniques, tools, and language extensions into existing performance libraries and frameworks to ensure real world impact.
Key facts
What you'll do
Design and execute performance research experiments to evaluate new optimization techniques for the Fabric dataflow architecture.
Develop and maintain AI assisted tools that automate performance analysis and guide optimization decisions across the software stack.
Collaborate with the architecture team to propose, model, and evaluate design changes that improve energy efficiency and compute throughput.
Work with the compiler team to evaluate code generation quality, suggest new intrinsics, and influence the instruction set architecture.
Experiment with new language extensions and programming models, including CUDA, HIP, and other parallel frameworks, to unlock maximum efficiency.
Drive the integration of novel research findings into production performance libraries and real world application workloads.
Model and simulate hardware behavior using HW simulation environments to validate performance hypotheses before silicon.
Partner with cross functional engineering teams to align on performance goals, lead discussions, and drive consensus on optimization strategies.
Apply expertise in linear algebra, machine learning, image processing, video processing, signal processing, audio processing, SDR, realtime programming, or robotics to solve domain specific challenges.
Conduct performance profiling, benchmark design, and comparative hardware analysis to quantify improvements and regressions.
Write, debug, and maintain low level C and C++ systems code, designing clean interfaces and modular components for scalability.
Actively leverage AI tools to generate, optimize, and debug code for performance critical paths across diverse workloads.
Familiarize with low level programming interfaces such as PTX, LLVM IR, and MLIR to enable advanced optimization and code generation.
Continuously document complex systems, lead technical discussions, and communicate findings to both technical and non technical stakeholders.
Requirements
Hands on software development experience working closely with hardware, including exposure to at least two RISC, DSP or GPU platforms.
A passion for understanding and addressing performance issues that are unique to the Fabric dataflow architecture and its workload characteristics.
Experience with framework and library design, particularly within resource constrained and realtime environments where efficiency is critical.
Experience with CUDA, HIP and or other parallel programming models to enable porting and optimization of existing workloads.
The ability to work independently and lead initiatives while owning complex performance problems from start to finish.
A collaborative spirit with the ability to work effectively and influence multiple engineering teams across hardware, software, and research.
Demonstrated ability to write, debug, and maintain low level C and C++ systems level code, as well as design clean interfaces and modular code structures.
Actively uses AI tools to generate, optimize, and debug code, integrating these workflows into the development process.
Familiarity with low level programming interfaces such as PTX, LLVM IR, and MLIR to support advanced optimization efforts.
Experience working with HW simulation environments to model performance and validate architectural changes before implementation.
Domain expertise in three or more of the following areas: linear algebra, machine learning, image processing, video processing, signal processing, audio processing, SDR, realtime programming, or robotics.
Background in performance profiling, benchmark design, or comparative hardware analysis to objectively measure and compare system behavior.
Excellent written, verbal, analytical and technical communication skills, with the ability to clearly document complex systems and lead discussions across diverse teams.
Minimum Master's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field, with a PhD preferred.
Equivalent work experience may be considered in lieu of a degree for candidates with exceptional portfolios.
Nice to have
Some experience working on compiler development to contribute to optimization passes and code generation improvements.
Some experience working with high performing hardware architecture teams to align on design constraints and performance targets.
Practical notes
This role is a full time position based in either San Jose, California or Pittsburgh, Pennsylvania.
Applicants must be eligible to work in the country where the position is located without sponsorship at this time.
No specific hours, travel, visa, or application deadline details are provided in the source material.