Optimization Engineer
Job description
About the role
The owns the full lifecycle of performance-critical software targeting our novel dataflow processor architecture. You will analyze, optimize, and benchmark complex kernel code to extract maximum throughput and efficiency from the Fabric. This role requires you to own entire application pipelines, transforming abstract requirements into highly tuned, production-grade implementations. You will collaborate tightly with the compiler team to provide feedback on code generation, ensuring our toolchain delivers peak performance for target workloads. A core part of this position involves pushing the performance boundaries of compute on our platform, validating results against real-world use cases. You will also work directly with embedded teams to integrate and benchmark third-party device interactions. Ultimately, you will build the optimized libraries and applications that define the efficiency advantage of our next-generation hardware.
Key facts
What you'll do
Architect and implement performance-critical kernel code specifically for Efficient's programmable processor to solve complex computational problems.
Analyze and maximize software performance on a unique dataflow architecture, identifying bottlenecks and designing low-level optimizations.
Write, debug, and maintain sophisticated low-level systems-level code in C and C++ for high-efficiency execution.
Utilize AI tools strategically to generate, optimize, and debug code, improving development speed and kernel performance.
Benchmark and optimize libraries and applications, ensuring they meet stringent performance and energy-efficiency targets.
Collaborate with compiler engineers to provide actionable feedback on code generation, debuggability, and optimization opportunities.
Work closely with embedded software teams to integrate and validate functionality on third-party devices and platforms.
Drive the definition and execution of performance benchmarks that validate hardware capabilities and software improvements.
Document complex system interactions and optimization strategies clearly for both technical and non-technical stakeholders.
Influence the product direction by contributing insights from field performance data back to hardware and software design teams.
Explore and adapt to parallel programming models to accelerate workloads on the heterogeneous compute fabric.
Investigate low-level programming interfaces such as PTX, LLVM IR, and MLIR to unlock deeper performance optimizations.
Apply domain expertise in areas like signal processing, audio processing, image processing, or machine learning to tailor solutions.
Continuously refine performance profiles and design benchmarks that accurately reflect real-world customer scenarios.
Requirements
Possess hands-on software development experience working closely with hardware, including exposure to at least one RISC, DSP, or GPU platform.
Demonstrate a passion for analyzing and maximizing software performance on a unique dataflow architecture, with a focus on energy efficiency.
Cultivate a collaborative spirit, with the ability to work with and influence multiple engineering teams across disciplines.
Show proven ability to write, debug, and maintain low-level systems-level code written in C/C++ in demanding performance environments.
Actively leverage AI tools in your workflow to generate, optimize, and debug code, integrating them into your development process.
Exhibit excellent written, verbal, analytical, and technical communication skills, with the ability to clearly document complex systems and trade-offs.
Hold a minimum Bachelor's degree in Computer Science, Computer Engineering, Electrical Engineering, or a related technical field.
Maintain a strong work ethic and the ability to manage multiple priorities in a fast-paced, deadline-driven environment.
Demonstrate meticulous attention to detail and a methodical approach to debugging and performance validation.
Bring intellectual curiosity and a commitment to learning new low-level programming techniques and hardware concepts.
Nice to have
Acquire some experience with CUDA, HIP, and/or other parallel programming models to broaden the scope of optimization efforts.
Develop some exposure to low-level programming interfaces, including PTX, LLVM IR, and MLIR, for advanced code transformation tasks.
Gain domain expertise in one or more of the following areas: Linear Algebra, Machine Learning, Image Processing, Video Processing, Signal Processing, Audio Processing, Software-Defined Radio, realtime programming, or Robotics.
Build some background in performance profiling, benchmark design, or comparative hardware analysis to inform optimization strategies.
Practical notes
This role operates on a full-time schedule. Travel may be required for team meetings, customer visits, or technical conferences. Employment eligibility to work in the United States is required for this position. No visa sponsorship is currently available for this role. The compensation range specified generally applies to positions located in the United States.