AI Infra Engineer
Black Sesame Technologies IncUSA3d ago
AIEngineeringremotecurated-jd
Job description
AI Infra Engineer at Black Sesame Technologies Inc.
About the role
We are seeking a driven AI Infra Engineer to join our NPU Hardware Architecture Team. This position focuses on the integration of AI model development with silicon architecture, collaborating with both hardware and algorithm teams to optimize and deploy advanced models on our proprietary NPU platform. This role is essential for engineers interested in enhancing the synergy between model design and hardware execution.
Key facts
What you'll do
- Collaborate with algorithm teams to gain insights into model architectures, operator patterns, and deployment needs, guiding them toward designs that are compatible with NPU.
- Evaluate AI models from a hardware perspective, pinpointing issues related to compute, memory access, data movement, and parallelism.
- Lead initiatives that combine hardware and algorithm design to enhance model efficiency, performance, and energy consumption on our NPU.
- Establish and advocate for best practices in model design that align with NPU capabilities, including operator selection and memory-efficient structures.
- Work alongside hardware architects, compiler engineers, and algorithm researchers to facilitate comprehensive optimization across the technology stack.
- Assess new AI models and workload trends to identify enhancements for future NPU capabilities through informed algorithm guidance.
- Act as a liaison between algorithm innovation and hardware implementation, ensuring advanced models can be effectively deployed in production environments.
Requirements
- Master's degree or higher in Computer Science, Electrical Engineering, Computer Engineering, Applied Mathematics, or a related discipline.
- Minimum of 3 years of industry experience, ideally in AI accelerators, NPU/GPU architecture, deep learning systems, or hardware/software co-design.
- Strong foundation in deep learning principles and familiarity with modern model architectures such as CNNs and Transformers.
- Comprehensive understanding of AI accelerator architecture, including compute engines, memory hierarchy, and parallel execution.
- Proven track record in analyzing model behavior and identifying performance bottlenecks related to architecture.
- Experience with AI frameworks and deployment tools like PyTorch and ONNX, along with profiling and optimization tools.
- Strong analytical skills with the ability to navigate across algorithm, software, and hardware domains.
- Excellent communication skills and the ability to work collaboratively across various engineering teams.
Nice to have
- Experience with in-house NPU/ASIC development or AI compiler stacks.
- Knowledge of model optimization methods such as quantization, operator fusion, and graph optimization.
- Practical experience in optimizing workloads in fields like computer vision or large language models.
- Familiarity with workload characterization and performance analysis techniques.
- Proven ability to influence model design based on hardware performance considerations.
What We Value
- A holistic understanding of how model behavior relates to architectural design.
- A strong desire to advance both algorithm efficiency and hardware capabilities.
- A pragmatic engineering mindset that prioritizes efficient and competitive model performance on our platform.
- The ability to thrive in a collaborative environment where architecture, software, and algorithms develop in tandem.
Practical notes
This position is open to candidates requiring visa sponsorship. Benefits include competitive salary, health insurance, and opportunities for professional development. Applications will be accepted until the position is filled.