
Machine Learning Research Engineer
Job description
About the role
Tenstorrent is building advanced AI systems that expand what is possible in model training, inference, and distributed computing at scale. The ML Models team bridges AI research with high-performance hardware, bringing modern machine learning models to Tenstorrent's custom AI accelerators. This position spans the full stack, from training large language models to scaling inference performance in production environments. You will own the design and execution of experiments that validate novel model architectures and training strategies on our hardware platforms. You will collaborate closely with hardware engineers to ensure software workloads maximize the capabilities of Tenstorrent's compute fabrics. You will translate abstract research ideas into robust pipelines that integrate seamlessly into the company's development workflows. You will contribute to the technical leadership of the ML Models team by driving best practices in reproducibility and performance engineering. You will play a key role in defining the direction for bringing state-of-the-art AI models to market on Tenstorrent's infrastructure.
Key facts
What you'll do
- Drive research and development work centered on LLM training and inference optimization across diverse model scales.
- Train, evaluate, and fine-tune AI models on Tenstorrent hardware to establish clear performance and accuracy baselines.
- Apply performance techniques including speculative decoding, quantization, kernel fusion, flash attention, and distributed training methods to remove bottlenecks.
- Identify system bottlenecks by analyzing telemetry and profiling data, then work with cross-functional teams to resolve issues.
- Convert recent ML research findings into scalable, production-grade implementations that meet strict reliability standards.
- Build and maintain tools, libraries, and benchmarks that enable efficient experimentation and rapid iteration cycles.
- Partner with product and hardware teams to define test scenarios that validate functionality under real-world conditions.
- Document methodologies, results, and architectural decisions to ensure clarity and continuity for the broader engineering organization.
- Mentor junior engineers and researchers by providing code reviews, technical guidance, and constructive feedback.
- Stay current with the latest advances in machine learning to proactively suggest improvements to the research roadmap.
Requirements
- 4+ years of industry or academic experience in ML research and LLM development, demonstrating a track record of shipped contributions.
- Strong Python and PyTorch skills for building and training deep learning models with attention to code quality and maintainability.
- Solid grasp of ML architectures, LLM training methods, and inference optimization strategies across different deployment scenarios.
- Practical experience training large-scale machine learning models, including data loading, checkpointing, and fault tolerance.
- Proven ability to debug complex numerical issues and performance regressions in training and inference workloads.
- Experience working in a fast-paced environment where priorities shift based on hardware capabilities and research findings.
- Excellent written and verbal communication skills to articulate technical concepts to both technical and non-technical stakeholders.
- Eligibility for U.S. export control regulations as defined by the EAR, which may restrict collaboration based on nationality or location.
Nice to have
- PhD degree with a strong publication record in relevant machine learning venues.
- Published research that demonstrates innovation in model architecture, training procedures, or optimization techniques.
- Direct experience with speculative decoding, including implementation and integration with existing inference frameworks.
Skills & tools
- Python
- PyTorch
- LLM training and inference
- Quantization
- Kernel fusion
- Flash attention
- Distributed training
- Custom AI accelerators
Practical notes
- Candidates at various experience levels are welcome; leveling will be determined during interviews and may differ from the posted title.
- Employment is contingent on eligibility to access U.S. export-controlled technology under EAR regulations, which may require citizenship, permanent residency, or prior license approval from the U.S. Commerce Department depending on nationality (applies to EAR Country Groups D:1, E1, and E2).
- Equal opportunity employer.
The role requires a commitment to working in environments where hardware constraints directly influence software design decisions. You will be expected to manage multiple concurrent experiments, analyzing results to determine the most efficient path forward. The successful candidate will have a high tolerance for ambiguity and a methodical approach to problem solving. Close collaboration with the engineering team is essential, as you will often need to align technical tradeoffs with business objectives. Travel may be required to support hardware validation or to participate in industry events that showcase the company's technology. The position is based in either Boston, Massachusetts, or Toronto, Ontario, with a hybrid work model that balances in-office collaboration with remote flexibility. Security clearance considerations related to export control regulations may impact project assignments and team placement. This position is full-time and open to candidates who are authorized to work in the United States or Canada. Deadlines for submission of application materials are not explicitly published; however, early application is encouraged given the volume of interest in roles at the intersection of machine learning and systems engineering.