Staff, Reliability Engineer
Tenstorrent
- Location
- Toronto, Ontario, Canada
- Work model
- On-Site
- Level
- Staff
- Posted
- 12h ago
About this role
Tenstorrent is leading the industry on cutting-edge AI technology, revolutionizing performance expectations, ease of use, and cost efficiency. With AI redefining the computing paradigm, solutions must evolve to unify innovations in software models, compilers, platforms, networking, and semiconductors. Our diverse team of technologists have developed a high performance RISC-V CPU from scratch, and share a passion for AI and a deep desire to build the best AI platform possible. We value collaboration, curiosity, and a commitment to solving hard problems. We are growing our team and looking for contributors of all seniorities.
Join Tenstorrent as a Staff Reliability Engineer and help define the reliability strategy behind the next generation of AI computing systems. In this highly visible technical leadership role, you'll drive reliability from architecture through production, partnering across hardware, software, and manufacturing teams to build high-performance AI platforms that set the standard for uptime, durability, and quality. If you're passionate about solving complex engineering challenges and influencing products at scale, you'll have the opportunity to shape technology powering the future of AI.
This role is hybrid, based out of Toronto, Canada.
We welcome candidates at various experience levels for this role. During the interview process, candidates will be assessed for the appropriate level, and offers will align with that level, which may differ from the one in this posting.
Who You Are
• You've spent 8+ years in reliability engineering, ideally in high-performance computing, AI hardware, or data center systems.
• You're comfortable with the statistical side of the job, HALT, HASS, ALT, MTBF, Weibull analysis, and FMEA are all familiar territory.
• You can work through a technical problem in a thermal lab and then explain the risks and trade-offs clearly to leadership.
• You're good at bringing people together, mechanical, electrical, thermal, and software teams, especially when timelines are tight.
What We Need
• Someone to set the reliability strategy for our next-generation AI computing systems, across both data center and workstation products.
• A strong problem-solver who can lead root-cause investigations on failures and follow through with fixes across engineering and the supply chain.
• Someone with real experience validating advanced cooling systems, vapor chambers, heat pipes, direct-to-chip liquid cooling.
• A steady point of contact for our manufacturing partners and suppliers, making sure reliability standards hold up through the NPI process.
• Someone who can mentor other engineers and lead design reviews that raise the bar for the team.
What You Will Learn
• How to shape the reliability strategy for some of the industry's most advanced AI computing platforms built around Tenstorrent's cutting-edge silicon.
• How reliability engineering influences every stage of product development, from architecture and validation through manufacturing and field deployment.
• How to collaborate with world-class experts across silicon, thermal, mechanical, electrical, and software engineering to solve large-scale system challenges.
• How to validate emerging cooling technologies, including advanced air cooling and direct-to-chip liquid cooling, for next-generation AI infrastructure.
• How to influence the future of AI hardware by helping build highly reliable systems that power tomorrow's largest AI workloads.
Tenstorrent offers a highly competitive compensation package and benefits, and we are an equal opportunity employer.
This offer of employment is