AMD收购Talus:押注模型定制芯片的未来
AMD收购了Talus,一家为特定AI模型定制芯片、实现每秒17,000 tokens的初创公司。这标志着AI推理走向非可编程、模型专用芯片的方向。
核心论点 · 点时间戳可跳到原声
Talus的颠覆性理念
Talus由Tenstorrent创始人Ljubisa Bajic创立,其核心概念是将AI模型直接烧录到芯片中,实现非可编程的定制化芯片。这一理念在发布时遭到质疑,因为模型迭代快,定制芯片可能很快过时。但Talus认为,即使芯片只适用于特定模型,其100倍的速度提升也足以弥补更新换代的成本。
— 主持人Hardcore One芯片性能
Talus的首款芯片Hardcore One基于TSMC 6纳米工艺,接近reticle尺寸,拥有530亿晶体管,功耗至少300瓦。在Llama 3.1 8B模型上,其宣称的每秒token数达到17,000,远超Nvidia H200的230、Groq的600和Cerebras的2,000。这一数字在演示中得到了验证,实际达到14,200 tokens/s。
— 主持人AMD收购细节
AMD已签署协议收购Talus,团队将向AMD AI负责人Vamsi Boppana汇报,并保持独立团队运作。Ljubisa曾在AMD工作,与Jim Keller相识,此次收购可视为回归。交易金额未披露,预计年底完成。
— 主持人AMD的AI战略契合
AMD CEO Lisa Su强调AI部署是多种芯片的混合。Talus的芯片适合agentic AI中的小型模型调用,与AMD的Instinct GPU、EPYC CPU等形成互补。AMD已与Cerebras合作,但Talus的模型专用特性提供了另一种选择。
— 主持人行业分析师观点
Sally Ward Foxton认为,Talus与Groq类似,采用全SRAM架构,需要多芯片支持大模型,但适合与GPU配合处理解码部分。她认为收购并不意外,因为Talus的软件栈需求低,且技术独特。
— Sally Ward Foxton收购价格猜测
Sally猜测收购价格可能在数十亿美元级别,因为Nvidia收购Groq的20亿美元交易推高了整个芯片初创公司的估值。Talus此前融资仅数千万美元,因此数十亿美元的估值是合理的。
— Sally Ward Foxton原话 · 已逐字校验
The point about Talus, when they came out of stealth, was to say, "Hey, we've got an idea which is going to break the internet, break how fast you can run AI tokens. And it's going to be this concept of baking the AI model into the chip."
Talus在走出隐身模式时的观点是:我们有一个想法,将打破互联网,打破AI token运行速度的极限。这个概念就是将AI模型烧录到芯片中。
主持人2:05
It says, "Chat Jimmy, how can I help you today?" And let's ask it a question. Say, "What is in the bubbles in pop?" Generated in 33 milliseconds, the answer came at a speed of 14,200 tokens per second.
它说:“Chat Jimmy,今天我能帮你什么?”让我们问它一个问题:“汽水里的气泡是什么?”生成耗时33毫秒,回答速度达到每秒14,200个token。
主持人6:11
But, is the speed up from 200 tokens per second to 14,200 in this demo, or based on their numbers, 17,000 tokens per second, worth the effort for a company that has an established workflow, who is going to deploy to millions of users? Yeah. You would buy a new infrastructure set every 6 months if your speed up goes from 200 to 17,000.
但是,从每秒200个token提升到14,200(或他们声称的17,000),对于一个已有工作流程、将部署给数百万用户的公司来说,值得吗?是的,如果速度提升从200到17,000,你每6个月就会购买一套新的基础设施。
主持人7:11
Thing is, if a tool chain is fixed and it knows what models are being called, this is where Talus can really help. This is where on their first generation chip, if they're only doing 17,000 tokens per second, who knows what they're going to do with second generation chips, third generation chips.
问题是,如果工具链是固定的,并且知道调用哪些模型,这正是Talus能真正帮助的地方。在第一代芯片上,如果他们每秒只处理17,000个token,谁知道第二代、第三代芯片会做什么。
主持人10:16
I mean, there's no doubt in my mind that it works. They also, because they only run one model at a time per chip, they don't really need much of a software stack, which is the other sticking point for startups usually in this kind of scenario.
我的意思是,我毫不怀疑它能工作。而且,因为每个芯片一次只运行一个模型,他们不需要太多的软件栈,这通常是初创公司在这种情况下的另一个痛点。
Sally Ward Foxton13:18
数字与实体
| Talus Hardcore One芯片的晶体管数 | 530亿 | 5:09 |
| Talus宣称的每秒token数 | 17,000 | 5:09 |
| 演示中实际达到的每秒token数 | 14,200 | 6:11 |
| Nvidia H200的每秒token数 | 约230 | 5:09 |
| Groq的每秒token数 | 600 | 5:09 |
| Cerebras的每秒token数 | 2,000 | 5:09 |
| Talus芯片功耗 | 至少300瓦 | 5:09 |
| Talus融资额 | 数千万美元 | 14:18 |
术语
- SRAM静态随机存取存储器
- 一种高速存储器,用于芯片缓存,速度快但容量小。
- reticle size掩模版尺寸
- 光刻中掩模版的最大尺寸,接近此尺寸的芯片面积很大。
- agentic AI智能体AI
- 能够自主决策和执行任务的AI系统,通常涉及多个模型调用。
收听指南
AI芯片创业者、投资者、数据中心架构师,以及关注推理性能与成本优化的技术决策者。
前3分钟关于AI加速器背景的介绍可跳过,直接进入Talus介绍。