AMD Buys Talus: Burning AI Models Into Silicon, 17,000 Tokens a Second
AMD has acquired Talus, a startup that builds custom chips for one specific AI model each and reaches 17,000 tokens per second. It marks a turn in AI inference toward non-programmable, model-specific silicon.
The argument · tap a timestamp to hear it
Burning a model into silicon only pays if the speedup is 100x
Talus was founded by Ljubisa Bajic, the founder of Tenstorrent, and its core idea is to burn an AI model directly into the chip — a non-programmable, custom part. The idea drew skepticism when it was announced, because models iterate quickly and a custom chip could be obsolete before long. Talus's answer is that even if a chip only runs one specific model, a 100x speed improvement is enough to cover the cost of replacing it.
— HostHardcore One's claimed 17,000 tokens a second dwarfs every rival on the list
Talus's first chip, Hardcore One, is built on TSMC 6nm, comes close to reticle size, carries 53 billion transistors, and draws at least 300 watts. On Llama 3.1 8B its claimed throughput reaches 17,000 tokens per second, far above Nvidia H200's 230, Groq's 600, and Cerebras's 2,000. That figure held up in the demonstration, which actually reached 14,200 tokens/s.
— HostSelling to AMD is a homecoming for founder Ljubisa Bajic
AMD has signed an agreement to acquire Talus. The team will report to AMD's AI head, Vamsi Boppana, and will keep operating as an independent group. Ljubisa worked at AMD before and knows Jim Keller from that time, so the acquisition can be read as a return. The price was not disclosed; the deal is expected to close by the end of the year.
— HostLisa Su's mixed-silicon thesis is exactly where a model-specific chip fits
AMD CEO Lisa Su has stressed that AI deployment is a mix of several kinds of chip. Talus's part suits the small-model calls inside agentic AI, complementing AMD's Instinct GPUs and EPYC CPUs. AMD already works with Cerebras, but the model-specific character of Talus offers a different kind of option.
— HostAn all-SRAM design like Groq's still needs a GPU beside it
Sally Ward Foxton argues that Talus resembles Groq: an all-SRAM architecture that needs multiple chips to hold a large model, but that pairs well with a GPU to handle the decode portion. She does not find the acquisition surprising, given that Talus's software-stack requirements are low and the technology is distinctive.
— Sally Ward FoxtonNvidia's Groq deal reset the price of every chip startup
Sally guesses the acquisition price could be in the billions of dollars, because Nvidia's $2 billion purchase of Groq pushed up valuations across the whole class of chip startups. Talus had previously raised only tens of millions of dollars, which makes a multi-billion-dollar valuation a reasonable read.
— Sally Ward FoxtonIn their own words · checked verbatim
The point about Talus, when they came out of stealth, was to say, "Hey, we've got an idea which is going to break the internet, break how fast you can run AI tokens. And it's going to be this concept of baking the AI model into the chip."
Host2:05
It says, "Chat Jimmy, how can I help you today?" And let's ask it a question. Say, "What is in the bubbles in pop?" Generated in 33 milliseconds, the answer came at a speed of 14,200 tokens per second.
Host6:11
But, is the speed up from 200 tokens per second to 14,200 in this demo, or based on their numbers, 17,000 tokens per second, worth the effort for a company that has an established workflow, who is going to deploy to millions of users? Yeah. You would buy a new infrastructure set every 6 months if your speed up goes from 200 to 17,000.
Host7:11
Thing is, if a tool chain is fixed and it knows what models are being called, this is where Talus can really help. This is where on their first generation chip, if they're only doing 17,000 tokens per second, who knows what they're going to do with second generation chips, third generation chips.
Host10:16
I mean, there's no doubt in my mind that it works. They also, because they only run one model at a time per chip, they don't really need much of a software stack, which is the other sticking point for startups usually in this kind of scenario.
Sally Ward Foxton13:18
Figures
| Transistor count of the Talus Hardcore One | 53 billion | 5:09 |
| Tokens per second Talus claims | 17,000 | 5:09 |
| Tokens per second actually reached in the demonstration | 14,200 | 6:11 |
| Tokens per second on Nvidia H200 | about 230 | 5:09 |
| Tokens per second on Groq | 600 | 5:09 |
| Tokens per second on Cerebras | 2,000 | 5:09 |
| Power draw of the Talus chip | at least 300 watts | 5:09 |
| Funding raised by Talus | tens of millions of dollars | 14:18 |
Glossary
- SRAM
- High-speed memory used for on-chip cache: fast, but small in capacity.
- reticle size
- The maximum mask size in lithography; a chip approaching it has a very large die area.
- agentic AI
- AI systems that decide and act on their own, usually across multiple model calls.
How to listen
AI chip founders, investors, data-center architects, and technical decision-makers working on inference performance and cost optimization.
The first three minutes of background on AI accelerators can be skipped; go straight to the introduction of Talus.