The world is too loud. Read what matters.

TechTechPotato

AMD Buys Talus: Burning AI Models Into Silicon, 17,000 Tokens a Second

AMD has acquired Talus, a startup that builds custom chips for one specific AI model each and reaches 17,000 tokens per second. It marks a turn in AI inference toward non-programmable, model-specific silicon.

AI chipsAMDTalusModel customizationInference acceleration
The episode identifies a new direction in AI silicon: chips built for a single model. For readers following AI infrastructure and chip investment, it is the key material for understanding where inference cost and performance inflect next.

The argument · tap a timestamp to hear it

1:05

Burning a model into silicon only pays if the speedup is 100x

Talus was founded by Ljubisa Bajic, the founder of Tenstorrent, and its core idea is to burn an AI model directly into the chip — a non-programmable, custom part. The idea drew skepticism when it was announced, because models iterate quickly and a custom chip could be obsolete before long. Talus's answer is that even if a chip only runs one specific model, a 100x speed improvement is enough to cover the cost of replacing it.

— Host
5:09

Hardcore One's claimed 17,000 tokens a second dwarfs every rival on the list

Talus's first chip, Hardcore One, is built on TSMC 6nm, comes close to reticle size, carries 53 billion transistors, and draws at least 300 watts. On Llama 3.1 8B its claimed throughput reaches 17,000 tokens per second, far above Nvidia H200's 230, Groq's 600, and Cerebras's 2,000. That figure held up in the demonstration, which actually reached 14,200 tokens/s.

— Host
8:13

Selling to AMD is a homecoming for founder Ljubisa Bajic

AMD has signed an agreement to acquire Talus. The team will report to AMD's AI head, Vamsi Boppana, and will keep operating as an independent group. Ljubisa worked at AMD before and knows Jim Keller from that time, so the acquisition can be read as a return. The price was not disclosed; the deal is expected to close by the end of the year.

— Host
10:16

Lisa Su's mixed-silicon thesis is exactly where a model-specific chip fits

AMD CEO Lisa Su has stressed that AI deployment is a mix of several kinds of chip. Talus's part suits the small-model calls inside agentic AI, complementing AMD's Instinct GPUs and EPYC CPUs. AMD already works with Cerebras, but the model-specific character of Talus offers a different kind of option.

— Host
12:17

An all-SRAM design like Groq's still needs a GPU beside it

Sally Ward Foxton argues that Talus resembles Groq: an all-SRAM architecture that needs multiple chips to hold a large model, but that pairs well with a GPU to handle the decode portion. She does not find the acquisition surprising, given that Talus's software-stack requirements are low and the technology is distinctive.

— Sally Ward Foxton
14:18

Nvidia's Groq deal reset the price of every chip startup

Sally guesses the acquisition price could be in the billions of dollars, because Nvidia's $2 billion purchase of Groq pushed up valuations across the whole class of chip startups. Talus had previously raised only tens of millions of dollars, which makes a multi-billion-dollar valuation a reasonable read.

— Sally Ward Foxton

In their own words · checked verbatim

The point about Talus, when they came out of stealth, was to say, "Hey, we've got an idea which is going to break the internet, break how fast you can run AI tokens. And it's going to be this concept of baking the AI model into the chip."

Host2:05

It says, "Chat Jimmy, how can I help you today?" And let's ask it a question. Say, "What is in the bubbles in pop?" Generated in 33 milliseconds, the answer came at a speed of 14,200 tokens per second.

Host6:11

But, is the speed up from 200 tokens per second to 14,200 in this demo, or based on their numbers, 17,000 tokens per second, worth the effort for a company that has an established workflow, who is going to deploy to millions of users? Yeah. You would buy a new infrastructure set every 6 months if your speed up goes from 200 to 17,000.

Host7:11

Thing is, if a tool chain is fixed and it knows what models are being called, this is where Talus can really help. This is where on their first generation chip, if they're only doing 17,000 tokens per second, who knows what they're going to do with second generation chips, third generation chips.

Host10:16

I mean, there's no doubt in my mind that it works. They also, because they only run one model at a time per chip, they don't really need much of a software stack, which is the other sticking point for startups usually in this kind of scenario.

Sally Ward Foxton13:18

Figures

Transistor count of the Talus Hardcore One53 billion5:09
Tokens per second Talus claims17,0005:09
Tokens per second actually reached in the demonstration14,2006:11
Tokens per second on Nvidia H200about 2305:09
Tokens per second on Groq6005:09
Tokens per second on Cerebras2,0005:09
Power draw of the Talus chipat least 300 watts5:09
Funding raised by Talustens of millions of dollars14:18

Glossary

SRAM
High-speed memory used for on-chip cache: fast, but small in capacity.
reticle size
The maximum mask size in lithography; a chip approaching it has a very large die area.
agentic AI
AI systems that decide and act on their own, usually across multiple model calls.

How to listen

Who it's for

AI chip founders, investors, data-center architects, and technical decision-makers working on inference performance and cost optimization.

Skip

The first three minutes of background on AI accelerators can be skipped; go straight to the introduction of Talus.