The world is too loud. Read what matters.

科技这碗饭

Google Invented the Transformer, Then Gave the LLM Era to OpenAI

Google held the Transformer, the TPU and the most compute at once, yet the burden of search advertising and two rival AI teams fighting each other handed the most critical race — language models — to OpenAI.

GoogleDeepMindTransformerLLMsOrganizationAI Chips

The video won't play here. Listen to the audio instead:

High information density, and a detailed 15-year organizational history of Google AI — especially the back half, where two teams fought over compute and over who got to ship, details you rarely find elsewhere.

The argument · tap a timestamp to hear it

5:05

Google's underlying problem, solved once and for all

In 2000, Google's index had not updated in five months, and engineers spent days looking for the cause without finding it. Jeff Dean and Sanjay put a batch of seemingly unrelated errors side by side and found that a single bit in machine memory was flipping at random — to save money, Google's servers used cheap memory without error correction. They did not ask for more expensive machines; they added a layer of insurance in software, automatically quarantining whichever machine went bad. This is the most typical way the two of them work: not patching a failure at the surface, but going down to the bottom layer and solving that entire class of problem at once, so that no other engineer ever has to think about whether the machine itself might fail.

11:44

Building a chip specifically for neural networks

Jeff Dean used speech alone to make an estimate, put it into a two-page deck and showed it to management: if every Android user talked to their phone for three minutes a day, processing just that speech would force Google to expand its global data centers two- to threefold. His answer was to build a chip specifically for neural networks and worry about nothing else. This went against his long-standing view of chips — specialized hardware is usually a bad idea, because by the time you build it the algorithms have moved on — but neural networks were the exception: image recognition and speech dictation both rest on the same kind of matrix multiplication. To save time, the chip was built as a card that slotted into the server space originally reserved for hard drives; from project kickoff to installation in a data center took only 15 months.

27:25

The phone call Hassabis made to Hinton

Hinton ran an auction by email from a hotel room, with buyers including Google, Microsoft, Baidu, and a fourth: Hassabis's DeepMind. DeepMind had little cash on hand, so Hassabis bid $10 million in company stock, roughly a fifth of the entire company. The price quickly passed what he could afford, and after dropping out he called Hinton and said that although he had only meant to offer ten million for them, in his heart he believed their true value was $50 million. Hinton took that in. Microsoft executives privately raised their bid by $10 million and he wavered for a moment, but in the end he declined. Google ultimately paid $44 million for Hinton and two of his students.

35:31

Others look at the hood; he wanted to see the engine

Google's due diligence team flew to London, Jeff Dean among them, tasked with checking whether what DeepMind claimed was actually true. At the time DeepMind had only 50 people and no revenue, yet the price on the table was twice what Google normally paid for a company of that size; plenty of people at the negotiating table thought it was absurd, and Jeff Dean thought it was a bit expensive too. After the deck was presented he made one request: he wanted to see the source code. A demo can be dressed up; code is hard to fake. The researcher responsible for writing it panicked, because when you are doing research you write code only to get experiments running as fast as possible, and it is usually messy and hard to read. Hassabis later called this moment crossing the Rubicon — once the code was handed over, there was no way back.

49:52

The more AlphaGo succeeded, the less it could be independent

DeepMind wanted to use the Alphabet reorganization to spin out of Google, and the talks ran more than a year, with terms even committed to paper. But Pichai said on the phone that the Other Bets column was reserved for moonshots far from Google's homepage, and AI no longer belonged in that category — it was becoming Google's most core technology. Hassabis's past strength was winning a battle big enough, then trading that victory for leverage in the next stage, but this time the leverage moved in the opposite direction: the more AlphaGo succeeded, the more important AI became inside Google, the closer it sat to the core of the homepage, and the less possible it was to put it in Other Bets. In April 2021, Hassabis announced at an all-hands that the push for independence was over.

1:07:01

Throw out the entire search index — that idea was too crazy

Not long after the Transformer paper was written, one of its authors, Noam Shazeer, began thinking about its future and brought Google leadership an almost insane proposal: throw out the index Google Search had used for years, train a giant neural network on the new architecture instead, relearn and reorganize all the information Google held, and simply answer whatever users asked. Even the collaborators who had written the paper with him thought the idea was too crazy, and the proposal went nowhere. Around 2018 he and colleagues trained a chat model called Meena that could discuss philosophy, ramble about TV shows, and even make puns, and the more the two of them worked on it the more they felt it had something. Over the following years they filed three separate requests — opening it to outside researchers, plugging it into the voice assistant, doing a public demo — and all three were rejected for the same reason: it did not fit the company's AI principles.

1:27:12

Same company, yet guarding folders from each other

DeepMind had no data center of its own; Google's data center was one giant shared compute pool, with money-making teams like Search, YouTube and Google Cloud and money-burning teams like Google Brain and DeepMind drawing from the same pool, and anyone wanting to train a model had to queue for an allocation. DeepMind and Google Brain competed year-round for talent and for results, and what they competed over most was machines. At the end of 2017, just to teach a model to play chess by playing itself, DeepMind occupied 5,000 TPUs at once; the next year Google Brain used 64 in total to train BERT. After GPT-3 came out, both teams' goal became language models, and DeepMind had to rename its folders so Google Brain would not know the parameter scale it was aiming for — so 280B became "mole".

2:02:29

The rival holds one sharp knife; Google has a drawer of them

In the second half of 2025 the biggest AI business shifted to the enterprise market, and what enterprises spent the most on was not chat but code writing; half of that market sits with Anthropic, the other half with Cursor. Google's problem was not that it failed to see the coding direction, but that it did too much, scattered across departments: on the external product side alone, five departments had built six or seven coding tools, with overlapping features. Google's internal rule was that in principle you could only write code with its own AI tools, and using a third party required special approval — but there was a loophole: if you could prove the business needed it, you could apply for an exception. Several former employees said that inside Google DeepMind, every team training Gemini applied for Cloud Code, on the grounds that "the best engineers should use the best tools."

In their own words · checked verbatim

We are going to make the biggest invention in human history. He will redo every product we have today from scratch. If you insist I pick one product, of course I can — but the fact that you ask the question that way shows you still do not understand what we are doing.

我们要做出人类历史上最大的发明 他会把今天所有的产品 全部都重新做一遍 你如果非要我挑一个产品出来 当然可以 但你会这么问问题 就说明你还没有明白 我们到底在做什么

Once the code is handed over, there is no way back. The people who understood this field better than anyone in the world had gone through his hand, card by card.

代码一旦交出去 就再也没有回头之路了 全世界最懂这一行的一群人 把他的底牌看了一个遍

If the people you bring in are both powerful and able to understand the technology, they will not be content to sit on the sidelines and give you advice forever. The so-called advisor very often ends up being your competitor.

如果你请来的是既有权势 又看得懂这项技术的人 那么他们不会甘心一直坐在场边给你提意见 所谓的顾问 最后很有可能就是你的对手

The question was never whether they saw the path at the time. It is that they clearly saw it, and in the end did nothing.

问题从来不是当时他们有没有看到这条路 而是明明已经看见了 为什么最后什么都没有做

At twelve, Hassabis decided that the most worthwhile thing he could do with his life was apply his mind to science. At forty-six, he stood in front of a room full of scientists and said: we cannot only do science anymore.

12岁那年 哈萨比斯决定 他这一辈子最值得做的事情 就是把脑力用在科学上 46岁这一年 他站在一屋子的科学家面前说 我们不能只做科学了

He wanted to have again a small team of barely a dozen people, focused on doing just one thing.

他想要重新拥有一支 只有十来个人 只专注做一件事情的小团队

Figures

Processor cores in Google Brain's cat-face experiment160009:07
Google's in-house chip, from kickoff to data center15 months14:11
Price Google paid for DeepMindabout $600 million-plus36:32
Global viewers of AlphaGo vs. Lee Sedol200 million41:34
DeepMind's losses over 2018 and 2019nearly £1 billion1:28:12
TPUs used by Google Brain to train BERT641:28:12
TPUs DeepMind occupied to play chess50001:28:12
What Google paid to buy back Shazeer's teamabout $2.7 billion1:52:23

Glossary

Transformer
The neural network architecture proposed in Google's 2017 paper, replacing word-by-word recurrence with attention and enabling large-scale parallel training.
TPU
Google's in-house chip designed specifically for neural networks, installed in data centers in 2015.
Reinforcement learning
A method that lets an agent learn by trial and error through action, adjusting its strategy based on feedback; AlphaGo used it to learn Go.
Sparse model
Splitting a model into many small networks and calling only one or two of them at a time, so total parameters are large but per-inference compute is small.
Chain of thought
Having a model demonstrate its reasoning process before giving an answer, which can raise accuracy severalfold.
NeoLab
Industry shorthand for a new lab founded by people leaving big companies that puts research ahead of product.

How to listen

Who it's for

Founders and investors watching the LLM competitive landscape, especially anyone trying to understand why technical leadership is not product leadership, and how internal friction eats a first-mover advantage.

Skip

19:16 to 24:23, the section on Hassabis's childhood and the founding of DeepMind, can be fast-forwarded; it is less dense than the back half.