The world is too loud. Read what matters.

张小珺·商业访谈录

The model is an oil refinery: the winner of the first phase is not the winner of the last

In business history, the companies that run out ahead in the first wave are almost never the winners of the next phase. Model companies are refineries that extract base oil; the value is in the chemical plants and the car companies.

Paper deep-readModel paradigmsInfraMultimodalAI history

The video won't play here. Listen to the audio instead:

A four-hour walkthrough of papers, dense in information but slow in pace — good for jumping around like a reference book, not for listening straight through on a commute.

The argument · tap a timestamp to hear it

19:35

GPU's victory was not foresight, it was persistence

In 1999 the GeForce 256 pulled 3D computation out of the CPU, and in 2001 programmable shaders were introduced, which is what made the GPU programmable. In 2004 Brook for GPU abstracted away the underlying graphics hardware and became the predecessor of CUDA. The key judgment: Jensen Huang did not see AI coming — Nvidia's 2012 vision for the GPU basically did not include the things of modern AI, and ATI also built something like CUDA at the time, it just didn't stick with it and exited the market. In CUDA's early days, graphics cards saw no gain in gaming performance and BOM cost nearly doubled, and Jensen Huang persisted for roughly ten years before scientific computing paid off.

— Xie Jinchi
52:35

Before ResNet, bigger models actually got worse

In 2015 people found that stacking a model past 100 layers didn't make it better, it made it worse — this is called model degradation, meaning that what everyone today knows as scaling actually did not work back then. ResNet introduced residuals, changing the learning target from "turn X directly into Y" to "how much to add to or subtract from X to get Y"; the latter is easier to learn, and it also alleviates vanishing and exploding gradients, essentially dissolving the degradation problem. Only after this could networks of hundreds or even thousands of layers be trained; before that they were usually limited to a few dozen layers. This paper has close to 300,000 citations, higher than Transformer.

— Xie Jinchi
1:12:48

AlphaGo Zero inspired the thinking mode

AlphaGo Zero used only reinforcement learning and gave the model no human prior knowledge at all; the model knew only the Go board and the rules, not game records, liberties or eyes. It used fewer cards than the version that beat Lee Sedol, and surpassed it after training for only 36 hours. The more profound impact is that it ran 1,600 MCTS searches on every move, which inspired OpenAI to do scaling at test time — that is, O1's thinking mode. A counterintuitive point: if a model doesn't do test-time computing, its capabilities are actually not as strong as we imagine.

— Xie Jinchi
1:22:51

COT shifted the center of scaling from pretraining to post-training

After 2017 everyone scaled models desperately, but the gains in arithmetic, common sense and symbolic reasoning were not obvious — not every domain has a scaling law. At the same time, using SFT to solve reasoning problems meant building reasoning datasets at very high cost, requiring many PhDs. COT found that simply showing the model the intermediate steps of a derivation greatly improves performance on reasoning tasks — this is the origin of "please think step by step." It made the whole industry realize that many capabilities are already contained in the model, just not activated, and the center of scaling shifted from pretraining to post-training, which also influenced the birth of later thinking models.

— Xie Jinchi
1:56:19

OpenAI and DeepMind disagreed about scaling

Two scaling law papers found a log-linear relationship between language model training and loss, so small-model experiments could predict the effect after scaling up and avoid spending three months opening a blind box. But the two reached different conclusions: OpenAI recommended training as large a model as possible under a given compute budget, even if that means stopping training early; DeepMind held that this leaves many models undertrained and argued for scaling model size and training token count in equal proportion — double the parameters, double the data. DeepMind also proved that training a smaller model on more data is better than training a larger model on less data, because small models have lower inference cost, which is why many small models today are trained on excessive data.

— Xie Jinchi
2:13:24

Only three companies in the world have ten-thousand-card training experience

Between 2020 and 2022, only about three companies in the world had experience connecting ten thousand cards to train: OpenAI, Google and DeepSeek. When DeepSeek built Firefly No. 2 it was already a very large cluster, more than they could use, and they even let university teachers and students apply to use it on academic grounds. Connecting ten thousand cards to act as one card is a challenge in training efficiency and training stability: more cards is not necessarily better, and if you get it wrong ten thousand cards can be less efficient than a thousand; GPUs fail at the ten-thousand-card scale, from physical damage to bit flips, requiring a deeply observable system to monitor, diagnose, attribute, automatically locate faults and automatically recover.

— Xie Jinchi
2:18:28

DeepSeek designed right up against the H800's bandwidth limit

DeepSeek's cards are H800s, whose characteristic is relatively narrow bandwidth. With narrow bandwidth, doing tensor parallelism means the model can't be passed through — for example, if bandwidth is only 80G and the model to be passed is 20G, you can only split across 4 cards; try to split across 5 and it jams. Their parallelism design sits right against the H800's limits, achieving an almost exact balance of compute and communication, so compute doesn't have to wait for communication. In training today MFU is only 50% or even less, meaning nearly half the GPU compute is left idle; in theory, if you could reach 100%, you could buy half as many cards.

— Xie Jinchi
3:00:00

A tuned 1.3B model is more obedient than 175B

GPT-3 was powerful but not usable: it generated untrue, toxic and unhelpful content, and its instruction-following ability was poor. These problems did not improve as scaling improved — scaling up from GPT-1's 0.1B by ten thousand times did not fully solve them either. InstructGPT used reinforcement learning from human feedback, hiring more than 40 contractors to build SFT data and ranking data, amplifying the good-versus-bad signal by one or two orders of magnitude. The result: a 1.3B GPT-3 tuned this way was better at following instructions than the 175-billion-parameter GPT-3, a 100-fold reduction in parameter count. This made the industry realize that beyond simply enlarging models, optimizing post-training methods also brings very large gains.

— Xie Jinchi

In their own words · checked verbatim

Then Mai said we don't provide any explanation of the model, any explanation of why the model works. If it works, we attribute it to the grace of God.

那麦说我们不提供对模型解释 为什么模型work的任何解释 如果它work 那我们把它归为神的仁慈

Xie Jinchi1:00:42

So finding a good question, a core question, is crucial. And you'll find — you'll find — if we look back, these good questions don't seem that hard.

所以发现一个好问题 核心问题很关键 对而且你会发现 你会发现 如果我们回过头来看 这些好问题 好像没有那么的难

Xie Jinchi1:38:04

He said we've gotten 70 years of intelligence out of it, and found that general methods that exploit computation ultimately prove most effective, and by a significant margin. He said the root cause lies in a certain law, or more broadly, the continuous exponential decline in the cost per unit of computation.

他说我们从能够智能搞到70年 发现利用计算的通用方法 最终最为有效 且优势显著 他说根本原因在于某个定律 或者说更广义的计算单位成本的 持续指数级下降

Xie Jinchi1:43:08

So AI researchers try to encode knowledge into their agents. This always works in the short term and is personally satisfying; it has returns, but in the long run it falls into a plateau and even hinders progress. Breakthroughs ultimately come from the opposite path: scaling computation through searching and learning.

所以AI研究者 试图将知识编码入其智能体 这在短期内总是有效 且令人的个人满足 它有收益 但长期会陷入平台期 甚至阻碍进展 而突破性的进展 最终来自相反路径 通过搜索 searching和learning 搜索和学习 实现计算规模扩展

Xie Jinchi1:46:11

You'll find that if you don't train with labeled data, the model loses its ability to understand human body structure.

你会发现 如果你不用标签的数据来训练的话 模型会失去对人体结构的理解能力

Xie Jinchi2:09:24

Figures

AlphaGo Zero training time36 hours1:12:48
AlphaGo Zero MCTS searches per move1,6001:12:48
ResNet citationsclose to 300,00054:38
GPT-2 parameter count1.5 billion2:54:57
Number of contractors hired by InstructGPTmore than 403:02:06
Number of image-text pairs in LAION-5B5B2:07:24

Glossary

MFU / Model FLOPs Utilization
The proportion of GPU compute actually being used, which today is usually only 50% or even less.
MOE / Mixture of Experts
A very large-parameter model activates only a small subset of parameters per inference, lowering inference cost.
COT / Chain of Thought
Having the model show the intermediate steps of a derivation, which greatly improves performance on reasoning tasks.
LoRA / Low-Rank Adaptation
Fine-tuning by attaching a small matrix alongside the model, leaving the original model unchanged; it can be merged at inference time.
RLHF / Reinforcement Learning from Human Feedback
Training a reward model on human ranking data, then using reinforcement learning to adjust the model's behavior.
Optical flow
The motion shadow obtained by subtracting adjacent frames, containing the motion information in a video.

How to listen

Who it's for

Engineers and product managers who want to systematically catch up on AI papers but don't know where to start; suited to jumping around by chapter like a reference book.

Skip

The Infra and data section from 1:52:58 to 2:21:28 — the guest admits to knowing less about it, so you can fast-forward.