The world is too loud. Read what matters.

张小珺·商业访谈录

17-Year-Old AI Native: The Reason to Study Math Has Become Proving to Universities That I Can Endure Hardship

After AI arrived, his reason for studying math degraded into proving to universities that he can endure hardship: the path he could once see — university, PhD, advancing mathematics — has narrowed, and the classmates around him who want to work in aerospace are also losing their upward motivation.

AI nativesAI and educationmodel companiespretrainingteenagers
The technical increment is limited, but a 17-year-old AI native's judgment about the landscape of model companies — and about AI eroding the next generation's motivation — is more direct than most adult interviews.

The argument · tap a timestamp to hear it

11:13

Pretraining feels magical, post-training feels like shaping a body

His reason for doing pretraining isn't papers, it's feel: post-training is more like shaping a body, while pretraining is making your specific material better so it can withstand the shaping that comes later — he says this is a bit magical. The cost is real money burn: this paper cost him 30,000 yuan in total, and the largest model in the paper was only 0.05B; to scale up, he calculated it would take roughly another 6,000 yuan. That night he was extremely nervous, wondering whether to spend the money, because even if the paper was good, submitting to a top conference could still get it rejected and the money would be wasted. In the end his older brother and his parents both supported him, and only then did he take the step — and the scaling step worked.

— Su Tinghao
13:13

Three models didn't save, 1,500 burned for nothing

He describes training as an expensive game. Once he forgot that training periodically saves checkpoints, trained three models at the same time, none of them saved, and the parameters weren't tuned well either — 1,500 lost. He says he cried, was extremely, extremely sad, and kept thinking about why he hadn't looked at that part again earlier. There is nothing profound in this: the pit a high school sophomore stepped into is the most everyday pit for anyone doing pretraining — money goes into compute, and one low-level mistake zeroes it out. His mother comforted him by saying that once you pay tuition you remember the lesson, and he never made that mistake again. He also admits his training code is very messy, only he can use it, and training efficiency has big problems.

— Su Tinghao
16:16

He cut the width of attention in half

The paper's starting point is Value Residual Learning: adding the first layer's attention value to later layers as a residual. He found that the low-layer value is doing two things — being used in its own layer, and being added to later layers as a residual — so he doubled the width of the first layer's attention projection, cutting half for its own layer and the other half dedicated to the backward residual, then added RMSNorm, and extended the approach from value to key and query, which worked better. He knows this looks very similar to Kimi's attention residual, but he says they are two completely different lines of thinking: what Kimi residuals is the hidden state, what he residuals is the attention projection. He himself is clear about which one works better — Kimi's path has already been proven at 2.5T scale, his is a small-scale result.

— Su Tinghao
18:19

The GPT framework only needs high school math; the hard part is intuition

His judgment is that the current big framework of Transformer and GPT doesn't require very advanced math or probabilistic thinking — it's all matrix multiplication, and high school math is about enough to understand it; what really takes time is cultivating intuition. The example is himself: a year or so ago he watched Andrej Karpathy's roughly twenty-hour video series and could only write whatever Karpathy wrote, with no understanding at all, and then these things settled in his head for half a year to a year before he suddenly understood some of it. So his aha moment wasn't learning a formula, but one day suddenly figuring out QKV — Q represents what this character of mine wants to look for, K represents what this character of mine is so others can use it to find me, V is what information you actually take away after finding me.

— Su Tinghao
22:20

The top 20 happiest moments of his life were playing chess

On the first evening of ICML, from seven to nine, there was a chess social event; the room was full of boards, and you could find someone and play. He calls this one of the top twenty happiest times of his life: walking up and asking what your research direction is, talking about AI while playing chess, meeting many people from different places over a few days, and still getting to keep playing chess. During those days he also did something very much like himself — he asked every person he met whether AI would lead to human extinction, asking roughly 30 to 35 people over that period. He also says that after coming back from Korea this time he has changed: his personality used to lean introverted, and now it leans a bit more extroverted.

— Su Tinghao
29:22

Application companies earn more, but they can only listen to upstream

His division of labor is stated bluntly: model companies are responsible for making the model good, and application companies ultimately just add something on top, can't change the specific model, and also make money by relying on the model, so they can only listen to model companies. But he also gives a counterintuitive judgment — as things stand, application companies make more money than model companies, because right now it seems that apart from those few, model companies are all burning money and have no profit. And he pushes the endgame even harder: if some company really builds AGI, AGI can write application software itself and directly replace application companies. He also says every company that wants to do AGI must say it wants to do AGI, because only then can it raise funding, and the ultimate goal is still to make money.

— Su Tinghao
38:26

Studying math is left with only proving to universities you can endure hardship

He describes the change AI has brought very concretely: before, when you studied physics and chemistry, you could see that you were better than some people, and you could see all the way to university, a PhD, advancing that discipline; now even people who haven't studied it can ask AI with one click and get a better answer than yours. So his motivation changed — studying math is no longer for the subject, but to prove to universities that he is someone who can endure hardship and can learn. The classmates around him who are very interested in aerospace physics are also thinking that in the future AI might do better than they can. He uses ‘it has narrowed, it has become harder’ to describe this path; what is lost is part of the sense of mission to contribute to a subject, to humanity.

— Su Tinghao
55:33

From ICML to teaching in the countryside, he deliberately turns AI off

Four or five days after coming back from ICML in Korea, he went to teach in a rural area, and the contrast is one he tells himself: at ICML everyone talks about the most advanced AI, whose model is how good; at the teaching site, every day he thinks about what lesson to prepare for tomorrow, and in the evening whether to walk in the farm or go catch insects nearby. He deliberately makes himself leave AI during this time, talk more with classmates, with the people who came with him, and with the kids, and play games. He has gone to teach for several consecutive years, doing the same thing: local kids use Chinese characters to remember English pronunciation, and he spent more than 500 yuan on Taobao buying gifts and designed a points system where kids answer questions in class and join games to earn points; registration only took 30 people, and in the end people could only get in by having the village chief call.

— Su Tinghao

In their own words · checked verbatim

Post-training is more like shaping a body, while pretraining is more like you make that specific material better to withstand the shaping of post-training that comes later. Pretraining is a bit magical, a bit like magic.

后训练更像塑造一个身体吧 预训练更像是你把那个具体的材料 变得更好来承受之后的后训练的塑造 预训练有点那个magical 有点像魔幻一样

Su Tinghao11:13

If AI can be controlled by humans, and it's very strong, then some of the people who control AI can become that god, because they can control an omnipotent AI, and then other people might become ants.

要是AI能被人类控制的话 而且很强的话 一些控制AI的人 就能变为那个神了 因为它能控制 无所不能的AI 然后其他人 就可能会变成蝼蚁了

Su Tinghao23:21

As things stand now, application companies earn more money than model companies, because right now apart from, like, those few companies, they're all burning money, none of them have that profit.

现在来看这个 应用公司是比模型公司 挣过更多的钱 因为现在除了好像 安斯拉北家公司都在烧钱 都没有那个profit

Su Tinghao29:22

With today's AI, you might feel that I'm actually studying math just to prove to universities that I'm someone who can endure hardship and can learn, rather than studying for this subject.

现在的AI 你可能觉得 我其实学数学 我只是为了 给大学证明 我是一个能吃苦 能学习的人 而不是为了这个科目 而学习

Su Tinghao38:26

The day KimiK3 appeared, I was happy for a whole day, because at that time I thought domestic or open source models would need a few months to catch up, but it turned out it didn't take that long to catch up.

KimiK3出现那一天 我高兴了一天 因为呢 我那时候以为 那个国内 或者Open Source模型 需要那个几个月 才能追上的 但是发现 其实没那么的长的时间 就可以追上了

Su Tinghao49:30

If I could tell my earlier self, I'd say seize this time, get away from these things, just don't overthink it, talk more with your classmates, with the people who came this time, with the kids this time, play some games, like Werewolf, hide-and-seek, just deliberately tell yourself, don't overthink these AI things.

我若是应该 之前告诉自己 抓紧这个时间 能离开这些东西 就是别多想 跟你同学 跟你这次来的人 跟你这次的小孩 就多交流 玩一些游戏 比如浪人杀 捉迷藏 就故意告诉自己 别多想这些AI的东西

Su Tinghao56:34

AI's intelligence — when it improves, all AI's intelligence improves, but humans are generation by generation, and it all goes through this slow growth period, so our cultivation is actually much slower than AI's.

AI的智商 当它提升 所有AI的智商都提升 但是人类都是一代人 一代人 它都经历 这个缓慢的成长期 所以我们的培育 其实是比AI慢很多的

Su Tinghao1:06:43

Figures

Papers accepted to ICML 202613:07
Total cost of training for the paper30,000 yuan RMB11:13
Estimated cost to scale upabout 6,000 yuan RMB11:13
Loss from three models not saving checkpoints1,500 (the source does not specify the currency)13:13
Fastest time to hand-write a Transformer from scratchabout 25 minutes14:13
Challenge of reading one paper a day30 consecutive days5:08
Taobao gift spending for the teaching activitymore than 500 yuan RMB53:31

Glossary

Attention Projection Mixing with Exogenous Anchors
The title of the paper he got accepted to ICML 2026; it changes the residual method of the attention projection.
Value Residual Learning
An existing method that adds the first layer's attention value to later layers as a residual; the starting point of his paper.
fineweb edu
The public open-source pretraining dataset he uses for free to train models, used up to 20B.
UBI
The idea he mentions in his ten-year prediction: most people don't need to work, and eat and drink at home.

How to listen

Who it's for

Investors watching AI's impact on education and talent, people building AI education products, and engineers who want to know how the post-2005 generation actually uses AI.

Skip

The chess, olympiad math and competition experience from 8 to 11 minutes can be fast-forwarded.