World Models Are Not Language Models: Xie Saining Says LMs Will Eventually Wither
Xie Saining argues that language models are excellent tools but not the foundation of general intelligence — they will eventually wither. The real foundation is vision and representation learning, which is why he left Silicon Valley and co-founded AMI with Yann LeCun.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Rejecting Ilya twice, the disagreement is over vision
In 2018 Ilya called, and Xie Saining rejected the OpenAI offer without saying anything. Ilya was angry and asked why he wouldn't even discuss it — was the money not enough? The second time was July 2024, when SSI had just been founded and Ilya emailed to invite him to work together. He rejected again. This time the two mainly discussed ‘how to give future AI the capacity for love’, but Xie Saining finally asked: what do you think about multimodality, about computer vision, about general perception models? Ilya's answer was ‘he felt this problem had already been solved quite well’. Xie Saining says SSI has its own language-based路线, and that路线 is at least for now not the路线 he wants to design.
— Xie SainingLMs will eventually wither, but they are not without value
Xie Saining's judgment on language models is ‘LMs will eventually wither’, but he immediately adds ‘LMs will never die’ — old soldiers never die, they just wither away. What he means is: this thing will certainly have its value, it is a very good tool, he himself uses LMs every day, but it is not the cornerstone on which we build a general intelligence system, it is not the foundation of the world-model skyscraper. This judgment is the premise of all his later choices: why do vision, why do world models, why start a company with Yann LeCun.
— Xie SainingThe baseline determines the ceiling of your research
Kaiming taught Xie Saining one thing: the ceiling of your research actually depends on how good your baseline is. If the baseline is very poor, you can easily fool yourself and end up producing nothing. This is counterintuitive, because everyone always says that if the baseline is a bit worse, the performance gain will be a bit bigger, and it will be easier to publish a paper. But Kaiming's idea was: push the baseline as high as it can possibly go, and then do something new on that foundation — that is what counts as ground breaking. Any improvement made under a weak baseline may just be a watered-down paper. So Kaiming single-handedly built an entire infrastructure from scratch on TPU, and Moco, MAE, and DiT all happened on top of it.
— Xie SainingThe Diamond Sutra and research taste
The book Kaiming gives to new hires is the Diamond Sutra, not a book teaching how to do research. Xie Saining says this touches on the question of research taste. The Diamond Sutra says ‘all appearances are illusory; if you see all appearances as non-appearances, you see the Tathagata’, which resonates with Kant's thing-in-itself and Schopenhauer's The World as Will and Representation: what you see is not the essence of the thing, and the world you see is not the substance. So when you read a paper, the important thing is to break the illusion the paper gives you and ask what substantive thing is actually hidden behind it. The source of research taste lies in whether you can cast aside these illusory appearances and keep seeking along the path to truth.
— Xie SainingKaiming finished his papers a month before the deadline
All of Kaiming's papers were finished a month before the deadline. While others were still pulling all-nighters for the deadline and getting enormous satisfaction from it, Kaiming had already finished a month earlier, then began polishing over and over, watching you all rush for the deadline. That means he started writing two months in advance. Xie Saining says this influenced him and gave him OCD: no line in a paper may have less than 60% text; if a line is mostly empty, it doesn't look good, so you have to fill the line or fill it to sixty or seventy percent, so the paper looks more elegant and uniform. Kaiming's idea was: a paper is not for you to read, it is for others to read, so you have to care about others' perception.
— Xie SainingDiT was rejected by CVPR for insufficient novelty
DiT was submitted to CVPR and rejected, on the grounds of insufficient novelty — you don't have long stretches of math, you don't have long stretches of complex structure, you made a very simple structure, and although it gets good results, the reviewers didn't buy it. Xie Saining says that by that moment he had slowly come to his senses: this matter of research papers, within this enormous random process, whether it gets in or not doesn't matter at all. So they then submitted to another conference, changed nothing, and got an oral paper. This again proves it is a purely random process. Later DiT was better than UNet-based systems on every dimension, but after it was published, although it got a lot of attention, nobody actually used it to do anything.
— Xie SainingSaw the Perplexity demo at Blue Bottle
Xie Saining says he may have been the first or second person to see the Perplexity demo. Aravind had come out of OpenAI, and at the Blue Bottle coffee shop in Palo Alto — which is also where many things in Silicon Valley happen — he held a laptop and showed him a browser, saying ‘we are going to revolutionize Google’. Xie Saining thought to himself, what is this thing, isn't this just GPT wrapped in a shell, why are you doing this. Then the other person asked if he wanted to join, and he said he still rather enjoyed being at NYU and continuing to do research. He says his understanding of entrepreneurs has also changed somewhat over the past few years; this thing really is different from research, there are many things in common, but also some different points.
— Xie SainingThe resource predicament of American academia
Xie Saining says academia in North America is in a very difficult position, mainly because resources are insufficient. The American funding system has barely grown over the past few decades; although there has been high inflation, everything has become very expensive, student tuition has become very expensive, but government funding and corporate funding programs remain at a very low level. For government agencies like the NSF, the total funding they can give to a single PI is roughly at the level of 500,000 dollars, over five years, about 100,000 per year. And industry funding opportunities are becoming fewer and fewer; once there is a funding opportunity, they generally give you 100,000 to 150,000 dollars, but roughly 100 teachers from 100 schools compete for that 100,000 dollars. 100,000 dollars can support one student for a year as tuition, or buy half an H100 cluster, or buy three to four cards.
— Xie SainingIn their own words · checked verbatim
LMs will eventually wither. No — LMs will never die, but they will eventually wither. Old soldiers never die, they just wither away.
LM终将凋零 不对 LM LM永不会死 但终将凋零 就老兵不死 终将凋零
Xie Saining2:24:32
The ceiling of your research actually depends on how good your baseline is. If your baseline is very poor, you can easily fool yourself, and you won't be able to produce anything.
你的research的上限 其实取决于你 based on的好坏 好 就如果你的based on很差的话 你可能很容易自欺欺人 你是做不出来什么东西的
Xie Saining2:33:37
All of Kaiming's papers were finished a month before the deadline — at least that was the case at FAIR.
凯明所有的论文都是在deadline前一个月做完的 只要在fair的时候是这样的
Xie Saining2:47:48
This matter of research papers, within this enormous random process, whether it gets in or not doesn't matter at all.
这个research paper这件事情 其实在这个巨大的随机过程里面 重或不重 一点都不重要
Xie Saining3:06:04
Figures
| AMI Labs team size | 25 people | 0:00 |
| Number of top-conference papers Xie Saining published during his PhD | five or six | 44:26 |
| Number of internships Xie Saining did during his PhD | five | 52:32 |
| Number of TPU cores rented by FAIR | about 5,000 | 2:32:35 |
| Total NSF funding to a single PI | about 500,000 dollars, over five years, about 100,000 per year | 3:23:14 |
| Amount of a single industry grant | 100,000 to 150,000 dollars | 3:23:14 |
| Number of teachers competing for industry grants | about 100 teachers from 100 schools | 3:23:14 |
Glossary
- research taste
- The ability to judge which problems are important and which directions are worth pursuing; Xie Saining believes it comes from casting aside appearances and asking about substance.
- representation learning
- Mapping data into a vector space with good properties, making downstream tasks more likely to achieve good results.
- contrastive learning
- In representation space, bringing similar objects closer and pushing different objects farther apart; Moco was the first framework that truly worked.
- anti-fragile
- Taleb's concept: a system that gains more than it loses from random shocks; Xie Saining believes research must be anti-fragile.
- DiT / Diffusion Transformer
- Replacing UNet with a ViT architecture for diffusion models, work Xie Saining and Bill Peebles produced in their last month at FAIR.
How to listen
AI researchers and founders following the debate between world models and the vision-versus-language路线; engineers who want to know how top researchers choose research topics and build baselines.
The first two hours of childhood, schooling, and internship chronology can be skipped; start from the research taste discussion at 2:43.