The world is too loud. Read what matters.

张小珺·商业访谈录

Sora Is a GPT-3 Moment, Not a ChatGPT Moment

What Sora shows is an emergent capability, not a usable product; it pushes the large-model war to a new level, but catching up in a year is too optimistic, because between being able to do it and actually doing it lies a great deal of know how.

SoraLarge ModelsMultimodalityInvestmentBubble
Two investors sampled a dozen-plus groups of people to piece together the puzzle of how Sora is implemented, and also give their judgments on the large-model elimination race, the bubble, and applications. Information density is high, but the second half is more valuable.

The argument · tap a timestamp to hear it

4:07

Sora is a GPT-3 moment, not a ChatGPT moment

Ji Yichao likens Sora to GPT-3: it is a model that has emerged with certain capabilities, not a product. ChatGPT was a product innovation that aligned GPT into a Chat form, letting users converse and generate more input and feedback; Sora, by contrast, generates a video and then it's over, with no interaction with the user. So it still has a long way to go before becoming a product, and what emergence shows is only a spark. This distinction determines how to price Sora: it proves that a known model architecture can produce Sora, but that does not mean it can start making money right away.

— Ji Yichao
23:45

Model size is about 6B, compute one or two thousand H100s

From the sampling, the model parameter count is roughly around 6B, based on the 4x and 16x compute scaling examples in the Tech Report and extrapolated from the DiT paper as a baseline. Training compute may be on the scale of one or two thousand H100s, that is, tens of millions of dollars, not something that requires a billion dollars to do. But running experiments and doing the final training run are two different things; OpenAI's already-deployed cards are estimated at two hundred thousand plus, so it can run experiments in parallel with more cards. Data is what everyone is most focused on now, because model size has not increased significantly, and progress generally lies in the data and how it is processed.

— Ji Yichao
28:54

Between being able to do it and actually doing it lies know how

Dai Yusen thinks catching up to Sora in 6 to 12 months is too optimistic. The general direction is known, but how to tune the specific parts is not, and that is exactly what has to be figured out. After GPT-4 came out last year, many companies said they would catch up by the end of the year; now it is 2024, and the only one that has truly reached GPT-4 level across the board may be Google's Gemini 1.5, and it has not yet been used in an actual product. Ji Yichao adds that OpenAI's starting point is much higher than that of domestic companies; Sora used DALL·E 3's recaptioning, and the language conditioning part may have used GPT's weights, so even with Sora here, it will still take a long time of investment to get close.

— Dai Yusen
34:37

Sora pushes the war to a new level, and invites regulation

The impact of Sora's release on the global large-model battle is, first of all, to push this war to a new level. First-tier companies all knew multimodality would see a breakthrough like this, it is just that the timing came earlier, and many original plans will have to change. Unlike ChatGPT, it is not an interactive model; as a pure generative model, what happens after generation is still unknown. But it presents such a good video so intuitively that the impact on startups, application companies, the entertainment and film industries, and government departments is huge, and it will also invite more worries about AI regulation; departments such as the broadcasting regulator are paying very close attention to Sora's appearance.

— Dai Yusen
37:06

World simulators have three layers: physical, social, long-duration

Ji Yichao divides world simulation into different layers: the first stage is relatively static scenery like aerial empty shots; the second stage is physical causality, such as a cup filled with water falling and shattering, but Sora's result on this example is very poor, the cup does not know how to shatter; the next layer is social causality, such as a shark appearing on a beach causing surprise, a dolphin causing cuteness, a crocodile causing flight, or a baby punching its father versus a strong man punching a thin person meaning completely different things; further still is the echo between the beginning and end of a film. He believes what Sora shows now is still an emergent capability, and it will take quite a long time before it is truly usable.

— Ji Yichao
46:22

A unified model has not yet found a viable path

Language is discrete, while images and audio are natural continuous signals, and putting all different inputs into one feature space is itself a difficulty. What is harder is the training objective: language models predict the next word, diffusion models learn to create or predict noise, and how to design a loss or task that lets it achieve both multimodal understanding and multimodal output at the same time is very hard and worth researching. There is currently no definitive evidence that training multiple modalities together can give the model a higher breakthrough in capability; the main capability of multimodal models still comes from the language backbone, and the more likely outcome is that progress advances in a stitched-together or plug-in form.

— Ji Yichao
56:19

The bubble is not scary; it brings infrastructure

Dai Yusen says the bubble will bring important infrastructure, and infrastructure lays the foundation for future applications; the 99% that die in the bubble will die, but the 1% that remains may be great companies. After the internet bubble burst, what remained was Amazon and Google, and Yahoo also survived for a long time at the time. Now we are far from the point where the bubble is relatively crazy, because the internet's true peak came when the first wave of internet-native applications truly landed and truly went public. Why not wait for the bubble to burst before investing? Because not investing at all means missing Amazon and Google, and the industry knowledge gained along the way is very important; it is hard to eat only the fifth bun.

— Dai Yusen
1:14:27

AI applications become useful more slowly, but spread faster

Ji Yichao uses mobile internet as an analogy: after the iPhone moment there was still the App Store moment and the iPhone 4 moment; now GPT has a GPT store but it cannot be compared to the App Store at all, and we may be at the point of 0708. Dai Yusen thinks that under the existing paradigm, the point at which AI applications become useful will be slower than mobile internet, because the model needs to reach a certain level of capability before it can emerge from useless into useful; but once it becomes useful, its spread may be far faster than mobile internet applications, because the internet and mobile internet were processes where both software and hardware had to spread, whereas for AI, as long as the device does not change and it runs on a phone, once useful it could sweep across hundreds of millions of people in one or two years.

— Dai Yusen

In their own words · checked verbatim

Sora is now also a model that has emerged with certain capabilities, but for it to become a product, there is actually still a long process in between.

Sora现在也是一个涌现出来了某些能力的一个模型 但是它要变成一个产品 其实这个过程中还有很远的过程

Ji Yichao4:07

Because I think being able to do it and actually being able to do it, there is actually a lot of know how in between.

因为我觉得能做和真能做出来 其实中间隔很多know how

Dai Yusen29:16

The bubble is not scary; our bubble will bring important infrastructure, infrastructure will lay the foundation for future applications, the 99% that die in the bubble will die, but the 1% that remains may be great companies.

泡沫不可怕 我们泡沫会带来重要的基建 基建会为未来的应用打下基础 就泡沫中死掉的99%会死掉 但是1%留下来的可能就是伟大公司

Dai Yusen56:19

We must avoid a trap: over-embellishing when the technology is not yet good enough.

我们要避免一个陷阱 就是在技术还不够好的时候 过分雕花

Dai Yusen1:09:39

Under the existing paradigm, AI applications, the point at which they become useful may be slower than mobile internet, because it needs the model to reach a certain level of capability before it can emerge from useless into useful, but once it becomes useful, its spread may be far faster than mobile internet applications.

在现有的范式下 AI应用 它可能应用出现 有用的时间点 会比移动互联网要慢 因为它需要模型 到达一定的能力程度 它才能从没用 涌现成有用 但是当它一旦变得有用之后 它的扩散速度 可能会远快于 移动互联网的应用

Dai Yusen1:14:27

If you hand over the thinking of your current prime years to a future AI, then to a certain extent you can achieve a kind of digital immortality.

如果你把你现在 年富力强的时候的思路 交给未来的一个AI的话 其实你一定程度上 能获得一个数字的一个永生

Ji Yichao1:30:45

Figures

Estimated Sora training computeone or two thousand H100s28:54
Estimated Sora training costtens of millions of dollars28:54
Estimated number of cards OpenAI has deployedtwo hundred thousand plus23:45
Inference time for Sora to generate a one-minute videoabout 20 minutes in reality33:18
Sora generated video length60 seconds27:16
Scale of domestic large-model financingtens of billions of dollars52:32
Scale of global large-model financinghundreds of billions of dollars52:32
ChatGPT user countover a hundred million people have used it1:32:45

Glossary

DiT / Diffusion Transformer
A model architecture that uses the Transformer architecture for diffusion generation, the backbone of Sora.
recaptioning
Using GPT-4 to write more detailed text descriptions for training material, enhancing generation realism.
native resolution
The model supports various resolutions in training and inference, directly outputting the resolution suited to the device.
cherry picking
Selecting the best from multiple generation results for display, with the outside world unable to see the failure cases.
alignment
Making model behavior conform to human intent and values, involving the question of whom to align with.

How to listen

Who it's for

Investors watching the large-model battle, engineers working on video generation or multimodality, founders trying to judge the timeline for catching up to Sora.

Skip

From 1:31:43 onward the talk about Jike and valuation methods is more casual chat and can be skipped.