The world is too loud. Read what matters.

张小珺·商业访谈录

NVIDIA's VP of Research: The Endgame for Models Is Ecosystem, Not Beating Rivals

Ming-Yu Liu: model capabilities will eventually converge, and NVIDIA's open-source world model is not competition but groundwork; physical AI's GPT moment has not arrived, but the conditions are converging.

World modelsCosmosNVIDIAJensen HuangPhysical AIRobotics
Ming-Yu Liu is NVIDIA's first, and still rare, case of a researcher promoted all the way to VP. He unpacks the judgment behind how Cosmos was started, converged and open-sourced, and offers a first-hand view of how Jensen Huang manages.

The argument · tap a timestamp to hear it

4:40

At NVIDIA, world models are a product, not a research project

Liu gives his title as VP of Research and VP of Cosmos Lab. The team reporting directly to him is roughly eighty or ninety people; adding the other teams inside the company doing Cosmos-related work brings the total to 200 to 300. He is the owner of the entire mission and the person who decides its direction. He stresses that this is a very significant internal program at NVIDIA, and that committing this many people is simply what ‘producing good results’ requires. That headcount is the tell: world models have been elevated inside NVIDIA to a level they have never held before — this is not a research project but an engineering organization with an explicit product mission.

— Ming-Yu Liu
1:24:21

A manager's recognition has to be aligned with the company's interest

Liu says he moved up a level every year in his first few years at NVIDIA, and that he may be the first person there to go from researcher to VP. The reason is that he joined early, before the company had AI talent, and helped build NVIDIA's standing in the AI community while explaining internally how AI would affect the business — and at the same time he liked shipping things, which earned trust step by step. Today the team he leads runs up a compute bill of several million dollars a day, and the fact that the company will hand him an operation of that size is what long-term alignment buys. He is emphatic that alignment with the company matters enormously: a pure researcher wants recognition, but a manager's recognition has to be consistent with the company's interest.

— Ming-Yu Liu
1:41:09

Before Sora, persuading Jensen Huang was hard; after it, easy

The decision to build Cosmos came around March 2024, right after Sora. Liu says that before Sora appeared, persuading Jensen Huang was fairly hard, and that after Sora it was not hard at all, because Huang saw immediately what it meant for the company. He and a group of researchers working on generative models wrote Huang an email together, asking that NVIDIA build a world model. Huang pulled everyone into a meeting — and told each of them to go away and come up with a product name. Liu's own reaction was that Huang would produce a better name anyway, so why waste the effort; Huang insisted he do it, because the name is bound up with the product's positioning. They settled on Cosmos, and the goal was never creative work: it was to solve problems in the physical world.

— Ming-Yu Liu
1:54:47

22 models is not richer choice, it is slower iteration

In the Cosmos 1 era Liu did his own estimate and concluded that covering the range of customer needs might take 22 models. He reported that number in a meeting with Jensen Huang, and Huang answered with Costco: when they cut the number of items on display, sales went up, because too many choices confuse the shopper. Liu also found that he himself could not always tell when to use A and when to use B — never mind an outside user. On top of that, every model has to be maintained, so the more models there are, the slower iteration gets. That is why Cosmos 2 to 3 is such a large gap: a convergence from many models into one unified world model that folds understanding, prediction, generation and action together.

— Ming-Yu Liu
2:08:22

The dual tower is only a stage; one tower would be enough

Cosmos 3 uses a dual-tower structure: one tower handles discrete information and is responsible for understanding, the other handles continuous information and is responsible for generation. Liu explains that this is a very sensible choice at the current stage, because keeping discrete reasoning and continuous generation apart means customers doing post-training on one will not disturb the other. If everything were wired together, a customer adjusting the generation distribution could damage the understanding capability. But he says plainly that the next goal is something simpler: ‘why two towers — one tower is enough’. They are moving toward a single tower, while weighing at the same time whether users can get started with it easily; you cannot sacrifice usability in pursuit of engineering perfection. The final architecture is not settled, and will depend on which one has better deployment cost.

— Ming-Yu Liu
2:31:45

People who fear mistakes stop deciding, so leave room for error

The Cosmos 3 paper has over 290 authors, just under 300. Liu thinks the core tension in a large team is that you need everyone pointed at the same goal and, simultaneously, every individual able to make decisions on their own. If only a few people decide, progress is slow; if everyone decides for themselves but the goals are not shared, the work diverges or turns into internal horse races. His answer is to put a great deal of thought into how the team collaborates: keep information transparent so everyone has enough of it to make the right call, and provide room to make mistakes, because people who are afraid of being wrong will not make decisions at all. He also stresses that well before the paper existed he had told the main leads exactly what the model would look like, so that everyone could see the same picture — the way a basketball player pictures the ball going cleanly through the net before taking a free throw.

— Ming-Yu Liu
2:52:41

Not laying people off is slower, but it buys employees who speak up

Liu says NVIDIA has no forced ranking and does not do layoffs; in his ten years there he has never heard news of a layoff. His view is that employees will only say the things people keep quiet about in places where they fear being cut once they feel safe with the company and trust it; if everyone is scrambling to stand out, cohesion and collaboration get harder. The downside is that when a technology matures, employees need time to learn new things and retrain, which is slower than cutting them and hiring fresh people. What the company gets in exchange is trust and cohesion. Jensen Huang has watched a great many storms pass through Silicon Valley, and choosing this model is a considered decision. NVIDIA runs a marathon, and it does that by positioning early — CUDA, for instance, was a ten-year bet.

— Ming-Yu Liu
3:07:40

You cannot blame the setup for falling behind; nothing was ever fair

Liu calls Jensen Huang a ‘hexagonal all-rounder’: he can go down to register-level detail in computer architecture and also talk marketing and pricing, covering every side. He recalls once explaining a gap against a competing product by saying the ‘setting was unfair’, and Huang came straight back with ‘Are you kidding me? Actually, nothing is fair’ — which taught him that doing research means creating the conditions your goal requires, not finding reasons. The core of Huang's management is that Mission is the boss: the company behaves like an organism, and when it decides to move in a direction, the skills of different teams align by themselves. Liu also notes that Huang reads an enormous volume of email every day, hunting for signals and making priority calls, and that this has shaped how Liu now works — switching topics every half hour.

— Ming-Yu Liu

In their own words · checked verbatim

When we finished cosmos 1, I asked jason whether we should keep going. ... And then janson said to me: you just keep going until cosmos97.

我们做完cosmos一的时候,问那个jason要不要继续往下做。…然后janson跟我说,你就做到cosmos97。

Ming-Yu Liu1:36:54

We believe a good model of the physical world also has to treat action as a first citizen. ... That is why we put language, uh, video and action together to build this cosmo.

我们认为一个好的物理世界的 model也要把action当成first citizen。…这是为什么我们把language哦video跟action合在一起建这个cosmo站。

Ming-Yu Liu2:02:11

Do you need your robot to be able to solve Olympiad math problems? Do you need a robot that can solve coding problems?

你需要你的机器人会觉奥数问题吗? 你需要有机器人会解这个co定的问题吗?

Ming-Yu Liu2:19:21

If you can work together with us and make use of what we have built, then you can go faster and further.

如果你可以跟我们一起合作啊,利用我们做的东西,那你可以走的更快更远。

Ming-Yu Liu2:43:18

Every single day he behaves as if the company only has 30 days of cash flow — uh, as if in 30 days it goes bankrupt.

他每天都觉得公司只能要30天的现金流,嗯,30天后就要破产。

Ming-Yu Liu2:50:34

When there is no such thing as cutting the bottom performers, people also do not get, well, do not get so worried about being cut that they fight to stand out.

当没有所谓的末位淘汰的时候,大家也不会就是真的这么的啊就是这么担心被淘汰而抢出头。

Ming-Yu Liu2:53:48

You do not need to beat your rivals. I do not like treating them as rivals; I like treating them as partners.

你不需要击败你的对手。我不喜欢把他们当成是对手,喜欢把他当成是伙伴。

Ming-Yu Liu3:04:25

What he said to me at the time was: Are you kidding me? Actually, nothing is fair.

他那时候给我句话说,are you quietba.实际上没有什么是公平的。

Ming-Yu Liu3:08:41

Figures

Headcount across Cosmos-related teams200 to 300 people4:40
Models estimated as necessary in the Cosmos 1 era221:54:47
Authors on the Cosmos 3 paperover 290, just under 3002:31:45
Daily compute cost of Liu's teamseveral million dollars1:24:21
Cosmos release cadence, and time to reach generation 97one generation every six months; getting to generation 97 would still take over 40 years2:29:45
NVIDIA headcount when he joinedunder 20,000; more than 2x that now2:50:34

Glossary

Physical AI
AI that can perceive the physical world and act to change its state — robots and autonomous driving, for example.
World Foundation Model
How Cosmos positions itself: a lower-level foundation model that serves every physical AI application.
Multimodal Transformer
A Transformer architecture that handles text, video, audio and action signals in one place; the core of Cosmos 3.
Dual-tower
A two-tower structure that processes discrete understanding signals and continuous generation signals separately; Cosmos 3's choice for this stage.

How to listen

Who it's for

Founders, investors and engineers following world models, robotics and NVIDIA's technical roadmap — especially anyone trying to understand why NVIDIA builds open-source models at all.

Skip

The first 20 minutes on his education and career path can be played at speed; the details on Jensen Huang's management and on how Cosmos evolved come in the second half.