The world is too loud. Read what matters.

Machine Learning Street Talk

A world model doesn't predict the future, it ranks robots for validation

Cosmos 3 stuffs forward dynamics, inverse dynamics and policy into one model, but the first thing it ships isn't training — it's using a neural simulator to rank a pile of checkpoints and filter most of them out before they ever touch real hardware.

World modelsRoboticsNVIDIASimulationEmbodied AI

The video won't play here. Listen to the audio instead:

The technical details and deployment path of NVIDIA Cosmos 3, laying out clearly what world models do first on robots and what comes later. Medium information density; the engineering judgment in the back half is worth more than the architecture description in the front half.

The argument · tap a timestamp to hear it

2:28

One model does three jobs, and an information bottleneck forces them to cooperate

Cosmos 3's architecture is two towers: an autoregressive vision-language model handling understanding and reasoning, and a bidirectional diffusion generator tower handling generation of video, action and audio. Inside the generator tower every token attends to every other token, so signals within a chunk are coherent. The key design choice is putting forward dynamics (given actions, predict the future), inverse dynamics (given visual change, infer the action) and policy (given a task, decide what to do) into the same model, and imposing an information bottleneck — capacity is limited, so the three must share a representation of the correlation between observations and actions. The paper's results show the three are synergistic, one helping another. Ming-Yu says this is essentially different slices of the same thing.

— Ming-Yu Liu
6:24

Signals at different frequencies must be squeezed onto one timeline

Video has different frame rates, audio has different hertz, and actions have their own frequency. To have one model handle all of these at once, you have to solve the problem of inconsistent time scales. Cosmos 3's approach is temporal position embedding, normalizing all signals onto the same axis so that each token represents one time segment of a signal, and the model knows both which tokens belong to the same moment and the relative distance between different moments. Ming-Yu says this is a key part of making the whole model work. Without this step, audio, video and action can't be aligned in the same attention space.

— Ming-Yu Liu
9:28

Human video is plentiful and robot data is scarce, but what transfers is patterns, not actions

The training data is severely asymmetric: first-person human video is abundant, robot video far scarcer. But Ming-Yu thinks transfer is feasible because of a shared vocabulary — the visual pattern of a human hand manipulating an object and the visual pattern of a robot gripper manipulating an object are very similar in the correlation between observation and action. Even if the action spaces don't correspond exactly, once a set of actions is aligned, the patterns are similar, and generalizing to another embodiment becomes easier. Training with multiple embodiments also helps generalization to an embodiment never seen before. He admits the model doesn't know when it shouldn't transfer, and that transfer can be harmful, but stresses that what matters is the pattern, not the exact action space.

— Ming-Yu Liu
12:53

The neural simulator does passive validation first, not training

Asked about the sim-to-real gap and policies gaming the simulator, Ming-Yu admits it can happen. But the deployment order he gives is: the neural simulator is used for passive validation first. Model training produces a huge number of checkpoints, and ideally you'd deploy every policy to real hardware and real environments to measure completion rate, but that costs too much. Use the world simulator in place of real hardware and a real operator, let the policy interact directly with the simulator, and measure success rate after the rollout ends. The simulator's success rate doesn't need to be accurate, only its ranking needs to hold — if A beats B in the simulator, it probably does on real hardware too. That quickly shrinks the number of checkpoints needing real-hardware validation and speeds up development. He says that even if the world model isn't mature enough to feed training data, it can still do policy validation.

— Ming-Yu Liu
16:47

Cosmos as a starting point: data, environments and initial weights

Cosmos wants to help the ecosystem with three things: better data, better environments, a better starting point. Ming-Yu says learning a shared representation by connecting video, action and text is a good starting point for building policy models, and this isn't theory — post-training a Cosmos model on a shared dataset produced pick and place policy results. The intuition: if you can predict pixel dynamics well and learn the correlation between pixels and actions, that helps you build policy models. As Cosmos models advance and bring in more embodiments, its quality as a starting point will keep rising, and the work needed to adapt to a new embodiment will keep shrinking. The repo provides post-training recipes so developers can reproduce the results.

— Ming-Yu Liu
19:53

Navigation is good enough already; manipulation is the real hard problem

Ming-Yu splits robot tasks into two categories. In navigation tasks you don't want the physical device to touch anything, which is relatively simple, and he thinks one model is good enough today. But manipulation tasks demand all kinds of interaction, and where there's contact there's compliance and potential deformation, so they're harder. He's optimistic about the whole field's enthusiasm for world models, the steady progress of deep learning research and the rise in compute, and thinks we'll reach a state where neural simulation handles complex manipulation tasks well. At that point you can test robot policies with a world simulator. This judgment draws Cosmos's capability boundary clearly — not all robot tasks are equally mature.

— Ming-Yu Liu
20:57

For humanoids entering the home, safety validation is the only way

Ming-Yu uses autonomous driving as an analogy: cars already move through human space, so today you need a lot of policy validation to guarantee they're safe enough. When humanoid robots enter many people's homes, safety matters even more — you don't expect a car in your house, but you do expect a humanoid robot around you, your children, your pets. How do developers guarantee that every policy iteration, and the whole system, is safe enough? You need policy validation. But you won't have enough space to build every kind of kitchen and every kind of task. So he thinks neural simulators are the only way to give you development speed. This passage pushes the motivation for Cosmos Dreams from technical to deployment scenario.

— Ming-Yu Liu
22:04

The world model should live where the robot lives

Cosmos comes in three sizes: Super is the frontier model, highest fidelity, the choice when you need maximum accuracy and have the compute; Nano sits in the middle, smaller, easier to post-train on all kinds of GPU resources; Edge has to run on edge devices like Jetson Thor, Orin and DGX Spark. The reason for Edge is to let the world model live where the robot lives, so taking an action doesn't require a round trip to the data center. Ming-Yu says in real deployment you can't count on the network connection being good enough to complete all tasks, and some scenarios are safety-critical, so the task must be completed on the spot. Edge models are smaller, need less compute, and are easy to fine-tune — they have a recipe that fine-tunes Cosmos Edge within a day to improve visual understanding. He also envisions the three sizes working together in the future: hard tasks call the strong model, simple tasks are handled on the spot by the edge model.

— Ming-Yu Liu

In their own words · checked verbatim

I think a world model is going to be the same. So I think a world model is a collection of useful tools. We model something because we are trying to achieve some goal.

Ming-Yu Liu5:22

Yeah, so I think that these three, they are trying to fundamentally trying to capture the correlation between the observation and action. Right? It's just a different slice, different perspective.

Ming-Yu Liu8:26

So I think it's more about our shared vocabulary for different embodiments, right?

Ming-Yu Liu9:28

If the success rates of the world simulator, the new world simulator correlates with the real-world, you know, testing, the ranking preserve, you don't need them to be precise. You just need to know if policy A is better than that. and policy B in neural simulator.

Ming-Yu Liu14:42

So even before it's mature enough to feed the policy model, the training data to train, wall simulator can be used as a policy verification.

Ming-Yu Liu15:43

You don't expect to see a car, you know, in your house. But once human-knowing is everywhere, you're going to expect to see human-knowing surrounding you and maybe your kid, your pet, right? So you do want them to be safe enough.

Ming-Yu Liu20:57

So the reason we built this one is that we want the world model lives where a robot lives, right? So that you don't need to have a round trip to data center when you want to take a certain action.

Ming-Yu Liu23:01

Figures

Cosmos model size tiersThree tiers: Super, Nano, Edge22:04
Target devices for the Edge modelJetson Thor, Orin, DGX Spark22:50
Cosmos Edge fine-tuning timeFine-tuned within a day to improve visual understanding23:01
Cosmos model and data hostingModels and data on Hugging Face, code on GitHub24:24

Glossary

world model
A family of models that model environment dynamics and can predict future observations or infer actions.
forward dynamics
Given a starting state and an action, predict what happens next.
inverse dynamics
Given a visual state transition, infer what action was performed.
sim-to-real gap
The phenomenon where a policy trained in a simulator loses performance when transferred to real hardware.
embodiment
A robot's physical body configuration, such as a human hand, a gripper, or different arm lengths.
temporal position embedding
Normalizing signals of different frequencies onto the same timeline so the model can align moments.

How to listen

Who it's for

Engineers and researchers working on robot policy, autonomous driving simulation, and embodied intelligence; investors evaluating the deployment path of world models.

Skip

The opening demo and project intro from 0:00-2:28 can be fast-forwarded; technical content starts at 2:28.