Runway Bets 1,000 A100s: World Models Skip the Digital Twin
Runway signed for a 1,000-A100 cluster while still at Series B, betting the video scaling law would hold; now its world model can roll out policies from a single photo of an environment, bypassing the old path of building a digital twin for every task.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Content will eventually be generated, so tools must be rebuilt
Runway's starting point was not building a tool but an extrapolation: after seeing early generative models in 2016 and 2017, they assumed resolution and quality would improve predictably over time, and concluded that ‘most content will eventually be generated’, so creative tools had to be redesigned. Anastasis calls this ‘a matter of when, not if’. The thesis later proved to apply far beyond creative tools themselves — once they built a research organisation they found the models were equally useful in real-world settings, and so they came full circle back to the starting point.
— AnastasisIt signed for 1,000 A100s while still at Series B
In mid-2022, Runway judged that the scaling law would act on image and video generation the way it acts on language, and signed an agreement to build a 1,000-A100 cluster — at the time the company was still at Series B. Anastasis himself calls this ‘slightly irrational’. The goal they set was: what does the video version of the latent diffusion / Stable Diffusion moment look like. The best video model at the time was CogVideo, only 256×256 and very poor quality. They did video-to-video before text-to-video because the conditioning is stronger and the problem easier, and Gen-1 shipped in January 2023.
— AnastasisGen-2 was a two-stage pipeline hacked together over a weekend
Gen-2 shipped before Gen-1 was generally available, only two months apart. The reason was that text-to-video simply could not be done directly, so Anastasis stitched two stages together in a weekend project: first have the model generate a depth map from text, then use Gen-1 to turn the depth map into RGB. He admits this two-stage pipeline makes the video structure occasionally look wrong, because depth has to come out first. Interestingly the idea later came back: Reve's text-to-image uses a planner model to generate bounding boxes first and then feeds them to a diffusion transformer, and Ideogram shipped the same innovation on the same day.
— AnastasisCamera control was the seed of the world-model research
Anastasis says camera control was the first time it felt like you were not ‘generating a video’ but ‘moving through a world’ — and that became the seed of the world models research direction. Runway set up a dedicated research group for world models at the end of 2023. Their argument: to predict video well you must simulate the world with ever greater capability; if the scaling law holds for video, the model will become ever more able to predictably simulate physics, human action and dynamics. He pushes back on ‘generating cute videos is not the same as understanding the world’: video models can hide physics flaws behind continuous cuts, so don't be fooled by current performance.
— AnastasisThree months after Sora it scaled the model up 10x
Gen-3 arrived in 2024, a few months after Sora launched. One of Sora's biggest changes was replacing the convnet with a diffusion transformer, and Runway realised it had to invest in model parallelism and distributed training infrastructure, which it spent the whole autumn of 2023 building, with plenty of false starts and failures along the way. Sora came out in February 2024 and Twitter was full of ‘Runway is finished’. Anastasis says those few hours were an existential crisis, but the team spent three months scaling the model size and training compute up 10x, produced a better model than Sora, and shipped Gen-3 that summer.
— AnastasisThe interface world model: generate pixels directly
Runway's interface world model replaces the entire software front end: no HTML, CSS or React, the interface pixels are output directly by a real-time video generation model, with clicks, drags and scrolls as inputs. You can use a prompt to describe the behaviour of each element — what should happen when this button is pressed. It is also a video-audio generation model, able to describe the visual result and the sound effect after a click. Anastasis's phrasing is ‘why generate code that produces pixels? Generate the pixels’. Right now it is more expensive, because the compute needed to run a real-time video model is far higher than rendering HTML.
— AnastasisWorld models skip the digital twin
Traditional simulators need to describe rigid-body physics precisely, and Isaac Sim or MuJoCo can do that, but they struggle with complex manipulation such as cloth or slippery surfaces, and you have to build a digital twin for every environment and task, which is extremely time-consuming. A world model only needs the first frame and you can roll out policies inside it — take one photo of an environment and you can test policy performance. Runway used the same scenes, settings and embodiment from the Roborina benchmark to measure how well action-model performance in the world model correlates with performance in the real world, and got very good correlation. This is the first signal that their model is useful for robotics.
— AnastasisThird-person video data is the biggest source of all
Anastasis sorts robot data into a few categories: teleoperation data, humie data where humans do tasks with a robot arm, and egocentric data from a GoPro on someone's head. He says all of them are useful, but there is three orders of magnitude more third-person video data in the world than egocentric data. Humans also learn to do things mainly by watching others, not first-person. So their argument is: after video pretraining you only need a few hundred hours of robot data to fine-tune, whereas pretraining a robot model from scratch takes on the order of 100,000 to a million hours. He states plainly that ‘the most abundant data source will eventually win’.
— AnastasisIn their own words · checked verbatim
There’s gonna be a point where most of content will be generated, and that was maybe the initial thesis of Runway was we will need, as a result of those generative models, rethink how creative tools are made.
Anastasis0:00
we signed this deal to build a cluster of a thousand A100s, which at the time we were a Series B startup. That was a almost, slightly irrational decision maybe, but we really believed that if we trained a video model at a large scale, we would get, like, a great model at the end.
Anastasis8:12
There were a lot of, a lot of chatter on Twitter about Runway. Runway’s done. like, there is no way Runway will catch up.
Anastasis34:58
It’s the end-to-end philosophy applying applied to front ends.
Anastasis45:18
the most plentiful source of data will ultimately wins.
Anastasis1:01:55
I call it the lucid dream test.
Anastasis1:07:47
Like one nice thing about image and video models is you can immediately tell with your eyes like what feels good from an aesthetic standpoint.
Anastasis1:23:12
And now my sense is increasingly people are gonna care about latency more and more. As those models become better, the ability to iterate very quickly becomes more important.
Anastasis1:26:08
Figures
| Gen-1 release | January 2023 | 8:12 |
| CogVideo resolution | 256×256 | 8:12 |
| Gen-3 scale and compute increase over previous models | 10x | 34:58 |
| character model autoregressive video generation length | up to 30 minutes | 53:40 |
| GWM Worlds open-ended world exploration generation length | on the order of a few minutes | 53:40 |
| Grok video context | 10 to 20 seconds | 53:40 |
| Genie generation length cap | up to 1 minute | 53:40 |
| Max number of references a model can accept | 50 | 1:24:48 |
| Runway film festival frequency | every May or June | 1:35:45 |
Glossary
- lucid dream test
- If you cannot tell whether what you are looking at is real pass-through or generated rendering, the world model is good enough.
- step distillation
- Compressing a diffusion model's multi-step sampling into few steps, to cut serving cost and speed up generation.
- error accumulation
- In autoregressive generation, each frame's small error is fed back into the model and keeps amplifying over time.
- counterfactual generation
- Making the model generate an outcome that did not happen, such as a goal not being scored, used to simulate failure and for online RL.
How to listen
Founders and engineers working on video generation, world models and robot foundation models, plus investors who want to understand Runway's technical roadmap and compute bet.
After 1:33:40, the discussion of AlphaFold-style specialised architectures is fairly general and can be skipped.