The world is too loud. Read what matters.

The a16z Show

Video generation doesn't lack speed anymore — it lacks the director's control

FAL post-trains open-source video models to be an order of magnitude faster and an order of magnitude cheaper than the originals, then concludes: speed is no longer the bottleneck — what's next is controllability of camera, lighting, character and motion.

Video generationPost-trainingInference optimizationHollywoodControllability

The video won't play here. Listen to the audio instead:

The first half is hardcore inference-system optimization detail, the second half is Hollywood workflows and a controllability roadmap — high information density, good for forming a judgment.

The argument · tap a timestamp to hear it

3:02

Open weights are the precondition for post-training

FAL had run inference for other model labs before, but never had permission to layer capabilities onto the model. Minimax H3 was the first video model that was genuinely open-source and belonged to the latest generation, could accept reference images, and had an architecture similar to other video models — that's what made FAL decide to go all in on post-training. Gorka says the biggest reason all these optimizations could come together this time is that H3 is the first genuinely next-generation open-source video model — no weights, no post-training business.

— Gorka Mirdsevin
4:02

The whole industry has been compute-constrained since April

Gorka offers a criterion he calls token market fit: whether one person can productively consume a large amount of tokens, on the order of 10K. Demand for video generation is big enough — people in this line of work sit at their computers generating all day and spend thousands of dollars. But since April, the whole industry and FAL itself have been compute-constrained: add however much compute, grow by that much. So FAL keeps hunting for efficiency, releasing the saved compute to other models, or letting H3 Max generate more tokens.

— Gorka Mirdsevin
5:04

System optimization hits a roofline; post-training punches through it

Batuan says the system-level optimization of the past three or four years can at most make the same model 2 to 3 times faster, because the architecture and constraints haven't changed — you're just squeezing the chip drier, and that roofline exists. New post-training plus system-model co-design can cross that roofline by an order of magnitude. FAL had practiced on open-source image models before, doing versions of ideogram and flux, and built up post-training infrastructure experience; this time, adding a frontier video model, speed went up by an order of magnitude.

— Batuan Tashkaya
7:10

Raise quality first, then cut steps — the order can't be reversed

Compressing a diffusion model from 50 steps to 20 loses quality, so FAL's approach is to first run post-training and the L pipeline to raise model quality, then layer on the optimization stack, so the final result is equal or better in quality while being an order of magnitude faster. Beyond post-training, there are kernels and systems engineering pulling hardware utilization from the 30-40% typical of inference workloads up to a theoretical MFU of 70-80%. And a video model isn't a single call — underneath is a pipeline: first a large LLM expands the prompt, then generation happens in latent space, then decoding back to pixels, plus possibly upscaling depending on the load, and by default none of these components is optimized.

— Batuan Tashkaya
10:19

The Turbo version does a five-second video in 1.5 seconds

A week after H3 Max launched, the team found it could run at 2x speed with quality still in the 97th percentile, so they released H3 Max Turbo: generating a five-second video takes just 1.5 seconds, at half the cost. Gorka's judgment is that today these models are already cheap enough and fast enough — an order of magnitude cheaper than frontier models and more than an order of magnitude faster — and users don't need faster or cheaper. His bet is that what comes next is raising quality, and controllability — which is exactly the direction of the past two or three weeks.

— Gorka Mirdsevin
16:43

Continuous video runs on two minutes of memory, not the last frame

The livestreams Rehan and levels io each made are essentially independent clips: one segment ends, its last frame is fed to the next segment, and the second segment remembers nothing but that last frame. FAL's internal ML team built something else: transitions are more seamless, there are two minutes of memory, when someone walks into the room everyone looks at him, and the scene is continuous. It remembers a highly compressed representation of the raw video, and beyond two minutes a continuously evolving system prompt maintains overall coherence, so it can remember the last four to eight scenes and stream indefinitely.

— Gorka Mirdsevin
29:00

Blender plus a video model equals near-100% controllability

One of the most popular approaches in professional workflows: first render a low-resolution scene with a non-AI technique like Blender, then feed that video as a reference to the AI model, and you basically get near-100% controllability. A week after H3 Max launched, people started generating scenes in Blender with GPT Astra and handing them to H3 Max, because it's fast enough to try many in parallel, opening up an entirely new pipeline. Gorka says the next month or two will be all-in on controllability, so studios and professional creators can get content that fits their use case exactly.

— Gorka Mirdsevin
33:03

Hollywood wants small point solutions, not generation from scratch

Gorka says Hollywood is FAL's fastest-growing segment, with usage a year ago at nearly zero. The NARA tool Amazon MGM Studios released runs mainly on FAL's infrastructure. What Hollywood wants isn't to generate everything from scratch, but to extend a video a bit, change camera control, change lighting — these small point solutions. There's a disconnect between what research labs are building and what creators actually need, and FAL wants to close that gap with post-training projects. The other half of the problem is legal and data residency; Seedance now has US hosting too, and Gorka thinks Hollywood's AI usage will grow 10x, 100x in the coming months.

— Gorka Mirdsevin

In their own words · checked verbatim

the biggest reason why everything came together for this particular moment was because H3 was the first truly next generation video model that's open-source.

Gorka Mirdsevin3:02

since around April, the whole industry and FAL itself, we've been compute constrained. We are growing as much as we are adding compute.

Gorka Mirdsevin4:02

this new set of post-training-related optimizations with system-slash-model co-design enables us to go beyond that roof line by an order of magnitude.

Batuan Tashkaya5:04

we have a version called H3 Max Turbo that's public that can generate like a five second video in like 1.5 seconds, which is like insane.

Gorka Mirdsevin10:19

my bet today is we just need to improve quality more than like the speed at these speeds

Gorka Mirdsevin11:25

everyone's waiting for a large consumer moment in AI. Now it's like good enough and cheap enough that like a truly novel social AI experience can be built on top of it.

Gorka Mirdsevin25:55

Hollywood is our fastest growing segment. And there's a lot of noise about how AI might disrupt Hollywood, but Hollywood usage was non-existent a year ago.

Gorka Mirdsevin34:03

Figures

H3 Max speedup over the original Minimax H3 endpoint35x6:09
Hardware utilization improvement rangefrom 30-40% to 70-80%7:10
Time for H3 Max Turbo to generate a five-second video1.5 seconds10:19
H3 Max Turbo cost reduction2x less10:19
Maximum continuous video length for H3 Max Directorup to 60 minutes19:47
H3 Max Director memory window2 minutes18:45
H3 Max usage lead on the FAL platformmore than twice the second place22:54
Hardware improvement from Hopper to Blackwell2 to 3x9:15

Glossary

token market fit
FAL's definition: whether one person can productively consume a large amount of tokens, on the order of about 10K.
MFU
A metric for how much of a chip's compute is actually utilized; 70-80% is already near the theoretical ceiling.
roofline
Under the same model architecture, the performance ceiling that system optimization can squeeze out.
LoRa
A small fine-tune layered onto a base model, used to add specific capabilities like camera angles or styles.
VAE
The component in a video pipeline that decodes the latent representation back to pixels.

How to listen

Who it's for

Engineers and founders working on video generation, inference infrastructure or AI tooling, plus investors watching the intersection of Hollywood and generative media.

Skip

The first two minutes of host introductions and the podcast promo at the end can be skipped.