The world is too loud. Read what matters.

The a16z Show

AI power concentration is a product of today's technology, not the endgame

Transformers only get smart by being fed massive data and compute, so power sits with the big companies that can afford data centers; but the Transformer isn't even ten years old, and the human brain is proof that small-data expertise exists.

Open-source modelsComputeTransformerAI power landscape

The video won't play here. Listen to the audio instead:

An 8-minute solo interview, medium information density, but worth hearing for the Transformer author's argument that concentration is a property of the technology rather than of AI itself.

The argument · tap a timestamp to hear it

0:00

Concentration is a property of the current technical path

Lucas Kaiser attributes AI power concentration to the current state of the technology rather than to the nature of AI: a Transformer only becomes ‘reasonably smart’ when fed the entire internet's data, and once you let it learn only some small domain, it's ‘just stupid’. So what big companies can do is not wait for a research breakthrough but go bigger, bigger, bigger — and bigger costs billions of dollars and requires scraping data from every corner of the internet. That path is inherently concentrating. But he stresses remembering that this is only the current state, and the Transformer isn't even ten years old.

— Lucas Kaiser
1:00

Without a research breakthrough, you can only pile on scale

Kaiser distinguishes two paths: one is waiting for a research breakthrough that lets models get smart on less data; the other is not waiting and just piling on scale. The problem is that research breakthroughs ‘sometimes they come, sometimes they don't’ — it isn't a business proposal, you can't put it on a roadmap. Piling on scale, by contrast, can be executed immediately, at the cost of billions of dollars and scraping the whole internet. That asymmetry — the executable path happens to be the most concentrating one — is the mechanism behind the current landscape, not big companies deliberately monopolizing.

— Lucas Kaiser
2:00

Humans are the existence proof for small-data expertise

Asked whether there is an algorithmic breakthrough that would let smaller players compete, Kaiser's answer is ‘we know it exists. We are the proof’: humans aren't generalists who know everything, but within their own fields, experts are sometimes stronger than very large-scale models. So that capability must be technically achievable, we just haven't found how to do it yet. He even guesses the answer may still involve an attention layer, because what determines the result isn't only the model architecture but also the loss, the data, the training method — too many variables to know where exactly to look.

— Lucas Kaiser
3:00

OpenAI went from a research lab to a product company

Kaiser observes that the lab's research focus is declining: when he joined, OpenAI was a pure research lab, and now ‘it's a research lab a bit’, but at the same time a big company that has to ship products. He sees this as an opportunity window for the open-source movement and academia — while big companies are distracted building products, the research space is left to universities and independent researchers. He himself recently started doing research independently.

— Lucas Kaiser
3:00

One consumer GPU beats the old eight-GPU machine

Kaiser bought a 5090 RTX GPU, with more compute than the eight-GPU machine his team used for the Transformer research back then. His conclusion is very concrete: you can't train a big LLM, but you can do research and experiments, and he thinks a lot of people should be doing exactly that. This is the hardest comparison in the whole piece — research that once required a team and a machine room can now be started by one person with one card.

— Lucas Kaiser
4:02

The optimum may be many distributed models

Kaiser's optimism comes from an analogy: humans are the most astonishing computers in the world, the brain still hasn't been surpassed by models, and humans aren't generalists, nor is there ‘one giant human brain’. From this he speculates that, given a fixed amount of data, the best way to learn may be to have a large number of distributed models — each strong on its own, stronger together. He mentions there are reasons in basic ensemble research supporting this.

— Lucas Kaiser
5:02

Data getting expensive will pull research back to fundamentals

Kaiser points out that machine learning research didn't stop in 2017. Because the Transformer is so useful, everyone focused on the big-data path; but now that path has become so expensive that it may instead push attention back to more fundamental research. His judgment: we know a better algorithm exists that can learn from less data and perform well in the domain it learns, we just haven't found how to do it on a computer yet, but we will.

— Lucas Kaiser
6:02

In the future everyone has their own model and their own sense of humor

The endgame Kaiser paints is very concrete: everyone can have their own model, learning from less data, becoming an expert in different domains like a person, even ‘one tells a joke this way, another tells it another way’. He complains that asking a large language model for a joke today always gets the same answer, always about atoms or something. He attributes this to the stage of the technology — when research catches up, the technology will change.

— Lucas Kaiser

In their own words · checked verbatim

the current state of the technology is a bit concentrating, but we should remember that it's just the current state. Sure, transformers are great, but they're not even 10 years old.

Lucas Kaiser1:00

But the big companies are like, okay, but you can do things without research breakthrough, which is to go bigger, bigger, bigger. Now that is very concentrating.

Lucas Kaiser1:00

Like transformers are really good if you train them on the whole internet and then they're reasonably smart. But if you try to say, you know, you just learn about this, then they're just stupid, right?

Lucas Kaiser2:00

Well, I mean, we know it exists. We are the proof, right? Humans are not experts in everything.

Lucas Kaiser2:00

So I bought a 5090 RTX GPU. It has more power than the eight GPU machines we used as a team to do the Transformers research.

Lucas Kaiser3:00

If you look at the world, humans are the most amazing computers in the world, right? Our brains are still unmatched by the models. And we are not generalists.

Lucas Kaiser4:02

So, yes, transformers are great when you feed them all of the Internet, but we know there is an algorithm that's even better, and it can learn from much smaller data and be amazing at the things it's learning. We are the proof.

Lucas Kaiser5:02

It's a step in the technology, but it's a step towards something much more fun where everyone can have their own model, maybe.

Lucas Kaiser6:02

Figures

GPU model Kaiser bought himself5090 RTX3:00
The machine used for the Transformer research back then8 GPUs3:00
Year the Transformer paper was published20175:02

Glossary

Attention is All You Need
The 2017 paper that introduced the Transformer architecture; Kaiser is a co-author.
ensemble
A method of combining multiple models to get a stronger result.

How to listen

Who it's for

Founders and investors watching open-source models, compute costs and the AI power landscape; anyone trying to judge whether concentration is inevitable.

Skip

The closing subscription and disclaimer section (after roughly 7:05) can be skipped.