AI power concentration is a product of today's technology, not the endgame
Transformers only get smart by being fed massive data and compute, so power sits with the big companies that can afford data centers; but the Transformer isn't even ten years old, and the human brain is proof that small-data expertise exists.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Concentration is a property of the current technical path
Lucas Kaiser attributes AI power concentration to the current state of the technology rather than to the nature of AI: a Transformer only becomes ‘reasonably smart’ when fed the entire internet's data, and once you let it learn only some small domain, it's ‘just stupid’. So what big companies can do is not wait for a research breakthrough but go bigger, bigger, bigger — and bigger costs billions of dollars and requires scraping data from every corner of the internet. That path is inherently concentrating. But he stresses remembering that this is only the current state, and the Transformer isn't even ten years old.
— Lucas KaiserWithout a research breakthrough, you can only pile on scale
Kaiser distinguishes two paths: one is waiting for a research breakthrough that lets models get smart on less data; the other is not waiting and just piling on scale. The problem is that research breakthroughs ‘sometimes they come, sometimes they don't’ — it isn't a business proposal, you can't put it on a roadmap. Piling on scale, by contrast, can be executed immediately, at the cost of billions of dollars and scraping the whole internet. That asymmetry — the executable path happens to be the most concentrating one — is the mechanism behind the current landscape, not big companies deliberately monopolizing.
— Lucas KaiserHumans are the existence proof for small-data expertise
Asked whether there is an algorithmic breakthrough that would let smaller players compete, Kaiser's answer is ‘we know it exists. We are the proof’: humans aren't generalists who know everything, but within their own fields, experts are sometimes stronger than very large-scale models. So that capability must be technically achievable, we just haven't found how to do it yet. He even guesses the answer may still involve an attention layer, because what determines the result isn't only the model architecture but also the loss, the data, the training method — too many variables to know where exactly to look.
— Lucas KaiserOpenAI went from a research lab to a product company
Kaiser observes that the lab's research focus is declining: when he joined, OpenAI was a pure research lab, and now ‘it's a research lab a bit’, but at the same time a big company that has to ship products. He sees this as an opportunity window for the open-source movement and academia — while big companies are distracted building products, the research space is left to universities and independent researchers. He himself recently started doing research independently.
— Lucas KaiserOne consumer GPU beats the old eight-GPU machine
Kaiser bought a 5090 RTX GPU, with more compute than the eight-GPU machine his team used for the Transformer research back then. His conclusion is very concrete: you can't train a big LLM, but you can do research and experiments, and he thinks a lot of people should be doing exactly that. This is the hardest comparison in the whole piece — research that once required a team and a machine room can now be started by one person with one card.
— Lucas KaiserThe optimum may be many distributed models
Kaiser's optimism comes from an analogy: humans are the most astonishing computers in the world, the brain still hasn't been surpassed by models, and humans aren't generalists, nor is there ‘one giant human brain’. From this he speculates that, given a fixed amount of data, the best way to learn may be to have a large number of distributed models — each strong on its own, stronger together. He mentions there are reasons in basic ensemble research supporting this.
— Lucas KaiserData getting expensive will pull research back to fundamentals
Kaiser points out that machine learning research didn't stop in 2017. Because the Transformer is so useful, everyone focused on the big-data path; but now that path has become so expensive that it may instead push attention back to more fundamental research. His judgment: we know a better algorithm exists that can learn from less data and perform well in the domain it learns, we just haven't found how to do it on a computer yet, but we will.
— Lucas KaiserIn the future everyone has their own model and their own sense of humor
The endgame Kaiser paints is very concrete: everyone can have their own model, learning from less data, becoming an expert in different domains like a person, even ‘one tells a joke this way, another tells it another way’. He complains that asking a large language model for a joke today always gets the same answer, always about atoms or something. He attributes this to the stage of the technology — when research catches up, the technology will change.
— Lucas KaiserIn their own words · checked verbatim
the current state of the technology is a bit concentrating, but we should remember that it's just the current state. Sure, transformers are great, but they're not even 10 years old.
Lucas Kaiser1:00
But the big companies are like, okay, but you can do things without research breakthrough, which is to go bigger, bigger, bigger. Now that is very concentrating.
Lucas Kaiser1:00
Like transformers are really good if you train them on the whole internet and then they're reasonably smart. But if you try to say, you know, you just learn about this, then they're just stupid, right?
Lucas Kaiser2:00
Well, I mean, we know it exists. We are the proof, right? Humans are not experts in everything.
Lucas Kaiser2:00
So I bought a 5090 RTX GPU. It has more power than the eight GPU machines we used as a team to do the Transformers research.
Lucas Kaiser3:00
If you look at the world, humans are the most amazing computers in the world, right? Our brains are still unmatched by the models. And we are not generalists.
Lucas Kaiser4:02
So, yes, transformers are great when you feed them all of the Internet, but we know there is an algorithm that's even better, and it can learn from much smaller data and be amazing at the things it's learning. We are the proof.
Lucas Kaiser5:02
It's a step in the technology, but it's a step towards something much more fun where everyone can have their own model, maybe.
Lucas Kaiser6:02
Figures
| GPU model Kaiser bought himself | 5090 RTX | 3:00 |
| The machine used for the Transformer research back then | 8 GPUs | 3:00 |
| Year the Transformer paper was published | 2017 | 5:02 |
Glossary
- Attention is All You Need
- The 2017 paper that introduced the Transformer architecture; Kaiser is a co-author.
- ensemble
- A method of combining multiple models to get a stronger result.
How to listen
Founders and investors watching open-source models, compute costs and the AI power landscape; anyone trying to judge whether concentration is inevitable.
The closing subscription and disclaimer section (after roughly 7:05) can be skipped.