The world is too loud. Read what matters.

Training Data

The Model Is the Refinery; the Value Is in the Chemical Plants and the Car Companies

In business history, the companies that break out in the first wave are almost never the winners of the next stage. Model companies are refineries that distill base oil; the real money is in the application layer that wires intelligence into the workflows of banks, pharma and government.

Enterprise AIApplication LayerOpen WeightsAgentsAI Diffusion
Levie breaks the "application layer vs. model layer" bet into concrete mechanisms: why subsidies must end, where open weights will win, and why coding diffuses fast while other knowledge work diffuses slowly.

The argument · tap a timestamp to hear it

2:02

Between the model and the workflow lies a trillion-dollar gap

Levie says the market's past two years of dismissing the "model wrapper" were wrong, because what enterprises actually need is the bridge from model capability to real workflows. That bridge is either very limited or very large — roughly a trillion dollars rides on which of those two outcomes it is. He judges the bridge is obviously large: however smart the model, the workflow still has to connect to other data systems, keep a human in the loop, wait on latency, and deal with legacy systems — five to ten things that are all operational grunt work, and the last thing a research-oriented organization wants to do.

— Aaron Levie
5:03

The infrastructure-and-application script has already played out once

Levie uses the cloud analogy to rebut "models will eat everything": AWS, GCP and Azure created trillions in market value, but there are also trillions in software value that exists only because of them. He says that if you had looked ten years ago at what AWS and GCP were building, you would never have predicted Snowflake or Databricks would appear — you would have thought the infrastructure already did this, so why pay ten billion dollars of revenue for a product that "just lets you process data." Intelligence will take the same path: models will be extremely valuable, but bringing models into the workflows of banking, life sciences, healthcare and government is itself a lot of software.

— Aaron Levie
9:07

Subsidizing tokens cannot go on forever

Levie thinks "the fox guarding the henhouse" is the structural reason the application layer must exist: when you hand a task to an agent system, you want the cost optimized with accuracy unchanged, which by definition should be done by a company that does not care which model is used. The countervailing force is that model companies can subsidize their own models, but he judges that cannot last — once these companies go public, they are bound by the same laws of capitalism; and the subsidies may already have crossed a threshold once they reached the scale of tens of billions in capex, and in the end they still have to pay for training runs. He adds an equilibrium problem: if the API is a high-margin business subsidizing the application product, and all the customers get migrated onto the application product, the API revenue is gone.

— Aaron Levie
18:42

Code is a tool; content is a proxy for character

Levie splits apart the difference in tolerance for "work slop": for most people in the world, code is a utility — it just has to automate something, let someone press a button to get to the next step — so having an agent generate an entire backend or frontend is not only acceptable, it is preferable. Content is different — when you receive a deck, you are still judging "can I trust this person to execute on this," and work slop destroys that judgment. He admits he uses AI for brainstorming himself, but when he receives someone else's AI output he is equally suspicious; it is a collective trap, and it may take three to five years to get through it.

— Aaron Levie
29:25

Open-weight adoption is higher than people think, lower than enterprises want

Levie gives a three-part judgment: open-weight adoption in enterprises is "higher than people think, lower than enterprises actually want, and much lower than it will be in five years." He estimates that 30% or more of it is attributable to the novelty of "wanting to try GLM," not to cost — he has seen Fortune 500 CIOs say they are playing with open source, but he knows that at that cost tier Gemini or Muse would have been good enough anyway. The actual obstacles today are concrete: sometimes token efficiency is lower, and sometimes it randomly starts speaking Chinese in the middle of a chain, which is very strange for a bank. Long term, mature use cases will be peeled off onto open weights, provided it really is cheaper, or post-training can deliver extra performance.

— Aaron Levie
32:34

Burning context into the weights runs into the enterprise permissions wall

Levie is interested in continual learning, but points out that the real friction in the enterprise world is permissions and access control. He says researchers often imagine the world runs their way — I have full access, so a model trained only on my world would be great — but a lawyer has an extremely narrow access point to five projects, because a colleague one door down is working on a competitor's project, and no document can pass between the two; it has to be a hard partition. Even if you could train a model for a single user, what about the fact that this person is added to or removed from projects every day, bringing in new and important context? He thinks this only gets truly interesting once the cost curve comes down and open weights get smaller and faster.

— Aaron Levie
39:29

A system of record must do both ends; doing one end loses

Levie says that if you are a SaaS platform with data and customer workflows, there are two things you must do at the same time, and doing only one means you lose. First, you have to have an agent that is absurdly strong on your own product, provably 10 to 20 points better than a general agent — not by crippling the competition, but by maxing out evals and tuning for the workflow. Second, you have to be headless, exposing APIs to Claude, ChatGPT and every other platform, letting external agents call you through MCP or deterministic APIs. He gives his own example: because he can get into Salesforce through Claude or ChatGPT MCP, he now uses Salesforce roughly ten times as often as before.

— Aaron Levie
48:00

Coding diffuses fast because it hits all six conditions

Levie lists six properties of why coding ran out ahead: the value of code is almost 100% represented by the generated text; models are trained on massive amounts of code; every AI lab treats coding as a competitive benchmark and runs its own evals daily; the users are the most technical population in history, so when they hit a bug or an MCP connection failure they fix it themselves instead of calling; and it is a high-paying vertical, so a 10% to 20% productivity gain automatically has value. By contrast, a sales rep's value creation is persuading an external customer to buy software or a Caterpillar truck, constrained by whether the customer replies, whether there is budget, whether they can meet next Tuesday — this is a continuum between "rate-limited by external factors" and "able to sit at a computer and type all day."

— Aaron Levie

In their own words · checked verbatim

I guarantee you would not have predicted Snowflake or data bricks existing you would have been like the infrastructure just already does that why would you pay another $10 billion of revenue to all these other products that are just making it so you can work with your data.

Aaron Levie6:05

I don't know how long that lasts though because when when these companies become public, I think they will be held to basically the same laws of of capitalism that everybody else is.

Aaron Levie10:07

For most of the world, code is a utility.

Aaron Levie20:12

I'm doing work slop for some of my you know brainstorms and decisions but when I get it from somebody else I'm like hm should I trust you?

Aaron Levie21:14

sometimes like randomly like I've heard stories like randomly it'll just like speak Chinese like like mid midchain so you're like okay well that'll be weird for a bank.

Aaron Levie30:24

in five years from now I would bet like 90% of all tokens in the enterprise are things that a user never kicked off and they just see a result.

Aaron Levie47:32

You're either like a year ahead or a year behind simply based on your feed and my feed is like so so wired in.

Aaron Levie55:39

Figures

Capital riding on the "model-to-workflow bridge"roughly one trillion dollars2:02
Box annualized revenue1.3 billion dollars52:37
Number of files in Levie's personal Box accounttens of millions25:19
Number of tests Box runs on each modelhundreds26:19
Levie's estimate of the share of open-weight adoption attributable to "novelty"30% or more29:25
Number of customers Levie talks to directly each yearabout two hundred40:30
Number of questions Levie asks AI systems each day20 to 3057:39

Glossary

work slop
Work output mass-generated with AI and lacking any trace of personal thought, such as board materials or presentations.
harness
The engineering layer that wraps the model, connects it to search and the file system, and orchestrates multi-step retrieval and reranking.
continual learning
The direction of letting model weights continuously adapt to a specific user or organization through use, rather than relying only on context retrieval.
open-weight model
A model whose weights are publicly downloadable and deployable, so enterprises can host it themselves to push down inference costs.
MCP
A protocol that lets external agents call a given system's API in a standard way.

How to listen

Who it's for

Founders building enterprise software and AI applications, investors betting on the application layer, and CTOs agonizing over whether they should train their own models.

Skip

The opening jokes about root canals and Doug Leone, roughly 0:30 to 1:55.