The Middle Layer of the AI Factory: Models Are Becoming a Resource That Has to Be Managed
The least glamorous layer of the five-layer cake is deciding who gets to sell AI to regulated enterprises — because neither the model weights nor the enterprise data want the other side to see them.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Models are no longer applications, they are resources
Hallak says that in the five-layer cake, models don't actually sit on top of the software infrastructure — over time, models have become a resource, just as GPUs are a resource and SSDs are a resource, and resources need to be managed. So what this launch does is model management: acting as an operating system that knows which model is good at writing code, which is good at generating images, which is expensive and which is cheap, and routes tasks to the right model. With only four or five models today that's fine, but once reinforcement learning and fine-tuning scale up, hundreds or thousands of agents are running, and models keep spawning new versions, this gets much harder.
— Renen HallakConfidential computing means neither side has to show its hand
Model developers don't want to hand over weights, and regulated enterprises don't want to hand over data — that's the deadlock blocking AI from entering industries like finance and healthcare. VAST's approach is to have inference happen inside the enterprise's own environment (a physical data center or the enterprise's own VPC) rather than on the model provider's cloud; weights stay encrypted the whole way, and using Nvidia's encrypted memory capability, encrypted data stays encrypted from memory to GPU, accessible only to the inference program. Hallak says the underlying technology isn't new — confidential computing on processors has existed for over a decade — what's new is that it now works on GPUs, and it takes multiple parties in the ecosystem cooperating to assemble a complete solution.
— Renen HallakShare everything, don't shard
Old-style scale-out systems are based on sharding: each node owns a piece, more nodes means more performance, but the number of connections between nodes grows quadratically, returns start diminishing past a hundred nodes, and when one component fails everyone has to participate in recovery. VAST's DASE architecture does the opposite — SSDs aren't attached directly to the local processor but sit on the other side of the network, and with new protocols like NVMe over Fabrics and new media, every node can see all the data as if it were locally attached, so nodes no longer need to talk to each other. Hallak says one customer already has a single cluster at multiple exabytes, tens of TB per second, with tens of thousands of nodes today.
— Renen HallakStorage is where startups go to die
Taking a storage company to VC in 2016 was an uphill battle; storage was considered the least attractive part of the whole stack. Hallak says that became today's moat: when they started in 2016, the underlying components their architecture depended on didn't exist yet and only became available in 2017, 2018 and 2019, so early on the board repeatedly asked him "what if this component never shows up," and his answer was that there was no plan B — which they did not like. Competitors who started 30 years or 3 years earlier couldn't make that kind of bet on the future, and today there's almost no one else at this layer, partly because it isn't sexy enough — building a model company or an application company is far more appealing than building systems.
— Renen HallakZero churn comes from customers being able to leave anytime
Asked whether data gravity makes customers worry about lock-in, Hallak's answer is: because all the interfaces are standard, it's as easy for customers to move off us as onto us, so our job is to keep them happy. Every morning he looks at exactly one chart — the customer satisfaction chart, with only green, yellow and red — and anyone in the company can push a customer from green to yellow or yellow to red, but very few can push them back, and the only metric that really counts is the customer saying "we're happy now." His result: no customer has actively decided to stop using VAST, gross retention isn't 100% only because some customers went bankrupt, and the average customer's data volume on their system doubles or even triples every year.
— Renen HallakAn agent inherits its principal's permissions, and its principal's risk
For enterprise data to be usable by models there are three paths: the context window, the KV cache (the model's memory), and RAG. VAST's approach is to hook into the authentication and authorization mechanisms the enterprise already has — Active Directory, access control lists — first working out who can see what, then extending those attributes to agents: an agent's permissions are inherited from whoever deployed it, or from another agent that spawned it. Beyond data access, you also have to manage tool calls and communication between agents and between agents and people; those conversations flow through the streaming service they provide, so they can be queried and audited. Hallak says most model providers don't let you run inference in your own environment — you can only send them your data and pray — which significantly slows down AI adoption.
— Renen HallakDemand is so large it scares him
Hallak says he wasn't persuaded by demand, he's sometimes scared by it. One AI cloud customer came a quarter ago to do three-year planning and said it would need roughly 500 PB; last week it came back saying it needed another 2 EB; he expects that in another quarter or two they'll come back saying they need double-digit exabytes. What he sees is everything beating expectations across the board: small and large environments alike are going from "we thought we'd need tens of exabytes" to "we're talking three digits." He thinks the current constraint is physical — land, power, chips — with buildout far slower than demand, and some AI clouds have already stopped selling capacity because the next year and a half is sold out. He isn't sure how long this can last, but thinks most industries are only just beginning their migration to AI and that at the current pace it can run for at least another five to ten years.
— Renen HallakNeoclouds win by building early, not by raising early
Asked which of the hundreds of neoclouds will survive, Hallak gives three conditions: get power, get equipment, get money. But what really separates winners is the timing of building capacity — the most successful are those who anticipated demand and built ahead of it, not those who lag demand and rent capacity to catch up. His reasoning is blunt: end users want compute right now; if you have it, they pay, and if you say "let's build together, we'll give it to you next year," they go find someone else. As for the popular argument that "neoclouds are essentially financial instruments for Nvidia GPUs and hyperscalers will eventually win with a lower cost of capital," he thinks reality is the opposite: hyperscalers still haven't overtaken them, which is the innovator's dilemma; the AI clouds actually doing the work learn every day, while the side sitting on cheap capital and not doing the work isn't learning, and the skills gap only widens.
— Renen HallakIn their own words · checked verbatim
Over time, models have become a resource, just like a GPU is a resource or a solid state drive is a resource, and resources need to be managed.
Renen Hallak13:16
You cannot have more than 100 nodes in one of these systems. Uh, to build AI, we need systems with many millions of nodes.
Renen Hallak23:13
I kept telling them that we didn't have a plan B, that they didn't like it, and so, um, I wasn't their favorite person at the time.
Renen Hallak26:13
Our gross retention rate is not 100% because some of our clients have gone bankrupt over the years, but no one has actively decided to stop using Vast.
Renen Hallak32:15
Most of these model builders don't allow you to make inferences in your environment, which means you just have to send them your data and hope for the best. I think this significantly slows down the adoption of AI.
Renen Hallak37:17
I'm not sure if it reassures me, and sometimes it scares me, uh, how big the demand is and how it's accelerating
Renen Hallak47:23
And with each passing day, these AI clouds that are actually in the trenches doing the work, um, they're learning. And the neo-clouders, who are sitting on very cheap funding but not yet, um, hyperscalers, who are sitting on very cheap funding but not yet doing this work, are not learning this. And so the skills gap just, um, keeps getting bigger.
Renen Hallak56:29
So, I think we're going to see more of a difference in the next 10 years than we've seen in the last thousand years, if this continues to develop the way it seems to be developing now.
Renen Hallak1:07:43
Figures
| VAST Data valuation | $30 billion | 0:01 |
| VAST single-cluster scale | multiple EB, tens of TB per second, tens of thousands of nodes | 23:13 |
| Rack power density change | from 10 kW to 500 kW | 2:05 |
| AI growth rate | roughly 10x every 2 years | 52:26 |
| Number of agents used to solve Navier-Stokes | 10,000 | 41:18 |
Glossary
- DASE / Disaggregated Shared-Everything Architecture
- VAST's architecture: SSDs sit on the other side of the network, and every node sees all the data as if it were locally attached.
- KV cache
- The context memory kept during model inference, which determines which GPU already holding that model a request should be routed to.
- DataEnclave
- A confidential computing approach that completes inference inside the enterprise's own environment with weights encrypted the whole way.
- neocloud
- A compute cloud built specifically for AI workloads, as opposed to the legacy stack of traditional hyperscalers.
- NVMe over Fabrics
- A network protocol that lets remote SSDs be accessed as if they were locally attached.
How to listen
Founders and investors focused on AI infrastructure, storage and the data stack, plus enterprise technical leaders evaluating building their own AI factory and worried about data leakage.
16:12 to 20:25 on P vs NP and personal math history — fast-forward.