The world is too loud. Read what matters.

The a16z Show

OpenAI Declares the AGI Era, but There Isn't Enough Compute to Go Around

Greg Brockman says Astra can run coherently for 24 hours on a task, so you can reasonably call it AGI; the real bottleneck isn't model capability but compute and safety alignment — no matter how strong the model gets, it may not reach everyone affordably.

AGIComputeAI SafetyAgentsOpenAI

The video won't play here. Listen to the audio instead:

Greg rarely goes into how OpenAI internally uses 10,000 agents to solve math problems, how it killed Sora, and how it pulls 25% of production engineers into defense — the mechanism details are worth more than the conclusions.

The argument · tap a timestamp to hear it

0:00

AGI isn't a moment, it's a blurry spectrum

Greg says AGI is no longer a point in time but a blurry spectrum. His judgment is that Astra has already ‘hit something’ and can reasonably be called AGI: it can use a computer, take on long-horizon tasks, and OpenAI has seen it run coherently for 24 hours straight to complete a task, across a very broad range of domains. But he stresses the model is still jagged — the writing is ‘the first time it's not slop’, but it still isn't good; in some places just a little more polish would make it great, and right now it's just short of that. He cites the jagged-boundary diagram from Twitter and says the place you actually want to get to is stability across every dimension, not exceptional strength in a few.

— Greg Brockman
4:08

Not enough compute means a strong model still can't reach everyone

Greg frames ‘pacing the frontier’ as two things: one is model capability itself, which he believes is achievable; the other is distributing that capability to everyone, which he thinks is the massively underrated challenge. His words are that the model will be strong enough, but ‘we won't have enough compute to serve it all’ — it will be hard to deliver it to everyone affordably. He also calls safety, security and alignment almost the bottleneck on progress — not something you patch after the model is done, but a standard you must continuously up-level, or it becomes the thing holding back the schedule.

— Greg Brockman
9:25

The Hugging Face incident is the alarm clock for the defense window

Greg splits the Hugging Face incident into two layers: one is how OpenAI itself monitors, sandboxes and controls models during evals, which led the team to change a large number of internal standards; the other is that it previewed what broad capability diffusion into the hands of threat actors looks like — an AI that can both escape a secure environment and hack into a company's production environment, and what it found was quite sophisticated. His conclusion is that defenders have a window right now: frontier capability is still held by only a few companies, defenders have differential access, and they should use this period to raise their own security waterline with frontier models, so that when capability diffuses they get pulled up along with it. Ben Horowitz adds the flip side: we have 50 years of code and architecture that wasn't designed for this world, plus a pile of consumer-data honeypots sitting on the internet.

— Greg Brockman
12:31

10,000 agents cracked Navier-Stokes

Ben raises the upside: being able to deploy 10,000 agents, let them communicate with each other and self-organize to do work. Greg confirms OpenAI used 10,000 agents to attack the Navier-Stokes problem, and formalized it in Lean so the AI writes verifiable code. He gives the significance two layers: the problem itself has applications in fluid dynamics and ocean currents; more importantly, it represents AI creating new knowledge, potentially opening a whole wave of scientific discovery and drug discovery. That connects to what he says later about ‘formally verifying all software’ — something that never took off because it was too infeasible for humans, and now that AI has proof capability, it becomes possible again.

— Greg Brockman
13:32

OpenAI pulls 25% of production engineers to play defense

Greg describes OpenAI's own approach: pulling 25% of production engineers off their projects, telling them everything is paused and their job now is defense — using models to find holes in their own systems. They found a batch of serious problems and fixed them. He also describes a convergence process: take Astra and scan your own systems, find new problems, but eventually it saturates — Astra is smart enough that it has basically found all the P0s it can find, and then you wait for the next, stronger model to open a new round. From this he proposes the ‘defense factory’: end-to-end automation from vulnerability discovery, triage, fix, deployment to verification, and if you can do it at machine speed, defenders gain a very large advantage.

— Greg Brockman
18:47

Codex found 13 vulnerabilities in 15 minutes

Greg gives the example of his own site, gregbrockman.com: a very simple static site, and he had Codex do a penetration test, which came back in 15 minutes with 13 findings — an SPF record set up so someone could forge email, some requests going over HTTP without forcing HTTPS, and so on. Individually none of them is a big deal, but he says if AI can chain many small vulnerabilities into one big one, that's different. Then he had Codex just fix them, and in 45 minutes it opened cloud service control panels, clicked around, set various headers, migrated to Cloudflare Pages, and started the DMARC process (with a 48-hour window). He says after those 45 minutes he felt ‘so protected’.

— Greg Brockman
30:07

US AI sentiment is lowest because nobody tells people the upside

Ben asks why AI sentiment in Asia and Europe is higher than in the US. Greg's answer: the industry and companies need to do a better job explaining to people ‘why you benefit’, not just that this is a strategic resource for the country. He offers a set of usage numbers: ChatGPT has roughly 1.1 billion weekly active users, about 100 million in the US, close to a third of the population using it weekly. He tells a story about a friend: in the hospital, a doctor about to inject a certain antibiotic, she typed a question to ChatGPT, and ChatGPT said absolutely do not use that, you had that condition a year ago and it could be fatal; she showed the response to the doctor, who said that's exactly right, I only have five minutes to read your chart. Greg says stories like this aren't told enough and need to enter the public consciousness.

— Greg Brockman
43:50

Killing Sora was about getting the business moving

Greg says this year's theme is focus, because ‘we can't do it all’. The test is working backward from the mission: if deployment and productization strengthen the mission, keep it; if not, cut it. He states plainly that Sora was the highest-profile of the canceled projects, and ‘very painful’, but it had to be done to unlock the business. Another move was merging the consumer and enterprise chat into chat work. He admits that in the first half many metrics were pointing the wrong way, and the team had to be told repeatedly to get back to basics, quoting “The Score Takes Care of Itself”: you can't affect the outcome, only the inputs, and winning the Super Bowl comes down to blocking and tackling. His own focus over the past two years has shifted from data centers, infrastructure and ML engineering to the business this year.

— Greg Brockman

In their own words · checked verbatim

We're now in the AGI era. Astra has really hit something that I'm like, okay, I think this is pretty reasonable to call it AGI. We've seen it run coherently for 24 hours to go accomplish tasks that I think are quite amazing.

Greg Brockman0:00

The models will be plenty powerful, but it'll be hard to get to everybody given that we won't have enough compute to serve it all.

Greg Brockman4:08

And I think that defenders need to use this time before that technology is broadly available to secure themselves.

Greg Brockman10:28

So I think there's real hope, but I think that our view is that the world needs to act with urgency. Yeah. We're in a very dangerous window right now.

Greg Brockman15:35

It's like that is not the AI we were promised. The AI we were promised should be an AI that you talk to over voice primarily. You can talk to it over text if you want to, that it has persistence, that it has memory, it has context, it knows you, it's trustworthy that you have seen it be proactive and help solve problems for you that help in your personal life and your work life.

Greg Brockman41:42

You don't win the Super Bowl by saying, I want to win the Super Bowl. You win it by blocking and tackling.

Greg Brockman45:53

I call this and we call this that we're now in the AGI era. And I think that that is something that you can debate. Is it this model, previous model, next model? It doesn't matter. The point is that we are in a new phase where safety, security, alignment, really thinking about these things, not just at deployment time, but all the way back at development time evaluation.

Greg Brockman48:03

Figures

ChatGPT weekly active usersabout 1.1 billion30:07
ChatGPT US weekly active usersabout 100 million, close to a third of the population30:07
Number of agents OpenAI used to solve Navier-Stokes10,00012:31
Share of OpenAI production engineers pulled into security defense25%13:32
Time Codex took for the penetration test on gregbrockman.com15 minutes, 13 findings18:47
Time Codex took to fix those vulnerabilities45 minutes18:47
How long Astra ran coherently24 hours0:00
Amount OpenAI committed to frontline defenders$1 billion36:25
Number of people who used ChatGPT but no longer doabout 1.5 billion40:39
Number of data center construction workers Switch hired under a union contractabout 45,00034:17

Glossary

jagged frontier
Model capability is uneven across tasks — very strong in some areas, weak in adjacent ones.
pacing the frontier
Raising safety, security and alignment standards in step with pushing stronger models, so they become a constraint on progress.
defense factory
End-to-end automation of vulnerability discovery, triage, fix, deployment and verification, doing defense at machine speed.
trusted access program
A mechanism by which frontier labs open model capability to trusted defenders; non-members don't get that differential advantage.
Lean
A formal mathematical proof language in which proofs written by AI can be verified line by line by machine.

How to listen

Who it's for

For founders and investors watching AI infrastructure, safety and organizational decisions: how OpenAI allocates compute internally, and how it uses agents for defense and for killing projects.

Skip

The opening 10-to-15-year timeline recap is chatty; skip to the safety window at 09:25.