The world is too loud. Read what matters.

Latent Space

Claude Code and Codex are fundamentally the same architecture: the next windfall is in replacing it

The author argues that mainstream programming agents like Claude Code, Codex, and Pi share nearly identical design choices when deconstructed; the real quantum leap comes from novel harness designs like RLM, which offload context to disk and enable composable sub-agents.

AI agentHarness designRLMGPU kernel optimizationAgent swarm

The video won't play here. Listen to the audio instead:

For readers wanting to understand agent system design rather than tool benchmarks—information-dense but technical, with higher barriers around GPU kernels and open-endedness.

The argument · tap a timestamp to hear it

6:33

Most AI-generated GPU kernels fail to survive production deployment

AI already writes runnable GPU kernels, and most new solutions on the GPU Mode leaderboard come from AI. But the team found a counterexample: Gauners, the community's top human kernel engineer, also uses AI for code but is the only top-10 leaderboard solution remaining stable in real end-to-end systems. Other pure AI solutions have shorter code and higher benchmark scores, but suffer serious "reward hacking"—they fool unit tests but collapse under real system integration. This shows validation methodology is the bottleneck in kernel optimization, and human expert value is shifting from code generation to understanding validation and steering models toward correct directions.

— Alex Zhang
18:55

The PhD advantage is betting on ideas no one believes in

The author argues PhDs have one genuine advantage over industrial labs: they can bet recklessly on ideas with no consensus backing. Examples: SWE-bench was ignored until Devin appeared; work on STaR, Quiet-STaR, and ReAct seemed "simple to the point of obvious" when first circulated; RLM itself was dismissed as "just sub-agents." When an idea is mocked as too simple or too obvious upon release, that's often a signal—showing most people haven't actually thought through its value. The academic problem runs deeper: too many students choose topics to please industrial reviewers rather than betting on unfashionable but story-rich problems.

— Alex Zhang
18:55

Language models need not be autoregressive decoders

Using the contested new model GEV to redefine language modeling's boundary. Critics say it's not a language model because it changes the output space, trading slow autoregressive decoding for fast calibrated binary classification. The author's rebuttal: language modeling means modeling language, not being an autoregressive decoder. Just as recurrent transformers do, any model that reshapes the tradeoff between output space and inference latency opens new design territory blocked by "we already accept the default architecture." Can you recurse only part of the model? Can harness scheduling logic move inside the model itself? These are the real questions about GEV worth discussing.

— Alex Zhang
36:38

Training on short tasks auto-generalizes to tasks 8–30× longer

The hardest-core finding: Train RLM (a harness using only code as a tool, offloading all context to disk) on short tasks, and it auto-generalizes to tasks 8 to 30 times longer—not because it memorized specific answers, but because it learned the meta-strategy itself: decomposition, sub-agent invocation, iteration. That meta-strategy is nearly the same program in both short and long settings; you just change a length variable. More powerful: this recursively applies across domains that seem unrelated. Competitive programming and GPU kernel optimization, the author argues, can share the same meta-strategy—list candidates, spawn sub-agents to judge, validate, iterate. Harness design determines how much generalization you extract from data, not pure data volume.

— Alex Zhang
56:10

Externalizing context to disk is the single biggest innovation

Prime Agent (co-developed with Prime Intellect for RLM) removes the "one-shot sub-agent" assumption in standard agent loops. Sub-agents used to be disposable; Swyx noted this was Codex's biggest pain point, and OpenAI explicitly discouraged long-task use. Prime Agent persists all traces and context to disk; sub-agents can be resumed later and re-engaged in conversation. It also designs a protocol for sub-agents to communicate with each other—all call relationships expressed in code. That combination—externalized to disk plus code-expressed orchestration—is what both sides call the single most powerful trick in the RLM suite.

— Alex Zhang
1:04:20

Burning money buys swarm convergence, not harness design

OpenAI claimed to solve a hard Navier–Stokes variant with "just throw a model at it." Published scale: 10,000 agents, 88 hours, 130 billion output tokens (~$40 million at public rates). Accounting for all context-passing between agents, actual token consumption runs roughly 2× that figure. The author's takeaway: specific harness design details barely matter here—what's nearly impossible and can't be taken for granted is making a large swarm actually converge on a single answer rather than pure token wastage. That's what the money buys.

— Alex Zhang
1:16:59

Kimi's collective intelligence hasn't learned to converge

Comparing Kimi's swarm (able to generate tables, "interesting but solving no genuinely new problems") with OpenAI's (convergent proof generation). OpenAI's earlier "dynamic workflows" failed—expensive and underperformed expectations. What actually worked: OpenAI explicitly trained the model for swarm coordination. Agent convergence isn't free. The gap isn't "whether you have swarm topology" but "whether your agents can truly coordinate toward one answer instead of one-agent-per-task parallelism."

— Alex Zhang
1:20:52

Models are ready for simple reliable work spanning a month

A bold claim: current frontier models, if recast as a capable 18-year-old, could handle straightforward but month-long reliable work. No current harness lets them do it—a "skill issue," not an intelligence gap. Jagged intelligence as a corollary: models are disproportionately strong at coding and math, so in theory should transfer that strength to more "simple but sustained" tasks across domains. The bottleneck is harness design, not model capability.

— Alex Zhang

In their own words · checked verbatim

we found that, like, his kernel was, like, basically the only one in, like, the top 10 that was actually stable in, like, actual, like. end-to-end systems.

Alex Zhang6:33

like, you just have to take big bets, and, like, a lot of them will fail? Like, that’s just. it’s, it’s natural.

Alex Zhang18:55

It’s why is it not just some stupid NLP classifier that, like, we’ve, we’ve been doing, back in our intro ML classes or something? I think what’s really interesting about Jev is that it kind of opens up this question of, are language models correct? Like, in the form that they’re in, can we consider a different design space other than text-to-text?

Alex Zhang18:55

one thing you’ll kind of observe is that the. as you just naively train your RLM on these tasks, they naturally learn to generalize, for example, to longer tasks, because the strategy is basically the same. You’re just modifying, like, a length variable.

Alex Zhang36:38

This is my number one pain with Codex right now. They, their subagents are just very ephemeral And they actively discourage you from using it for long-running things.

Swyx56:10

While there’s a lot in there, I do wanna say, the amount that OpenAI spent is semi-public. It’s, 10,000 agents in 88 hours.

Nothing comes so easily for free. For example, like, the agent swarm design is not something you can just take for granted.

Alex Zhang1:16:59

I think that it genuinely is a skill issue of you can get a model to be as good as, let’s say, like, just some 18-year-old high school kid

Alex Zhang1:20:52

Figures

OpenAI math-proof swarm scale10,000 agents over 88 hours1:04:20
OpenAI swarm tokens and cost130 billion output tokens, ~$40 million1:04:20
RLM task-length generalization8–30× the length of training tasks36:38

Glossary

RLM (Recursive Language Model)
A harness design that uses code as its sole tool, recursively instantiates sub-agents, and persists context to disk and external storage.
GEV
A novel model that modifies output space, using rapid calibrated classification for low latency, challenging whether language models must be autoregressive decoders.
harness tax
The argument that most agent harness design choices in practice have minimal impact on ultimate outcomes.
locally in-distribution
Ensuring each sub-agent invocation falls within distributions the model has encountered, even when the composite task is novel.
continual harness
A framework allowing agents to self-modify their capabilities, sub-agents, and prompts during execution.

How to listen

Who it's for

Engineers designing agent and harness architectures, plus founders and investors tracking RLM, agent swarm economics, and research methodology.

Skip

Open-endedness research and the Sakana Japan market discussion are tangential; they can be skipped.