The world is too loud. Read what matters.

The a16z Show

AI Coding Tools Code Faster, But Software Itself Hasn't Gotten Smarter in a Decade

Coding agents let people code faster, but they produce the same kind of software. Diogo Almeida argues what we should really do is make AI a new primitive in software, automating things that previously couldn't be automated.

AI codingCoding agentsSaaSAI reliabilityAGI narrativesProbabilistic programming

The video won't play here. Listen to the audio instead:

A founder shares his journey from early-stage AGI disillusionment at OpenAI to building Jev, clarifying the fundamental difference between AI-native programming and traditional coding agents.

The argument · tap a timestamp to hear it

3:02

Coding agents and Jev solve fundamentally different problems

Cloud Code and Codex represent "just in time software"—writing code in natural language produces output equivalent to code written by humans, with unchanged expressiveness. Diogo's vision isn't faster coding but a new AI primitive embedded in software itself: not automating "software engineering" as a job, but automating tasks that should be automated but currently aren't. That changes what software itself can do.

— Diogo Almeida
7:11

Jev is a classifier, but that's not a limitation

Diogo explicitly describes Jev as a classifier—this is precise technical language, not dismissive. It's more akin to a library: you describe desired behavior in natural language, pair it with a state machine, and Jev selects branches with confidence scores. He estimates current Jev likely already outperforms what a traditional ML engineering team in 2019 would ship after months of dataset collection and model tuning—and ordinary developers can now "write" this capability on the spot.

— Diogo Almeida
12:19

We're building products, not creating deity

Ben Horowitz cited TypeSafe's phrase "we build prod, not God" as its most compelling insight. Other large-model labs, even privately, revel in the weighty narrative that "we are creating potentially uncontrollable technology." TypeSafe inverts this—AI optimism is simply put on the table. Diogo traces much of the industry's pessimism to the "single brain rules everything" myth: if you believe only one ever-larger centralized model represents AI's future, it's hard to grasp why developers get excited about a programmable AI primitive.

— Ben Horowitz
18:25

He believed GPT was AGI until his entire worldview collapsed

Diogo participated in releasing early RLHF models at OpenAI and hadn't anticipated instruction-tuning's generalization strength. The team deliberately tested with ‘why should you eat socks before meditating’—a question guaranteed never to appear on the internet—and the model produced sensible, human-like answers. He genuinely believed this model probably was AGI. When it wasn't, he describes it as a moment his ‘entire worldview collapsed,’ and the inflection point toward his later conviction that AI should be genuinely useful, not merely appear intelligent.

— Diogo Almeida
19:26

Disbelieves recursive self-improvement but sees AGI through a different lens

Diogo explicitly states that for various complex reasons, he doesn't believe AI follows a recursive self-improvement path—neither in the past nor now. However, he does endorse OpenAI's simpler definition of AGI: automating most economically valuable work globally. This is entirely feasible since much valuable economic work is actually basic and repetitive. The real problem emerged post-RLHF: industry focused energy on satisfying a "human rater" proxy metric. GPT-3 was initially well-calibrated, but with humans scoring outputs, the model increasingly produced eloquent text without proportionally improving at actual task automation.

— Diogo Almeida
24:39

Customer service stalls at the long tail, not at intelligence limits

Diogo points to OpenAI's customer service automation efforts since 2020 as a counterexample: intelligence already suffices—so why hasn't it worked? Martin offers a precise observation from early customer service company data: they'd claim ‘we resolve 95% of tickets,’ which sounds impressive until you deduplicate by content and find only ~50% are unique problems; the rest are password resets and other high-frequency repeats. The bottleneck isn't model smartness. The real world naturally follows long-tail distribution—standard chatbots can handle the head, not the exceptions in the tail.

— Diogo Almeida
26:42

Reliable systems don't need to produce identical outputs every time

Diogo breaks reliability into three layers. Layer one approximates traditional availability and SLA metrics. Layer two approaches "determinism," but he stresses this matters for unit tests, not real systems—for instance, include a UUID in a prompt and outputs are functionally equivalent but literally different. For the third layer he hasn't yet settled on terminology; call it "robustness"—not requiring identical answers, but requiring each answer to be a judgment that would make sense to a human. He contends that each additional "9" in reliability unlocks an entirely new category of applications previously too risky for AI deployment.

— Diogo Almeida
33:49

The SaaS doomsday thesis overlooked a key detail: major PRs are tiny

When coding agents arrived, markets feared SaaS displacement and stock prices fell accordingly. Jev's emergence prompted collective SaaS celebration—because while "software is cheap" is true, "therefore easily replaced" is wrong. Enormous value lies in invisible engineering details. Ben Horowitz adds a calibrating observation: a typical major pull request at a large company contains ~10 lines of actual code. Coding agents automate writing those 10 lines, not adding new capability. Jev provides a new primitive that lets software grow functionality that didn't previously exist. This explains why AI-written code is faster but software itself hasn't improved—and why less review now means more security risk.

— Diogo Almeida

In their own words · checked verbatim

Where the fuck is all the automation? AI is so unbelievably smart, and yet it's so useless at all other stuff.

Diogo Almeida0:00

But TypeSafe is making AI for software. You know, we want to make AI powerful, not just for humans in the loop, but to actually build real software.

Diogo Almeida3:02

But like intelligence per dollar is my North Star right now. And it could be wrong, just to be clear.

Diogo Almeida8:15

And my favorite thing that you guys say is we build prod, not God.

Ben Horowitz12:19

And, you know, my favorite query was, why is it important to eat socks before meditating?

Diogo Almeida18:25

I am not, for nuanced reasons, I don't think we are on the path of RSI.

Diogo Almeida19:26

OpenAI has been trying to automate customer service since 2020.

Diogo Almeida23:35

You know, like, imagine if all technology just did what you mean. That is, like, that's not sci-fi.

Diogo Almeida41:57

Figures

OpenAI customer service automation start year202023:35
Support tickets claimed resolvable vs unique problems after deduplication95% claimed; ~50% unique after deduplication24:39
Lines of code in typical major PR~1033:49

Glossary

RSI / Recursive Self-Improvement
A speculative AI development pathway where systems iteratively improve their own capabilities, potentially accelerating beyond human control; frequently cited in AGI risk scenarios.
SaaS Apocalypse / SaaSpocalypse
The market narrative that emerged when AI coding tools gained prominence: that AI makes software cheap and easy to replicate, destroying traditional SaaS company valuations.
4GL / Fourth-Generation Programming Language
Experimental programming languages from the 1980s intended to express intent closer to natural language; an earlier precursor to higher-level abstractions.
Probabilistic Programming
A programming paradigm where code directly models uncertainty and probability distributions; a dormant field since the 1970s that recent AI developments are reviving.

How to listen

Who it's for

Engineers building AI applications or AI infrastructure, product leaders, and investors concerned about how AI coding tools affect SaaS valuation logic.

Skip

Around 9 to 11 minutes contains personal growth anecdotes (math competitions, Kaggle experience) with low information density and can be skipped.