AI Coding Tools Code Faster, But Software Itself Hasn't Gotten Smarter in a Decade
Coding agents let people code faster, but they produce the same kind of software. Diogo Almeida argues what we should really do is make AI a new primitive in software, automating things that previously couldn't be automated.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Coding agents and Jev solve fundamentally different problems
Cloud Code and Codex represent "just in time software"—writing code in natural language produces output equivalent to code written by humans, with unchanged expressiveness. Diogo's vision isn't faster coding but a new AI primitive embedded in software itself: not automating "software engineering" as a job, but automating tasks that should be automated but currently aren't. That changes what software itself can do.
— Diogo AlmeidaJev is a classifier, but that's not a limitation
Diogo explicitly describes Jev as a classifier—this is precise technical language, not dismissive. It's more akin to a library: you describe desired behavior in natural language, pair it with a state machine, and Jev selects branches with confidence scores. He estimates current Jev likely already outperforms what a traditional ML engineering team in 2019 would ship after months of dataset collection and model tuning—and ordinary developers can now "write" this capability on the spot.
— Diogo AlmeidaWe're building products, not creating deity
Ben Horowitz cited TypeSafe's phrase "we build prod, not God" as its most compelling insight. Other large-model labs, even privately, revel in the weighty narrative that "we are creating potentially uncontrollable technology." TypeSafe inverts this—AI optimism is simply put on the table. Diogo traces much of the industry's pessimism to the "single brain rules everything" myth: if you believe only one ever-larger centralized model represents AI's future, it's hard to grasp why developers get excited about a programmable AI primitive.
— Ben HorowitzHe believed GPT was AGI until his entire worldview collapsed
Diogo participated in releasing early RLHF models at OpenAI and hadn't anticipated instruction-tuning's generalization strength. The team deliberately tested with ‘why should you eat socks before meditating’—a question guaranteed never to appear on the internet—and the model produced sensible, human-like answers. He genuinely believed this model probably was AGI. When it wasn't, he describes it as a moment his ‘entire worldview collapsed,’ and the inflection point toward his later conviction that AI should be genuinely useful, not merely appear intelligent.
— Diogo AlmeidaDisbelieves recursive self-improvement but sees AGI through a different lens
Diogo explicitly states that for various complex reasons, he doesn't believe AI follows a recursive self-improvement path—neither in the past nor now. However, he does endorse OpenAI's simpler definition of AGI: automating most economically valuable work globally. This is entirely feasible since much valuable economic work is actually basic and repetitive. The real problem emerged post-RLHF: industry focused energy on satisfying a "human rater" proxy metric. GPT-3 was initially well-calibrated, but with humans scoring outputs, the model increasingly produced eloquent text without proportionally improving at actual task automation.
— Diogo AlmeidaCustomer service stalls at the long tail, not at intelligence limits
Diogo points to OpenAI's customer service automation efforts since 2020 as a counterexample: intelligence already suffices—so why hasn't it worked? Martin offers a precise observation from early customer service company data: they'd claim ‘we resolve 95% of tickets,’ which sounds impressive until you deduplicate by content and find only ~50% are unique problems; the rest are password resets and other high-frequency repeats. The bottleneck isn't model smartness. The real world naturally follows long-tail distribution—standard chatbots can handle the head, not the exceptions in the tail.
— Diogo AlmeidaReliable systems don't need to produce identical outputs every time
Diogo breaks reliability into three layers. Layer one approximates traditional availability and SLA metrics. Layer two approaches "determinism," but he stresses this matters for unit tests, not real systems—for instance, include a UUID in a prompt and outputs are functionally equivalent but literally different. For the third layer he hasn't yet settled on terminology; call it "robustness"—not requiring identical answers, but requiring each answer to be a judgment that would make sense to a human. He contends that each additional "9" in reliability unlocks an entirely new category of applications previously too risky for AI deployment.
— Diogo AlmeidaThe SaaS doomsday thesis overlooked a key detail: major PRs are tiny
When coding agents arrived, markets feared SaaS displacement and stock prices fell accordingly. Jev's emergence prompted collective SaaS celebration—because while "software is cheap" is true, "therefore easily replaced" is wrong. Enormous value lies in invisible engineering details. Ben Horowitz adds a calibrating observation: a typical major pull request at a large company contains ~10 lines of actual code. Coding agents automate writing those 10 lines, not adding new capability. Jev provides a new primitive that lets software grow functionality that didn't previously exist. This explains why AI-written code is faster but software itself hasn't improved—and why less review now means more security risk.
— Diogo AlmeidaIn their own words · checked verbatim
Where the fuck is all the automation? AI is so unbelievably smart, and yet it's so useless at all other stuff.
Diogo Almeida0:00
But TypeSafe is making AI for software. You know, we want to make AI powerful, not just for humans in the loop, but to actually build real software.
Diogo Almeida3:02
But like intelligence per dollar is my North Star right now. And it could be wrong, just to be clear.
Diogo Almeida8:15
And my favorite thing that you guys say is we build prod, not God.
Ben Horowitz12:19
And, you know, my favorite query was, why is it important to eat socks before meditating?
Diogo Almeida18:25
I am not, for nuanced reasons, I don't think we are on the path of RSI.
Diogo Almeida19:26
OpenAI has been trying to automate customer service since 2020.
Diogo Almeida23:35
You know, like, imagine if all technology just did what you mean. That is, like, that's not sci-fi.
Diogo Almeida41:57
Figures
| OpenAI customer service automation start year | 2020 | 23:35 |
| Support tickets claimed resolvable vs unique problems after deduplication | 95% claimed; ~50% unique after deduplication | 24:39 |
| Lines of code in typical major PR | ~10 | 33:49 |
Glossary
- RSI / Recursive Self-Improvement
- A speculative AI development pathway where systems iteratively improve their own capabilities, potentially accelerating beyond human control; frequently cited in AGI risk scenarios.
- SaaS Apocalypse / SaaSpocalypse
- The market narrative that emerged when AI coding tools gained prominence: that AI makes software cheap and easy to replicate, destroying traditional SaaS company valuations.
- 4GL / Fourth-Generation Programming Language
- Experimental programming languages from the 1980s intended to express intent closer to natural language; an earlier precursor to higher-level abstractions.
- Probabilistic Programming
- A programming paradigm where code directly models uncertainty and probability distributions; a dormant field since the 1970s that recent AI developments are reviving.
How to listen
Engineers building AI applications or AI infrastructure, product leaders, and investors concerned about how AI coding tools affect SaaS valuation logic.
Around 9 to 11 minutes contains personal growth anecdotes (math competitions, Kaggle experience) with low information density and can be skipped.