世界太吵,来原声听播客

80,000 Hours

AGI 时间线提前一年,但脏活仍是瓶颈

2026 年 AI 进展证据:收入爆炸、编码智能体成熟,但脏活任务仍差;时间线缩短约一年,但长期仍存变数。

AGIAI 智能体收入增长脏活任务推理扩展时间线
Rob Wiblin 系统梳理 2026 年 AI 关键证据,区分真实进展与炒作,对判断 AGI 时间线有直接参考价值。

核心论点 · 时间戳为按文稿位置估算

1:29

从悲观到狂热的反转

2025 年 10 月 AI 情绪悲观,GPT-5 被认为失败,智能体不可靠。两个月后 Claude Opus 4.5 改变一切,程序员发现智能体真的可用,Karpathy 等名人转变态度。关键不是突破性技术,而是 scaling 和 RL 继续奏效,模型变聪明后长链任务出错率降低。Epoch Capabilities Index 显示 2024 年 9 月推理模型出现后,进展速度加快 2-3 倍,且持续至今。

— Rob Wiblin
4:46

AI 收入爆炸式增长

OpenAI 和 Anthropic 过去六个月年化收入增长 700%,近三个月达 1600%。Anthropic 过去 5.5 个月年化增长 8400%,5 月年化收入达 470 亿美元。毛利率从 38% 升至 70% 以上,说明不是亏本销售。收入增长证明真实价值,投资者继续投入,行业进入正循环。

— Rob Wiblin
9:30

METR 时间线翻倍加速

METR 任务完成时间线显示,AI 能完成需人类 12-24 小时的任务,且翻倍速度从每 7 个月加快到每 4 个月。但需谨慎:仅限编码等清晰任务,反馈密度高;AI 公司重点投入;基准比较的是新程序员而非熟悉代码者;50% 成功率是任意阈值,实际需更高可靠性。

— Rob Wiblin
15:38

Mythos 跳升或非新趋势

Anthropic 的 Mythos 模型带来 ECI 指数跳升,相当于 3 个月完成 6 个月进展。但后续两个月进展恢复原速,说明跳升可能源于训练了更大模型(参数约为前代 5 倍),而非新加速趋势。未来每 1-2 年可能重复此类跳升,但非持续加速。

— Rob Wiblin
17:37

Anthropic 内部 AI 加速研究

Anthropic 报告 80% 合并代码由 Claude 编写,员工代码产出提升 8 倍,代码质量与人类相当,预计一年内超越。Claude 完成开放任务成功率从 10% 升至 75%。但证据多来自编码,且员工调查可能高估 AI 贡献。

— Rob Wiblin
22:29

脏活任务仍是短板

AI 在编码、数学等清晰任务上表现优异,但在运行企业等脏活任务上仍差。GDPval 显示 AI 在特定任务上超越人类,但 Vending-Bench 2 中 AI 赚 1 万美元 vs 人类 6 万美元;AI Village 项目成果有限;Andon Labs 的咖啡馆和商店虽运营但亏损。

— Rob Wiblin
33:46

数学证明:原创但非突破

OpenAI 未发布模型解决了单位距离问题猜想,证明其错误。方法人类化,但依赖跨领域知识和耐心。数学家认为人类未发现是因相信猜想为真、缺乏跨领域知识、证明繁琐。Rob 认为这未更新时间线,因为数学是 AI 优势领域,且未见 AI 有灵感闪现。

— Rob Wiblin
38:54

推理扩展成本低于预期

Redwood Research 分析显示,AI 完成任务成本约为人类的 3%,且未上升。尽管思维链变长,但任务也更大,且 token 成本下降。Rob 之前担心推理扩展导致 AI 不经济,现认为不是大问题。

— Rob Wiblin

原话 · 已逐字校验

I’ve never felt this much behind as a programmer. The profession is being dramatically refactored as the bits contributed by the programmer are increasingly sparse… Clearly some powerful alien tool was handed around… the resulting magnitude-9 earthquake is rocking the profession.

我从未感到作为程序员如此落后。这个职业正在被剧烈重构,程序员贡献的部分越来越稀疏……显然有某种强大的外星工具被分发……由此引发的 9 级地震正在撼动这个职业。

Andrej Karpathy1:29

Over the past six months, the combined revenues of OpenAI and Anthropic have been rising at an annualised rate of 700%.

过去六个月,OpenAI 和 Anthropic 的合计收入年化增长率达 700%。

Rob Wiblin4:46

In fact, the most recent data indicates that the gross margins on their inference infrastructure have increased from 38% to over 70% over the year to May 2026.

事实上,最新数据显示,截至 2026 年 5 月的一年里,他们推理基础设施的毛利率从 38% 升至 70% 以上。

Rob Wiblin6:09

Since AI agents took off in December, it’s risen fully 50%.

自 12 月 AI 智能体起飞以来,它上涨了整整 50%。

Rob Wiblin7:36

For a task that AI models have a 50% chance of succeeding at, METR’s internal rule of thumb is that the AIs will complete a third of them 100% of the time, they’ll successfully do a third of them some of the time, and a third of them they’ll manage to complete 0% of the time.

对于 AI 模型有 50% 成功率任务,METR 的经验法则是:AI 能 100% 完成其中三分之一,有时能完成三分之一,完全无法完成三分之一。

Rob Wiblin13:54

In a recent post, titled “When AI builds itself,” the company reported that 80% of their merged code is now written by Claude, and each staff member is now shipping eight times as much code as they were 18 months ago.

在最近一篇题为「当 AI 构建自身」的文章中,公司报告称 80% 的合并代码现在由 Claude 编写,每位员工的代码产出量是 18 个月前的 8 倍。

Rob Wiblin17:37

For the outdoor seating application, it simply completely hallucinated the café’s floor plan. It impersonated an Andon Labs employee to try to obtain an alcohol licence. And after being busted doing that and told to stop, it just tried again using a different employee’s name.

对于户外座位申请,它完全幻觉了咖啡馆的平面图。它冒充 Andon Labs 员工试图获取酒牌。被揭穿并被告知停止后,它又用另一个员工的名字再次尝试。

Rob Wiblin29:05

We are now around the 90th percentile of my AI timelines from 2021 [and from 2019]. I think we’ve now seen enough signs that there could be a relatively fast takeoff starting very soon. We no longer have strong quantitative indicators that transformative AI isn’t coming soon, and are flying blind. 1–4 year timelines are consistent with existing trends.

我们现在处于我 2021 年(和 2019 年)AI 时间线的第 90 百分位左右。我认为我们已经看到足够迹象表明可能很快出现相对快速的起飞。我们不再有强有力的定量指标表明变革性 AI 不会很快到来,我们正在盲目飞行。1-4 年的时间线与现有趋势一致。

Rob Wiblin43:22

数字与实体

OpenAI 和 Anthropic 过去六个月年化收入增长率700%4:46
OpenAI 和 Anthropic 近三个月年化收入增长率1600%4:46
Anthropic 过去 5.5 个月年化收入增长率8400%4:46
Anthropic 5 月年化收入运行率470 亿美元4:46
Anthropic 推理基础设施毛利率38% 升至 70% 以上6:09
H100 租金自 12 月以来涨幅50%7:36
2026 年 AI 资本支出约 6000 亿美元7:36
METR 时间线翻倍速度每 4 个月9:30
Anthropic 合并代码由 Claude 编写比例80%17:37
Anthropic 员工代码产出提升8 倍17:37

术语

RLVR可验证奖励强化学习
利用自动验证反馈训练模型的方法。
METR模型评估与威胁研究
评估 AI 任务完成时间线的组织。
ECI能力指数
综合 40 个基准衡量 AI 进展的指标。
GDPvalGDP 评估
OpenAI 的评估套件,测试 AI 在专业任务上的表现。
Vending-Bench 2自动售货机基准 2
模拟经营自动售货机的 AI 评估环境。
AI VillageAI 村庄
测试 AI 在真实世界项目管理能力的项目。

收听指南

谁该听

关注 AGI 时间线的创业者、投资人和 AI 工程师,尤其是需要判断 AI 能力边界和投资时机的人。

可跳过

第 33-37 分钟数学证明部分可略读,对时间线判断影响不大。