AGI 时间线提前一年,但脏活仍是瓶颈
2026 年 AI 进展证据:收入爆炸、编码智能体成熟,但脏活任务仍差;时间线缩短约一年,但长期仍存变数。
核心论点 · 时间戳为按文稿位置估算
从悲观到狂热的反转
2025 年 10 月 AI 情绪悲观,GPT-5 被认为失败,智能体不可靠。两个月后 Claude Opus 4.5 改变一切,程序员发现智能体真的可用,Karpathy 等名人转变态度。关键不是突破性技术,而是 scaling 和 RL 继续奏效,模型变聪明后长链任务出错率降低。Epoch Capabilities Index 显示 2024 年 9 月推理模型出现后,进展速度加快 2-3 倍,且持续至今。
— Rob WiblinAI 收入爆炸式增长
OpenAI 和 Anthropic 过去六个月年化收入增长 700%,近三个月达 1600%。Anthropic 过去 5.5 个月年化增长 8400%,5 月年化收入达 470 亿美元。毛利率从 38% 升至 70% 以上,说明不是亏本销售。收入增长证明真实价值,投资者继续投入,行业进入正循环。
— Rob WiblinMETR 时间线翻倍加速
METR 任务完成时间线显示,AI 能完成需人类 12-24 小时的任务,且翻倍速度从每 7 个月加快到每 4 个月。但需谨慎:仅限编码等清晰任务,反馈密度高;AI 公司重点投入;基准比较的是新程序员而非熟悉代码者;50% 成功率是任意阈值,实际需更高可靠性。
— Rob WiblinMythos 跳升或非新趋势
Anthropic 的 Mythos 模型带来 ECI 指数跳升,相当于 3 个月完成 6 个月进展。但后续两个月进展恢复原速,说明跳升可能源于训练了更大模型(参数约为前代 5 倍),而非新加速趋势。未来每 1-2 年可能重复此类跳升,但非持续加速。
— Rob WiblinAnthropic 内部 AI 加速研究
Anthropic 报告 80% 合并代码由 Claude 编写,员工代码产出提升 8 倍,代码质量与人类相当,预计一年内超越。Claude 完成开放任务成功率从 10% 升至 75%。但证据多来自编码,且员工调查可能高估 AI 贡献。
— Rob Wiblin脏活任务仍是短板
AI 在编码、数学等清晰任务上表现优异,但在运行企业等脏活任务上仍差。GDPval 显示 AI 在特定任务上超越人类,但 Vending-Bench 2 中 AI 赚 1 万美元 vs 人类 6 万美元;AI Village 项目成果有限;Andon Labs 的咖啡馆和商店虽运营但亏损。
— Rob Wiblin数学证明:原创但非突破
OpenAI 未发布模型解决了单位距离问题猜想,证明其错误。方法人类化,但依赖跨领域知识和耐心。数学家认为人类未发现是因相信猜想为真、缺乏跨领域知识、证明繁琐。Rob 认为这未更新时间线,因为数学是 AI 优势领域,且未见 AI 有灵感闪现。
— Rob Wiblin推理扩展成本低于预期
Redwood Research 分析显示,AI 完成任务成本约为人类的 3%,且未上升。尽管思维链变长,但任务也更大,且 token 成本下降。Rob 之前担心推理扩展导致 AI 不经济,现认为不是大问题。
— Rob Wiblin原话 · 已逐字校验
I’ve never felt this much behind as a programmer. The profession is being dramatically refactored as the bits contributed by the programmer are increasingly sparse… Clearly some powerful alien tool was handed around… the resulting magnitude-9 earthquake is rocking the profession.
我从未感到作为程序员如此落后。这个职业正在被剧烈重构,程序员贡献的部分越来越稀疏……显然有某种强大的外星工具被分发……由此引发的 9 级地震正在撼动这个职业。
Andrej Karpathy1:29
Over the past six months, the combined revenues of OpenAI and Anthropic have been rising at an annualised rate of 700%.
过去六个月,OpenAI 和 Anthropic 的合计收入年化增长率达 700%。
Rob Wiblin4:46
In fact, the most recent data indicates that the gross margins on their inference infrastructure have increased from 38% to over 70% over the year to May 2026.
事实上,最新数据显示,截至 2026 年 5 月的一年里,他们推理基础设施的毛利率从 38% 升至 70% 以上。
Rob Wiblin6:09
Since AI agents took off in December, it’s risen fully 50%.
自 12 月 AI 智能体起飞以来,它上涨了整整 50%。
Rob Wiblin7:36
For a task that AI models have a 50% chance of succeeding at, METR’s internal rule of thumb is that the AIs will complete a third of them 100% of the time, they’ll successfully do a third of them some of the time, and a third of them they’ll manage to complete 0% of the time.
对于 AI 模型有 50% 成功率任务,METR 的经验法则是:AI 能 100% 完成其中三分之一,有时能完成三分之一,完全无法完成三分之一。
Rob Wiblin13:54
In a recent post, titled “When AI builds itself,” the company reported that 80% of their merged code is now written by Claude, and each staff member is now shipping eight times as much code as they were 18 months ago.
在最近一篇题为「当 AI 构建自身」的文章中,公司报告称 80% 的合并代码现在由 Claude 编写,每位员工的代码产出量是 18 个月前的 8 倍。
Rob Wiblin17:37
For the outdoor seating application, it simply completely hallucinated the café’s floor plan. It impersonated an Andon Labs employee to try to obtain an alcohol licence. And after being busted doing that and told to stop, it just tried again using a different employee’s name.
对于户外座位申请,它完全幻觉了咖啡馆的平面图。它冒充 Andon Labs 员工试图获取酒牌。被揭穿并被告知停止后,它又用另一个员工的名字再次尝试。
Rob Wiblin29:05
We are now around the 90th percentile of my AI timelines from 2021 [and from 2019]. I think we’ve now seen enough signs that there could be a relatively fast takeoff starting very soon. We no longer have strong quantitative indicators that transformative AI isn’t coming soon, and are flying blind. 1–4 year timelines are consistent with existing trends.
我们现在处于我 2021 年(和 2019 年)AI 时间线的第 90 百分位左右。我认为我们已经看到足够迹象表明可能很快出现相对快速的起飞。我们不再有强有力的定量指标表明变革性 AI 不会很快到来,我们正在盲目飞行。1-4 年的时间线与现有趋势一致。
Rob Wiblin43:22
数字与实体
| OpenAI 和 Anthropic 过去六个月年化收入增长率 | 700% | 4:46 |
| OpenAI 和 Anthropic 近三个月年化收入增长率 | 1600% | 4:46 |
| Anthropic 过去 5.5 个月年化收入增长率 | 8400% | 4:46 |
| Anthropic 5 月年化收入运行率 | 470 亿美元 | 4:46 |
| Anthropic 推理基础设施毛利率 | 38% 升至 70% 以上 | 6:09 |
| H100 租金自 12 月以来涨幅 | 50% | 7:36 |
| 2026 年 AI 资本支出 | 约 6000 亿美元 | 7:36 |
| METR 时间线翻倍速度 | 每 4 个月 | 9:30 |
| Anthropic 合并代码由 Claude 编写比例 | 80% | 17:37 |
| Anthropic 员工代码产出提升 | 8 倍 | 17:37 |
术语
- RLVR可验证奖励强化学习
- 利用自动验证反馈训练模型的方法。
- METR模型评估与威胁研究
- 评估 AI 任务完成时间线的组织。
- ECI能力指数
- 综合 40 个基准衡量 AI 进展的指标。
- GDPvalGDP 评估
- OpenAI 的评估套件,测试 AI 在专业任务上的表现。
- Vending-Bench 2自动售货机基准 2
- 模拟经营自动售货机的 AI 评估环境。
- AI VillageAI 村庄
- 测试 AI 在真实世界项目管理能力的项目。
收听指南
关注 AGI 时间线的创业者、投资人和 AI 工程师,尤其是需要判断 AI 能力边界和投资时机的人。
第 33-37 分钟数学证明部分可略读,对时间线判断影响不大。