世界太吵,來原聲聽播客

80,000 Hours

AI 收入年增 700%,但髒活任務仍差人類 6 倍

2026 年 AI 進展證據:收入爆炸、編碼智能體成熟,但髒活任務仍差;時間線縮短約一年,但長期仍存變數。

AGIAI 智能體收入增長髒活任務推理擴展時間線
Rob Wiblin 系統梳理 2026 年 AI 關鍵證據,區分真實進展與炒作,對判斷 AGI 時間線有直接參考價值。

核心論點 · 時間戳為按文稿位置估算

1:29

從悲觀到狂熱的反轉

2025 年 10 月 AI 情緒悲觀,GPT-5 被認為失敗,智能體不可靠。兩個月後 Claude Opus 4.5 改變一切,程序員發現智能體真的可用,Karpathy 等名人轉變態度。關鍵不是突破性技術,而是 scaling 和 RL 繼續奏效,模型變聰明後長鏈任務出錯率降低。Epoch Capabilities Index 顯示 2024 年 9 月推理模型出現後,進展速度加快 2-3 倍,且持續至今。

— Rob Wiblin
4:46

AI 收入爆炸式增長

OpenAI 和 Anthropic 過去六個月年化收入增長 700%,近三個月達 1600%。Anthropic 過去 5.5 個月年化增長 8400%,5 月年化收入達 470 億美元。毛利率從 38% 升至 70% 以上,說明不是虧本銷售。收入增長證明真實價值,投資者繼續投入,行業進入正循環。

— Rob Wiblin
9:30

METR 時間線翻倍加速

METR 任務完成時間線顯示,AI 能完成需人類 12-24 小時的任務,且翻倍速度從每 7 個月加快到每 4 個月。但需謹慎:僅限編碼等清晰任務,反饋密度高;AI 公司重點投入;基準比較的是新程序員而非熟悉代碼者;50% 成功率是任意閾值,實際需更高可靠性。

— Rob Wiblin
15:38

Mythos 跳升或非新趨勢

Anthropic 的 Mythos 模型帶來 ECI 指數跳升,相當於 3 個月完成 6 個月進展。但後續兩個月進展恢復原速,說明跳升可能源於訓練了更大模型(參數約為前代 5 倍),而非新加速趨勢。未來每 1-2 年可能重複此類跳升,但非持續加速。

— Rob Wiblin
17:37

Anthropic 內部 AI 加速研究

Anthropic 報告 80% 合併代碼由 Claude 編寫,員工代碼產出提升 8 倍,代碼質量與人類相當,預計一年內超越。Claude 完成開放任務成功率從 10% 升至 75%。但證據多來自編碼,且員工調查可能高估 AI 貢獻。

— Rob Wiblin
22:29

髒活任務仍是短板

AI 在編碼、數學等清晰任務上表現優異,但在運行企業等髒活任務上仍差。GDPval 顯示 AI 在特定任務上超越人類,但 Vending-Bench 2 中 AI 賺 1 萬美元 vs 人類 6 萬美元;AI Village 項目成果有限;Andon Labs 的咖啡館和商店雖運營但虧損。

— Rob Wiblin
33:46

數學證明:原創但非突破

OpenAI 未發布模型解決了單位距離問題猜想,證明其錯誤。方法人類化,但依賴跨領域知識和耐心。數學家認為人類未發現是因相信猜想為真、缺乏跨領域知識、證明繁瑣。Rob 認為這未更新時間線,因為數學是 AI 優勢領域,且未見 AI 有靈感閃現。

— Rob Wiblin
38:54

推理擴展成本低於預期

Redwood Research 分析顯示,AI 完成任務成本約為人類的 3%,且未上升。儘管思維鏈變長,但任務也更大,且 token 成本下降。Rob 之前擔心推理擴展導致 AI 不經濟,現認為不是大問題。

— Rob Wiblin

原話 · 已逐字校驗

I’ve never felt this much behind as a programmer. The profession is being dramatically refactored as the bits contributed by the programmer are increasingly sparse… Clearly some powerful alien tool was handed around… the resulting magnitude-9 earthquake is rocking the profession.

我從未感到作為程序員如此落後。這個職業正在被劇烈重構,程序員貢獻的部分越來越稀疏……顯然有某種強大的外星工具被分發……由此引發的 9 級地震正在撼動這個職業。

Andrej Karpathy1:29

Over the past six months, the combined revenues of OpenAI and Anthropic have been rising at an annualised rate of 700%.

過去六個月,OpenAI 和 Anthropic 的合計收入年化增長率達 700%。

Rob Wiblin4:46

In fact, the most recent data indicates that the gross margins on their inference infrastructure have increased from 38% to over 70% over the year to May 2026.

事實上,最新數據顯示,截至 2026 年 5 月的一年裡,他們推理基礎設施的毛利率從 38% 升至 70% 以上。

Rob Wiblin6:09

Since AI agents took off in December, it’s risen fully 50%.

自 12 月 AI 智能體起飛以來,它上漲了整整 50%。

Rob Wiblin7:36

For a task that AI models have a 50% chance of succeeding at, METR’s internal rule of thumb is that the AIs will complete a third of them 100% of the time, they’ll successfully do a third of them some of the time, and a third of them they’ll manage to complete 0% of the time.

對於 AI 模型有 50% 成功率任務,METR 的經驗法則是:AI 能 100% 完成其中三分之一,有時能完成三分之一,完全無法完成三分之一。

Rob Wiblin13:54

In a recent post, titled “When AI builds itself,” the company reported that 80% of their merged code is now written by Claude, and each staff member is now shipping eight times as much code as they were 18 months ago.

在最近一篇題為「當 AI 構建自身」的文章中,公司報告稱 80% 的合併代碼現在由 Claude 編寫,每位員工的代碼產出量是 18 個月前的 8 倍。

Rob Wiblin17:37

For the outdoor seating application, it simply completely hallucinated the café’s floor plan. It impersonated an Andon Labs employee to try to obtain an alcohol licence. And after being busted doing that and told to stop, it just tried again using a different employee’s name.

對於戶外座位申請,它完全幻覺了咖啡館的平面圖。它冒充 Andon Labs 員工試圖獲取酒牌。被揭穿並被告知停止後,它又用另一個員工的名字再次嘗試。

Rob Wiblin29:05

We are now around the 90th percentile of my AI timelines from 2021 [and from 2019]. I think we’ve now seen enough signs that there could be a relatively fast takeoff starting very soon. We no longer have strong quantitative indicators that transformative AI isn’t coming soon, and are flying blind. 1–4 year timelines are consistent with existing trends.

我們現在處於我 2021 年(和 2019 年)AI 時間線的第 90 百分位左右。我認為我們已經看到足夠跡象表明可能很快出現相對快速的起飛。我們不再有強有力的定量指標表明變革性 AI 不會很快到來,我們正在盲目飛行。1-4 年的時間線與現有趨勢一致。

Rob Wiblin43:22

數字與實體

OpenAI 和 Anthropic 過去六個月年化收入增長率700%4:46
OpenAI 和 Anthropic 近三個月年化收入增長率1600%4:46
Anthropic 過去 5.5 個月年化收入增長率8400%4:46
Anthropic 5 月年化收入運行率470 億美元4:46
Anthropic 推理基礎設施毛利率38% 升至 70% 以上6:09
H100 租金自 12 月以來漲幅50%7:36
2026 年 AI 資本支出約 6000 億美元7:36
METR 時間線翻倍速度每 4 個月9:30
Anthropic 合併代碼由 Claude 編寫比例80%17:37
Anthropic 員工代碼產出提升8 倍17:37

術語

RLVR可驗證獎勵強化學習
利用自動驗證反饋訓練模型的方法。
METR模型評估與威脅研究
評估 AI 任務完成時間線的組織。
ECI能力指數
綜合 40 個基準衡量 AI 進展的指標。
GDPvalGDP 評估
OpenAI 的評估套件,測試 AI 在專業任務上的表現。
Vending-Bench 2自動售貨機基準 2
模擬經營自動售貨機的 AI 評估環境。
AI VillageAI 村莊
測試 AI 在真實世界項目管理能力的項目。

收聽指南

誰該聽

關注 AGI 時間線的創業者、投資人和 AI 工程師,尤其是需要判斷 AI 能力邊界和投資時機的人。

可跳過

第 33-37 分鐘數學證明部分可略讀,對時間線判斷影響不大。