The world is too loud. Read what matters.

老石谈芯

In the age of agents, GPUs are perpetually jammed—CPUs are the traffic control system

Agents transform AI work from parallel probability prediction to serial task execution; GPUs, lacking branch prediction and out-of-order execution, frequently idle, making CPU's logical scheduling a new scarce resource in data centers.

CPUAI agentsAI infrastructureIntelx86 ecosysteminference compute
Three layers of logic link seamlessly: agents' serial needs, AI infrastructure supply-side restructuring, x86's competitive moat—high information density, but skip transcription errors around 4:03-6:03 and 9:05-10:07.

The argument · tap a timestamp to hear it

1:01

Chat is a single order, agents are planning a banquet

Chatbots work in a question-answer loop: input a prompt, the model predicts the next token, like ordering beef noodles from a cafeteria—you ask, the server just ladles one bowl. This single-step parallel probability calculation is perfect for GPUs, which use hundreds of compute cores and TensorCores to excel at matrix multiplication. Agents work differently. Give one a research task and it must decompose steps itself, search the web, write code, debug, revise—even spawn sub-agents for division of labor, like having the cafeteria manager write the menu and orchestrate the kitchen to serve a mapo-tofu hotpot spread plus BBQ all at once. The work pattern shifts from single-step parallel to multi-step serial.

— Lao Shi
3:03

GPUs are jammed, CPUs are the traffic lights

Agent tasks are filled with conditionals and exception handling, requiring constant logic switching and serial execution chain-link style—precisely where GPUs hit an architectural wall. GPUs excel at doing thousands of simple repetitive operations in parallel but lack the ability to handle complex logic; once the pipeline breaks, vast compute units sit idle waiting—what we call the compute-traffic-jam phenomenon. CPUs have fewer cores, but each one packs massive branch-prediction circuits and out-of-order execution engines designed to handle conditionals and task scheduling, even predicting the next steps ahead of time.

— Lao Shi
6:03

Feeding GPUs requires stacking CPU core density

When thousands of agents run simultaneously and need seamless data feeds to GPUs, infrastructure providers stack CPU core density in data centers—like hiring more prep cooks to keep the GPU chef always busy. Intel's Xeon 6 Clearwater Forest packs up to 288 efficient cores per chip; a dual-socket server delivers 2,304 cores; a single cabinet reaches 46,080 cores. Cloud Arrow's full cabinets (36U, 6 CPUs per U max) with Xeon 6 Plus support 60,000+ cores, theoretically running 240,000 lightweight agents simultaneously.

— Lao Shi
8:05

When inference exceeds training, CPUs become the bottleneck

In 2025, large-model inference already demands half of total AI compute; by 2026 that rises to 2/3. But per-request compute isn't heavy at inference—the real bottleneck shifts to latency-sensitive chores: serialization, KV cache management, prompt assembly, retrieval-augmented generation routing. CPUs handle all the gritty work. Gary summarizes the difference this way: GPUs solve for ‘intelligence quality,’ CPUs solve for ‘intelligence density."

— Gary
11:08

An AI factory isn't one machine, it's three

AI infrastructure is shifting from a crude compute warehouse toward an ‘intelligent factory’ with CPU as the orchestration brain: one machine is the GPU server generating tokens; one is a dense CPU cluster hosting thousands, even millions, of CPU-resident agents; one is a high-performance storage server caching KV residuals from GPU inference—next time you hit a similar problem, pull the pre-computed result instead of recalculating. The common denominator across all three: raw processor power.

— Gary
12:09

CPU handles compute, storage, connectivity, and reliability—all four

Gary breaks down the AI factory's hard requirements into four pillars: compute—CPUs feed data to GPUs; Xeon 6 Plus uses AMX for direct AI data preprocessing and vector acceleration; storage—when KV cache balloons beyond HBM capacity, it offloads to CPU memory or storage, like Dumbledore extracting memories into a Pensieve; reliability—CPU RAS features can warn before memory bits fail, even repair errors; connectivity—new E835 NICs provide 200G RDMA links, letting CPU clusters access GPU and storage clusters at top speed.

— Gary
18:12

CPU-to-GPU ratios are approaching 1:1

Pile compute, storage, connectivity, and reliability together and the conclusion is clear: the more complex AI gets, the more agents run, the harder CPUs work. That's why data-center CPU-to-GPU ratios are evolving from 1:8 to 1:4, 1:2, even approaching 1:1—not because CPUs grew powerful enough to replace GPUs, but because maintaining full data flow without GPU idle cycles requires scaling CPUs in parallel.

— Lao Shi
19:12

What scares enterprises about changing architecture is migration, not performance

By 2030, x86 still claims 80% of global server inventory, growing from 63 million units in 2025 to 76 million. Gary highlights an overlooked fact: Intel is always first to complete compatibility with SSDs, NICs whenever new standards like PCIe 5.0 arrive—system makers find x86 design twice as fast, other architectures require shuttling between vendors. For enterprises, swapping architectures isn't just swapping chips—it means recompiling and retesting a decade of accumulated business systems, open-source libraries, private databases. That trial cost is x86's true moat, earned through decades of ecosystem building.

— Gary

In their own words · checked verbatim

But have you considered whether this CPU boom from agents is a real trend or a US-centric demand signal?

但你有没有想过 这智能体带来的这波CPU的爆发 到底是真趋势还是美需求呢

Lao Shi0:00

It's like a city with ultra-wide highways but no traffic lights or traffic control systems—when you hit a complex intersection, all vehicles get jammed. We call that the compute-traffic-jam phenomenon of the agent era.

就好像一个城市只有极宽的高速公路 却没有红绿灯和交通指挥系统 当遇到复杂的交叉口的时候 所有的车辆都会死死的堵在路上 我们就把它叫做智能体时代出现的算力交通管制现象

Lao Shi3:03

To generate tokens in this factory, imagine three massive machines: one is what we traditionally call a GPU server.

那么为了产生Token 在这个抽机工厂里头 你可以想象有三台机器 三台巨大的机器 一台机器就是我们传统的 所谓GPU的服务器

Gary11:08

It's like in Harry Potter—Dumbledore would extract memories from his mind, put them in a vial, then place them in the Pensieve to review when needed.

这有点像哈利波特里 邓比利多总会把一些记忆从脑子里给抽出来 然后放到小瓶子里 需要的时候再把它放到冥想盆里去查看

Gary14:11

What enterprises really buy is never just a cold chip, but an entire ecosystem that ensures stable business operations.

企业真正购买的从来都不只是一块冷冰冰的芯片 而是一整套能够确保业务稳定运行的生产环境

Lao Shi20:14

Figures

IDC forecast for active agents in 2031350 million5:03
IDC forecast for China's active agents growth rate 2026-2027200%5:03
Xeon 6 Clearwater Forest efficient cores per chip2886:03
Intel solution CPU cores per cabinet46,0806:03
Cloud Arrow cabinet CPU cores60,000+7:05
Cloud Arrow cabinet lightweight agent capacity240,0007:05
Large-model inference share of AI compute in 20262/38:05
Data center CPU:GPU ratio evolutionfrom 1:8 to 1:4, 1:2, approaching 1:118:12

Glossary

KV Cache
Intermediate computation results cached during inference, the model's working memory or notes for later lookup.
RAS
CPU reliability features: high reliability, high availability, high serviceability.
AMX
Intel CPU instruction set for accelerating AI matrix operations.
PD Separation
Inference architecture that splits input processing (prefill) and output generation (decode) into separate deployments.
RDMA
Remote Direct Memory Access; server-to-server data transfer bypassing CPU for low-latency links.

How to listen

Who it's for

Anyone tracking AI infrastructure, semiconductor investment, or data-center architecture; practitioners seeking to understand how the AI agent boom is driving CPU demand.

Skip

Transcription errors around 4:03-6:03, 9:05-10:07 (repetitive word stacking), and 13:11 (mixed Chinese-English); skippable without loss of meaning.