Chinese superpods now out-compute NVIDIA in total, but 5x the power buys no ecosystem
Huawei's CloudMatrix 384 spends roughly 560 kW to deliver 1.7x the total compute of an NVL72, but what decides the fate of domestic superpods is not the spec sheet — it is manufacturing capacity and the software ecosystem.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
No superpod on your booth means no reason to exhibit
More than 10 superpod products were on display at this year's WAIC. Huawei, Alibaba and Baidu, chip makers including Biren (壁仞), Moore Threads (摩尔线程), MetaX (沐曦) and Enflame (燧原), and even server vendors such as ZTE and Sugon (曙光), all put their superpods in the most prominent spot on the booth. Brokerages are calling 2026 year one of the domestic superpod. Behind this me-too competition is intense FOMO: without a superpod, you almost feel too embarrassed to show up. The superpod evolved from the eight-card box into a rack form factor holding dozens to more than a hundred cards, and its core purpose is to solve the communication wall in model training and inference.
— YaxianChinese chips cannot close the single-die gap, so they compete as systems
The gap between domestic AI chips and overseas ones at the single-chip level is obvious and cannot be closed quickly, so performance has to be raised at the system level. Kimi's K3 model has 2.8T total parameters and 896 experts, with each token routing to 16 experts; the official recommendation is a superpod of 64 or more accelerators to solve capacity, bandwidth and latency. At the same time, interconnect technology is now good enough to productise superpods, and vendors would rather sell a whole rack than a single card, because the value per unit is on a completely different scale. Stack these factors together and the superpod becomes the mandatory path for domestic compute to catch up.
— Zhang HaijunEvery scale-up protocol is proprietary, so substitution is an ecosystem problem
Almost every scale-up protocol in the world's superpods is proprietary: NVIDIA has NVLink, Huawei has UB, and Alibaba and the domestic chip companies each have their own. Overseas, UALink was launched by AMD, Marvell, Broadcom and others, but AMD is the main force pushing it; Marvell makes the switch silicon yet lacks motivation, and Broadcom walked out to do its own protocol, which leaves the open standard loose. Huawei's UB can unify scale-up and scale-out, whereas NVIDIA's NVLink only handles interconnect inside the rack, with IB between racks. Fragmented protocols mean ecosystem lock-in: domestic substitution is not only a hardware problem.
— Zhang HaijunCloudMatrix 384 beats NVL72 on total compute by brute-forcing power
NVIDIA's NVL72 is built from 18 compute trays and 9 NVSwitch trays, 72 B200s in total, and the full rack draws over 120 kW; early GB200 shipments were also mostly air-cooled rather than fully liquid-cooled. Huawei's CloudMatrix 384 is China's first superpod: 16 racks holding 384 910Cs. Its single-chip BF16 compute is only 30% of a B200's, but total system compute is 1.7x that of an NVL72. The price is roughly 560 kW for the whole system, close to 5x NVIDIA's. This is buying total compute by piling on hardware, and it explains why China has to take the road of system-level innovation.
— Zhang HaijunServer OEMs refuse to stay assemblers, but only two ship superpods at volume
Traditional server OEMs such as ZTE and H3C (华三) brought superpod designs of their own this time. ZTE drew on its networking expertise to develop its own switch chip, with a protocol called Olink that it says is compatible with several domestic chips, though no details have been published. Lightelligence (曦智科技) is betting on optical interconnect, positioning itself as a supplier of optical communication and optical switching components. Today only Huawei and Sugon actually have superpods in volume production in China; Huawei's customers are spread across carriers and cloud providers, while Sugon is tied to Hygon (海光) and targets the xinchuang (信创) domestic-substitution market. Other vendors' "zero-adaptation" claims are still a long way from large-scale commercial deployment.
— YaxianBeating NVIDIA means winning on an open market, not on spec sheets
Zhang Haijun (张海军) argues that domestic superpods have already passed NVIDIA on system specifications: CM384's total compute is 1.7x that of an NVL72. But his definition of passing is this: if NV were free to sell into the market and cloud providers, on equal terms, still preferred the domestic superpod, that would count as winning. That is hard today, because the CUDA ecosystem runs too deep — domestic engineers are used to CUDA, and DeepSeek has even found bugs in NVIDIA chips — so switching to domestic hardware and software means relearning; NV is also more stable. But the delay to NVIDIA's next-generation product opens a window, and the domestic software ecosystem has improved markedly over the past year, so the gap is narrowing. From 0 to 1 China is weak; from 1 to 100 China is stronger.
— Zhang HaijunCapacity, not the number of chip vendors, is the real bottleneck
A lot of chip vendors does not mean a lot of capacity; the real chokepoint is AI chip capacity. The price of a B300 eight-card box rose from 3.7-3.8 million yuan in Q3 last year to 12-13 million yuan this year, which shows that overseas parts cannot get in and are scarce. One important reason the big internet companies are buying more domestic chips this year than last is that domestic capacity is relatively reliable. Because the software ecosystems do not interoperate, no large buyer can adapt to every chip at once, so a second-tier vendor that still has not won a big-platform order by next year will have a very hard time afterwards. The industry may shake out the way autonomous driving did, leaving five or six players at the end. Huawei and Cambricon (寒武纪) hold stable positions, and big-tech in-house efforts such as T-Head (平头哥) and Enflame also have a shot.
— Zhang HaijunPower, liquid cooling and switch chips may monetise before superpods do
Superpod power draw rises sharply, and NVIDIA started pushing 800V high-voltage DC power distribution last year; liquid cooling has gone from optional to standard, with GB200 later switching to liquid cooling and GB300 fully liquid-cooled, and Google's figures show that 2027 liquid cooling demand is worth 4x what it was before. On the network side, China previously relied mainly on Broadcom switch chips, and domestic switch silicon can now match Broadcom. Scale-out is still dominated by Broadcom, while scale-up stays fragmented across each vendor's proprietary protocol. The upgrade of these supporting technologies is quite likely to release commercial opportunity earlier than the superpod itself.
— Zhang HaijunIn their own words · checked verbatim
Going from the eight-card box to the superpod, the main change is really about solving the communication wall.
从巴卡机到超节点,它的一个主要变化,其实是为了解决通信墙的问题。
Zhang Haijun6:35
A single die may have only a thirtieth of its compute. But once you put them together into a superpod, the total compute is roughly twice what it has.
单片算力是可能只有它的多少30分之1。但是合在一起会做成超点点以后,其实总算力是它大概两倍。
Zhang Haijun23:21
It is possible for us to pass NV. Because our characteristic is that from 0 to 1 we may be a bit weaker than the American side. But from 1 to 100, I think we are definitely stronger than the American side.
我们超NV是有可能的。因为我们的特点是说,我们可能从0到1可能会比美国那边要弱一些。但是我们从1到100,我觉得肯定是要强于美国那边的。
Zhang Haijun32:54
Capacity is going to be a bottleneck now, meaning whoever has capacity probably has the bigger advantage.
现在产能会是一个瓶颈,就是说谁油产能,可能谁的优势就会更大一些。
Zhang Haijun37:02
If you still have not landed a big order from one of these major internet companies, then I think it gets fairly hard from there.
你还没有从这几家大的互联网厂商拿到一个大的订单的话,那我觉得后面可能就比较难了。
Zhang Haijun41:06
Figures
| Superpod products on display (WAIC 2025) | more than 10 | 4:26 |
| K3 total parameters | 2.8T | 11:51 |
| K3 expert count | 896 | 11:51 |
| NVL72 rack power | over 120 kW | 22:10 |
| 910CA BF16 compute vs B200 | about 30% | 23:21 |
| CM384 total compute vs NVL72 | 1.7x | 23:21 |
| B300 eight-card box price (Q3 last year) | 3.7-3.8 million yuan | 38:02 |
| Google's 2027 liquid cooling demand value | 4x the prior level | 44:23 |
Glossary
- Scale-up
- Combining many chips into a single giant GPU to raise the compute available to one task.
- Scale-out
- Expanding cluster scale by adding more racks or nodes.
- UALink / Ultra Accelerator Link
- An open accelerator-interconnect protocol launched by AMD, Broadcom and others; participation so far is loose.
- CloudMatrix / Huawei's superpod line
- Huawei's multi-rack interconnected superpod, such as the 384-chip configuration.
How to listen
Cloud procurement teams, chip startups and compute-sector investors tracking the substitution of domestic AI compute, who need to read both the technology roadmap and the inflection point in orders.
The opening walk-the-show grumbling (0:00-4:30) is skippable; start at 5:28 with the rundown of superpod vendors.