The world is too loud. Read what matters.

硅谷101

General LLMs hit 99 points, real e-commerce tasks only 61% pass rate

Zhang Kuo said general large models score nearly perfect on benchmarks, but their tests of 107 real e-commerce tasks show even the strongest models only achieve 61% pass rates without human intervention—the gap is not in model intelligence, but in task complexity, difficulty of verification, and long feedback cycles.

E-commerce agentsModel evaluationMulti-model routingCross-border e-commerceAI cost

The video won't play here. Listen to the audio instead:

Combines specific evaluation methodology and data with practical multi-model routing techniques that cut execution costs to one-third of competitors—high information density.

The argument · tap a timestamp to hear it

2:00

Coding saturation limits AI adoption; business agents are just beginning

Zhang Kuo estimates AI penetration in coding is already about 99%—it's hard to find a developer who doesn't use AI to write code. The reason: code is 100% digitized, all preserved in repositories like GitHub, and correctness can be verified instantly by compilation. Commercial scenarios are entirely different: data is more heterogeneous, there are no unified workflows, different companies have different standards for what constitutes a "good result," and verification cycles can last a week, a month, or even a year. This fundamental difference explains why Cowork-type products advance much more slowly in complex scenarios like e-commerce.

— Zhang Kuo
9:03

Supplier comparison squeezed from months to days

Without agents, merchants maintain a 20-column Excel spreadsheet comparing a dozen suppliers' production credentials, pricing, payment terms, and logistics solutions. Back-and-forth communication typically takes months before they can get a sample. Accio Work first breaks down an idea into a manufactureable product and required capabilities, matches suppliers, negotiates on their behalf, and delivers quotes—what once took a month now returns results in 24 hours. Zhang Kuo notes that suppliers have remarked, "I've never seen such a professional design pack," because historically buyers mostly just sketched requirements on paper.

— Zhang Kuo
22:09

Benchmarks near perfect; real tasks barely pass

Many models launch with benchmarks already at 99 points (out of 100), yet the team doesn't feel an equivalent improvement in daily operations. So they constructed 107 e-commerce tasks from millions of real operational data points, spanning seven major categories: quotations, disputes, shipping, and others. They tested whether models could complete these without human intervention. Result: even the most frontier models only achieve about 61% pass rate. Zhang Kuo emphasizes that high coding and general intelligence scores show no strong correlation with successfully completing e-commerce tasks.

— Zhang Kuo
24:10

Multi-model routing cuts costs to one-third of competitors

Accio Work doesn't run every task on the most expensive model. Instead, it routes problems by Pareto optimality: simple L3-and-below tasks run on small, cost-effective models, while complex problems go to frontier models. It leverages Accio Work's own harness and context capabilities to extract better results from the same model. Zhang Kuo illustrates this with Qwen's latest small model, which through post-training achieves results approaching frontier top-tier performance. For the same 107 tasks at equivalent pass rates, Accio Work's estimated cost is $3.69, while Codex and Claude Code both exceed $9—making Accio Work's cost one-third of the other two.

— Zhang Kuo
27:13

Commercial AGI means earning more than humans in identical conditions

Zhang Kuo defines commercial AGI concretely: under the same budget and supply chain constraints, an agent making day-to-day operational decisions must earn more money than a human. That's what counts. However, he warns that models and agents undergo a corruption phase after iterating for a time—cases that worked previously may fail later. Benchmarking therefore isn't one-time; the evaluation of 107 tasks should be re-run with each version to prevent the illusion that a new version represents progress when it may actually regress on certain tasks.

— Zhang Kuo
29:13

High cross-platform revenue doesn't mean real profit

Chinese merchants typically operate across multiple platforms—Douyin, Alibaba, JD.com, Pinduoduo. Which platform is actually profitable today isn't determined by gross flow volume but by reconciling investment, revenue, and returns data across all platforms. Historically merchants relied on manually switching between backend dashboards and assembling data in Excel spreadsheets, often with errors. Zhang Kuo believes reconciling such cross-platform real-time operational insights and routing them to AI to interpret should be quite easy now—and it's also one of the most common requests merchants make as task difficulty scales from L1 (Q&A) up to L4 (help me decide how to manage the business).

— Zhang Kuo
39:24

Multiple agents give conflicting advice; someone must decide

When merchants simultaneously use multiple platforms' own agents and receive contradictory recommendations, Zhang Kuo's answer is straightforward: that's normal, because Accio Work's positioning is to optimize from the merchant's single source of truth—a unified view—rather than optimizing for any one platform's interests. Whose advice to follow depends on the owner's own business goals: pursuing margin, scale, or speed. "There's no way to use a single unified agent to reconcile all agents' opinions and give users better feedback," he says. "That's not our product design intent—final commercial decisions must remain with humans."

— Zhang Kuo
44:26

Small merchants scale to tens of millions via AI

The CoCreate conference registered approximately 15,000 attendees. Zhang Kuo shared two startup case studies using Accio Work: a 15-year-old student in the UK used it to source suppliers and built a series of pet cooling towels that now generate about £1 million in annual sales; Christian Reed, an MIT graduate, completed everything from idea to supply chain using Alibaba.com and Accio Work, now with customers including Schneider Electric and SpaceX, scaling his business to several tens of millions USD this year. Zhang Kuo regards these as typical examples of AI helping small enterprises scale.

— Zhang Kuo

In their own words · checked verbatim

I estimate about 99% because it's hard to find a developer around who doesn't use some form of AI to write code.

我估计有99% 因为很难看到身边有一个研发人员 不用任何AI去写代码

Zhang Kuo1:00

The cycle is usually measured in months, and in the end you get a sample; if you have an agent, we can compress this cycle to days.

周期一般可能是以月计 到最后你能拿到一个样品 如果有Agent 我们可以把你这个周期 缩短到以天计

Zhang Kuo8:02

So many benchmarks have already hit 99 points out of 100—after the decimal point there's room to improve—but when we use it ourselves in daily operations, we don't feel like it's actually that good.

甚至非常多的榜单 都已经达到99分了 满分100 99小数点之后去进步 但我们自己在日常去使用的时候 又没有感觉到像说得这么好

Zhang Kuo21:08

In fact, the most frontier models, in our benchmark tests, only achieve a pass rate of just over 61%—reaching even a passing grade is a bit tough right now.

事实上可能现在 最frontier的模型 在我们这个基准测试下面 能够通过率就是61%多一点点 它想要达到及格水平 现在都有点费劲

Zhang Kuo22:09

With Accio Work, our estimated cost is $3.69—a figure I see you provided—Codex and Claude Code are both above $9, so Accio Work's cost is one-third of the other two.

用Accio Work的估算成本 是3.69美元 我看到是你们给到的一个数字 Codex和Claude Code 都是在9美元以上 所以就是Accio Work的成本 是另外两家的三分之一

Yiwen33:18

There's no way to use a single unified agent to reconcile all agents' opinions and give users better feedback—that's not the design intent of our product.

这个事情没有办法 用一个统一Agent 来统一一下所有Agent的意见 然后给用户一个更好的反馈 这不是我们这个产品设计的初衷

Zhang Kuo39:24

We probably estimate about 15,000 registered on-site attendees. We definitely made our registration booth too small—there was a queue all the way until lunch and people still weren't fully through.

我们可能最终估计 注册到场的有15000人 我们肯定前面那个摊位 注册摊搞小了 一直排队到中午吃饭的时候 人还没有完全进来解锁

Zhang Kuo43:25

Figures

Coding scene AI penetration rate~99%1:00
Overall search vs multimodal search growthOverall +20%, multimodal +130%4:02
Alibaba International's self-built e-commerce benchmark tasks107 tasks (seven categories)21:08
Frontier model pass rate without human intervention~61%22:09
Accio Work estimated cost to complete equivalent tasks$3.6933:18
Codex/Claude Code cost to complete equivalent tasks$9+33:18
CoCreate conference registered attendees~15,00043:25
UK 15-year-old entrepreneur's annual sales~£1 million43:25
US entrepreneur Christian Reed's business scale this yearSeveral tens of millions USD+44:26

Glossary

Harness
An engineering layer connecting tools, managing memory and permissions—determines whether an agent can execute tasks end-to-end.
Ontology
Structured understanding of different types of data and business semantics within an enterprise.
Verifiable
Whether task results can be directly judged as correct or incorrect, such as whether code compiles successfully.

How to listen

Who it's for

People working on cross-border e-commerce, B2B procurement, or enterprise AI agents who want to know whether model benchmark scores translate to real business results.

Skip

10:13-14:11 US-China SaaS ecosystem comparison overlaps with previous episode content; can skip ahead.