The fatal wound in custom AI outsourcing: capabilities accumulate in people, not corporate moats
Custom AI outsourcing embeds capabilities in individual people, who disperse them when they leave; real moats can only accumulate in the enterprise's own model weights.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
FDE is just outsourcing with a new name
Zhang Fan argues that FDE is essentially enterprise custom outsourcing with a cooler name. Capability embeds itself in specific people, not in the enterprise. When people leave, the capability disperses with them and can't be replicated at scale. His reading of Palantir's surging margins also diverges from mainstream narrative—not organizational or business model innovation, but AI compressing what once took a team into work one person can handle (cost drops from 100 to 10). The value leap comes from AI itself. This lead is bound to Pentagon-tier unique customers and elite engineers, hard for typical FDE firms to replicate.
— Zhang FanThe clash is intelligence-first versus enterprise-first
Two opposing commercial philosophies. Model companies worship AGI: if intelligence gets strong enough, enterprises will pay any price. Enterprises ask the opposite: do we want to chain ourselves to one model vendor? Zhang Fan cites Kimi recruiting ecosystem partners for FDE work—if the deliverable ties to Kimi's model, enterprises may not accept. This gulf between intelligence-centric and enterprise-centric is rarely made explicit, yet it's enormous.
— Zhang FanSynthetic data works when you flip the math
Zhang Fan once rejected synthetic data as self-delusion—if 100 rules generate 10,000 data points, it's guaranteed overfitting. Then the math flipped for him: if 10,000 rules generate 100 data points, overfitting disappears. Large models convert human knowledge into trillions of parameters, precisely enough to scatter trillions of rules evenly across that space. He's now writing an alternate-history game as proof: synthetic data is the lever that lets the narrative evolve without getting trapped by real history.
— Zhang FanThe real threshold in FDE is 85 points out of 100
Reaching 60 in enterprise AI is easy. But 70 is 10 times harder, 80 is 10 times harder still, 90 another order of magnitude. Yet the real pass line lands at 85. Below 85, you have scrap—zeros. Above 85, the value suddenly exists. FDE work demands mastery of everything: research chops, hands-on with hundreds of GPUs in training, shipped in real business. Those people cost over 1 million a year. Entry-level grads with grit? Can't carry this.
— Zhang FanBuild the boat, not the lighthouse
The base model is like an ever-rising sea level. Build a fixed lighthouse and it drowns. Build a boat that rises with the sea. Toss any problem at the model: the answer lands at the center of the normal curve—sensible middle, reasonable nonsense. What pays is the distribution tail, the edge cases. Musk and Jobs are scarce because they twisted the public curve. An enterprise's only moat is the special distribution stacked on top of the base model, independent of it. But the sea rises 100 meters a year. Now it's 100 meters a month.
— Zhang FanWeights matter more than workflows
Knowledge takes three forms: rules (tens to hundreds of dimensions), text (thousands), weights or neurons (trillions). Rules are most controllable but weakest in expressiveness. The stronger the form, the harder to control. An enterprise's master craftsperson's intuition can't be written as rules or workflows. It only sediments into weights. That's why Zhang Fan endorses the insight ‘it's not your weight, it's not your product’from a Sequoia partner—because only weights have ultimate expression. Workflows are interim tools, each one generated on the fly and discarded after use.
— Zhang FanEnterprises learn like AlphaGo Zero, not AlphaGo
Use the chess analogy for how enterprises should learn. Deep Blue beat Kasparov with rules and brute force. AlphaGo beat Lee Sedol, trained on 30 million human games. AlphaGo Zero didn't use human games—just played itself for three days—and swept AlphaGo 100–0, while compute plummeted from 1,000+ GPUs to 4 TPUs. This suggests human experience is heavy with noise. Enterprises shouldn't summarize experience into training data. Define your board and rules, then let the system play and evolve inside that frame.
— Zhang FanAI reverses the industrial revolution's organizing logic
The industrial revolution's core: human muscle could scale infinitely but brains couldn't, so muscle bowed to brain, spawning standardization and fine division of labor. Ford's Model T comes in black only—that compromise. AI lets brains scale infinitely too. Muscle and mind rebalance. Result: organizations, interfaces, business models all become liquid. But here's the key: SaaS ships at its peak and decays from there. Models ship broken and grow. Enterprises need to rewire from decay logic to growth logic.
— Zhang FanIn their own words · checked verbatim
Whenever the work is custom, capability accumulates in people instead of the enterprise. The moment it sits in people, there's no way to scale it.
只要是定制的活 它就把能力沉淀在人 而不是沉淀在企业 你只要沉淀在人 这件事就没办法规模化
Zhang Fan6:03
The two philosophies come down to whether we're centered on the enterprise or centered on intelligence.
两种哲学在于我们到底是以企业为中心 还是以智能为中心
Zhang Fan16:09
Why does it have to be rules fewer than data? If I had 10,000 rules generating 100 data points, wouldn't that avoid overfitting?
我为什么一定是规则要比数据少呢 如果我要是1万个规则 生成100条数据 它不就不overfitting了吗
Zhang Fan25:16
So 60 is 1,000 times worse than 90, right? But here's what matters—you have to know where the kill line is. It's probably 85. Below 85, everything is scrap, total zero. Above 85, it suddenly flips to one.
所以60比90差1000倍 对吧 但是这1000倍来讲 但是你要知道 这个里面的斩杀线在哪呢 可能是在85 就是85以下都是废材 全是零 85以上一下变成一了
Zhang Fan30:18
But there's one problem—even though you've made it stable and clear, you must know the sea level rises 100 meters a year. Now it's 100 meters a month.
但是你发现这里有一个缺点 虽然你稳了 你清晰了 但是你要知道海平面 每年上涨100米 现在是每个月涨100米
Zhang Fan48:24
‘It's not your weight, it's not your product.’Right—I think that's exactly it. Because only this has ultimate expression.
its not your weight its not your product 对 我觉得这个就非常对 就因为只有这个的表达力是极致
Zhang Fan1:02:27
You find that when training AlphaGo, they used roughly 1,000 GPUs, but when training AlphaGo Zero, they used 4 TPUs. In short, the computational scale shrank drastically.
你发现他训AlphaGo 的时候 他大概用了1000块GPU 但是训AlphaGo Zero的时候 用了4块TPU 总之算力是极大的缩小
Zhang Fan1:18:33
So what was the old one? The old organization ran on decay logic. The new organization runs on growth logic.
所以一个原来的是什么 原来的组织叫做腐化逻辑 现在的组织叫做 生长逻辑
Zhang Fan1:37:43
Figures
| Harvey valuation | 10+ billion USD | 18:12 |
| Anthropic's cutting-edge premium model revenue share | 11% | 41:22 |
| AlphaGo training compute | approximately 1,000 GPUs | 1:18:33 |
| AlphaGo Zero training compute | 4 TPUs | 1:18:33 |
| Top-tier FDE research team annual cost | over 1 million | 30:18 |
Glossary
- FDE
- Front-line Deployment Engineer; an engineer stationed at a customer site to customize and deliver AI capabilities, a delivery model originating from Palantir.
- Ontology
- A knowledge-modeling method that reduces unstructured data into structured triples; a traditional approach to knowledge graphs.
- RSI
- Recursive self-improvement; a mechanism where systems refine their own capabilities through self-learning and feedback loops.
- MoE
- Mixture of Experts; a model architecture where multiple expert subnetworks compose the full model, with only some activated per inference.
- Harness
- The scaffolding that wraps a model and chains together tool calls and capabilities; the runtime framework for agents.
How to listen
Founders, investors, and technical leaders commercializing enterprise AI, designing FDE delivery models, or thinking through AI product moats.
Minutes 36–39: commentary on AI hype cycles (shrimp farming, horses, etc.); low information density, skippable.