Nvidia Has No Monopoly, Only First-Mover Advantage — Take 90% of the Money and Someone Will Come Dig Up Your Foundation
Nvidia's moat is the CUDA software layer plus a hardware lead, but it has not formed a monopoly. The higher the gross margin, the stronger the incentive for model makers to build their own chips — swinging that shovel is partly about answering to the capital markets.
The video won't play here. Listen to the audio instead:
The argument · tap a timestamp to hear it
Taking 90% of the money means everyone wants to dig you up
Nvidia's moat is software plus hardware: the hardware is the GPU, the software is the CUDA middle layer that bridges human programming and the GPU. But it has not formed a monopoly, only a leading first-mover advantage. The problem is that when you run especially fast and take 90% of the world's money, the price is that everyone wants to settle that money with you, with a crowd of people chasing you holding guns. And the higher the gross margin, the greater everyone's incentive to chase. Model makers themselves have an interest in raising gross margins, whether through cost cuts or efficiency gains, and once the pressure reaches a certain point, that knife will naturally swing at Nvidia's ecosystem — even if it is only a test swing.
Swinging this shovel is about answering to capital
Why do model makers insist on developing their own chips even knowing how hard it is? This is not just a technology-selection question, it is even a question of how to answer to the capital markets. I want to swing this shovel to tell the capital markets that I am trying to do this, that I have not raised my hands in surrender and been monopolized by a single ecosystem. Amazon and Google have already proven that building your own chips is feasible, and the bigger the scale, the more money you save. So OpenAI also announced it would make its own chip, and the evaluation score is called Mexican chili. But this takes too much money — OpenAI says it will spend 45 billion US dollars to rent compute; can 45 billion build its own? Probably not enough.
An eight-card server is 14 million RMB in China spot
Take the fairly advanced B300 server as an example: 4T of memory, eight cards per machine, spot price in China over 14 million RMB. The same machine costs only just over 500,000 US dollars in the US, about six to seven million RMB, three times the price. There are multiple reasons for the high cost: US government export controls — the so-called defense is about preventing you from shipping it out. You can order on Amazon in the US and have it delivered to a home in Silicon Valley, but it cannot leave the country — it needs KYC, it cannot be China, and it needs IDC proof. When the machine is routed through layer after layer of forgery to Southeast Asia, say Thailand, the price basically doubles, and then breaking through layer after layer of defense from Thailand into the mainland, going through both sides, doubles it again.
Delivering 512 units means plugging in and lighting up 512 times
Spot trading is like housing, generally sold in powers of two: two units, four, eight, 128, 256. Someone ran into 512 units quoted at over 7 billion. The transaction process is: first you screenshot your bank account to prove you really have the money, and the other side has to shoot video proving it really has the goods, saying the company name and holding that day's newspaper. On delivery it is one unit at a time: plug in and light up to show it is good, wire the money, haul one away, then the next, repeated 512 times, which takes about a day — because lighting up one server with cards pulled takes half a day. The 80 billion Hong Kong dollars Alibaba raised is only enough to buy 5,000 units.
Futures are all fake, and 70% of spot is brokers
In the flood of futures inquiries, an 8 million quote is cheaper than the 14 million spot price, but nobody dares do it; now it takes full-payment letters of credit. A letter of credit is a bank guarantee, understandable as an ancient paper version of Alipay, used for bulk commodity trade, and it also needs a third-party warehouse: goods go into the third-party warehouse, the warehouse receipt goes to the bank, the bank pays on seeing the receipt, and the buyer takes the receipt to the warehouse to collect the goods. But the current play is: the buyer issues a full letter of credit, the seller takes the letter of credit and the buyer's documents to the bank as collateral, gets the money and goes to buy the goods. So futures are 100% all fake — as long as the goods arrive, there is no compensation risk, and you can sell spot directly, who wouldn't do it? Spot is also estimated to be 70% fake, with 70% of sellers being brokers.
What does a 14 million machine have to run to pay for itself
Do the math: 14 million buys eight cards; for training, one or two machines is definitely not enough, you need at least several hundred; for inference, how much does buying tokens cost? When does it pay back? It used to be said ten months to pay back, and now that has already come down. Some people saw the news of DeepSeek's IPO and said model makers have pretty high margins, seemingly reaching 60-70% — but that does not count the investment, does not count the servers; with hardware assets counted by amortization, plus the data center, power and maintenance costs, you might figure a 60-70 gross margin. The key is whether there is an endless stream of writable content to keep the cards fully loaded. Many businesses do not make money because there is too much idle time — just like a restaurant, if it can be full 24 hours a day, its margin is also very high.
A domestic GPU runs a top-tier model, Nvidia falls then rises
Zhipu released GLM 5.3 Flash; its score is not first, but it basically caught up with Opus 4.8, and it claims to have entered the top-tier ranks. The most important thing is not the score, but the line in the article: our compute runs on domestic GPUs, and after optimization, the computational efficiency of domestic GPUs has caught up with that of overseas GPU products. That sentence went viral. The timing of the news was also bad, landing during evening development hours. Nvidia fell only a little yesterday and rose over 6% today. The market is now either extremely optimistic or extremely pessimistic; one piece of news seriously affects the direction, and two days later it comes back.
Models split in two: top price for creativity, cheapest for repetition
Blind tests of models can no longer tell the difference, so people have started choosing on price. But on the enterprise side it actually splits into two scenarios: creativity-driven scenarios, such as writing code or doing long-horizon analysis and producing reports, where no matter how cheap it is you will not use it — you will definitely use the top tier, because the sunk cost is too high and redoing it after a bad job is not worth it, wasting both time and money. The other is work like content analysis, where the output directly determines enterprise revenue, so you must choose the model with a more reasonable ROI. It is like hiring people: for a cleaner, even one at 10,000 a month feels unnecessary no matter how good, so you find a passing one among the cheapest; for an engineer, you are willing to pay a higher salary to buy creativity. It is two ends — stand in the middle and you have neither side.
The official website is for AI to read; more words is an advantage
One report says 70% of internet traffic is already AI. If the official website is for AI to read, you must write it more complexly and give complete information to gain an advantage from the GEO angle. SEO used to be about stuffing keywords, but AI does not look at those; if it cannot summarize them it skips them, so you must write it clearly with a relatively complex and complete narrative logic, not afraid of too many words, but afraid you did not cover something. Technically, the old dynamic pages no longer work; AI scraping on the client side only gets an empty page, so you must use SSR backend rendering and put all the content in HTML returned in one go. There is another approach: the large model sends header information about acceptable content types, and if it is TXT plus markdown, you return pure markdown with nothing in it that could cause hallucination.
In their own words · checked verbatim
This is not a question of technology selection; it is even a question of how to answer to the capital markets.
这不是一个技术选型的问题,这甚至是一个如何向资本市场交代的问题。
So in this market right now, first, futures are all fake, 100% all fake, it is simply impossible for them to be real. Because he just turns around and sells spot.
所以现在这个市场第一期货全是假的,百分之百全是假的,就是不可能是真的。因为他转手就拿现货卖了。
A lot of businesses do not make money because they show too much time — just like a restaurant, if you can be full 24 hours a day, your rate is also guaranteed to be very high, very high.
很多的生意不赚钱,是因为他显的时间太多了,就跟饭馆一样,你要二十四小时丢人都能坐满车,你的率也保证的很高,很高很高。
Actually it is two ends, two ends, you just cannot stand in the middle; if you stand in the middle you have neither side — you have to be either more badass or cheaper.
其实就两头就是两头,你就不能站中间,你站中间就两边不靠,要比牛逼要便宜。
In the future, if your content or your website's interface is not agent-friendly, you will be eliminated from this set of game rules.
将来如果你这个内容或者你这个网站的接口对agent不友好,你就会被从这个游戏规则里干掉。
All product logic, all your engineering logic, will be flattened by AI.
所有的产品逻辑,你的工程逻辑都会因为ai就抹平了。
Figures
| B300 eight-card server spot price in China | over 14 million RMB | 6:01 |
| Same-configuration server price in the US | just over 500,000 US dollars (about six to seven million RMB) | 6:01 |
| OpenAI compute rental amount | 45 billion US dollars | 4:59 |
| Alibaba fundraising amount | 80 billion Hong Kong dollars | 10:10 |
| Quote for 512 machines | over 7 billion | 10:10 |
| Model makers' gross margin | 60-70% | 19:21 |
| Share of internet traffic that is AI | 70% | 39:14 |
| Nvidia's single-day gain | over 6% | 24:44 |
Glossary
- CUDA
- Nvidia's software middle layer connecting programming and the GPU, the core of its moat.
- KYC
- Identity verification under export controls, requiring proof of a non-Chinese buyer and an IDC location.
- Letter of credit
- A bank-guaranteed paper version of Alipay used in bulk trade, requiring a third-party warehouse and warehouse receipt.
- SSR
- The backend renders content into HTML and returns it in one go, making it easy for AI client-side scraping.
- GEO
- Content optimization for AI generation engines, replacing traditional SEO keyword strategy.
- MCP
- A scheme that wraps data services into a protocol for AI to call.
How to listen
Founders and engineers watching the compute supply chain, the GPU black market, enterprise AI deployment and product design logic.
The model evaluation and blind-test section in the middle (27:00-31:00) can be fast-forwarded.