Training Data with Shaun Maguire & Sonya Huang · 2026-06-30 · 双语整理

Hardware-Software Co-Design Is AI's Real 100x

软硬件协同设计才是 AI 真正的 100x —— SemiAnalysis 创始人 Dylan Patel
从住在汽车旅馆、在自家加油站给顾客"人肉预测"香烟牌子的小孩,到管着 $50M+ 捐赠硬件、被传年收入过亿美金的半导体研究机构创始人。Dylan Patel 的核心论点:芯片快一点、内核优一点、模型改一点,各是一个 2x;真正的爆发在把模型、内核、硅片放在一起协同设计——"instead of being multiplicative to 8x, it's actually 100x." 这解释了 DeepSeek 为什么长成 Hopper 的形状、TPU 为什么跑不动它、CUDA 护城河为什么从来不是 CUDA,以及 Jensen 为什么要自掏腰包制造一个多极世界。
来源:YouTube · Sequoia Capital · Training Data · 2026-06-30 · 70 min · 约 15,800 词 · 18 章 · 逐字双语对照
TL;DR · 速读

Co-Design, Living Benchmarks & a Multipolar World

  1. 跨层协同设计才是真正的 100x

    "Instead of being multiplicative to 8x, it's actually 100x, because you've optimized across all three layers."

    "The real breakthrough innovation is when you leapfrog a few layers, you co-optimize and co-design them."

    硬件、内核、模型各自优化是 2x×2x×2x=8x;把三层放在一起协同设计,增益直接跳两个数量级。

  2. DeepSeek 的 expert 是照着 Hopper 长的

    "If you look at the shapes of all the experts in DeepSeek V3, they were all optimized for Hopper."

    "Despite the fact that TPUs are objectively an amazing chip... TPUs suck at running DeepSeek, but they are really really great at running other kinds of models."

    模型架构会长成硬件的形状:V4 已经在为 Blackwell 和华为的芯片优化。硬件的选择反过来塑造模型。

  3. CUDA 护城河从来不是 CUDA

    "What people call the CUDA moat is not actually anything to do with CUDA."

    "DeepSeek, Kimi, Zhipu, Alibaba, Tencent — their models are co-designed for GPUs, and therefore if I want to run them on TPUs, in some cases they don't run really well."

    模型已经会自己写 kernel,软件护城河被部分拆掉;真护城河是整个开源模型生态都为 Nvidia 硬件协同设计。

  4. 同等质量的推理成本,一年降 60x

    "We've seen model cost drop for equivalent quality by like 60x a year."

    "The update cycle for most of these libraries is twice a week... it's a relentless breakthrough after breakthrough that keeps driving efficiency and cost down."

    所以基准测试必须"活着":InferenceX 每天在 $50M 捐赠硬件、约 15 种芯片上重跑所有最新模型,配置全公开。

  5. OpenAI 和 Anthropic 正被拖向不同硬件

    "The way OpenAI's models are headed, it would be a terrible decision for them to use TPUs."

    "And the way that Anthropic and Google's models are headed, it's actually a terrible decision potentially for them to train with GPUs."

    OpenAI 更稀疏、Anthropic 更稠密;matmul 单元尺寸、NVLink(72 卡有交换机)vs ICI(8000 卡无交换机)的拓扑差异,反过来塑造两家的模型架构。

  6. 一切都是吞吐-交互性曲线的下游

    "Most things in hardware, infrastructure, model, application layer — everything is downstream of that curve."

    "We see this with Anthropic, right? Claude Code fast mode costs way more than regular mode. Same with OpenAI's priority queue thing."

    同一块硬件,batch 100 慢速 vs 单用户极速是 4x 成本差;愿为速度付 4x 的,是"用 token 的人更贵"的场景。

  7. 推理会是比石油更大的市场

    "Inference of AI will be many percentage points of the GDP."

    "Use of tokens is going to be the biggest market, and the value that's created from tokens is going to be the biggest market... much bigger than oil."

    他预测 2030 年仅 OpenAI + Anthropic 合计就超 100 吉瓦;2040 年是太瓦量级,且过半增量算力会上太空。

  8. 算力荒的根源:模型扩 TAM 快过算力增长

    "The TAM for Fable 5 is not just 2x that of Opus... and yet compute in the world did not double in that same time frame."

    "But the demand for useful tasks that can be done by AI — the number of useful tasks and the value of them — has."

    只要模型能做的有价值工作扩张快过算力供给,价格就涨、crunch 就一直在。这是需求侧的结构性短缺。

  9. Anthropic 的飞轮:每块 GPU 立刻变成正毛利 token

    "Every GPU I rent, I can immediately turn around and sell tokens on it at a positive margin."

    "Anthropic in Q2 is profitable... their margins on an Opus 4.8 token is like north of 80% for the API price."

    毛利 75% 时算力成本翻倍也还剩 50%——所以他们敢按任何价格租 GPU,NOI 照样上涨。

  10. Jensen 在自费制造一个多极世界

    "A world where OpenAI, Anthropic and Google models are the only models is one in which he's screwed."

    "There's a reason he's blowing money on random AI labs... he wants to create a multipolar world. That's why he loves Chinese labs."

    今天 GPU 卖给谁都一个价;五年后 Crusoe / CoreWeave / NeoLab 的存在,意味着 TPU 和 Trainium 更弱、Nvidia 更强。

  11. AI 云把超大规模云的老本行变成负资产

    "The mechanics of the GPU rental market meant that a lot of the expertise of the hyperscalers fell away."

    "No one rents a single GPU in a 72-GPU rack. They rent the whole rack, and in fact they rent many of the racks... everyone has these long-term contracts."

    Nitro 网卡、租户隔离、自研网络这些 CPU 云的王牌,在 AI 云里要么无用要么直接拖性能——NeoCloud 的窗口因此存在。

  12. Gigawatt 之间并不平等

    "A gigawatt given to Anthropic is objectively worth more revenue than a gigawatt given to OpenAI."

    "Trainium sells at sub-$10 billion per gigawatt rental rate... GPUs usually went around $12 to $13 billion per gigawatt."

    数据中心层是二元的(有或没有);算力层同样一个 GW,租金和产出能差一倍以上。Google 甚至在 1GW 机房里塞 1.5GW 硬件"晃功率"。

  13. 人人都会做 ASIC,但小心局部最优

    "Some people will race to a local minima, and then the question is: how do you scoot back over to the absolute minima?"

    "You talk to people at labs — they literally don't know what architecture they're going to be doing in a year. They have many research bets."

    专用芯片赌的是架构不变。Google 同时养三条不同架构的 TPU 设计线、还按 $11/GPU/小时租 xAI 的卡,对冲的就是这个。

  14. "AI 没有 ROI"是他的触发词

    "The line has been up and to the right in terms of capabilities this entire time."

    "They're like, look, this benchmark didn't improve — that's cuz it's at 90%. Look at the new benchmarks — they're skyrocketing."

    用饱和的 benchmark 证明模型停滞,是他最受不了的论调;更搞笑的是"事实全对、结论全错"的人。

  15. 能源瓶颈有土办法:百万台柴油卡车引擎

    "Take the millions of diesel engines for trucks that the US has the capacity to make — you can very trivially convert them to be using gas."

    "You can just pull people out of car mechanic shops and have them run around and repair truck engines."

    "能源难"被过度神化:卡车引擎产线改燃气发电机 + 汽修工当运维,是被忽视的暴力解。

Chapter 01

The $100M Research Shop

工程师 × 对冲基金的斗兽场
00:00 — 01:58 · 冷开场 · SemiAnalysis 文化 · 营收传闻
Dylan Patel00:00:00

I think it's really fun inside of SemiAnalysis, because we have 90 people, and like a big chunk of them are technologists, engineers across the whole supply chain.

我觉得 SemiAnalysis 内部特别好玩,因为我们有 90 个人,其中一大块是技术专家——覆盖整条供应链的工程师。

Um, and then a big chunk is people who are formerly at hedge funds.

呃,另一大块是以前在对冲基金干过的人。

And you see these arguments — people are like, "Oh, well that doesn't matter." And then someone's like, "Well, but cost." And then the engineers are like, "No, no, no, but this technology is the coolest."

然后你就会看到这种争论——有人说"哦,那个不重要",接着有人说"可是成本呢",然后工程师们说"不不不,这个技术才是最酷的"。

And you see this organically, like, fight it out.

你会看到这一切自发地、真刀真枪地吵出结果。

Um, and we're pretty informal, and you know, given the fact that I was a forum moderator, you can imagine — the enjoying it.

呃,我们还挺不拘小节的,而且你想想我以前是论坛版主,你可以想象——大家乐在其中。

Shaun Maguire00:00:29

You don't wrestle with a pig because a pig enjoys it, right?

别跟猪摔跤,因为猪乐在其中,对吧?

Dylan Patel

Exactly.

正是。

Shaun Maguire00:00:50

We're here in the SemiAnalysis office with Dylan Patel. You know, I'm Shaun from Sequoia. My partner, Sonya Huang.

我们现在在 SemiAnalysis 办公室,和 Dylan Patel 坐在一起。我是 Sequoia 的 Shaun,这位是我的合伙人 Sonya Huang。

It's pretty insane what you've done.

你做成的这些事相当疯狂。

Semis five years ago were not very sexy in the West. They were sexy in the East, but people here in the West had kind of forgotten about them.

五年前,半导体在西方根本不性感。在东方它很性感,但西方这边的人几乎把它忘了。

You did not forget about them though. You went very long.

但你没忘。你重仓做多了它。

You created probably the premier research company in the space that's been educating the world on, you know, the state of the art — from very technical details to supply chain to the bigger picture.

你创办了这个领域大概是最顶级的研究公司,一直在给全世界科普最前沿的东西——从极技术性的细节,到供应链,再到大图景。

Um, there's rumors that SemiAnalysis recently passed 100 million of revenue. I don't know how accurate those are.

呃,有传闻说 SemiAnalysis 最近营收过了 1 亿美金。我不知道有多准。

Whatever the numbers are, you guys are crushing it.

不管数字是多少,你们都在碾压式地赢。

Dylan Patel00:01:30

It's as accurate as the information is. Yeah.

传闻的准确度,跟消息本身的准确度一样。哈。

Cool. You know, you never know.

酷。你懂的,谁说得准呢。

Shaun Maguire00:01:34

Um, there's also rumors that you might start a venture fund — like, you know, I hear all the time in the ecosystem people wanting, you know, affiliation with SemiAnalysis.

呃,还有传闻说你可能要做一支风险基金——我在圈子里总是听到有人想跟 SemiAnalysis 攀上关系。

You've built this trusted brand, and so whatever you do, it's working.

你建立起了一个被信任的品牌,所以不管你做什么,它都在起作用。

This is clearly, like, just the beginning of the journey for you. Congratulations on all of that.

这显然只是你旅程的开始。为这一切恭喜你。

Chapter 02

Motel Kid Origins

加油站里的第一个"神经网络"
01:58 — 03:11 · 汽车旅馆 · 家族生意 · 香烟推荐系统
Shaun Maguire00:01:51

But how did this happen? Like, my first question is: what is the background? How did you kind of get to where you are now?

但这一切是怎么发生的?我的第一个问题是:你的背景是什么?你是怎么走到今天这一步的?

Dylan Patel00:01:57

Well, well, when I was a young boy, you know, coming out of the womb… No.

嗯,当我还是个小男孩,刚从娘胎里出来的时候……开玩笑。

So, okay. So, I grew up in like a small business. My parents had a motel. We lived in the motel. We also had our gas station.

好,是这样。我在小生意家庭里长大。我父母开了一家汽车旅馆,我们就住在旅馆里,我们还有一个加油站。

So, you know, I was selling — you know, I joke a lot of times, the first neural network I trained was racially and visually profiling people based on when they enter the gas station, which cigarette to pick.

所以我从小就在卖东西——我经常开玩笑说,我训练的第一个神经网络,就是顾客一进加油站,靠人种和长相来"画像",预测他要买哪种烟。

Right, basically, you know, the cigarettes were all strewn across the top, and I was too short to actually, like, you know, reach them — and technically it wasn't legal to sell cigarettes at that age, but whatever.

基本上是这样:香烟全都摆在货架顶上,我个子太矮够不着——严格说那个年纪卖烟也不合法,但管它呢。

I had to move the step stool over to the right area.

我得提前把小板凳挪到正确的位置。

Shaun Maguire00:02:34

I started working my first job before it was legal, too. So — but it's good experience.

我的第一份工作也是在不到法定年龄时开始干的。所以——不过这是很好的历练。

Dylan Patel00:02:37

Well, I didn't get paid, right? It's a family business.

但我可没有工资,对吧?家族生意嘛。

Shaun Maguire

Same.

我也一样。

Dylan Patel00:02:39

Same. Um, but yeah, we had our motel, and then across the street was our gas station.

一样。呃,总之我们有汽车旅馆,马路对面就是我们的加油站。

So, you know, sometimes someone would walk in, and so like, if an old white lady with curly hair walked in, I'd move the ladder or the step stool over to where the Camels are.

有时候有人走进来——比如一位卷发的白人老太太进门,我就把梯子或小板凳挪到骆驼牌香烟那边。

And if, you know, different age, demographic, profession, you know, race, etc., I would move the step stool over.

不同的年龄、人群、职业、人种等等,我就把小板凳挪到对应的位置。

And I joke this is the first neural network I trained — because if I waited for them to tell me, I'd have to, like, move it over and then step up, versus just being ready.

我开玩笑说这是我训练的第一个神经网络——因为如果等他们开口,我还得再挪板凳、再爬上去;而我可以提前就位。

Um, so you know, menthols versus, you know, 100s, slims and all these things — you know, I joke that's the first neural network I trained.

呃,薄荷烟、100s、细支烟这些——我开玩笑说,那就是我训练的第一个神经网络。

Chapter 03

The Red Ring of Death

Xbox 红环、Reddit 版主与互联网学徒时代
03:11 — 06:42 · 硬件启蒙 · 12 岁当版主 · 星际争霸宗师
Dylan Patel00:03:11

But I grew up in family businesses. Um, lived in a motel.

总之我是在家族生意里长大的。呃,住在汽车旅馆里。

And it all really goes back to when it was, like, my 8th birthday. Um, my birthday's in May, and it was April when the Xbox 360 was announced.

而这一切真正的起点,是我 8 岁生日那年。我生日在五月,而 Xbox 360 是四月发布的。

Um, for my birthday, I didn't ask for a birthday gift. My parents asked what I wanted — I asked for it for Christmas.

那年生日我没要生日礼物。父母问我想要什么——我说我想把它当圣诞礼物。

Uh, we celebrated Christmas, but there was no way — at least at the time I thought there was no way — they would give me the Xbox 360 for Christmas. And so I asked on my birthday for it for Christmas.

我们家过圣诞,但我觉得——至少当时觉得——他们不可能圣诞节送我 Xbox 360。所以我在生日那天,预定了这份圣诞礼物。

Anyways, Christmas comes around, I get it.

总之,圣诞节到了,我拿到了。

Um, you know, fast forward a couple months, my cousin who lives in Alabama — they also lived in a motel — was going to come over for spring break, um, and we were going to just hang out at my house.

快进几个月,我住在阿拉巴马的表兄弟——他们家也住汽车旅馆——春假要过来玩,我们打算就在我家待着。

And he's in between me and my older brother in age. Brother's a bit more jockey, um, so he didn't really care too much about the Xbox. He played sometimes, but he didn't really care.

他年纪介于我和我哥之间。我哥更像运动型的,所以不怎么在乎 Xbox,偶尔玩玩,但真不上心。

Um, but my cousin — you know, I wanted him to think I was cool, right? You know, so I bragged many times on the phone. I was like, "Yeah, I got an Xbox."

但我那个表兄弟——我想让他觉得我很酷,对吧?所以我在电话里吹了好多次牛:"对,我有 Xbox。"

And then the Xbox broke. There was a hardware defect called the Red Ring of Death.

然后 Xbox 坏了。有个著名的硬件缺陷,叫"红环之死"。

Um, but long story short, I had to open it up and, you know, short the temperature sensor, and it fixed it.

呃,长话短说,我只好把它拆开,把温度传感器短接,就修好了。

Um, but there were many other tricks I tried first, and none of them worked.

在那之前我还试了很多别的偏方,全都没用。

Um, and so that's sort of how I got into hardware. It was like opening Pandora's box.

于是我就这样入了硬件的坑。就像打开了潘多拉魔盒。

By the time I was 12, I was, like, on these forums a lot, reading, uh, posting a lot — and this is around the time when Reddit ate all other forums.

到 12 岁时,我已经泡在各种论坛上,大量地读、大量地发帖——那正是 Reddit 把其他论坛都吞掉的年代。

And so I became a moderator of, you know, Android and Apple and Google, as well as, like, hardware — and was, you know, looking at Intel, Nvidia and AMD and all these other forums, right? Build-a-PC.

于是我成了 Android、Apple、Google 这些版块的版主,还有硬件版——同时盯着 Intel、Nvidia、AMD 那些论坛,对吧?还有装机版。

All these forums I was watching, reading, posting a lot, but some of them I was moderating a lot.

这些论坛我都在看、在读、在发帖,其中一些我还深度参与管理。

Um, and so, you know, smartphones — watching smartphones develop from, like, very simple to speed racing to being technologically more advanced than PCs, um, in many ways architecturally. And same with, like, you know, GPUs — just tracking and watching that, reading every comment.

然后是智能手机——看着智能手机从极简陋,到军备竞赛,再到在架构上很多方面比 PC 还先进。GPU 也一样——就这么追踪着、看着,读每一条评论。

Um, always having the economic tinge, because I grew up in a small business. So I was always looking at the economics, right?

而且永远带着一层经济视角,因为我在小生意家庭长大。所以我总是在看经济账,对吧?

There was a time where all the — I'd say neckbeards — on the internet loved AMD GPUs. And like, I personally had bought an AMD GPU too, because price-performance.

有段时间,网上所有——姑且叫"技术宅"吧——都爱 AMD 显卡。我自己也买过一块 AMD 显卡,因为性价比。

But then when it came down to, like, what's technically better, I'd always be like: no, no, no, Nvidia is better — because they use a smaller chip to get, you know, better performance at better power efficiency, and their margins are better.

但真论技术上谁更强,我总是说:不不不,Nvidia 更好——因为他们用更小的芯片做出更好的性能、更优的能效,而且利润率更高。

And so, like, I would always, like, talk about how Nvidia's margins were better than AMD's in the GPU landscape. And so it's, like, very fun.

所以我总在聊 Nvidia 在 GPU 市场的利润率怎么比 AMD 高。特别好玩。

Sonya Huang00:05:32

And you were 12 at the time?

你那时才 12 岁?

Dylan Patel00:05:32

I started moderating when I was 12, but this is all through my teenage, tween-age and high school years, right?

我 12 岁开始当版主,但这贯穿了我整个少年、青春期和高中时代。

Sonya Huang00:05:39

Do you have any other weird hobbies, or was it just semis?

你还有别的奇怪爱好吗,还是只有半导体?

Dylan Patel00:05:42

I played a ton of StarCraft. At one point, I was Grandmaster on the North American ladder. StarCraft 2.

我打了海量的星际争霸。一度打到北美天梯宗师段位。星际争霸 2。

Shaun Maguire

Very serious.

非常硬核。

Sonya Huang00:05:48

So you've gotten just obsessively good at multiple things.

所以你在好几件事上都练到了痴迷级的水平。

Dylan Patel00:05:51

Yeah. I mean, it's — obsession is good.

对。我是说——痴迷是件好事。

Sonya Huang00:05:53

How were your grades?

你成绩怎么样?

Dylan Patel00:05:54

Um, they were decent. Um, I would say, like, I had mostly A's, but there are classes that I, like, thought were really boring or, you know, I just didn't enjoy.

呃,还不错。基本都是 A,但有些课我觉得特别无聊,或者纯粹不喜欢。

Um, like Spanish, I got, like, not the greatest grades. Um, you know — but I speak fluent Spanish, by the way, so it's really dumb.

比如西班牙语,分数就不太好看。不过顺便说一句,我现在西班牙语说得很流利,所以这事挺荒唐的。

Shaun Maguire00:06:17

Maybe that's why you didn't get a good grade.

说不定这就是你分数不高的原因。

Dylan Patel00:06:18

I didn't learn Spanish till later, to be fair.

公平地讲,我是后来才学会西班牙语的。

But yeah, so sort of my grades were fine, right? Like, I mean, they were fine enough for Asian parents.

总之我成绩还行,对吧?至少在亚洲父母的标准里过得去。

I was better than most of school, but you know, it wasn't like, you know, try-hard maxing for, like, you know, all A's.

我比学校里大多数人强,但也没到那种为了全 A 拼命卷的程度。

Chapter 04

From Quant to Founder

被 doxx 出来的 SemiAnalysis
06:42 — 09:16 · 被抢功的 quant · 2020 谷底 · 24 岁生日发博客
Sonya Huang00:06:33

Okay. So you're very much a student of the internet, then — this is how you developed this expertise.

好。所以你完全是互联网教出来的学生——你的专业能力就是这么练出来的。

At what point did you decide to start SemiAnalysis, and what's been the biggest surprise since starting the company?

你是什么时候决定创办 SemiAnalysis 的?创业以来最大的意外是什么?

Dylan Patel00:06:42

Yeah, so I went to school, I got a few degrees in stuff that wasn't related to semiconductors. Um, was a quant for two years at a small quant risk firm.

嗯,我上了大学,拿了几个跟半导体无关的学位。然后在一家小型量化风险公司做了两年 quant。

Um, and then basically, you know, there was a culmination of events that happened, right?

然后基本上,一连串事情叠加在一起爆发了。

One was that I got screwed out of a bonus. I'd made my company many millions of revenue — of risk-free revenue — because I exploited, like, a risk thing in the market. Um, you know, I think well over 10 million.

第一件:我的奖金被坑了。我利用市场上的一个风险机制,给公司赚了好几百万的无风险收入——我记得远超 1000 万。

And then someone else took credit for my work and all this sort of stuff. But eventually I did get right-sized. But you know, I lost the social contract with the company I was working with.

结果别人把我的功劳抢走了,诸如此类。最后待遇虽然补正了,但我和这家公司之间的"社会契约"已经破裂了。

Um, add some — you know, my grandparents grew up in my house with us, right? Were in the motel with us. Uh, they lived with us, and so, you know, I was very close with them.

再加上——我的祖父母一直和我们同住,就住在汽车旅馆里。他们跟我们一起生活,所以我和他们非常亲。

And my grandmother got dementia, and she forgot who I was, and she fell down some stairs and had, like, a tragic accident and passed away.

我祖母得了失智症,忘了我是谁,后来从楼梯上摔下来,出了不幸的意外,去世了。

So all of that happened in early 2020. Um, additionally there were some, like, you know, girl things.

这些全发生在 2020 年初。呃,另外还有一些感情上的事。

And so, you know, there were a few things that happened that made me, like, kind of very sad. Um, and so all of those things sort of culminated.

所以好几件事凑在一起,让我非常低落。所有这些事叠加到了顶点。

Then COVID happened, and my brother's like, "Dude, just come stay with me." He lived in Nashville, so I came and stayed with him in Nashville.

然后新冠来了,我哥说:"兄弟,来跟我住吧。"他住在 Nashville,我就搬过去和他一起住。

We were like, "Oh, lockdowns will be a few weeks. You can stay with me while they happen, and then you can go back home, and you know, whatever." Famous last words. Lockdowns lasted much longer.

我们当时想:"封控也就几周,你先住我这儿,结束了再回家呗。"经典的 flag。封控持续了长得多。

But, you know, living with my brother for a few months — you know, it was sort of like, okay, I didn't know what I was doing. I was now at my brother's home. Um, everything was his rules.

跟我哥住的那几个月——那种状态就是:我不知道自己在干嘛,住在哥哥家,一切都得按他的规矩来。

You know, him and his fiancée at the time — now wife — you know, were there. And so, like, I basically had to tiptoe around. But I didn't care about my job.

他和他当时的未婚妻——现在的妻子——都在,所以我基本得踮着脚过日子。而且我已经不在乎那份工作了。

And so I was, like, posting even more than normal. I'd always been posting a lot on the internet. I'd always been trading stocks a lot.

于是我发帖比平时更凶了。我一直是网上的高产发帖人,也一直在大量炒股。

But like, I made a lot of money shorting COVID and going long in COVID and, like, all this stuff. Semiconductor shortages happened around then too.

新冠期间我做空、再做多,赚了不少钱。半导体短缺也差不多是那时候发生的。

And anyways, I was, like, very much obsessed with posting and things like that.

总之,我当时对发帖这件事非常痴迷。

And eventually, um, around that time, I got into an argument with someone on the internet, and they doxxed me, right? They publicly revealed my identity for my anonymous account.

终于,大概在那段时间,我在网上跟人吵架,对方把我 doxx 了——公开曝光了我匿名账号背后的真实身份。

And at the time I was like, "Oh no, I'm scared." I stopped posting for, like, three weeks. And then I was like, "What am I doing? Why do I care?"

当时我吓坏了,停更了大概三周。然后我想:"我在干嘛?我为什么要在乎这个?"

So then I just started posting under — I had, like, blogs and stuff as well — I made a real blog, SemiAnalysis, and on my 24th birthday, I posted, um, you know, two blogs.

于是我开始用真名发——我之前也写过一些博客——我做了一个正经博客,就叫 SemiAnalysis,在我 24 岁生日那天发了两篇文章。

And then from there it just — it was not a newsletter, but I got so much traction, because now instead of posting under an anonymous name, it was a real name, and I put a lot more effort into those two posts than I usually did.

从那以后——它当时还不是 newsletter,但增长势头特别猛,因为不再是匿名马甲,而是真名实姓,而且那两篇比我平时用心得多。

Instead of, like, posting on the internet, it was, like, real effort into the blog.

不再是随手网上灌水,而是真正下功夫写博客。

Um, you can actually go back and read those if you want. They're not that great, but you know, they were good for the time. They were the best stuff you could find on the internet about semis.

你们现在还能翻回去读那两篇。写得不算多好,但在当时算好的了——那是当时互联网上能找到的关于半导体最好的内容。

Um, and I just kept posting, posting, posting. I started getting a lot of consulting business.

然后我就不停地发、发、发。咨询生意开始源源不断地来。

Chapter 05

The Homeless Research Roadtrip

国家公园、40 场会议与 SPIE 深渊
09:16 — 14:04 · $30 汽车旅馆 · 一年 40+ 会议 · 供应链黑话
Dylan Patel00:09:16

You know, 2020, I also sort of — I was again crashing out, didn't know what I wanted to do.

2020 年,我又一次处在崩溃边缘,不知道自己想干什么。

So I, uh, packed everything up — or sort of, I took my truck, I bought a tent that fits on the back of the truck, um, bought an air mattress, whatever, and would, like, drive around all these national parks all around America.

于是我收拾好一切——开上我的皮卡,买了一顶能架在车斗上的帐篷、一张气垫床什么的,开着车把全美国的国家公园逛了个遍。

And so, like, two or three or four days of the week, I'd stay in a random motel where I negotiated the price to be, like, $30 a night for a room, and I would work on stuff.

一周里有两三四天,我会住进某家随便找的汽车旅馆——我把房价砍到一晚 30 美金——然后干活。

And then the weekends, I'd read books — and oftentimes read textbooks — um, while in some random national park, or hiking, and listen to audiobooks, um, about semiconductors, about AI, about all the things that I cared a lot about.

周末就在某个国家公园里看书——经常是啃教科书——或者边徒步边听有声书,内容全是半导体、AI,以及所有我特别在乎的东西。

And got way more educated over these six months where I'm just, like, going to every national park.

在这逛遍国家公园的六个月里,我的知识储备暴涨。

Um, and the whole time I was alone. The whole time I was posting blogs. Um, everyone was like, "Dude, what the f— are you doing?"

全程一个人。全程在发博客。所有人都在问:"哥们,你到底在搞什么?"

Shaun Maguire

Pre-Starlink, or the very early days of Starlink?

那是 Starlink 之前,还是 Starlink 刚起步的时候?

Dylan Patel

Pre-Starlink, pre-Starlink. Um, yeah, so it was very much like, "What are you doing?"

Starlink 之前,完全在那之前。所以大家真的都在问:"你在干嘛?"

Um, I traveled around Latam — again, like, for a year initially with my friend, and then with my ex, you know, for about a year.

后来我又去拉美转了一圈——先是和朋友玩了一年,然后和前任又是差不多一年。

And then — end of '21, '22, '23 and '24 — I'm still completely homeless since mid-2020, right? Um, but I'm traveling around to every conference in the world.

然后是 21 年底、22、23、24 年——从 2020 年年中起我一直处于完全"无家可归"的状态,但我在满世界跑会。

I go to 40-plus conferences a year, no matter where in the supply chain it is. I'm like, "Oh, that looks interesting. I guess I'll go to that."

我一年参加 40 多场会议,不管它在供应链的哪个环节。我就是:"哦,这个看着有意思,那去呗。"

And I'm like — I went to one conference, like, "Wow, this is amazing. You get to talk to the experts, and they're going to talk to you." And then you're so excited.

我去了第一场会议就觉得:"哇,太棒了。你可以直接跟专家聊,而且他们真的愿意跟你聊。"然后你就特别兴奋。

And in the case of semiconductors, everyone's a boomer. So it's like — they don't see young people who are, like, excited about it, so they're really happy to tell you stuff. And so you just have to ask.

而且半导体这行全是老一辈。他们很少见到对这行兴奋的年轻人,所以特别乐意倾囊相授。你只需要开口问。

Shaun Maguire00:10:51

On this — was there, like, a part of the supply chain, or one of these conferences, that, you know, particularly changed your view of the semi world, or that you felt then or feel now is particularly underrated?

说到这个——供应链里有没有哪个环节、或者哪场会议,特别地改变了你对半导体世界的看法?或者你当时觉得、现在也觉得被特别低估的?

Dylan Patel00:11:06

I think the trade shows and conferences range really widely.

我觉得展会和学术会议的差异跨度极大。

Um, obviously some of the ones I have the most fun at, you know, include NeurIPS. Why is that? Because it's 20,000 AI researchers, and they're generally in my distribution of age range. So it's, like, a lot of fun — but they're also, like, leading AI researchers, and it's a lot of fun and you learn a lot. Um, there's also a lot of parties.

显然我玩得最开心的会议里有 NeurIPS。为什么?因为那是两万名 AI 研究员,而且年龄段基本和我重合,所以特别好玩——他们又都是顶尖的 AI 研究员,好玩之余还能学到很多。派对也很多。

And then it ranges all the way to, like, you know, there's a random chemical conference in Japan where it's 300 Japanese dudes. It's like 20 guys from ASML, 20 guys from TSMC, 20 guys from Intel, and those are the only people who speak English. Uh, everyone else speaks only Japanese. And you're like, huh — I guess they're still pretty interesting and fun.

而光谱的另一端,是日本某场没什么人知道的化学会议:300 个日本大叔,ASML 来 20 人、TSMC 来 20 人、Intel 来 20 人——只有这些人说英语,其他人只讲日语。你心想:嗯,好像也挺有意思的。

I think, like, one thing that I have, like, a skill set of, is I'm able to bond with anyone regardless of their background and, like, who they are. I'm able to talk to them, find something interesting to talk about. Oftentimes it's the tech stuff.

我觉得我有一项技能:不管对方什么背景、什么身份,我都能跟他们建立联系,找到有意思的话题聊。多数时候聊的就是技术。

And so I think, like, the most interesting conferences are oftentimes, like, you know, the really big ones, because that's where the biggest stuff is happening.

最有意思的会议往往是那些特别大的,因为最大的事都发生在那里。

Um, but I think the niches that are really, really exciting — it's like, you know, SPIE. Um, so there's IEEE, which is international electrical engineering something, um, and there's SPIE, which is another ecosystem.

但真正让人兴奋的是那些小众领域——比如 SPIE。有 IEEE,国际电气工程什么的;还有 SPIE,是另一个生态。

SPIE conferences are super, super deep in details. Every single one that I went to — especially, like, SPIE Advanced Lithography or SPIE Photomask — I went to them the first time, I didn't even understand 90% of what I heard.

SPIE 的会议在细节上深到不可思议。我去过的每一场——尤其是 SPIE 先进光刻或 SPIE 光掩模——第一次去,听到的内容 90% 都听不懂。

And then I read, read, read — I had made some context, of course — and then next time I went, I understood, like, half of what I went to. Third time I went, I understood, like, 75%.

然后我就读、读、读——当然也积累了些上下文——第二次去能听懂一半,第三次能听懂 75%。

Even now, I went and I was like, I still don't understand everything that's going on.

即使现在再去,我仍然没法完全听懂所有内容。

Whereas, like, you go to, like, NeurIPS, you know, a couple times, you can understand — okay, what's neurosymbolic reasoning, okay, what's this, what's that — like, you can kind of get a mapping of what everything is pretty quickly.

相比之下,NeurIPS 去个几次,你就能搞明白——哦,什么是神经符号推理,这是什么,那是什么——很快就能在脑子里画出一张全景图。

But some parts of the supply chain are so arcane and so deep and so technical, it takes a lot of time for you to even understand what's happening.

但供应链的某些环节是如此晦涩、如此深、如此技术化,光是搞清楚"正在发生什么"就要花很长时间。

And you know, you go to a conference for a few reasons, right? You understand the research — it's all the research that's being published — but what you really care about is understanding: how does that research intersect with technology? Also, how does that research differ from what's there today?

去开会有几个目的:了解正在发表的研究是一方面,但你真正在乎的是——这些研究和产业技术怎么交汇?它和今天已有的东西差在哪?

And none of these research papers tell you what's happening today. But then you just ask people, and you build contacts, and you learn.

论文可不会告诉你产业界今天在发生什么。但你只要去问人、积累人脉,你就学到了。

And then you, like, learn about the supply chain — and oh, this company supplies this company, even though it's not publicly stated anywhere. Or, like, you know, you learn that this chemical costs about this much, and a tool uses about this much.

然后你就摸清了供应链——哦,原来这家公司在给那家公司供货,虽然任何公开渠道都没写。或者你了解到这种化学品大概什么价、一台设备大概用多少。

Shaun Maguire00:13:28

You hear the horror stories, of like: this chemical had a shortage, and it totally threw off this part of the supply chain, and then it turns out there's only three companies in the world that make that chemical.

你还会听到那些恐怖故事:某种化学品短缺,把供应链的某个环节彻底打乱,然后你才发现全世界只有三家公司生产它。

Dylan Patel00:13:41

My favorite one is — I learned, uh, from a Japanese guy at that specific Japanese conference that I went to where almost no one spoke English — in very broken English, he told me about how his father worked in this industry in the 1980s.

我最爱的一个故事:就在那场几乎没人说英语的日本会议上,一位日本人用非常蹩脚的英语告诉我,他父亲上世纪 80 年代就在这个行业干。

The only factory in the world that built this chemical, uh, burned down, and that caused memory prices to, like, double or triple.

当年全世界唯一生产某种化学品的工厂烧毁了,直接导致存储芯片价格翻了两三倍。

And I was like, wow — not too different from today.

我当时就想:哇,跟今天也没差多少嘛。

Shaun Maguire00:14:02

Not at all. Crazy.

一点没差。疯狂。

Chapter 06

Bigger Than Oil: InferenceX

比石油更大的市场,与一个活着的基准
14:04 — 18:14 · tokenomics · $50M 捐赠硬件 · 60x/年降本
Shaun Maguire00:14:05

Inference going to be the biggest market on Earth, biggest market beyond Earth — agree or disagree?

推理会成为地球上最大的市场、地球之外最大的市场——同意还是不同意?

Dylan Patel00:14:11

Um, I mean, obviously use of tokens is going to be the biggest market, um, and the value that's created from tokens is going to be the biggest market.

呃,显然,token 的使用会是最大的市场,token 创造出的价值也会是最大的市场。

But I think tokenomics — sort of the use of tokens, adoption of AI — sort of is the most important thing that's happening.

但我认为 tokenomics——token 的使用、AI 的普及——差不多是当下正在发生的最重要的事。

And inference, whether it's open models or closed models, will be, like, one of the biggest markets in the world. Much bigger than oil, I think — much bigger than, like, you know, many other parts.

而推理,不管是开源模型还是闭源模型,都会是世界上最大的市场之一。我认为比石油大得多——比很多其他行业都大得多。

Like, inference of AI will be, you know, many percentage points of the GDP. Yeah, right.

AI 推理会占到 GDP 的好几个百分点。真的。

Shaun Maguire00:14:35

What you've done with InferenceX, I think, is, you know, industry standard.

你们做的 InferenceX,我认为已经是行业标准了。

Maybe say a word on why you started it, what it does, and, you know, what do people misunderstand about, uh, performance benchmarking on inference?

要不讲讲你为什么做它、它是干什么的,以及大家对推理性能基准测试有哪些误解?

Dylan Patel

Yeah. So, to zoom back, right — like, SemiAnalysis, uh, we do a lot of stuff that's, like, you know, a lot of it is research for institutional clients and our subscriptions versus products.

好,先拉远一点看——SemiAnalysis 做的很多事,是给机构客户的研究,以及我们的订阅和产品。

But a lot of it is also, like: hey, you know, this would just be cool to figure out. Let's figure out how to figure it out, and just post it publicly. And that gets, you know, more and more scale.

但还有很多是:"嘿,把这事搞明白会很酷。那我们就想办法搞明白,然后直接公开发出来。"这类东西的规模越滚越大。

And so we've done this with a lot of GPU benchmarking and testing, and training performance and inference performance.

我们在 GPU 基准测试上做了很多这样的事,训练性能、推理性能都测。

But you know, ultimately we saw, like, inference benchmarking was, like, point-in-time. You know, you test it, and you take some time, you release it, and it's, like, slow and arcane and outdated — because models change all the time.

但我们最终发现,推理基准测试都是"时点式"的:你测一轮、花些时间、发出来——又慢又晦涩,而且发出来就过时了,因为模型一直在变。

Every — I feel like every week there's a new model, whether it's a Chinese model, or, you know, today Mythos 5, Fable dropped. And new models are coming out all the time.

我感觉每周都有新模型,要么是中国模型,要么像今天,Mythos 5 和 Fable 就发布了。新模型源源不断。

Um, on the software layer, uh — PyTorch, vLLM, SGLang, um, new drivers, new something drops, you know. In fact, the update cycle for most of these libraries is twice a week.

软件层也一样——PyTorch、vLLM、SGLang,新驱动、新东西不停地发。实际上这些库大多数的更新周期是一周两次。

So you basically have the software updating all the time, and therefore performance changing.

所以软件基本上一直在更新,性能也就一直在变。

Um, you know, new inference optimizations are coming out, and those get updated. And so I feel like it's a relentless breakthrough after breakthrough after breakthrough that keeps driving efficiency and cost down.

新的推理优化不断冒出来、不断被更新进去。我感觉这是一场无休止的突破接突破,持续把效率推高、把成本压低。

Which is why we've seen, you know, model cost drop for equivalent quality by, like, 60x a year. It's incredible.

这就是为什么我们看到,同等质量的模型成本一年能降 60 倍。难以置信。

Um, but to stay on top of that, you can't have point-in-time benchmarking. You need to have benchmarks be living and breathing — i.e., you know, constantly running on the latest hardware, on the latest models.

要跟上这个节奏,时点式的基准测试就不行了。基准必须是活的、会呼吸的——也就是持续跑在最新的硬件、最新的模型上。

And so we embarked on a project, and we got a lot of buy-in from the ecosystem.

于是我们启动了这个项目,并拿到了生态里大量的支持。

This was only possible because we had, you know, enough aura with some of the ecosystem, where we were able to get CoreWeave and Crusoe and Nebius and Oracle and Microsoft and Amazon and Google and OpenAI to contribute to us, um, compute.

这事能成,是因为我们在生态里攒了足够的"光环"——CoreWeave、Crusoe、Nebius、Oracle、Microsoft、Amazon、Google、OpenAI 都愿意给我们捐算力。

And then we were able to work with SGLang and vLLM — and now Radix Arc and InRact, uh, which are the private companies who are sort of leading those efforts, um, the open-source efforts — to collaborate with us.

然后我们又拉来了 SGLang 和 vLLM——现在还有 Radix Arc 和 InRact,就是在背后主导这些开源项目的私人公司——跟我们合作。

We were able to get Nvidia and AMD and Google and Amazon now — because we're adding TPUs and Trainium — uh, to collaborate.

Nvidia、AMD,还有 Google 和 Amazon——因为我们正在加入 TPU 和 Trainium——现在也都参与协作。

Now we've got all these people collaborating. We've got over $50 million of hardware, uh, donated to us. Um, once we launch TPUs and Trainium, it actually should be over $100 million of hardware.

现在所有这些人都在协作。捐给我们的硬件已经超过 5000 万美金。等 TPU 和 Trainium 上线,实际会超过 1 亿美金。

Um, you know, maybe about, like, 15 different chip types, all running these benchmarks every single day on all the latest models, right?

大概 15 种不同的芯片,每一天都在最新的模型上跑这些基准。

The best model from Moonshot, the best model from Alibaba — um, there's about five different Chinese models, the best open-source models, the best Chinese labs there. We run benchmarks on their models every day. And then also the best US open-source models — um, GPT-OSS, Nemotron, etc.

Moonshot 最好的模型、Alibaba 最好的模型——大概五个中国模型,最好的开源模型、最好的中国实验室都在里面,我们每天都在测。还有美国最好的开源模型——GPT-OSS、Nemotron 等等。

So we're running these benchmarks every day, um, in an automated fashion, and they run on these servers that are dedicated to us for inference benchmarking. And we sweep across so many different configurations and optimization types.

这些基准每天自动化地跑,跑在专门给我们做推理基准的服务器上。我们会扫过非常多不同的配置和优化类型。

And then what it creates is — and all the results are public, and all the configurations are public — so now we have the Pareto optimal curve.

它产出的东西是——所有结果公开、所有配置公开——于是我们就有了帕累托最优曲线。

Because a lot of, you know, times when people are comparing inference performance, they're, like, taking a suboptimal curve or point for someone else and comparing it to their optimal one.

因为很多时候人们比较推理性能,是拿别人的次优曲线或次优点,来对比自己的最优点。

And it's like — well, yeah, if I drove a Porsche versus, like, some race car driver, obviously I'd drive it slower. The same thing with inference benchmarking.

这就好比——同一辆保时捷,我开肯定比职业车手开得慢。推理基准测试也是一个道理。

And so what we did is we created open-source, uh, basically containers for the optimal points across every, uh, point on the interactivity — i.e., how fast is it responding to me — versus, you know, batch size — i.e., how many users am I simultaneously serving — curve.

所以我们做了一件事:把"交互性(它回我有多快)vs 批大小(我同时服务多少用户)"曲线上每个最优点,都做成了开源的容器。

And so now anyone who wants the optimal point can just go to InferenceX, download it, and run that as the optimal point. And they can check every day if they want, or they can even auto-download the most optimal point for that model — and their inference performance will be near peak.

现在任何人想要最优点,直接上 InferenceX 下载,跑起来就是最优配置。愿意的话可以每天查,甚至自动拉取该模型的最优点——推理性能就能接近峰值。

Chapter 07

The Curve Everything Is Downstream Of

吞吐-交互性曲线:AI 基建的中心图形
18:14 — 20:28 · batch vs latency · fast mode · 4x 价差
Sonya Huang00:18:12

Is that curve, like, the most important curve in your opinion? The throughput-interactivity curve is the most important one?

在你看来,那条曲线是最重要的曲线吗?吞吐-交互性曲线是最重要的那条?

Dylan Patel

Yeah, I think, um, most things in hardware, infrastructure, uh, model, application layer — everything is downstream of that curve, right?

对。我认为硬件、基础设施、模型、应用层的大多数东西——一切都是那条曲线的下游。

Is it something that needs to be super, super fast, super low latency? Um, and I don't really care about the cost — so I make batch size very low, and I use techniques like speculative decoding or multi-token prediction heavily. And there's so many, you know, possible techniques there.

这个任务是不是必须超级快、超低延迟?如果我不在乎成本——那我就把批大小压得很低,重度使用投机解码、多 token 预测这些技术。那边可用的技术非常多。

Or is it something where, actually, I'm batch processing a ton of documents, and I don't really care about all these things?

还是说,这个任务其实是批量处理一大堆文档,我根本不在乎那些?

I don't use these techniques that actually are worse on cost efficiency but help you with speed for an individual user — because I just want to pack a bunch of users. I don't care if the document takes all night to process, right?

那我就不用那些牺牲成本效率来换单用户速度的技术——因为我只想把一堆用户打包塞进去。文档处理一整晚我都无所谓。

Um, and right now, the way we treat AI infrastructure — it's, like, one-size-fits-all.

而现在我们对待 AI 基础设施的方式,是"一刀切"。

But over time, we're going to get to the point where, you know, there's stuff where you have batch workloads, or, you know, you need instant response — and there's the whole curve that's going to matter for, uh, users.

但随着时间推移,我们会走到这一步:有的场景是批处理负载,有的需要即时响应——对用户来说,整条曲线都会变得重要。

And so we see this with Anthropic, right? Claude Code fast mode costs way more than regular mode. Um, and same with OpenAI's priority queue thing.

我们已经在 Anthropic 身上看到了:Claude Code 的 fast mode 比常规模式贵得多。OpenAI 的优先队列也是一样。

Sonya Huang00:19:17

Sorry, dumb question — how does cost factor into the chart?

抱歉,问个蠢问题——成本是怎么进入这张图的?

Dylan Patel

So if — let's say, imaginary example — I have a batch size of 100, okay? And I can do 10 tokens per second per user. So in total, I'm doing a thousand tokens per second, uh, off of that one piece of compute. That's one side of the curve — super slow, 10 tokens per second.

假设——举个想象中的例子——我的批大小是 100,每个用户每秒出 10 个 token。那这一块算力总共每秒产出 1000 个 token。这是曲线的一端——对单个用户超慢,每秒 10 个 token。

Um, you know, the other side is: I have, uh, 500 tokens per second, but I only have one user. And so maybe 250 tokens per second, one user.

另一端是:我能跑到每秒 500 个 token,但只有一个用户。或者说每秒 250 个 token,单用户。

And then there's points in the middle that are more Pareto optimal, right? The average person actually wants, like, 50 or 100 tokens a second, and maybe, you know, the number of users I can batch together.

中间还有一些更接近帕累托最优的点。普通人实际想要的是每秒 50 到 100 个 token,再配上一个我能凑起来的批量用户数。

So the curve is: okay, a thousand tokens total per second, or 250 tokens total per second, depending on how many users I batch. And there's a curve in the middle.

所以曲线就是:总吞吐每秒 1000 个 token,还是每秒 250 个,取决于我把多少用户打包在一起。中间是一条连续的曲线。

And so ultimately, some workloads will actually want the 4x cost decrease — because the same unit of hardware can do a thousand versus 250.

最终,有些负载就是想要那 4 倍的成本下降——因为同一块硬件能跑 1000 而不是 250。

And some users — I'll pay 4x more, because I don't care about the price, I care about time. Because the person using the tokens is expensive, or the feedback loop that I have here is expensive.

而另一些用户——我宁愿多付 4 倍,因为我不在乎价格,我在乎时间。因为用这些 token 的人本身很贵,或者这里的反馈闭环很贵。

Chapter 08

Space & Intelligence per Watt

太空数据中心与每瓦智能
20:28 — 23:20 · 2030 100GW · 2040 太空过半 · 40x/年
Shaun Maguire00:20:18

If you had to guess — you choose the time frame, 10 years or 15 years — what percent of inference compute do you think will happen in space? Can be 0%, 50%...

如果让你猜——时间范围你自己选,10 年或 15 年——你觉得会有百分之多少的推理算力发生在太空?可以是 0%,可以是 50%……

Dylan Patel

Shaun — 99%!

Shaun——99%!

Um, this is a tough one. Um — you choose the time frame, like, whatever time frame.

呃,这题不好答。时间范围随便你定。

So I think the non-consensus — or at least against-SpaceX — thing… you know, I love SpaceX by the way, and I totally would buy the IPO if I could buy stocks.

我要说一个非共识的——或者至少是"不利于 SpaceX"的判断……顺便说,我爱 SpaceX,如果我能买股票,IPO 我肯定会买。

Shaun Maguire

Not investment advice.

不构成投资建议。

Dylan Patel00:20:48

Not investment advice. Thank you. Thank you. Not investment advice — from either of us.

不构成投资建议。谢谢,谢谢。我们俩说的都不构成投资建议。

Um, I don't think that space data centers will really matter in the next, um, you know, 3 to 5 years.

呃,我不认为太空数据中心在未来 3 到 5 年内真的重要。

Um, with that said, I think in, you know, 20 years, I think the vast majority of compute will be going in space.

但话说回来,我认为 20 年后,绝大多数算力会上太空。

Um, and so the real factor there is sort of, you know: what's the cost — it's the time frame, it's the cost of building power on terrestrial land, and how much power you're going to be able to do on terrestrial land.

真正的决定因素是:成本——时间范围、在地面建电力的成本,以及地面到底还能扩出多少电力。

And I think, obviously, my views of where inference — you know, how many gigawatts or terawatts are devoted to inference — it's a crazy curve for me personally.

显然,在"多少吉瓦、多少太瓦会投给推理"这个问题上,我个人的预期是一条疯狂的曲线。

Shaun Maguire00:21:25

What's your forecast — how many gigawatts?

你的预测是多少——多少吉瓦?

Dylan Patel00:21:27

Um, yeah, I think by, you know, 2030, just OpenAI and Anthropic will have over 100 gigawatts combined.

嗯,我认为到 2030 年,仅 OpenAI 和 Anthropic 两家合计就会超过 100 吉瓦。

Um, and then you'll add, you know, Meta and Google and, you know, so on and so forth. It's a humongous amount of compute that will be dedicated to inference.

再加上 Meta、Google 等等。将有一个庞大到夸张的算力体量专门用于推理。

Um, and by, like, 2040, it'll be terawatts, right? Um, the curve of, like, productivity that we're going to get — and so, you know, inference deployments are going to be huge.

到 2040 年左右,就是太瓦量级了。我们将获得的生产力曲线——推理部署的规模会非常巨大。

And so if you look at, like, 2040, I think, like, you know, probably more than half of the incremental compute will be going in space. But if you look at 2030, I think it's sub-1%.

所以看 2040 年,我认为超过一半的增量算力大概会上太空。但看 2030 年,我认为不到 1%。

Sonya Huang00:21:56

Do you think intelligence per watt has been increasing?

你觉得"每瓦智能"一直在提升吗?

Uh, and then it seems like there's still a giant gap between where we are on intelligence per watt versus, like, human biology. And so, like, do you think we are going to close that gap? And if so, where is that gain going to come from?

而且看起来,我们现在的每瓦智能和人类生物大脑之间还有巨大差距。你觉得这个差距能补上吗?如果能,增益从哪来?

Dylan Patel

Yeah, I think it often depends on what you're doing, too, right? Like, a TI-84 is way more intelligence per watt in terms of doing math than us, and it's, like, 30 years old, right? Obviously this is, like, a dumb, dumb, you know, sort of—

嗯,这往往取决于你在做什么,对吧?论做数学,一台 TI-84 计算器的每瓦智能远超人类,而它已经 30 年历史了。当然这是个很蠢的类比——

Sonya Huang

General intelligence.

我说的是通用智能。

Dylan Patel00:22:20

Yeah. But general intelligence-wise, um — so one of the things InferenceX does is we also measure the power and cost of all of this hardware.

好,那说通用智能——InferenceX 做的事情之一,就是同时测量所有这些硬件的功耗和成本。

And so we offer not just, you know, throughput versus interactivity — we offer cost versus interactivity, we offer power versus interactivity.

所以我们提供的不只是吞吐 vs 交互性,还有成本 vs 交互性、功耗 vs 交互性。

And so, as far as has intelligence per watt been increasing — um, I mentioned, you know, it's been a 60x cost decrease for the same benchmark level. Um, we've also seen the same on intelligence per watt. It's not been exactly 60x — it's been closer to, like, 40x. Uh, some of the efficiencies are non-power ways.

那么每瓦智能有没有在涨——我前面说过,同等基准水平下成本降了 60 倍。每瓦智能上我们看到了同样的趋势,不完全是 60 倍,更接近 40 倍。有一部分效率提升不走功耗这条路。

But there's been a humongous improvement in intelligence per watt on an annual basis — at least so far this year, last year, year before, year before. And I expect that to continue.

但每瓦智能每年都有巨大的提升——至少今年、去年、前年、大前年都是如此。我预计还会继续。

As far as where we are from the human brain — we're many orders of magnitude away. Thankfully, it doesn't really matter. We can devote a lot of power to computers. Much easier to power computers than human brains.

至于和人脑的差距——还差好几个数量级。所幸这其实无关紧要:我们可以给计算机堆非常多的电力。给计算机供电比给人脑"供电"容易多了。

Like, you know, we have sickness, disease, and, like, food preferences…

毕竟人类有生病、疾病,还有挑食这些问题……

Sonya Huang

Sleep.

还要睡觉。

Dylan Patel

Uh, yeah, exactly.

对,正是。

Chapter 09

Co-Design Is the Real 100x

三层协同:2x × 2x × 2x ≠ 8x
23:20 — 29:02 · 三层之争 · DeepSeek 形状 · 中西差异
Shaun Maguire00:23:18

Let me just ask one more question on the, like, on the general theme — in my opinion, in terms of, like, you know, intelligence per watt, or intelligence per dollar, like, any of these metrics — I think there's kind of three levels of input.

让我在这个大主题上再问一个问题——在我看来,不管是每瓦智能还是每美元智能,这类指标背后有三个层次的输入。

You can get hardware improvements, where the hardware is more efficient. You can get low-level systems optimizations — like kernel-level, you know, improvements, matrix multiplication libraries, things like that. Or you can get, like, high-level, like, model-level algorithmic improvements, you know, at the highest level.

一是硬件改进,硬件本身更高效;二是底层系统优化——kernel 级的改进、矩阵乘法库这类;三是最上层的、模型层面的算法改进。

It seems, like, to me — it seems like in the last three years, most of the gains have come from the hardware level, and, you know, some from the model level.

在我看来,过去三年大部分增益来自硬件层,一部分来自模型层。

Like, do you agree with that? Do you think that's what it'll look like in the future? Like, do you think there's a bunch of juice to squeeze in, say, like, the kernel level?

你同意吗?你觉得未来还会是这个格局吗?比如 kernel 层,你觉得还有很多油水可榨吗?

Dylan Patel

Yeah, Shaun, I completely disagree with you, by the way.

嗯,Shaun,顺便说一句,我完全不同意你。

Shaun Maguire00:24:19

Great, great — that's why I'm asking this question.

很好,很好——所以我才问这个问题。

Dylan Patel00:24:22

Um, okay, so I think, you know, one way is to look at it as these three different layers.

呃,好。一种视角确实是把它看成这三个层次。

Um, and in that sense, like, okay — from Hopper to Blackwell, which is all we've had over the last three years, roughly a 30x improvement on DeepSeek on the most optimized deployment — which is, you know, you can see on InferenceX, there's about a 30x improvement.

从这个角度看——从 Hopper 到 Blackwell,这就是过去三年硬件的全部,在最优化部署下跑 DeepSeek 大约有 30 倍提升——InferenceX 上能看到,大约 30 倍。

But you know, over the last three years, um, we've had way more improvement in intelligence per watt — a lot of that coming from the model layer, right?

但过去三年,每瓦智能的提升远不止这个数——其中很大一块来自模型层。

If you look back three years, it's GPT-4. Now it's, like, you know, maybe, like — one of the smaller Qwen models that's, like, you know, 27B parameters total and, like, 2 billion active, is, like, way better.

往回看三年,那时是 GPT-4。而现在,也许某个小号的 Qwen 模型——总共 27B 参数、激活 2B——就已经好得多了。

Um, and so you've got this huge improvement on the model layer, you've got this pretty sizable improvement on hardware — but it's that co-design layer, and I think that's what's important, right?

所以模型层有巨大提升,硬件层有可观提升——但关键是那个"协同设计层",我认为那才是重点。

If you look at the architecture of, you know, any of these models — but DeepSeek is the most famous one, at least, uh, that's public and people have seen…

看看这些模型中任何一个的架构——DeepSeek 是最出名的,至少它是公开的、大家都看过的……

Shaun Maguire00:25:12

Yeah, DeepSeek got huge efficiency gains from, like, co-optimization, or kernel-level optimizing memory.

对,DeepSeek 靠协同优化、kernel 级的内存优化拿到了巨大的效率增益。

Dylan Patel

Yes — I think it's, like, kernels, of course, but it's actually: you build the model architecture for the chip.

对——当然有 kernel 的部分,但更本质的是:你是照着芯片来构建模型架构的。

So if you look at the shapes of all the experts in DeepSeek, uh, V3 — they were all optimized for Hopper. And if you look at V4, they're optimized for Blackwell and Huawei's chip.

你去看 DeepSeek V3 里所有 expert 的形状——全是为 Hopper 优化的。再看 V4,它们是为 Blackwell 和华为的芯片优化的。

And what's interesting is, despite the fact that TPUs are objectively an amazing chip — you know, and they run all of DeepMind, and they do all the training, uh, for Anthropic as well, on the pre-training side at least — TPUs suck at running DeepSeek.

有意思的是,尽管 TPU 客观上是块惊人的芯片——DeepMind 全跑在上面,Anthropic 的训练(至少预训练侧)也全在上面——TPU 跑 DeepSeek 却烂得很。

But they are really, really great at running other kinds of models that don't run well on Nvidia.

但那些在 Nvidia 上跑不好的其他类型模型,TPU 又跑得非常非常好。

There is some level of such deep optimization that has been done — um, whether it be shapes, uh, network IO, uh, patterns, you know, how you do the collectives, how you do, um, things around, you know, the arithmetic intensity of the attention mechanism.

这里面存在一层极深的优化——不管是形状、网络 IO、通信模式、集合通信怎么做,还是注意力机制的算术强度怎么安排。

All these different things are co-optimized between the model and the hardware and the infra software in between. And it's hard to say you can disentangle the gains.

所有这些都是在模型、硬件,以及夹在中间的基础设施软件之间协同优化的。你很难把增益拆开来单独归因。

Shaun Maguire00:26:15

Do you think that, like — my understanding is that, like, China has done this a lot better than the West the last few years. Like, DeepSeek was one of the first models to really, like, do this.

你觉得——我的理解是,过去几年中国在这件事上做得比西方好得多。DeepSeek 是最早真正这么做的模型之一。

Dylan Patel00:26:27

I don't necessarily think so. I think it's more so that the West doesn't tell people what they do, right?

我不这么认为。更多是因为西方不告诉别人他们在做什么。

Like, OpenAI didn't tell people that, you know, GPT-4o was — how sparse it was, what the shape size was, all these things.

比如 OpenAI 从没告诉过大家 GPT-4o 有多稀疏、形状尺寸是什么,这些统统没说。

But GPT-4o is roughly the same size — slightly smaller than DeepSeek V3 — and 4o came out, you know, a little bit earlier, right? If I recall correctly.

但 GPT-4o 的规模差不多——比 DeepSeek V3 略小——而且 4o 发布得还更早一点,如果我没记错的话。

Shaun Maguire

So is your view that, like, all three of these things have been happening simultaneously at, like, roughly the same rate, and the biggest gains are when you just co-optimize?

所以你的观点是:这三层一直在以差不多的速度同时推进,而最大的增益来自协同优化?

Dylan Patel00:26:57

I would say there's been more gains on the model layer than on the sort of software infrastructure layer and the hardware layer. Um, but there's been innovations on every layer.

我会说模型层的增益比基础设施软件层和硬件层都多。但每一层都有创新。

And really, the biggest gain — and the beauty of the best labs — is when they co-optimize all three.

而真正最大的增益——也是最好的实验室的美妙之处——在于把三层一起协同优化。

You know, and that's what, like — you know, when Anthropic, even though they use many different kinds of hardware: they don't really inference too much on TPUs, they mostly train on TPUs, um, and they inference a lot on Trainium and GPUs — and GPU is more a jack of all trades.

Anthropic 就是这样——尽管他们用很多种硬件:TPU 上基本不做推理,主要用来训练;推理大量跑在 Trainium 和 GPU 上——GPU 更像个多面手。

But they've optimized their hardware, they've optimized their model, they've optimized everything so they can do that.

但他们优化了自己的硬件、优化了模型、优化了一切,所以才能这么玩。

Whereas OpenAI — their prior models were optimized for Hopper more; now they're more optimized for Blackwell. And you step forward through time with these labs.

OpenAI 那边——之前的模型更多为 Hopper 优化,现在更多为 Blackwell 优化。沿着时间轴看这些实验室都是如此。

And the same with Google, right? Gemini 2 was really optimized for the TPU v6e, Gemini 3 was — and then the next Gemini that's coming out is really optimized for TPU v7.

Google 也一样:Gemini 2 是深度为 TPU v6e 优化的,Gemini 3 也是——接下来要出的下一代 Gemini,则是深度为 TPU v7 优化的。

Um, and so, sort of like, a lot of these things are being co-optimized. And actually, when you pull that model and put it — run it on the old hardware — it's really not that great.

所以很多东西都在被协同优化。而实际上,你把那个模型拽出来放到老硬件上跑,效果真的不怎么样。

Um, and so I think a lot of this co-optimization is the most important thing. It's called software-hardware co-design.

所以我认为这种协同优化是最重要的事。它的名字叫软硬件协同设计(software-hardware co-design)。

And that's what's, like, really exciting about, like, you know, sort of what my day-to-day is like: great, you get to look at one layer — there's all these innovations happening here, there's all these innovations happening on every layer.

这也是我日常工作里真正让人兴奋的地方:你看任何一层——这里在发生一堆创新,每一层都在发生一堆创新。

The real breakthrough innovation is when you leapfrog a few layers — you co-optimize and co-design them. And now, all of a sudden, you've taken what could have been a 2x here, 2x here, 2x here — and instead of being multiplicative to 8x, it's actually 100x, because you've optimized across all three layers.

而真正的突破性创新,是当你跨越几层——把它们协同优化、协同设计。突然之间,原本这里 2x、那里 2x、再来个 2x——不是相乘得 8x,而是直接 100x,因为你是横跨三层一起优化的。

And so that's what's really exciting about sort of what you see at the labs. Which you see at, like, a company like Nvidia — who's not co-optimizing on the model layer per se, but a little bit — from the model layer all the way downstream to, you know, silicon.

这就是实验室里最让人兴奋的东西。你在 Nvidia 这样的公司也能看到——他们虽然不直接做模型层的协同优化,但也有一点——从模型层一路向下游打通到硅片。

Or you look at a company like TSMC — they're co-optimizing not just, you know, fabrication, but all the way from the components and the consumables and the tools, all the way upstream to what the designs — their chips, the customers — are telling them.

再看 TSMC——他们协同优化的不只是制造本身,而是从零部件、耗材、设备,一路向上游打通到客户的芯片设计给他们的反馈。

This is co-optimization across many layers of the abstraction stack.

这是横跨抽象栈许多层的协同优化。

Chapter 10

The Bottlenecks He's Watching

内存墙、1W/mm² 与柴油发电机
29:02 — 33:06 · DRAM 40 年没变 · 功率密度 · 能源土办法
Shaun Maguire00:28:59

There will always be bottlenecks somewhere in that optimization, though, that are, like, lagging behind and then need to get pulled forward, you know — and band-aids to attack.

不过这套优化里永远会有瓶颈——某个地方落在后面,需要被拽上来,或者先打个补丁顶着。

If you had to predict, like — at any level of the stack, it can be literally anywhere — what are some of the bottlenecks you're, kind of, tracking most acutely the next year?

如果让你预测——栈的任何一层都行,真的哪儿都可以——未来一年你盯得最紧的瓶颈是哪些?

And not necessarily in the supply chain, not in, like, scale — but in terms of the actual — and it can be in the supply chain too, but just, like, you know: is it memory improvements? Is it just, like, scaling?

不一定是供应链、不是规模问题——当然供应链也可以——比如:是内存的改进吗?还是单纯的 scaling?

Dylan Patel00:29:35

So memory — memory is an easy one that everyone's talked about. But I'm not going to talk about it from a supply chain angle; I'm talking about it from a technology angle, right?

内存——内存是人人都在谈的显而易见的答案。但我不打算从供应链角度谈,我从技术角度谈。

Memory, um, capacity and bandwidth have been improving very slowly.

内存的容量和带宽一直提升得非常慢。

The NAND cell was invented, like, 25 years ago. The DRAM cell was invented, like, 40 years ago. And there's been no major breakthrough in cell — like, you know, what a NAND cell is. Obviously NAND is, like, a very simple gate, or the DRAM cell.

NAND 单元是大约 25 年前发明的,DRAM 单元是大约 40 年前发明的。在"单元"本身上一直没有重大突破——NAND 本质上就是个很简单的门电路,DRAM 单元也是。

There is stuff that could come down the pipeline that could be hugely innovative.

管线里倒是有一些可能极具创新性的东西在路上。

But even over the last, you know, five years, all we've really done is make the HBM, you know, more stacks, faster.

但即使过去五年,我们真正做到的也只是把 HBM 堆得更高、跑得更快。

But actually, there's, like, new innovations coming in the next few years, where instead of, you know, stacking the HBM separately from the chip, you stack the memory directly on the chip — and that makes your bandwidth explode.

而实际上,未来几年会有新的创新:不再把 HBM 和芯片分开堆叠,而是把内存直接堆叠在芯片上——这会让带宽爆炸式增长。

Um, and so there's interesting companies in that space, and interesting things that companies are trying to do there. I think, like, memory bandwidth is one of the biggest.

这个方向上有一些有意思的公司,大家在尝试一些有意思的事。我认为内存带宽是最大的瓶颈之一。

Another one is, um — for the history of, like, silicon, basically, for the last two decades at least, you know, how many watts a chip is can be easily predicted just by looking at it. For a data center or desktop chip, it peaks out at one watt per millimeter squared.

另一个是——纵观硅片的历史,至少过去二十年,一块芯片多少瓦,看一眼面积就能算出来:数据中心或桌面芯片,功率密度顶格在每平方毫米 1 瓦。

And so if a chip is 100 millimeters squared, generally the power consumption is around 100 or a little bit less.

如果芯片是 100 平方毫米,功耗一般就在 100 瓦上下,或略少。

Um, and if you look at the newest Nvidia silicon, the newest TPU silicon — it's still in that range of one watt per millimeter squared.

看最新的 Nvidia 硅片、最新的 TPU 硅片——仍然在每平方毫米 1 瓦这个区间。

So, you know, chips are now getting to, you know, 1,400 watts. Next generation is 2,000 watts for Nvidia, um, with Rubin and such. Uh, and you move forward to Rubin Ultra — it's going to be, like, 4,000 watts or something like that. But really, they're increasing the amount of silicon.

所以芯片现在到了 1400 瓦。Nvidia 下一代 Rubin 是 2000 瓦,再往后 Rubin Ultra 大概是 4000 瓦这个量级。但本质上,他们只是在增加硅片面积。

What's exciting is we're now finally doing things — and it's in development right now — where you actually can pump the amount of power into the silicon, uh, to be way more than one watt per millimeter squared.

让人兴奋的是,我们终于开始做一件事——现在正在研发中——真的把泵进硅片的功率提高到远超每平方毫米 1 瓦。

And now, that all of a sudden means you need less silicon. Obviously, it's running at higher power, it's less efficient in some cases — but you reduce the amount of silicon.

这一下就意味着你需要的硅片变少了。当然它跑在更高功率上,某些情况下效率更低——但硅片用量降下来了。

Shaun Maguire00:31:29

Like, over thermal issues?

比如散热问题?

Dylan Patel00:31:31

Thermal issues, um, there's, uh, interference — like, electrical interference issues. There's all sorts of different issues, uh, that crop up, and that's why it's a hard engineering problem. That's why we've been stuck at about one.

散热问题,还有干扰——电气干扰问题。各种各样的问题会冒出来,所以这是个很难的工程问题,也是我们一直卡在 1 瓦/平方毫米左右的原因。

But what's exciting is the world is trying to change these things.

但让人兴奋的是,整个世界都在试图改变这些。

I think, interesting, like, in a different part of the supply chain — it's sort of like, you know, people will talk about, like, energy is hard, and you know, we have energy bottlenecks. And it's like, yeah — but there's actually, like, very simple solutions one could think of, right?

供应链的另一个环节也很有意思——大家都说能源难、我们有能源瓶颈。确实,但其实有一些很朴素的解法是想得到的。

Um, take the millions of diesel engines for trucks that the US has the capacity to make. Um, you can very trivially convert them to be using gas, uh, in the assembly line.

比如,美国有年产数百万台卡车柴油发动机的产能。你可以在产线上极其简单地把它们改成烧燃气。

And then stick them up to an electrical motor, like, back-driving it — so the electrical motor generates electricity, rather than the electrical motor causing the rotation of the wheel, for example, but doing it the opposite direction.

然后接上一台电机,反向驱动——让电机发电,而不是像平常那样由电机带动车轮转,方向反过来用。

And now you've generated electricity by pumping gas into something that the US can make millions of.

于是,往一个美国能造几百万台的东西里灌燃气,你就发出了电。

Um, and then — okay, well, that sounds like a pain in the ass to, uh, service, right? Because now you have to have hundreds of these on a data center site.

然后——好吧,这听起来维护起来很要命,对吧?因为现在一个数据中心场地上要摆几百台这玩意。

Well, actually, you can just pull people out of car mechanic shops and have them run around and repair truck engines.

但实际上,你直接从汽修店里把人拉出来,让他们跑来跑去修卡车发动机就行了。

Actually, it's actually pretty trivial to — no, I don't want to say it's trivial. I couldn't do it.

这其实相当简单——不,我不该说简单。反正我自己是干不来的。

Shaun Maguire00:32:41

I think you're making a really good point, which is that, like, because the West wasn't really thinking about semis — even hardware more broadly — the last 20, 30 years, we didn't have much innovation. We didn't have the best minds, like, thinking about how do you improve these.

我觉得你说到了一个要点:因为过去二三十年西方根本没在想半导体——甚至更广义的硬件——所以我们没什么创新,最聪明的头脑没有在想怎么改进这些东西。

Dylan Patel00:32:57

Why would you want to go work in hardware when you can, uh, make ads to serve ads?

能去做广告、投广告赚钱,谁还愿意去搞硬件呢?

Shaun Maguire

Yeah, exactly.

对,正是。

Chapter 11

Nvidia vs TPU & the CUDA Moat

稀疏与稠密的分岔,护城河的真相
33:06 — 38:47 · NVLink vs ICI · 开源生态 · 大厂都 fork 了 PyTorch
Sonya Huang00:33:03

Um, okay, I'm dying to ask: Nvidia versus TPU — what are your thoughts?

呃,好,我憋不住要问了:Nvidia 对 TPU——你怎么看?

Dylan Patel00:33:10

Um, I think, like, everyone wants to pick one or the other for this. But it's really, like, a function of, like — look, you know, you look two years from now: Google's going to make 10-plus million TPUs, and through their supply chain, and Nvidia is going to make, you know, many more million — tens of millions — of GPUs.

呃,我觉得所有人都想二选一。但这其实是个体量问题——看两年后:Google 会通过自己的供应链造出 1000 万颗以上的 TPU,Nvidia 会造出多得多的、数千万颗的 GPU。

And both are going to be 100-plus billion dollar — you know, well, Google's going to be 100-plus billion dollars, you know, of TPU created a year, and Nvidia will be, you know, 500-plus, or, you know, whatever. I'm not making a specific estimate.

而且两边都会是千亿美金级——Google 每年造出的 TPU 会超过 1000 亿美金,Nvidia 会是 5000 亿以上,或者随便多少。我不是在做具体预测。

Sonya Huang

This is not a revenue forecast. This is just a thought experiment.

这不是营收预测,只是个思想实验。

Dylan Patel

Yeah. Or research.

对。或者说"研究"。

Shaun Maguire00:33:37

You've been media trained.

你是受过媒体训练的。

Dylan Patel00:33:39

Absolutely. You know, getting ready for the SpaceX IPO.

当然。为 SpaceX 的 IPO 做准备呢。

Um, are you guys big in SpaceX? Okay, so that makes sense.

呃,你们重仓 SpaceX 吗?好,那就说得通了。

Shaun Maguire00:33:45

We're very lucky to be very large investors.

我们很幸运,是非常大的投资方。

Dylan Patel

Awesome. Awesome.

棒。棒。

Um, so I would say, um, in the case of, sort of, like, Google TPUs versus, uh, Nvidia GPUs — they both have, like, points that are really, like, in their favor, right?

我会这么说:Google TPU 对 Nvidia GPU——双方都有真正对自己有利的论点。

You know, Nvidia will be like, "Oh, well, we have switches, and we're general purpose." And TPUs will be like, "Well, we're more optimized, actually more energy efficient, and our network is actually more, um, optimized for certain types of network architectures."

Nvidia 会说:"我们有交换机,而且我们是通用的。"TPU 会说:"我们更深度优化、实际上更省电,而且我们的网络对某些网络架构更友好。"

And so you have, like, these counterpoints that both would really, uh, get into. And you know, I could, with a straight face, argue with you that GPUs are way better than TPUs, or TPUs are way better than GPUs.

双方都能就这些对立论点吵起来。而我可以面不改色地跟你论证 GPU 远胜 TPU,也可以论证 TPU 远胜 GPU。

But it comes down to hardware-software co-design.

但归根结底还是软硬件协同设计。

So actually, the way OpenAI's models are headed, it would be a terrible decision for them to use TPUs, potentially. And the way that Anthropic and Google's, uh, models are headed, it's actually a terrible decision, potentially, for them to train with GPUs.

实际上,照 OpenAI 模型的走向,用 TPU 对他们来说可能是个糟糕透顶的决定;而照 Anthropic 和 Google 模型的走向,用 GPU 训练对他们来说可能同样糟糕透顶。

Shaun Maguire00:34:33

I mean, it'd be fun to — what's the fundamental difference there?

有意思——那本质差别在哪?

Dylan Patel00:34:38

There's various things, right? Like, the size of the matrix multiply unit is different, as a very simple thing. And therefore the shape of the matrix multiply you do, the attention mechanism you use, uh, the way that attention mechanism is structured, the way the experts are structured.

有很多方面。最简单的一条:矩阵乘法单元的尺寸不同。于是你做的矩阵乘法形状、你用的注意力机制、注意力机制的结构方式、expert 的结构方式,都跟着不同。

Shaun Maguire00:34:51

So you think — so OpenAI and Anthropic are converging to very different model architectures?

所以你认为——OpenAI 和 Anthropic 正在收敛到非常不同的模型架构?

Dylan Patel00:34:56

I think they have quite different model architectures. In fact, um, you know, OpenAI's are much more sparse, um, and that has benefits. And then Anthropic's are — you know, they're still sparse, but more dense in general, and that has different benefits.

我认为他们的模型架构已经相当不同。事实上,OpenAI 的模型稀疏得多,这有它的好处;Anthropic 的——依然稀疏,但总体更稠密,这又有另一套好处。

And there's many other things, right? The network topology, right? Nvidia — all of their chips are connected to switches, NVLink switches.

还有很多别的。比如网络拓扑:Nvidia 的所有芯片都接在交换机上,NVLink 交换机。

For Google, they have no switch. Um, but what they've done is — you know, Nvidia, the NVLink can only connect 72 GPUs; for Google, their ICI can connect 8,000 chips at super high bandwidth — but you have to pass through other chips to get there, because there's no switch.

Google 则没有交换机。但他们做到的是——Nvidia 的 NVLink 只能连 72 块 GPU;Google 的 ICI 能以超高带宽连 8000 颗芯片——只是因为没有交换机,数据得途经其他芯片中转。

And so there's, like — there's trade-offs there. There's positives and negatives, and that influences the model architecture.

所以这里有取舍,有利有弊,而这些会影响模型架构。

It's not necessarily that you should, uh, you know, claim one is better than the other. Because at the end of the day, how do you say that this is better than that, when you can't measure them in isolation? Because it also extends up to the model layer, right?

你不该断言谁比谁强。因为说到底,当你没法把它们隔离开单独测量时,怎么说这个比那个好?毕竟影响一路延伸到模型层。

Shaun Maguire00:35:49

But I remember, for a long time, thinking, you know — one, the programmability of Nvidia, and just CUDA as such a big moat.

但我记得很长一段时间里,大家都认为 Nvidia 的可编程性、CUDA 本身,是一条巨大的护城河。

It seems to me that narrative has kind of changed — at least in my mind — for the last three, six months. Like, model companies no longer care about — if we have to write custom kernels for, you know, this other chip, so be it. We'll work with four or five chips if we have to.

但在我看来,这个叙事在过去三到六个月变了。模型公司不再在乎——要给另一块芯片写定制 kernel?那就写呗。必要的话我们可以同时伺候四五种芯片。

Um, Claude and Codex are actually quite good at doing a lot of that optimization work.

而且 Claude 和 Codex 干这类优化活其实已经相当在行。

And then, you know, it's not like there's 10,000 model companies that each need programmability — there's on the order of tens, maybe, model companies.

再说,也不是有一万家模型公司各自需要可编程性——模型公司的数量级也就是几十家。

And so it seems to me that, like, the fundamental premise of, like, tens of thousands of big customers that need CUDA compatibility — it seems that kind of thesis is changing in the last…

所以在我看来,"成千上万个大客户需要 CUDA 兼容性"这个底层前提,最近正在瓦解……

Dylan Patel

Yeah. I mean, certainly the CUDA moat — the software moat — is at least partially, uh, disentangled. Because, you know, models are just great at coding, and all software gets commoditized in that case.

对。CUDA 护城河——软件护城河——确实至少被部分拆解了。因为模型写代码太强了,那种情况下所有软件都会被商品化。

I do think there is some level of, like, open source, and — you know, what people call the CUDA moat is not actually anything to do with CUDA.

但我确实认为还有一层开源的东西——大家嘴里的"CUDA 护城河",其实跟 CUDA 一点关系都没有。

It's, like, the fact that DeepSeek, Kimi, and Zhipu, and Alibaba, and Tencent — all these companies; Xiaomi had an awesome model recently — their models are co-designed for GPUs. And therefore, if I want to run them on TPUs, actually, in some cases, they don't run really well on TPUs.

真相是:DeepSeek、Kimi、智谱、阿里、腾讯——所有这些公司,小米最近也出了个很棒的模型——他们的模型都是为 GPU 协同设计的。所以我要是想在 TPU 上跑,某些情况下真的跑不好。

Now Google just has to create their own open-source model ecosystem, or open-source models themselves — so they have the Gemma models.

于是 Google 只能自建开源模型生态,或者自己开源模型——所以他们有 Gemma 系列。

And so you end up with, like — well, that's not really CUDA as a moat. It's that the downstream product is more optimized for Nvidia.

最后你会发现——护城河根本不是 CUDA,而是下游产品都更深地为 Nvidia 优化了。

And in these cases, these companies are just open-sourcing them — or, like, Nemotron is just open-sourcing it.

而这些公司就这么把模型开源出来——比如 Nemotron 就直接开源。

And then the users of it — for example, the inference API providers, the RL companies that are trying to take open models and customize them for companies' business use cases — all these different companies are downstream of the fact that, like: okay, well, I guess I need to use Nvidia, because the ecosystem uses Nvidia.

然后它的使用者们——比如推理 API 提供商、想拿开源模型为企业业务场景做定制的 RL 公司——所有这些公司都被这个事实框住了:好吧,那我只能用 Nvidia,因为生态用的就是 Nvidia。

Even though I don't particularly care about writing CUDA kernels — because the models are great at that — but it's, like, the shape of, like: well, this expert, the d_model is this, and, you know, the hidden dimension blah blah blah is this, right?

哪怕我根本不在乎写 CUDA kernel——模型已经很会写了——但架不住形状问题:这个 expert 的 d_model 是这样,hidden dimension 是那样,等等等等。

And so, therefore, it's better to run on Nvidia GPUs than it is on TPUs — and vice versa, right?

所以它在 Nvidia GPU 上就是比在 TPU 上跑得好——反过来也成立。

If Google were to actually open-source really good models, you know, this would be the same thing, right? People would take their models and they'd be like, "Oh wow, these don't run that well on Nvidia GPUs. Um, I should actually just rent TPUs, or buy TPUs, and do it on there."

如果 Google 真把很强的模型开源出来,同样的事就会反过来发生:大家拿到模型一看——"哦豁,这在 Nvidia GPU 上跑不太行啊,那我还是去租 TPU、买 TPU,在那上面跑吧。"

For small teams, you're going to want to use all the open-source software — like vLLM and SGLang, um, PyTorch, all that stuff.

小团队肯定会用全套开源软件——vLLM、SGLang、PyTorch 这些。

But the big labs, they don't necessarily need to use all that, right? OpenAI forked PyTorch long ago. And, you know, Anthropic and all these other people don't necessarily rely heavily on the open-source implementation of, you know, these things — they forked things or built it on their own already.

但大实验室不一定需要这些。OpenAI 很早就 fork 了 PyTorch。Anthropic 和其他几家也并不重度依赖这些东西的开源实现——他们要么 fork 了,要么早就自研了。

And so they don't need to rely on the open source. And therefore, now it's more like: you know, I'll choose the best hardware, and I'll co-design my model and infrastructure software through and through for that hardware, uh, that is the best and most cost-efficient.

所以他们不需要依赖开源。于是逻辑变成:我选最好的硬件,然后为这块最好、最具成本效率的硬件,把我的模型和基础设施软件从头到尾协同设计一遍。

Shaun Maguire00:38:43

And, you know, I'll have AI help me write all that software.

而且,我还会让 AI 帮我把那些软件全写了。

Chapter 12

Cerebras & the Price of Speed

快即市场:fast mode 经济学
38:47 — 42:17 · dark GDP · 按天盯 token 账单 · 皮卡类比
Sonya Huang00:38:46

What do you think of Cerebras?

你怎么看 Cerebras?

Dylan Patel

I think Cerebras is a really innovative company. Um, I think in some spots of the market, they're really, really good.

我认为 Cerebras 是家非常有创新性的公司。在市场的某些位置上,他们真的非常强。

Um, very fast inference — I think that's a big market. Uh, we use fast mode almost exclusively at SemiAnalysis.

超快推理——我认为那是个大市场。在 SemiAnalysis,我们几乎清一色地用 fast mode。

Shaun Maguire00:39:02

By the way, I love how disciplined you've been about accounting for — I don't know if that was one exhibit you did, or if you do it consistently — but accounting for the dollars spent and the ROI on each task.

顺便说,我特别欣赏你们在核算上的纪律性——不知道那只是你们做过的一张图表,还是一直在做——把每个任务花的钱和 ROI 都记下来。

Awesome analysis.

很出色的分析。

Dylan Patel

Yeah, we do it pretty diligently, and so — thank you. That was the "dark GDP" article that we wrote.

是的,我们做得挺勤的——谢谢。那是我们写的"dark GDP"那篇文章。

Um, and also, like, I track everyone's token spend by day. And if someone's, like, spiked up, I'm like, "What did you do?" And it's like — okay, thank you for telling me that. That seems worth it. Cool. On with my day.

另外,我按天盯每个人的 token 开销。谁的开销突然飙了,我就问:"你干了啥?"然后——好,谢谢你告诉我,看起来值。行,我继续忙了。

I think fast mode is obviously worth a lot for high-end tasks, right? I could just see so many different use cases where, you know, super fast tokens are worth it.

fast mode 对高价值任务显然非常值。我能想到太多"超快 token 值回票价"的场景了。

I can also see the flip side, where there's a lot of use cases where super fast tokens aren't needed, and therefore, uh, the market won't pay for them, and they'll use GPUs and TPUs instead.

但我也看得到另一面:很多场景根本不需要超快 token,市场不会为它付钱,大家会转头用 GPU 和 TPU。

I think the big risk for Cerebras is — I mostly think the best models are the ones that you want to use fast mode on, and small models you necessarily might not use fast mode on.

我认为 Cerebras 最大的风险在于:我基本认为,你想上 fast mode 的都是最强的模型;小模型你未必会用 fast mode。

I could see that being wrong with, you know, financial markets, maybe, or something like that — like a Jane Street, high-frequency trading or something like that, um, or medium-frequency trading.

这个判断在金融市场里可能不成立——比如 Jane Street、高频交易之类,或者中频交易。

Um, but ultimately, you know, running really large models at really long context is very difficult on SRAM-based chips like Cerebras, like Groq.

但归根结底,在 Cerebras、Groq 这类基于 SRAM 的芯片上,以超长上下文跑超大模型是非常困难的。

And so now, all of a sudden, it's like — you know, what happens then, if, like, the models get too big, right?

于是问题突然变成:如果模型变得太大,会怎么样?

If OpenAI's model is not, you know, on the order of, uh, hundreds of billions of parameters, or, you know, low trillions of parameters — but it's actually 10-plus trillion parameters — now, all of a sudden, I don't think that will fit on Cerebras, right?

如果 OpenAI 的模型不是几千亿参数、一两万亿参数的量级,而是 10 万亿参数以上——那我不觉得 Cerebras 还装得下,对吧?

And then if that — with a long context length, right? If you have a million context length, now that makes it really difficult to justify.

再叠加长上下文——如果是一百万的上下文长度,那就真的很难算得过账了。

And as — so far, we've seen the bulk of revenue and usage at the labs be on their best model. Even when the model price has gone up, we've seen that.

而且到目前为止,各实验室营收和用量的大头都在他们最强的模型上。即使模型涨价了,依然如此。

Um, there's some data that shows that even though Fable just released today, they've had incredible amounts of people switch to Fable and Mythos — sort of that next-tier model — even though it's way more expensive.

有数据显示,尽管 Fable 今天才发布,已经有惊人数量的用户切换到 Fable 和 Mythos——那个更高一档的模型——哪怕它贵得多。

Sonya Huang00:40:53

And that's volume by dollars, totally. But was that volume by tokens?

那是按美元算的量,没问题。但按 token 算也是吗?

Dylan Patel00:40:57

Well, I guess — who cares about volume by tokens? It's about the dollars.

呃,这么说吧——谁在乎按 token 算的量?重要的是美元。

Sonya Huang

Fair enough.

有道理。

Dylan Patel00:41:00

Right? If I don't care that there's, you know — I don't know — 200,000 Mini Coopers or Toyota Camrys sold, if, uh, you know, F-150s are 5x the ASP and they sell only half as much.

对吧?如果 F-150 的均价是 5 倍、销量只有一半,那我才不在乎卖了 20 万辆 Mini Cooper 还是丰田凯美瑞。

And therefore, the most lucrative market is pickup trucks in America. Right? Mostly being facetious, but, like…

照这个算法,美国最赚钱的市场就是皮卡。对吧?我大半是在开玩笑,但……

Sonya Huang00:41:19

I do think this is one of the things that you've done so well, and differentiates you from almost everyone else — is that you care so much about the economics, in addition to the technology. And I think very few people bridge those two things well.

我真心觉得这是你们做得最出色、也是把你们和几乎所有人区分开的一点——你们在技术之外,还如此在乎经济账。能把这两件事桥接好的人非常少。

Dylan Patel00:41:26

I think it's really fun inside of SemiAnalysis, because we have 90 people, and, like, a big chunk of them are technologists, engineers across the whole supply chain. Um, and then a big chunk is people who are formerly at hedge funds.

SemiAnalysis 内部真的很好玩:90 个人,一大块是覆盖整条供应链的技术专家、工程师;另一大块是从对冲基金出来的人。

And you see these arguments — like, people are like, "Oh, well, that doesn't matter." And it's like — then someone's like, "Well, but cost." And then the engineers are like, "No, no, but this technology is the coolest." You see this organically, like, fight it out.

于是你会看到这种争论——有人说"哦,那不重要",有人接"可是成本呢",工程师们喊"不不,这技术才是最酷的"。你能看到这一切自发地吵出结果。

Um, and we're pretty informal. And, you know, given the fact that I was a forum moderator — you can imagine the enjoying it.

我们也很不拘小节。而且想想我以前是论坛版主——你可以想象那种乐在其中。

Shaun Maguire00:42:00

You don't wrestle with a pig, because a pig enjoys it.

别跟猪摔跤,因为猪乐在其中。

Dylan Patel

Exactly.

正是。

Chapter 13

Trigger Topics

"AI 没有 ROI"最令人恼火
42:17 — 44:23 · 否认模型进步 · 事实全对结论全错
Sonya Huang00:42:07

Just on this topic, before going to the next question — are there, like, trigger topics in semis for you?

趁着这个话题,进入下一个问题之前——半导体圈里有没有什么话题是你的"触发词"?

You know, like, if someone's like — which is, like, such a meme — you think "this person must be a…" Like, if it's, like, "Oh, you like — memory is the bottleneck."

就是那种一听到就想"这人肯定是个……"的梗——比如有人张口就是"内存才是瓶颈"。

Dylan Patel00:42:25

I mean, it's true, but, like — um, I think, moreover, the one that really gets me is people are like, "AI has no ROI."

这话倒是没错,但——真正让我上头的,是那些说"AI 没有 ROI"的人。

It infuriates me, right? Like, there's, like, "What's the ROI?" Or, like, denying model progress, right?

这让我火冒三丈。张口就是"ROI 在哪呢?"——或者否认模型在进步。

There's these people that are like: models aren't getting better, they're not reasoning, they can't think, they're going to dead-end and plateau.

有一群人整天说:模型没有在变好,它们不会推理、不会思考,马上就要撞死胡同、进入平台期了。

And it's like — bro, the line has been up and to the right in terms of capabilities this entire time.

拜托——能力那条线从头到尾都是一路向右上方走的。

And they're like, "Look, this benchmark didn't improve." That's cuz it's at 90%! Look at the new benchmarks — you saturated it; now they're skyrocketing, right?

他们说:"你看,这个 benchmark 没涨。"那是因为它已经 90% 了!去看新的 benchmark——旧的饱和了,新的正在飙升,对吧?

It's like — I think that's more so the issue and challenge.

我觉得这才是真正的问题和挑战所在。

Like, I think semis are really complex, and I don't fault people for, um, lacking, like, understanding of it.

半导体真的很复杂,大家理解不了,我不怪他们。

Like, I learn stuff every day about the semiconductor supply chain from people. And I've been studying it for, you know, arguably 18 years — since I started moderating the forums when I was 12, right? Like, you know, arguably been studying it for that long. But even then — and it's, like, live-breathed, and that's all I care about — but there's so many layers of the abstraction stack.

我自己每天都还在从别人那里学到半导体供应链的新东西。而我研究它,严格说有 18 年了——从我 12 岁当论坛版主算起。就算这样——我朝夕浸在里面,这是我唯一在乎的事——抽象栈的层数还是多到学不完。

Like, I learned about a new chemical that does, like, a hundred million dollars of sales, like, yesterday. And I'm like, whoa — didn't know this one existed, and what process it did.

比如就在昨天,我才知道有种化学品一年卖大约一亿美金。我心想:哇,居然还有这个,它是用在哪道工艺上的。

And it's like — you know, you learn about things all the time. It's like, okay, a hundred million dollars of sales in a, you know, couple-hundred-billion-dollar industry is whatever, but…

你就是会不停地学到新东西。好吧,在一个几千亿美金的行业里,一亿美金的销售额不算什么,但是……

Shaun Maguire00:43:44

But it's essential.

但它是不可或缺的。

Dylan Patel00:43:45

It's essential — and it's like, actually, every chip requires it. It's like, wow, I guess there are a thousand process steps.

不可或缺——而且实际上每块芯片都需要它。这就是那种"哇,原来有一千道工艺步骤"的感觉。

And you know, it's like, "Oh yeah? You like semiconductors? Name every process step." It's like — no, come on.

就像有人说:"哦是吗?你喜欢半导体?那把每道工艺步骤都报出来。"——不是,拜托。

What I think is the most funny is when people have all the facts in front of them, and then they get the conclusion completely wrong.

我觉得最好笑的,是那种所有事实都摆在眼前,结论却错得离谱的人。

Sonya Huang00:44:01

That happens in our job all the time, too.

这在我们这行也天天发生。

Shaun Maguire00:44:04

Yeah. I mean — I think my attitude is not to be mad that you do that; it's to do it as fast as possible.

是啊。我的态度不是为"结论错了"生气,而是尽可能快地把这个过程跑完。

Chapter 14

Ten-Year Bets

太空、共封装光学与 Naveen 的全栈豪赌
44:23 — 46:30 · CPO 时间表 · 模拟计算 · 在网上钓到 Naveen
Sonya Huang00:44:10

I think the industry — because it's so — it's just, like, AI is the most important thing in the world right now, and there are so many near-term bottlenecks. We talk a lot about the near-term.

这个行业——因为 AI 是当下世界上最重要的事,而且有太多近期的瓶颈,我们聊的也大多是近期。

Are there longer-term things that you're really excited about — like, say, on a 10-year time frame?

有没有更长期的、真正让你兴奋的东西——比如放到十年的时间尺度上?

We talked about orbital data centers. But, like — like silicon photonics: do you think they're underrated or overrated on a 10-year time frame? Are there other things on a 10-year time frame?

轨道数据中心我们聊过了。但比如硅光子:十年尺度上你觉得它被低估还是被高估?还有别的十年尺度的东西吗?

Dylan Patel00:44:32

Yeah, I mean, I think — on space: I think space is, like, super crazy awesome in the 10-year time frame — for space data centers, and all these sort of mining asteroids and all these things — which is, you know, super excited about the vision of SpaceX, right? Um, again, not investment advice, before you hop in.

嗯,先说太空:十年尺度上太空真的酷到疯狂——太空数据中心、小行星采矿这些——我对 SpaceX 的愿景超级兴奋。再说一遍,不构成投资建议,别急着进场。

Um, I think on the semiconductor side — tremendous market movements and tremendous, like, things can happen just when, like, things happen one year later or sooner.

半导体这边——某件事早一年或晚一年落地,就能引发巨大的市场波动和巨大的连锁反应。

And so that's all, like, technology that, like — you know, in terms of, like, co-packaged optics: like, well, everyone knows it's going to happen by the end of the decade. The debate is, like, '27, '28, '29, 2030 — but at some point along there, it's going to happen.

这类技术都是如此——比如共封装光学(CPO):所有人都知道十年内它一定会发生,争的只是 27、28、29 还是 2030 年——但总归会在那个区间里发生。

I think the more interesting thing is, like, there's companies like — um, did you guys invest in Naveen Rao's company?

我觉得更有意思的是像——对了,你们投了 Naveen Rao 的公司吗?

Shaun Maguire

We did.

投了。

Dylan Patel00:45:12

Okay. Yeah. So I think, like, he's trying to innovate on, like, the silicon layer, on the software abstraction layer, and the model layer simultaneously. And he fully understands that it's not, like, a "we're going to do this in a few years" thing.

好。我认为他在同时对硅片层、软件抽象层、模型层做创新。而且他完全清楚,这不是一个"几年内搞定"的事。

Shaun Maguire00:45:23

It's not a two-year time frame.

这不是一个两年尺度的事。

Dylan Patel00:45:24

Yeah, it's not a few-year time frame. It's a long-term bet.

对,不是几年的事,是一个长期赌注。

Um, and, like, stuff like that is, like — okay, we're going to bring, like, potentially, like, analog compute with energy-based models, and, like, all this crazy, all at once. It's like — that's exciting. Probably won't work, but, you know, that's exciting, and I really look forward to it.

这类东西就是——好,我们可能要把模拟计算和基于能量的模型这些疯狂的东西一次性全押上。这才叫激动人心。大概率成不了,但激动人心,我非常期待。

Shaun Maguire00:45:41

Definitely won't work quickly.

至少肯定不会很快成。

Dylan Patel00:45:42

Yeah — "definitely won't work quickly" is what I should say. I believe in Naveen.

对——我该说的是"肯定不会很快成"。我相信 Naveen。

And, like, you know, I met him very — you know, I think he's one of the first people I met in the industry, um, funnily enough — like, in 2020 or 2021. Um, actually 2020. Yeah.

说来有趣,他是我在这个行业里最早认识的人之一——2020 还是 2021 年,应该是 2020 年。

Shaun Maguire00:45:55

It says something about him. I think — he's someone, in my experience, he's always trying to help the younger generation. He's trying to identify talent.

这很能说明他的为人。以我的经验,他一直在帮助年轻一代,一直在发掘人才。

He was also so ahead of his time with Mosaic. I remember getting pitched.

而且他做 Mosaic 也太超前了。我还记得被 pitch 的场景。

Dylan Patel00:46:00

I baited him on the internet! I baited him on the internet.

我是在网上"钓"到他的!在网上钓到的。

No, it was 2019 — I was still anonymous then, actually. I baited him on the internet, and he started replying, and then I just took it to DMs, and then took it to a call.

不对,是 2019 年——那时我其实还是匿名的。我在网上钓他,他开始回复,然后我把话题转进私信,再约了通话。

And, like, that was the first person who's, like, really important that I talked to in the entire semiconductor industry. Funny.

他是我在整个半导体行业里聊过的第一个真正重要的人物。很有意思。

Um, but yeah — sorry to interrupt.

呃,抱歉打断了你。

Shaun Maguire

That's funny.

这挺逗的。

Chapter 15

The End State of Silicon

人人 ASIC,与局部最优陷阱
46:30 — 50:51 · 三条 TPU 设计线 · $11/时租 xAI · 通用算力的意义
Sonya Huang00:46:27

What do you think is the end state of the ecosystem? Like, do you think every lab, every hyperscaler just has its own chips?

你觉得这个生态的终局是什么?每个实验室、每个超大规模云都有自己的芯片?

Like, Trainium seems like it's now working, right? So do you think we end up with: every lab, every hyperscaler has its own chips, at least for inference — and then maybe for training, you go to Nvidia or whoever? Or what do you think is the end state?

Trainium 现在看起来跑通了,对吧?终局会不会是:每家实验室、每家超大规模云至少在推理上都用自研芯片,训练也许再去找 Nvidia 或别家?你觉得终局什么样?

Dylan Patel00:46:44

I think everyone will try, and some will stop trying. I think, ultimately, um, you know, supply chains matter. What technology you can bring in matters. And more and more, as the industry gets bigger, supply chain diversification happens.

我认为每家都会去试,而有些会中途放弃。最终,供应链很重要,你能引入什么技术很重要。而且随着行业越来越大,供应链的多元化会不断发生。

Um, you know, right now, everyone's chip more or less looks the same: it's a big logic compute die in the center, and there's some HBM on the right and left, and on the top and bottom — top side is networking, and then the bottom side is PCIe and other IO.

现在所有人的芯片长得都差不多:中间一块大的逻辑计算 die,左右是 HBM,上下两侧——上面是网络,下面是 PCIe 和其他 IO。

Um, and that is the exact same structure for Trainium, TPU, Nvidia chips. Um, and most of the startups — not Groq and Cerebras, they are doing weird, but that's cool, you know.

Trainium、TPU、Nvidia 的芯片,结构一模一样。大多数创业公司也是——Groq 和 Cerebras 除外,他们在搞怪路线,不过那很酷。

Um, I think, like, as you step forward, we're going to get more bifurcation of hardware architecture and model architecture. And therefore, people are going to co-optimize them.

往前走,硬件架构和模型架构会出现更多分岔。于是大家会去做协同优化。

And, you know, some of them will end up in local minima, right? You know, as we're — you know, if this is, like, gradient descent — like, people are trying to go to the most optimized solution. Some people will race to a local minima. And then the question is, like: how do you leap — how do you scoot back over to, like, the absolute minima?

而其中一些会掉进局部最优。如果把这看成梯度下降——大家都在往最优解走,有些人会一头冲进某个局部最优。然后问题就来了:你怎么跳出来,挪回到全局最优去?

And to some extent, like — Nvidia will always be more general purpose than anyone else's chip, in general, um, at least on a parallel AI compute basis. Because they have so many customers who care about different things, who will always give them feedback in the design.

某种程度上——Nvidia 永远会比其他任何人的芯片更通用,至少在并行 AI 计算这个范畴里。因为他们的客户太多、关心的东西各不相同,会持续在设计上给他们反馈。

You know, the minima will always be better than them — but is that minima a local minima?

专用方案的"最优点"永远会比他们更优——但那个最优点是不是局部最优?

Like, is the TPU, or Trainium, or Groq, or Cerebras, or whoever's design optimized awesomely for here — but in the end state, actually, you've got to go over here, and so they're the wrong…

TPU、Trainium、Groq、Cerebras,随便谁的设计,可能为"这里"优化得漂亮至极——但终局其实在"那里",于是他们押错了……

Um, and maybe they make — they're great for a little bit of time, but then they end up being wrong. It's like — that's the real question.

也许他们风光一阵子,最后却是错的。这才是真正的问题。

Um, and so I think there will be a big market for general-purpose AI compute. Um, because you talk to people at labs — they don't even know what architecture they're going to be doing in a year. Like, right — like, they literally don't know what architecture they're going to be doing in a year.

所以我认为通用 AI 算力会有一个很大的市场。因为你去跟实验室的人聊——他们连一年后自己会用什么架构都不知道。真的,他们真不知道一年后的架构长什么样。

They have bets. They have many research bets, and that's this exciting thing — but they don't know where it's going.

他们有押注,有很多研究方向的押注,这正是激动人心的地方——但他们不知道最终会走向哪。

Generally, they, like, know what hardware they have, and they're trying to co-optimize. But ultimately, like, if a new breakthrough happens on model architecture — it's like, just replace the attention mechanism with something else, right? Who knows.

一般来说,他们知道自己手里有什么硬件,然后努力做协同优化。但如果模型架构上出现新突破——比如干脆把注意力机制换成别的东西,谁说得准呢。

Or, you know, all of a sudden, you know, something happens — the best hardware will change. And therefore, like, are people going to make five-year investments on hardware solely on, you know, an ASIC that is more specialized? Or are they going to have some bucket of more general-purpose compute?

或者突然发生点什么,"最好的硬件"就换人了。那么,大家会把五年期的硬件投资全押在一颗更专用的 ASIC 上吗?还是会留一桶更通用的算力?

And so you see this with, like — Google's paying $11 an hour per GPU to xAI for GPUs, right? Like, that's insane, right?

你已经能看到这种事了——Google 在按每块 GPU 每小时 11 美金的价格,向 xAI 租 GPU。这很疯狂,对吧?

It's a very high amount — obviously compute is limited, and so on and so forth — but it's, like, very, like, insane. But at the same — you know, despite the fact that they have TPUs. And so there's, like, some questions there: like, why do they do that?

这个价非常高——当然算力紧缺等等——但真的很疯狂。关键是,他们自己有 TPU。那问题来了:他们为什么这么干?

Um, Google actually has three different design programs for TPUs. They're making a TPU with Broadcom; that's a different architecture than the TPU with MediaTek; that's a different TPU than the architecture that is — you know, I won't disclose — you know, by research.

实际上 Google 有三条不同的 TPU 设计线。他们和 Broadcom 做一款 TPU;和 MediaTek 做的又是另一种架构;而第三种架构又不一样——具体我不披露——这是研究得来的。

Um, but, you know, they're making different architectures. It's not just, like, "Oh, they're making TPUs with a couple vendors, it's the same architecture." It's different architectures. And the third one is a very different architecture from the first two.

重点是他们在做的是不同的架构。不是"哦,他们找了几家供应商做同一种 TPU"——是不同的架构。而且第三种跟前两种差别非常大。

And so I think people recognize that the local minima can happen. And therefore, um, I think everyone will have their own ASIC program. I think everyone will deploy billions of dollars of their own ASICs — tens of billions of dollars. In the case of Google, hundreds of billions of dollars a year of their own ASICs.

所以我认为大家都意识到局部最优是会发生的。因此每家都会有自己的 ASIC 项目,每家都会部署数十亿、上百亿美金的自研 ASIC——Google 的话,一年几千亿美金的自研 ASIC。

But ultimately, they're also going to have workloads that don't use TPUs, right? Some of the Google bets that are not Gemini/DeepMind actually primarily use GPUs — they don't use TPUs. Um, some of them also primarily use TPUs, right?

但同时,他们也会有不用 TPU 的工作负载。Google 一些非 Gemini/DeepMind 的业务押注,其实主要用 GPU,不用 TPU;当然也有些主要用 TPU。

It's a bit of a broad thing, but, like, you know — maybe for drug discovery, or for Waymo, you might not want to use TPUs. I won't say which one it is. But, like, you know, there's different architecture bets and different paths for AI.

说得笼统一点——比如药物发现,或者 Waymo,可能就不想用 TPU。具体是哪个我不说。但 AI 有不同的架构押注、不同的路径。

AI for science may have different algorithmic patterns than, than general intelligence, AGI models.

AI for science 的算法模式,可能和通用智能、AGI 模型不一样。

Um, and so I think we'll see diversity continue to proliferate. Yeah. And because the market has gotten so big, niches will be carved out. And so that makes it possible for companies to have their niche and actually make money — even if the majority of the pie goes to Nvidia and TPU and Trainium.

所以我认为多样性会持续扩散。而且因为市场已经这么大,各种利基会被切出来——公司守住自己的利基就真能赚到钱,哪怕蛋糕的大头归 Nvidia、TPU 和 Trainium。

Chapter 16

Compute Crunch & the Profit Flywheel

算力荒与 Anthropic 的盈利飞轮
50:51 — 56:55 · 今年 20GW · Fable 的 TAM · 80% 毛利 · 杠杆之辩
Sonya Huang00:50:47

Okay, love that. Can we talk about the data center buildout?

好,我喜欢这个答案。我们能聊聊数据中心建设潮吗?

Like — one, it seems like — I mean, by all accounts, if you look at the charts, like dollars per compute hour — we are in the middle of a crazy compute crunch.

首先,从各方数据看——看每算力小时的美元价格那些图——我们正处在一场疯狂的算力荒中间。

Um, and it seems like it's both a demand and supply side crunch, right? Demand for long agents skyrocketing; supply — all these data center buildouts are delayed.

而且看起来是需求和供给两头一起紧:长时 agent 的需求在飙升;供给端,这些数据中心建设全在延期。

Um, do you think we're in a compute crunch for the foreseeable future, or do you think it alleviates at some point?

你觉得可见的未来我们都会处在算力荒里,还是某个时点会缓解?

Dylan Patel00:51:10

Yes — every quarter we're deploying vastly more compute than the prior quarter, and there's more data centers built than the prior quarter.

是这样——每个季度我们部署的算力都远超上个季度,建成的数据中心也多于上个季度。

Um, this year there's going to be 20 gigawatts — even accounting for the delays. And next year there's going to be more than 30 gigawatts, accounting for the delays.

今年会有 20 吉瓦——已经把延期算进去了。明年会超过 30 吉瓦,同样算上延期。

Um, of course delays happen on everything, right? Anything hardware can have a delay. That's just the reality of life.

当然,万事都会延期。任何硬件都可能延期,这就是现实。

Are we going to have a compute crunch for the rest of our lives? It depends on what happens with models.

我们会一辈子处在算力荒里吗?这取决于模型接下来怎么走。

But, like — the TAM for Mythos, you know, Mythos 5, Fable 5, is not just, like, 2x that of Opus, right? The model is so much better, and it can do so many more tasks, that the TAM for it is way larger than that.

但你看——Mythos 5、Fable 5 的 TAM,不只是 Opus 的 2 倍。模型好太多了、能做的任务多太多了,它的 TAM 远大于 2 倍。

And yet compute in the world did not double in the last, you know, six months, right? From, you know — maybe, like, seven or eight months since Opus 4.5 launched — to now.

然而全世界的算力在过去六个月里并没有翻倍——从 Opus 4.5 发布到现在,大概七八个月。

You know, 4.6, 4.7, 4.8 were improvements. But Fable and Mythos were, like, a huge step-function improvement.

4.6、4.7、4.8 是渐进式改进,但 Fable 和 Mythos 是一次巨大的阶跃式提升。

The world's compute did not double — or quadruple, or whatever — in that same time frame. But the demand for useful tasks that can be done by AI — the number of useful tasks and the value of them that can be done by AI — has.

同一时间段里,全世界的算力没有翻倍、更没有翻四倍。但"AI 能完成的有价值任务"——任务的数量和价值——翻了。

And so now the question is: what happens?

于是问题变成:接下来会怎样?

Well, obviously Anthropic in Q2 is profitable — their net income profitable, um, excluding stock-based compensation. Um, and I think by Q3 they may even be profitable including stock-based compensation. That's, like, how profitable they're getting.

显然,Anthropic 二季度已经盈利了——净利润为正,不含股权激励。我认为到三季度,他们把股权激励算进去可能都盈利。他们的赚钱能力已经到这个程度了。

And their margins on an Opus token — at least an Opus 4.8 token — is, like, north of 80% for the API price.

他们在 Opus token 上的毛利——至少 Opus 4.8 的 token——按 API 价格算超过 80%。

They've got a lot of deals where their total corporate gross margins get clawed down a little bit, uh, because of, like, how they do Bedrock deals and Vertex deals and things like that. But ultimately, their per-token margin is so high.

他们有不少交易会把公司整体毛利往下拽一点——比如 Bedrock、Vertex 那类分销协议的结构。但归根结底,单 token 毛利高得吓人。

Well, then — they have the capability to pay, ultimately, for every GPU they buy, at above market rate.

那么——他们最终有能力为买到的每一块 GPU 支付高于市价的钱。

You know, they also bought GPUs at above market rate from SpaceX — which is below the rate of Google, but that's because they signed earlier.

他们也确实从 SpaceX 那里按高于市价买了 GPU——虽然低于 Google 付的价,但那是因为他们签得早。

Um, you know, it's something that, you know, other companies — maybe a venture-backed company, or a company that's not really got positive, uh, margins — can't necessarily do, right?

这是其他公司——比如靠融资续命的公司、毛利还没转正的公司——未必做得到的。

What is the cost-benefit ratio? Like: every GPU I rent — because I'm out of compute capacity — I can immediately turn around and sell tokens on it. Or every TPU, or every Trainium — I can immediately sell tokens on it at a positive margin.

成本收益比是什么?我因为算力不够而租进的每一块 GPU,都能立刻转手在上面卖 token;每一块 TPU、每一块 Trainium,都能立刻以正毛利在上面卖 token。

And if I'm running 75% gross margin and I double the cost of the compute — it's fine, I'm still running 50% gross margin.

如果我毛利 75%,算力成本翻倍——没关系,我还有 50% 的毛利。

And spinning up more compute nodes is not really necessarily a human-requiring task for them, if they're renting them.

而且如果是租来的,扩容更多计算节点对他们来说都算不上一件需要人力的事。

And so, ultimately, it's like — well, my NOI still goes up, right? And so I'm going to rent GPUs at whatever price. At some level, whatever price I want to pay, I can pay.

所以最终就是——我的净营业收入照样在涨。那我就按任何价格去租 GPU。某种程度上,我想付什么价就付得起什么价。

Sonya Huang00:53:47

I have almost the reverse question: of, like, at some point does this compute buildout go bump in the night?

我的问题几乎是反过来的:这场算力建设潮会不会在某个时点半夜撞鬼?

Earlier today, I think there was a tweet — like, Crusoe publicly said one of their customers had asked to halt construction on one of their data center buildouts.

今天早些时候好像有条推文——Crusoe 公开说,他们有个客户要求暂停某个数据中心项目的施工。

Like, it seems like everybody in the ecosystem is so levered right now, to, like: we got to build, we got to go build, we got to build.

感觉现在生态里所有人都杠杆拉满,一门心思:得建、快去建、接着建。

High leverage, high growth, to me, is, like — makes me very, very nervous as an investor.

高杠杆加高增长,作为投资人,这让我非常非常紧张。

Dylan Patel00:54:08

Wait, hold on. High leverage, high growth means a small amount of equity has huge upside. You're not a debt investor — you're an equity investor, right? Let's go!

等等。高杠杆高增长,意味着一小笔股权有巨大的上行空间。你又不是债权投资人——你是股权投资人,对吧?冲啊!

Um — look, you've got to go to the school of private equity. Levered buyouts only.

呃——你得去私募股权学院进修一下。只做杠杆收购。

Sonya Huang00:54:23

I actually come from the school of private equity.

我还真是私募股权出身的。

Dylan Patel

Oh, awesome.

哦,厉害。

Shaun Maguire00:54:26

She forgot the school. It's been a VC for too long.

她把那个学院忘光了。当 VC 太久了。

Sonya Huang00:54:28

No, I just do revenue multiples.

不,我现在只看营收倍数。

No, but — do you see any signs of that? Are you worried about that?

说正经的——你看到那种迹象了吗?你担心吗?

Dylan Patel00:54:34

I see what you mean, right. And that sort of goes back to the model point, right?

我明白你的意思。这又绕回模型那个论点了。

Obviously, if the models expanding the total economically valuable work — sort of the "dark GDP", uh, report that we did, that you mentioned earlier — if the work that these models can do does not expand faster than the compute capacity, then that tide turns, right?

显然,如果模型扩张"总的经济有价值工作"——就是我们做的、你刚提到的那份 dark GDP 报告——如果模型能干的活扩张得没有算力产能快,那潮水就会掉头。

And over the last six months, that tide has been, you know, very much levered in this direction of, um — you know, the models can do more work, or are expanding their TAM of work they can do, faster than the compute is increasing. And so prices go up.

而过去六个月,潮水一直强烈地偏向这一边:模型能干的活、能干的活的 TAM,扩张得比算力增长更快。所以价格上涨。

It's very possible that, all of a sudden, model progress stops.

模型进步突然停滞,完全是有可能的。

You talk to anyone at Anthropic or OpenAI — maybe they're drinking the Kool-Aid — but you talk to basically all of them, they're like: "No, no, no, no. Model progress still goes up."

你去问 Anthropic 或 OpenAI 的任何人——也许他们是在自灌迷魂汤——但基本上所有人都会说:"不不不不,模型还会继续进步。"

Um, and so, you know, ultimately, you know, current methods could stall somewhere. I'm not sure where that would be.

当然,现有方法总可能在某处失速。我说不准会是哪里。

It seems like we have line of sight to model improvement — rapid model improvement.

但看起来,模型改进——快速的模型改进——是有清晰视线的。

And in fact, models are improving faster than they were six months ago or a year ago. Because there's — I wouldn't call it recursive self-improvement, but basically, the models are helping write all the infra and launch the next model sooner and sooner and sooner.

事实上,模型现在改进的速度比六个月前、一年前还快。因为——我不想叫它递归自我改进,但本质上,模型在帮忙写全部的基础设施,让下一代模型越来越早地发布。

So you've got this, like, pseudo recursive self-improvement loop going, and so the models are getting better and better and better, faster.

所以这里有一个"伪递归自我改进"的循环在转,模型变好的速度本身也在加快。

Um, but ultimately, you know, capital is a big problem — which is why Google raised capital.

但归根结底,资本是个大问题——这就是 Google 融资的原因。

You know, they've got an ungodly amount of SpaceX, right? They own, like, 5% of the company.

要知道,他们手里握着多到离谱的 SpaceX 股份——大约占公司 5%。

Shaun Maguire00:55:57

I think a little more, but yeah. Yeah, maybe.

我记得还要多一点,不过差不多。

Dylan Patel00:55:59

I think at one point they had, like, 10%. Larry Page invested a billion dollars at a $10 billion valuation, got 10% of the company, it got diluted — like, all this. But that was one of the greatest investments of all time. Good job, Larry. The guy.

我记得他们一度有 10%。Larry Page 在 100 亿美金估值时投了 10 亿,拿了公司 10%,后来被稀释了之类。但那是有史以来最伟大的投资之一。干得漂亮,Larry。这家伙。

So they know they have, like, a hundred billion dollars in the bank that they can sell in, you know, nine months or whatever, from the lockup.

所以他们清楚自己银行里躺着差不多一千亿美金——锁定期一过,九个月后就能变现。

And they have all the gross profit they do. And yet they still modeled it, and they were like, "We need to raise capital." And so they did an offering. And it's like — that's insane. So that tells you how much they think they need to spend.

再加上他们那么高的毛利。然而他们建完模型还是得出结论:"我们需要融资。"于是做了增发。这太疯狂了——这告诉你,他们认为自己要花的钱有多大。

But capital is, like, really — you know, Meta did announce that they're going to do a raise. Stock tanked. People don't like it.

但资本真的是——Meta 也宣布了要融资,股价应声跳水,大家不喜欢这事。

But, you know, all these companies are going to raise capital, whether it be debt or equity.

但所有这些公司都会融资,债权也好,股权也好。

At some point, the money spigots will have to, you know, slow down.

总有一天,资金的龙头会不得不拧小。

But right now, every GPU that Amazon adds — they're making higher revenue. Or every TPU or Trainium, you know, whoever, anyone adds — is making gross profit.

但眼下,Amazon 每加一块 GPU,营收就更高;任何人每加一块 TPU 或 Trainium,都在产生毛利。

Chapter 17

Not All Gigawatts Are Equal

石油纯度类比、功率魔法与 GW 定价
56:55 — 64:17 · Trainium <$10B/GW · $/kW/月 · Google 塞 1.5GW
Shaun Maguire

I'll do a little bit of a tee-up on this, to turn it into a question for you.

我先铺垫一下,再把它变成一个问题抛给你。

But, like, as we talk about this, for me, the thing that's going through my head — that is almost an alternative hypothesis for, like, the Crusoe example — I'm going to use an analogy in oil.

聊到这儿,我脑子里转的东西——差不多是对 Crusoe 那个例子的另一种假说——我用石油打个比方。

Like, in oil, Saudi Arabia has way lower cost per barrel to produce oil than a lot of other countries. There's also, like, the purity of the oil — a lot of, you know, Saudi has generally, like, very low contaminants in their oil, which makes refining easier, all of this.

在石油行业,沙特每桶油的开采成本比很多国家低得多。还有油的纯度——沙特的油杂质普遍很低,炼化更容易,诸如此类。

The question for me is, like: when you look at, for every gigawatt that's being put in the ground — of, call it the 20 gigawatts coming online today — like, how much homogeneity do you see in those gigawatts?

我的问题是:落地的每一个吉瓦——就说今天上线的这 20 吉瓦——在你看来,它们之间有多同质?

Is it something like — and you can tell me whatever metric you think is right — but, like, are Google's gigawatts two times more valuable than, say, most NeoClouds'? Because they have optical switches, and they have, like — they've been doing it for a long time, and, like, they know how to do power smoothing.

比如说——用什么指标你说了算——Google 的吉瓦会不会比大多数 NeoCloud 的值钱两倍?因为他们有光交换机、干这行干了很久、懂怎么做功率平滑。

Because I think this could be the alternative hypothesis: that some of the people that are — it's like, the people that are good at building data centers, they should just do it to the max, because there's so much demand and they're so much better at it. But then maybe we're starting to see the early signs of the people that are, like, not as good at it kind of getting hit a little.

这可能就是那个替代假说:擅长建数据中心的人应该开足马力干到底,因为需求太大、他们又强得多;而我们也许正开始看到早期迹象——那些没那么擅长的人,开始挨打了。

So, like, I don't know the reality here. I'm just curious how you think about this.

我不知道真实情况如何,只是好奇你怎么想。

Dylan Patel00:58:18

So, so far, um, there are metrics for this, right?

到目前为止,这事是有指标可查的。

So, uh, Trainium sells at sub-$10 billion per gigawatt rental rate, uh, to Anthropic and to OpenAI.

Trainium 卖给 Anthropic 和 OpenAI 的租用价,折合每吉瓦不到 100 亿美金。

GPUs — at least before the craziness of the last six months — usually went around $12 to $13 billion per gigawatt. So the rental rate — and this is from a NeoCloud, versus Amazon even. And now, when Amazon sells GPUs, they'd also be 13 or so.

GPU——至少在过去六个月的疯狂之前——通常在每吉瓦 120 到 130 亿美金上下。这是租用价——NeoCloud 的价,甚至 Amazon 也一样;现在 Amazon 卖 GPU,也差不多 130 亿。

Shaun Maguire00:58:42

And my understanding of that also is that those numbers, like — Amazon subsidized that a little bit. So it's like — I actually think the disparity was even more.

而且据我了解,那些数字——Amazon 是补贴过一点的。所以我其实觉得差距比这还大。

Dylan Patel00:58:52

It's less than 10. It's less than 10, but there's, like, some weird — basically…

是低于 100 亿。低于 100 亿,但里面有些奇怪的结构……

Shaun Maguire

And, like, look — my understanding, obviously, like — Anthropic played a big role in making Trainium useful, in terms of, you know, writing all the libraries, etc.

而且据我了解——Anthropic 在让 Trainium 变得可用这件事上出了大力,比如把那些库全写了,等等。

Everything I hear is that Trainium's really freaking good hardware, and it's getting, like, way better. And obviously Anthropic is now using it a lot. So hopefully we would see that price go up.

我听到的所有说法都是:Trainium 是真他妈好的硬件,而且还在快速变强。Anthropic 现在显然在大规模用它。所以但愿我们能看到它的价格涨上去。

Dylan Patel00:59:18

You know, like — per the deal they did, there was actually, like, a floor mechanism. And, like, if it didn't do well, it would be, like, cheaper — to the point where it's cancelable. And, you know, if it did really well, the price is kind of higher.

按他们签的那份协议,其实有一个保底机制:如果硬件表现不好,价格会更低——低到可以直接取消;如果表现很好,价格就相应更高。

Um, but effectively, um, less than 10, right, is where Trainium shakes out at.

但实际落点上,Trainium 就是低于每吉瓦 100 亿。

Whereas GPUs — I mean, the SpaceX deal, again, was, like, 25 or something crazy — billion dollars per gigawatt, or $25 million per megawatt, right, a year, rental rate, with Google. I was like — that's a crazy divergence.

而 GPU——SpaceX 和 Google 那笔交易,是疯狂的每吉瓦 250 亿美金上下,也就是每兆瓦每年 2500 万美金的租用价。我当时就想——这个价差太疯狂了。

Now, obviously, if Amazon was selling Trainium today, it'd probably be more expensive than 10, because of the compute shortages.

当然,如果 Amazon 今天再卖 Trainium,大概会高于 100 亿,因为算力短缺。

But you do see this already, in the sense of, uh — with data centers, oftentimes the rental price of a data center, if you're doing colocation, right — not compute in there, but just power, "here's the data center" — um, you price it generally on dollars per kilowatt per month.

但这种分化已经能看到了——数据中心这边,做托管(colocation)的话——里面不带算力,只有电力,"数据中心在这儿"——租金一般按每千瓦每月多少美元定价。

And so they used to be $60 per kilowatt per month. And now you see things transacting at anywhere from, like, 120 to 160.

以前是每千瓦每月 60 美金。现在的成交价在 120 到 160 之间。

Um, but different quality data centers — actually, I've seen data centers go as high as 200, um, when the customer is not such a great credit rating and the data center is a pretty good one.

而数据中心的质量参差——我见过高到 200 的:客户信用评级不太行、数据中心本身又相当好的时候。

And I've seen stuff go as low as 100 still — or, in like, India, go as low as 80 — because the grid's not reliable, the internet connection's not great, and it's a pretty mid data center. But at least it's a data center.

也见过低到 100 的——在印度甚至低到 80——因为电网不稳、网络一般、数据中心本身也就中不溜。但好歹是个数据中心。

Um, and so you see this huge discrepancy there already.

所以巨大的价差已经存在了。

Um, in the case of, like, data center construction — usually the pitfalls: they just fail.

至于数据中心建设——常见的坑就是:直接失败。

There's a lot of people who, you know, claim — they're, like, four guys, they're like, "Yeah, we — here, I bought some turbines. I put the money down for them. I'm going to build a data center." And then they get delayed, delayed, delayed, and fail.

很多人号称要干——四个人的小团队,说"看,我买了几台燃气轮机,定金都付了,我要建数据中心"。然后延期、延期、再延期,黄了。

Um, so you have to, like, probability-weight, time-weight, time-lag the teams that suck versus don't.

所以你得给团队做概率加权、时间加权、时滞修正——分清哪些团队不行、哪些行。

Um, and sort of, you know, our data center model does that. We kind of track every data center, uh, and try and do this for every single one, based on, you know, equipment that they're using and all these things.

我们的数据中心模型就在做这件事。我们几乎追踪每一座数据中心,根据他们用的设备等等,对每一座都做这种测算。

One of the things you mentioned about Google: you know, in a gigawatt data center, they'll actually put, like, 1.5 gigawatts of hardware.

你刚提到 Google——他们真的会在一个 1 吉瓦的数据中心里,塞进 1.5 吉瓦的硬件。

And because they have such understanding all the way from workload to — you know, they're able to slosh the power around.

因为他们从工作负载开始把整条链路都吃透了——他们能把功率"晃"来晃去地调配。

And so instead of, you know — a gigawatt of compute typically runs at, like, 60 or 70% utilization in terms of power consumption — not utilization of the hardware; someone's always renting it — um, they're now running it at, like — you know, that 60 to 70% means it's at a gigawatt, and you're using the full gigawatt.

通常 1 吉瓦的算力,功耗层面只跑到 60%-70% 的利用率——不是硬件没人用,硬件总是有人租的——而他们能做到:那个 60%-70% 恰好顶满 1 吉瓦,把整个吉瓦全用掉。

Um, you see people doing deals with — including Google — with utilities, where they're like: "Oh, well, I know this grid can sustainably take a gigawatt. But, you know, except for three days of the year, you can actually do two gigawatts. So give me two gigawatts, and then just tell me to turn off." And so they'll do that.

你还会看到有人——包括 Google——跟电力公司做这种交易:"我知道这张电网可持续承载 1 吉瓦。但一年里除了三天,其实能给到 2 吉瓦。那就给我 2 吉瓦,到那三天你喊我关机就行。"他们真会这么干。

And so these sorts of tricks — and then you need to have supreme management of workload, backup power, all these things, um, generators on site, to figure out how to actually keep it two gigawatts sustainably.

这类花招——再配上对工作负载、备用电力这些东西的顶级管理,加上现场发电机,才能真正把 2 吉瓦可持续地稳住。

When people do this, they're able to charge more. Whether it be: I'm actually selling two gigawatts despite only having one gigawatt, because those three days I'm able to deal with via battery, gas, etc. Or: I figured out how to build power on site — now I have a gigawatt where no one else does, and so I'm able to do it quickly.

做到这些的人就能收更高的钱。要么是:我明明只有 1 吉瓦,却在卖 2 吉瓦,因为那三天我能靠电池、燃气顶过去;要么是:我搞定了现场发电——在别人都没有的地方我有 1 吉瓦,而且能快速交付。

Um, it's not necessarily transacting for a higher price — it's that I'm selling more gigawatts. And sometimes there are levers where you're selling more gigawatts, where each gigawatt is selling at a different price.

这不一定体现为单价更高——而是我卖出了更多的吉瓦。有时候的杠杆是:你卖出更多吉瓦,而每个吉瓦的售价还各不相同。

Um, I think it's more — on the data center and energy layer, it's more about just having it versus not, and then that being delayed or not. It's more binary.

所以在数据中心和能源这一层,更多是"有还是没有"、"延不延期"的问题。它更接近二元。

But on the compute side, I do think there's a lot more, um, interesting work there. Right? A gigawatt given to Anthropic is objectively worth more revenue than a gigawatt given to OpenAI.

但算力这一层,有意思的东西就多得多了。同样一个吉瓦,给 Anthropic 客观上能产出比给 OpenAI 更多的收入。

And it seems that both of them could sell every gigawatt that they have right now — uh, given rate-limit problems and token max limits and all these sorts of things at OpenAI and Anthropic. Uh, especially since Codex 5.5 came out — it's much better.

而且看起来,他们俩手里的每一个吉瓦现在都能卖光——看看 OpenAI 和 Anthropic 的限流问题、token 上限这些就知道了。尤其 Codex 5.5 发布之后——它强了很多。

And then, likewise, if you gave a gigawatt to SpaceX, you know, they'd turn…

同样地,如果你给 SpaceX 一个吉瓦,他们会……

Shaun Maguire01:03:04

My guess — like, my suspicion — is that they probably make better use of the, you know, hardware than most people.

我猜——我的直觉是——他们对硬件的利用大概比大多数人都好。

Um, just like — I think people underestimate how much networking experience they have from Starlink in particular, and also how much just, like, power management experience they have via — from Tesla.

我觉得大家低估了他们从 Starlink 积累的网络经验,也低估了他们从 Tesla 带来的电力管理经验。

Dylan Patel01:03:26

Yeah — people like Brett Mayo are, like, incredible, like…

对——像 Brett Mayo 这样的人,厉害到离谱……

Shaun Maguire

Pretty good.

相当强。

And so I think that, like, for me, that's actually — I think probably the thing that might — I don't actually know the answer, but I think that might be missing from the analysis a lot of people are doing.

所以对我来说——我其实不知道答案——但我觉得这可能正是很多人的分析里缺掉的那块。

Dylan Patel01:03:37

I think it's also the fact that, when CoreWeave builds a gigawatt — even though their GPU compute is objectively better than Amazon or Google or Microsoft's in terms of performance; we've tested the performance and reliability —

我觉得还有一个事实:CoreWeave 建一个吉瓦的时候——尽管他们的 GPU 算力在性能上客观优于 Amazon、Google 或 Microsoft;性能和可靠性我们都实测过——

— um, the problem is, Google sells it six months before they have it up, and they need to turn around and take that paper that they signed, to get debt, uh, with that credit backing, and then turn around so they can actually pay for the PO that they've already issued, you know, for the order that they've already issued.

——问题在于,他们(按上下文应指 CoreWeave)得在机房上线前六个月就把它卖掉,再拿着签好的合同当信用背书去举债,然后才能付得起自己已经开出去的采购订单。

Whereas SpaceX was like: "No, no, no — this is running now. Buy it."

而 SpaceX 是:"不不不——这套现在就在跑。直接买吧。"

Right — and it's a big discrepancy when you have a balance sheet to do that, versus not. And that also helps your revenue per megawatt, like, be much higher.

有没有一张能这么玩的资产负债表,差别是巨大的。这也会让你的每兆瓦收入高出一大截。

Chapter 18

Why NeoClouds Exist & Jensen's Multipolar World

NeoCloud 的窗口,与 Jensen 的棋局
64:17 — 70:15 · Amazon Cloud Crisis · 撒饵喂鱼 · Thinking Machines
Sonya Huang01:04:15

Why does the NeoCloud opportunity even exist? Because if you had asked me five years ago, I would have said the hyperscalers are going to own this.

NeoCloud 这个机会到底为什么会存在?五年前你要问我,我会说这块一定被超大规模云拿下。

And, you know, you mentioned just now, CoreWeave has better performance than the hyperscalers. Like, why does this opportunity exist — maybe at the macro level, and then at the execution level?

而你刚才还说,CoreWeave 的性能比超大规模云更好。这个机会为什么存在——先说宏观层面,再说执行层面?

Dylan Patel01:04:29

Yeah. So in 2023, I wrote a report that had, uh, Amazon really hate me. Um, it was called "Amazon Cloud Crisis."

好。2023 年我写了一份让 Amazon 恨透我的报告,标题叫《Amazon Cloud Crisis》。

So I talked about how Amazon was the best cloud, because they had their Nitro NICs, which offered, like, tenant isolation — all the hypervisor ran on the NIC, and then you could sell all the cores.

我在里面讲了 Amazon 为什么曾是最好的云:他们有 Nitro 网卡,提供租户隔离——整个 hypervisor 跑在网卡上,于是所有 CPU 核都能拿去卖钱。

And they had, you know, custom SSDs that they made — they'd buy the raw NAND, and they'd have lower cost because they'd buy the raw NAND and build their own SSDs.

他们还有自研的 SSD——直接采购裸 NAND 颗粒自己造 SSD,成本更低。

Um, and you know, they had their custom Graviton CPUs, and that drove down cost per core.

还有自研的 Graviton CPU,把单核成本压了下来。

And so they had all these things that enabled them to sell more cores, have better security, good networking, uh, better storage — but this was all for the traditional CPU, you know, cloud world.

所以他们有这一整套东西:卖更多核、更好的安全性、优秀的网络、更好的存储——但这一切都是为传统 CPU 云世界打造的。

But in the AI cloud, a lot of this stuff hurt performance, right? These Nitro NICs — bad for performance. Still are worse performance. Although they've caught up a lot, because they've had a couple iterations to, like, you know, improve them — but they're still worse for performance.

而在 AI 云里,这套东西很多反而拖累性能。Nitro 网卡——对性能就是坏事,至今仍然更差。虽然经过几轮迭代已经追上不少,但性能上还是吃亏。

Um, a lot of the security stuff doesn't matter — because it's not like I'm time-slicing users, or splicing a socket into many users, right?

那些安全能力也大多派不上用场——因为我又不是在给用户做分时切片,或者把一颗插槽切给一堆用户。

It's like — no one rents a single GPU in an 8-GPU server. No one rents a single GPU in a 72-GPU rack. They rent the whole rack. And in fact, they rent many of the racks.

现实是——没人在 8 卡服务器里租单卡,没人在 72 卡机柜里租单卡。人家租整柜,而且一租就是很多柜。

And then there's no, like, "Oh, I rent for six hours and I give it back." It's — everyone has these long-term contracts.

也不存在"我租六小时就还回去"这种事——所有人签的都是长期合同。

So the mechanics of the GPU rental market meant that a lot of the expertise of the hyperscalers fell away.

所以 GPU 租赁市场的运行机制,让超大规模云的很多看家本领直接失效了。

Um, and a lot of the expertise that they did have — actually, some of it was detrimental, right? Network performance for Google and Amazon: they had custom networks that were better for traditional CPU work and for the stuff that they were doing, but actually worse for AI.

而他们真正拥有的那些本领——有一部分甚至是负资产。比如 Google 和 Amazon 的网络性能:他们的自研网络对传统 CPU 业务、对他们原有的场景更好,但对 AI 反而更差。

Um, and then in other cases, it's like — well, you know, Microsoft would save money by building their own data centers. But their data center teams actually were not that great.

还有别的例子——Microsoft 自建数据中心是能省钱,但他们的数据中心团队其实并不怎么样。

And so when it was predictable building, it was, like, fine. When it came time to, like, actually double your forecast for the year — it's like, they fell on their face, and they had to go get a bunch of NeoCloud capacity.

按部就班的可预测建设,没问题。可到了要把全年预测翻倍的时候——他们直接摔了个大跟头,只好跑去买一堆 NeoCloud 的容量。

I think — so, performance, I think, you know, is one. I think time to market's another one, right?

所以我认为——性能是一条,上市速度(time to market)是另一条。

You know, these massive organizations — no one's getting rich from building this data center faster, right?

在那些庞然大物般的组织里——没有谁会因为把这座数据中心建得更快而发财。

But you look at Crusoe, for example — Chase, and all the other people on the team; you know, I was going to name some people on the team, I'd rather not — you know, all these people are getting rich if they deliver this compute faster. They're, you know — they're hyper-levered equity owners.

但你看 Crusoe——Chase,还有团队里其他人;我本来想点几个名,还是算了——这些人只要把算力更快交付出来,就真的会发财。他们是超高杠杆的股权持有者。

Shaun Maguire01:06:48

Hey, look — they're also all coming from Bitcoin. And, you know — you're not supposed to say that.

嘿,你看——他们还都是比特币圈出来的。这话本来不该说的。

Dylan Patel01:06:52

Uh — I mean, a lot of the data center, like — their main data center guy came from Microsoft.

呃——其实数据中心那块——他们数据中心的头儿是从 Microsoft 来的。

Shaun Maguire01:06:55

I don't know, I'm just teasing. But it's, uh — you know, it's like — you learn a lot when you're in a very high-fluctuation, you know, market.

我不确定,开个玩笑而已。不过——在一个剧烈波动的市场里摸爬滚打,人确实能学到很多。

Sonya Huang01:07:03

What do you think — is Jensen playing 4D chess?

你觉得——Jensen 是在下四维棋吗?

Dylan Patel01:07:07

Jensen absolutely hates a world where all the hyperscalers have all the power.

Jensen 绝对痛恨一个"所有权力都握在超大规模云手里"的世界。

There's a reason he's, like, blowing money on, like, random AI labs — that, like, I don't even know if it makes sense to. But, like, you know, he's blowing money and pumping them up, and going to, you know, everyone around the world and saying, "You should invest in this company" — because he wants to create a multipolar world.

他往各种 AI 实验室里砸钱是有原因的——有些我甚至看不出砸得有没有道理。但他就是在砸钱、把它们抬起来,满世界跟人说"你该投这家公司"——因为他想创造一个多极世界。

That's why he loves Chinese labs — because he wants to create a multipolar world.

这也是他喜爱中国实验室的原因——他要一个多极世界。

A world where OpenAI, Anthropic and Google models are the only models, is one in which he's screwed. Yep.

一个只剩 OpenAI、Anthropic 和 Google 模型的世界,就是他完蛋的世界。没错。

Right. Um, a world in which, you know, the hyperscalers are the only ones building compute, is one he's screwed in. Yeah.

同样,一个只有超大规模云在建算力的世界,也是他完蛋的世界。

And so, you know, of course he needs to point the allocation gun at NeoClouds, help backstop their clusters, do anything and everything.

所以他当然要把"配额之枪"指向 NeoCloud,帮忙给他们的集群兜底,能做的全做。

Because while today, a GPU sold to Crusoe, and a GPU sold to, um, CoreWeave, and a GPU sold to Google and Amazon are all the same price for him — five years from now, Crusoe and CoreWeave existing means Google TPU will be weaker, and means Amazon Trainium will be weaker. And more inference being done with, you know, non-closed model labs is better for him.

因为今天,GPU 卖给 Crusoe、卖给 CoreWeave、卖给 Google 和 Amazon,对他来说价格都一样——但五年后,Crusoe 和 CoreWeave 的存在,意味着 Google TPU 更弱、Amazon Trainium 更弱。而更多推理跑在非闭源模型实验室那边,对他更有利。

So I think, you know, the NeoCloud ecosystem — it's, you know, these people that are wild west. These NeoLabs as well — a lot of them have investments from Nvidia. It's the wild west.

所以 NeoCloud 生态——这帮人就是狂野西部。NeoLab 们也一样——很多都拿了 Nvidia 的投资。就是狂野西部。

Some will fail — many will fail — but, you know, some will emerge as really great teams.

有些会失败——很多会失败——但也会有一些冒出来,成为真正伟大的团队。

Whether it be, you know, oddly, Crusoe — who's a bunch of crypto guys who then started building data centers and doing flared gas stuff. Or, you know, CoreWeave — who initially was a bunch of New York hedge fund guys…

不管是——说来奇妙——Crusoe:一帮加密货币出身的人,后来开始建数据中心、搞燃除气发电;还是 CoreWeave:最初是一帮纽约对冲基金的人……

Shaun Maguire01:08:30

They were also — and then they were doing —

他们也是——而且后来他们在搞——

Dylan Patel01:08:31

— and crypto guys. But then they, like, built — you know, there were a lot of people who didn't bubble up like them, who started around the same time, and just failed. Right?

——也是加密圈的。但他们真的建起来了——要知道,同期起步、却没能像他们那样冒出来的人有一大把,全失败了。

Shaun Maguire01:08:40

I gotta say, both those teams are phenomenal. They deserve a lot of credit — and it's like, that's your point.

必须说,这两个团队都是现象级的,值得大大的赞——这正好印证了你的观点。

Dylan Patel01:08:45

Yeah — I mean, my point is, like: you throw a bunch of, like, bait into the water, and the best fish will figure it out and survive, right?

对——我的意思是:你往水里撒一把饵,最强的鱼自然会想明白怎么活下来。

Um, and sort of the same way with the NeoClouds. And he hopes the NeoLabs as well. We'll see if any of the NeoLabs really bubble up.

NeoCloud 就是这么回事,他希望 NeoLab 也一样。NeoLab 里到底能不能真冒出来谁,拭目以待。

But, like, you know, Thinking Machines has a few hundred million dollars of ARR, right? That's pretty impressive — even though they've had, you know, in the media, it's like, "Oh, they've lost all this talent."

不过你看,Thinking Machines 已经有几亿美金的 ARR 了。相当了不起——尽管媒体上都在说"哦,他们人才流失了一大批"。

It's like — well, but Tinker is doing a few hundred million of ARR. Like, that's pretty impressive for, out of the gate, a product that's less than six months old or whatever.

可事实是——Tinker 在做着几亿美金的 ARR。一个上线不到六个月的产品,开局就到这个量级,相当了不起。

Um, and we hope the same happens to other NeoLabs. And so, um, you know — he wants a multipolar world.

我们希望其他 NeoLab 也能这样。总之——他要的是一个多极世界。

Shaun Maguire01:09:20

Truly, congratulations on the success.

真心恭喜你取得的成功。

Just the last thing I'll say is — I've seen a little bit of this — I think the public, they can probably tell from listening to you how hard you work. But, like, it's clear you've just been working your ass off for more than a decade, and it, you know, led to the last few years of being in the right place, right time.

最后说一句——这我多少亲眼见过——公众光是听你说话,大概就能感觉到你有多拼。但事实很清楚:你玩命干了十几年,才换来这几年的"天时地利"。

But, like, it's unbelievable what you've accomplished, and I know it's just the beginning.

你的成就令人难以置信,而我知道这只是开始。

Dylan Patel01:09:41

Thank you so much.

太感谢了。

Thank you for doing this.

谢谢你们来做这期节目。

Shaun Maguire

Awesome.

太棒了。