协同设计把 8x 变成 100x
"Instead of being multiplicative to 8x, it's actually 100x because you've optimized across all three layers."
"The real breakthrough innovation is when you leapfrog a few layers, you co-optimize and co-design them, and now all of a sudden you've taken what could have been a 2x here, 2x here, 2x here."
全篇题眼:硬件、系统软件、模型架构各自 2x 相乘本是 8x,跨三层协同优化能叠成 100x——真正的性能跃升不在单颗芯片。
TPU 客观优秀,却跑不动 DeepSeek
"TPUs suck at running DeepSeek, but they are really really great at running other kinds of models that don't run well on NVIDIA."
"Despite the fact that TPUs are objectively an amazing chip, and they run all of DeepMind and they do all the training for Anthropic as well on the pre-training side at least."
DeepSeek 的专家形状是为 Hopper/Blackwell 塑形的,芯片和模型深度绑定——协同设计的代价是换硬件就跑不好。
CUDA 护城河从来不真是 CUDA
"What people call the CUDA moat is not actually anything to do with CUDA."
"It's the fact that DeepSeek, Kimi and Zhipu and Alibaba and Tencent — all these companies, Xiaomi had an awesome model recently — their models are co-designed for GPUs and therefore if I want to run them on TPUs, in some cases they don't run really well on TPUs."
真正的锁定来自下游开源模型都为 GPU 协同设计,换 TPU 跑不好。模型公司愿同时适配四五款芯片、Claude/Codex 能写 kernel,软件层护城河在松动,但架构路径依赖更深。
稀疏 vs 稠密,决定各自的芯片
"The way OpenAI's models are headed, it would be a terrible decision for them to use TPUs potentially."
"OpenAI's are much more sparse, and that has benefits. And then Anthropic's are still sparse, but more dense in general, and that has different benefits."
OpenAI 走稀疏、Anthropic/Google 偏稠密,配上 NVLink 只连 72 片 vs Google ICI 能连 8000 片的网络差异,把两条模型路线拉向不同硬件。
推理将成为地球最大市场,远超石油
"Inference, whether it's open models or closed models, will be like one of the biggest markets in the world — much bigger than oil."
"Obviously use of tokens is going to be the biggest market and the value that's created from tokens is going to be the biggest market. Inference of AI will be many percentage points of the GDP."
token 的使用与其创造的价值会占到 GDP 好几个百分点,是全球最大市场之一。
同质量模型成本一年降 60x
"We've seen model cost drop for equivalent quality by like 60x a year."
"It's a relentless breakthrough after breakthrough after breakthrough that keeps driving efficiency and cost down. To stay on top of that you can't have point-in-time benchmarking."
软件库一周更新两次、优化不停迭代,所以 InferenceX 要做"活基准":15 种芯片、$50M 捐赠硬件,每天自动跑最新模型出帕累托最优曲线。
2040 年过半增量算力上太空
"If you look at like 2040, I think probably more than half of the incremental compute will be going in space. But if you look at 2030, I think it's sub 1%."
"Just OpenAI, Anthropic we'll have over 100 gigawatts combined by 2030, and then you'll add Meta and Google. By like 2040 it'll be terawatts."
近 5 年太空数据中心无关紧要,但地表能源与土地约束下,20 年后绝大多数算力会入轨;推理部署从 100GW 级涨到 terawatt 级。
按美元算,别按 token 算
"Who cares about volume by tokens? It's about the dollars."
"If Ford F150s are 5x ASP and they sell only half as much — the most lucrative market is pickup trucks in America."
收入集中在最好的大模型上,即便更贵用户也在迁移。衡量市场要看营收而非用量——这也是 Cerebras 的风险:SRAM 芯片扛不住 10 万亿参数加长上下文。
"AI 没有 ROI"最让他上头
"People are like 'AI has no ROI' — infuriates me. The line has been up and to the right in terms of capabilities this entire time."
"They're like 'look this benchmark didn't improve' — that's 'cause it said 90%. Look at the new benchmark, you saturated, now they're skyrocketing."
Dylan 的触发词是"AI 没 ROI"和"模型见顶":旧 benchmark 不涨是因为刷到饱和,换新的分数就在飞涨。
算力紧缺:TAM 扩张快过算力增长
"If the work that these models can do does not expand faster than the compute capacity, then that tide turns."
"Over the last six months that tide has been very much levered in this direction — the models can do more work, expanding their TAM of work faster than the compute is increasing. And so prices go up."
世界的算力半年没翻倍,但 AI 可做任务的价值翻倍了,所以价格涨、紧缺持续——除非模型进步某天停下。
Anthropic 单 token 毛利超 80%
"Their margins on an Opus 48 token is like north of 80% for the API price. If I'm running 75% gross margin and I double the cost of the compute, it's fine, I'm still running 50%."
"Every GPU I rent because I'm out of compute capacity I can immediately turn around and sell tokens on it at a positive margin. Anthropic in Q2 is profitable."
毛利够高,即便算力成本翻倍仍赚钱,所以能高于市价抢 GPU——这是"能出任意价买算力"的底气。
同一吉瓦,给 Anthropic 更值钱
"A gigawatt given to Anthropic is objectively worth more revenue than a gigawatt given to OpenAI."
"Google in a gigawatt data center will actually put 1.5 gigawatts of hardware and they're able to slosh the power around — so they're selling more gigawatts."
每一吉瓦的成色不同:软件与工作负载决定营收,Google 靠功率腾挪超卖、Trainium 每 GW 租金比 GPU 低一档——数据中心不是同质大宗商品。
Jensen 花钱买一个多极化世界
"Jensen absolutely hates a world where all the hyperscalers have all the power. That's why he loves Chinese labs — he wants to create a multipolar world."
"He needs to point the allocation gun at NeoClouds, help backstop their clusters. Five years from now Crusoe and CoreWeave existing means Google TPU will be weaker and Amazon Trainium will be weaker."
一个只有 OpenAI/Anthropic/Google 模型、只有超大厂建算力的世界,英伟达就完了。所以他砸钱扶持 neocloud/neolab、偏爱中国实验室——撒饵让最强的鱼活下来。
I think it's really fun inside of SemiAnalysis because we have 90 people and like a big chunk of them are technologists engineers across the whole supply chain.
我觉得 SemiAnalysis 内部特别好玩,因为我们有 90 号人,其中一大批是技术出身的工程师,横跨整条供应链。
Um, and then a big chunk is people who are formerly at hedge funds.
呃,然后还有一大批是以前在对冲基金干过的人。
And you see these arguments like people are like, "Oh, well that doesn't matter."
你就会看到这种争吵,有人说"嗐,那个根本不重要"。
And it's like, then someone's like, "Well, but cost."
然后又有人说"可是,成本呢"。
And then someone the engineers like, "No, no, no, but this technology is the coolest."
接着工程师那边又说"不不不,可是这个技术最酷了"。
And you see this you see this organically like fight it out.
你就看着他们自发地这么互掐。
Um, and we're pretty informal and you know, given the fact that I was a forum moderator, you can imagine what the the enjoying it.
呃,而且我们挺随意的,你想想我以前是论坛版主,就能想象我有多享受这个。
You don't wrestle with a pig because a pig enjoys it, right?
你不会去跟猪摔跤,因为猪乐在其中,对吧?
Exactly.
没错。
We're here in the SemiAnalysis office with Dylan Patel.
我们现在就在 SemiAnalysis 的办公室,跟 Dylan Patel 在一起。
You know, I'm Shaun from Sequoia.
我是 Sequoia 的 Shaun。
My partner Sonya Huang.
这是我的合伙人 Sonya Huang。
It's pretty insane what you've done.
你做出来的这些东西,真的挺疯狂的。
Semi semis 5 years ago were not very sexy in the west.
半导体在五年前的西方,并不怎么性感。
They were sexy in the east but uh people here in the west had kind of forgotten about them.
在东方是挺火的,但呃西方这边的人差不多把它们给忘了。
You did not forget about them though.
你却没忘。
You went very long.
你重仓做多了它。
You created probably the premier research company in the space that's been educating the world and you know the state of the art from very technical details to supply chain you know to the bigger picture.
你打造了这个领域大概最顶尖的研究公司,一直在给全世界做科普,从最技术性的细节到供应链,再到整个大局。
Um there's rumors that SemiAnalysis recently passed 100 million of revenue.
呃,有传言说 SemiAnalysis 最近营收过了一个亿。
I don't know how accurate those are.
我不知道这些传言有多准。
Whatever the numbers are you guys are crushing.
不管数字是多少,你们都杀疯了。
It's it's as accurate as the information is.
这传言准不准,就看信息本身准不准了。
Yeah.
嗯。
Cool.
好。
You know, you never you never know.
你懂的,谁也说不准。
Um there's also rumors that you might start a venture fund like you know I I hear all the time in the ecosystem people wanting you know affiliation with SemiAnalysis.
呃,还有传言说你可能要开个风险基金,我在这个圈子里老是听到有人想跟 SemiAnalysis 扯上关系。
You you've built this trusted brand and so whatever you do it's working.
你打造了这么一个受信任的品牌,所以你干什么都成。
This clearly like just the beginning of the journey for you.
对你来说,这显然只是这段旅程的开头。
Congratulations all of that.
这些都恭喜你。
But how did this happen?
但这一切是怎么发生的?
Like how did you first question is like what is the background?
第一个问题是,你的背景是什么?
How did you kind of get to where you are now?
你是怎么一步步走到今天的?
Well, well, when I was a young boy in the, you know, coming out of the womb.
这个嘛,我还是个小男孩的时候,呃,刚从娘胎里出来那会儿。
No.
开玩笑。
So, so, okay.
好,好吧。
So, I grew up in like a small business.
我是在小生意的环境里长大的。
My parents had a motel.
我父母开了家汽车旅馆。
We lived in the motel.
我们就住在那家汽车旅馆里。
We laid our gas station.
我们还经营着一家加油站。
So, you know, uh I was selling.
所以,呃,我从小就在卖东西。
You know, I joke a lot of times the first neural network I trained was uh racially and and visually profiling people based on when they enter the gas station, which cigarette to uh pick.
我经常开玩笑说,我训练的第一个神经网络,就是在客人走进加油站时,呃靠种族和外貌去判断,该拿哪种烟给他。
Right.
对吧。
Basically, you know, the cigarettes were all extruded across the top and I was too short to actually like, you know, reach them and technically it wasn't legal to sell cigarettes at that age, but whatever.
基本上,烟都一排排码在最上面,我个子太矮根本够不着,而且严格说那个岁数卖烟是违法的,但管它呢。
I I had to move the step stool over to the right area.
我得先把那个小凳子挪到对的位置。
I started working my first job was before it was legal, too.
我第一份工作也是在合法年龄之前就开始干了。
So, but it's good experience.
不过,那也是很好的历练。
Well, I didn't get paid, right?
我可是没工资的,对吧?
It's a family business.
这是家族生意。
Same.
一样。
Same.
一样。
Um, but yeah, we had our motel and then across the street was our gas station.
呃,不过对,我们有汽车旅馆,街对面就是我们的加油站。
So, you know, sometimes, you know, you know, someone would walk in and so like if a old white lady with curly hair walked in, I'd move the ladder or the step stool over to where the camels are.
所以有时候,有人一走进来,比如一个卷发的白人老太太进来,我就会把梯子或者小凳子挪到 Camel 骆驼牌那边。
And if you know and you know different different age, demographic, profession, you know, race, etc. I would move the step stool over and I joke this is the first neural network I trained because if I waited for them to tell me I'd have to like move it over and then I'd step up versus like just being ready.
碰上不同年龄、不同人群、不同职业、不同种族等等,我就会把小凳子挪过去,我开玩笑说这是我训练的第一个神经网络,因为如果我等他们开口告诉我,我就得再去挪凳子、再踩上去,而不是提前就准备好。
Um so you know menthols versus, you know, 100 slims and all these things, you know, I joke that's the first neural network I trained.
呃,所以薄荷烟啊、100 细支啊这些区别,我开玩笑说,这就是我训练的第一个神经网络。
But I grew up in family businesses.
总之我是在家族生意里长大的。
Um lived in a motel and um it all really goes back to when I was like, you know, it was my 8th birthday.
呃,住在汽车旅馆里,呃,一切真正的起点是在我八岁生日那会儿。
Um, my birthday's in May.
呃,我生日在五月。
Um, and it was April when the Xbox 360 was announced.
呃,Xbox 360 是四月发布的。
Um, for my birthday, I didn't ask for the Xbox or I didn't ask for a birthday gift.
呃,生日的时候,我没要 Xbox,我压根没要生日礼物。
My parents asked what I wanted.
我父母问我想要什么。
I asked for it for Christmas.
我说圣诞节再要。
Uh, we celebrated Christmas, but there was no way, at least at the time, I thought there was no way they would ask would give me the Xbox 360 for Christmas and so I got it for I asked for my birthday for tab for Christmas.
呃,我们过圣诞节,但当时我觉得他们绝不可能圣诞节送我 Xbox 360,所以我生日不要、留到圣诞节再要。
Anyways, Christmas comes around, I get it.
总之,圣诞节到了,我拿到手了。
Um, you know, fast forward a couple months, my cousin who lives in Alabama, they also lived in a motel, was going to come over for spring break, um, for his spring break and we were going to just hang out at my house and he's in between me and my older brother in age, brother's a bit more jockey.
呃,快进几个月,我住在 Alabama 的表哥——他们家也住在汽车旅馆——要来过春假,呃,来过他的春假,就在我家一起玩,他年龄夹在我和我哥中间,我哥更爱运动一点。
Um, so he didn't really care too much about the Xbox.
呃,所以我哥对 Xbox 没太大兴趣。
He played sometimes, but he didn't really care.
他偶尔玩,但真不怎么在意。
Um, but my cousin, you know, I wanted to think I'm him to think I was cool, right?
呃,但我表哥嘛,我想让他觉得我很酷,对吧?
You know, so I bragged many times on the phone.
所以我在电话里吹了好多次牛。
I was like, "Yeah, I got an Xbox."
我就说"对啊,我有台 Xbox"。
And then the Xbox broke.
结果 Xbox 坏了。
There was something there's a hardware defect called the red ring of death.
有个硬件缺陷,叫"红圈死机"(red ring of death)。
Um, but long story short, I had to open it up and, you know, short the temperature sensor and it fixed it.
呃,长话短说,我得把它拆开,把温度传感器短接一下,就修好了。
Um, but I there was many other tricks I tried first and none of them worked.
呃,但我一开始试了很多别的招,全都没用。
Um, and so that's sort of how I like got into hardware.
呃,所以这算是我入坑硬件的方式。
I was like open Pandora's box.
就像打开了潘多拉的盒子。
By the time I was 12, I was like on these forums a lot reading uh posting a lot and this is around the time when Reddit ate all other forums and so I became a moderator of you know Android and Apple and Google as well as like hardware and was watch you know looking at Intel, Nvidia and AMD and all these other forums right?
到我 12 岁的时候,我就泡在各种论坛上大量看帖、呃大量发帖,那正好是 Reddit 把其它所有论坛都吞掉的时候,于是我成了 Android、Apple、Google 以及硬件这些板块的版主,一直在盯着 Intel、Nvidia、AMD 还有所有这些论坛,对吧?
I was build a PC.
我在装机(build a PC)。
All these forums I was watching, reading, posting a lot, but some of them I was moderating a lot.
所有这些论坛我都在盯着、看着、大量发帖,其中一些我还当着版主。
Um, and so, you know, smartphones, watching smartphones develop from like very simple to speed racing to being technologically more advanced than PCs, um, in many ways architecturally and and same with like, you know, all the in GPUs like just tracking and watching at that, reading every comment.
呃,所以智能手机,我眼看着它们从很简单一路飙到在技术上比 PC 还先进,呃,在很多架构层面都是,GPU 也一样,我就一直追踪、盯着这些,每条评论都读。
Um, always having the economic tinge because I grew up in a small business.
呃,而且我看什么都带着经济的味道,因为我是在小生意里长大的。
So, I was always looking at the economics, right?
所以我总是盯着经济账,对吧?
There was a time where all the like I'd say neck beards on the internet loved AMD GPUs and like I personally had bought an AMD GPU too because price performance but then when it came down to like what's technically better I'd always be like no no no Nvidia is better because they use a smaller chip to get you know better performance at better power efficiencies and their margins better and and so like I would always like talk about how Nvidia's margins were better than than AMD's in the GPU landscape and so it's like very fun.
有段时间,网上那些我称之为"油腻宅男"(neckbeards)的人都爱 AMD 的 GPU,我自己也买过一块 AMD 的 GPU,因为性价比,但一到"技术上到底谁更强"这个问题,我总会说不不不,Nvidia 更强,因为他们用更小的芯片、更好的能效跑出更好的性能,毛利率还更高,所以我老是聊 Nvidia 在 GPU 格局里的毛利率比 AMD 高,这特别好玩。
And you were 12 at the time.
那时候你才 12 岁。
I started moderating when I was 12, but this is all through my teenage tween age and high school years, right?
我 12 岁开始当版主,但这些事贯穿了我整个少年、青春期一直到高中,对吧?
Do you have any other weird hobbies or was it just semis?
你还有别的什么奇怪爱好吗,还是只有半导体?
I played a ton of Starcraft.
我打了超多《星际争霸》(Starcraft)。
At one point, I was grandmaster on the North American ladder.
有段时间,我在北美天梯上是宗师段位。
Starcraft 2.
《星际争霸 2》。
Very serious.
很硬核啊。
So, you've gotten just obsessively good at multiple things.
所以你是那种,好几件事都痴迷到玩得特别精。
Yeah.
对。
I mean, it's it's it's obsession is good.
我是说,痴迷是好事。
How were your grades?
你成绩怎么样?
Um, they were decent.
呃,还行。
Um, I would say like I had mostly A's, but they're classes that I like were thought were really boring or, you know, I just didn't enjoy.
呃,大部分是 A,但有些课我觉得特别无聊,或者就是不喜欢。
Um, like Spanish I got like not the greatest grades.
呃,比如西班牙语,我成绩就不怎么样。
Um, you know, but but it was like I speak fluent Spanish by the way, so it's really dumb.
呃,可顺便说一句,我西班牙语说得很流利,所以这事挺蠢的。
But like it's just sort of
但就是那种——
maybe that's why you didn't get a good grade.
也许这正是你成绩不好的原因。
I didn't learn Spanish till later to be fair.
说句公道话,我西班牙语是后来才学的。
But yeah, so sort of my grades were fine, right?
但对,我成绩算过得去,对吧?
Like I mean I like they were fine enough for Asian parents.
我是说,对亚洲父母来说勉强还行。
I was better than most of school but you know it wasn't like you know tryh hard maxing for like you know all A's.
我比学校里大多数人都强,但也不是那种拼了命去卷全 A 的人。
Okay.
好。
So you're very much a student of the internet then this is how you how you develop this expertise.
所以你基本上是互联网教出来的学生,你的这身本事就是这么练出来的。
At what point do you decide to start semi analysis and what's been the biggest surprise since starting the company?
你是在什么节点决定创办 SemiAnalysis 的?创办公司以来最大的意外又是什么?
Yeah so I went to school I got a few degrees in stuff that wasn't related to semiconductors.
对,我上了大学,拿了几个跟半导体毫不相关的学位。
Um was a quant for two years at a small quant risk firm.
呃,在一家小型量化风控公司当了两年 quant(量化研究员)。
Um and then basically, you know, there's a culmination of events that happened, right?
呃,然后基本上,一堆事凑到一块儿爆发了,对吧?
One was that my um you know, sort of like I got screwed out of a bonus.
一件是,呃,我的奖金被人给坑掉了。
I I'd made my company many millions of revenue of risk-free revenue because I exploited like a risk, you know, thing in the market.
我给公司挣了好几百万的无风险收入,因为我利用了市场里的一个风险漏洞。
Um you know, I think well over 10 million and they then someone else took credit for my work and all this sort of stuff.
呃,我觉得远超一千万,然后别人把我的功劳抢走了,诸如此类。
But eventually I did get rightsized.
但最后我还是被"优化"掉了。
But you know, I lost a social contract with the company I was working with.
总之,我跟这家公司之间的那份社会契约破裂了。
Um add some you know, my grand my grandparents grew up in my house with us, right? are in the motel with us.
呃,再加上,我的祖父母是跟我们一起住在家里的,对吧?跟我们一起住在汽车旅馆里。
Uh they lived with us and so you know very close with them and my grandmother got dementia and she forgot who I was and she she fell down some stairs and had like a tragic accident and passed away.
呃,他们和我们同住,所以我跟他们很亲,我奶奶得了痴呆症,不记得我是谁了,后来她从楼梯上摔下去,出了场惨剧,去世了。
So all of that happened in early 2020.
这一切都发生在 2020 年初。
Um additionally there were some like you know girl things and so you know there's a few things that happened that made me like kind of very sad.
呃,另外还有一些感情上的事,所以有好几件事凑一起,让我特别难过。
Um and and and so all of those things sort of culminated.
呃,所以这些事全堆到了一起,到了顶点。
Then COVID happened and my brother's like dude just just come stay with me.
然后 COVID 来了,我哥说,哥们儿你就来跟我住吧。
He lived in Nashville so I came and stayed with him in Nashville.
他住在 Nashville,所以我就去 Nashville 跟他一起住。
We were like, "Oh, lockdowns will be a few weeks.
我们当时想,"哦,封控也就几周吧。
You can stay with me while they happen and then you can go back home and you know, whatever."
封控期间你就住我这儿,完事你再回家,随便啦。"
Famous last words.
典型的乌鸦嘴。
Lockdowns lasted much longer.
封控持续的时间长得多。
But, you know, living with my brother for a few months, you know, was like sort of like, okay, didn't know what I was doing.
但跟我哥住了几个月,那感觉就是,好吧,不知道自己在干嘛。
I was now at my brother's home.
我现在住在我哥家里。
Um, everything was his rules.
呃,一切都按他的规矩来。
You know, sort of like, you know, him and him and his fiance at the time, now wife, you know, were like there.
就是他,还有他当时的未婚妻、现在的老婆,都住那儿。
And so, like, I basically had to tiptoe around, but I didn't care about my job.
所以我基本上得小心翼翼、蹑手蹑脚,但我已经不在乎我那份工作了。
And so, I was like posting even more than normal.
于是我发帖比平时还多。
I'd always been posting a lot on the internet.
我一直都在网上大量发帖。
I'd always been trading stocks a lot, but like I made a lot of money shorting COVID and long in COVID and like all this stuff.
我一直也炒股炒得很凶,做空 COVID、做多 COVID 这些操作,让我赚了不少钱。
Semiconductor shortages happened around then too.
半导体短缺也差不多在那时候发生了。
And anyways, I was like very much obsessed with posting and and things like that.
总之,我那时特别痴迷于发帖之类的事。
And eventually um around that time someone I got into an argument with someone on the internet and they doxed me, right?
终于,呃,大概那段时间,我在网上跟人吵了一架,对方把我人肉了,对吧?
They they publicly revealed my identity for my anonymous account.
他们公开曝光了我匿名账号背后的真实身份。
And at the time I was like, "Oh no, I scared.
当时我就想,"糟了,我怕了。
I stopped posting for like three weeks and I was like, what am I doing?
我停了差不多三周没发帖,心想,我这是在干嘛?
Why do I care?"
我干嘛要在乎这个?"
So then I just started posting under I had like had like blogs and stuff as well.
所以之后我干脆开始用真名发,我本来也有博客那些东西。
I made a real blog, SemiAnalysis, and on my 24th birthday, I posted um you know, two blogs and and then from there it just like it was not a newsletter, but I got so much traction because now instead of posting on an anonymous name, it was a real name and I put a lot more effort into those two posts than I usually did.
我做了个正经的博客,SemiAnalysis,在我 24 岁生日那天,呃,发了两篇博文,从那儿开始,它虽然还不是 newsletter,但一下就起势了,因为这回不是匿名发,而是用真名,而且这两篇我下的功夫比平常大得多。
Instead of like posting on the internet, it was like real effort into the blog.
不是随手在网上发帖,而是真正花心思写博客。
Um you can actually go back and read those if you want.
呃,你要愿意的话其实可以翻回去看那两篇。
They're they're not that great, but you know, they were they were good for the time.
写得不算多好,但在当时算不错了。
They were the best stuff you could find on the internet about semis.
那是当时网上关于半导体你能找到的最好的内容。
Um, and and I just kept posting, posting, posting.
呃,我就一直发、一直发、一直发。
I started getting a lot of consulting business.
我开始接到大量咨询业务。
You know, 2020, I also sort of I was again crashing out.
2020 年,我又一次有点崩了。
Didn't know what I wanted to do.
不知道自己想干什么。
So, I uh packed everything up or sort of I I I took my truck, I bought a tent that fits on the back of the truck, um, bought a air mattress, whatever, and would like and drove around all these national parks all around America.
所以我,呃,把所有东西打包好,开上我的皮卡,买了个能架在皮卡车斗上的帐篷,呃,买了张充气床垫之类的,然后就开着车跑遍了全美各地的国家公园。
And so, like two or three or four days of the week, I'd stay in a random motel where I negotiated the price to be like $30 a night for a room.
所以一周里有个两三四天,我会随便找家汽车旅馆住,把房价砍到差不多 $30 一晚。
And I was work on something else stuff.
同时我还在鼓捣别的一些事。
And then the weekends I'd read books and oftentimes read textbooks um while in some random national park or hiking and listen to audiobooks um about semiconductors about AI about all the things that I cared a lot about and got way more educated over these six months where I'm just like going to every national park.
然后周末我会读书,经常是读教科书,呃,一边待在某个随便挑的国家公园里或者徒步,一边听关于半导体、关于 AI、关于我特别在乎的那些东西的有声书,在这跑遍每一个国家公园的六个月里,我学到的东西多了太多。
Um and the whole time I was I was alone the whole time I was posting blogs.
呃,而且这整段时间我都是一个人,一直在发博客。
Um everyone was like DD what the fuck are you doing
呃,大家都在问,DD 你他妈到底在干嘛
pre-Starlink or the very early days of Starlink?
那是 Starlink 之前,还是 Starlink 最早期那会儿?
Pre-Starlink, pre-Starlink.
Starlink 之前,Starlink 之前。
Um yeah, so it was like very much like what are you doing?
呃,对,所以那感觉就是,你到底在干嘛?
Um I travel around LatAm again like for for a year initially with my friend and then with my ex you know for you know about a year and then I go then 22 23 24 end of 21 22 23 and 24 I'm completely I'm still completely homeless since mid-2020 right um but I'm traveling around to every conference in the world.
呃,我又跑遍了 LatAm(拉美),一开始跟我朋友跑了大概一年,然后跟我前任,又大概一年,再往后是 22、23、24 年,从 21 年底到 22、23、24 年,我彻底——从 2020 年年中起我到现在都还是彻底居无定所的状态,对吧,呃,但我在满世界跑遍每一场大会。
I go to 40 plus conferences a year no matter where in the supply chain it is.
我一年参加 40 多场会,不管它在供应链的哪个环节。
I'm like oh that looks interesting.
我就想,哦这个看着挺有意思。
I guess I'll go to that.
那我就去这个吧。
And I'm like I went to one conference like wow this is amazing.
然后我去了一场会,心想,哇这太棒了。
you get to talk to the experts and they just like they they're they they're going to talk to you because and then you're so excited and in the case of semiconductors everyone's a boomer so it's like it's great to like you know they're like they don't see young people who are like excited about it so they're really happy to tell stuff and so you just have to ask on this
你能跟专家聊上,他们也愿意跟你聊,你又那么兴奋,而半导体这行大家都是老一辈,所以特别妙,他们平时见不到对这行这么兴奋的年轻人,所以特别乐意跟你讲东西,你只要开口问就行。
was there like a part of the supply chain or one of these conferences that you know particularly changed your view of the semi-world or that you felt then or feel now is particularly underrated.
有没有哪个供应链环节、或者哪场这样的会,特别改变了你对半导体世界的看法,或者你当时觉得、现在依然觉得被严重低估了?
I think I think the trade shows like R&D conferences range really widely.
我觉得,这些展会、还有研发类的会,差别特别大。
Um obviously some of the you know the ones I have the most fun at you know include NeurIPS.
呃,显然我玩得最开心的其中之一是 NeurIPS。
Why why is that?
为什么呢?
Because it's 20,000 AI researchers and they're generally in my distribution of age range.
因为那儿有两万名 AI 研究员,而且他们基本都在我这个年龄段。
So it's like a lot of fun but they're also like leading AI researchers and it's a lot of fun and lot you learn a lot.
所以特别好玩,而且他们还是顶尖的 AI 研究员,又好玩又能学到很多。
Um there's also a lot of parties and then it ranges all the way to like you know there's random chemical conference in Japan where it's 300 Japanese dudes.
呃,派对也多,然后另一端能一直延伸到——日本有个不知名的化学品大会,现场是 300 个日本大叔。
It's like 20 guys from ASML, 20 guys from TSMC, 20 guys from Intel, and those are the only people who speak English.
大概 20 个 ASML 的人、20 个 TSMC 的人、20 个 Intel 的人,而且只有这些人会说英语。
Uh, everyone else speaks only Japanese, and you're like, h, I guess they're still pretty interesting and fun.
呃,其他人只说日语,你就想,嗯,我猜这些会照样挺有意思、挺好玩的。
I think I think like one thing that I have like a skill set of is like I'm able to bond with anyone regardless of their background and like who they are.
我觉得我的一项本事是,不管对方什么背景、是什么样的人,我都能跟他们拉近关系。
I'm able to talk to them, find something interesting to talk about.
我能跟他们聊起来,找到有意思的话题。
Oftentimes, it's the tech stuff, but, you know, it's it's and so I think like the most interesting conferences are oftentimes like, you know, the really big ones because that's where the biggest stuff is happening.
通常是技术话题,不过,所以我觉得最有意思的会往往是那些超大型的,因为最大的事都在那儿发生。
Um but I think the niches that are really really exciting is like you know SPIE um so there's IE which is international electrical engineering something um and there's SPIE which is another ecosystem.
呃,但我觉得真正特别让人兴奋的小众领域,比如 SPIE,呃,有个 IE,就是国际电气工程什么的,呃,还有个 SPIE,那是另一个生态圈。
SPIE conferences are super super deep in details.
SPIE 的会在细节上钻得特别特别深。
Every single one that I went to, especially like SPIE advanced lithography or SPIE photomask, I went to them the first time I didn't even understand 90% of what I heard.
我去的每一场,尤其是 SPIE 先进光刻或者 SPIE 光罩,第一次去,我听到的东西 90% 都听不懂。
And then I read, read, read, I had made some context, of course, and then next time I went, I understood like half of what I went to.
然后我读、读、读,当然攒了点背景知识,等下次再去,我大概能听懂一半。
Third time I went, I understood like 75% of what I went to.
第三次去,我大概能听懂 75%。
Even now, I went and I was like, I still don't understand everything that's going on.
即便到现在,我去了还是会觉得,我依然没搞懂里头所有的事。
Whereas like you go to like NeurIPS, you know, a couple times you can understand, okay, what's neurosymbolic reasoning?
相比之下,像 NeurIPS 你去个几次,就能搞明白,好,什么是神经符号推理(neurosymbolic reasoning)?
Okay, what's this?
好,这是什么?
What's that?
那又是什么?
like you can you can kind of get a mapping of what everything is pretty quickly but some parts of the supply chain are so arcane and so deep and so technical.
你能挺快地把各种东西大致梳理出一张图,但供应链里有些环节实在太晦涩、太深、太技术了。
It takes a lot of times for you to even understand what's happening in and you know on everything right um for every research paper doesn't necessarily mean you didn't you know you go to a conference for a few reasons right you understand the research you understand but like it's all the research that's being published but what you really care about is understanding how does that research intersect with technology also how does that research differ from what's there today and none of these research papers tell you what's happening today but then you just ask people and you you build contacts and you learn and then you like learn about the supply chain and oh this company supplies this company even though it's not publicly stated anywhere or like you know you learn that the the this chemical is like cost about this much and a tool uses about this much and you
你得花很多次,才能弄明白里头到底在发生什么,对吧,呃,你去参加一场会是出于几个原因,对吧,你去理解研究,你理解的是所有已经发表出来的研究,但你真正在乎的是搞清楚这些研究怎么跟技术交叉,还有这些研究跟今天现有的东西有什么不同,而这些论文没一篇会告诉你今天正在发生什么,于是你就直接去问人,建立人脉,学习,然后你会摸清供应链,哦原来这家公司供货给那家公司,哪怕这事在任何地方都没公开写出来,或者你会打听到,这种化学品大概值多少钱、一台设备大概要用掉多少,而你——
hear you hear the horror stories of like this chemical had a shortage and it totally threw off this part of the supply chain and then it turns out there's only three companies in the world that make that chemical and it's like
你会听到那些吓人的故事,比如某种化学品闹短缺,把供应链的这一整块彻底搅乱了,然后发现全世界只有三家公司生产这种化学品,那种感觉就是——
my favorite one is I learned uh a Japanese guy at that specific Japanese uh conference that I went to where no almost no one spoke English in very broken English he told me about how uh his father worked in this in in in this industry in the 1980s that the the only factory in the world that built this chemical uh burned down and that caused memory prices to like double or triple and I was like wow not too different from today
我最喜欢的一个是,呃,在我去的那场几乎没人说英语的日本大会上,有个日本人用磕磕巴巴的英语跟我讲,呃,他父亲在 1980 年代干过这行,当时全世界唯一一家生产这种化学品的工厂,呃,烧毁了,结果导致存储器价格翻了一到两倍,我当时就想,哇,跟今天没差多少。
not not not at all
一点、一点都没变
crazy
太疯狂了
um
呃
inference going to be the biggest market on earth biggest market beyond earth agree or disagree
推理会成为地球上最大的市场、也是地球之外最大的市场——同意还是不同意?
um I mean obviously use of tokens is going to be the biggest market
呃,我是说,很明显,token 的使用会成为最大的市场
um and the value that's created from tokens is going to be the biggest market
呃,而 token 创造出来的价值也会是最大的市场
but I think tokenomics sort of the use of tokens adoption of AI sort of is the most important thing that's happening
但我觉得 tokenomics,某种意义上说 token 的使用、AI 的采用,才是当下正在发生的最重要的事
and inference whether it's open models or closed models will be like one of the biggest markets in the world much bigger than oil
而推理,不管是开源模型还是闭源模型,都会成为全世界最大的市场之一,比石油大得多
I think much bigger than like you know many other parts like inference of AI will be you know many percentage points of the GDP yeah right
我觉得比很多其他领域大得多,推理这块,AI 推理会占到 GDP 的好几个百分点,对吧
what you've done with inference X I think is you know industry standard
你用 InferenceX 做的这套东西,我觉得已经是行业标准了
maybe say a word on why you started it what it does
要不讲讲你当初为什么做它、它到底是干嘛的
and you know what do people misunderstand about uh performance benchmarking on inference
还有,大家在推理的性能基准测试上,呃,通常会误解什么
yeah So, so to zoom back, right, like semi analysis, uh, we do a lot of stuff that's like, you know, a lot of it is like research for institutional clients and and our subscription versus products, but a lot of it is also like, hey, you know, this would just be cool to figure out.
对,那,往回倒一下哈,就说 SemiAnalysis,呃,我们做很多事情,你懂的,其中很大一部分是给机构客户做研究、还有我们的订阅和产品,但还有很大一部分是那种,嘿,你懂的,这东西搞明白了会挺酷的。
Let's figure out how to figure it out and just post it publicly.
那我们就琢磨怎么把它搞清楚,然后直接公开发出来。
And that gets, you know, more and more scale.
然后它就会,你懂的,规模越铺越大。
And so we've done this with a lot of GPU benchmarking and testing and training performance and inference performance, but you know, ultimately we saw like inference benchmarking was like point in time.
我们在很多 GPU 基准测试、测试、训练性能、推理性能上都是这么干的,但你懂的,说到底我们发现,推理基准测试是个"时间点快照"式的东西。
you know, you test it and you take some time, you release it and it's like slow and arcane and out outdated because models change all the time.
你懂的,你测一遍,花上一段时间,发出来,结果又慢、又晦涩、又、又过时,因为模型一直在变。
Every I feel like every week there's a new model whether it's a Chinese model or you know today mythos 5, Fable dropped and new models are coming out all the time.
我感觉每、每周都有新模型,不管是中国的模型,还是像今天 Mythos 5、Fable 就发布了,新模型一直在往外冒。
Um on the software layer uh PyTorch, VLM, SG lang um new drivers, new new something drops, you know, in fact the update cycle for most of these libraries is twice a week.
呃,在软件层,呃,PyTorch、vLLM、SGLang,呃,新驱动、新的、新的什么东西一直在发,你懂的,其实这些库大多数的更新周期是一周两次。
So you basically have the software updating all the time and therefore performance changing.
所以软件基本上一直在更新,性能也就一直在变。
Um you know new inference optimizations are coming out and those get updated and and so I feel like it's a relentless breakthrough after breakthrough after breakthrough that keeps driving efficiency and cost down which is why we've seen you know model cost drop for equivalent quality by like 60x a year.
呃,你懂的,新的推理优化不断出来,然后被更新进去,所以我觉得这就是一场不停歇的突破接着突破接着突破,不断把效率往上推、把成本往下压,这也是为什么我们看到,你懂的,同等质量下模型成本一年降大约 60x。
It's incredible.
太夸张了。
Um but to stay on top of that you can't have point in time benchmarking.
呃,但要跟上这个节奏,你就不能用时间点快照式的基准测试。
You need to have benchmarks be living and breathing i.e. you know constantly running on the latest hardware on the latest models.
你得让基准测试是活的、会呼吸的,也就是说,你懂的,在最新的硬件、最新的模型上持续不断地跑。
And so we embarked on a project and we got a lot of buyin from the ecosystem.
于是我们启动了一个项目,拿到了整个生态圈的很多支持。
This was only possible because we had you know enough aura with some of the ecosystem where we're able to get coreweave and cruso and nebus and Oracle and Microsoft and Amazon and Google and OpenAI to contribute to us um compute
这之所以能成,是因为我们,你懂的,在生态圈里攒了足够的"气场",才能让 CoreWeave、Crusoe、Nebius、Oracle、Microsoft、Amazon、Google 还有 OpenAI 给我们贡献,呃,算力
and then we were able to work with SG Lang and VLM and now Radix Arc and InRact uh which are the private companies who are sort of leading those efforts um the open source efforts um to collaborate with us.
然后我们又能跟 SGLang、vLLM,现在还有 Radix Arc 和 InRact 合作,这些是牵头这些工作、呃,牵头那些开源工作的私营公司,呃,跟我们一起合作。
We're able to get Nvidia and AMD and Google and Amazon now because we're adding TPUs and Trainium uh to collaborate.
我们现在还能拉到 Nvidia、AMD、Google 和 Amazon 一起合作,因为我们要加上 TPU 和 Trainium,呃,一起搞。
Now we've got all these people collaborating.
现在这么一大堆人都在跟我们合作。
We've got over $50 million of hardware uh donated to us.
我们拿到了超过 $50M 的硬件,呃,是捐给我们的。
Um once we launch TPUs and trainum it actually should be over $100 million of hardware.
呃,等我们把 TPU 和 Trainium 上线,这个数其实应该会超过 $100M 的硬件。
Um you know maybe about like 15 different chip types all running these benchmarks every single day on all the latest model, right?
呃,你懂的,差不多有 15 种不同的芯片,每一天都在所有最新模型上跑这些基准,对吧?
the best model from Moonshot, the best model from Alibaba, the best model from um there's about five different Chinese models, the best open source models, the best Chinese labs there.
Moonshot 最好的模型、Alibaba 最好的模型、呃还有,大概有五个不同的中国模型、最好的开源模型、那边最好的中国实验室。
We run benchmarks on their models every day and then also the best US open source models um GPT-OSS, Nemotron, etc.
我们每天在他们的模型上跑基准,然后还有美国最好的开源模型,呃,GPT-OSS、Nemotron 等等。
So we're running these benchmarks every day um in an automated fashion and they run on these these servers that are dedicated to us for inference benchmarking and we sweep across so many different configurations and optimization types
所以我们每天跑这些基准,呃,是自动化跑的,它们跑在这些、这些专门给我们做推理基准的服务器上,我们扫遍了特别多不同的配置和优化类型
and then what it creates is and all the results are public and all the configurations are public.
然后它产出的是——所有结果都是公开的,所有配置也都是公开的。
So now we have the pareto optimal curve because a lot of you know times when people are comparing inference performance they're like taking a suboptimal curve or point for someone else and comparing it to their optimal one.
所以现在我们有了帕累托最优曲线,因为很多时候,你懂的,人们在比推理性能的时候,是拿别人一条次优的曲线或次优的点,去跟自己最优的比。
And it's like, well, yeah, I can make I can I can stick, you know, if I drove a Porsche versus like some some race car driver, obviously I'd drive it slower.
那当然,就好比,你懂的,如果我开一辆保时捷,去跟某个、某个赛车手比,那我肯定开得更慢。
The same thing with inference benchmarking.
推理基准测试也是一个道理。
And so what we did is we created open- source uh basically containers for the optimal points across every uh point on the interactivity, i.e. how fast is it responding to me versus you know batch size, i.e. how many users am I simultaneously serving curve?
所以我们做的是,开源了,呃,基本上就是一个个容器,对应交互性曲线上每一个、呃,每一个点上的最优点,交互性就是它给我响应有多快,而另一头是批大小,也就是你懂的,我同时在服务多少用户,这么一条曲线。
And so now anyone who wants the optimal point can just go to inference X download it and run that as the optimal point
所以现在任何人想要那个最优点,直接去 InferenceX 把它下载下来,当成最优点来跑就行
and they can check every day if they want or they can even autod download the most optimal point for that model and and their inference performance will be near peak.
他们愿意的话可以每天查一次,甚至可以自动下载那个模型最优的点,那他们的、他们的推理性能就会接近峰值。
Um
呃
is that curve like the most important curve in your opinion?
在你看来,那条曲线算是最重要的一条曲线吗?
The throughput interactivity curve is the most important one.
吞吐-交互性这条曲线是最重要的。
Yeah, I think I think um most things in hardware infrastructure uh model application layer everything is downstream of that curve, right?
对,我觉得、我觉得,呃,硬件基础设施里的大部分东西、呃,模型应用层,一切都是那条曲线的下游,对吧?
Is it is it something that needs to be super super fast, super low latency?
它是不是、是不是那种需要超级超级快、超低延迟的东西?
Um, and I don't really care about the cost, so I make batch size very low and I use techniques like speculative decoding or multi-token prediction heavily and and there's so many, you know, possible techniques there.
呃,那我根本不在乎成本,所以我把批大小压得很低,大量用像投机解码或者 multi-token prediction 这样的技术,而且,你懂的,那里面可用的技术太多了。
Or is it something where actually I'm batch processing a ton of documents and I don't really care about all these things.
还是说,它其实是我在批量处理一大堆文档,那这些东西我根本不在乎。
I don't use these techniques that actually are worse on cost efficiency but help you with speed for an individual user because I just want to pack a bunch of users.
我不会用那些其实在成本效率上更差、但能帮单个用户提速的技术,因为我只想把一大批用户塞进去。
I don't care if the document takes all night to process, right?
我才不在乎这文档是不是要处理一整晚,对吧?
Um, and right now the way we treat AI infrastructures, it's like one-sizefits-all.
呃,现在我们对待 AI 基础设施的方式,是那种一刀切、一个尺寸通吃。
But over time, we're going to get to the point where, you know, there's stuff where you you have batch workloads or, you know, you need instant response and there's there's the whole curve that's going to matter for uh users.
但随着时间推移,我们会走到那么一步:你懂的,有些是批处理的工作负载,或者,你懂的,有些是你需要即时响应,整条、整条曲线都会对,呃,用户变得重要。
And so we see this with entropic, right?
我们在 Anthropic 身上就看到这个了,对吧?
Cloud code fast mode cost way more than regular mode.
Claude Code 的 fast mode 比普通模式贵得多。
Um, and same with open eyes priority Q thing.
呃,OpenAI 那个优先队列(priority queue)也是一个道理。
Um,
呃,
sorry, dumb question.
抱歉,一个蠢问题。
How does cost factor into the chart?
成本是怎么体现进这张图里的?
So if if I let's say imaginary example, I have 100 I have a batch size of 100,
比如说、比如说,举个假想的例子,我有 100,我有一个 100 的批大小,
okay?
好?
And I can do 10 tokens per second per user.
然后我每个用户能做到每秒 10 个 token。
So in total I'm doing a thousand tokens per second uh off of that one piece of compute.
所以总共,呃,我用那一块算力在做每秒一千个 token。
That's one side of the curve.
这是曲线的一端。
Super slow, 10 tokens per second.
超级慢,每秒 10 个 token。
Um you know, other side is I have uh uh 500 tokens per second, but I only have one user.
呃,你懂的,另一端是我有,呃、呃,每秒 500 个 token,但我只有一个用户。
And so maybe 250 tokens per second, one user.
所以也许是每秒 250 个 token,一个用户。
And then there's points on the middle that are more fraal optimal, right?
然后中间还有一些点,是更接近最优的(fraal optimal),对吧?
the average person actually wants like 50 or 100 tokens a second and maybe you know the this the the number of users I can batch together.
普通人其实想要每秒差不多 50 或者 100 个 token,然后也许,你懂的,这个、这个、这个我能一起打包的用户数量。
So the curve is okay a thousand tokens total uh per second or 250 tokens total per second depending on how many users I batch and there's a curve in the middle
所以这条曲线就是:好,总共每秒一千个 token,呃,或者总共每秒 250 个 token,取决于我打包多少用户,中间还有一条曲线
and so ultimately some workloads will actually want the 4x cost decrease because the same unit of hardware can do a thousand versus 250
所以最终有些工作负载其实会想要那个 4x 的成本下降,因为同一块硬件能做一千对比 250
and some users I'll pay 4x more because I don't care about the price I care about time because the person using the tokens is expensive or the feedback loop that I have here is expense is is expensive.
而有些用户,我愿意多付 4x,因为我不在乎价格,我在乎时间,因为用这些 token 的人很贵,或者我这里的反馈循环很贵、很、很贵。
If you had to guess, you choose the time frame 10 years or 15 years.
如果让你猜,时间跨度你自己选,10 年还是 15 年。
What percent of inference compute do you think will happen in space?
你觉得会有百分之多少的推理算力,发生在太空?
Can be 0% 50%
可以是 0%,50%。
Sean 99% like
Sean 觉得 99% 之类的。
this is a tough one.
这个问题不好答。
Um you choose the time frame like 10 whatever time frame and you're
呃,你挑一个时间跨度,比如 10 年,随便什么跨度,然后你就……
so I think I think the non-consensus or at least against SpaceX thing, you know, I love SpaceX by the way and I totally would buy the IPO if I could buy stocks.
所以我觉得,我觉得那个非共识的、或者至少是唱反 SpaceX 的观点——顺便说,我超爱 SpaceX,要是能买股票,IPO 我肯定砸进去。
not investment.
不构成投资……
Not investment advice. Thank you. Thank you.
不构成投资建议。谢谢,谢谢。
not invested ice um from either um I don't think that space data centers will really matter in the next um you know 3 to 5 years
not invested ice(不投冰),呃,反正也没投,呃,我不觉得太空数据中心在接下来的、呃、你懂的 3 到 5 年里会真的有多大意义。
um with that said I think in you know 20 years I think the vast majority of compute will be going in space
呃,话虽如此,我觉得 20 年后,绝大多数算力都会跑到太空里去。
um and so the real real factor there is sort of you know what's the cost it's the time frame it's the cost of building power on terrestrial land and how much power you going to be able to do on terrestrial land
呃,所以这里真正真正的关键因素,某种程度上就是,你懂的,成本是多少、时间跨度是多少、在地面上建电力的成本是多少,以及你在地面上到底能搞出多少电力。
and I think obviously my views of where inference you know you know how many gigawatts or terawatts are devoted to inference is it's a crazy curve for me personally
而且我觉得,很明显,我对推理走向的看法——你懂的,有多少 gigawatt 或者 terawatt 投在推理上——在我个人看来,那是一条疯狂的曲线。
what's your forecast how many gigawatts or
你的预测是多少?多少 gigawatt,还是……
um yeah I think I think by you know 2030 just OpenAI, Anthropic we'll have over 100 gigawatts combined
呃,是的,我觉得,我觉得到 2030 年,光是 OpenAI、Anthropic 两家加起来就会超过 100 gigawatt。
um and then you'll add you know meta and Google and you know so on and so on so forth
呃,然后你再把 Meta、Google 之类的都加上,等等等等。
it's it's a humongous amount of compute that will be dedicated to inference
这是一个,这是一个巨量的算力,全都会投到推理上。
um and by like 2040 it'll be terawatts right
呃,到 2040 年左右,就会是 terawatt 级别了,对吧。
um the the curve of like productivity that we're going to get and so you know inference deployments is going to be huge
呃,我们将获得的那条生产力曲线,所以你懂的,推理部署会大到吓人。
And so if you look at like 2040, I think like you know probably more than half of the incremental compute will be going in space.
所以你要是看 2040 年,我觉得,你懂的,大概超过一半的增量算力都会跑到太空里去。
But if you look at 2030, I think it's sub 1%.
但你要是看 2030 年,我觉得这个比例还不到 1%。
Do you think intelligence per watt has been increasing?
你觉得每瓦特智能(intelligence per watt)一直在提升吗?
Uh and then it seems like there's still a giant gap between where we are intelligence per watt versus like human biology and so like if we are do you think we are to close that gap?
呃,而且看起来我们现在的每瓦特智能和人类生物体之间还有一道巨大的鸿沟,那么,如果我们……你觉得我们能不能追平这道鸿沟?
And if so, where is that game going to come from?
如果能,那这个突破会从哪里来?
Yeah, I think I think it often depends on what you're doing too, right?
对,我觉得,我觉得这也往往取决于你在干什么,对吧?
Like a TI84 is way more intelligence per watt in terms of doing math than us and it's like 30 years old, right?
比如一台 TI84 计算器,论算数学,它的每瓦特智能可比我们高多了,而且它都 30 年老古董了,对吧?
Obviously this is like a dumb dumb you know sort of
当然,这是那种很蠢很蠢的,你懂的,某种意义上的……
general intelligence.
通用智能。
Yeah. But general intelligence wise um so one of the things inference X does is we also measure the power and cost of all of these this hardware.
对。但就通用智能而言,呃,所以 InferenceX 做的其中一件事,就是我们还会测量所有这些硬件的功耗和成本。
And so we offer not just you know throughput versus interactivity we offer cost versus interactivity.
所以我们提供的不只是,你懂的,吞吐量对交互性,我们还提供成本对交互性。
We offer power versus interactivity.
我们还提供功耗对交互性。
And so as far as has you know intelligence per watt been increasing?
所以说到,你懂的,每瓦特智能有没有在提升?
Um I mentioned you know it's been a 60x cost decrease for same benchmark level.
呃,我之前提过,你懂的,在同一基准水平上,成本已经降了 60x。
Um we've also seen the same on on uh intelligence per watt.
呃,我们在,呃,每瓦特智能上也看到了同样的情况。
Um it's not been it's not been exactly 60x. It's been closer to like 40x.
呃,不是,不是正好 60x,更接近 40x 左右。
Uh some of the efficiencies are nonpower ways, but there's been a humongous improvement in in intelligence per watt on an annual basis at least so far this year, last year, year before, year before.
呃,有些效率提升是走非功耗的路子,但每瓦特智能按年来看已经有了巨大的进步,至少到目前为止,今年、去年、前年、大前年都是如此。
And I expect that to continue as far as where we are from the human brain.
而且我预计这会持续下去。至于我们离人脑还有多远——
We're we're many orders of magnitude away.
我们还差好几个数量级。
Thankfully, doesn't really matter.
好在,这其实没太大关系。
We can devote a lot of power to computers.
我们可以往计算机上砸大量电力。
Much easier to power computers than human brains.
给计算机供电比给人脑供能容易多了。
Like you know, we have sickness, disease, and like food preferences. sleep.
你懂的,人有生病、有疾病,还有,比如挑食的毛病,还要睡觉。
Uh yeah, exactly.
呃,对,正是如此。
Let me just ask one more question on the like on the general theme in my opinion in terms of like you know intelligence per watt or intelligence per per dollar like any any of these metrics.
让我就这个大主题再问一个问题,在我看来,就是那种,你知道的,每瓦智能、或者每美元智能,随便哪个指标都行。
I think there's kind of three levels of input.
我觉得输入大概分三个层次。
You can get hardware improvements that are where the hardware is more efficient.
你可以靠硬件进步,也就是硬件本身变得更高效。
You can get lowlevel systems optimizations like kernel level you know improvements matrix multiplication libraries you know things like that or you can get like highlevel like model level algorithmic improvements you know at the highest level.
你可以靠底层的系统优化,比如 kernel 层的改进、矩阵乘(matrix multiply)库之类的东西,或者你可以靠高层的、模型层的算法改进,你知道的,也就是最高那一层。
It see like to me it seems like in the last three years most of the gains have come from hardware level and you know and some from the model level like do you think that that is what do you agree with that do you think that's what look like in the future do like do you think there's a bunch of juice to squeeze in a say like kernel level like
在我看来,过去三年大部分收益来自硬件层,还有,你知道的,一些来自模型层——你怎么看,你同意吗?你觉得未来也会是这样吗?你觉得在比如 kernel 层还有很多油水可以榨吗?
yeah Sean I completely disagree with you by the way great great that's why I'm asking this question
是啊 Sean,顺便说一句,我完全不同意你的看法——很好很好,所以我才问这个问题。
um okay so I I think you know one way is to look at as these three different layers
呃,好,我我觉得,你知道的,一种看法就是把它拆成这三个不同的层。
Um and in that sense like okay from hopper to blackwell which is all we've had over the last three years roughly 30x improvement on DeepSeek on the most optimized deployment which is you know you can see on inference there's about a 30x improvement
呃,从这个角度看,好,从 Hopper 到 Blackwell——这就是过去三年我们经历的全部——在 DeepSeek 上、在最优化的部署方案上,大概有 30x 的提升,你知道的,在推理这块你能看到差不多 30x 的提升。
but you know over the last three years um we've had way more improvement intelligence per watt a lot of that coming from the model layer right
但你知道,过去三年,呃,我们在每瓦智能上的提升要大得多,其中很多来自模型层,对吧。
if you look back three years it's GPT-4 now it's like you know you know maybe like Qwen one of the smaller Qwen models that's like you know 27B parameters total and like 2 billion active is like way better.
如果你回看三年前,那时是 GPT-4,现在呢,你知道的,你知道的,可能是 Qwen,某个更小的 Qwen 模型,总参数大概 27B、激活参数(active parameters)大概 20 亿,却要强得多。
Um, and so you've got this huge improvement on model layer, you've got this pretty sizable improvement on hardware, but it's that co-design layer and I think that's that's what's important, right?
呃,所以你在模型层有巨大的提升,在硬件层有相当可观的提升,但真正关键的是那个 co-design(协同设计)层,我觉得那才是重点,对吧?
If you look at the architecture of, you know, any of these models, but deepseek is the most famous one at least, uh, that's public and people have seen.
如果你去看这些模型里任何一个的架构,你知道的——不过至少 DeepSeek 是最出名的那个,呃,公开的、大家都见过的。
Yeah, Deepseek got huge efficiency gains from like co-op optimization or kernel level optimizing memory.
对,DeepSeek 靠协同优化、或者说 kernel 层对内存的优化,拿到了巨大的效率提升。
Yes, I I think it's it's it's like kernels of course, but it's actually you build the hardware architecture for the chip.
对,我我觉得,当然有 kernel 的成分,但其实是你围绕芯片的硬件架构来搭建(模型)。
So if you look at the shapes of all the experts in in DeepSeek, uh V3, they were all optimized for Hopper.
所以你去看 DeepSeek 呃 V3 里所有专家(expert)的形状(shape),它们全都是针对 Hopper 优化过的。
And if you look at for V4, they're optimized for Blackwell and Huawei's chip.
再看 V4,它们是针对 Blackwell 和华为的芯片优化的。
And what's interesting is despite the fact that TPUs are objectively an amazing chip, you know, and and they run all of deep mind and they do all the training uh for anthropic as well on the pre-training side at least.
有意思的是,尽管 TPU 客观上是一款了不起的芯片,你知道的,DeepMind 全靠它跑,而且 Anthropic 的训练——至少 pre-training 这块——也全在上面做。
TPUs suck at running DeepSeek, but they are really really great at running other kinds of models that don't run well on NVIDIA.
但 TPU 跑 DeepSeek 就很烂,可它们跑另一些在 NVIDIA 上表现不好的模型却非常非常强。
there is some level of such deep optimization that has been done um whether it be shapes uh network IO uh patterns you know how you do the collectives how you do um things around you know the the arithmetic intensity of the attention mechanism all these different things are co-optimized between the model and the and the and the hardware and the infrasoftware in between
这里做了某种极深的优化,呃,不管是形状(shape)、呃 网络 IO 呃 的模式、你知道的怎么做集合通信(collective)、怎么处理呃 你知道的注意力机制(attention mechanism)的算术强度(arithmetic intensity)之类的——所有这些不同的东西,都在模型、硬件、以及夹在中间的基础软件之间做了协同优化。
and it's it's hard to say you can disentangle the gains
所以很难说你能把这些收益拆解开来。
do you think that like my understanding is that like China has done this a lot better than the west the last few years like in the DeepSeek was one of the first models to really like do this.
你觉得——我的理解是,过去几年中国在这件事上做得比西方好得多,比如 DeepSeek 就是最早真正这么干的模型之一。
I don't necessarily think so.
我不一定这么认为。
think it's more so that the west doesn't tell people what they do right like open AI didn't tell people that you know GPT-4o was uh how sparse it was what the shape size was all these things
我觉得更多是因为西方不告诉别人他们干了啥,对吧,比如 OpenAI 就没告诉大家,你知道的,GPT-4o 有多稀疏(sparse)、shape 大小是多少这些事。
but GPT-4o is roughly the same size slightly smaller than DeepSeek v3 and 4o came out you know a little bit earlier right if I recall correctly
但 GPT-4o 的规模差不多、比 DeepSeek V3 还略小一点,而且 4o 出得,你知道的,还早一点,对吧,如果我没记错的话。
so is your is your view that like all three of these things have been happening simultaneously at like roughly the same rate and the most the biggest gains are when you just co-optimize
所以你的、你的看法是不是:这三件事一直在同时发生、速度大致相当,而最大的收益出现在你把它们协同优化的时候?
I would say I would say there's been more gains on the model layer than on that co-op than than on the sort of software infrastructure layer and the hardware layer.
我会说,我会说,模型层的收益比那个协同(co-op)层、比软件基础设施层和硬件层都要多。
Um, but there's been innovations on every layer and and and really the biggest gain and the beauty of the best labs is when they co-optimize all three, you know,
呃,但每一层都有创新,而真正最大的收益、最顶尖实验室的高明之处,就在于他们把三层一起协同优化,你知道的。
and and and that's what like you know when Anthropic is is, you know, even though they used many different kinds of hardware, they don't really inference too much on TPUs.
就像你知道的,Anthropic,呃,你知道的,虽然他们用了很多种不同的硬件,但他们其实不怎么在 TPU 上做推理。
They mostly train on TPUs.
他们主要在 TPU 上做训练。
um and and they inference a lot on Trainium and GPUs and GPU is more a jack of all trades but they've optimized their hardware they're optimized their model they've optimized everything so they can do that
呃,他们大量在 Trainium 和 GPU 上做推理,而 GPU 更像是个万金油,但他们把硬件优化了、模型优化了、什么都优化了,所以才做得到。
whereas open AI they're you know prior models were optimized for hopper more now they're more optimized for blackwell
而 OpenAI 呢,你知道的,他们早先的模型更多是针对 Hopper 优化的,现在则更多针对 Blackwell 优化。
and you know you you step forward through time these these um these labs and and and the same with Google right
而且你知道,你顺着时间往前走,这些呃 这些实验室——Google 也一样,对吧。
they've they've they've optimized you know Gemini 2 was really optimized for the TPU uh v uh v6e or tp Gemini 3 was
他们优化了,你知道的,Gemini 2 是真正针对 TPU 呃 v 呃 v6e 优化的,Gemini 3 也是。
and then Gemini uh you know the next Gemini that's coming out is really optimized for TPU v7.
然后 Gemini 呃 你知道的,接下来要出的那个 Gemini 是真正针对 TPU v7 优化的。
Um and so sort of like a lot of these things are being co-optimized and actually when you pull that model and put it run it on the old hardware it's really not that great.
呃,所以差不多这么说,很多东西都在被协同优化,而实际上,当你把那个模型拿出来、放到老硬件上跑,它其实就没那么厉害了。
Um and so I think a lot of this co-optimization is is the most important thing.
呃,所以我觉得这种协同优化很大程度上才是最重要的东西。
It's called software hardware co-design and that's what's like really exciting about like you know sort of what what you know I think my day-to-day is like you know great you get to look at one layer there's all these innovations happening here there's all these innovations happening on every layer.
这叫软硬件协同设计(co-design),这也是,你知道的,真正让人兴奋的地方,差不多就是,我觉得我的日常就是,你知道的,很棒,你能盯着某一层,这里在发生各种创新,每一层都在发生各种创新。
The real breakthrough innovation is when you leapfrog a few layers, you co-optimize and co-design them, and now all of a sudden you've you've taken what could have been a 2x here, 2x here, 2x here, and instead of being multiplicative to 8x, it's actually 100x because you've optimized across all three layers.
真正突破性的创新,是当你跨越好几层、把它们协同优化、协同设计,于是突然之间,你把本来这里 2x、那里 2x、再一个 2x——本来相乘是 8x——变成了实际上的 100x,因为你在三层上都做了优化。
And so that's what's really exciting about sort of like what you see at the labs, which you see at like a company like Nvidia who's not co-optimizing on the model layer per se, but a little bit from the model layer all the way downstream to, you know, silicon.
所以这就是真正让人兴奋的,差不多就是你在这些实验室里看到的,你也在像 Nvidia 这样的公司里看到——它本身并不在模型层做协同优化,但从模型层一路往下、一直到,你知道的,硅片,都沾一点边。
Or you look at a company like TSMC, they're co-optimizing not just, you know, fabrication, but all the way from the components and the consumables and the tools all the way upstream to what the designs, their chips are, the customers are telling them is this co-optimization across many layers of the abstraction stack.
或者你看像 TSMC 这样的公司,他们协同优化的不只是,你知道的,制造工艺,而是从元器件、耗材、工具一路往上游、一直到那些设计、他们的芯片长什么样、客户告诉他们要什么——这是一种跨越抽象栈很多层的协同优化。
There will always be bottlenecks somewhere in that optimization though that are like lagging behind and then need to get pulled forward,
不过在那种优化里,总会有某个地方存在瓶颈,就是那种拖在后头、然后需要被拉上来的部分,
you know,
你知道的,
and band-aids to toact.
还有那种创可贴式的临时补丁。
If you had to predict like what are at any level of the stack, it can be literally anywhere.
如果让你预测,在这个栈的任何一层——它可以真的在任何地方——
What are some of the bottlenecks you're most like you're kind of tracking most acutely the next year?
未来一年,你最、你算是最紧盯着的一些瓶颈是哪些?
And not necessarily in the supply chain, not in like scale, but in terms of the actual um and it can it can be in the supply chain too, but just like you know, is it memory improvements?
不一定是供应链方面的,也不是那种规模扩张,而是就实际的呃——它也可以是供应链方面的,但就,你知道的,是内存的改进吗?
Is it is it that like just like scaling?
还是说就是那种单纯的规模扩张(scaling)?
So memory memory is memory is an easy one that everyone's talked about, but I'm not going to talk about from a supply chain angle.
那内存内存是,内存是个容易的例子,大家都聊过,但我不打算从供应链的角度来讲。
I'm talking about from a technology angle, right?
我是从技术的角度来讲,对吧?
Memory um capacity and bandwidth have been improving very slowly.
内存的呃 容量和带宽一直提升得非常慢。
The NAND cell was invented like 25 years ago.
NAND cell 大概是 25 年前发明的。
The DRAM cell was invented like 40 years ago and there's been no major breakthrough in in cell like you know how what a NAND cell is.
DRAM cell 大概是 40 年前发明的,而在单元(cell)层面一直没有重大突破,你知道的,就是 NAND cell 到底是个啥。
Obviously NAND is like a very simple gate or DRAM cell.
显然 NAND 就是个非常简单的门,DRAM cell 也一样。
There there is stuff that could come down the pipeline that could be hugely innovative.
是有一些可能正在酝酿的东西,可能会带来巨大的创新。
But even over the last, you know, five years, all we've really done is make the HBM, you know, more stacks, faster, but actually there's like new innovations coming in the next few years where instead of, you know, stacking the HBM separately from the chip, you stack the memory directly on the chip and that makes your bandwidth explode.
但即便是过去,你知道的,五年,我们真正做的也就是把 HBM,你知道的,堆更多层、做得更快,可实际上接下来几年会有新的创新:不再,你知道的,把 HBM 跟芯片分开堆,而是把内存直接堆在芯片上,这会让你的带宽爆炸式增长。
Um, and so there's interesting companies in that space and interesting PC's that companies are trying to do there.
呃,所以这个领域有一些有意思的公司,也有一些公司在那儿尝试的有意思的 PC(此处原文存疑)。
I think like memory bandwidth is one of the biggest.
我觉得内存带宽是最大的瓶颈之一。
Another one is um for the history of like silicon basically for the last two decades at least you know how many watts a chip is can be easily predicted just by looking at it for for a data center or desktop chip it it peaks up at one watt per millimeter squared
另一个是,呃,纵观硅片的历史,基本上至少过去二十年,你知道的,一颗芯片有多少瓦,只要看一眼就能轻松预测——对数据中心或桌面芯片来说,它的峰值大约是每平方毫米 1 瓦。
and so if a chip is 100 millimeter squared generally the power consumption is around 100 or a little bit less
所以如果一颗芯片是 100 平方毫米,一般功耗就在 100 瓦左右、或者略低一点。
um and if you look at the newest Nvidia silicon the newest TPU silicon it's still on that range of one watt per millimeter squared so you know chips are now getting to you know, 1400 watts.
呃,而如果你看最新的 Nvidia 硅片、最新的 TPU 硅片,它仍在每平方毫米 1 瓦这个区间,所以你知道的,芯片现在已经到了,你知道的,1400 瓦。
Next generation is 2,000 watts for Nvidia.
Nvidia 下一代是 2000 瓦。
Um, with Rubin and such.
呃,就是 Rubin 那一代之类的。
Uh, and and you move forward to Rubin Ultra, it's going to be like 4,000 watts or something like that.
呃,再往前走到 Rubin Ultra,大概会是 4000 瓦左右。
But really, there's increasing the amount of silicon.
但说到底,这是在增加硅片的用量。
What's exciting is we're now finally doing things and and it's in development right now where you actually can pump the amount of power into the silicon uh to be way more than one watt per millimeter squared.
让人兴奋的是,我们现在终于在做一些事——而且正在研发中——你其实可以把灌进硅片的功率呃 提到远超每平方毫米 1 瓦。
And now that all of a sudden means you need less silicon.
于是这一下子就意味着你需要的硅片更少了。
Obviously, it's running at higher power.
显然,它是在更高的功率下运行。
It's less efficient in some cases, but you reduce the amount of silicon and you're able to like
某些情况下它效率更低,但你减少了硅片用量,而且你能够,就像——
like over thermal issues,
比如说,搞定散热问题之类的,
thermal issues.
散热问题。
Um there's uh interference of like electrical interference issues.
呃,还有呃 那种干扰,像电气干扰的问题。
There's all sorts of different issues uh that crop up and that's why it's a hard engineering problem.
有各种各样呃 冒出来的问题,所以这才是个难搞的工程问题。
That's why we've stuck at about one.
所以我们才一直卡在大约 1(瓦每平方毫米)。
But what's exciting is the world is trying to change these things.
但让人兴奋的是,整个业界都在试图改变这些。
I think interesting like in a different part of the supply chain it's sort of like you know people people will talk about like energy is hard and you know we have energy bottlenecks and it's like yeah but there's actually like very simple solutions you know one could think of right
我觉得有意思的是,在供应链的另一个环节,差不多就是,你知道的,大家会说能源很难搞、你知道的我们有能源瓶颈,然后就像,是啊,但其实有一些非常简单的解决办法,你知道的,你能想到的,对吧。
um take the millions of diesel engines for trucks that the US has the capacity to make um you can very trivially convert them to be using for gas uh in the assembly line and then stick them up to a electrical motor like back driving it so the electrical motor generates electricity rather than the electrical motor causing the the rotation of the wheel, for example, but doing it the opposite direction.
呃,拿卡车上那几百万台柴油发动机来说——美国有能力造出这么多——呃,你可以非常轻松地在装配线上把它们改成烧天然气(gas),呃,然后接到一台电动机上,像反向驱动它一样,让电动机去发电,而不是让电动机去带动车轮转动,举个例子,就是反过来做。
And now you've generated electricity by pumping gas into something that us can make millions of.
这样一来,你就通过往一个美国能造几百万台的东西里灌天然气,把电给发出来了。
Um, and then, okay, well, that sounds like a pain in the ass to uh service, right?
呃,然后,好吧,这听起来伺候起来呃 挺麻烦的,对吧?
Because now you have to have hundreds of these on a data center site.
因为现在你得在一个数据中心场地上摆上几百台这玩意儿。
Well, actually, you can just pull people out of car mechanic shops and have them run around and repair truck engines.
不过其实呢,你直接从汽车修理铺里把人拉出来,让他们满场跑着去修卡车发动机就行。
Actually, it's actually pretty trivial to not I don't want to say it's trivial, I couldn't do it.
其实,这其实挺简单的——不,我不想说它简单,反正我是干不了。
Um
呃——
I think you're making a really good point which is that like because the west wasn't really thinking about semiconductors even hardware more broadly the last 20 30 years we didn't have like much innovation we'd have the best minds like thinking about how do you improve these
我觉得你说得非常对,就是因为过去二三十年西方并没有真正在琢磨半导体、乃至更广义的硬件,所以我们没什么创新,我们不会让最顶尖的头脑去琢磨怎么改进这些东西。
why why would you why would you want to go work in hardware when you can uh make ads to ads
你干嘛,你干嘛要跑去搞硬件,明明可以呃 去做广告(赚广告钱)呢。
yeah exactly
对,正是如此。
um okay I'm dying to ask Nvidia versus TPU what are your thoughts
呃好,我特别想问一下 Nvidia 对 TPU,你怎么看?
um I think I think like everyone wants to pick one or the other for this, but it's really like a function of like look, you know, you look two years from now, Google's going to make 10 plus million TPUs and through their supply chain and Nvidia is going to make, you know, many more million tens of millions of GPUs
呃我觉得,我觉得大家都想在这两者里选一个站队,但这其实取决于——你看,两年之后,Google 通过它的供应链会造出一千多万片 TPU,而 Nvidia 会造出几千万片 GPU,数量还多得多。
and both are going to be 100 plus billion dollar, you know, well, Google's going to be 100 plus billion dollars, you know, of TPU created a year and and Nvidia will be, you know, 500 plus or, you know, whatever.
两边都会是一千多亿美元的量级——嗯,Google 一年造出的 TPU 会值一千多亿美元,Nvidia 会是五千多亿,或者随便多少吧。
I'm not making a specific estimate.
我不是在给一个具体的估算。
This is not revenue forecast. This is just a thought experiment.
这不是营收预测,这只是个思想实验。
Yeah. Or research.
对。或者叫研究。
You've been media trained.
你是被媒体培训过的啊。
Absolutely. you know, getting ready for the SpaceX idea.
那当然。你懂的,为 SpaceX 那个点子做准备嘛。
Um, are you guys big in SpaceX? Okay, so that makes sense. Um,
呃,你们在 SpaceX 上投得多吗?好,那就说得通了。呃,
we're very lucky to be very large investors.
我们很幸运,是它非常大的投资方。
Awesome. Awesome. Um, so I would say um the the case of sort of like Google TPUs versus uh Nvidia GPUs, they both have like points that are really like in their favor, right?
牛。牛。呃,所以我会说,呃 Google TPU 对 Nvidia GPU 这场对决,两边都各自有真正站得住脚的点,对吧?
You know, Nvidia will be like, "Oh, well, we have switches and we're general purpose."
你懂的,Nvidia 会说:"哦,我们有交换机,而且我们是通用的。"
And and TPUs will be like, "Well, we're more optimized. actually more energy efficient and our network is actually more um optimized for certain types of network architectures.
而 TPU 这边会说:"嗯,我们更优化,其实更省电,而且我们的网络其实针对某些类型的网络架构更、呃更优化。"
And so you have like these counterpoints that both would really uh get into and you know I could with a straight face argue with you like that GPUs are way better than TPUs or TPUs are way better than GPUs but it comes down to hardware software co-design.
所以你会有这一来一回的对呛,两边都能真的、呃杠进去,而且我可以面不改色地跟你论证 GPU 比 TPU 强得多,或者 TPU 比 GPU 强得多——但归根到底,这取决于硬件软件的 co-design(协同设计)。
So actually the way OpenAI's models are headed, it would be a terrible decision for them to use TPUs potentially.
所以其实按 OpenAI 模型的走向,他们去用 TPU 有可能是个糟糕透顶的决定。
And the way that Anthropic and Google's uh models are headed, it's actually a terrible decision potentially for them to train with GPUs.
而按 Anthropic 和 Google、呃模型的走向,他们用 GPU 来训练反倒有可能是个糟糕透顶的决定。
I mean, it'd be fun to what's the what's the fundamental difference there?
我是说,这挺有意思的——那到底、那底层的差别是什么?
There's various things, right?
有好几方面,对吧?
Like the size of the matrix multiply unit is different as a as a very simple thing.
举个最简单的,matrix multiply(矩阵乘)单元的大小就不一样。
And therefore, the shape of the matrix multiply you do, the attention mechanism you use, uh the way that attention mechanism is structured, the way the experts are structured.
因此,你做的矩阵乘的形状、你用的 attention mechanism(注意力机制)、呃这个注意力机制怎么搭、experts(专家)怎么搭,都不一样。
So you think so open AI and anthropic are converging the very different model architectures. I think they're I think they have quite different model architectures.
所以你觉得 OpenAI 和 Anthropic 在往这套很不一样的模型架构上收敛。我觉得他们,我觉得他们的模型架构相当不一样。
In fact, um, you know, OpenAI's are much more sparse, um, and that has benefits.
事实上,呃,你懂的,OpenAI 的要 sparse(稀疏)得多,呃,这有它的好处。
And then anthropics are, you know, they're still sparse, but more dense in general, and and that has different benefits.
然后 Anthropic 的呢,你懂的,他们还是 sparse,但整体上更 dense(稠密)一些,而这又有另一套好处。
And there's many other things, right?
还有很多别的因素,对吧?
The network topology, right?
网络拓扑,对吧?
Nvidia, all of their chips are connected to switches, NVLink switches.
Nvidia,他们所有芯片都连到交换机上,NVLink 交换机。
For Google, they have no switch.
Google 呢,他们没有交换机。
Um, but what they've done is they've been able to, you know, Nvidia, the NVLink can only connect 72 GPUs.
呃,但他们做到的是——你懂的,Nvidia 那边,NVLink 只能连 72 片 GPU。
for Google, their ICI can connect 8,000 chips at super high bandwidth, but you have to pass through other chips to get there because there's no switch.
而 Google 这边,他们的 ICI 能以超高带宽连 8,000 片芯片,但你得穿过别的芯片才能到,因为没有交换机。
And so there's like there's trade-offs there.
所以这里就有,这里有取舍。
There's positives and negatives and that influences the model architecture.
有正面有负面,而这会反过来影响模型架构。
It's not necessarily that you should uh you know claim one is better than the other because at the end of the day, how do you say that this is better than that when you can't measure them in isolation because it also extends up to the model layer, right?
你未必该、呃你懂的、断言哪个比哪个强,因为说到底,当你没法孤立地衡量它们时,你怎么说这个比那个好?因为它还会一路延伸到模型层,对吧?
Um
呃
but I remember for a long time thinking you know one the programmability of Nvidia and just CUDA as as such a big moat.
但我记得有很长一段时间,我一直觉得——你懂的,Nvidia 的可编程性,还有 CUDA 本身,是一条那么宽的 moat(护城河)。
It seems to me that narrative has kind of changed at least in my mind for the last three six months like model companies no longer care about if we have to write custom kernels for you know this other chip so be it.
在我看来,这套叙事有点变了,至少在我脑子里,过去这三到六个月里——模型公司不再在乎"要是我们得为、你懂的、另一款芯片写自定义 kernel,那就写呗"。
We'll work with four or five chips if we have to.
要是必须的话,我们会同时搞四五款芯片。
Um Claude and Codex are actually quite good at doing a lot of that optimization work.
呃 Claude 和 Codex 其实相当擅长干很多这类优化的活。
And so it seems like some of the and then and then it's you know it's not like there's 10,000 model companies that are each you know each need programmability.
所以看起来有些、然后、然后就是——你懂的,又不是有一万家模型公司,每一家都、你懂的、每一家都需要可编程性。
There's on the order of tens maybe model companies and so it seems to me that like if you the fundamental premise of like tens of thousands of big customers that need CUDA compatibility like it seems that kind of thesis is is changing in the last
模型公司也就几十家这个量级,所以在我看来,如果你——那种"有几万家大客户需要 CUDA 兼容"的根本前提,这套论点好像在最近正在变。
Yeah. I mean I mean certainly the CUDA moat and software moat is at least partially uh disentangled because you know models are just great at coding and all software gets commoditized in that case.
对。我是说,我是说,CUDA 护城河和软件护城河当然至少部分被、呃拆解掉了,因为你懂的,模型就是特别擅长写代码,那样一来所有软件都会被商品化。
I do think there is some level of like open source and you know what people call the CUDA moat is not actually anything to do with CUDA but it's like the fact that DeepSeek Kimi and and Zhipu and and Alibaba and Tencent all these all these companies Xiaomi had an awesome model recently their models are co-designed for GPUs
我确实觉得有某种程度上的开源,而且你懂的,人们所谓的 CUDA 护城河其实跟 CUDA 一点关系都没有,它其实是这么回事:DeepSeek、Kimi、还有 Zhipu、还有 Alibaba、Tencent,所有这些、所有这些公司——小米最近也出了个很棒的模型——他们的模型都是针对 GPU 做 co-design 的。
and therefore if I want to run them on TPUs actually in some cases they don't run really well on TPUs now
因此,如果我想在 TPU 上跑它们,其实在有些情况下它们在 TPU 上跑得并不好。
Google just has to create their own open source model ecosystem or open source models themselves so they have the Gemma models
Google 就只能去搞自己的开源模型生态,或者自己把模型开源,所以他们有了 Gemma 系列。
and and so you end up with like well that's not really CUDA as a moat it's that the downstream product is more optimized for Nvidia
所以你最后得到的结论是:嗯,那其实不是 CUDA 当护城河,而是下游产品对 Nvidia 更优化。
and in these cases these companies are just open sourcing them or like Nemotron is just open sourcing it
在这些情况下,这些公司只是把模型开源,或者像 Nemotron 就是把它开源。
and then the users of it for example the open you know the inference uh API providers the RL companies that are trying to take open models and customize them for company's business use cases all these different companies are downstream of the fact that like okay well I guess I need to use Nvidia because the ecosystem uses Nvidia even though I don't partic particularly care about writing CUDA kernels because the models are great at that
然后用它的人——比方说开源的、你懂的、那些 inference(推理)、呃 API 提供商,那些想拿开源模型、把它们为各家公司的业务场景做定制的 RL 公司——所有这些不同的公司都处在这么一个事实的下游:好吧,我大概得用 Nvidia,因为生态就是用 Nvidia 的,哪怕我并不、特、特别在乎自己写不写 CUDA kernel,因为模型干这个特别在行。
but it's like the shape of like well this expert the the demod is this and you know the hidden dimension blah blah blah is this right
但问题是那个形状——嗯,这个 expert、那个 demod 是这样,还有你懂的,hidden dimension 之类之类是那样,对吧。
and and so therefore it's better to run on Nvidia GPUs than it is on TPUs and vice versa right
所以因此,在 Nvidia GPU 上跑就比在 TPU 上跑更好,反过来也一样,对吧。
if Google were to actually open source really good models you know this would be the same thing right
如果 Google 真的把很好的模型开源,你懂的,道理是一模一样的,对吧。
people would take their models and they'd be like oh wow these don't run that well on Nvidia GPUs um I should actually just rent TPUs or buy TPUs and do it on there
人们会拿他们的模型,然后就会:哦哇,这些在 Nvidia GPU 上跑得没那么好,呃我其实应该直接租 TPU 或者买 TPU,在上面跑。
for small teams you're going to want to use all the open source software like vLLM SGLang um PyTorch all that stuff
对小团队来说,你会想把所有开源软件都用上,比如 vLLM、SGLang、呃 PyTorch,诸如此类。
but the big labs they don't necessarily need to use all that right
但大厂实验室,他们未必需要用那一整套,对吧。
OpenAI forked PyTorch long ago and you know anthropic and all these other people don't necessarily rely heavily on the open- source implementation of you know these things
OpenAI 老早就 fork 了 PyTorch,而且你懂的,Anthropic 还有其他这些人,未必重度依赖这些东西的开源实现。
they forked things or built it on their own already and so they don't need to rely on the open source
他们要么 fork 了,要么早就自己造好了,所以他们不需要依赖开源。
and therefore now it's more like you know I'll choose the best hardware and I'll co-design my model and infrastructure software through and through for that hardware uh that is the best and most cost efficient
因此现在更像是——你懂的,我会挑最好的硬件,然后把我的模型和基础设施软件从头到尾为那款硬件做 co-design,呃就是那款最好、最划算的硬件。
and you know I'll have AI help me write all that software.
而且你懂的,我会让 AI 帮我把那些软件全写出来。
What do you think of Cerebras?
你怎么看 Cerebras?
I think Cerebras is a really innovative company.
我觉得 Cerebras 是家特别有创新力的公司。
Um I I think in in some spots of the market they're really really good.
呃我我觉得在市场的某些点上,他们真的真的很强。
Um very fast inference.
呃就是超快的推理。
I think that's a big market.
我觉得那是个大市场。
Uh we use fast mode almost exclusively at SemiAnalysis.
呃我们在 SemiAnalysis 几乎全程都用 fast mode。
Um
呃
by the way I love how disciplined you've been about accounting for I don't know if that was one exhibit you did or if you do it consistently but accounting for the dollar spent and the ROI on each task.
顺便说一句,我特别欣赏你在核算这件事上有多自律——我不知道那只是你做的一次展示,还是你一直这么干,但你会把每个任务花的钱和 ROI 都算清楚。
Awesome analysis.
分析得太棒了。
Yeah. Yeah. We we we uh we do it pretty diligently and so thank you.
对。对。我们我们我们呃我们做得挺勤快的,所以谢谢你。
That was the dark GDP article that we wrote.
那就是我们写的那篇「暗 GDP」文章。
Um and so and and also like track everyone's token spend by day and if someone's like spiked up I'm like what did you do?
呃所以还还会按天追踪每个人的 token 花销,要是谁突然飙上去了,我就会问:你干什么了?
It's like okay thank you for telling me that that seems worth it. Cool. On with my day.
然后就是:好,谢谢你跟我说,这看起来挺值的。行。继续过我的一天。
I think fast mode is obviously worth a lot for high-end tasks, right?
我觉得 fast mode 对高端任务显然很有价值,对吧?
I could just see so many different use cases where you know super fast tokens are worth it.
我能想到特别多不同的场景,超快的 token 是值这个钱的。
I can also see the flip side where there's a lot of use cases where super fast tokens aren't needed and and therefore uh the market won't pay for them and they'll use GPUs and TPUs instead.
我也能看到反面:有很多场景根本不需要超快的 token,所以呃市场不会为它买单,他们会改用 GPU 和 TPU。
I think the big risk for Cerebras is I mostly think the best models are the ones that you want to use fast mode on and small models you necessarily might not use fast mode on.
我觉得 Cerebras 最大的风险在于——我基本认为,你想用 fast mode 的是那些最好的模型,而小模型你不一定会用 fast mode。
I could see that being wrong with you know financial markets maybe or something like that like a Jane Street high frequency trading or something like that um or medium frequency trading.
这个判断也可能错,比如在金融市场或者类似的地方,像 Jane Street 那种高频交易之类的,呃或者中频交易。
Um but ultimately you know running really large models at really long context is very difficult on SRAM based chips like Cerebras like Groq.
呃但归根结底,在基于 SRAM 的芯片上——像 Cerebras、像 Groq——跑超大模型加超长上下文是非常难的。
and so now it all of a sudden is like you know what happens then if like the models get too big right if OpenAI's model is not you know on the order of uh you know hundreds of billions parameters or you know low trillion parameters but it's actually 10 plus trillion parameters now all of a sudden I don't think that that will fit on Cerebras right
所以现在突然就变成:那要是模型变得太大了会怎样,对吧,如果 OpenAI 的模型不是呃几千亿参数、或者一两万亿参数这个量级,而是实际有 10 万亿以上参数,那一下子我就不觉得它塞得进 Cerebras 了,对吧。
and then if that doesn't with a long context length right if you have a million context length now that makes it really difficult to justify you know,
然后如果再配上长上下文,对吧,要是你有一百万的上下文长度,那就更难说得过去了,你懂的,
and and as all so far we've seen the bulk of revenue and usage at the labs be on their best model.
而且到目前为止,我们看到各家实验室的绝大部分收入和用量都集中在他们最好的模型上。
Even when the model price has gone up, we've seen that.
哪怕模型价格涨了,我们也看到是这样。
Um there's some data that shows that even though Fable just released today, they've had incredible amounts of people switch to Fable and Mythos, sort of that next tier model, even though it's way more expensive.
呃有些数据显示,即便 Fable 今天才发布,已经有惊人数量的人切换到 Fable 和 Mythos——也就是那种更高一档的模型——哪怕它贵得多。
And so um
所以呃
is that and that's volume by dollars totally.
那这个,那完全是按美元算的量吧。
But was that volume by tokens?
但那是按 token 算的量吗?
Well, I guess who cares about volume by tokens?
呃,我觉得谁在乎按 token 算的量啊?
It's about the dollars.
重点是美元。
Fair enough.
有道理。
Right. If I don't care that there's, you know, uh, you know, I don't know, 200,000 Mini Coopers or Toyota Camry sold if if, uh, you know, I don't know, Ford F150s are 5x ASP and they sell only half as much.
对。我才不在乎卖了,你懂的,呃,反正,不知道,20 万辆 Mini Cooper 或者 Toyota Camry,如果如果,呃,你懂的,不知道,Ford F150 的 ASP 是它们的 5 倍,而销量只有一半。
Okay. Right.
好。对。
And then and and therefore the most lucrative market is pickup trucks in America.
那么所所以美国最赚钱的市场就是皮卡。
Right.
对吧。
Mostly being facetious, but like
基本上是开玩笑啦,但就是
I do think this is one of the things that you've done so well and differentiates you from almost everyone else is that you you care so much about the economics in addition to the technology.
我确实觉得,这是你做得特别好、几乎让你和所有人拉开差距的一点——你除了懂技术,还这么在意经济账。
And I think very few people bridged those two thing things well.
我觉得很少有人能把这两件事很好地连起来。
And so
所以
I think I think it's really fun inside of SemiAnalysis because we have 90 people and like a big chunk of them are technologist engineers across the whole supply chain.
我觉得我觉得在 SemiAnalysis 内部特别好玩,因为我们有 90 个人,其中很大一块是横跨整条供应链的技术型工程师。
Um and then a big chunk is people who are formerly at hedge funds and you see these arguments like people are like oh well that doesn't matter and it's like then someone's like well but cost and then someone the engineers like no no but this technology is the coolest.
呃然后还有一大块是以前在对冲基金干过的人,你会看到这种争论,有人说「哦这个不重要」,然后有人说「可是成本啊」,然后有个工程师说「不不,可是这技术最酷了」。
You see this you see this organically like fight it out.
你会看到这个,你会看到他们很自然地就吵起来。
Um and and were pretty informal and you know given the fact that I was a for moderator is you can imagine what the the enjoying it
呃而且我们挺随意的,而且你懂的,鉴于我以前是个论坛版主,你可以想象那那有多享受
you don't wrestle with a pig because a pig enjoys it.
你别跟猪摔跤,因为猪乐在其中。
Exactly.
正是。
Just on this topic before going to the next question.
在进入下一个问题之前,就着这个话题说一下。
Are there like trigger topics in semis for you?
在半导体里有没有那种一说就能戳到你的话题?
You know like if someone's like which is like such a meme you think this person must be a like if you know if it's like oh you like memory is the bottleneck.
你懂的,就像要是有人说——这已经是个梗了——你就觉得这人肯定是个,比如要是他说「哦你懂的,内存才是瓶颈」。
I mean it's true but like um I think I think moreover the one that really gets me is people are like AI has no ROI
我是说这话没错,但就是呃我觉得我觉得更让我上头的是那些说「AI 没有 ROI」的人
infuriates me right like there's like what's the ROI or like denying model progress right there's these people that are like models aren't getting better they're not reasoning they can't think they're going to deadend and plateau
这让我火大,对吧,就是那种「ROI 在哪」的说法,或者否认模型在进步,对吧,有那么一批人说「模型没在变好、它们不会推理、不会思考、迟早撞死胡同然后见顶」
and it's like bro the line has been up and to the right in terms of capabilities this entire time
然后就是:哥们,能力这条线这一整段时间都是往右上方走的啊
and they're like look this benchmark didn't improve that's cuz it said 90% look at the new benchmark you saturated now they're skyrocketing, right?
然后他们说「看,这个 benchmark 没提升」,那是因为它已经到 90% 了,你看新的 benchmark,你把旧的刷爆了,现在它们在猛涨,对吧?
It's like I think that's more so the issue and challenge.
就是我觉得问题和挑战更多是出在这儿。
Like I think semis are really complex and I don't fault people for um lacking like understanding of it.
比如我觉得半导体真的很复杂,我不会因为呃有人不太懂它就怪他们。
Like I learn stuff every day about the semiconductor supply chain from people and I've been studying it for you know arguably 18 years since I started moderating the forums when I was 12 right like you know arguably been studying it for that long
比如我每天都在从别人那儿学到半导体供应链的新东西,而我研究它,你懂的,可以说有 18 年了,从我 12 岁开始当论坛版主算起,对吧,你懂的,可以说研究了这么久
but even then like and it's like live breathed and that's all I care about but there's so many layers of the abstraction stack
但即便这样,就是,我天天活在里面、呼吸着它,我在乎的就这一件事,可这个抽象栈还是有太多层了
it's like like I learned about a new chemical that does like a hundred million dollars of sales like yesterday and I'm like whoa didn't know this one existed and what process it did
就好比我昨天才知道有一种新化学品,一年能做到大概一亿美元的销售额,我就想:哇,不知道还有这玩意儿,也不知道它是干哪道工序的
and it's like but it's like you know you learn about things all the It's like okay hundred billion dollar sales in a you know couple hundred billion dollar industry is whatever but like you know it's like
然后就是,但就是你懂的,你一直在学新东西,就是,好吧,一千亿美元销售额搁在一个你懂的几千亿美元的行业里也就那样,但就是你懂的,就是
but it's essential
但它是不可或缺的
it's essential and it's like actually every chip requires it.
它是不可或缺的,而且就是,实际上每颗芯片都离不开它。
It's like wow I guess there are a thousand process steps and you know it's like oh yeah you like semiconductors name every process step.
就是,哇,我猜有一千道工序,你懂的,就是,哦对,你像半导体那样把每道工序都点一遍名。
It's like no come on.
就是,不,得了吧。
What what I think is the most funny is when people have all the facts in front of them and then they get the conclusion completely wrong.
我我觉得最好笑的,是人们明明所有事实都摆在眼前,结论却完全搞错了。
Um and that's
呃而这
that happens in our job all the time too.
这种事在我们这行也天天发生。
Yeah.
对。
Yeah. I mean, I can't I I get I I think my attitude is not to be mad that you do that.
对。我是说,我没法,我我懂,我我觉得我的态度不是因为你犯这错就生气。
It's to do it as fast as possible.
而是尽可能快地把它做出来。
I think the industry because it's so it's just like AI is the most important thing in the world right now and there's so many near-term bottlenecks.
我觉得这个行业,因为它就是,AI 现在就是全世界最重要的东西,而且短期的瓶颈太多了。
We talk a lot about the near-term.
我们聊了很多短期的事。
Are there longer term things that you're really excited about?
有没有更长期、让你特别兴奋的东西?
Like say on a 10-year time frame?
比如说放到 10 年的时间尺度上?
We talked about orbital data centers, but like like siliconics, you think they're underrated or overrated on a 10-year time frame?
我们聊过在轨数据中心,但像 siliconics 这种,你觉得在 10 年尺度上是被低估还是被高估了?
Are there other things that on a 10-year time frame?
还有别的、放在 10 年尺度上值得关注的吗?
Yeah, I mean I think on SP I think space is like super crazy awesome in the 10-year time frame that I'm you know for space data centers and all these sort of mining asteroids and all these things which is you know super excited about the vision of SpaceX right um again not investment advice before you hop in um
嗯,我是说,我觉得在太——我觉得太空在 10 年尺度上简直疯狂地酷,你懂的,太空数据中心、开采小行星这一类东西,我对 SpaceX 的这个愿景超级兴奋,对吧,呃,再说一遍,你上车之前这不构成投资建议,呃
I think I think on the semiconductor side tremendous market movements and tremendous like things can happen just when like things happen one year later or sooner and so that's all like technology that like you know in terms of like co-packaged optics like well like everyone knows it's going to happen by the end of the decade
我觉得,我觉得在半导体这边,会有巨大的市场变动,巨大的,那种事随时能发生,就是有些事晚一年或早一年而已,这些都是技术层面的东西,你懂的,比如 co-packaged optics,大家都知道它会在这个十年结束前落地
the the debate is like '27, '28, '29, 2030 but some point along there it's going to happen.
争论只是在 '27、'28、'29 还是 2030,但反正就在这几年里某个时间点会发生。
I think the more interesting thing is like there's companies like um I did you guys invest in Naveen Rao's company?
我觉得更有意思的是,有些公司,比如,呃,你们投了 Naveen Rao 的公司吗?
We did.
投了。
Okay. Yeah.
好。嗯。
So I think like he's trying to innovate on like the silicon layer on the software abstraction layer and the model layer simultaneously
我觉得他是想同时在硅片层、软件抽象层和模型层上一起做创新,
and he fully understands that it's not a like a you know we're going to do this in a few years.
而且他非常清楚这不是那种,你懂的,"我们几年内就搞定"的事。
It's not a two-year time frame.
这不是两年能搞定的。
Yeah. It's not a few year time frame.
对。这也不是几年的时间尺度。
It's a long-term bet.
这是个长期赌注。
Um, and like stuff like that is like, okay, we're going to bring like potentially like analog compute with energy based models and like all this crazy all at once.
呃,像这种事就是,好,我们可能要一口气把模拟计算(analog compute)和能量模型(energy based models)这些疯狂的东西全搞上。
It's like that's exciting.
这就很让人兴奋。
Probably won't work, but you know, that's exciting and I I like really look forward to
大概率不会成,但你懂的,这很兴奋,我我真的很期待——
definitely won't work quickly.
肯定不会很快就成。
Yeah, definitely won't work quickly is what I should say.
对,我应该说"肯定不会很快成"。
I believe in Naveen and like, you know, I I I met him very, you know, I think he's one of the first people I met in the industry um, funnily enough, like in 2020 or 2021.
我信 Naveen,而且,你懂的,我我我很早就认识他了,我觉得他是我在这个行业里最早认识的人之一,呃,说来好笑,大概是 2020 或 2021 年。
Um, actually 2020.
呃,其实是 2020。
Yeah. It says something about him.
对。这挺能说明他这个人的。
I think he's someone in my experience. He's always trying to
我觉得,以我的经验,他一直在——
I baited him on the internet. I baited him on the internet.
我在网上钓他上钩的。我在网上钓他上钩的。
That's
那——
He's always trying to help the younger generation.
他一直想帮年轻一代。
He's trying to identify talent. And
他一直在发掘人才。而且——
he was also so ahead of his time with Mosaic.
他做 Mosaic 的时候也太超前于时代了。
I remember getting pitched.
我还记得当时被人推介过(来投)。
No, it was 2019.
不,是 2019。
I was still I was still anonymous then actually.
那时候我我其实还是匿名的。
I I baited him on the internet and he started replying and then I just took it to DMs and then took it to a call
我我在网上钓他,他就开始回复,然后我把话题拉到私信,再拉到通话,
and like that was the first person who's like really important that I talked to in the entire semiconductor industry.
他算是我在整个半导体行业里聊过的第一个真正重要的人。
funny.
挺好玩的。
Um, but yeah, sorry to interrupt.
呃,不过,抱歉打断了。
That's funny.
挺有意思的。
What do you think is the end state of the ecosystem?
你觉得这个生态的终局是什么样?
Like do you think every lab, every hyperscaler just has its own chips?
比如你觉得每个实验室、每个 hyperscaler 都会有自己的芯片吗?
Like train seems like it's now working, right?
像训练芯片这块现在看起来跑通了,对吧?
So do you think we end up with every lab, every hyperscaler has it own chips at least for inference and then maybe for training you go to Nvidia or whoever or what do you think is the end state?
所以你觉得我们最终会走到:每个实验室、每个 hyperscaler 至少在推理上都用自己的芯片,然后训练也许还是去找 Nvidia 之类的?还是说你觉得终局是什么样?
I think everyone will try and stop trying.
我觉得所有人都会去试,然后又停下来。
I think ultimately um you know supply chains matter.
我觉得归根结底,呃,你懂的,供应链很重要。
what technology you can bring in matters and more and more as the industry gets bigger supply chain diversification happens.
你能引进什么技术很重要,而且随着行业越来越大,供应链的多元化会越来越多地发生。
Um you know right now everyone's chip more or less looks the same.
呃,你懂的,现在大家的芯片长得差不多都一样。
It's a big logic compute die in the center and there's some HBM on the right and left and on the top and bottom top side is networking and then the bottom side is PCIe and other IO.
中间是一大块逻辑计算裸片(logic compute die),左右两边放一些 HBM,上下也有,顶上是网络,底下是 PCIe 和其他 IO。
Um and that is the exact same structure for tranium TPU Nvidia chips.
呃,而 Trainium、TPU、Nvidia 的芯片结构一模一样。
Um and most of the startups um not Grock and Fris are doing weird but that's cool you know um
呃,而且大多数创业公司,呃,除了 Groq 和 Fris 在搞些奇怪的东西——但那挺酷的,你懂的,呃
I think like as you step forward we're going to get more bifurcation of hardware architecture and model architecture and therefore people are going to co-optimize them
我觉得往前走,硬件架构和模型架构会出现更多分叉,于是大家会去协同优化它们,
and you know some of them will end up in local minimas right
然后你懂的,其中有些会落进局部最小值(local minima),对吧,
you know as we're you know if this is like gradation gradient descent like people are like trying to go to the most optimized solution some people will race to a local minima
你懂的,如果把这想成梯度下降(gradient descent),大家都想走到最优解,但有些人会一路冲进某个局部最小值,
and then the question is like how do you leap how do you scoot back over to like the absolute minima
那问题就是:你怎么跳出去,怎么挪回到那个绝对最小值(absolute minima),
and some to some extent like a general more Nvidia will always be more general purpose than anyone else's chip in general um at least on a parallel AI compute basis
而且某种程度上,Nvidia 总体上永远会比别家的芯片更通用,呃,至少在并行 AI 计算这个维度上,
because they have so many customers who care about different things who will always give them feedback in the design
因为他们有太多在乎不同东西的客户,这些客户会一直在设计上给他们反馈,
you know the minima will always be better than them
你懂的,那个(专用的)最小值总会比他们更好,
but is that minima a local minima like is is the TPU or tranium or grock or cerebras or whoever's design optimized awesomely for here but in the end state actually you got to go over here and so they're the wrong
但问题是那个最小值是不是局部最小值——就是 TPU、Trainium、Groq、Cerebras 或者随便谁的设计,在"这儿"优化得超棒,可到了终局你其实得去到"那儿",那他们就押错了——
um and Maybe they make a great time, they're great for a little bit of time, but then they end up being wrong.
呃,也许他们风光一阵,好上一小段时间,但最后押错了。
It's like that's the real question.
这才是真正的问题。
Um, and so I think I think there will be a big market for general purpose AI compute.
呃,所以我觉得,我觉得通用 AI 算力会有一个巨大的市场。
Um, because you talk to people at labs, they don't even know what architecture they're going to be doing in a year.
呃,因为你去问实验室里的人,他们连一年后自己要做什么架构都不知道。
Like, right, like they literally don't know what architecture they're going to be doing in a year.
就是,对吧,他们真的完全不知道一年后要做什么架构。
They have bets.
他们有一堆押注。
They have many research bets and and that's this exciting thing, but they don't know where where it's going.
他们有很多研究上的押注,这正是激动人心的地方,但他们不知道最后会走到哪。
generally they like know what hardware they have and they're trying to co-optimize
一般来说他们知道自己手上有什么硬件,然后想去做协同优化,
but ultimately like if a new breakthrough happens on model architecture it's like just replace the attention mechanism with something else right who knows
但归根结底,如果模型架构上出了个新突破,比如把注意力机制换成别的什么,对吧,谁知道呢,
or you know all of a sudden you know something happens the best hardware will change
或者你懂的,突然出了点什么事,最好的硬件就变了,
and therefore like are people going to make fiveyear investments on hardware solely on you know an an asich that is more specialized or are they going to do so they're going to have some bucket of more general purpose compute
所以人们会不会单押一款更专用的 ASIC、做一笔五年的硬件投资?还是说他们会留一桶更通用的算力?
and so you see this with like Google's paying $11 an hour per GPU to XAI for G for GPUs, right?
你看 Google 就是这样,为了拿 GPU,按每张 GPU 每小时 $11 的价钱付给 xAI,对吧?
Like that's insane, right?
这也太疯狂了,对吧?
It's a very high amount of uh obviously compute is limited and and so on and so forth, but it's like very like insane,
这价格非常高,呃,显然算力是稀缺的,等等等等,但真的很疯狂,
but at the same, you know, despite the fact that they have TPUs and so there's like some questions there like why do they do that?
但与此同时,你懂的,他们明明自己有 TPU,所以这里就有些疑问,他们为什么要这么干?
Um Google actually has three different design programs for TPUs.
呃,Google 其实有三套不同的 TPU 设计项目。
They're making a TPU with Broadcom.
他们和 Broadcom 一起做一款 TPU。
That's a different architecture than the TPU with MediaTek.
那和跟 MediaTek 一起做的 TPU 是不同的架构。
That's a different TPU than the architecture that is, you know, I won't disclose, you know, by research.
那又和另一套架构不同,那套,你懂的,出于我们的研究我就不披露了。
Um but, you know, they're they're making different architectures.
呃,但你懂的,他们做的是好几种不同的架构。
It's not just like, oh, they're making TPUs with a couple vendors.
不是那种"哦,他们就是找几家供应商一起做 TPU"那么简单。
It's the same architecture.
(不是)同一套架构交给不同厂商。
It's different architectures.
而是不同的架构。
And the third one is a very different architecture from the first two.
而且第三套架构跟前两套差别非常大。
And so, I think people recognize that the local minima can happen.
所以我觉得大家都意识到,局部最小值这种事是会发生的。
And therefore, um, I think everyone will have their own ASIC program.
因此,呃,我觉得每一家都会有自己的 ASIC 项目。
I think everyone will deploy billions of dollars of their own AS6, tens of billions of dollars.
我觉得每一家都会砸几十亿美元做自研 ASIC,甚至数百亿美元。
In the case of Google, hundreds of billions of dollars a year of their own AS6.
拿 Google 来说,每年会砸上千亿美元在自研 ASIC 上。
But ultimately, they're also going to have workloads that don't use TPUs, right?
但归根结底,他们也会有一些不用 TPU 的工作负载,对吧?
Some of the Google bets that are not Gemini Deepbind actually primarily use GPUs.
Google 有些不属于 Gemini / DeepMind 的押注,其实主要用 GPU。
They don't use TPUs.
它们不用 TPU。
Um, some of them also primarily use TPUs, right?
呃,其中有些也主要用 TPU,对吧?
It's a bit of a broad thing, but like, you know, maybe for drug discovery or for Waymo, you might not want to use TPUs.
这说得有点笼统,但你懂的,也许做药物发现、或者做 Waymo,你就不会想用 TPU。
I won't say which one it is, but like, you know, there's there's there's there's different architecture bets and different paths for AI.
我不说具体是哪个,但你懂的,AI 有不同的架构押注,也有不同的路径。
AI for science may have different algorithmic patterns than than general intelligence AGI models.
科学用 AI(AI for science)的算法模式,可能跟通用智能的 AGI 模型不一样。
Um, and so I think we'll see we'll see diversity continue to proliferate.
呃,所以我觉得我们会看到,多样性会继续扩散。
Yeah.
对。
and and and because the market has gotten so big, niches will be carved out
而且因为市场已经变得这么大,会切分出一个个细分领域,
and so that's makes it possible for companies to have their niche and actually make money even if the majority of the pie goes to Nvidia and TPU and tranium.
所以哪怕蛋糕的大头被 Nvidia、TPU 和 Trainium 拿走,公司们照样能占住自己的细分领域,并且真正赚到钱。
Okay. Love that.
好,我喜欢这个。
Can we talk about the data center buildout?
我们能聊聊数据中心建设吗?
Like one, it seems like I mean by all accounts if you look at the charts like dollars per compute hour, we are in the middle of a crazy compute crunch.
就是,第一,看起来——我是说,不管从哪个角度看,如果你看那些图表,比如每算力小时的美元成本,我们正处在一场疯狂的算力紧缺当中。
Um and it seems like it's both a demand and supply side crunch, right?
呃,而且看起来这是需求侧和供给侧双重紧缺,对吧?
demand for long agents skyrocketing, supply, all these data center buildouts are delayed.
长时程 agent 的需求在暴涨,供给这边,所有这些数据中心建设又都在延期。
Um, do you think this we're in a compute crunch for the foreseeable future or do you think it alleviates at some point?
呃,你觉得我们会在可预见的未来一直处于算力紧缺,还是觉得它到某个时候会缓解?
Yes, every quarter we're deploying vastly more compute than the prior quarter and there's more data centers built than the prior quarter.
会,每个季度我们部署的算力都远远超过上一个季度,建成的数据中心也比上一季度多。
Um, this year there's going to be 20 gigawatts uh even accounting for the delays and next year there's going to be more than 30 gigawatts accounting for the delays.
呃,今年会有 20 gigawatts,呃,就算把延期算进去也有;明年把延期算进去也会超过 30 gigawatts。
Um, of course delays happen on everything, right?
呃,当然,什么东西都会延期,对吧?
Anything hardware can have a delay.
任何硬件都可能延期。
That's that's just the reality of life.
这就是,这就是现实。
Are we gonna have a compute crunch for the rest of our lives?
我们这辈子会不会一直算力紧缺?
It depends on what happens with models.
这取决于模型会怎么发展。
But like the TAM for Mythos, you know, Mythos 5, Fable 5 is not just like 2x that of Opus, right?
但比如说 Mythos 的 TAM,你懂的,Mythos 5、Fable 5 的 TAM 可不只是 Opus 的 2 倍,对吧?
The model is so much better and it can do so many more tasks that the TAM for it is way larger than that.
模型好太多了,能做的任务多太多了,所以它的 TAM 要比那大得多。
And yet compute in the world did not double in the last, you know, six months, right?
然而全世界的算力在过去,你懂的,这半年里并没有翻倍,对吧?
From, you know, Opus or maybe like seven or eight months since Opus 45 launched to now.
从,你懂的,Opus,或者说大概是从 Opus 45 发布到现在的七八个月里。
huge you know 46 47 48 were improvements but fable and Mythos were like a huge step function improvement
巨大的——你懂的,46、47、48 都是改进,但 Fable 和 Mythos 是那种巨大的阶跃式提升。
the world's compute did not double in that or or quadruple or whatever in that same time frame
而同一时间段里,全世界的算力并没有翻倍,或者翻四倍,或者别的什么。
but the demand for useful tasks that can be done by AI the number of useful tasks and the value of them that can be done by AI has
但 AI 能做的有用任务的需求——AI 能完成的有用任务的数量和价值——确实翻上去了。
and so now the question is what happens
所以现在的问题是接下来会怎样。
well obviously anthropic in Q2 is profitable their net income profitable um excluding stockbased compensation
嗯,很明显,Anthropic 在 Q2 是盈利的,他们的净利润是盈利的,呃,不含股权激励的话。
um And and I think by Q3 they may even be profitable including stockbased compensation.
呃,而且我觉得到 Q3,他们甚至可能连股权激励一起算都盈利。
That's like how profitable they're getting.
他们就是赚到了这个程度。
And their margins on a on a on an Opus token, at least Opus 48 token is like north of 80% for the API price.
而他们在每个,每个,每个 Opus token 上的毛利率,至少 Opus 48 的 token,按 API 价格算是 80% 以上。
They've got a lot of deals where their total corporate gross margins gets clawed down a little bit uh because of like how they do bedrock deals and vertex deals and things like that.
他们有很多合作,会把整体公司层面的毛利率往下拉一点,呃,因为他们做 bedrock 那种合作、vertex 那种合作之类的方式。
But ultimately their their per token margin is so high.
但归根到底,他们每 token 的毛利率高得离谱。
Well, then if you don't have the cap, they have the capability to pay ultimately every GPU they buy at above market rate.
那么,如果你没有那个上限——他们有能力,归根到底,为他们买的每一块 GPU 付高于市场价的钱。
You know, they also bought GPUs at above market rate from SpaceX, which is below the rate of Google, but that's because they signed earlier.
你懂的,他们也从 SpaceX 那儿用高于市场价买了 GPU,那个价比 Google 的低,但那是因为他们签得早。
Um, you know, it's it's something that, you know, other companies, maybe a ventureback company or company that's not really got positive uh margins can't necessarily do, right?
呃,你懂的,这是那种,你懂的,别的公司——比如一家靠风投输血的公司,或者一家其实没有正毛利的公司——不一定做得到的事,对吧?
What is the cost benefit ratios like every GPU I rent because I'm out of compute capacity I can immediately turn around and sell tokens on it or every TPU or every Trainium I can immediately sell tokens on it at a positive margin
成本收益比是怎样的呢:我因为算力不够而租的每一块 GPU,我可以马上掉头在上面卖 token,或者每一块 TPU、每一块 Trainium,我都能马上在上面以正毛利卖 token。
and if I'm running 75% gross margin and I double the cost of the compute it's fine I'm still running 50% gross margin
而如果我跑着 75% 的毛利,就算我把算力成本翻倍也没关系,我还有 50% 的毛利。
and spinning up more compute nodes is not really necessarily a human requiring task for them if they're renting them
而且多起几个算力节点,对他们来说不一定是需要人力的活儿,如果是租的话。
and so ultimately it's like well my NOI still goes up right
所以归根到底就是,嗯,我的 NOI 照样往上走,对吧。
and and so I'm going to rent GPUs at whatever price at some level whatever price I want to pay I can pay.
所以我会去租 GPU,不管什么价——某种程度上,不管我想付什么价,我都付得起。
I have almost the reverse question of like at some point does this compute build out go bump at night?
我几乎有个相反的问题:这场算力建设会不会在某个时候半夜里就出岔子?
Earlier today I think there was a tweet like Crusoe publicly said one of their customers had asked to halt construction on one of their data center buildouts.
今天早些时候我记得有条推特,好像 Crusoe 公开说他们有个客户要求停掉他们某个数据中心的建设。
Like it seems like everybody in the ecosystem is so levered right now to like we got to build, we got to go build, we got to build.
就是,看起来整个生态里的每个人现在都重仓押在"我们必须建、我们得去建、我们得建"上。
High leverage high growth to me is like makes me very very nervous as investor.
高杠杆高增长这套,对我来说,作为投资人让我非常非常紧张。
Like
就是——
wait hold on.
等等,打住。
High leverage high growth means small amount of equity has huge upside.
高杠杆高增长的意思是,一小笔股权有巨大的上行空间。
You're not a debt investor.
你又不是债权投资人。
You're a credit you're an equity investor, right?
你是做信贷——你是股权投资人,对吧?
Let's go.
来吧。
Um,
呃,
look, you got you got to go to the school of private equity.
听着,你得,你得去上上私募股权那一派的课。
Levered buyouts only.
只做杠杆收购那种。
I actually come from the school of private equity.
我其实就是私募股权那一派出身的。
Oh, awesome.
哦,厉害。
She forgot the school.
她把这派给忘了。
It's been a VC for too long.
当 VC 当太久了。
No, I just do revenue multiples.
才没有,我只不过是用营收倍数估值罢了。
No, but are you do you see any signs of that?
不过说真的,你——你有看到那种迹象吗?
Are you worried about that?
你担心这个吗?
I I I see what you mean.
我,我,我懂你的意思。
Right.
对。
And that sort of goes back to the model point, right?
这某种程度上又回到了模型那个点上,对吧?
Obviously if the models expanding the total economic valuable like work sort of the dark GDP uh report that we did and the you mentioned earlier um if the work that these models can do does not expand faster than the compute capacity then that tide turns right
很明显,如果模型在扩大总的、有经济价值的那种工作——就像我们做的那份 dark GDP,呃,报告,还有你之前提到的——呃,如果这些模型能做的工作,扩张速度没有快过算力的产能,那么潮水就会掉头,对吧。
and over the last six months that tide has been you know very much levered in this direction of um you know the models can do more work or can is exp or expanding their TAM of work they can do faster than the compute is increasing
而过去这半年,那股潮水,你懂的,非常明显地压在这个方向上:呃,你懂的,模型能做更多工作,或者说在扩张它们能做的工作的 TAM,速度比算力增长还快。
And so prices go up.
所以价格就往上走。
It's very possible that all of a sudden model progress stops.
很有可能模型的进展会突然停下来。
You talk to anyone at Anthropic or OpenAI, maybe they're drinking the Kool-Aid, but you talk to basically all of them, they're like, "No, no, no, no. Model progress still go up."
你去问 Anthropic 或 OpenAI 的任何人,也许他们是在给自己灌迷魂汤,但你基本上问遍他们所有人,他们都说:"不不不不,模型进展还会继续往上走。"
Um, and so, you know, ultimately, you know, current methods could stall somewhere.
呃,所以,你懂的,归根到底,你懂的,现有的方法可能会在某个地方卡住。
I'm not sure where that would be.
我不确定那会是在哪儿。
It seems like we have line of sight to model improvement, rapid model improvement.
看起来我们对模型的提升是有清晰视线的,而且是快速的提升。
And in fact, models are improving faster than they were six months ago or a year ago because there's I wouldn't call it recursive self-improvement, but basically the engineer the models are helping write all the info and and launch the next model sooner and sooner and sooner.
而且事实上,模型改进的速度比半年前、一年前还快,因为有——我不会把它叫作递归式自我改进,但基本上,工程——模型在帮着写所有的东西,并且把下一代模型越来越快、越来越快、越来越快地推出来。
So you've got this like pseudo recursive self-improvement loop going
所以你就有了这么一个伪递归式自我改进的循环在转。
and so the models are getting better and better and better faster.
所以模型变得越来越好、越来越好、越来越好,而且越来越快。
Um and so but ultimately, you know, capital is a big problem which is why Google raised capital.
呃,所以——但归根到底,你懂的,资本是个大问题,这就是为什么 Google 去融资。
You know, they they've got an ungodly amount of SpaceX, right?
你懂的,他们持有多得离谱的 SpaceX 股份,对吧?
They own like 5% of the company.
他们大概持有这家公司 5% 的股份。
I think a little more, but yeah.
我觉得还要多一点,不过对。
Yeah, maybe.
对,也许吧。
I think at one point they had like 10%.
我记得他们一度大概有 10%。
Larry Page invested a billion dollars at a $10 billion valuation, got 10% of the company, it got diluted, like all this.
Larry Page 在 100 亿美元估值时投了 10 亿美元,拿到了公司 10% 的股份,后来被稀释了,诸如此类。
But that was one of the greatest investments of all time.
但那是史上最伟大的投资之一。
Good job, Larry.
干得漂亮,Larry。
The guy.
这哥们儿。
So, they know they have like a hundred billion dollars in the bank that they can sell in, you know, nine months or whatever from the lockup.
所以,他们知道自己有大概一千亿美元躺在账上,可以在,你懂的,锁定期结束后九个月左右卖掉。
And they have all the gross profit they do, and yet they still modeled that.
而且他们赚着他们赚的那么多毛利,结果他们还是测算出了那个结论。
and they were like we need to raise capital and so they did an offering and it's like that's insane.
然后他们就觉得"我们需要融资",于是就做了一次发行,这简直是,那太疯狂了。
So that tells you how much they think they need to spend.
所以这就告诉你,他们觉得自己需要花多少钱。
But capital is like really, you know, you know, Meta's do Meta did announce that they're going to do a raise.
但资本真的是,你懂的,你懂的,Meta 也——Meta 确实宣布了他们要做一轮融资。
Stock tanked.
股价暴跌。
People don't like it, but you know, that's all these companies are going to raise capital, whether it be debt or equity.
大家不喜欢,但你懂的,这些公司都会去融资,不管是债权还是股权。
At some point, money spigots will have to, you know, slow down.
到某个时候,资金的水龙头总得,你懂的,拧小一点。
But right now, every GPU that Amazon adds, they're making higher revenue or every TPU or Trainium, you know, whoever anyone adds is is making is making gross profit.
但眼下,Amazon 每加一块 GPU,他们就多赚一份营收,或者每一块 TPU、每一块 Trainium,你懂的,不管谁加,都在赚,都在赚毛利。
I do a little bit of a tee up on this to turn into a question for you.
我在这上面稍微铺垫一下,好把它变成一个问给你的问题。
But like
就是——
as we talk about this for me, the thing that's going my through my head that's that is almost an alternative hypothesis for like the Crusoe example.
我们聊到这个,我脑子里一直转的,几乎是对 Crusoe 这个例子的一个另类假设。
I'm going use an analogy in oil like in oil Saudi Arabia has way lower cost per barrel to produce oil than a lot of other countries.
我打个石油的比方,就像在石油里,沙特阿拉伯每桶的开采成本比很多国家低太多了。
There's also like the purity of the oil.
还有原油的纯度问题。
A lot of you know Saudi has generally like very low contaminants in their oil which makes refining easier all of this.
你知道,沙特的原油杂质普遍很低,炼化起来更容易,诸如此类。
The question for me is like when you look at for every gigawatt that's being put in the ground if call it the 20 gigawatts coming online today like how much like how much homogeneity do you see in those gigawatts?
我的问题是,当你看每一个正在落地的 gigawatt,就说今天上线的这 20 gigawatts,这些 gigawatt 之间你觉得有多同质?
Is it something like and I don't you can tell me whatever metric you think is right but like are Google's gigawatts two times more valuable than say most Neoclouds because they have optical switches and they have like they've been doing it for a long time and like they know how to do power smoothing because I think this could be the alternative hypothesis
是不是那种——你用你觉得对的任何指标来说都行——Google 的 gigawatt 是不是比大多数 neocloud 值钱两倍,因为他们有光交换机,他们做这个做了很久,懂得怎么做功率平滑,我觉得这可能就是那个另类假设。
that some of the people that are it's like the people that are good at at building data centers they they should just do it to the max because there's so much demand and there's so much better than it, but then maybe we're starting to see the early signs of the people that are like not as good at it kind of getting hit a little.
就是说那些擅长建数据中心的人,他们就该顶格猛干,因为需求这么大、他们又做得这么好,但也许我们已经开始看到早期迹象:那些做得没那么好的人开始有点挨打了。
So I like I don't know the reality here.
所以我其实不知道真实情况到底是怎样。
I'm just curious how you think about this.
我只是好奇你是怎么看这件事的。
So so far um there there are metrics for this, right?
呃,目前是有一些指标能衡量这个的,对吧?
So uh Trainium sells at sub-10 billion per gigawatt rental rate uh to Anthropic and to OpenAI.
呃,Trainium 卖给 Anthropic 和 OpenAI 的租金是每 gigawatt 不到 100 亿美元。
GPUs at least before the craziness of the last six months usually went around 12 to $13 billion per gigawatt.
GPU,至少在过去六个月疯狂涨价之前,通常是每 gigawatt 120 到 130 亿美元左右。
So the rental rate and this is from a neocloud versus Amazon even and now when Amazon sells GPUs they'd also be 13 or so
所以这个租金——这还是拿 neocloud 跟 Amazon 比——现在 Amazon 卖 GPU 也是 130 亿左右。
and my understanding of that also is that those number like Amazon subsidized that a little bit so that it's like I actually think the numbers were even like I think the disparity was even more
而且据我了解,那些数字里 Amazon 其实还补贴了一点,所以我觉得实际的数字——我觉得这个差距其实还要更大。
it's less than 10.
是不到 100 亿。
It's less than 10 but there's like some weird basically
不到 100 亿,但基本上里面有些奇怪的门道。
and like look I my understanding obvious like Anthropic played a big role in making Trainium useful in terms of you know writing all the libraries etc and and so like I
而且你看,据我了解,很明显 Anthropic 在把 Trainium 变得好用这件事上出了大力,你知道,写各种库之类的,所以我——
everything I hear is that Trainium's really freaking good hardware and it's getting way like way better and obviously Anthropic now using it a lot so hopefully we would see that price go up
我听到的所有说法都是,Trainium 是真他妈好的硬件,而且还在变得越来越好,现在 Anthropic 显然用得很多,所以但愿我们会看到那个价格涨上去。
you know like per the the deal they did was actually like there was a floor mechanism them and like it if it didn't do well it would be like cheaper and then to the point where it's cancelceable and you know if it if it did really well the price is kind of higher
你知道,他们签的那笔交易其实有个保底机制,就是如果它表现不好,价格就会更便宜,便宜到一个可以取消的地步,而如果它表现真的好,价格就会高一些。
um but effectively um less than 10 right is is where Trainium shakes out at whereas GPUs I mean this the SpaceX deal again was like 25 or something crazy billion dollars per gigawatt or $25 million per megawatt right a year rental rate with Google I was like that's a crazy divergence
呃,但实际上 Trainium 最后大概落在不到 100 亿,而 GPU——我是说 SpaceX 那笔交易——又是每 gigawatt 250 亿还是多少那种疯狂数字,也就是每 megawatt 每年 2500 万美元的租金,是跟 Google 签的,我当时就觉得这差距太夸张了。
now obviously if if if Amazon was selling Trainium today it' probably be more expensive than 10 because the comput shortages,
当然,如果 Amazon 今天卖 Trainium,价格可能会高于 100 亿,因为算力紧缺。
but you you do see this already in the sense of uh with data centers oftent times a rental price of a data center if you're doing collocation, right? Not compute in there, but just power.
但你其实已经能看到这一点了,呃,就数据中心而言,数据中心的租金——如果你做的是托管(collocation),对吧?里面不含算力,只卖电力。
Here's the data center. Um you you price it generally on a uh dollars per kilowatt per month.
就是"这是数据中心",呃,你一般是按每千瓦每月多少美元来定价。
And so they used to be $60 per kilowatt hour per month, and now you see things transacting at anywhere from like 120 to 160.
以前是每千瓦时每月 60 美元,现在你看到成交价从 120 到 160 都有。
Um but different quality data centers, this actually you've I've seen data centers go as high as 200. um when the customer is not such a great credit rating and then the data center is a pretty good one.
呃,但不同品质的数据中心——我其实见过有的高到 200,呃,那是客户信用评级不太好、而数据中心又相当不错的情况。
And I've seen stuff go as low as 100 still or in like India go like as low as 80 because the grid's not reliable, the internet connection's not great and it's a pretty mid data center but at least it's a data center.
我也见过低到 100 的,或者像在印度低到 80 的,因为电网不可靠、网络连接不太好、数据中心也就中等水平,但至少它是个数据中心。
Um and so you you see this huge discrepancy there already.
呃,所以你已经能看到那里存在巨大的价差了。
Um in the case of like data center construction, usually the pitfalls they just fail.
呃,就数据中心建设来说,通常的坑就是——他们直接失败了。
There's a lot of people who fail, you know, claim they're g they're like they're like four guys they they're like, "Yeah, we here I bought some turbines. I put the money down for them. I'm gonna build a data center."
有一大堆人会失败,你知道,他们号称——就四个人——他们说:"对,我们买了几台涡轮机,我付了定金,我要建一个数据中心。"
And then they get delayed, delayed, delayed, and fail.
然后就是一拖再拖再拖,最后失败。
Um, so you have to like probability, weight, time, weight, time lag, the teams that suck versus don't.
呃,所以你得给它们做概率加权、时间加权,算上时间滞后,区分烂团队和不烂的团队。
Um, and and sort of, you know, our data center model does that.
呃,某种程度上,你知道,我们的数据中心模型就是干这个的。
We kind of track every data center uh and try and do this for every single one based on, you know, equipment that they're using and all these things.
我们基本上会追踪每一个数据中心,呃,并根据他们用的设备等等这些东西,对每一个单独做这套评估。
One of the things you mentioned about Google is you know in a gigawatt data center they'll actually put like 1.5 gigawatts of hardware and because they have such understanding all the way from workload to u you know they're able to slosh the power around
你刚提到 Google 的一点是,你知道,在一个 gigawatt 的数据中心里,他们其实会塞进大概 1.5 gigawatt 的硬件,因为他们对从工作负载往下的一切都理解得很透,呃,你知道,他们能把功率来回腾挪。
and so instead of you know constantly you know a gigawatt of compute which typically runs at like 60 or 70% utilization in terms of power consumption not utilization of the hardware someone's always renting it um they're now running it at like you know you know that 60 to 70% means it's at a gigawatt and you're using the full gigawatt
所以,不是像通常那样,一 gigawatt 的算力在功耗上一般只跑到 60% 或 70% 的利用率——注意是功耗利用率,不是硬件利用率,硬件总有人在租——呃,他们现在跑成这样:那 60% 到 70% 就意味着它已经到一 gigawatt 了,而你用满了整整一 gigawatt。
um you see people doing deals with including Google with utilities where they're like, "Oh, well, I know this grid can sustainably take a gigawatt, but you know, except for three days of the year, you can actually do two gigawatts, so give me two gigawatts and then just tell me to turn off."
呃,你会看到有人——包括 Google——跟电力公司谈这种交易:"哦,我知道这个电网能持续承受一 gigawatt,但你知道,除了一年里那三天,其实能上到两 gigawatt,所以给我两 gigawatt,到时候你让我关掉就行。"
And so they'll do that.
于是他们就这么干。
And so these sorts of tricks and then you need to have supreme management of workload, backup power, all these things, um, generators on site to figure out how to actually keep it 2 gigawatts sustainably.
所以就是这类手段,然后你还得有一流的工作负载管理、备用电源这一切,呃,现场的发电机,才能搞清楚怎么把它可持续地维持在两 gigawatt。
When people do this, they're able to charge more.
当有人能做到这一点,他们就能收更多钱。
Whether it be I'm actually selling two gigawatts despite only having one gigawatt because those three days I'll be able to deal with via battery, gas, etc. or I figured out how to build power on site.
不管是"我明明只有一 gigawatt 却在卖两 gigawatt,因为那三天我能靠电池、燃气等等扛过去",还是"我搞清楚了怎么在现场自建电力"。
Now I have a gigawatt where no one else does and so I'm able to do it quickly.
现在我在别人没有的地方有了一 gigawatt,所以我能很快搞定。
Um it's not necessarily transacting for a higher price. It's that I'm selling more gigawatts.
呃,这未必是成交单价更高,而是我卖出了更多 gigawatt。
And sometimes there are levers where you're selling more gigawatts is where where each gigawatt is selling at a different price.
而有时候会有这样的杠杆:你卖更多 gigawatt 的同时,每一 gigawatt 又是按不同的价格在卖。
Um I think it's more on the data center and energy layer. It's more about just having it versus not and then that being delayed or not. It's more binary.
呃,我觉得在数据中心和能源这一层,更多是"有没有"以及"延不延期"的问题,更像是二元的。
But on the compute side, I do think there's a lot more um interesting work there. Right?
但在算力这一侧,我确实觉得那里有多得多的、呃,有意思的门道,对吧?
A gigawatt given to Anthropic is objectively worth more revenue than a gigawatt given to OpenAI.
同样一 gigawatt,给 Anthropic 客观上能产生比给 OpenAI 更多的营收。
And it seems that both of them could sell every gigawatt that they have right now.
而且看起来他们俩现在手上每一 gigawatt 都能卖得出去。
Uh given rate limit problems and token max limit and all these sorts of things at OpenAI and Anthropic.
呃,考虑到 OpenAI 和 Anthropic 那边的 rate limit 问题、token 上限之类的这些事。
Uh especially since Codex 5.5 came out, it's much better.
呃,尤其是 Codex 5.5 出来之后,它好多了。
And then likewise, if you gave a gigawatt to SpaceX, you know, they turn
同样地,如果你给 SpaceX 一 gigawatt,你知道,他们会——
my my guess, like my suspicion is that they're they probably make better use of the, you know, hardware than most people.
我的猜测,或者说我的直觉是,他们对硬件的利用大概比大多数人都更好。
Um, just like I think people underestimate how much networking experience they have from Starlink in particular and also how much just like power management experience they have via from Tesla.
呃,就像我觉得人们低估了他们从 Starlink——尤其是 Starlink——积累的网络经验,也低估了他们通过 Tesla 积累的电源管理经验。
Yeah. people like Brett Mayo are like incredible like
对。像 Brett Mayo 那样的人简直厉害得不行,就是——
pretty good.
相当强。
And so I I think that like for me that's actually I think probably the thing that might I don't actually know the answer but I think that might be missing from the analysis a lot of people are doing.
所以我觉得,对我来说,这其实可能就是——我并不真的知道答案——但我觉得这可能正是很多人在做的分析里漏掉的那块。
I think I think it's also the fact that when CoreWeave builds a gigawatt even though their GPU compute is objectively better than Amazon or Google or Microsoft's in terms of performance.
我觉得还有一点是,当 CoreWeave 建一 gigawatt 的时候,尽管他们的 GPU 算力在性能上客观来说比 Amazon、Google 或 Microsoft 的都要好。
We've tested the performance and reliability.
我们测过性能和可靠性。
Um, the problem is Google sells it six months before they have it up and they need to turn around and take that paper that they signed to get debt uh with that credit backing and then turn around so they can actually pay for the PO that they've already issued you know for the order that they've already issued.
呃,问题是 Google 在东西上线前六个月就把它卖掉了,然后他们得拿着签好的那份合约去,呃,靠那个信用背书借债,再回头去付他们已经开出的采购单——你知道,就是他们已经下的那笔订单。
Whereas SpaceX was like no no no this is running now buy it right and it's it's a big discrepancy when you have a balance sheet to do that versus not and that also helps your revenue per megawatt like be much higher.
而 SpaceX 是那种"不不不,这个现在就在跑,买吧",对吧,你有没有那样一张能这么干的资产负债表,差别很大,而这也让你每 megawatt 的营收高得多。
Why does the Neocloud opportunity even exist?
neocloud 这个机会为什么会存在?
Because if you had asked me five years ago, I would have said the hyperscalers are going to own this.
因为要是你五年前问我,我会说这块会被 hyperscaler 全占了。
And you know, you mentioned just now core weight has better performance than than the hyperscalers.
而你知道,你刚才提到 CoreWeave 的性能比 hyperscaler 还好。
Like what why does this opportunity exist maybe at the macro level and then in the execution level?
那到底为什么会有这个机会——也许先从宏观层面讲,再讲执行层面?
Yeah. So in 2023 I wrote a report that had uh Amazon really hate me.
好。2023 年我写过一份报告,呃,让 Amazon 特别恨我。
Um it was called Amazon cloud crisis.
呃,那份报告叫《Amazon cloud crisis》(亚马逊云危机)。
So I talked about how Amazon was the best cloud because they had their Nitro NICs which offered like tenant isolation. all the hypervisor ran on the nick and then you could sell all the cores and they had you know custom SSDs that they made and they'd buy the raw NAND and they'd have lower cost because they'd buy the raw nand and build their own SSDs
我在里面讲了为什么 Amazon 是最好的云:因为他们有自己的 Nitro NIC,能做租户隔离,整个 hypervisor 都跑在网卡上,这样你就能把所有核心都卖出去;他们还有自己造的定制 SSD,他们买原始 NAND 颗粒,成本更低,因为他们买原始 NAND 自己造 SSD。
um and you know they had their custom Graviton CPUs and that drove down cost for for per core and so they had all these things that enabled them to sell more cores have better security good networking for but this was all for the traditional CPU better storage for the traditional you know cloud world
呃,你知道,他们还有自己定制的 Graviton CPU,把每核成本压了下来,所以他们有这一整套东西,让他们能卖更多核心、有更好的安全性、好的网络——但这全是为传统 CPU、为传统的、你知道的那种云世界准备的更好的存储。
but in the AI cloud a lot of this stuff hurt performance right these Nitro NICs bad for performance, still are worse performance.
但在 AI 云里,这里面很多东西反而拖累性能,对吧,这些 Nitro NIC 对性能有害,现在依然性能更差。
Although they've caught up a lot because they've had a couple iterations to like, you know, improve them, but they're still worse for performance.
虽然他们追上来不少了,因为迭代了好几代,你知道,把它们改进了,但性能上还是更差。
Um, a lot of the security stuff doesn't matter because it's not like I'm time splicing users or splicing a socket into many users, right?
呃,很多安全方面的东西根本不重要,因为又不是说我在对用户做时间分片,或者把一个插槽切给很多用户,对吧?
It's like no one buy rents a single GPU and an 8GPU server.
就好比没人会在一台 8 卡 GPU 服务器里去租单张 GPU。
No one rents a single GPU in a 72GPU rack.
没人会在一个 72 卡 GPU 的机架里去租单张 GPU。
They rent the whole rack and in fact, they rent many of the racks.
他们租的是整个机架,实际上还是很多个机架一起租。
And so, and then and then there's no like, oh, I rent for six hours and I give it back. it's everyone has these long-term contracts.
所以,也不存在那种"哦,我租六小时然后还回去",每个人签的都是长期合约。
So, the mechanics of the GPU rental market meant that a lot of the expertise of the hyperscalers fell away.
所以,GPU 租赁市场的运作机制,意味着 hyperscaler 的很多专长都失效了。
Um, and a lot of the expertise that they did have were actually some of them were detrimental, right?
呃,而他们确实有的很多专长里,有一些其实是有害的,对吧?
Network performance for Google and Amazon. It was they had custom networks that were better for traditional CPU and for the stuff that they were doing, but actually worked for AI.
比如 Google 和 Amazon 的网络性能。他们有为传统 CPU、为他们当时在做的那些事优化得更好的定制网络,但对 AI 却[反而不适用]。
Um and then in other cases it's like well you know Microsoft would save money by building their own data centers but their data center teams are actually were not actually that great and so when it came time to run you know when it was predictable building it was like fine when it came time to like actually double your forecast for the year it's like they fell on their face and they had to go get a bunch of NeoCloud capacity.
呃,在另一些情况里,就好比 Microsoft 想靠自建数据中心省钱,但他们的数据中心团队其实没那么强,所以到了真要跑的时候——你知道,在可预测、按部就班地建的时候还好,一到要把年度预测翻倍的时候,他们就栽了跟头,不得不去搞一大堆 neocloud 的产能。
I think so performance I think you know I think time to market's another one right ne you know these massive organizations no one's getting rich from building this data center faster right
所以我觉得性能是一方面,我觉得,你知道,上市速度是另一方面,对吧,在这些庞大的组织里,没人会因为把这个数据中心建得更快而发财,对吧。
but you look at Crusoe for example Chase and and and and all the other people at the team you know I was going to name some people at the team I'd rather not you know these all these people are getting rich if they deliver these this compute faster they're they're you know they're lever they're hyperlevered equity owners
但你看 Crusoe 就是个例子,Chase 还有团队里所有其他人,你知道,我本来想点几个团队里的人的名,还是算了吧,你知道,这些人只要把这批算力交付得更快就能发财,他们,你知道,他们是加了杠杆的、是极度加杠杆的股权持有者。
hey look they're also all coming from Bitcoin And and you know they you're not supposed to say that.
嘿你看,他们其实也都是从 Bitcoin 那边过来的,而且,你知道,这话是不该说出来的。
Uh I mean a lot of the data center like their main data center guy came from Microsoft.
呃,我是说,很多数据中心方面——他们主管数据中心的那个人是从 Microsoft 来的。
I don't know. I'm just I'm just teasing.
我也不知道。我就是逗你玩的。
But it's uh you know it's like you learn you learn a lot when you're in a very high fluctuation you know market.
但你知道,呃,就是当你身处一个波动极大的、你知道的那种市场时,你会学到很多。
What do you think was Jensen playing 4D chess?
你觉得 Jensen 是在下一盘什么样的 4D 大棋?
Jensen absolutely hates a world where all the hyperscalers have all the power.
Jensen 绝对讨厌一个所有权力都攥在 hyperscaler 手里的世界。
There's a reason he's like blowing money on like random AI labs that like I don't even know if like it makes sense to but like you know he's blowing money and pumping them up and going to you know everyone around the world and saying you should invest in this company because he wants to create a multipolar world.
他往一些随便什么 AI 实验室里砸钱是有原因的,那些实验室我甚至都不确定砸进去合不合理,但你知道,他就是在砸钱、在把它们捧起来,跑去跟全世界的人说"你应该投这家公司",因为他想造出一个多极化的世界。
That's why he loves Chinese labs because he wants to create a multipolar world.
这就是为什么他偏爱中国的实验室,因为他想造出一个多极化的世界。
A world where open Anthropic and Google models are the only models is one in which he's screwed.
一个只有 OpenAI、Anthropic 和 Google 的模型的世界,是一个让他完蛋的世界。
Yep.
没错。
Right. Um a world in which you know the hyperscalers are the only ones building compute is one he's screwed in. Yeah.
对。呃,一个只有 hyperscaler 在建算力的世界,同样是个让他完蛋的世界。是的。
And so, you know, of course he needs to point the allocation gun at NeoClouds, help back stop their clusters, do anything and everything because while today a GPU sold to Crusoe and a GPU sold to um CoreWeave and a GPU sold to Google and Amazon are all the same price for him,
所以,你知道,他当然得把"分配权这把枪"对准 neocloud,帮它们的集群兜底,想尽一切办法,因为虽然今天卖给 Crusoe 一张 GPU、呃,卖给 CoreWeave 一张 GPU、卖给 Google 和 Amazon 一张 GPU,对他来说都是同一个价,
five years from now Crusoe and CoreWeave existing means Google TPU will be weaker and means Amazon Trainium will be weaker and more inference being done with you know non-clos model labs is is better for firm.
但五年后,Crusoe 和 CoreWeave 的存在意味着 Google TPU 会更弱、意味着 Amazon Trainium 会更弱,而更多推理由那些、你知道的、非闭源大厂的模型实验室来跑,对他更有利。
So I think you know the neocloud ecosystem is you know it's these people that are wild west these neolabs as well a lot of them have investments from Nvidia it's the wild west some will fail many will fail but you know some will emerge as really great teams
所以我觉得,你知道,neocloud 这个生态,你知道,就是这帮野蛮生长的人,这些 neolab 也一样,其中很多都拿了 Nvidia 的投资,这就是一片蛮荒西部,有些会失败,很多都会失败,但你知道,总有一些会冒出来成为真正出色的团队。
whether it be you know oddly Crusoe who's a bunch of crypto guys who then started building data centers and doing flared gas stuff or you know CoreWeave who initially was a bunch of New York hedge fun guys
不管是,你知道,说来奇怪的 Crusoe,一帮搞加密货币的人,后来开始建数据中心、搞伴生气(flared gas)那套东西,还是,你知道,CoreWeave,一开始是一帮纽约对冲基金的人。
they were also and then they were doing
他们也是,然后他们在做——
and crypto guys but then they they they like built you know there were a lot of people who didn't bubble up like them started around the same time just failed Right.
也是搞加密货币的,但后来他们、他们就、你知道,建了起来,当时有很多没能像他们那样冒出来的人,差不多同时起步,结果就是失败了。对吧。
And so I think you know um
所以我觉得,你知道,呃——
I gotta say both those teams are phenomenal.
我得说,这两个团队都非常了不起。
They deserve a lot of credit and it's like that's your point. But
他们值得很多称赞,而这也正是你的观点。不过——
yeah, I mean my point is like he he you know you throw it's like you throw a bunch of like bait into the water and the best fish will figure out and survive, right?
对,我是说,我的意思是,他、他,你知道,你撒下去——就像你往水里撒一大把饵,最会来事的那些鱼会想明白怎么活下来,对吧?
Um and and sort of the same way with the Neoclouds and and and he hopes the Neols as well.
呃,neocloud 大体也是这个路数,而且、而且他也希望 neolab 也能这样。
We'll see if any of the Neolabs really bubble up, but like you know Thinking Machines has a few hundred million dollars of ARR, right?
neolab 里到底有没有哪家能真正冒出来,还得看,但你知道,Thinking Machines 有好几亿美元的 ARR,对吧?
That's pretty impressive even though they've had you know in the media it's like oh they've lost all this talent.
这挺厉害的,尽管他们经历了,你知道,媒体上说的那种"哦,他们流失了这么多人才"。
It's like, well, but Tinker is doing a few hundred million of ARR.
但问题是,Tinker 正在做出好几亿美元的 ARR。
Like, that's pretty impressive for out of the gate a product that's less than six months old or whatever.
对一个刚上线、还不到六个月还是多久的产品来说,这相当厉害了。
Um, and and you, you know, we hope the same happens to other Neolabs and and so um you know, he wants a multipolar world.
呃,而且,你知道,我们希望其他 neolab 也能有同样的结果,所以,呃,你知道,他想要一个多极化的世界。
Truly, congratulations on the success.
真心的,恭喜你取得这样的成功。
Thank you.
谢谢。
Just the last thing I'll say is I've seen a little bit of this.
我最后想说的一点是,这些我算是亲眼见过一点。
I think the public, they can probably tell from listening to you how hard you work, but like it's clear you've just been working your ass off for more than a decade
我觉得公众大概光听你说话就能感觉到你有多拼,但很明显,你已经他妈拼命干了十几年。
and it, you know, led to the last few years of being in the right place, right time.
而这,你知道,才带来了过去这几年天时地利的局面。
But like it's unbelievable what you've accomplished and I know it's just the beginning. So,
但你所取得的成就真是难以置信,而且我知道这才只是开始。所以——
thank you so much.
太感谢了。
Thank you for doing this.
谢谢你来上节目。
Awesome.
太棒了。