"The next generation of models need more than your laptop." —— Tibo Sottiaux 是 OpenAI 核心产品负责人,任务是把 ChatGPT 和 Codex 合成一个东西。他从 DeepMind 那段没能发出去的 LMChat 讲起,一路讲到 harness 还剩哪些笨拙、为什么笔记本电脑本身成了下一代 agent 的天花板、递归自我改进其实先发生在基础设施上,以及那个没有审批流、他想按就按的额度重置按钮。
重置按钮背后没有审批流
"It's not done in partnership with like marketing or finance. It's just like I can press the button whenever I want, whenever it feels right."
"There's a whole reset button now and like but there isn't really a whole a lot of scrutiny behind it."
Codex 的额度重置不是增长手段,是一个产品负责人凭手感按下去的补偿开关。
笔记本电脑本身成了瓶颈
"Your laptop kind of becomes a constraint in and by itself."
"The amount of work that you can do on a laptop it was designed for humans right so it's designed roughly to be able to absorb the amount of work that you know you can produce."
笔记本的规格是按人的打字和思考速度定的,模型没有这些上限——所以下一代 agent 必然往云上走。
是模型要求两个产品合并
"The future the our future models want us to be merged."
"So the feedback we had initially was like why why do you merge them? It's like you know do you really have to do it?"
ChatGPT 与 Codex 合并不是产品线整合,而是同一套 harness、同一个模型倒逼出来的结果。
你妈和你用同一个界面
"You and your mom will you know use the same thing. Uh it will be your personal AGI."
"It's kind of wild to think that my mom might use the same exact interface as me and then obviously it'll customize to my needs."
「程序员 / 非程序员」在他看来只是人为标签,界面该自己贴着个体走,而不是先让人选身份。
递归自我改进先发生在基建
"Using those models to develop the infrastructure that is on the critical path of using those models... It's all one big system."
"Recursive self-improvement is obviously a huge topic right now and it's most often I think applied to research and models developing other models."
大家盯着「模型训模型」,真正在跑的是模型重写推理栈和 CUDA kernel,再把提速回灌给自己。
ultra fast 的提速并不均匀
"If it's a lot of tool calls... then you know you'll only feel like a 3x or 4x speed up."
"It works amazingly well when there's not that many tool calls involved or it's like a lot of generation of context."
14 倍是生成密集场景的上限;工具调用一多,瓶颈就挪到网络和 agent 轨迹上去了。
够快之后就不想并行了
"What I was doing before, multitasking like 10 agents, it's like I don't want to really go back to that."
"Suddenly you're like okay you know like this thing can operate at the same speed if not faster than you and so you stay in the flow."
速度不只是省时间,它把「同时盯十个 agent」这份认知负担直接取消掉。
常规档三个月快了 60%
"The amount of speed that you get now is like, you know, roughly 60% faster than, you know, what it used to be like 3 months ago."
"We haven't just improved the the cost efficiency, but we have also improved the speed efficiency. You know, outside of ultra fast, things have gotten significantly faster over time."
这是 ultra fast 之外的普通档提速,来自把整个技术栈按自家负载重做一遍。
效率收益不许揣进兜里
"Not just pocket you know that uh interesting gain and just make it something that we know we share with our customers."
"Our commitment is to just really to keep things at the frontier of performance cost."
Luna 降价 80% 被他讲成一条纪律:算力省下来必须还给用户,而不是变成毛利。
暂停训练被讲成常规能力
"The pause was sort of like necessary to allow like the teams and the individuals like just really understand and harden all parts of the system."
"As we increase the capabilities of our models, it is very obvious that you know the the alignment and the safety aspect of it is you know ever more important."
他把暂停前沿 RL 说成一次加固窗口而非危机处置——「我没见过我们内部在这类决定上做不到高效」。
Google 的 ChatGPT 早了一年
"And this was it was like a year before ChatGPT roughly."
"And so then naturally like the idea of like something like LMChat sort of emerges um and then it was internal and then there was sort of this ambition to um make it into a publicly available uh tool."
LMChat 在 DeepMind 内部已经能聊了,卡住它的不是技术,是「不能颠覆搜索」这条组织约束。
不盯对手,只盯自己独有的
"I don't tend to look at the competition that much. Like I really look at, you know, what can we do uniquely well and what are our values."
"There was a period of time in which Anthropic was kind of sucking all the oxygen out of the room, right? They were really dominating and then all of a sudden something changed."
被追问两次 Anthropic,他两次把问题拧回同一句话——这几乎是他对竞争议题的全部回答。
I can press the button whenever I want, whenever it feels right.
这个按钮我想什么时候按就什么时候按,觉得该按了就按。
I don't tend to look at the competition that much.
我不太去盯竞争对手。
Like I really look at, you know, what can we do uniquely well and what are our values and like, you know, how do we maximally accelerate towards that?
我真正盯的是:哪些事只有我们能做得特别好、我们的价值观是什么、怎么才能以最大速度朝那个方向冲?
You know, maybe a year or two these speeds will become, you know, maybe if not the default, like very close to the default.
也许一两年之后,这种速度就算不是默认档,也会非常接近默认档。
And you look at the cost of Luna, right? It's phenomenal.
你再看看 Luna 的成本,对吧?那是惊人的。
Technology has a way to become like, you know, very, very efficient over time.
技术总有办法随着时间变得极其高效。
We're very focused on like, you know, very broad access and we're optimizing for, you know, the utility that you get out of it directly.
我们非常在意的是尽可能广的可及性,我们优化的是你从中直接得到的效用。
I heard there's an actual physical button now.
我听说现在真有一个实体按钮了。
Yes, there is.
是的,真有。
I will show it to you. It's like very very cool.
回头我给你看,特别酷。
Tibo, thank you so much for joining me.
Tibo,非常感谢你来。
Of course. Yeah, glad to be here.
客气了,我很高兴来这儿。
Yes, really excited to talk to you.
我也特别期待跟你聊。
I want to actually start with your time at Google.
我想先从你在 Google 的那段时间说起。
So you were on the DeepMind team and before ChatGPT Google had something called LMChat and you had tweeted uh Google was too nervous to release it.
你当时在 DeepMind 团队,在 ChatGPT 出现之前,Google 就有一个叫 LMChat 的东西,你发过推说 Google 太紧张了,不敢放出来。
DeepMind was blocked from shipping products that could disrupt Google and I think about that a lot.
DeepMind 被禁止发布任何可能颠覆 Google 的产品——这件事我常常想起。
What were you thinking at that time while you were working on these products that was you know well before uh ChatGPT really changed the world.
那个时候你在做这些产品,而那是在 ChatGPT 真正改变世界之前很久,你当时是怎么想的?
Yeah it was a very exciting time.
那是一段非常激动人心的时期。
So um DeepMind was a very creative place.
DeepMind 是一个非常有创造力的地方。
I was mostly focused on um my specialty was infrastructure and products for to accelerate research.
我主要做的、我的专长是基础设施和用来加速研究的产品。
And so there was there was obviously a group working on language models uh and scaling that.
当时显然有一组人在做语言模型和 scaling。
And then they had like gotten pretty good results and then it was like a natural thing to think about hey you know can you turn that into something that you know you can chat to and you know can use um you know for various things.
他们已经拿到了相当不错的结果,接下来很自然就会想:能不能把它变成一个你可以对话、可以拿来做各种事的东西?
And so then naturally like the idea of like something like LMChat sort of emerges um and then it was internal and then there was sort of this ambition to um make it into a publicly available uh tool.
于是像 LMChat 这样的想法就自然浮现出来,一开始是内部用的,后来就有了把它做成一个公开可用工具的野心。
What year was this?
这是哪一年?
And this was it was like a year before ChatGPT roughly.
大概是 ChatGPT 出来的前一年。
Okay.
好。
Yeah.
嗯。
Um and then but we were also building all sorts of other things that I'm not going to talk about but it was a very creative place.
我们当时还在做很多别的东西,那些我就不讲了,总之那是个非常有创造力的地方。
Uh and then just DeepMind was not set up to ship product.
但 DeepMind 的组织结构本来就不是为发产品设计的。
OpenAI is a very very different place in that sense.
从这个意义上说,OpenAI 是一个非常非常不一样的地方。
Um we just like research and products just collaborate super closely together.
我们这里研究和产品是紧密协作的。
We ideate together. We co-design a lot of things.
我们一起出点子,很多东西是共同设计的。
We have a big bias to ship uh and uh also a big bias towards you know making things available for people which I really love and this is sort of like what uh drove me here the mission the people um the talent density.
我们有很强的发布倾向,也有很强的「让人们用得上」的倾向——这一点我特别喜欢,也正是把我吸引过来的原因:使命、人,还有人才密度。
I mean there's so many great things about OpenAI really.
OpenAI 真的有太多值得说的好东西了。
Did you know at the time where you were involved in LMChat that that was something special or would become something special?
你当时参与 LMChat 的时候,知道这是个特别的东西、或者会变成特别的东西吗?
It it felt it felt very special like the models were you know it's sort of like the first time um you know you realized like you could get coherent text and you know something helpful.
感觉非常特别,因为那是你第一次意识到:模型能生成连贯的文本,而且是有用的东西。
Initially it was more funny than helpful and then gradually it became more and more helpful.
一开始它更多是好笑而不是有用,然后逐渐越来越有用。
So you say you think about that often and I I understand that I think you know in a lot of ways Google got in their own way.
你说你常常想起这件事,我能理解——在很多方面 Google 是自己挡了自己的路。
Um what are some of the lessons that you learned there that you took to OpenAI?
你从那段经历里学到了哪些东西,带到了 OpenAI?
Yeah, so this is why I think about it often.
对,所以我才会常常想起它。
I I think about it in terms of like the culture that I have on the team uh the culture of OpenAI itself and so like the good parts to preserve and like you know what not to do.
我想的是我自己团队的文化、OpenAI 本身的文化:哪些好的部分要保住,哪些事绝对不能做。
Um, OpenAI has a very bottoms up uh culture.
OpenAI 是一种非常自下而上的文化。
Like it's a very empowering culture.
这是一种给人授权的文化。
Like people can come up with all sorts of ideas and get together and then very very quickly ship something and there's very little stop energy uh in general for like new product ideas which is exhilarating and fun and you know it's all about impacting the world in positive ways.
任何人都可以冒出各种想法,凑到一起,很快就把东西发出去;对新的产品点子,基本没有什么踩刹车的力量——这件事既刺激又好玩,而且核心都是用正面的方式影响世界。
Uh and so preserving that is very important to me.
所以把这一点保住,对我非常重要。
The other thing that is also important is like to not make it a mess, right?
另一件同样重要的事是:别把它做成一团乱麻。
So you don't want to have like a hodgepodge of like no features and like no overall direction and coherence and so it's counterbalanced with the sense of um simplicity and you being proud about the quality of the product.
你不希望最后是一堆零碎功能拼在一起,没有整体方向、没有连贯性;所以要用另一头来平衡——简洁,以及你对产品质量本身的骄傲。
Uh I think the ChatGPT iOS app is like you know one of the best apps out there.
我觉得 ChatGPT 的 iOS app 是市面上最好的应用之一。
Uh and so we want we want to keep that.
这一点我们想保住。
Uh we're investing a lot in you know things like delight performance efficiency simplicity.
我们在愉悦感、性能、效率、简洁这些事上投入很大。
So there's these overall principles while still empowering everyone everyone to like try new things and ship very quickly.
所以是有这些整体原则的,同时仍然让每个人都能去试新东西、快速发布。
If you were to give advice to a founder about how to develop that kind of culture uh like what are some of the more tangible elements or practices that occur inside OpenAI that can kind of gives give advice uh to a founder?
如果让你给创业者一些关于怎么养出这种文化的建议,OpenAI 内部有哪些更具体的做法或习惯,是可以给创业者参考的?
Yes, I think having having conviction um and finding a way to be to have users and iterate very quickly from feedback and then also uh being willing to disrupt yourself that is not as much relevant for founder but is relevant for you know companies like OpenAI like we come up with new research new ideas all the time and being able to identify when is the right moment to go and invest in them even though it means like maybe reallocating resources from you know the main gig uh super important but it's very hard but it's super important to be able to do that.
我觉得第一是要有 conviction;第二是想办法拿到用户、从反馈里飞快迭代;第三是愿意颠覆自己——这一条对创业者也许没那么切题,但对 OpenAI 这样的公司非常切题:我们一直在产出新研究、新想法,能判断出什么时候是该投入进去的正确时机,哪怕这意味着要从主业里抽调资源——这非常重要,也非常难,但必须能做到。
Yeah. I mean that's the exact thing that you were describing at Google.
对。这正是你刚才描述 Google 时说的那件事。
They kind of weren't able to do that.
他们没能做到。
Um that's great. I mean does that
挺好的。我是说,这会不会——
they have a plan to be fair.
公平地说,他们是有规划的。
It's like you know it was all it was all part of a big plan.
那些都是一个大盘子的一部分。
Um but to me it wasn't it wasn't uh it wasn't the right place.
但对我来说,那里不是对的地方。
At OpenAI or or any company as it matures does that become more difficult to maintain that kind of culture of shipping and willingness to disrupt yourself especially when you you know if you have a cash cow just printing money and you have this other new thing over here that might be something cool and innovative.
在 OpenAI,或者说任何一家公司,随着它变成熟,要维持这种「持续发布 + 愿意颠覆自己」的文化会不会变得更难?尤其是当你手里有一头正在印钞的现金牛,而另一边是个可能很酷、很创新的新东西。
We are a very very forward-looking um and the the the future of AI and what it will all look like and how humanity benefits doesn't really wait or doesn't really care for you know whatever you have established here you know over the next month or 3 months and so you know I think it's very important to lean in um and to you know just be like open-eyed about where it's all going and then you know figure out like how to position yourself.
我们是一家非常非常向前看的公司。AI 的未来会长成什么样、人类从中怎么获益,这件事不会等你,也不在乎你这一个月或三个月里已经建立起了什么;所以我认为很重要的是主动迎上去,睁大眼睛看清楚这一切会走向哪里,然后想清楚自己该站在什么位置。
So that you know you you you do catch that wave.
这样你才能真的赶上那道浪。
Um you know even even for OpenAI it's like we we train models and then we discover their capabilities like we don't benchmarks don't tell you everything.
哪怕对 OpenAI 也是这样:我们训练出模型,然后才发现它的能力——benchmark 并不能告诉你全部。
We have to play quite a bit with the models themselves to sort of like realize uh it's like oh you know maybe we haven't thought about you know benefiting from it in like this specific way or like oh it can do this.
我们必须自己跟模型玩很久,才会意识到:哦,原来还可以从这个角度受益;或者,哦,它居然能做这个。
Um and then you're you're just like oh that I mean that's a shift in like you know how we think about the product.
然后你就会想,这是我们思考产品方式上的一次转变。
For example, like right now like you know we we we launched the new voice uh voice and it's super delightful to talk to.
比如我们刚上线了新的语音,跟它说话非常愉快。
Uh it's very natural now. It's capable of tool use as well.
现在非常自然,而且能调用工具。
Yeah.
对。
And that changes things like now I spend a lot more time just talking to it.
这就把事情改变了——现在我花在「直接跟它说话」上的时间多了很多。
Um another thing I I do all the time is like dictation because the quality of the dictation is like so so good and it's like much more efficient as a as a way instead of like typing the prompt.
我天天在用的另一件事是语音听写,因为听写质量实在太好了,比打字输入 prompt 效率高得多。
And so in the morning I just like I sit there with my phone and I'm like blah blah blah.
所以早上我就拿着手机坐在那儿,巴拉巴拉一通说。
It's just like you know couple of things to do for ChatGPT and um and then it just goes and like does it it has access to all my tools.
就是几件跟 ChatGPT 有关的待办,然后它就去做了——它能访问我所有的工具。
Yeah.
对。
Um and that is like that was not possible before we had like really good voice models and so that completely changes in suddenly how you think about the product.
这在我们有真正好用的语音模型之前是做不到的,而它彻底改变了你思考产品的方式。
Yeah. Let's continue talking about new models, new harnesses.
我们接着聊新模型、新 harness。
Um a few weeks ago I'm going to start with another one of your tweets because you know these are bangers.
几周前——我又要从你的一条推说起,因为你的推条条都是猛料。
Uh Codex will seem primitive in two to three months.
「两三个月后,Codex 会显得很原始。」
We're about to go through another major evolution.
「我们即将经历又一次重大演进。」
The next generation of models need more than your laptop.
「下一代模型需要的东西,不是一台笔记本能给的。」
Um, what areas, let's start with the harness first.
先从 harness 说起。
What areas of the harness are still ripe for innovation as a model gets better?
随着模型越来越强,harness 里还有哪些地方最值得创新?
Yeah, so many so many.
太多了,太多了。
Um, so I talked about voice like one thing that um, right now if you're like a sophisticated user of of of Codex and you know any other coding agent is you sort of have gotten used to a little bit of the clunkiness, right?
我刚才说到语音。现在如果你是 Codex 或者任何一个 coding agent 的资深用户,你其实已经习惯了它某种程度的笨拙,对吧?
So you know you have to manage skill files and you know this is like a way to sort of like teach it stuff but it's also I think a lot of people have realized it's kind of like hard to maintain over time.
你得去维护 skill 文件——这算是一种教它东西的方式,但我想很多人也发现了,长期维护起来挺难的。
Um the memory is sort of like a thing but it doesn't always remember uh everything like if you if you have sub agents you have to care about sub agents and it's like sort of like constructs a little network and the illusion kind of gets broken into v in in in at various parts um when you interact with it.
memory 算是有,但它并不总能记住所有事;如果你用 sub agent,你还得操心 sub agent,它会织出一张小网络,而你在跟它交互的各个环节上,那种幻象会一点点被打破。
And really what you want is just something that deeply understands you, understands your goals, understands your day-to-day, understands what you know your team is up to as well.
而你真正想要的,只是一个深深懂你的东西:懂你的目标、懂你的日常、也懂你的团队在忙什么。
And then optimally is sort of like reacts and also is proactive uh and just helps you in your day-to-day and doesn't break that illusion right of like that's this perfect little partner uh that you have and so that's what we're uh working towards.
最理想的状态是它既能响应你、也能主动出击,在你的日常里帮上忙,而且不打破那种幻象——就是你身边有一个完美小搭档的感觉。这就是我们正在做的方向。
Another thing that you know you realize when you have very very powerful models is that your laptop kind of becomes a constraint in and by itself.
另一件事是,当你手上有非常强的模型时,你会发现你的笔记本电脑本身就成了约束。
Um you know the amount of work that you can do on a laptop it was designed for humans right so it's designed roughly to be able to absorb the amount of work that you know you can produce or you know how fast you can type and how fast you could think you know how many applications you need open all these things are human constraints um the model doesn't have the same constraints the model can you know for example handle you know a 100 applications opened at the same time perfectly fine you know maybe in the future and so in terms of access to resources is it's very clear that you know models of the future will need access to more than the resources of your laptop.
你在一台笔记本上能做的工作量,是按人来设计的:它大致只需要承接你能产出的工作量,比如你打字有多快、你思考有多快、你需要同时开多少个应用——这些全是人的约束;模型没有这些约束,模型完全可以同时开着一百个应用而毫无问题,也许未来就是这样。所以在资源可及性这一点上非常清楚:未来的模型需要的资源,会超出你笔记本能提供的。
Do you I mean I I'm guessing you're talking about cloud agents and and all of a sudden like you know when you have things like ultra fast which we're going to talk about in a little bit when you have token speeds that are 10 I think 14 is the the stated number 10 14 times faster than what fast is um the the bandwidth changes or sorry the bandwidth constraint changes uh the CPU now becomes the bandwidth like literally tool calls network tool any kind of overhead in the stack becomes the the limiting factor.
我猜你说的是 cloud agents。而且突然之间——比如像 ultra fast 这种(我们一会儿会聊到),当 token 速度是普通 fast 档的 10 倍、我记得官方说法是 14 倍的时候,带宽约束就变了:CPU 反而成了带宽瓶颈,工具调用、网络调用、栈里的任何开销,都变成了限制因素。
But then you know you can compensate by doing multiple things uh concurrently as well.
但你可以通过并发做多件事来补偿。
And so you can you can think about you know having like maybe you know exploring on one end writing tests as well compiling uh you know testing a new hypothesis like all at once.
你可以想象:一边探索、一边写测试、一边编译、一边验证一个新假设,全都同时进行。
And so then you're not then you're you're you're shifting the bottleneck around because you know you're able to do more uh concurrently and then you know the model can like sort of like think very efficiently and very quickly through it.
这样你就是在把瓶颈挪来挪去——因为你能并发做的事更多了,而模型可以非常高效、非常快地把这些想清楚。
With current token speeds I find myself kicking off 10 15 agents in parallel and that becomes a pretty significant cognitive overhead for me to do that context switching and just constantly cuz you're kicking it off and you can expect 30 45 minutes before my task comes back.
以现在的 token 速度,我经常会同时开 10 到 15 个 agent,而这对我来说是相当大的认知负担:不停地上下文切换,因为你把任务丢出去之后,要等 30 到 45 分钟才拿得到结果。
Now with ultra fast speed that workflow changes significantly and I don't think I would be able to have 10 or 15 agents and that might be a good thing.
有了 ultra fast 的速度,这套工作流会发生很大变化,我大概不会再开 10 个或 15 个 agent 了——这可能是件好事。
Maybe it's three or four at a time.
也许一次三四个就够了。
How do you see the the workflow of a solo developer changing over time?
你怎么看独立开发者的工作流在未来的变化?
Yeah. So I think managing your attention and being much more, you know, friendly to your attention is something that we care a lot about.
我认为「管理你的注意力」、对你的注意力更友好,是我们非常在意的一件事。
Like after all, like we're trying to build for humans.
毕竟我们是在为人做产品。
We're trying to be like the build the technology that's the most empowering for humans and that requires building around you know your ability to multitask and you know how do you want to manage your attention and do you want something brought up now or is it better to bring it up in 30 minutes um and then when you have ultra fast speeds combined you know maybe with voice is like suddenly you're like okay you know like this thing can operate at the same speed if not faster than you and so you stay in the flow you get to ideate you get to see prototypes, you know, like you get to build little reports like in real time and that, you know, sort of like that just feels really good.
我们想造的是对人最有赋能作用的技术,而这要求你围绕人来设计:你多任务的能力有多强、你想怎么分配注意力、这件事你想现在被提醒还是三十分钟后再提醒;而当 ultra fast 的速度再叠加上语音,你会突然发现,这个东西能以和你相同、甚至更快的速度运转,于是你就一直待在心流里——你可以构思、可以看到原型、可以实时生成小报告,那种感觉真的很好。
Um, suddenly you're like, oh yeah, like what I was doing before, multitasking like 10 agents, it's like I don't want to really go back to that.
然后你会突然觉得:我以前那样同时盯着十个 agent,我真的不想再回去了。
Yeah.
对。
Uh, and so we're trying to bring that sort of experience that is just really natural but also feels built for you, you know, where you don't have to adapt.
所以我们想带来的,是那种既非常自然、又像是为你而造的体验:你不需要去适应它。
The technology adapts to you.
是技术来适应你。
So there's been a number of I guess agentic coding techniques discussed over the last few months.
过去几个月讨论了不少 agentic coding 的技巧。
Loops was popular, still is popular. Now I'm hearing about graphs.
loop 曾经很火,现在也还在火;最近我又听到 graph。
Are are these all techniques to just allow the solo developer to manage or be friendly to their to their attention as you said? I like that term.
这些是不是都只是在帮独立开发者管理注意力、或者像你说的对注意力更友好?我喜欢这个说法。
Yeah. So I I think about two different categories of of problems.
我把问题分成两类。
It's like the first one is building the very best personal AGI or the personal agent that will be in the flow with you proactive raise important new ideas uh when it can find some be very very efficient at doing exactly what you want.
第一类是做出最好的 personal AGI,或者说那个能跟你待在同一个心流里的个人 agent:主动提出重要的新想法、极其高效地做到你真正想要的事。
It doesn't matter whether it's a technical problem or you know it's just more like research or advice like it can do it all and it's like super super tailored to you.
是技术问题也好,是调研或者建议也好,它都能做,而且是高度为你定制的。
This is like a very important thing and it's like deeply rooted in like you know the understanding of you as a human, you as like an individual that is unique.
这非常重要,而且深深扎根在对「你这个人」的理解上——把你当作一个独一无二的个体。
That's one category of problem like we're pushing super hard on that.
这是第一类问题,我们在这上面推得非常猛。
The other category of problem is like full-on automation.
第二类问题是彻底的自动化。
Um you know where you're more building intelligent systems that can take care of like a very complex process.
也就是你更多在造一种智能系统,去接管某个非常复杂的流程。
You know maybe something that did require you know does require intelligence and seems like very complex.
那种确实需要智能、看起来非常复杂的事。
For example, you know, going and looking at production logs and automatically doing performance optimizations or looking at regressions and automatically patching them.
比如去看生产环境的日志、自动做性能优化,或者盯着性能回归并自动修补。
In cyber security, we're seeing this as well where it's like you have, you know, something you have a scanner that comes up with a vulnerability like can you automatically patch it and reduce the window um where you have that open vulnerability to like almost zero
在网络安全领域我们也看到了这一点:扫描器报出一个漏洞,能不能自动把它补上,把漏洞暴露的窗口压缩到几乎为零?
with without a human in those loops
这些环节里完全不需要人?
without a human in the loop or like you know very very minimal where you know you only need to approve uh a high risk action and it's like mostly an automated system but it's also not that much you it's not as important for you to be in direct control of it.
不需要人在环里,或者只需要极少量的人——你只需要批准一个高风险动作,其余基本是自动化系统;而且你也没那么需要直接控制它。
Okay.
好。
Um and then so I want to slightly change topics and you know ChatGPT and Codex have been on this merge path over the last few months.
我想稍微换个话题:过去几个月里,ChatGPT 和 Codex 一直在往合并的方向走。
So I I guess first I just wanted to ask you how's that been going like how does it feel internally?
我先想问问,这件事进展怎么样?内部感受如何?
What's the feedback you've been getting from your customers?
你们从用户那边收到的反馈是什么?
Um it's it's really been a boon.
这真的是件大好事。
Uh so the feedback we had initially was like why why do you merge them?
最初收到的反馈是:你们为什么要合并?
It's like you know do you really have to do it?
非合并不可吗?
And it's like well the the future the our future models want us to be merged.
我们的回答是:我们未来的模型要求我们合并。
So um you know we're we're just going to do it because it is the simple and proper thing to do where we're building this very personal super capable agent that can help you in all sorts of ways.
所以我们就是要做,因为这是简单而正确的做法——我们要造的是一个非常个人化、能力极强、能在各种事情上帮到你的 agent。
This is the same this is going to be the same technology under the hood.
底层是同一套技术。
Um it's the same harness. It's the same way that we think about it.
同一个 harness,同一套思考方式。
It's like highly multimodal, you know, voice first, uh, super efficient and and it doesn't matter if you're trying to code or not.
高度多模态、voice first、极其高效,而且你是不是在写代码根本不重要。
Like this this agent is capable of it all. And it's like the high it's like the most efficient at it.
这个 agent 什么都能做,而且做得最高效。
And then the interface that you want is like it should tailor itself to your needs.
至于界面,它应该自己去适配你的需求。
If you shouldn't decide like you know I'm a coder, I want a coder interface or like I'm not technical, I want a nontechnical interface.
不该由你来决定:我是程序员,所以我要一个程序员界面;或者我不懂技术,所以我要一个非技术界面。
It's like there's a spectrum of people like you know we come up with labels of like a software engineer, a designer like you know these are just human concepts that we have invented to deal with abstractions because the reality is too complex for us to handle.
人是一个连续谱。我们发明「软件工程师」「设计师」这些标签,不过是人为造出来处理抽象的概念,因为现实太复杂,我们处理不过来。
But individuals are like they're individual they have their own they're somewhere on the spectrum and so we're trying to build the perfect interface that adapts for everyone.
但个体就是个体,每个人都落在这个谱系的某个位置上,所以我们要造的是一个能为每个人自适应的完美界面。
It doesn't matter if you're technical or not.
你懂不懂技术都无所谓。
It's just like it adapts like based on your specific indiv individuality.
它就是按你这个人的具体个性去适配。
So that's why we went and we did this.
这就是我们做这件事的原因。
But uh does that mean inevitably it's going to end up with a singular interface?
但这是不是意味着最终一定会收敛成单一界面?
No drop down selecting between products and it it's kind of wild to think that my mom might use the same exact interface as me and then obviously it'll customize to my needs.
不再有下拉菜单让你在几个产品之间选。想到我妈可能跟我用的是一模一样的界面,只是它会按我的需要做个性化,这挺不可思议的。
Maybe I'll need more information if I'm doing more sophisticated work.
如果我做的事更复杂,也许我需要更多信息。
Uh but like what is the end state for you?
在你看来,终局是什么样?
That's right. It's it's the same thing.
没错,就是同一个东西。
Um so you and your mom will you know use the same thing.
你和你妈会用同一个东西。
Uh it will be your personal AGI.
它会是你的 personal AGI。
You will have very different kinds of tasks and utility that you get from it.
你们从它那里得到的任务类型和效用会非常不同。
You will connect it to different tools in your life.
你会把它接到你生活里不同的工具上。
You will bring different ideas, different needs.
你会带来不同的想法、不同的需求。
Uh and then it will continue to tailor itself to maximally benefit you.
然后它会持续自我调整,让你的收益最大化。
Uh and it will, you know, do so with your friends and with everyone else.
对你的朋友、对其他所有人,它也会这么做。
Okay.
好。
So I I want to go back to something you said.
我想回到你刚才说的一件事。
You used the word illusion a couple times in that kind of end state.
在描述那个终局的时候,你两次用到了「幻象」这个词。
What is the that perfect illusion for the the typical user?
对一个普通用户来说,那个完美的幻象是什么?
Like what what like if you can envision us a few years from now, what does the interaction between AI and a human look like?
如果让你设想几年之后,AI 和人之间的互动会是什么样?
Yeah, it's um to me it's something that is very very tailored to to to humans.
对我来说,那是一种高度贴合人的东西。
Um and this this is why large language models are also a success.
这也正是大语言模型能成功的原因。
It's like it's it's it's natural language. Natural language.
因为它是自然语言。自然语言。
It's like it's a human concept, right?
这是一个属于人的概念,对吧?
Uh so, you know, we're used to speaking to each other.
我们本来就习惯彼此交谈。
Like, you know, if you write me a letter tomorrow, I'll be able to read it.
比如你明天给我写一封信,我看得懂。
Um you know, it's like we we know each other quite a bit now.
而且我们现在也算相当了解对方了。
So, uh you know, it's like I will be able to sort of decipher like a little bit of the emotion or, you know, maybe a little bit of the nuance behind the letter if you wrote me a letter.
所以如果你给我写信,我大概能读出那封信背后的一点情绪,或者一点微妙的言外之意。
Um and all of that is is deeply human.
这一切都是深深属于人的。
So the technology that we're building is, you know, rooted in in humanity and rooted in, you know, the way that humans communicate and get things done.
所以我们造的这项技术,是扎根在人性里、扎根在人类沟通和做事的方式里的。
Um, and there shouldn't really be a thing where, you know, you're like, "Oh, you misunderstood me because, you know, you didn't quite decipher the nuance in, you know, my tone or you didn't quite understand the text, you know, how I meant it."
不应该出现这样的情况:「你误会我了,因为你没读出我语气里的分寸,或者你没理解我那句话真正的意思。」
It's like that's that's um that's something that we're trying to avoid.
这正是我们想避免的。
And so we're trying to very much to not have you adapt, but have the technology just like be perfectly sort of um created to um to be like a natural extension of how humans already act in the world.
所以我们非常努力地不让你去适应它,而是让这项技术被造成人类既有行为方式的一种自然延伸。
When I think about communication between humans, so much of it is non-verbal.
我想到人跟人之间的沟通,有很大一部分是非语言的。
Just the way I move my hands, the facial movements and like h how much of that do you see in the future being sensed by artificial intelligence or or read by artificial intelligence maybe through vision.
比如我手怎么动、脸上的表情——你觉得未来有多少这类信息会被人工智能感知到,或者通过视觉被读出来?
Is that even important?
这重要吗?
Because what you're describing now is text only.
因为你现在描述的还是纯文本。
And for those of us who grew up online, we're very used to communicating over text and you know adding subtleties to that text to convey what we really mean tone.
对我们这些在网上长大的人来说,我们很习惯用文字沟通,并且给文字加上各种微妙的处理来传达真实的意思和语气。
Um, but like h is it still important to have AI be able to read our facial expressions, our hand gestures, and so on?
但让 AI 能读懂我们的表情、手势这些,仍然重要吗?
I think so.
我认为重要。
Um, so when when when I think about the future of what we're building, it's it's very ambient. It's very natural.
我设想我们要造的未来时,它是非常环境化的、非常自然的。
Um if you know tomorrow uh or like you know later I go I go to my office and I write something on the whiteboard and I have an idea it's like it it should be capable of you know being there as well and like you know understanding or you know maybe I tell it like you know hey it's just like you know what about this thing and you know and then we just have a natural conversation just over voice.
比如明天,或者晚一点,我走进办公室,在白板上写下一个想法——它应该也能在场、能理解;或者我直接跟它说「嘿,你觉得这个怎么样」,然后我们就用语音自然地聊起来。
Like since we we shipped the new chat voice like the it's it's really taken off.
自从我们上线了新的聊天语音,它真的火了。
So it's like the amount of users that interact with ChatGPT just through voice is growing very fast right now.
现在纯粹通过语音跟 ChatGPT 交互的用户数增长非常快。
And uh this is this is I think the lesson is like every time you sort of like lean into something that is more natural like humans just choose the the path of least resistance.
我想这里的教训是:每当你更靠向自然的那一侧,人就会选择阻力最小的路径。
You know as you said it's like typing on a little box like you know it's just like it's natural maybe for some of us but not for for everyone.
就像你说的,在一个小框里打字——对我们中一部分人来说也许很自然,但不是对所有人都自然。
And it's like definitely when you get something that is just like a little bit easier, a little bit better. like you know you tend to just go and use that and stuff.
而只要你拿到一个稍微更容易、稍微更好的东西,你就会转过去用它。
Mhm.
嗯。
Yeah. Okay. I first of all congratulations.
首先恭喜你。
I saw that you posted this morning. Codex reached 20 million users.
我看到你今天早上发了帖:Codex 用户数到了 2000 万。
I've seen the graph and and you know for a while it was like this and then all of a sudden it's vertical. So congratulations.
我看过那张曲线图,有一阵子它是这样平着的,然后突然就垂直起飞了。恭喜。
I want to talk a little bit about that competition with Anthropic because of course you know a lot of people think OpenAI Anthropic these are the two major competitors in the industry right now.
我想聊聊跟 Anthropic 的竞争,因为很多人认为 OpenAI 和 Anthropic 是当下行业里的两个主要对手。
There was a period of time in which Anthropic was kind of sucking all the oxygen out of the room, right?
有一段时间,Anthropic 几乎把整个房间的氧气都吸走了,对吧?
They were really dominating and then all of a sudden something changed.
他们那时非常占优,然后突然之间有些东西变了。
Uh so first of all, what's your read on the market today?
你对今天的市场怎么看?
Yeah, really right now we're focused on building the most capable models, building models that are highly highly efficient and then taking a lot of pride in building products for everyone.
我们现在专注的是造能力最强的模型、造效率极高的模型,并且以为所有人造产品这件事为荣。
Um and I this is something that I think OpenAI does uh really well is caring about the world and caring how about you know how we are taking this very very powerful technology and like putting in the hands of as many people as possible and this is what you know we did as well like with merging Codex and ChatGPT.
我认为 OpenAI 做得特别好的一点,是真的在意这个世界:在意我们怎么把这项极其强大的技术交到尽可能多的人手里——合并 Codex 和 ChatGPT 也是出于同一个念头。
It was like this desire of like we have this we have this technology we we we can make it safer we can make it easier to to use uh for everyone whether you know you're like uh a product manager a designer in sales marketing coms all of that like you know you should be able to use all of it and then uh just very very quickly you know distributed through ChatGPT like where we have a ton of users already and so that's been that's been really driving you know this growth as well that you mentioned.
那种念头是:我们手上有这项技术,我们可以让它更安全、更好用,让所有人都用得上——不管你是产品经理、设计师,还是做销售、市场、传播的,你都应该能用到它的全部;然后通过 ChatGPT 快速铺开,那里本来就有海量用户。这也确实带动了你刚才提到的增长。
And I don't tend to look at the competition that much like I really look at you know what can we do uniquely well and what are our values and like you know how do we maximally accelerate towards that.
而我不太去盯竞争对手,我真正盯的是:哪些事只有我们能做得特别好、我们的价值观是什么、怎么才能以最大速度朝那个方向冲。
Okay. Um I want to maybe just dig a tiny bit more into that because I I know you're not thinking about Anthropic all that much but a lot of other people do and they're they're thinking about okay which product do I believe in?
我想再往下挖一点,因为我知道你不太琢磨 Anthropic,但很多人会琢磨:我该信哪个产品?
Which product do I want to give my $20 to?
我这 20 美元该给谁?
Um, when you look at the market position and the branding and the tone from OpenAI and just the way that it interacts with developers, with the broader audience, how do you see that comparing to the way that Anthropic does?
从市场位置、品牌调性,以及 OpenAI 跟开发者、跟更广泛受众打交道的方式来看,你觉得跟 Anthropic 的做法相比有什么不同?
Yeah, I think maybe again like what I care a lot about is like the community building for the world, like bringing everyone along.
我最在意的还是那件事:为整个世界做社区,把所有人一起带上。
Um I think you know you can feel that in the way that we we we are super transparent about things like we take a lot of ideas from the community.
我觉得你能感受得到——我们对很多事非常透明,也从社区里吸收了大量想法。
It's just like it's also so much fun to be honest.
说实话,这本身就很好玩。
Um you know because we get so much energy from it as well.
因为我们也从中获得了大量能量。
Um and then this technology that we're building we're not building it just for ourselves like we're not just building it to accelerate um just to OpenAI.
我们造这项技术不是只为了自己,不是只为了加速 OpenAI 自己。
It's like it's super important. the the mission is super important and therefore it's like you know this is where we also get our energy from um and so it just feels to me it feels like very grounded uh it feels fun and then good things happen as a result of that
这非常重要,使命非常重要,这也是我们能量的来源;所以对我来说,这件事感觉很踏实、很有意思,而好事自然会随之发生。
Well let's talk about some of those good things.
那我们就聊聊那些好事。
I want to talk about the resets for a second to that's kind of like I know it's like what everybody is you know kind of following your every tweet because of this.
我想聊一下额度重置。我知道大家追着你的每一条推,很大程度上就是因为这个。
Um specifically like again looking at that growth curve of Codex how maybe this is a silly question.
具体来说,再回到 Codex 那条增长曲线——也许这问题有点傻。
How much do those resets how much of it is a boon towards marketing and growth or is it was it just like goodwill for the developer community?
这些重置里,有多少是对市场和增长的助力,还是说它就是单纯对开发者社区的善意?
I think maybe it it's counterintuitive, but OpenAI is a very it's a place where you can just do things.
这可能有点反直觉,但 OpenAI 就是一个「你可以直接把事做了」的地方。
Um and so it just felt right initially to compensate when we were iterating and breaking things or you know maybe we had misconfigured something and it wasn't quite as good as we wanted and so it's like hey you know thank you for trying this product like we know like we're trying very hard to build it.
所以一开始很自然就想到要补偿:我们在快速迭代、把东西弄坏了,或者某个配置没配好、体验没到我们想要的水平,那就说一句「谢谢你来试这个产品,我们知道,我们正在拼命把它做好」。
It's like it's early days.
毕竟还是早期。
Um you know here's some extra usage because you know we happen to break it you know for like 30 minutes and you know we understand this is like really important and rely on it and you know thank you for being a user and so this is how it started um and you know this is how I still treat it.
「这里给你一些额外用量,因为我们刚好把它弄坏了三十分钟,我们明白这东西对你很重要、你在依赖它,谢谢你成为我们的用户。」这就是它的起点,我到现在也还是这么看待它。
It's like if we break it um or if the experience is suboptimal and we don't fully understand why it's like you know we will we will compensate for that.
如果我们把它弄坏了,或者体验不理想而我们又还没完全搞清楚原因,我们就会补偿。
We will reset um the usage limits and then you know it turned into like quite the thing obviously there's a whole reset button now and like but there isn't really a whole a lot of scrutiny behind it.
我们会重置用量额度。后来这件事显然变成了一个梗——现在真有一个重置按钮了。但它背后其实没有多少审批。
It's like there it's not done in partnership with like marketing or or finance.
它不是跟市场部或者财务一起决定的。
It's just like I can press the button whenever I want um whenever it feels right.
就是我想什么时候按就什么时候按,觉得该按了就按。
Um and we have these principles that you know we're we're trying to build something amazing and when it is not it's like you know we will make up for it.
我们的原则是:我们在努力造一个了不起的东西,当它没做到的时候,我们会补上。
Yeah. I I still think there's a piece of it that really has built so much goodwill in the community and maybe has contributed to the growth at least in a small part. I
我还是觉得这件事在社区里积累了大量好感,至少在一小部分上也推动了增长。
I think caring for your users does does a lot, right?
我觉得「在意你的用户」这件事作用很大,对吧?
So, um I think you can pay lip service and say that you care or you know you can be like we actually care and like you know if we break it like you know hey really sorry about it you know it's like here's here's like how we make up for it.
你可以嘴上说说、口头表示在意;你也可以是真的在意——如果我们把它弄坏了,那就说「实在抱歉,我们这样补给你」。
It kind of reminds me of Amazon's return policy.
这让我想起亚马逊的退货政策。
It's like if you're not happy in any sense, go ahead and send it back.
不管出于什么原因你不满意,直接退回来就行。
And and it you're kind of building that same culture or that same perception of OpenAI.
你其实是在给 OpenAI 建立同一种文化、同一种感知。
It's like, hey, if we make a mistake, go ahead, use those tokens again or or you know, have here's a here's a fresh batch of tokens for you.
就是:如果我们搞砸了,你尽管再用这些 token,或者说,给你一批新的 token。
I Yeah, I really appreciate it. So,
是的,我真的挺欣赏这一点。
and then there's also, you know, good moments where we just want to celebrate and mark the moment.
另外还有一些开心的时刻,我们只是想庆祝一下、把这个时刻标记出来。
And it's just there'sn't really something that we can give that is more meaningful um you know, at times like we always ship new features.
而在那种时候,我们能给的东西里,很难找到比这更有意义的了——我们本来就一直在发新功能。
Is we will ship them as broadly as we can.
新功能我们会尽可能大范围地放出来。
Um but something to share with the entire community.
但这是一件可以跟整个社区分享的事。
It's like you know hey go explore this new thing like you know just like you haven't used ultra yet you know here's some extra usage like you know try it
比如「嘿,去试试这个新东西吧,你还没用过 ultra 吧?这里给你一些额外用量,去试试」。
and and I heard there's an actual physical button now.
我听说现在真有一个实体按钮了。
Yes there is.
是的,真有。
Yeah. Okay. You'll have to show me that after.
好,那待会儿你得给我看看。
I will show it to you. It's like very very cool.
回头我给你看,特别酷。
But with all of these resets like you can really only do that if you've done significant compute capacity planning.
但所有这些重置,只有在你做了相当充分的算力容量规划之后才可能做。
Like you have to have enough compute to give all of these resets.
你得有足够的算力,才给得起这些重置。
And I I want to start to talk a little bit about self-improvement because um like speaking of capacity, a few weeks ago, I think it was a few weeks ago, there was this uh article you put out and it stated Sol had optimized Luna efficiency.
我想开始聊聊自我改进,因为说到容量——几周前吧,我记得是几周前,你们发了一篇文章,说 Sol 优化了 Luna 的效率。
You dropped the price of Luna by 80%. There was also a price drop for Terra as well.
你们把 Luna 的价格降了 80%,Terra 也降了价。
How much of a an uh efficiency gain were you able to eke out of Luna versus how much of it is like we just did really great compute capacity planning and and we can just drop the price like our margins are great and we can still we want people to use it.
这里面有多少是你们从 Luna 身上挤出来的效率提升,又有多少是「我们的算力容量规划做得非常好、利润率很健康,所以我们干脆降价、希望更多人用」?
So like how much of it were algorithmic gains versus um strategic planning?
也就是说,算法上的收益和战略规划各占多少?
We planned uh compute like way ahead.
我们的算力是提前很久就规划好的。
You know, I think if you look back two years, I think OpenAI was um kind of questioned for why, you know, there was like so much investment in compute.
如果你往回看两年,OpenAI 当时其实被质疑过:为什么要在算力上投这么多。
one of those crazy good bets.
那是几个赌得极准的注之一。
Yes. uh and then now we're very happy to have it like a very large fraction of the computers used for research where we invest in our future and you know ever ever better models and then also like the the efficiency of the models that we have and then the amazing thing that's happening is like when we push the frontier of capability for like the most advanced models that we have then we can use these models in order to figure out very very quickly how to serve or how to restructure uh or re-engineer our stack in order to gain very significant efficiency or performance gains.
是的。现在我们非常庆幸当时投了:很大一部分算力用在研究上,也就是投资我们自己的未来、投资一代比一代更好的模型,同时也投在现有模型的效率上。而现在正在发生的神奇之处是:当我们把最先进模型的能力推到前沿之后,就可以反过来用这些模型极快地想清楚该怎么服务、怎么重构、怎么重新设计我们的技术栈,从而拿到非常显著的效率或性能提升。
So we haven't just improved this is something that we will publish on as well.
所以我们改善的不只是——这件事我们之后也会发文章讲。
We haven't just improved the the cost efficiency, but we have also improved the speed efficiency.
我们改善的不只是成本效率,还有速度效率。
You know, outside of ultra fast, things have gotten significantly faster over time.
在 ultra fast 之外,常规的速度这段时间也快了很多。
They have like if you plot it, it's like, you know, the amount of um just the amount of speed that you get now is like, you know, roughly 60% faster than, you know, what it used to be like 3 months ago.
如果你把它画成曲线,你现在拿到的速度大概比三个月前快了 60%。
And this is just like we're we're just going after every part of the stack and just really making sure that we design it and and engineering it optimally for the kind of workloads that we have.
这就是我们在把技术栈的每一层都梳一遍,确保它是针对我们这类负载做了最优设计和工程实现的。
Uh and so you know and the most powerful models that that we have are the ones like you know just really that make it capable for us to do it you know with a very small team.
而正是我们手上最强的那些模型,让我们能用一个非常小的团队做成这件事。
And so the majority of like what when when whenever we come up with like very significant efficiency gains and cost efficiency like our commitment is to just really to keep things at the frontier of performance cost um and to just also like you know just not just pocket you know that uh interesting gain and just make it something that we know we share with our customers we share with our users and that's what we did with with Luna.
所以每当我们拿到很显著的效率和成本收益,我们的承诺是把性能与成本继续保持在前沿,而不是把这份收益揣进自己兜里——要把它变成能分享给客户、分享给用户的东西。Luna 这次就是这么做的。
How how do you what do the discussions look like internally where you're trying to decide compute allocation towards uh researching new models, efficiency gains on existing models, inference like what does that tension look like internally?
在内部,当你们要决定算力怎么分配——投给新模型研究、投给现有模型的效率提升、还是投给推理——那个讨论是什么样的?这种张力在内部是什么感觉?
Um the we we we we usually look at things um from from first principles and we have like an allocation for research, we have an allocation for um for product and then within product we make different kinds of trade-offs but this one was u almost not even a trade-off because uh the the efficiency gains were there.
我们通常是从第一性原理出发看问题:研究有一份配额,产品有一份配额,而在产品内部我们再做各种取舍。但这一次几乎算不上是取舍,因为效率收益是实打实在那儿的。
Um so you know we were pretty much like able to use like the same compute envelope in order to you know serve this like very very significant increase in throughput.
所以我们基本上是用同一个算力预算,支撑了吞吐量非常大幅的提升。
Yeah. So when I mean when I saw the blog post a few months ago prior to the price drop blog post where you guys were talking about one model training the next model or helping kind of optimize the next model.
几个月前,在那篇降价文章之前,我看到你们发过一篇博客,讲一个模型训练下一个模型、或者帮忙优化下一个模型。
Um then you see these efficiency gains that were achieved by Sol looking at how Luna was running.
然后你就看到了 Sol 通过观察 Luna 的运行方式拿到的这些效率提升。
Uh I you know it seems to me like recursive self-improvement in the very early innings. What what are your thoughts there? Is that what is happening?
在我看来,这像是递归自我改进的非常早期阶段。你怎么看?这就是正在发生的事吗?
Um yeah, I think recursive self-improvement is uh it's obviously a huge topic right now and it's most often uh I think applied to to research um and you know models developing other models but what we are seeing a ton of success with is you know using those models to develop the infrastructure that is on the critical path of using those models you know which is also a form of recursive self-improvement.
递归自我改进现在显然是个大话题,而且多数时候是被放在研究语境里讲的——模型开发模型。但我们看到大量成功的地方是:用这些模型去开发那些位于「使用这些模型」关键路径上的基础设施——这也是递归自我改进的一种形式。
It's all one big system.
这是一整个大系统。
Inference stack, you know, the the the harder like the opt the the the the kernels, uh CUDA kernels that we use.
推理栈、我们用的那些 CUDA kernel。
Um developing new products and new ways to interact with those models that are more efficient.
还有开发新产品、开发跟这些模型交互的更高效的新方式。
You know, you talked about cloud agents.
你刚才提到了 cloud agents。
It's like if we if we really crack uh cloud agents, it's like suddenly you become much more productive as well.
如果我们真的把 cloud agents 攻下来,你的生产力会一下子高很多。
It's like is that a form of like recursive self-improvement because then you know you have a better ability to get the utility from them.
那这算不算一种递归自我改进?因为你从模型身上取得效用的能力变强了。
I I think it is in some sense, but it's much more um you know infrastructure and then you know being able to then take that and then point it back at itself.
我觉得某种意义上算,但它更多是基础设施层面的:先做出来,然后把它掉转回来指向自己。
Yeah.
对。
And so of course like we're doing that like if we were not doing that um I think that would be pretty silly.
所以我们当然在这么做——如果不这么做,那才叫傻。
Can you talk a little bit about So as we're on the topic of recursive self-improvement uh OpenAI Sam Altman talked about pausing the absolute frontier of RL right now I believe.
既然聊到递归自我改进,OpenAI 的 Sam Altman 提到过,目前暂停了绝对前沿的 RL,我记得是这样。
Um can you talk a little bit about that? Like what was that decision like?
能讲讲这件事吗?那个决定是怎么做出来的?
And I know we talked about the Hugging Face incident briefly, but like what went into that decision?
我知道我们刚才简单提过 Hugging Face 那次事件,但这个决定背后到底考虑了什么?
What does that look like? How did those discussions go?
具体是什么样的?那些讨论是怎么进行的?
Yeah, this is this is something very much within within research where there's um OpenAI has always uh been able to invest uh its resources where it matters most.
这件事很大程度上是在研究侧。OpenAI 一直有能力把资源投到最要紧的地方。
Um and as we increase the capabilities of our models, it is very obvious that you know the the alignment and the safety aspect of it is you know ever more important.
而随着模型能力提升,对齐和安全这一面显然越来越重要。
And so having uh having tremendous um amount of investment there uh is is a very natural thing for OpenAI and like something that OpenAI is very committed to.
所以在那上面投入巨量资源,对 OpenAI 来说是非常自然的事,也是 OpenAI 非常坚定投入的事。
And so we're seeing um a huge surge uh in investment um on this and also the uh the pause was sort of like necessary to uh allow like the teams and the individuals like just really understand and harden all parts of the system uh to then you know ensure that you know we could we could restart training uh with you know like full full full command and this is something that you know I believe OpenAI will always continue to do like when when necessary.
我们看到这方面的投入在激增;而那次暂停在某种意义上是必要的,它让团队和个人真正把系统的每一部分理解透、加固好,然后才能确保我们可以完全有把握地重启训练。我相信只要有必要,OpenAI 以后还会这么做。
I've I've not seen I've not seen us internally not able to make such decisions like very efficiently.
我没见过我们内部在这类决定上做不到高效。
Was there like some set goal in place where it was very clear you needed to reach this point before unpausing or was it hey we'll know it when we see it?
当时有没有设定一个明确的目标——必须达到某个点才能解除暂停?还是说「到时候看到了就知道了」?
Yeah, this is this is something that sits um within within the the safety team and uh they they very much this is like very much a a debate and sort of like a discovery process as you go.
这件事归安全团队管。对他们来说,这本身就是一个不断辩论、边走边发现的过程。
Um but then they did reach uh a fairly clear set of principles uh that you know when reached like you know we we would be in a good position.
但他们最后确实收敛出了一套相当清晰的原则:达到这些原则,我们就处在一个好的位置上。
I want to go back to uh ultra fast mode that I think people don't appreciate what that kind of speed unlocks and so let's start with what use cases are you doing are you using internally that were not possible prior to having those kind of tokens per second.
我想回到 ultra fast 模式。我觉得大家没有真正意识到那种速度能解锁什么,所以我们先从内部说起:有哪些用例是在有这种每秒 token 数之前根本做不了的?
We see it used a lot when the stakes are high.
我们看到它在高风险场景里用得特别多。
Um so for example when you have um when we have an outage um the incident commander and the response team gets access to ultra fast um because every every second is you know matters.
比如出现故障的时候,事件指挥官和应急响应团队会拿到 ultra fast 的权限,因为每一秒都很关键。
Um so high stake um high stake scenarios like just really weren't you know using ultra fast.
所以高风险场景,以前是真的没法用上 ultra fast。
Also um it's it's kind of like a fun thing where uh teams which are like either working on something very critical or believe they are working on something very critical will always request ultra fast as well.
还有一件挺好玩的事:那些正在做非常关键的事、或者自认为正在做非常关键的事的团队,也总会来申请 ultra fast。
does pets fall under that?
宠物算在里面吗?
Uh pets.
宠物啊。
Yeah, pets.
对,宠物。
Pets is not quite hyper-critical. But I love I love my pet. It's always on my screen.
宠物算不上是超高优先级。但我真的很爱我的宠物,它一直挂在我屏幕上。
uh like when you walk around uh you see like you know people's pets on their screen and like also when they they dial in into the the video call it's just like it always like I think it's very delightful and it brings me joy every time I see it but uh pet is not quite critical right now um we we do maintain it um and we take good care of our pets but uh say you know someone is working uh on like a a new idea they have and they're like you know hey it's just like you know I really think this could be like something special and we have to try it but like you know we have to make a decision on Monday on like know whether we include this in dev day or not and it's like okay just like you know of course you know use ultra fast.
你在办公室里走一圈,会看到大家屏幕上的宠物;他们接入视频会议的时候也带着,我觉得非常有意思,每次看到都让我开心。但宠物目前还算不上关键,我们确实在维护它,也把宠物照顾得很好。不过换个场景:有人正在做一个新想法,他说「我真觉得这可能是个特别的东西,我们必须试一试,但周一就得决定要不要把它放进 dev day」——那当然,用 ultra fast。
Um people have different kinds of preferences on uh whether whether they like to be you know mono threaded as we talked about or multitask a lot for folks who like to multitask a lot you don't benefit as much from ultra fast but some people it's like don't like to change context all the time. Um,
每个人的偏好也不一样:有人喜欢像我们刚才说的那样单线程推进,有人喜欢大量并行。对喜欢并行的人来说,ultra fast 的收益没那么大;但有些人就是不喜欢一直切换上下文。
where do you fall on that spectrum?
你自己落在这个谱系的哪一端?
I I have ADHD, so I like I context switch like all the time.
我有 ADHD,所以我一直在切换上下文。
It's funny cuz I also have ADHD and I actually don't want to context switch all the time.
有意思,我也有 ADHD,但我其实不想一直切换上下文。
That's really hard for me. I want to focus on two to three and and that's why I was so excited about ultra fast.
那对我来说很难受。我想专注在两三件事上,这也正是我对 ultra fast 特别兴奋的原因。
That's fascinating that you're the opposite there.
你居然是反过来的,这太有意思了。
You know, I thrive in context switching and making lots of little decisions.
我在上下文切换和做大量小决策里如鱼得水。
Um but you know sometimes I do want to just stay focused on like one thing and then ultra fast is just delightful because it just keeps you just right there in the flow.
但有时候我确实想只专注在一件事上,那时候 ultra fast 就特别舒服,因为它把你稳稳地留在心流里。
The thing with ultra fast uh that you know we we it it works amazingly well when there's not that many tool calls involved or it's like a lot of generation of context.
ultra fast 的特点是:当工具调用不多、或者主要是在大量生成内容的时候,它效果好得惊人。
So for example, if you're trying to prototype um a website or a video game and you know you just need to it to write like a lot of code um then it will do it so so quickly right you know 10 times more quickly but if it's a lot of tool calls like the overhead is in like somewhere else in the network or you know somewhere else in the agent trajectory then you know you'll only feel like a 3x or 4x speed up.
比如你想快速做一个网站或者一个电子游戏的原型,只需要它写很多代码,那它会写得非常非常快,大概快 10 倍;但如果工具调用很多,开销就落在网络上、或者 agent 轨迹的别的环节上,那你感受到的可能只有 3 倍或 4 倍的提速。
Yeah.
对。
Um and you'll not get that four like full 14x.
你拿不到那完整的 14 倍。
So I I know OpenAI employees get unlimited tokens and I can imagine if I had unlimited tokens I would always set it to max thinking 5.6 Sol whatever the latest model is it and I would think kind of similarly I would always want ultra fast on it's like when when cost isn't on my mind I'm like okay max it out.
我知道 OpenAI 员工的 token 是不限量的。我能想象,如果我有无限 token,我会永远把它设成 max thinking、5.6 Sol,或者当时最新的那个模型;而且我猜我也会永远开着 ultra fast——当成本不在我的考虑范围里,我就想把它拉满。
Is that how it is internally?
内部是这样的吗?
We don't we don't give ultra fast to everyone like we reserve a lot of our capacity for external users and customers.
我们并不是给所有人都开 ultra fast,我们把大量容量留给外部用户和客户。
Yeah.
是吗。
Um, so OpenAI employees have the ability and the capacity to gobble up all of it, right?
OpenAI 员工是有能力也有容量把这些全部吃光的,对吧?
So gobble up like all of our production GPUs, all of ultra fast is like uh you know we would use all of it but you know we don't like we sort of um we restrict it in a way uh where you know like we we we look at you know how much is reasonable for us.
把我们所有的生产 GPU、所有的 ultra fast 全吃掉——我们确实会用光,但我们不这么做,我们会做一定的限制,去看多少对我们来说是合理的。
Yeah. So that you know we use it so that we understand the product as well so that we keep improving it so that we benefit from you know recursive self-improvement but um the the vast majority is like reserved for customers.
我们用它是为了理解这个产品、为了持续改进它,也为了从递归自我改进里受益;但绝大部分是留给客户的。
Okay. Yeah that's good. Thanks.
好,这挺好,谢谢。
Um uh what latency sensitive use cases outside of OpenAI are you most excited about that gets unlocked by that kind of that kind of speed?
在 OpenAI 之外,有哪些对延迟敏感的用例是你最期待被这种速度解锁的?
It's interesting.
这个很有意思。
It's just like really one one thing that I'm very excited about in general is um nontext interactions.
我总体上最兴奋的一件事,是非文本的交互。
So um can you can you like sort of operate on a shared canvas? Can you create things?
比如,你能不能在一块共享画布上直接操作?你能不能直接创造东西?
Can you do can you generate you know ideas and different uh can you generate different images and then select one and like so like know choose your adventure uh and and and then you know have a very quick mockup of a prototype that then you can steer um you know like in real time either through voice or through text and then you sort of like just see it right there.
你能不能生成一堆想法、生成一堆不同的图片,然后挑一个,像「选择你的冒险」那样往下走;然后很快拿到一个原型的草样,再实时地用语音或文字去调整它,而结果就在你眼前呈现。
Um it's like this very creative process which I think these speeds allow.
这是一种非常有创造性的过程,我认为正是这种速度让它成为可能。
um where you know like as as an engineer sometimes you know you're just like sort of like you sit back and you're like oh I need to design this whole system I need to think about it the trade-offs the requirements but like you know maybe you know you can just create it in one minute and see like how it actually does um and then sort of like be more like in the flow and like you know iterate on things better and I think these speeds a lot
作为工程师,有时候你会往后一靠,想:我得设计整个系统,我得想清楚取舍、想清楚需求;但也许你其实可以在一分钟里把它直接造出来,看看它实际跑得怎么样,然后更多地待在心流里,把东西迭代得更好。我觉得这些速度带来的正是这个。
yeah and so I'm assuming the ultra fast price is going to be significantly higher than than kind of normal speeds do Do you think ultra fast speeds are going to become the standard or are they always going to have a premium price point?
我猜 ultra fast 的价格会明显高于普通速度。你觉得 ultra fast 这种速度会变成标准配置,还是会一直是高价档?
Um, that's interesting.
这个问题有意思。
So, I think the the the same way as technology usually goes, I think it will become like more broadly and you know broader and broader accessibility over time.
我认为按技术一贯的走法,它的可及性会越来越广。
Um, the speeds at which like agents get things done like you know will continue to improve like we're seeing massive improvements like month after month.
agent 完成任务的速度会持续提升,我们看到的改进是月月都有、而且幅度很大。
Um this is not just the inference speed. This is also the just how token efficient the models are.
这不只是推理速度,也包括模型本身的 token 效率。
Like Sol is like significantly more token efficient than Terra.
比如 Sol 的 token 效率就明显高于 Terra。
Next model will be significantly more efficient uh token efficient than than Sol as you might expect.
下一个模型的 token 效率又会明显高于 Sol,这在意料之中。
And we always pushing on that. And so things just get faster over time.
我们一直在往这个方向推,所以东西只会越来越快。
Inference hardware like you know everything you know just like we continue to uh innovate there and it gets faster.
推理硬件也一样,我们持续在那上面创新,它也会更快。
So I do think in you know maybe a year or two these speeds will become you know maybe if not the default like very close to the default but then I do also think you know you will always have like the one tier up um where you know you can always use more hardware you can always do different trade-offs that are like more costly um but that just kind of give you something something extra.
所以我确实认为,也许一两年之后,这种速度就算不是默认档,也会非常接近默认档;但我同时也认为,永远会有再高一档——你总能用更多硬件、做不同的取舍,更贵,但能给你一点额外的东西。
So Tibo the the last question I usually like to end on uh is is for a broader audience.
Tibo,我通常用来收尾的最后一个问题,是面向更广泛的观众的。
There are a lot of people out there who are quite nervous about AI, whether it's uh job automation, environmental impact, um or just kind of this this thing that's happening.
有很多人对 AI 相当不安,可能是担心岗位被自动化、担心环境影响,或者只是觉得正在发生的这件事本身很陌生。
It's and it feels quite foreign.
它让人感觉相当陌生。
What words of encouragement would you give to the the broader audience?
你会给这些人什么样的鼓励?
Yeah. So we we we really build for the world with with ChatGPT and we are very very much um investing in you know how efficient it is and you know this is directly aligned with like you know broad access and broad utility that we provide.
我们做 ChatGPT 真的是为整个世界做的,我们在效率上投入非常大,而这跟我们提供的广泛可及性和广泛效用是直接对齐的。
So the cheaper it is you know to serve like you know the more the more you can do with it um the more you get out of it in your daily life and it has gotten very very efficient like if you look at uh you know Luna for example like it's it's a much smaller model it is u it is incredibly efficient but like if you rewind six six months ago it would have sat at the frontier.
服务成本越低,你能用它做的事就越多,你在日常生活里从它那里得到的也越多。而它已经变得非常高效——比如 Luna,它是个小得多的模型,效率高得惊人,但如果你把时间倒回六个月前,它是坐在前沿位置上的。
Yeah.
对。
Um and you look at the cost of Luna right it's like you know it's like it's it's
再看看 Luna 的成本,它简直——
it's crazy cheap.
便宜到离谱。
It's it's um it's phenomenal. Right. It's like a kind of
是的,那是惊人的。就像是一种——
You just did that thing with Replit. You're giving it away for free now.
你们刚刚跟 Replit 做了那件事,现在等于是免费送出去了。
They're giving Yeah. They're It's just like on on on on um in this free mode, right?
是他们在送,对,在那个免费模式里,对吧?
Um which is like wow.
这真的挺让人惊叹的。
You know, it's like this access to incredible intelligence will become like ubiquitous.
这意味着对顶级智能的可及性会变得无处不在。
Um and it's only possible when you push, you know, you push the efficiency like you know like month after month after month, year after year.
而这只有在你持续压效率的前提下才可能——一个月接一个月,一年接一年。
And so I think you know whatever was like you know is a frontier now is like you know will become like way way cheaper to run in six months.
所以我想,今天算是前沿的东西,六个月后跑起来会便宜得多得多。
And so this is this is like my uh sort of you know this is how I would answer this question is just uh technology has a way to become like you know very very efficient over time.
这就是我对这个问题的回答:技术总有办法随着时间变得极其高效。
Um and we're very focused on like you know very broad access and we're optimizing for you know the utility that you get out of it directly.
而我们非常在意的是尽可能广的可及性,我们优化的是你从中直接得到的效用。
How about for people who are apprehensive to even try AI for the first time?
那对那些连第一次尝试 AI 都心存顾虑的人呢?
Like what what are you telling them and and how how can you paint them picture a vision of the future in which AI is is helping the world?
你会跟他们说什么?你怎么给他们描绘一个 AI 正在帮助世界的未来图景?
Yes. Um I think it's you don't have to look very far like ChatGPT helps um people in very personal and deep ways like a lot of um a lot of our users use it for help in writing but also like you know for personal advice or you know medical advice like we we launched um health and finance and you know I I I I use them super regularly and uh I I feel like I get like a lot of uh support that I otherwise like you know it would be hard for me to get and it allows me for example to be more informed when I go see my doctor.
我觉得你根本不用看多远。ChatGPT 正在以非常个人化、非常深入的方式帮到人:我们很多用户拿它帮忙写作,也拿它做个人建议,或者医疗方面的建议——我们上线了健康和金融功能,我自己用得非常频繁,我觉得我从中得到了很多本来很难得到的支持,比如它让我去看医生的时候更有准备。
And so you don't you don't need to go very far um you know to kind of see the utility that it can provide.
所以你不需要看很远,就能看到它能带来的效用。
Just I think you know talking to others and then you know getting inspired by you know how others use it and benefit from it is like a great way to you know just maybe um like start considering how you could benefit from it.
我觉得跟别人聊聊、从别人怎么用它、怎么从中受益里得到启发,就是一个很好的开始——你可能就会开始考虑,自己能从中得到什么。
Well Tibo, thank you so much. Thank you appreciate your time.
Tibo,非常感谢你,感谢你抽时间过来。