Lenny's Podcast with Lenny Rachitsky · 2025-02-09 · 双语整理

Why Soft Skills Are the Future of Work

Karina Nguyen · OpenAI Researcher · 从 Anthropic 到 OpenAI 的"前沿创造者"视角
"It's actually really really hard to teach the model how to be aesthetic with really good visual design, or how to be extremely creative in the way they write."
——所以正在被低估的不是 coding,而是品味、创造力、和让人愿意跟你协作的能力。
嘉宾 Karina Nguyen 在 OpenAI 主导 Canvas、Tasks、o1 的产品研究,此前在 Anthropic 负责 Claude 3 系列的后训练与评估,以及 100K 上下文窗口的 file upload。本期话题:模型怎么训练、Canvas/Tasks 的方法论、evals 为什么会变成 PM 的新工作、未来 3 年最值得投资的技能。
TL;DR · 速读

OpenAI 研究员怎么看 AI 时代的工作变化

  1. 模型训练更像艺术不像科学

    "Model training is more an art than a science. We think a lot about data quality — it's one of the most important things."

    "The way you debug models is actually very similar to the way you debug software."

    debug 模型和 debug 软件类似,但数据质量的判断需要长期积累的直觉,不是流程可以复现的。

  2. 没有数据墙,真正的墙是 evals

    "There will be infinite amount of tasks. We are literally hitting the wall in evals."

    "The scaling in post-training itself is not hitting the wall — we went from raw datasets to infinite tasks taught via reinforcement learning."

    o1 时代的后训练范式让数据不再稀缺,所有 benchmark 都被打满,瓶颈反而变成"我们写不出更难的 evals"。

  3. Canvas 本质上是教会模型三个 behavior

    "It came down to three main behaviors: when to trigger canvas, when to update the document, and how to make comments."

    每个 behavior 都用合成数据训练,边界判断、精准 edit、评论质量全都靠 evals 测量,产品力被压进训练目标里。

  4. 写 evals 正在变成 PM 的新核心工作

    "Product development might move from 'here's a spec, let's build it' to 'AI build this — here's what correct looks like.'"

    未来 PM 多数时间花在定义"什么叫做对",AI 去实现。spec/PRD 的重心从描述功能搬到描述 evals。

  5. Prompting 就是 PM/Designer 的新 prototyping

    "Prompting is a new way of product development or prototyping for designers and product managers."

    在 Anthropic 早期 Karina 就是边 prompt 边验证想法。今天 PRD + 设计稿 + 原型这三件事可以被一句 prompt 合在一起。

  6. 智能的边际成本在加速下降

    "Small models are becoming even smarter than larger models — because of distillation research."

    Claude 3 Haiku 已经比 Claude 2 还聪明。AI 进入"曾被 intelligence 卡住的所有事"都被解锁的阶段,医疗、教育首当其冲。

  7. 最难教模型的恰恰是品味和创意

    "It's actually really really hard to teach the model how to be aesthetic or how to be extremely creative in the way they write."

    "ChatGPT kind of sucks at writing — that's because it's bottlenecked by creative reasoning."

    aesthetic 数据稀缺、有 taste 的人不够、creative reasoning 还是开放研究问题。这条线短期内 AI 不会替代人。

  8. Soft skills 才是新的 hard skills

    "Prioritization, communication, management, empathy, collaboration — those are still humane."

    研究进展现在被"管理质量"限制——算力分配、项目优先级、人怎么配。Karina 直接说:研究团队的 mismanagement 限制了人类潜能。

  9. Strategy 也会被 AI 接走,这一点别幻想

    "Strategy is more like data analysis — connecting the dots across feedback, dashboards, and other inputs. Models are really good at that."

    Lenny 认为"AI 会做战略",Karina 同意——还指出 self-improving models 离我们已经不远。

  10. Anthropic 与 OpenAI 是两套节奏

    "Anthropic taught me real care and craft toward model behavior. OpenAI gives more research and product freedom — much more bottoms-up."

    Anthropic = focus + 极度 prioritization + craft;OpenAI = bottoms-up + risk-taking + 创意空间大。Karina 在两边都待过,直接对比。

Chapter 01

Why Models Get Confused

模型为什么会困惑 · self-knowledge 与 function-calling 的冲突
model training · debugging · self-knowledge · function calls

Lenny: "Today my guest is Karina, and Karina is an AI researcher at OpenAI, where she helped build Canvas, Tasks, the o1 chain-of-thought model and more."

Lenny:"今天我的嘉宾是 Karina。Karina 是 OpenAI 的 AI 研究员,她参与了 Canvas、Tasks、o1 chain-of-thought 模型等产品的研发。"

"Prior to OpenAI she was at Anthropic, where she led work on post-training and evaluation for the Claude 3 models, built a document upload feature with 100K context windows, and so much more."

"在 OpenAI 之前,她在 Anthropic 负责 Claude 3 系列的后训练和评估,做了 100K 上下文窗口的文档上传功能,还做了好多其他事情。"

"She was also an engineer at the New York Times, was a designer at Dropbox and at Square."

"她还曾是《纽约时报》的工程师、Dropbox 和 Square 的设计师。"

"It's very rare to get a glimpse into how someone working on the bleeding edge of AI and LLMs operates and how they think about where things are heading."

"能近距离观察一个工作在 AI 与 LLM 最前沿的人怎么做事、怎么判断方向,机会非常难得。"

"In our conversation we talk about how teams at OpenAI operate and build product, what skills she thinks you should be building as AI gets smarter, how models are created, why synthetic data will allow models to keep getting smarter, and why she moved from engineering to research after realizing how good LLMs are going to be at coding."

"对话里我们会聊到:OpenAI 的团队怎么运作、怎么做产品;在 AI 越来越强的当下,你应该在练哪些技能;模型是怎么被造出来的;为什么 synthetic data 会让模型持续变聪明;以及她为什么在意识到 LLM 写代码会有多强之后,从工程转去做研究。"

"If you enjoy this podcast, don't forget to subscribe and follow in your favorite podcasting app or YouTube. It's the best way to avoid missing future episodes and it helps the podcast tremendously."

"如果你喜欢这档节目,记得在喜欢的播客 app 或者 YouTube 上订阅关注。这能帮你不错过新一期,也能很大程度上帮到我们。"

"With that, I bring you Karina N."

"那么,有请 Karina N。"

中间略去 Interpret 和 Vanta 两段口播 sponsor。回到正文。

Lenny: "Karina, thank you so much for being here. Welcome to the podcast."

Lenny:"Karina,非常感谢你来,欢迎来到节目。"

Karina: "Thank you so much, Lenny, for inviting me."

Karina:"Lenny,非常感谢你的邀请。"

Lenny: "I'm very excited to have you here, because not only are you working at the cutting edge of AI and LLMs, you're actually building the cutting edge of AI and LLMs."

Lenny:"我特别期待这次聊天,因为你不只是工作在 AI 和 LLM 的最前沿——你本身就是在构建那个最前沿。"

"You recently launched this feature which is basically the first agent feature of OpenAI."

"你们最近发布的那个功能,基本上就是 OpenAI 的第一个 agent 类功能。"

"I also just did this survey — I don't know if you know about this — I did a survey of my readers and asked them what tools they use every day in their work and most use. ChatGPT was number one, above Gmail, above Slack, above anything else. 90% of people said they use ChatGPT regularly. It's absurd. And it wasn't around two years ago."

"我前阵子做了一份读者调研,不知道你看到了没——问大家'你每天工作里用得最多的工具是什么'。ChatGPT 排第一,在 Gmail、Slack 和其他所有工具之上。90% 的人说他们经常用 ChatGPT。这有点夸张。要知道两年前它根本还不存在。"

"Also, we're recording this the week that OpenAI announced Stargate — which is this half-trillion-dollar investment in AI infrastructure. So there's just a lot happening, constantly, in AI."

"另外,我们录这期节目正好是 OpenAI 宣布 Stargate 的那一周——一个 5000 亿美元规模的 AI 基础设施投资计划。所以 AI 领域一直都有大事发生,几乎没停过。"

"You have a really unique glimpse into how things are working, where things are going, how work gets done. So I have a lot of questions for you. I want to talk about how you operate and how you work at OpenAI, where you think things are going, what skills are going to matter more and less in the future, and also just where things are going broadly. So how does that sound?"

"你站在一个特别独特的视角上,能看到事情怎么在跑、往哪走、是怎么被做出来的。所以我有一堆问题想问你。想聊聊你的工作方式、OpenAI 是怎么运作的、你判断的方向、哪些技能会越来越值钱、哪些会变得不重要,以及对整个行业更宏观的看法。听起来怎样?"

Karina: "Sounds great. Thank you so much. Yeah, I was extremely lucky to join early days on Anthropic and kind of learned a lot of things there. And I joined OpenAI around like eight months ago, so yeah, I'm excited to chat more."

Karina:"听起来很棒,谢谢。我特别幸运,在很早期就加入了 Anthropic,在那学到了非常多东西。我大约 8 个月前加入了 OpenAI,所以很期待多聊聊。"

Lenny: "I'm definitely going to ask you about the differences between those, but I want to start more technical and just dive right in. I want to talk about model training. People always hear about models being trained, these big models, how much data it takes, how long it takes, how much money it costs, how we're running out of data — which I want to talk about."

Lenny:"我等会儿肯定会问你这两家公司的差异,但我想先从技术层面切入,直接进入正题。我想聊模型训练。大家总能听到'模型被训练','这些大模型要多少数据、多长时间、多少钱',以及'我们快没数据可用了'——这个我都想聊。"

"Let me just ask you this question — what do you think people most misunderstand about how models are created?"

"我先问你这个问题——关于模型是怎么被造出来的,你觉得人们最大的误解是什么?"

"Model training is more an art than a science."

Karina: "Model training is more an art than a science. And in a lot of ways, we as model trainers think a lot about data quality."

Karina:"模型训练更像艺术,而不是科学。从很多角度看,我们做模型训练的人,花了大量时间思考数据质量这件事。"

"It's one of the most important things in model training — how do you ensure the highest quality data for a certain interaction or model behavior that you want to create?"

"这是模型训练里最重要的事情之一——怎么保证'你想塑造出的那种交互、那种模型行为'背后用上的是最高质量的数据?"

"But the way you debug models is actually very similar to the way you debug software."

"但你调试模型的方式,其实和调试软件非常像。"

"One of the things that I've learned early days at Anthropic was — we discovered, especially with Claude 2 training, when you taught the model some of the self-knowledge of, hey, you actually don't have a physical body to operate in the physical world."

"我早期在 Anthropic 学到的一件事是——我们发现,特别是在训练 Claude 2 的时候,如果你教给模型一些自我认知,比如'嘿,你其实没有一个能在物理世界里行动的身体'……"

"But then at the same time, we had data that kind of taught the model some of the function calls — like, this is how you set an alarm."

"但与此同时,我们又喂了一些函数调用的数据,比如'这就是你怎么设置闹钟'。"

"And so the model would get extremely confused about whether it can set an alarm — but it doesn't have a body in the physical world."

"于是模型就极度困惑:我到底能不能设置闹钟?可是我又没有身体在物理世界里啊。"

"So it's like the model gets confused, and sometimes it over-refuses. So sometimes it says like, 'I don't know, sorry, I cannot help you.'"

"所以模型就懵了,有时甚至直接过度拒答,说'抱歉,我不知道,我没法帮你'。"

"There is always like a balanced tradeoff between — how do you make the model to be more helpful for users, but also not being harmful in other scenarios?"

"所以这里永远存在一个平衡的取舍——怎么让模型对用户更有帮助,同时又不至于在其他场景下变得有害?"

"It's always about — how do you make the model more robust and operate across a variety of diverse scenarios?"

"核心问题始终是——你怎么让模型更稳健,能在多种多样的场景下都正常运作?"

Lenny: "That is so funny. I never thought about that. Most of the data that's trained on is kind of like assuming it's a human describing the world and how they operate — and it assumes there's a body and you can do things. And the model told you it doesn't have a body."

Lenny:"太有意思了。我从来没想过这一层。模型读到的绝大多数数据,都是人类在描述这个世界、描述自己怎么活动——默认有身体、能做事。然后模型反过来告诉你:它没有身体。"

Chapter 02

No Data Wall, Only an Evals Wall

没有数据墙,只有 evals 墙 · synthetic data 与 RL 后训练
synthetic data · post-training · RLHF · benchmark saturation

Lenny: "Okay, I want to talk a little bit about data while we're on this topic. I know you have strong opinions here."

Lenny:"既然聊到这里,我想再展开聊聊数据。我知道你在这块有很强的观点。"

"There's kind of this meme that models are going to stop getting smarter because they're running out of data. They're trained in large part on the internet, and there's only one internet, and they've already been trained on it. What more can you show them about the world?"

"现在有个流行说法:模型马上就要停止变聪明了,因为数据快用完了。它们的训练数据大部分来自互联网,而互联网只有一个,而且已经被训练过一遍了。你还能再给它们看到关于世界的什么新东西呢?"

"And there's this trend of synthetic data — this term, 'synthetic data.' What is synthetic data? Why do you think this is important? Do you think it's going to work?"

"于是有了'合成数据(synthetic data)'这股浪潮。合成数据到底是什么?你为什么觉得它重要?你认为这条路能走通吗?"

Karina: "I think there are two questions here. We can unpack one at a time."

Karina:"我觉得这其实是两个问题,我们一个一个拆。"

"People say if you're hitting the data wall — I think people think more in terms of pre-trained large models, that are trained on the entire internet to predict the next token."

"人们说'撞到数据墙'的时候,脑子里想的多半是预训练阶段的大模型——它们在整个互联网上学预测下一个 token。"

"But what the model is actually learning during that process is — actually, how do you compress? The compression algorithm here is that the model learns to compress a lot of knowledge, and it learns how to model the world."

"但模型在这个过程里真正在学的,其实是'怎么压缩'。这里说的'压缩算法',指的是模型在学着把海量知识压缩成内部表征,并对世界建模。"

"So the next prediction of the word like, 'teach me how to drive' — and you only have a few words that will match that, like 'a car.' So the model actually learns about the world in itself. It's modeling human behavior, sometimes."

"比如下一句词预测——你说'教我怎么开',模型能想到匹配的就那么几个词,比如'车'。所以模型其实在学习这个世界本身。它有时候也在对人类行为建模。"

"And when you talk to pre-trained models, which are very, very large, they're actually extremely diverse and extremely creative — because you can talk to almost any kind of writer or user through a pre-trained model."

"你跟那些非常大的预训练模型说话时,它们实际上极其多样、极其有创造力——因为你能通过一个预训练模型,跟几乎任何风格的写作者或用户对话。"

"But I think what's happening right now is a new paradigm — the o1 series — where scaling in post-training itself is not hitting the wall."

"但我觉得现在正在发生的是一种全新的范式——o1 系列——它告诉我们,在后训练阶段做 scaling 这件事本身,根本没撞到墙。"

"And that's because basically we went from raw datasets from pre-trained models, to an infinite amount of tasks that you can teach the model in the post-training world via reinforcement learning."

"原因是:我们从'预训练阶段的原始语料数据集',跨到了'后训练阶段通过强化学习能给模型喂的无限多任务'。"

"So any task — for example, how to search the web, how to use the computer, how to write well — all sorts of tasks that you're trying to teach the model, all the different skills."

"任何任务都可以——比如怎么搜网页、怎么操作电脑、怎么把文章写得好——你想教模型的任何任务,任何技能,都可以丢进来。"

"And that's why I've been saying there's no data wall, or whatever — because there will be an infinite amount of tasks. And that's how the model becomes extremely super intelligent."

"所以我一直说没有所谓的'数据墙'——因为任务的数量是无限的。这就是模型走向极致超级智能的路径。"

"We are literally hitting the wall in evals."

"We are actually getting saturated in all benchmarks. So I think the bottleneck is actually in evaluations — that we don't have all the frontier evals."

"我们几乎把所有 benchmark 都打满了。所以我觉得真正的瓶颈在评估侧——我们没有足够的前沿 eval 题目可以用了。"

"Like, I don't know, GPQA — which is like a Google-proof question answering, PhD-level intelligence benchmark — is getting to like, I don't know, more than 60-70%, which is what PhDs get."

"比如 GPQA——一个号称 Google 都搜不到答案、专为 PhD 级智能设计的题库——现在模型已经能拿到 60-70% 分了,而这就是 PhD 本人的水平。"

"So we're literally hitting the wall in evals."

"所以我们其实是在 evals 上撞了墙。"

Lenny: "I want to follow both those threads. So the first is on this idea of synthetic data. Is a simple way to understand it that the models are generating the data that future models are trained on, and you ask it to generate all these ways of doing stuff, all these tasks as you described, and then the newer models are trained on this data that the previous model generated?"

Lenny:"这两条线我都想顺着追下去。先说合成数据。可以这样简单理解吗——模型自己生成数据,给下一代模型训练用?你让它生成各种'做事的方式'、各种你说的任务,然后新模型就用上一代模型生产的数据来训练?"

Karina: "Some tasks are synthetically curated. So this is like an active research area — how can you synthetically construct new tasks for the model to learn?"

Karina:"有些任务是合成构造出来的。这本身就是一个活跃的研究方向——你怎么用合成的方式构造出新任务给模型学?"

"Sometimes, you know, when you develop product, you get a lot of data from the product and user feedback. You can use that data too in this post-training world."

"有时候你在做产品的时候,会从产品和用户反馈里拿到大量数据,你也可以把这些数据放到后训练阶段用。"

"Sometimes you still want to use human data — because actually some of the tasks can be really, really hard to teach, like experts only know certain knowledge about some chemicals or biological knowledge."

"有时候你还是得用人类数据——因为有些任务真的非常难教,只有领域专家才掌握某些化学或生物知识。"

"So you actually need to tap into the expert knowledge a lot."

"所以你确实需要大量接入专家知识。"

"So yeah, I think — to me, synthetic data training is more for product. It's like a rapid model iteration for product outcomes. We can dive more into it."

"对我来说,synthetic data training 更偏向产品视角。它本质上是为'产品结果'服务的快速模型迭代手段。我们可以再深入聊。"

"The way we made Canvas and Tasks and new product features was mostly done by synthetic training."

"我们做 Canvas、Tasks 以及一系列新产品功能,主要都是靠合成数据训练完成的。"

Chapter 03

How Canvas Was Actually Built

Canvas 的三种 behavior · 触发、编辑、评论
canvas · synthetic training · behavior design · rapid iteration

Lenny: "Let's actually get into that. That's really interesting. I want to talk about evals, but let's follow that thread. So talk about how this helped you create Canvas."

Lenny:"我们就顺着这个聊吧,我觉得特别有意思。我本来还想聊 evals,但先把这条线走完。给我们说说合成数据是怎么帮你做出 Canvas 的。"

Karina: "So when I first came to OpenAI, I really had this idea of — okay, it would be really cool for ChatGPT to actually change the visual interface, but also change the way it is with people."

Karina:"我刚到 OpenAI 的时候有一个特别想做的事——让 ChatGPT 真的去改变它的视觉界面,同时也改变它和人协作的方式。"

"So going from being a chatbot to more of a collaborative agent — and a collaborator. It's a step towards more agentic systems that become innovators ultimately."

"从一个聊天机器人,变成一个协作型 agent——成为'协作者'。这是迈向'agentic 系统、最终变成创新者'的一步。"

"And so an entire team of applied engineers, designers, product, like research, kind of got formed in the air, almost out of nothing. It's just like a collection of people who just got together and we rapidly started iterating."

"于是整支团队——应用工程师、设计师、PM、研究员——几乎是凭空在空气里聚拢起来的。就是一群人凑到一起,开始飞速迭代。"

"Canvas is like — I would say like the first project of OpenAI where researchers and applied engineers started working together from the very beginning of the product development cycle."

"Canvas 算是 OpenAI 的第一个项目——研究员和应用工程师从产品开发周期最开始就并肩干。"

"And there's a lot of things that we have learned on the way, but I definitely came with the mindset of — we need to do a really rapid model iteration, such that it would be much easier for engineers to work with the latest model possible, but also learn from user feedback or early internal dogfood."

"一路上学到了不少东西。我自己带着一个明确的心智模型来——我们必须把模型迭代做到极快,工程师才能用上最新的模型,同时也能从用户反馈或者早期内部 dogfood 中学到东西。"

"How do we improve the model very rapidly? It's really hard to figure out — when you deploy a product — how people would be able to use it."

"怎么把模型迭代做快?这本身就很难——你部署一个产品,'用户实际会怎么用'这件事,是非常难提前预判的。"

"And so the way you synthetically train the model is basically figuring out what are the most core behaviors that you want this product feature to do."

"所以合成数据训练的方式,本质上就是想清楚——你希望这个产品功能拥有哪些最核心的行为。"

"For Canvas, it came down to three main behaviors."

"For Canvas, for example, it came down to three main behaviors."

"以 Canvas 为例,最后归结成三种核心行为。"

"It was — how do you trigger Canvas for prompts like 'write me a long essay,' when the user intention is mostly like iterating over long documents, or 'write me a piece of code'?"

"第一种:当用户说'帮我写一篇长文'或'帮我写一段代码',意图明显是'在长文档/代码上反复迭代'时,你怎么触发 Canvas?"

"Or when to not trigger Canvas — for prompts like 'can you tell me more about President X,' or general questions where the user intention is mostly getting an answer, not necessarily iterating on a long document."

"反过来,什么时候不该触发 Canvas——比如'跟我讲讲某位总统'这种,用户意图就是要个答案,而不是要在长文档上反复改。"

"The second behavior is — how do we teach the model to update the document when the user asks?"

"第二种行为是——当用户提出修改时,我们怎么教模型去更新文档?"

"So one of the behaviors is the model actually has some agency or autonomy to literally go to the document and select specific sections, and either delete it or edit, or highlight it and rewrite certain sections."

"这里其中一种行为是:模型其实拥有一定的自主性——它能真的进到文档里,选中特定段落,要么删掉、要么改、要么高亮再重写。"

"So sometimes the user would just say like, 'change the second paragraph to be something friendlier.' You would have to teach the model to literally find the second paragraph in the document and change it to a friendly tone."

"比如有时候用户就一句'把第二段改得更友好一点'。你得教模型真的在文档里定位到'第二段',再把它改成更友好的语气。"

"So basically you teach both how to trigger the edit itself, but also how do you teach the model to get higher quality edit for the document."

"所以你既要教它'怎么触发 edit',也要教它'怎么把 edit 改得质量更高'。"

"In case of coding, for example, there's also the question of how good the model is at completely rewriting the document versus having very specific targeted edits."

"比如代码场景里就有个问题:模型整段重写文档,和做非常精准、局部的修改,哪种更好?"

"That's another layer of decision boundary within edit itself — select the entire document then rewrite completely, or you want to have very targeted custom behavior."

"这就是 edit 本身的另一层决策边界——是'选中整个文档整段重写',还是'做精准的定制化修改'。"

"When we first launched the model, we would bias the model towards more rewrites because we thought the quality of the rewrites was much higher."

"我们刚上线那一版时,把模型偏向了'更多整段重写',因为我们判断重写的质量更高。"

"But over time, you're kind of shifting based on user feedback — that was the learning from iterative deployment."

"但随着时间推移,你会根据用户反馈调整方向——这就是迭代式上线带给我们的学习。"

"Lastly, the third behavior that we taught the model synthetically is how to make comments on any document."

"最后一种我们用合成数据教模型的行为,是怎么对任意文档做评论。"

"The way we used it is — we would use a o1 model to produce, to simulate user conversation. Let's say, 'write me a document about XYZ.' But then we used o1 to produce the document, and then we kind of injected a user prompt to be like, 'oh, make some comments, critique my piece of writing,' or 'critique this piece of writing that you just made.'"

"我们的做法是——用 o1 模型来模拟用户对话。比如'帮我写一份关于 XYZ 的文档',然后用 o1 把这份文档造出来,再注入一句用户提示'好,给我点评一下,挑挑这份文章的毛病'或者'挑挑你刚才写的这份的毛病'。"

"And then we taught the model to make comments on the document on very specific document sections. So it's also like — what kind of comments do you want the model to make? Do they make sense or not? How do you teach the quality of that?"

"接着我们就训练模型在文档中非常具体的位置上做评论。这里又涉及到:你希望模型做出什么类型的评论?这些评论合不合理?你怎么衡量评论的质量?"

"And it all came down to measuring progress via very robust evals. So yeah, this is how you use o1 for synthetic data generation for training."

"最后所有这些都回到一件事:用足够稳健的 evals 来度量进展。这就是我们怎么用 o1 做合成数据生成、用于训练。"

Lenny: "Okay, this is so interesting. So you talk about this idea of teaching the model, and you mention how it's using synthetic data to teach the model different behaviors."

Lenny:"等等,这个真的特别有意思。你提到'教模型',然后说你们是用合成数据去教模型不同的行为。"

"Is a simple way to think about it basically — that's where you do that by showing it what success looks like using basically evals? Here's what doing this successfully would look like, and that teaches it. Okay, I see — this is what I should do."

"是不是可以简单理解成——你通过 evals 给模型'示范'什么叫做成功?这就是它学会的方式:'哦,原来这就是我该做的样子'。"

Karina: "Yeah, great. Yeah, you got it."

Karina:"对,完全正确。你抓到了。"

Chapter 04

Evals Are the New PRD

Evals 是 PM 的新核心工作
evals · PM skills · spec design · model designer

Lenny: "I want to start unpacking what your day-to-day looks like as you're building these sort of things. Is it like you sitting there talking to some version of ChatGPT, crafting these evals?"

Lenny:"我想开始拆解一下你的日常工作——你做这些功能的时候,具体是什么样?是不是就是坐在那儿跟某个版本的 ChatGPT 对话,手写这些 evals?"

Karina: "Sometimes I do that. Sometimes I do sit — actually I think I learned this so much from Anthropic, is like people spend so much time just prompting models and quality eyeball-ing all the time."

Karina:"有时候是的。有时候我会坐下来——其实这件事我在 Anthropic 学到太多了:大家会花大量时间 prompt 模型,然后肉眼盯着看输出质量。"

"And you actually get a lot of new ideas of how do you make the model better. It's like, 'oh, this response is kind of weird, why is it doing this?' — and you start debugging, or you start figuring out new methods of how to teach the model to respond in a different way, like have better personality."

"这样你反而会冒出很多让模型变得更好的新想法。比如'咦,这条回答有点怪,它为什么这样回?'——然后你就开始 debug,或者琢磨出新方法教模型用不同方式回应,比如让它的人格更鲜明。"

"It's the same thing of how personality is made in the models. It's very similar methods."

"模型里'人格'是怎么造出来的,基本就是同一套方法。"

"But yes, I think my time has changed. When I first came I was mostly like research IC work. So I was building a lot, writing code, changing models, writing evals, working with PMs and designers to teach them how to even think about evaluations."

"但说实话,我的时间分配变化挺大的。刚来的时候我基本是研究 IC 的状态——写很多代码、改模型、写 evals,还要带着 PM 和设计师,教他们怎么去思考 evaluation 这件事。"

"I think that was really cool experience. And I think this is an adoption of how do we do this PM management of AI features or AI models."

"那是一段很酷的经历。它本质上就是在摸索:我们怎么用 PM 的方式来管理 AI 功能、AI 模型。"

"Now it's mostly like management and mentorship. I'm still doing IC research code after 4pm although. But yeah, it's just kind of changed."

"现在我做的事更多是管理和带人。不过下午 4 点之后我还是会自己写一些研究代码。整体的工作节奏已经变了。"

Lenny: "Alright, don't talk too much about being a manager because everyone's firing their managers — who needs managers anymore, that's what I hear now. Just kidding."

Lenny:"行,你别太多说管理的事——现在大家都在裁管理者,'还要管理者干嘛'是当下流行说法。开玩笑啦。"

"It's interesting that so much of your time was spent on teaching product teams how evals integrate and how important that is. I've heard this a few times and I haven't personally experienced it yet, so I think it's an important thread to follow."

"有意思的是,你之前大量时间都在教产品团队怎么把 evals 整合进流程、它有多重要。我已经听到好几次类似的说法,但我自己还没真正体验过,所以这条线值得跟下去。"

"How writing these evaluations is going to become increasingly an important part of the job of product teams, especially when they're building AI features and working with LLMs."

"也就是说,写 evaluations 会越来越成为产品团队工作的核心——尤其是当他们在做 AI 功能、和 LLM 一起干活的时候。"

"So can you just talk a bit more about what that looks like? Is it like sitting there with an Excel spreadsheet, basically showing 'here's the input, here's the output, here's how good the result was'? Talk about what that actually looks like very practically."

"能不能更具体地讲讲这是什么样?是不是就是开个 Excel,一列输入、一列输出、一列写'这个结果有多好'?你给我们一个非常具体的画面。"

Karina: "It certainly depends on what you're developing. But there are various types of evaluations."

Karina:"这要看你做的是什么。但 evaluation 其实有好几种类型。"

"So sometimes I do ask product managers — or there's also a new role that we have, like model designers — to kind of go through some of the user feedback, maybe, or think of various user conversations that should have triggered, under this scenario it should trigger Canvas."

"有时候我会让 PM——或者我们新设的'模型设计师(model designer)'这个角色——去翻一些用户反馈,或者构思各种用户对话,判断'在这种场景下该不该触发 Canvas'。"

"Then you have this ground-truth label of, okay, under this conversation it should trigger Canvas, under this conversation it should not trigger Canvas. And you have this very deterministic kind of eval that — for this behavior — is like this."

"然后你给每条对话打一个 ground-truth 标签:这条该触发 Canvas、那条不该。这样就形成一个非常确定性的 eval——这个 behavior 的判定标准就是这样。"

"When we were launching Tasks, for example — how do you make correct schedules is actually really hard for the model. But we built out some of the deterministic evaluations: if the user says like 7pm, the model should say 7pm."

"比如我们做 Tasks 的时候——让模型生成正确的日程其实非常难。我们就构建了一些确定性 eval:用户说 7pm,模型就得返回 7pm。"

"So if you can have different deterministic evals, whether it's pass or fail. So yeah, the way it works is — sometimes I ask PMs to just go create like a Google Sheet, have different tabs, like what's the current behavior, what's the ideal behavior, why, or some notes."

"于是你就有了很多条 pass/fail 的确定性 eval。具体怎么做——有时候我会让 PM 拉一个 Google Sheet,分不同标签页:当前行为、理想行为、原因、备注。"

"Sometimes you use it for evals, sometimes we use it for training. Because if you give the spreadsheet to an o1 model, it can probably figure out how to teach itself a good behavior."

"这表有时拿来做 eval,有时直接拿来训练。因为你把这张表丢给 o1,它大概率能自己琢磨出'怎么把这个行为学会'。"

"I think there are second type of evals that is kind of more prevalent — it's human evaluations."

"还有第二种 eval,其实更主流——人工评估。"

"You can have specific trainers, or you can have internal people. When you have a conversation, a prompt, and then various completions of models, you kind of choose the win rate — which model is the best, which model produced the highest-quality comment or edit."

"你可以请专门的 trainer,也可以发动公司内部的人。给定一句对话/prompt + 多个模型给出的回答,让人选哪个最好——哪个模型产出的评论/编辑质量最高,从而形成 win rate。"

"And then you can have continuous win rates. As you develop new models, it should always win over the previous models. So it depends on what you want to measure."

"这样你就有了持续的 win rate 指标。每次开发新模型,它都应该比上一代赢——具体怎么衡量,要看你想测什么。"

"Product development might move from 'here's a spec, let's build it' to 'hey AI, build this thing for me — and here's what correct looks like.'"

Lenny: "So interesting. Basically what I'm hearing, and something I'm learning as I talk to people, is — product development might move from this like 'here's a spec, PRD, let's build it together, then cool let's review it, are we happy with this' — from that to 'hey AI, build this thing for me, and here's what correct looks like. And I'm spending all my time on what does correct look like, on evals essentially.'"

Lenny:"太有意思了。我在和很多人聊的过程中也在学到这一点——产品开发的范式可能正在从'这是一份 spec/PRD,我们一起做,做完一起 review,大家都满意吗?',转变成'嘿 AI,把这个东西给我做出来,这里是正确长什么样'。然后 PM 的所有时间,都花在定义'什么叫对了'上,本质上就是写 evals。"

Karina: "You definitely want to measure progress of your model, and this is where eval is. Because you can have prompted model as a baseline already."

Karina:"你绝对希望能度量模型的进展,这就是 eval 的作用——你可以用一个 prompt 调得最好的模型当作基线。"

"And the most robust eval is the one where prompted baselines get the lowest score or something. Because then you know — oh, if you trained a good model, then it should hill-climb on that eval over time, while not also regressing on other intelligence evals."

"最稳健的 eval 是那种'prompt 基线得分最低'的——这样你才能确定:训出一个好模型后,它能在这条 eval 上爬坡,同时不会让其他智能维度退步。"

"So it's more like — that's what I'm saying — it's more an art than science. Like, okay, if you optimize the model with this behavior, you kind of don't want to brain-damage it in other areas of intelligence. And this is happening all the time in every lab, in every research team."

"所以我说它更像艺术,而不是科学。意思是:你为某个 behavior 做优化的时候,千万别把模型在其他智能维度上搞坏。每个实验室、每个研究团队每天都在和这件事博弈。"

"Prompting is also a way to prototype new product ideas."

"prompt 同时也是给新产品想法做原型的一种方式。"

Chapter 05

Prompting Is the New Prototyping

Prompting 就是新版 prototyping · Karina 在 Anthropic 的实验
prompting · prototyping · personalized starters · title generation

Karina: "Like, early days at Anthropic when I was working on the file uploads feature, I remember just prompting the model — and when we were launching like 100K context, I was just prompting this in my local browser."

Karina:"比如我早期在 Anthropic 做文件上传功能的时候,就是在本地浏览器里直接 prompt 模型——当时正赶上 100K 上下文上线,我就在那儿不停 prompt。"

"People really really loved it and they just wanted like API for file uploads or something. And then that's when it clicked to me — like, I also wrote a blog post a long time ago — it clicked to me like, prompting is a new way of product development or prototyping for designers and for product managers."

"用户特别喜欢,马上就有人问:有没有文件上传的 API?那一刻我突然意识到——我后来还写了一篇 blog 讲这事——prompting 其实是一种新的产品开发方式,也是设计师和 PM 的新型原型方式。"

"For example, one of the features that I wanted to do is — have a personalized recommended, like, personalized starter prompts. So whenever you come to Claude, it should recommend you starter prompts based on what your interests are."

"举个例子,我当时想做的一个功能是——个性化的、推荐式的起始 prompt。每次你打开 Claude,它就根据你的兴趣推荐起手 prompt。"

"And so you can literally do it by prompting for that experiment. Another feature was like generating titles for the conversations."

"你完全可以直接用 prompt 把这个实验跑起来。另一个功能是给对话自动生成标题。"

"It's a very small micro-experience, but I'm really proud of it. The way we did that was — we took the five latest conversations from the user, asked the model 'what's the style of the user,' and then for the next new conversation the generated title would be of the same style."

"这只是个很小的 micro-experience,但我挺为它骄傲。我们的做法是:取用户最近 5 段对话,问模型'这个用户的风格是什么',然后新对话的标题就沿用这种风格。"

"It's really little micro-experiences like this."

"就是这种很小、很微的体验细节。"

Lenny: "That's so cool. Did you do that at Anthropic or at OpenAI?"

Lenny:"太酷了。这是你在 Anthropic 还是 OpenAI 做的?"

Karina: "At Anthropic."

Karina:"在 Anthropic。"

Lenny: "Okay, cool. I love the file upload feature that Claude has, by the way. Oh, ChatGPT doesn't have that yet, is that right?"

Lenny:"挺好。顺带一提,我特别喜欢 Claude 的文件上传功能。哦,ChatGPT 还没这个功能?"

Karina: "I think it has — but I think the way it's implemented is very different though."

Karina:"应该是有的——但实现方式很不一样。"

Lenny: "Okay, maybe it's the PDF feature. Cuz I use it all the time with Claude. Okay, that's cool. Someone needs to get on that."

Lenny:"好吧,可能我说的是 PDF 那个功能。我在 Claude 里超频繁地用。OK,挺有意思的,得有人把这个补上。"

"Man, it's wild how many features you built that I use every day and that many people use every day."

"我用每天的功能里,有不少是你做的——这点真的挺神奇的。"

"Prompting is a new way of product development or prototyping for designers and product managers."

"This prototyping point you made is really important. It's something that comes up a ton on this podcast also — how that is maybe the way that AI has most impacted the job of product builders recently is just prototyping."

"你提到的这个'原型'的点,真的非常关键。这也是我节目里反复出现的话题——AI 对产品人最直接的冲击,可能就是把原型环节彻底改写了。"

"Instead of going from showing just 'here's a PRD, here's a design,' PMs more and more just say 'here's the prototype of the idea that I have and it's working — you can play with it.'"

"过去 PM 是'我给你 PRD,给你设计稿',现在越来越多人直接说:'我有一个想法,这是已经能跑的原型,你来玩玩看。'"

Karina: "Yeah, yeah."

Karina:"对,对。"

Chapter 06

How Tasks Went 0 → 1 in Two Months

Tasks 怎么 2 个月从 0 到 1 · staffing 与 spec 设计
tasks · 0-to-1 · staffing · tool spec · JSON schema

Lenny: "Okay, I want to spend a little more time on how you operate. So you talked about you built and launched this Tasks feature. Talk about how that emerged and let's better understand just how you collaborate with product teams and how OpenAI works in that way, whatever you can share there."

Lenny:"OK,我想再多聊聊你的工作方式。你提到你做了 Tasks 这个功能并上线了。给我们讲讲它是怎么诞生的——更具体一点,你怎么和产品团队协作?OpenAI 在这件事上是怎么运作的?能讲多少讲多少。"

Karina: "I think Canvas and Tasks are going into the bucket of like projects where it's like more short- or medium-term."

Karina:"Canvas 和 Tasks 我觉得都属于'中短期项目'这一类。"

"Actually the way Canvas and Tasks came about was — it started with one person prototyping and creating a spec. It's kind of like a PRD — creating a spec of the behavior of the model."

"实际上 Canvas 和 Tasks 的起点都是——一个人开始做原型,写一份 spec。这份 spec 有点像 PRD——但写的是模型行为的 spec。"

"I don't think Tasks is extremely groundbreaking necessarily. What makes it really cool is because the models are so general — the model can search, they can write sci-fi stories, they can search for stocks, they can summarize the news every day."

"我倒不觉得 Tasks 多么颠覆性。它真正酷的地方是——因为模型已经太通用了,可以搜索、写科幻故事、查股票、每天给你新闻摘要。"

"Because the models are so general, giving something familiar to people — like notifications is very familiar, like having reminders is very familiar — so creating a form factor for the people who very familiar things like models can build is very familiar. But then you add a magical AI moment and it becomes very powerful."

"正因为模型已经这么通用,你把它套进一个大家熟悉的 form factor——比如通知、提醒——立刻就好用。再叠加一个'魔法般的 AI 时刻',它就极有威力。"

"The way it comes usually operationally is — yeah, sometimes it's just a prototype. Like, a literally prompted prototype of how you would want the model to behave for Tasks, for example."

"实际操作上是怎么发生的呢?往往就是——一个原型。比如对 Tasks 来说,就是一个完全靠 prompt 跑出来的原型,展示'你希望模型在 Tasks 场景里如何表现'。"

"You kind of need to design a little bit of design systems, design thinking. It's like — okay, well, if the user says 'remind me to go to lunch at 8am tomorrow,' what kind of information does the model need to extract from that prompt in order to create a reminder?"

"这里要做一点设计系统、设计思维的工作。比如:用户说'明天早上 8 点提醒我吃午饭',模型需要从这句话里抽出哪些信息,才能生成一条提醒?"

"And so this is how you design a spec for a new feature, like a tool. Canvas and Tasks are all tools."

"这就是给一个新功能(或者说一个工具)设计 spec 的过程。Canvas 和 Tasks 都属于'工具(tool)'。"

"So it's how do you create the tool stack? And then it's mostly developing a JSON schema. So okay, from this prompt maybe the model should extract the time that the user requested. And then you think about which format you want the time to be in."

"所以这件事就是'怎么搭这个工具栈?'答案大多是写一份 JSON schema。比如:从用户的 prompt 里,模型应该抽出用户指定的时间;接着你考虑'时间应该是什么格式'。"

"And then how do you want the model to notify you? Basically the user should give instruction to the model, and then this instruction would fire off every day or something at that particular time."

"接着是'你希望模型怎么通知你'?基本逻辑是:用户给模型一条指令,然后这条指令会在你指定的时间(比如'每天'+'某个具体时刻')自动触发。"

"For example, if you say 'every day I want to learn about the latest AI news,' the model should write into 'search for the latest AI news' — and this task will get fired at that particular time that the user requested."

"比如你说'每天我都想了解最新 AI 新闻',模型就该把这件事翻译成'搜索最新 AI 新闻',然后在你指定的时间触发这个任务。"

"And then your design is like the tool spec. And actually — I don't know, sometimes it's through conversations. Like, people ask me to join the team, and they're like 'oh my god, we need researchers, or we need some support, we need to train the models.'"

"于是你的设计稿就是这份 tool spec。说实话——有时候这件事就是通过对话发生的。会有人来邀请我加入某个团队,说'我们需要研究员、需要支持、需要训练模型'。"

"Sometimes — like with Canvas — it was mostly like I just pitched the idea. It got staffed quite immediately."

"有时候——像 Canvas 那次——就是我把想法 pitch 出去,立刻就有人来站队、组队。"

"It depends on the project. And usually with staffing, it's mostly a product manager, a model designer, an actual product designer, a couple of researchers, and a bunch of applied engineers, depending on the complexity of the project."

"具体看项目。staffing 一般是这样:一个 PM、一个 model designer、一个真正的产品设计师、若干研究员、再加一批应用工程师,人数看项目复杂度。"

"For Tasks, it took roughly two months to go from zero to one. For Canvas, this was like four to five months."

"For Tasks, it took like — I don't know — two months or so to go from zero to one, basically."

"Tasks 大概用了两个月从 0 到 1。"

Lenny: "Oh wow."

Lenny:"哇。"

Karina: "For Canvas, this was like four to five months, I guess, to go from zero to one."

Karina:"Canvas 从 0 到 1 大概用了 4 到 5 个月。"

"And then you teach product managers how to build evals, and how do we not only ship a better feature but how do we think more longer term — what kind of cool features did you want Tasks to have? Like, I think it would be nice for Tasks to be a little more personalized. It'd be nice to have to create tasks via voice on the mobile, right?"

"接下来就是教 PM 怎么写 eval,以及怎么从更长线视角思考——'Tasks 接下来还能长出什么酷的功能?'比如我觉得 Tasks 应该更个性化,应该能在手机端用语音创建任务。"

"So you kind of need to — this is how you get research roadmaps right here. It's thinking about how the feature will be developed in the future. And then from there it's like you start creating datasets."

"研究的 roadmap 就这么生成出来——从'这个功能未来怎么演化'倒推出'现在该准备什么数据集'。"

"With evals, you want to make sure that goes well. And then you need to have a tradeoff between what methods you want to use."

"evals 这一块,你得保证它跑得稳。然后你还要在'要不要采用哪种方法'之间做取舍。"

"And the reason why I really love relying purely on synthetic data instead of collecting data from humans is because it's much more scalable. It's cheap. You literally sample from the model, and you teach the quality behaviors of the models, and that will generalize to all sorts of diverse coverage."

"我之所以特别偏爱纯合成数据、而不是从人类那里收集——因为它的可扩展性强得多,还便宜。你直接从模型里采样,教它好的行为,这种能力就会泛化到各种多样的场景。"

"When you launch the better feature, you learn so much from the users that all your synthetic sets can be shifted in the distribution of how the users behave on the product behavior. And this is how you improve."

"当你把更好的版本上线后,你会从用户身上学到大量东西——所有的合成数据集都可以朝着'用户在产品上的真实行为分布'去靠。改进就是这么发生的。"

"And this is what happened with Canvas, too — when we went from beta to GA."

"Canvas 上线时也是这个套路——我们就是这样从 beta 走到 GA 的。"

中间略去 Loom 的口播 sponsor。

Lenny: "Something that I want to help people understand — and I don't even 100% understand this — is what's the simplest way to understand the job of a researcher versus, say, a model designer and other folks involved? Like what's the simplest way to understand what researchers do at OpenAI?"

Lenny:"有一件我想帮听众弄清楚的事——其实我自己也没完全搞懂——'研究员'和'模型设计师'以及其他相关角色,最直观的差别是什么?在 OpenAI 里,研究员到底做什么?"

Karina: "The projects that I described are mostly like product-oriented. So research is mostly like product research."

Karina:"我刚才描述的那些项目主要是'产品导向'的,所以那一类研究本质上是'产品研究(product research)'。"

"Another component of my team is actually more like longer-term exploratory projects, and it's more about developing new methods, understanding those methods under a variety of circumstances."

"我团队的另一个组成,是更长线的探索性项目——重点是发明新方法,以及在各种条件下理解这些方法的边界。"

"To develop new methods, you kind of need to follow very similar kind of recipe of building evals, but it's much more sophisticated evals — you want to have out-of-distribution evals. If you want to measure generalization, you need to capture that."

"要做新方法,流程跟做 eval 类似,但 eval 本身更复杂——你需要'分布外(out-of-distribution)'的 eval。如果想衡量泛化能力,就必须把这一点捕捉到。"

"It's more sciencey in a way, where if we talk about synthetic data — one of the hardest things about synthetic data is how do you make it more diverse. Diversity in synthetic data is one of the most important questions right now, and just exploring ways to inject diversity as a general method that will work for all — that's one of the research explorations."

"这一块更偏'科学'。比如合成数据,最难的问题之一就是怎么让它更多样。'多样性'是当下 synthetic data 里最重要的问题之一——研究探索的目标之一,就是找到一种'放之四海而皆准'的多样性注入方法。"

"Other ones are more like developing new capabilities."

"另一类研究探索是发展全新能力。"

"I feel like it's all about — you work on a new method, you have signs of life that it's working, either you think of how to make it more general or how to make it very useful. And this is how longer-term projects become more like medium- or short-term projects."

"我的感受是——你做一个新方法,看到'有迹象在 work',然后就开始想'怎么让它更通用'或者'怎么让它更有用'。长线项目就是这样一步步变成中短期项目的。"

Lenny: "That makes sense. Essentially working on developing ways to make the model smarter — like o4, o5, o6. Like, o1 was a big breakthrough, right? The way it operates — it's not just 'here's your answer,' it actually thinks and takes time to think through the process of coming up with an answer."

Lenny:"懂了。本质上就是研究'怎么让模型变得更聪明'——o4、o5、o6 这种。o1 是一次很大的突破——它的工作方式不是'直接给你答案',而是真的会先思考、花时间把答题过程走一遍。"

Chapter 07

The Cost of Intelligence Is Collapsing

Intelligence 成本在崩塌 · 小模型变聪明,AI 普惠
distillation · small models · healthcare · education

Lenny: "Speaking of that, of thinking about the future where things are going — I want to spend some time on just this insight that basically you are building the cutting edge of AI. Like at the very bleeding edge of where AI is going and where it is."

Lenny:"既然聊到未来——我想多花点时间在一个洞察上:你本质上就在亲手造 AI 的最前沿。'AI 要去往哪里、现在到了哪'——你就站在那个边缘上。"

"I'm very curious to hear just your take on how you think things are going to change in the world, and how people work, based on where you see things are going."

"我特别想听你说说,你判断的世界变化方向是什么?基于你看到的趋势,人们的工作方式会怎么变?"

"I know it's a broad question, but let's say in the next three years — how do you see the world changing, how do you see people's way of working changing?"

"我知道这个问题很大。但就说接下来三年——你怎么看世界的变化、人的工作方式的变化?"

Karina: "It's a very humbling experience to be in both labs, I guess. To me, when I first came to Anthropic, I was like 'oh God, I really love front-end engineering.' And then the reason why I switched to research is because I realized at that time — oh my God, Claude is getting better at front-end. Claude is getting better at coding. I think Claude can develop new apps, or develop new features for the thing that I'm working on."

Karina:"在这两家实验室待过都让我觉得'非常 humbling'。我刚到 Anthropic 的时候心想'天啊,我真的太喜欢做前端工程了'。后来我转去做研究,就是因为那一刻意识到——Claude 在前端越来越强,在写代码上也越来越强。我觉得它能开发新的 app,或者开发我正在做的功能。"

"So it was kind of this meta realization where it's like — oh my God, the world is actually changing."

"于是有了一种'meta 觉醒'——这个世界真的在改变。"

"When we first launched 100K context, at that time obviously I'm thinking about form factors. It's like — yeah, file uploads were very natural, very familiar to people. But you could imagine we could just make infinite chats in the Claude AI app, right? As if it's like in 100K context."

"我们刚发布 100K context 的时候,我自然就开始想 form factor。文件上传对用户来说是非常自然、非常熟悉的形态。但你也可以想象——在 Claude AI app 里直接做无限长的聊天,只要 100K 上下文撑得住。"

"But because file uploads — it's like form follows function — the form factor of file uploads kind of enabled people to literally upload anything: books, reports, financials, and ask any task to the model."

"不过文件上传遵循'form follows function'的逻辑——它这种 form factor 让用户能上传几乎任何东西:书、报告、财务数据,然后给模型派任何任务。"

"I remember enterprise customers — like financial customers — were really interested in that. It was like, 'oh wow, actually this is one of the very common tasks people do in that setting.' It was kind of crazy to see how some of the redundant tasks were getting automated basically by these smart models."

"我记得当时企业客户——尤其是金融客户——对这件事特别感兴趣。'天哪,原来这就是这一行人天天在做的事。'看着那些重复性的工作被聪明的模型自动化掉,挺震撼的。"

"And we're entering an era where I actually don't know, for example, sometimes if o1 gives me the correct answer or not — because I'm not an expert in that field. I don't even know how to verify the output of the models — because only experts know that. They can verify this."

"现在已经进入了一个状态——很多时候 o1 给我答案,我自己都判断不了对不对,因为我不是那个领域的专家。我甚至不知道该怎么验证模型的输出——只有领域专家才能验证。"

"So basically there are trends that are going on. The first trend is the cost of reasoning and intelligence is drastically going down."

"所以这里有几个趋势。第一个趋势:推理与智能的成本正在断崖式下降。"

"I had a blog post about this. Maybe I should update on the latest benchmarks, because at that time MMLU, everybody was doing like one benchmark and then quickly saturating the benchmark — and now to do the same plot you have to use another frontier eval."

"我写过一篇 blog 讲这件事。也许该用最新的 benchmark 把它更新一下——之前大家都盯着 MMLU 这种 benchmark,很快就被打满了,现在画同样的曲线就得换更前沿的 eval。"

"But the cost of intelligence is going down because — it becomes much cheaper. Small models are becoming even smarter than large models, and that's because of the distillation research."

"但智能的成本确实在降——而且变得越来越便宜。小模型已经比一些大模型还聪明,这要归功于蒸馏方面的研究。"

"This happened with Claude 3 Haiku — I was working on like Claude 3 Haiku, and I realized it was much smarter than Claude 2, which was way bigger."

"Claude 3 Haiku 就是这样的例子——我参与了 Claude 3 Haiku,我意识到它比体量大得多的 Claude 2 还要聪明。"

"All the work that has been bottlenecked by intelligence will be unblocked."

"The power of small models becoming very intelligent and fast and cheap — we are moving towards that road. That has multiple implications. It means that people will have more access to AI. And that's really good — like builders and developers will have much better access to AI."

"小模型变得又聪明、又快、又便宜——我们正在走这条路。它有好几层含义。第一层是,人们将得到更强的 AI 接入——这非常好,builders 和开发者尤其受益。"

"But also it means all the work that has been bottlenecked by intelligence will be unblocked."

"另一层含义是:那些过去因为'智能不够'而被卡住的所有工作,都会被解锁。"

"So anyone like — I'm thinking about healthcare, right? Like, if I have — instead of going to the doctor, I can ask ChatGPT, or give ChatGPT a list of symptoms, and ask it like, 'oh, which would I have — a cold, flu, or something else?' Like, I can literally get the access to a doctor almost. And there has been some research studies around that."

"举个例子,医疗——我不一定非要去看医生,我可以问 ChatGPT,或者把症状列出来问它'我是感冒、流感还是别的什么?'这几乎就是在拿到一个医生级别的咨询。已经有研究在追踪这个现象。"

Lenny: "Yeah, there was a New York Times story about that, where they compared doctors to doctors using ChatGPT to just ChatGPT — and just ChatGPT was the best of them all. Like doctors made it worse."

Lenny:"对,《纽约时报》有篇报道,对比了三组——医生、医生+ChatGPT、纯 ChatGPT——结果纯 ChatGPT 是最好的。医生反而把效果拖低了。"

Karina: "Yeah, that's crazy. Like, right — education, I think. I would have dreamed if I had a tool like ChatGPT when I was young, and I would learn so much."

Karina:"对,挺吓人的。再一个就是教育——我小时候要是有 ChatGPT 这样的工具,我能学到的东西不知道多多少。"

"People can now learn almost anything from these models. They can learn new languages, they can learn how to build new apps. Like, anything that you want. And I'm so — it's humbling to launch Canvas, and bring that thing to people, enable them to do something else that they couldn't have ever before."

"现在人们可以从这些模型身上学几乎任何东西——学新语言、学怎么写 app,任何你想学的。所以做出 Canvas 并把它送到用户手里,让他们做以前做不到的事——这种感觉让我谦卑。"

"I think this is something magical around this experience."

"这件事本身带着一种魔力。"

"Education has massive implications. Scientific research, right? I think the dream of any AI researcher is to automate AI research. It's kind of scary, I'd say."

"教育这件事的影响是巨大的。再就是科研——任何 AI 研究员的终极梦想,就是让 AI 自动化 AI 研究。说实话,这也有点吓人。"

"Which makes me think that people management will stay. It's like one of the hardest things — it's emotional intelligence for the models, or creativity in itself. These are some of the hardest things."

"这也让我觉得'人的管理'会留下来——这是模型最难学的东西之一:模型的情感智能、创造力本身——这些是最难训的。"

"Writers, I don't think people should be worried as much. I think it elevates a lot of redundant tasks for people."

"作家这一类,我觉得没必要那么慌。AI 主要是把那些重复、无聊的部分抹掉。"

Lenny: "This is awesome. Okay, I want to follow this thread for sure. And it's funny that what you described is — you were an engineer at Anthropic, and you're like, 'okay, Claude is going to be very good at engineering, this isn't going to be a potentially career long term, so I'm going to move into research and AI is going to need me for a long time to build it, make it smarter.'"

Lenny:"太精彩了,这条线我一定要跟下去。有意思的是你刚才描述的——你之前是 Anthropic 的工程师,然后判断'Claude 在工程上会变得很强,这条职业线长期靠不住,所以我转研究,AI 这条线还会长期需要我去把它做更聪明'。"

Karina: "I would say we still have, I think the Canvas team has really cool front-end engineers — people who really care about interaction design, like interactive spirit. I don't think models are there yet. But we can get the models to like top 1% of front-end, for sure."

Karina:"补充一下,Canvas 团队现在也还有非常厉害的前端工程师——那种真心在乎交互设计、'交互的灵魂'的人。我不觉得模型已经到位了。但模型有望抵达前端工程师里 1% 的水平,这是肯定的。"

Chapter 08

Soft Skills Are the New Hard Skills

Soft skills 才是新的 hard skills · 创造、品味、管理
creativity · taste · management · prioritization · empathy

Lenny: "What I want to move on to next along these lines is just — and this is just speculation, but — what skills do you think will be most valuable going forward for product teams in particular?"

Lenny:"沿着这个线索我想往下问——纯粹是推测——你觉得对产品团队来说,接下来最值得练的技能是什么?"

"So folks are listening and they're like 'okay, this is scary, what should I be building now to help me stay ahead and not be in trouble down the road?' What skills do you think are going to be more and more important to build?"

"听众听到这里可能在想:'有点慌啊,我现在该练什么,才能不掉队、未来不至于陷入麻烦?'你觉得哪些技能会越来越重要?"

Karina: "Yeah, I think creative thinking. Like, you kind of want to come up — generate a bunch of ideas, and filter through them, and know — just like build the best product experience."

Karina:"我觉得是创造性思考。你要能生成一大堆想法,然后筛选,做出最好的产品体验。"

"Listening. You want to build something that the most general model will not replace. And oftentimes you build something and you make it really really good for a specific set of users."

"听用户。你想做'通用模型替代不了'的东西——通常是把某一类细分用户群伺候到极致。"

"And actually the moat is now in your user feedback. The moat is more in like — whether you listen to them, whether you can rapidly iterate. The moat is in here."

"今天的护城河就在用户反馈里——你听不听他们的话?能不能飞速迭代?护城河在那。"

"I don't think we are yet to think — there's so many ideas, there's an abundance of ideas that you can do, regardless. Like, I wouldn't be worried."

"我觉得我们其实还远没把所有点子都想完——可做的想法多到爆炸。我反而不太焦虑。"

"I feel like in fact, I do think people in AI field are — I wish they were a little more creative and like connecting dots across different fields or something like that — to develop really cool new generation, like a new paradigm of interactions with this AI."

"说实话,我反而希望 AI 圈的人能更有创造力一点——跨学科把点连起来,做出与 AI 交互的全新一代范式。"

"I don't think we've cracked this problem at all."

"这件事我们其实远远没解决。"

"A couple years ago I was telling some people — you kind of want to build for the future. So it doesn't necessarily matter whether the model is good or not good right now, but you can build product ideas such that by the time the models will be really good, it will work really well."

"几年前我就跟一些人说——你要为未来的模型造产品。当下模型好不好其实没那么重要,但你可以设计这样的产品:等模型真的强大,它就天然 work。"

"And I think it just happened naturally — for example, at Anthropic, right, the Claude artifact and I feel like early days of Canvas, like back in 2022. Like before ChatGPT, writing IDE was like unheard of."

"这种事其实在自然地发生。比如 Anthropic 的 Claude Artifact,以及 Canvas 早期想法——大概在 2022 年。在 ChatGPT 之前,'写作 IDE'这种东西基本没人听说过。"

"But I feel like Claude 1.3 model itself was not there to make really good high-quality edits, for example like coding."

"但当时 Claude 1.3 模型本身,在做高质量编辑这件事——比如写代码——还没到位。"

"And I feel like I see startups like Cursor, and it's doing super well — and that's because they iterate so fast, they invent new ways of training models, they move really fast, they listen to users, massive distribution. It's kind of cool."

"现在我看 Cursor 这种创业公司,跑得超级好——原因就是他们迭代极快、自创训练模型的方法、动作飞快、听用户、铺得开。挺酷的。"

Lenny: "That's really helpful actually. So what I'm hearing is that soft skills essentially are going to be more and more important — powerful. You talked about management, leading people, being creative, coming up with innovative insights, listening."

Lenny:"挺受用的。所以你的意思是——soft skills 会越来越重要、越来越有力。你提到了管理、带团队、创造力、独到洞察、倾听。"

"There's a post I wrote that I'll link to, where I tried to analyze what AI — how AI will impact product management. And we're actually very aligned. My sense was the same thing — that soft skills are going to become more and more important. And the things that are going to be replaced are the hard skills."

"我写过一篇 post,等会儿链上来——我尝试分析 AI 会怎么冲击产品经理这份职业。咱俩的结论挺一致——soft skills 会越来越重要,而被替代的恰恰是 hard skills。"

"Which is interesting because usually people value the hard skills — like coding, design, writing really well. And it's interesting that AI is actually really good at that because it's taking a bunch of data, synthesizing it, and writing/creating a thing."

"有意思的是——大家通常更重视 hard skills,写代码、设计、写得好这些。但 AI 恰好在这些上很强,因为它本质上就是吃数据、综合、产出东西。"

"Versus all these fuzzy things — what influences/convinces people to do things, aligning, and listening like you said, creativity. Anything along those lines come up as I say that?"

"反过来,那些模糊的东西——影响力、说服力、对齐、像你说的'倾听'、创造力——你听到这些有共鸣吗?"

"It's actually really, really hard to teach the model how to be aesthetic, or how to be extremely creative in the way they write."

Karina: "I think it's actually really really hard to teach the model how to be aesthetic, or to do really good visual design, or how to be extremely creative in the way they write."

Karina:"教模型审美、做出真正好的视觉设计,或者把写作做到极致创意——这件事真的非常非常难。"

"I still think ChatGPT kind of sucks at writing. And that's because it's bottlenecked by this creative reasoning."

"我现在还是觉得 ChatGPT 写作能力很一般。这是因为它被'创造性推理'卡住了。"

"I think prioritization is one of the most important. For managers, I feel like actually AI research progress is bottlenecked by management — research management — because you have a constrained set of compute, and you need to allocate the compute to the research bets that you feel most convinced about."

"我觉得 prioritization 是最重要的之一。对管理者来说,我认为 AI 研究的进展其实就是被'管理'卡住的——研究管理——因为你的算力是稀缺的,必须押在你最有信念的研究方向上。"

"You need to really have a really high conviction in the research bets to put the compute. It's more like return on investment kind of situation."

"你必须对你下注的研究方向有非常高的信念,才舍得砸算力进去。这本质上是 ROI 决策。"

"It's like, okay yeah — I'm thinking a lot about, like, okay, across all my projects, which projects are higher priority? Prioritization. And also on the lower level — which experiments are really important to run right now and which are not, and like cut through the line."

"我每天都在想——所有这些项目里哪几个优先级最高?这是 prioritization。在更细的层面也是——哪些实验现在该跑、哪些可以砍掉?"

"So I see prioritization, communication, management, people skills like empathy, like understanding people, like collaboration. I think Canvas wouldn't have been an amazing launch if it wasn't about people."

"所以我看重 prioritization、沟通、管理、人际能力(共情、理解他人)、协作。Canvas 之所以做成,不是因为模型有多强,而是因为人对了。"

"It's a wonderful, good group of people, and I got a chance to work with people like Lee Byron — who's like a co-creator of GraphQL — and some of the best Apple designers. It's so cool to see — how do you create this collaboration between people? It's just something that's still humane, I think."

"那是一群很棒的人,我有机会跟 Lee Byron(GraphQL 联合创始人)以及一群前苹果顶级设计师合作。'怎么把这种人和人之间的协作拧出来'——这件事依然是非常'人'的。"

Lenny: "Let me just follow this through a little bit, because I imagine people listening are like 'okay, but once we have AGI or SGI, it's like it'll do all this.' There's a world where, like, why isn't all this done? I think it's easy to just assume all that."

Lenny:"我再往下追一下。我猜听众会想:'OK,但等我们有了 AGI 或者 SGI,这些都会被 AI 做掉吧?'好像很容易就默认这种结局。"

"I'm curious — this idea of creativity and listening — why you think AI isn't good at it, other than 'it's just very hard to train it to do this well.' Is there anything there of just like why this is especially difficult for AI/LLMs to get good at?"

"我想问——创造力和倾听这两件事,你觉得为什么 AI 还不擅长?除了'训练这件事很难'这个回答之外,有没有更深层的原因解释为什么这对 AI/LLM 特别难?"

Karina: "I think currently it's difficult for many reasons. I think it's still an active research area, something my team is working on — okay, how do we teach the model to be more creative in writing?"

Karina:"目前难,有好几层原因。这本身还是个活跃的研究领域,我团队就在做——怎么教模型在写作上更有创造力。"

"Actually, I'm thinking this new paradigm — like models that think more — should actually lead to better writing in itself."

"实际上我觉得这种新范式——'让模型多思考'——应该会带来写作上的提升。"

"But when it comes down to idea generation or discriminating what is a good visual design and not — I think if it hasn't had learned examples from people to discriminate it very well — I do think it's because there are not that many people who are like actually really — it's not accessible to model to learn from these people, I guess."

"但落到'生成创意'或者'分辨什么是好的视觉设计'——如果模型没有从足够多真正有 taste 的人身上学过例子,它就分不出来。说实话,这种有 taste 的人本来就不多——模型也没机会大规模从这些人那学到东西。"

"So I guess that's why it sucks."

"所以它在这方面就是不行。"

Lenny: "Yeah, that makes sense. Basically there's not enough of you yet — researchers training, teaching it to do these things, slash people that have incredible taste and creativity that can teach these things. You could argue this will come, but we don't need to keep going down that thread."

Lenny:"懂了。本质上就是'你这种人'还不够——既缺会训练它的研究员,也缺真正有 taste 和创造力的人来教它。可以说这些将来会有,但我们不用继续往这边追。"

Chapter 09

Yes, Strategy Gets Automated Too

Strategy 也会被 AI 接走 · 别幻想
strategy · self-improving models · data analysis · scientific research

Lenny: "Let me ask you a specific question. In this post I wrote, I made this argument that a lot of people disagreed with — that strategy is something that AI tooling will become increasingly great at and take over."

Lenny:"我有个具体问题想问你。我写过一篇 post,里面提了一个被很多人反对的观点——战略,是 AI 工具未来会越做越好、最终接管的事情。"

"There's the sense that — that's a thing that people will continue to be much better at, and you can't offload to AI basically developing your strategy, telling you what to do to win."

"主流的感觉是——战略是人类一直更强的领域,你没法把'制定战略、告诉我怎么赢'这件事交给 AI。"

"My case is — isn't strategy just take all the inputs, all the data you have available, understand the world around you, and come up with a plan to win? Feels like AI would be — an LLM would be incredibly smart at this. What's your take?"

"我的论点是——战略不就是'接住所有输入、所有可用数据,理解周围世界,再想出一个赢的方案'吗?LLM 在这件事上理应会很强。你怎么看?"

Karina: "I think so too."

Karina:"我也这么觉得。"

"I think — again, like, you teach the model all sorts of tools and capabilities and reasoning, right? And it's like — when it comes down to, like, as is for Canvas right now, would have been very cool to the model to just aggregate all the feedback from users, like the top five most painful user flows, like user experiences."

"你教模型一堆工具、能力、推理。比如对当下的 Canvas,如果模型能自动把所有用户反馈聚合起来,直接告诉你'前 5 个最痛苦的用户路径、最糟糕的体验',那就太有价值了。"

"And then the model itself is very capable of, like, thinking of — knowing how it's being made, figuring out how to create the datasets for itself to train on it."

"而且模型自己其实非常擅长这个——它知道自己是怎么被造出来的、能想清楚要训哪类问题就该构造什么样的数据集。"

"I don't think we are far away from self-improving models."

"And I don't think we are far away from that kind of self-improvement models becoming, like, self-improved. Like then, part of development is basically they kind of self-improve. It's kind of like its own organism or something."

"我不觉得'自我改进的模型'还离我们很远。一旦它们能自我进化,产品开发就变成——它们像自己有机体一样,自己迭代自己。"

"Yeah. Again, like — strategy is more like data analysis and coming up with — like, I think what models are really good at is connecting the dots."

"战略本质上更像数据分析+方案生成——模型最强的能力之一就是'连点成线'。"

"I think it's like, okay — if you have user feedback from this source, but you also have an internal like dashboard with metrics, and then you have other kind of feedback or input — and then it can co-create a plan for you, like recommendations even."

"比如:你给它一份这边的用户反馈、那边的内部 dashboard 指标、再加一些别的反馈或输入——它就能和你一起生成一份计划,甚至直接给你推荐。"

"I think this is, like, one of the most common use cases for ChatGPT — coming up with this sorts of things."

"这其实已经是 ChatGPT 最常见的用法之一——做这种'综合+建议'。"

Lenny: "That makes sense. Essentially a human can only comprehend so much information at once, and look at so much data at once, to synthesize takeaways. And as you said, these context windows are huge now. Here's all the information — what's the most important thing I should do?"

Lenny:"完全对。人一次能消化的信息量是有限的,数据量也是有限的——而 context window 现在已经很大了。把所有信息丢给它:'我接下来最该做的是什么?'"

Karina: "Yeah. Same as like scientific research, because, like — ideally the model would be able to like suggest ideas, like new ideas, iterate on the experiment, or given the empirical results of the previous experiments — like how do you come up with new ideas or methods?"

Karina:"对。科研也一样——理想状态是,模型能提出新想法、迭代实验、根据上一次实验的结果反推出新方法或新思路。"

Lenny: "Oh man. Okay, so just to close the loop on this conversation, this part of the thread is — the skills you're suggesting people focus on building and leaning into are soft skills — like creativity, managing, influence, collaboration, looking for patterns. Is that generally where your mind is at?"

Lenny:"天哪。OK 把这条线收个尾——你建议人们专注练习的,是 soft skills——创造力、管理、影响力、协作、看出规律。这是你大致的判断对吗?"

Karina: "Yeah. I'm thinking a lot about — like, how do we make collaborations more effectively. I think this is most of like management I guess. Like, how do you organize like research teams or generally teams — like combine compose teams such that they will be at their maximally succeed, like at the maximum performance of what can possibly — like we can literally create like the next generation of computers."

Karina:"对。我现在花很多时间在思考——怎么让'协作'变得更高效。这其实就是管理的大部分——你怎么把研究团队、或者任何团队拼起来,让他们能跑出最高水平?他们有能力造出下一代计算机,前提是你把人组合对。"

"It's just like the matter of connection and the way you manage through that. It's like scaling organizations or scaling product research, I guess."

"这本质上就是'连接的艺术',和你怎么管理这件事。就是组织扩张,或者产品研究的扩张。"

Lenny: "Yeah. I think what you're basically building this thing — and not efficiently doing it is limiting the potential of the human species right now. Mismanagement within the research team at OpenAI, Anthropic, and some of these other models."

Lenny:"对。本质上你在造的这个东西——如果做得不够高效,就在限制人类物种当下的潜能。OpenAI、Anthropic 这些公司的研究团队,如果管理跟不上,人类的天花板就被压低了。"

Karina: "Yeah. It's kind of crazy to think about it. Holy moly."

Karina:"对,想想真挺夸张的。"

Chapter 10

Anthropic vs OpenAI, Operator, and Westworld

两套文化对比 + 早期 Anthropic 回忆 + Operator agent
anthropic vs openai · claude in slack · operator · multimodal · lightning round

Lenny: "Okay, so speaking of Anthropic and OpenAI — you've worked at both, very few people have worked at both companies and seen how they operate. I'm curious just what you've noticed about the differences between these two — how they operate, how they think, how they approach stuff. What can you share along those lines?"

Lenny:"既然提到 Anthropic 和 OpenAI——你两家都待过,这一点上几乎没人比你更有发言权。我想听你聊聊你看到的差异:它们怎么运作、怎么思考、怎么做事。能分享多少分享多少。"

Karina: "It's more similar than different. Obviously there is a lot of — there are some differences, also comes down to nuances. Coming to culture, I really love Anthropic and I have a lot of friends there. And I also love OpenAI and still have a lot of friends though. So it's not about enemies."

Karina:"相似远大于差异。当然两边也有差别,而且常常体现在细节里。说到文化——我真的很爱 Anthropic,那边有很多朋友;我也爱 OpenAI,那边也还有很多朋友。这件事不是'敌对阵营'。"

"I feel like there's like in the eye — all yeah, the competitors there like enemies — but it's actually like a fun big community of people doing the same thing."

"外人看上去好像两家是敌人,其实内部更像是一个有趣的大社区,大家在做同一件事。"

"I would have learned from Anthropic is this like real care and craft towards — it's like model behavior, model craft, model training. And I've been thinking a lot about, okay, like what makes Claude Claude than what makes ChatGPT ChatGPT."

"我从 Anthropic 学到的是对模型行为真正的'用心和工艺感'——model behavior、model craft、model training。我也在反复琢磨:Claude 之所以是 Claude、ChatGPT 之所以是 ChatGPT,差别到底在哪里?"

"This comes down to operational processes that kind of lead to the outputs — the outputted model."

"答案是:差异来自'最终模型输出'背后的整套运营流程。"

"It's like the reason why Claude has so much more personality and is more like a librarian. I don't know, like visualizing Claude being like a librarian, like a very nerdy or something."

"这也是为什么 Claude 拥有更多人格——它给人一种'图书管理员'的感觉,有点 nerdy。"

"It's because I feel like it's a reflection of the creators who made this model, and a lot of details around the character and personality, and whether the model should follow up on this question or not, like was the correct ethical behavior for the model in this scenario."

"我觉得 Claude 其实就是创作者的影子——还有大量围绕'角色与人格'的细节:模型在这个场景里要不要追问?道德上的正确反应是什么?"

"A lot of craft, and like 'read it like this and this is — read it like this.' And this is where I learned that part of art, I guess, at Anthropic."

"大量的工艺感,大量的'再读一遍、再改改'。我就是在 Anthropic 学到了这门艺术。"

"Anthropic was much smaller — 70 people when I joined, 700 when I left. Hardcore prioritization."

"Anthropic is like much smaller. Like when I joined it was like — what, like 70 people? When I left it was 700 people. So obviously the culture changed so much."

"Anthropic 当年小得多。我入职时是 70 人左右,离开时已经 700 人。整个文化变化非常大。"

"I really enjoyed doing like early days startup vibes, and like people knew each other as a family. But like the culture shifted."

"我很享受那种早期 startup 的氛围,大家像一家人一样彼此认识。后来文化变了。"

"I would say I learned from Anthropic that — much better at like focusing and prioritization — like very, very hardcore prioritization, I guess. And then need to do it, like."

"我从 Anthropic 学到的最重要的是——极致的聚焦与优先级管理,真的非常硬核的 prioritization。"

"But I think OpenAI is much more innovative and much more risk takers in terms of product or research."

"但 OpenAI 在产品和研究上更倾向于创新、更愿意冒险。"

"Actually, you know, your full-time job can be just like teaching the model how to be a creative writer. And there's some luxury in this — like research freedom that comes with scale, maybe, I don't know."

"你的全职工作可以就是'教模型怎么变成更好的创意写作者'。这种研究自由有种'奢侈'的味道——可能是规模带来的吧,我说不准。"

"But it gives you — I feel like I have much more creative product freedom to do almost anything I guess within OpenAI. Like, you lean into the direction that you want. It's more like — yeah, probably bottoms up, I guess."

"它给你——我感觉在 OpenAI 我有更大的产品创意自由,几乎想做什么都行。你想往哪个方向钻就往哪钻。整体更 bottoms-up。"

Lenny: "Yeah, that's how I was thinking about it. It feels like OpenAI is more bottoms-up — distributed, people bubble up ideas, try stuff, there's more — and that leads to more products launching, I imagine. More things just kind of being tried. Versus more of a 'let's just make sure everything we do is awesome and great and craft and thinking deeply about every right — every investment.' That's really interesting. I've never heard it described this way."

Lenny:"对,这也是我的感受。OpenAI 更 bottoms-up——分布式、人们自下而上冒想法、做尝试,可能因此上线更多产品、更多东西被试出来。而 Anthropic 更像'每件事都得做得极致、有工艺感、每一笔投入都深思熟虑'。挺有意思的对比,我之前没听人这么描述过。"

"Karina, we've covered so much ground. This is going to help a lot of people with so many ways of thinking about where the future is going. Before we get to our very exciting lightning round, I'm curious if there's anything else that you think might be helpful to share or get into."

"Karina,我们聊了好多。这一集会帮到很多人去理解未来的方向。在进入快问快答之前——你还有没有什么想分享或者还想展开聊的?"

Karina: "One of my regrets, I guess, when I was early days at Anthropic — was that I think there was some luxury of the time pre-ChatGPT to actually come in with a bunch of ideas and prototype like almost every day."

Karina:"我有个遗憾——早期在 Anthropic、ChatGPT 出来之前那段时间,我们其实拥有一种'时间上的奢侈',可以每天进来就 prototype 一堆想法。"

"And I think we did a lot of cool ideas, like Claude in Slack was actually one of the first tool-use-y products. It's like Claude could operate in your workplace. Now it's like — kind of like — you @Claude summarize the thread."

"我们做出了一些挺酷的东西,比如 Claude in Slack——其实是最早的'工具使用类'产品之一。意思是 Claude 能在你的工作场景里活动。今天看就是——@Claude 帮我总结这段对话。"

"So maybe you have an entire conversation with someone and then you want a summary, like what happened. You can say at-Claude summarize this. Also, it was really fun to iterate on the model itself — like when you just talk to the model in Slack forever."

"假设你和某人聊了很长一段,你想要个'到底发生了什么'的摘要,就 @Claude 总结。同时在 Slack 里直接对模型不停迭代——非常好玩。"

"It created some social element. It's kind of cool. It's kind of like — 'me, join me in this Discord' — like people learned so much about prompting and how to work with Claude."

"它带来了某种社交属性,有点像'来,加入我这个 Discord'——大家学到了很多 prompt 技巧,也学会了怎么和 Claude 协作。"

"I feel one of the features that was like early Tasks prototype was — every Monday Claude would just summarize the entire channel, or every Friday Claude would just summarize a bunch of channels and give the news about the organization or something."

"我觉得其中一个功能其实就是'早期版本的 Tasks'——每周一 Claude 自动总结一整个频道,或者每周五 Claude 自动总结多个频道,给你一份组织级别的新闻。"

"It's really cool, like form factor. I think thinking about form factor is a really important question in AI especially. We haven't even figured out how to create like an awesome product experience with like o-series models."

"这是一个很酷的 form factor。我觉得'form factor 该是什么'是 AI 里非常重要的问题。我们甚至还没真正搞清楚:o 系列模型应该长成什么产品形态才好。"

"It's like the paradigm between, like, synchronous, real-time, give-an-answer paradigm — into more asynchronous paradigm of agents working in the background."

"这是范式的迁移——从'同步、即时、要答案'的范式,转向'异步、agent 在后台干活'的范式。"

"But then now the question is, the agents should build trust with you, right? And trust is built over time, like with humans."

"但接下来问题是:agent 必须和你建立信任。信任是时间累积出来的——和人之间也是这样。"

"You start this collaboration — which is why this collaboration model between you and a model is so important — because you both trust, and the model learns from your preferences, so that it can become more personalized."

"所以'你和模型协作'的这个模式之所以重要——是因为你们之间形成了信任,模型学会了你的偏好,从而越来越个性化。"

"And it will start predicting the next action that you want to take on the computer or something. And it's like kind of like more predictive, much more — we went from like personal computer to like personal model, basically here. That's right."

"它会开始预测你下一步想在电脑上做什么,更具有'预测性'。本质上,我们正在从'个人电脑'走向'个人模型'。"

Lenny: "Why is it not a thing — that seems like such an obvious feature that every LLM should have, is a Slackbot version of them? Is that a thing I can install or is that not a thing right now?"

Lenny:"为什么没普及?好像每个 LLM 都该有一个 Slackbot 版本——这是显而易见的功能吧?我现在能安装这种东西吗?"

Karina: "I know that Claude in Slack was sunsetted in like 2023 or something. But that's because — I think after ChatGPT, the focus was mostly on consumer use cases or enterprise use cases."

Karina:"Claude in Slack 大概在 2023 年被下线了。原因是——ChatGPT 出来之后,大家的重点要么转到消费者场景,要么转到企业场景。"

"I think the form factor of Claude in Slack was kind of constrained a little bit. When you want new features — I want that. I know that ChatGPT had like a Slack bot too. So I don't know, maybe it'll come back some day."

"Claude in Slack 这个 form factor 还是受了不少限制。我也希望它回来——ChatGPT 之前也有 Slack bot。说不定哪天它会回归。"

Lenny: "Alright, I would pay for that. Any other memories from that time? Of early days, because that's a really special place to have been — as early days Anthropic. Any other memories or stories from that time that might be interesting to share?"

Lenny:"我愿意为这个付费。早期那段时间还有什么记忆?能在 Anthropic 早期就在那儿,是非常特别的。还有哪些故事可以分享?"

Karina: "I think the very first launch when we felt like 'click and use,' was like 100K context launch. It's when the models could input the entire book and give you a summary of the book, or the entire financial — like have multi-files financial reports — and then give you an answer to a very specific question."

Karina:"第一次让我们觉得'click and use'(一上来就有用)的发布,是 100K context。模型能把一整本书塞进去给你总结,或者读完多份财务报告再回答你一个非常具体的问题。"

"I think there was something in there that kind of, like, oh my God, this is a really cool new capability — not like model capability, but more like the capabilities that came from the product form factor itself, rather than the model capability as much."

"那一刻我们意识到——这是一个很酷的新能力,但它不是'模型能力'的延伸,更多是来自'产品 form factor 本身'催生出来的能力。"

"I think other prototypes that we were thinking about — like Claude workspaces. It's kind of the same idea — Claude and I would have this shared workspace, and that shared workspace like a documents, and you can like edit and I feel like sometimes the ideas like, partial ideas, lag and they lock for like two years, just like in this case."

"我们当时还想做 Claude workspaces——和这个思路差不多:Claude 和我共享一个工作空间,里面有共享文档,可以一起编辑。有时候一些'半成品想法'会被搁置两年之久——这个就是。"

Lenny: "It's interesting — there are these milestones that open up our view of what is happening and where things are going. ChatGPT was the first 'wow, this is much better than I would have thought.' You talked about 100K context windows where you could upload a book and ask questions and have it summarized. I actually use that all the time when I have interview guests and they wrote a book — I sometimes don't have time to read the whole book, so I use it to understand the most interesting parts."

Lenny:"挺有意思——这些 milestone 改变了我们对'现在到哪了、要往哪去'的认知。ChatGPT 第一次让我们觉得'天啊,这比我想象的强太多'。你说的 100K context——能上传一本书、问问题、要摘要——我现在就经常用。访谈嘉宾如果出过书,我没时间读完整本,就用它先抓最有意思的部分。"

"And then I don't know — maybe voice was another one, where you could talk to ChatGPT. Is there any other moments there that you're like 'wow, this is much better than I thought it was going to be'?"

"再就是语音——你能直接和 ChatGPT 对话。还有什么瞬间让你觉得'比我以为的强多了'?"

Karina: "Yeah, I think like — the computer-use agents — the model operating the desktop. You can essentially think of like, you know, a new kind of experience where the model can learn the way you browse, and from that preference it can just browse just like you. It's kind of like a simulation, simulated like persona."

Karina:"还有 computer-use agent——模型操作桌面。你可以想象一种新体验:模型学习你的浏览方式,基于这个偏好,它能像你一样上网。本质上是一种'模拟人格'。"

"It's actually very similar to the idea of, like, okay, maybe Sam Altman doesn't have a lot of time. Maybe I want to talk to like his simulator — his simulation — and ask. Or, like, I really appreciate some of the tech, like Naval — but he doesn't have a lot of time. So I really want to ask him these questions, like how he'd respond."

"和这个思路接近的还有——Sam Altman 没太多时间,我能不能跟他的'模拟体'对话?或者我很喜欢一些技术派比如 Naval 的思路,他也没空——我特别想问他几个问题,看他怎么回。"

"Let's simulate the environments like those — would be really cool."

"如果能模拟出这种环境,会非常酷。"

Lenny: "It's a great place to plug Lenny bot. I have one of those — it's trained on all of my podcasts and newsletters, and it sits on many models. I don't know which one exactly they use, but it's exactly that. And it's not even me — it's all the guests that have been on the podcast and newsletter I wrote. You can just ask it 'how do I grow my product, how do I do strategy' — and it's actually shockingly good."

Lenny:"这正好可以打个广告——我有一个 Lenny bot。它用我所有播客和 newsletter 训练过,跑在多个模型上,具体哪个我不清楚。它不只是我——还包括所有上过我节目和 newsletter 的嘉宾。你可以直接问它'怎么增长我的产品、怎么做战略'——效果好得让人吃惊。"

Karina: "Do you feel like it reflects who you are?"

Karina:"你觉得它能反映真实的你吗?"

Lenny: "The best part of it is — you can talk to it. There's an 11 Labs voice version that's trained on my voice now from this podcast, and it's actually very good. People have told me they sit there for hours talking to it."

Lenny:"最棒的一点是你可以直接和它对话。基于这档播客,11 Labs 用我的声音训了一个语音版本,效果非常好。有人告诉我他们一坐就和它聊好几小时。"

Karina: "Wow."

Karina:"哇。"

Lenny: "And somebody told it 'interview me like I am on Lenny's podcast, ask me questions about my career' — and he did a half-hour podcast episode with Lenny."

Lenny:"还有人跟它说'当我是 Lenny 播客的嘉宾来采访我,问我职业相关的问题'——结果他自己和 Lenny bot 录了一段半小时的播客。"

Karina: "Oh God, that's so fun. It's incredible. Future is wild."

Karina:"我天,太好玩了。这难以置信。未来真的太疯了。"

"Yeah, I think like — content transformation. I would imagine sometime, like, when you generate a sci-fi story in Canvas, you can transform this into like audio. Like where you have very natural content transformation, one media to another media."

"我能想象,未来你在 Canvas 里写了一篇科幻短篇,能直接转成音频。'一种媒介自然地转换成另一种媒介'。"

"I think one of my earliest inspirations is like — one of the last episodes of Westworld, where Dolores comes to her workspace and starts writing a story, and that story — like a 3D virtual reality — starts creating on the fly."

"我最早的灵感之一来自《Westworld》最后一季——Dolores 来到她的工作区,开始写一个故事,这个故事就被实时渲染成 3D 虚拟现实。"

"So I kind of want to create that. Kind of cool."

"我有点想做这种东西。挺酷的。"

Lenny: "Wow. Speaking of medium — I was wondering if I should go in this direction or not, but real quick. Kevin Weil, I don't know exactly how to pronounce his last name — the CPO of OpenAI — is it 'wile' or 'wheel'? I think it's 'wheel.' Okay, let's just say that."

Lenny:"哇。说到媒介——我犹豫了一下要不要跑题。Kevin Weil(我不太确定怎么念,Wile 还是 Wheel?好像是 Wheel)——他是 OpenAI 的 CPO。"

"He did a panel at the Lenny and Friends summit last year, and he made this really fascinating point — that chat is a really interesting interface for these tools. Because they're just getting smarter and smarter and smarter, and chat continues to work as a paradigm to just interact with them, similar to a human."

"他去年在 Lenny and Friends 峰会上做嘉宾时讲了一个很妙的观点——chat 是这类工具非常有意思的界面。因为模型一直在变聪明,而 chat 这种'像人一样对话'的范式始终都能用。"

"You could talk to Albert Einstein, you could talk to someone not very smart, and it's all conversation still. So it's a really flexible way to interact with increasingly good intelligence."

"你可以跟爱因斯坦聊,也可以跟一个不太聪明的人聊,本质上都是对话。这是一种非常灵活的、适配各种智能水平的交互方式。"

"At some point it'll not be so great, and you're talking about all these ways that you're adding additional ways to interact. But it's interesting that chat proved to be a really powerful layer on top of all the stuff."

"也许某天它会显得不够用,你们也在探索其他交互方式。但 chat 确实证明了——它是一个非常强大的、套在一切之上的交互层。"

Karina: "Yeah, that's really cool. I feel like chat also has like a social element which is very humane. It's like, yeah, you sometimes want to get into a group chat — and having conversations there. It's kind of like a group chat in itself, like messaging."

Karina:"对,挺酷的。chat 还有一个很'人性'的社交属性——你会想跟一群人开 group chat 在里面聊天。它本身就像 group chat,像 messaging。"

"This idea of how do you build features like this — I see Tasks as this general kind of feature that will scale very nicely as the models develop new capabilities."

"怎么去做这一类功能?我觉得 Tasks 就是一种特别通用的功能,随着模型能力提升,它会很好地 scale。"

"As models will be able to do better searches, create new, come up with more creative writing, render React apps and HTML preview apps — you can have every day a new puzzle for you, every day continue the story from the previous days. It scales very nicely."

"模型能搜得更好、能写得更有创意、能渲染 React 应用、HTML 预览应用——你就能每天有一个新谜题,每天接着昨天的故事写。这件事 scale 起来非常自然。"

Lenny: "You mentioned something as we were getting into this extra section that we ended up going down — this idea of your agents using a computer. I know this is actually something you're going to launch today, the day we're recording it, which will be out by the time this comes out, called Operator. Can you talk about this very cool feature that people will have access to?"

Lenny:"你刚才提到一个东西,我们顺势聊下去——agent 操作电脑这件事。我知道你们今天就要发布——也就是录制当天——一个叫 Operator 的功能,等节目上线时它已经发布了。能给大家讲讲吗?"

Karina: "Yeah, so unfortunately I did not work on that, but I'm really really excited about this launch."

Karina:"很遗憾我没参与开发,但我对这次发布非常兴奋。"

"It's basically an agent that can complete the task in its own virtual computer, in its own virtual environment. You can do literally any task — 'order me a book on Amazon.' Then ideally the model will either follow up with you, like 'which book do you want?' Or know you so well that it starts recommending — 'here's five books I might recommend you to buy.' And then you hit like 'yeah, help me buy.'"

"它本质上是一个能在自己的虚拟环境/虚拟电脑里完成任务的 agent。你可以让它做任何事——'帮我在亚马逊买一本书'。理想状态下,要么它会追问你'你要哪本',要么它已经太了解你,直接推荐 5 本,然后你点'好,帮我买'。"

"And then the model goes off into its own virtual little browser and completes the task and buys the book on Amazon. And then if you give the model like initial credit card — obviously it comes with a lot of trust and safety — then it will just complete the thing for you. That's a virtual assistance."

"接着模型进入它自己的小型虚拟浏览器,把这件事干完——在亚马逊把书买了。如果你给了它初始的信用卡,这当然涉及信任和安全的考量——它就帮你把事情办完。这就是虚拟助理。"

Lenny: "It's interesting how this just sounds like — obviously this should happen, why is this not other? — which is also mind-blowing that we're just assuming this should exist. Like, just some AI doing things for you on a computer you just ask it to do. It's absurd."

Lenny:"有意思的是,这听起来很'当然该有'——为什么以前没有?也很疯狂——我们已经默认这种东西应该存在了。'电脑里一个 AI 替你跑腿',想想其实挺荒诞的。"

Karina: "It's actually really hard. I think we're still cracking this. Do you feel like — I don't know if you used like Topel, it's a pair-programming product."

Karina:"这其实非常难做。我们还在攻关。你用过 Topel 没?那是个 pair programming 产品。"

Lenny: "No, but —"

Lenny:"没用过——"

Karina: "I don't remember if you love pair programming. So if you — oh yeah, Shopify uses this, I remember it came up on a podcast episode."

Karina:"我忘了你喜不喜欢 pair programming。哦对,Shopify 用这个,我记得在一期播客里听过。"

"It's a very cool product where you can call anyone at any time and share screen, and the other person can have access to the screen and start literally operating your computer. It's very real-time. The latency is very high quality."

"是个很酷的产品——你可以随时呼叫任何人,共享屏幕,对方就能直接操作你的电脑。延迟很低、实时性很好。"

"And I kind of want the same — I want to pair-program with my model. And the model should even talk to me, draw very specific section in my code in VS Code, and tell me — like, teach me. You can have different modes."

"我希望有同样的东西——我想跟模型 pair program。模型应该能开口讲话,在 VS Code 里圈出我代码里某一段,边讲边教我。可以设置不同模式。"

"It's like — right here, this is the product right here for you. I don't know — some people should build it. It sounds like a startup just got birthed."

"'这就是给你做的产品'——有人去做啊。听起来一个 startup 刚刚诞生。"

Karina: "Yes, from someone listening to this."

Karina:"对,可能就是听这期播客的某个人。"

Lenny: "You mentioned that it's very hard to do this agent controlling a computer as you and helping out. What makes it so hard, for however much you can explain briefly?"

Lenny:"你说'让 agent 像你一样操作电脑'很难。能简单讲讲为什么难吗?"

Karina: "Much of it is because right now the models are operating on pixels instead of language or whatnot. Pixels are actually really really hard for the models — because visual perception. I think there's still a lot of multimodal research that's going on."

Karina:"很大一部分原因是——现在模型操作的是像素,而不是语言。像素对模型来说非常难,因为它涉及视觉感知。multimodal 研究还有大量工作要做。"

"But I think language scaled so much easier compared to multimodal — because of that."

"语言的 scaling 比 multimodal 顺利得多,就是这个原因。"

"Another thing that my team is working on, that is like — how do you derive human intent very correctly?"

"我团队还在攻关另一件事——怎么正确地推断人类意图?"

"It's like sometimes — does a model know enough information to ask a followup question, or to complete the task? You kind of don't want an agent to like go off for 10 minutes and then come back with an answer that you didn't even want. That actually creates much worse user experience."

"有时候——模型有没有足够的信息去追问,或者去完成任务?你绝不希望一个 agent 闷头跑 10 分钟,然后给你一个你根本不想要的答案。那样体验比不做还糟。"

"And this comes with teaching the model like 'people skills' — like, what do people like — kind of like creating the mental model of the user, and like care about the user in order to ask certain questions."

"这就回到'教模型人际能力'——它得理解人喜欢什么、能在脑海里建立一个用户的心智模型,愿意问出对的追问题。"

"Like, actually that part is hard to — for the models."

"这部分对模型来说,确实非常难。"

Lenny: "That relates to what we talked about earlier, where this kind of the soft skill, people skill pieces — yeah, not where these models are strong yet."

Lenny:"这就回到我们前面聊的——soft skills、人际能力这一块,正是模型目前还不强的地方。"

"Okay, I'm going to skip the lightning round. I want to ask just one question from the lightning round, something fun."

"OK,正式的'快问快答'我跳过,只问其中一个有趣的问题。"

"Okay, so when AI replaces your job, Karina, I'm curious — and it gives you a stipend, gives you a monthly stipend, here's your salary for the month — what would you want to do? What do you want to spend your time on? What will you be doing in this future world?"

"假设有一天 AI 替代了你的工作,Karina——并且按月发你津贴,你的'工资'继续给。你最想做什么?把时间花在哪?在那种未来你的日常是什么样?"

Karina: "I've been thinking about this a lot of times. I have — I feel like I have a lot of job options."

Karina:"我经常想这件事。我感觉自己有很多职业选项。"

"I would love to be a writer. I think that would be super cool. You should like write like short stories, sci-fi stories, novels."

"我想当作家。短篇、科幻、长篇都行,这件事我觉得超酷。"

"I really like art history. So you know those like conservationists in the museums who just try to preserve art paintings — but just like painting through a lot of things — I think that would be really cool to do."

"我也很喜欢艺术史——那种博物馆里负责修复绘画的人(conservationist),他们的工作就是把油画修复回原本的样子。我觉得当这种人也特别酷。"

Lenny: "Yeah, that sounds beautiful. I don't know — what I'm hearing is you need to nerf these models to not get very good at writing so that you can continue. Although at that point, you don't need it for like — you don't need people to buy. You're just doing it for fun. So it doesn't even matter if they're incredibly good at writing or art, art conservation."

Lenny:"挺美的。我听到的是——你需要给模型在写作上做点'降智',这样你还有的写。不过到那时候,你也不靠卖书过日子了,纯粹是好玩。所以模型在写作或艺术修复上有多好,其实也不重要。"

"Oh man, what an episode — what a conversation. What a wild time we're living in."

"我天,这一期真精彩。这是一个多疯狂的时代。"

"Karina, thank you so much for being here. Two final questions — where can folks find you online if they want to reach out and follow up on anything, and how can listeners be useful to you?"

"Karina,非常感谢你来。最后两个问题:1)大家想联系你或继续追问,该去哪找你?2)听众怎么能帮到你?"

Karina: "You can find me — I'm on Twitter. You can also shoot me an email on my website."

Karina:"在 Twitter 上能找到我。也可以通过我个人网站发邮件给我。"

"And my team is hiring. So I'm looking for research engineers, research scientists, as well as machine learning engineers — like people who come from like product engineers who want to learn like model training."

"我们团队正在招人——research engineer、research scientist,以及 machine learning engineer——尤其欢迎从产品工程背景出身、想转去学模型训练的人。"

"Actually hiring for my team — my team is called Frontier Product Research, and we train models, we develop new methods, but for product-oriented outcomes."

"我团队叫 Frontier Product Research,我们训模型、做新方法——但都是为了产品导向的结果。"

Lenny: "What a place to work — holy moly. What's the best way for people to apply for these very lucrative roles?"

Lenny:"这工作太香了。大家要申请这些岗位,最合适的方式是什么?"

Karina: "I think you can shoot me a DM on Twitter."

Karina:"可以在 Twitter 上私信我。"

Lenny: "Okay."

Lenny:"OK。"

Karina: "Or — I'm yet to create a job description. Okay, this is the job description. Or you can apply into the post-training team."

Karina:"或者——我还没正式写出 JD。嗯,这一段就当 JD 了。你也可以直接申请 post-training 团队。"

Lenny: "Okay, this — you're going to get a flood of DMs. I hope you're prepared. Karina, thank you so much for being here. This was incredible."

Lenny:"OK,你准备好接收私信洪流吧。Karina,非常感谢你来,这一期真的非常棒。"

Karina: "Thank you so much, Lenny. Bye everyone. Fun, thank you so much for listening."

Karina:"谢谢 Lenny。各位再见,挺有意思的——谢谢大家收听。"