Lenny's Podcast with Lenny Rachitsky · 2026-07-26 · 双语整理

You Have to Sweat the Tokens

Anthropic 第一位技术 PM:evals 就是新的 PRD

“You have to sweat the tokens as much as you sweat the pixels. You have to be using the models to come up with good and great and better ideas, and there's no substitute for that.”

Dianne Penn 2023 年加入 Anthropic 时是公司第一位技术产品经理,产品团队只有五个工程师,整条 API 业务线只有一个人。三年后她管着 AI Research 和 Labs 两个团队的产品,Claude 2 到 Fable 每一代模型都经过她的手,Claude Code、MCP、skills、Claude Design 都是从这里孵出来的。这一集她讲的不是「AI 会改变一切」,而是一个产品组织在指数曲线内部具体怎么改工作方式——从写 PRD 改成写 evals,从走查像素改成读会话轨迹。

来源:YouTube · Lenny's Podcast · 2026-07-26 · 1:33:50 · 约 15,040 词 · 15 章 · 逐字双语对照
TL;DR · 速读

Anthropic 的产品工作,具体改了哪些地方

  1. evals 就是新的 PRD

    “So we do write some product documents and PRDs, but we actually have a saying on the team of evals are the new PRDs.”

    “And for my team as research product managers, the way to drive user value is to figure out the right user feedback, the evals that then can be a personification of that user need.”

    交付用户价值的载体,从「写一份文档」换成了「定义一组可测的失败样本」。

  2. 抠 token 得像抠像素

    “You have to sweat the tokens as much as you sweat the pixels.”

    “In the past, we might do a user interview. I think if you go deep enough, you might have the user walk you through their user flow, the pixels.”

    过去 PM 陪用户走查界面,现在要一条条读会话轨迹,判断这次失败是幻觉还是过度自信。

  3. 先有前沿产品,才有前沿模型

    “One thing we say a lot on the team is you need frontier products in order to have frontier models and for people to feel the magic of frontier models.”

    “I think what was magical about Opus 4.5 is we also now not just had a model, but a vehicle, which is a great product experience like Claude Code.”

    Opus 4.5 和 Claude Code 互为对方的放大器,少了任何一个,那个时刻都不会发生。

  4. 模型早会了,只是还没人用上

    “I think there's product overhang and user overhang, to maybe put it in our PM language, even on today's models.”

    “That's so interesting that you may have developed this AI brain that can do something you're not even aware of.”

    能力先于产品出现,所以「发现模型已经会做什么」本身就是一份工作。

  5. 实验未必是一个人的运动

    “Experimentation is not always necessarily a individual sport.”

    “And then within maybe 10 or so requests, there was something magical or potentially in a use case that emerges.”

    Anthropic 早期几乎全公司挤在一个 Slack 频道里试 Claude,用例是互相接力接出来的。

  6. 主题咬死,原型随时推翻

    “You can be very strongly held opinion about the theme or the area, and then more weakly held about the exact prototype.”

    “The thing that we really try to emphasize within the teams is, especially right now, there are so many things that could be built. What does it mean then to have a discontinuous bet?”

    Labs 的打法:方向不放手,具体形态说改就改,跑不通就等一两代模型再试。

  7. 「幻觉了」不是需求

    “If you bring that to a researcher and you say, "Please fix Claude from being hallucinated," it's not very actionable.”

    “So for example, we might get feedback on claude.ai. "Claude hallucinated." It's very vague.”

    PM 的活是把它拆成「该调工具没调」还是「文档找对了但抓错事实」,拆到研究员当天能动手。

  8. 用 Claude 8 检验今天

    “One thing I ask the team frequently or how I think about when we're building a product is let's say Claude 8 comes around. What changes in what users do?”

    “It is a question we challenge ourselves with, but the technology is moving so quickly, and so how do you make sure what you're building is actually forward compatible?”

    把「更有野心」落成一道能回答的题:模型再强几代,今天这个设计还站得住吗。

  9. 当管理者也得亲自发版

    “So I do feel pretty strongly that if you're a manager, you have to be hands-on; you have to spend a portion of your time actually shipping.”

    “And so even for folks that I hire who have more tenured PM experience, the onboarding plans are exactly the same as somebody who is more early career.”

    她自己每代模型都留一两条工作流亲自扛,为的是保住对模型进展的手感。

  10. 思考搭档不是只会点头的

    “If you have an AI that can be a thinking partner, a thinking partner doesn't just agree with you. It should add to you.”

    “And you should come away at the end of the day having better ideas because you worked with Claude. That should be the hero goal, not just making your ideas 10% better.”

    alignment 研究让 Claude 敢反驳,而敢反驳恰恰是它更有用的原因,不是代价。

  11. 谁签字比谁执笔重要

    “Maybe less around who's writing, but who's verifying who's signing off? That becomes more what matters than who's writing it.”

    “I want to get to a place where the monthly business review, the writing of that, is potentially asymmetrically less valuable than the thinking.”

    月度业务回顾可以整篇交给 Claude,她要做的是审稿和背书,而不是当写手。

  12. 判断力是攒出来的

    “I think judgment is an area where, it's accumulation of so much nuance and so much experience.”

    “There are so many things AIs can build. Which one are the things that an org like Labs should build? A lot of that requires human judgment, persistence, so proactivity.”

    能造的东西无限多,该造哪个仍然要人来判断——这是她认为人脑最后失守的地方。

  13. 休假回来不该欠一屁股债

    “It's not just that you can take PTO and you come back to 3X the amount of things to do.”

    “who the night before a launch, even if they're not the core DRI on that model, will stay up and help the DRI to review the blog post and make edits and come up with better demos”

    一年四个模型变成一个季度四个,靠的不是个人硬扛,是团队能替你做对决定。

  14. 上面永远还有一层

    “And my grandfather always says, no matter how far you go, there's always another level.”

    “So I was actually raised by my grandparents for the first 10 years of my life. And my parents were immigrant college and master's students in the US.”

    遇到全新、没有先例的事情时,她拿这句话当锚。今年上半年,这样的事特别多。

Chapter 01

Nobody Said Anthropic and Coding in the Same Sentence

早期的 Anthropic · 五个工程师和一座金门大桥
00:00 — 07:46 · 冷开场 · 五人产品团队 · 24 小时上线的 Golden Gate Claude
Dianne Penn00:00:00

In 2023 when I started, nobody said Anthropic and Claude and coding in the same sentence.

2023 年我刚来的时候,没人会把 Anthropic、Claude 和写代码放在一句话里说。

Lenny00:00:06

I want to go back to the beginning of Anthropic.

我想回到 Anthropic 最开始的时候。

I remember feeling, "Man, these guys have no chance.

我记得当时的感觉是:“天,这帮人没戏。

OpenAI is so far ahead."

OpenAI 领先太多了。”

Dianne Penn00:00:14

By the time I saw people were starting to use these models, not just for code autocomplete, but actually writing long-form code.

等我看到大家开始用这些模型,不只是做代码自动补全,而是真的在写成篇的代码——

Is that an opportunity for us to train Opus 3 to be better at?

这是不是个机会,让我们把 Opus 3 训练得更擅长这个?

That was the inflection.

那就是那个转折点。

Lenny00:00:26

I always think about Opus 4.5 a year later during winter break when everyone was home able to code.

我总想到一年后的 Opus 4.5,正赶上寒假,所有人都在家,都能写代码。

Dianne Penn00:00:31

What was magical about Opus 4.5 is we also now not just had a model, but a vehicle, a great product experience like Claude Code.

Opus 4.5 神奇的地方在于,我们那时不只有模型,还有一个载体,一个很好的产品体验,就是 Claude Code。

Opus 4.5 wouldn't have had that moment without a product like Claude Code and Claude Code wouldn't have had that type of adoption accelerated without Opus 4.5.

没有 Claude Code 这样的产品,Opus 4.5 不会有那个时刻;没有 Opus 4.5,Claude Code 的普及也不会被推得这么快。

Lenny00:00:50

I want to talk about how the product role is changing.

我想聊聊产品这个岗位正在怎么变。

Dianne Penn00:00:53

For my team, the way to drive user value is to figure out the right user feedback.

在我的团队,创造用户价值的方式,是找到对的用户反馈。

The evals, we actually have a saying on the team of evals are the new PRDs.

说到 evals,团队里其实有句话:evals 就是新的 PRD。

Lenny00:01:02

Something Garry Tan's been talking about.

Garry Tan 一直在讲的一件事。

If you are willing to spend $100,000 a year right now in tokens, you are living the way somebody in 2028 is going to live.

如果你现在愿意一年在 token 上花 $100,000,你过的就是 2028 年的人才会过的生活。

Dianne Penn00:01:10

You have to sweat the tokens as much as you sweat the pixels.

抠 token 得像抠像素一样抠。

You have to be using the models to come up with good and great and better ideas, and there's no substitute for that.

你必须靠这些模型想出好点子、很棒的点子、更好的点子,这件事没有替代品。

Lenny00:01:20

People need to be more ambitious with AI tools these days because they're just capable of so much.

现在大家用 AI 工具得更有野心一点,因为它们能干的事实在太多了。

Dianne Penn00:01:26

One thing I ask the team is let's say Claude 8 comes around.

我会问团队一个问题:假设 Claude 8 出来了。

What changes in what users do?

用户做的事会有什么变化?

What does that mean for how you're building today?

这对你今天怎么做产品意味着什么?

Lenny00:01:36

Today, my guest is Dianne Penn, head of product for the AI research and Labs teams at Anthropic.

今天的嘉宾是 Dianne Penn,Anthropic AI 研究团队和 Labs 团队的产品负责人。

She joined Anthropic as the first technical product manager over three years ago, which is a lifetime in AI time.

三年多前,她作为第一位技术产品经理加入 Anthropic——按 AI 的时间尺度,那是一辈子那么久。

When the product team was just five engineers, she's helped ship every model at Anthropic from Claude 2 through Fable.

当时产品团队只有五个工程师;从 Claude 2 到 Fable,Anthropic 的每一代模型她都参与上线。

She's also helped incubate and launch Claude Code, MCP, skills, Claude Design, and also core capabilities like computer use, tool use, and reasoning.

她还参与孵化和推出了 Claude Code、MCP、skills、Claude Design,以及 computer use、tool use、推理这些核心能力。

It is always such a treat and so mind-expanding to get to talk to someone who's at the very center of AI and product management, it's hard to imagine someone who has seen more of where things are going than the head of product for Anthropic's research and Labs teams.

能跟一个身处 AI 与产品管理正中心的人聊天,永远是种享受,也总能把脑子撑开;很难想象还有谁比 Anthropic 研究团队和 Labs 团队的产品负责人看到过更多事情正往哪儿走。

Before we get into it, don't forget to check out lennysproductpass.com for a year free of the hottest and most beautifully crafted AI products in the world available exclusively to Lenny's newsletter subscribers.

开始之前,别忘了去看看 lennysproductpass.com,免费用一年全世界最热门、做得最漂亮的 AI 产品,只对 Lenny's newsletter 的订阅者开放。

With that, I bring you Dianne Penn.

那么,有请 Dianne Penn。

Dianne, thank you so much for being here and welcome to the podcast.

Dianne,非常感谢你来,欢迎来到播客。

Dianne Penn00:02:41

Thank you, Lenny.

谢谢你,Lenny。

It's so nice to see you again.

很高兴又见到你。

Lenny00:02:43

I want to go back to the beginning of Anthropic, the early days.

我想回到 Anthropic 最开始的时候,那些早期的日子。

I remember when Anthropic first launched, this was, I don't know, the first model when it launched years ago, three years ago, something like that.

我记得 Anthropic 刚出来的时候,第一个模型发布,那大概是几年前,三年前吧。

Dianne Penn00:02:56

It works.

对得上。

Lenny00:02:56

Three years.

三年。

I remember just feeling that, man, these guys have no chance.

我记得当时就觉得,天,这帮人没戏。

OpenAI is so far ahead.

OpenAI 领先太多了。

They're just like, what are they thinking?

他们到底在想什么?

How is this possible?

这怎么可能?

OpenAI is one.

OpenAI 已经赢了。

It's too late.

太晚了。

Things are very different now.

现在情况完全不一样了。

The latest number I saw was Anthropic was making, I don't know, $50 billion in ARR.

我看到的最新数字是,Anthropic 的 ARR 做到了 $50 billion。

That's what companies used to go public at.

过去公司上市也就是这个体量。

Very successful companies went public at 50 billion in valuation.

非常成功的公司,上市时估值就是 50 billion。

Anthropic reportedly is making that every single year.

而据说 Anthropic 每年都能做到这个数。

You joined as one of the earliest PMs.

你是最早的那批 PM 之一。

There were something like five engineers when you joined.

你加入的时候大概只有五个工程师。

The model hadn't even launched when you joined.

你加入的时候模型都还没发布。

What was it like in those early days of Anthropic?

Anthropic 早期那段日子是什么样的?

What's something that might surprise people about what it was like at the beginning?

最开始是什么样子,有什么可能会让大家意外的?

Dianne Penn00:03:44

I think a big part of what's made Anthropic today actually has been very much the core of even the early days.

让 Anthropic 成为今天这样的东西,很大一部分在早期就已经是核心了。

So I joined in 2023.

我是 2023 年加入的。

Like you said, we had five product engineers.

像你说的,我们当时有五个产品工程师。

There was one engineer for the entirety of our API business, if you believe.

整个 API 业务只有一个工程师,信不信由你。

And I think a big portion of it was the culture was really strong.

很大一部分原因是文化非常强。

And I think this is something I emphasize for folks who are interested in the company, really do walk the walk of the mission and the culture and the values.

对想了解这家公司的人我一直强调这一点:使命、文化、价值观,我们是真的说到做到。

And the energy was very much like a startup.

那股劲头就是一家创业公司。

I think you're right, we were very much trying to find our identity in the early years.

你说得对,早几年我们一直在找自己的身份定位。

I think there's one piece around the technology, but how does that technology bring value to users, bring value to society, and what could it possibly be?

一块是技术本身,但这技术怎么给用户创造价值、给社会创造价值,它到底可能是什么?

And I think the early years were us exploring that in different ways.

早几年就是我们在用各种方式探索这件事。

We did start with Claude.ai, another chatbot, chat assistant, and evolving into things like tool use.

我们确实是从 Claude.ai 开始的,又一个聊天机器人、聊天助手,然后慢慢长出 tool use 这类东西。

I think one of the moments where really we started to get into our groove was shipping things like Golden Gate Claude.

真正开始进入状态的一个时刻,是上线 Golden Gate Claude 这样的东西。

I don't know if you remember that.

不知道你还记不记得。

Lenny00:05:02

No.

不记得。

Dianne Penn00:05:03

So this was actually up for about 24 hours or so.

这东西其实只上线了 24 小时左右。

We had just published one of our early interpretability research in early 2024.

2024 年初我们刚发了一篇早期的可解释性研究。

And one of the examples was essentially you could have what's called features of the model within the layers, which express certain types of thematics.

其中一个例子是,模型的各层里存在所谓的特征,它们表达某些类型的主题。

So one of the themes that the researchers was able to identify was, let's say, bullet point writing.

研究员识别出来的一个主题,比如说,是要点罗列的写法。

Another one was people and places.

另一个是人物和地点。

And one that really came up frequently that resonated was the Golden Gate Bridge.

还有一个出现得特别频繁、也特别有共鸣,是金门大桥。

And so when you actually essentially dialed up that feature, Claude would obsess about the Golden Gate Bridge.

所以你把那个特征调高,Claude 就会满脑子都是金门大桥。

So meaning in every one of its responses, it would come back and talk about the Golden Gate Bridge.

也就是说,它每一次回答都会绕回金门大桥。

So if you said like, "Give me a recipe for making spaghetti," it would say, " Here is a recipe and the orange color is just like international red that the Golden Gate Bridge looked like."

比如你说“给我一个做意面的菜谱”,它会说:“这是菜谱,那个橙色就像金门大桥的那种国际红。”

And so it was really quirky.

它真的很古怪。

And we very much wanted to, in that situation, just bring it to the masses and bring it to people who were starting to use Claude.

当时我们特别想把它推给大众,推给那些刚开始用 Claude 的人。

And so the entire experience actually we spun up on our Claude.ai website within 24 hours.

所以整个体验,我们是在 24 小时内在 Claude.ai 网站上搭起来的。

And that took engineering, product, design, our research teams all working together.

这需要工程、产品、设计和研究团队一起上。

And we were really, really proud of it.

我们真的非常为它自豪。

I think it maybe reached only 2,000 people, to be honest.

说实话,它可能只触达了 2,000 人。

But it made us feel like, oh, we can actually bring new user experiences, showcase our research in a way that's different and authentic to us.

但它让我们觉得,我们真的能做出新的用户体验,用一种不一样的、属于我们自己的方式展示我们的研究。

And in a very startup be like pace.

而且是非常创业公司的节奏。

That to me was one of those maybe hidden inflection points of we were starting to find our identity, that we could build products, build experiences that were different from what our competitors had seen, what was already out there.

对我来说,那算是一个隐性的转折点:我们开始找到自己的身份,我们能做出产品、做出体验,跟竞争对手做过的、跟市面上已有的都不一样。

And I think that obviously Labs, Claude Code, et cetera, we then started to identify ourselves as would we actually think the world, how to think about AI, how to bring that closer to the public.

后来显然就是 Labs、Claude Code 这些,我们开始把自己定位成去想这个世界、想怎么看待 AI、怎么把它拉得离公众更近。

But it was a very bottoms-up culture.

但那是一种非常自下而上的文化。

And so that entire experience was very bottoms up.

所以整件事也非常自下而上。

I see engineers, I see designers donating time to work on.

我看到工程师、看到设计师主动贡献时间来做。

And so I like to always use that as an example of what the early days were like.

所以我总喜欢拿它当例子,说明早期是什么样子。

But the culture and the values have very much, I think, stayed the same since those early days.

但文化和价值观,从那时候到现在基本没变。

Chapter 02

Two Inflections: Opus 3, and Opus 4.5 × Claude Code

两个转折点 · 要有前沿模型,先得有前沿产品
07:46 — 13:50 · WorkOS 口播 · 圣诞节的 Opus 3 · 模型与产品互为放大器
Lenny00:07:46

This episode is brought to you by our season's presenting sponsor, WorkOS.

本期节目由本季特约赞助商 WorkOS 支持播出。

What do OpenAI, Anthropic, Cursor, Vercel, Replit, Sierra, Clay, and hundreds of other winning companies all have in common?

OpenAI、Anthropic、Cursor、Vercel、Replit、Sierra、Clay,还有另外数百家赢家公司,有什么共同点?

They are all powered by WorkOS.

背后都是 WorkOS 在驱动。

If you're building a product for the enterprise, you felt the pain of integrating single sign-on, SCIM, RBAC, audit logs, and other features required by large companies.

如果你在做面向企业的产品,你一定体会过接入单点登录、SCIM、RBAC、审计日志这些大公司必备功能有多痛。

WorkOS turns those deal blockers into drop-in APIs with a modern developer platform built specifically for B2B SaaS.

WorkOS 把这些卡单的障碍变成了拿来就能用的 API,背后是一个专为 B2B SaaS 打造的现代开发者平台。

Literally every startup that I'm an investor in that starts to expand upmarket ends up working with WorkOS.

我投的创业公司,只要开始向上打大客户,最后毫无例外都用上了 WorkOS。

And that's because they are the best.

因为他们就是做得最好。

Whether you are a seed stage startup trying to land your first enterprise customer or a unicorn expanding globally, WorkOS is the fastest path to becoming enterprise ready and unblocking growth.

不管你是想拿下第一个企业客户的种子轮公司,还是正在全球扩张的独角兽,WorkOS 都是最快具备企业级能力、扫清增长障碍的那条路。

It's essentially Stripe for enterprise features.

说白了,它就是企业级功能领域的 Stripe。

Visit workos.com to get started or just hit up their Slack where they have actual engineers waiting to answer your questions.

去 workos.com 就能上手,或者直接去他们的 Slack,那边有真的工程师等着回答你的问题。

WorkOS allows you to build faster with delightful APIs, comprehensive docs, and a smooth developer experience.

好用的 API、齐全的文档、顺滑的开发者体验,WorkOS 让你做得更快。

Go to workos.com to make your app enterprise ready today.

现在就去 workos.com,让你的应用达到企业级。

What are some of the other big inflection moments as you think about just Anthropic going from just this lab that's trying to compete with this juggernaut of OpenAI at that point to what it is today?

还有哪些大的转折时刻?我是说,你回头看 Anthropic 从当年那个想跟 OpenAI 这个庞然大物掰手腕的实验室,一路走到今天这个样子。

What are some moments that stick out of like, wow, that really changed things?

有哪些时刻你现在还记得,哇,那一下真的把事情改变了?

Dianne Penn00:09:11

Definitely when we were training and testing Opus 3, I think that was the moment when the company...

肯定是训练和测试 Opus 3 的那段时间,那是公司……

I think we were less than 200 people still at that point.

当时公司还不到 200 人。

And it was very clear that we needed and wanted to create a frontier model.

而且很清楚,我们既需要、也想要做出一个前沿模型。

And that was very important in terms of our ability to reach users, consumers, and to showcase our research.

这对我们能不能触达用户、触达消费者,能不能展示自己的研究,非常重要。

And we were looking for ways for also why should somebody choose Claude?

我们也在找一个说法:为什么别人要选 Claude?

And that was a core question.

这是个核心问题。

And that was a core question we were getting asked in the early days.

早期一直有人拿这个核心问题问我们。

And I think with Opus 3, it launched, I think, early March 2024, but there was many, many months of various teams across inference, across research, fine-tuning, pre-training that rallied at different points and towards a common goal.

Opus 3 是 2024 年 3 月初发布的,但在那之前有很多很多个月,推理、研究、fine-tuning、预训练各条线的团队在不同节点上一起冲,朝着同一个目标。

And I think everybody that was involved was really proud.

参与过的每个人都非常自豪。

I remember being the PM, us, the research leads, myself, we were all in our...

我记得当时我做 PM,我们几个、研究负责人、还有我自己,都在各自的……

This was around December, so we were all at home in our various parents' homes and seeing everybody's background of their childhood room.

那大概是 12 月,所以我们都在家,在各自父母家里,视频背景里能看到每个人小时候的房间。

And everybody was working really hard to figure out what are we training the model for?

每个人都在拼命弄清楚:我们训练这个模型到底是为了什么?

Is it showing up the right way?

它呈现出来的样子对不对?

So I think that was really powerful in terms of just building a lot of trust.

所以那段经历在攒信任上,力量非常大。

And a lot of our research leads have actually, from that time, are now leading reinforcement learning, leading our character work and alignment work.

而且我们很多研究负责人,从那时候一路走到现在,正在带强化学习、带我们的性格研究和 alignment 工作。

So that foundational trust I think also helped us work well now with any of our production models across product and research because we were working just so much in the trenches together in the early days.

所以那份打底的信任,也让我们现在在任何一个生产模型上,产品和研究都配合得很好,因为早期我们真的长时间一起泡在战壕里。

And then I think there were things like identifying that coding was important.

再有就是像「看准编程很重要」这样的事。

In 2023 when I started, nobody said Anthropic and Claude and coding in the same sentence.

2023 年我入职的时候,没有人会把 Anthropic、Claude 和编程放在同一句话里。

I think competitor models like GPT-4 at the time was used a bit for coding, but it was one of many use cases.

当时像 GPT-4 这样的竞品模型有一点用在写代码上,但那只是众多用例之一。

And one thing that, for example, I saw was people are starting to use these models, not just for code, not just code autocomplete, but actually writing long-form code.

举个例子,我看到的一件事是:大家开始用这些模型,不只是写代码、不只是代码自动补全,而是真的在写长篇的代码。

And is that an opportunity for us to train Opus 3 to be better at?

那这对我们是不是一个机会——把 Opus 3 往这件事上训得更强?

And it ended up being a relatively smaller change from a training perspective, but it ended up helping us differentiate in the early days competitively for users and actually bring a lot of the very early Claude enthusiasts and developers because we were providing a value that they didn't really think was possible at the time.

从训练的角度看,这最后是个相对小的改动,但它让我们在早期竞争里对用户做出了差异化,也真的带来了一大批最早的 Claude 拥趸和开发者,因为我们提供了一种他们当时并不觉得可能的价值。

Lenny00:12:10

It's so interesting you talk about Opus 3 like that's so long ago and just it's hard to think that was a big inflection.

你这么说 Opus 3 太有意思了,那都过去那么久了,现在很难想到那会是一个大的转折点。

And so this is really interesting to hear that that was internally a big milestone.

所以听到那在内部是个大的里程碑,真的挺有意思。

It almost feels like this confidence y'all built that, wow, we could really ship a frontier model, which is now today so not great if you compare it to what we've got today.

几乎像是你们从那里建立起一种信心:哇,我们真的能上线一个前沿模型——而那个模型放到今天,跟现在手上这些一比已经很不行了。

What I always think about is Opus 4.5, which was an interestingly a year later also during winter break when everyone was home able to code.

我一直会想到的是 Opus 4.5,有意思的是那也是一年之后,同样在寒假期间,大家都在家、都能写代码。

Was that another big milestone?

那是不是另一个大的里程碑?

Dianne Penn00:12:38

Yeah.

对。

Opus 4.5 was definitely another large moment.

Opus 4.5 绝对是另一个重大时刻。

I think what was magical about Opus 4.5 is we also now not just had a model, but a vehicle, which is a great product experience like Claude Code.

Opus 4.5 神奇的地方在于,那时我们不只有一个模型,还有一个载体,也就是 Claude Code 这样出色的产品体验。

One thing we say a lot on the team is you need frontier products in order to have frontier models and for people to feel the magic of frontier models.

团队里我们常说一句话:要有前沿模型,先得有前沿产品,也才能让人感受到前沿模型的魔力。

And I think we felt the magic of Claude Code for many months before that.

而 Claude Code 的魔力,我们在那之前好几个月就感受到了。

But the fact that the model essentially got to a level of intelligence where at a very broad level, users can experience both frontier intelligence in new use cases, allow it to run things end to end in agentic manner, I think that was the inflection.

但模型确实到了这样一个智能水平:在非常宽泛的层面上,用户既能在新的用例里体验到前沿智能,又能让它以 agentic 的方式端到端跑完,我认为那就是那个转折点。

It was actually both.

其实两者都有。

I think Opus 4.5 wouldn't have had that moment without a product like Claude Code.

没有 Claude Code 这样的产品,Opus 4.5 不会有那个时刻。

And Claude Code, I think, wouldn't have had that type of adoption accelerated without Opus 4.5.

而 Claude Code 这边,没有 Opus 4.5,采用也不会加速到那个程度。

Chapter 03

Inside the Exponential

身在指数曲线里面 · 涌现能力没人预告
13:50 — 20:03 · 适应力 · scaling laws 的另一张图 · 产品 overhang
Lenny00:13:50

So speaking on this thread, Dario, interestingly, if you look back at all his predictions, he's just like, "Okay, coding's going to be solved 100% in a year."

顺着这条线说,Dario 挺有意思的,你回头看他所有的预测,他就是那种口气:“好,编程一年之内就 100% 解决了。”

Something like that.

差不多是这个意思。

He kept talking about how AI is going to do all our code.

他一直在说 AI 会把我们所有的代码都写了。

And I remember everyone being like, "There's no way.

我记得所有人都在说:“不可能。

This is way too complicated.

这也太复杂了。

How is AI ever going to get really good at this very complex thing that humans do?

人类做的这么复杂的事,AI 怎么可能真做到很在行?

No, this is going to be humans for a long time."

不,这件事很长时间里还得是人来做。”

He was completely right.

他完全说对了。

Something else that he talks a lot about is this exponential.

他还经常谈的另一件事就是这条指数曲线。

Now that we're on, that's the way he describes it, now we're like, we're on the exponential curve.

现在我们已经在上面了,他就是这么形容的,现在我们已经在指数曲线上了。

I remember not long ago, new models were being released and everybody was like, "Okay, we're done.

我记得不久之前,新模型发布出来,所有人都说:“好了,到头了。

There's no more upside.

没有上升空间了。

It's plateauing.

已经走平了。

It's over.

结束了。

There's no more room to grow."

没有再往上长的余地了。”

And now it's the opposite.

而现在正好相反。

Now we're inside.

现在我们在里面了。

If you think about the curve of the exponential, we're inside of the exponential now, which by definition means every improvement is a massive jump because we're on that hockey stick part.

你想想指数那条曲线,我们现在就在指数曲线的内部,按定义讲,这意味着每一次改进都是一次巨大的跳跃,因为我们正处在曲线陡起来的那一段。

What's it like just being on the inside of this crazy historic moment when AI is improving so fast, so much is being unlocked?

身处这个疯狂的历史时刻内部,AI 进步这么快,这么多东西一下子解锁出来,是什么感觉?

What is it like and how should people prepare for the coming acceleration of more and more improvement from AI?

那是什么感觉,人们又该怎么为接下来 AI 越来越快、越来越多的进步做准备?

Dianne Penn00:15:05

One thing I like to say on the team is most of us weren't actively working yet when the internet transitioned from this novelty to something that everyone can use.

我在团队里常说的一件事是,互联网从一个新鲜玩意儿变成人人都能用的东西的时候,我们大多数人还没进职场。

And it feels like that's just taking humans...

感觉那件事就是在把人类带向……

I think analogies are helpful.

类比会有帮助。

And so the analogy of that is I think a couple of things.

所以类比过来,大概有这么几点。

Number one is adaptability becomes very important.

第一,适应力变得非常重要。

I think we have evals, we have on the safety side, safety testing, red teaming on the capabilities and product side, new prototypes, products like Claude Code, Tag and others.

我们有 evals,安全这边有安全测试、红队测试,能力和产品这边有新的原型,有 Claude Code、Tag 这些产品。

But it's very hard to predict the exact moment or the exact model.

但很难预测确切是哪个时刻、确切是哪个模型。

And so the adaptability of when you're faced with new information, how do you then make better decisions versus keeping the same plan?

所以适应力就是,面对新信息的时候,你怎么做出更好的决策,而不是守着原来那个计划?

And so that agility is really important.

所以这种敏捷度真的很重要。

I think another piece is with that, how do you actually be thinking very first principles and reason through what's next?

另一块是,在这个基础上,你怎么真正用第一性原理去思考、去推演下一步是什么?

What's the so what?

那又说明什么?

How do we invest in new products?

我们怎么在新产品上投入?

How do we invest in explaining the differences to users?

我们怎么投入去向用户讲清楚这些差别?

So a lot of the experiences I think of being in that exponential is that pace, understanding how you operate and make better decisions, and then applying that first principle's thinking to then do something that maybe we pull up a plan that we're expecting a few months from now, but now the model can actually do and work on and actually bring that to users.

所以身处指数曲线里的很多体验,就是那个节奏,是搞清楚你怎么运作、怎么做出更好的决策,然后用第一性原理思考去做一些事——比如把原本预计几个月之后才做的计划提前拿上来,因为现在模型其实已经能做、能跑,可以真的把它交到用户手里。

So this is things like CoWork, skills, Tag.

CoWork、skills、Tag 就是这样的例子。

It's a very positive self-enforcing loop.

这是个非常正向的自我强化循环。

And I think a big part of it also is just having the trust in each other, making sure we're thinking through the right decision making.

还有很大一部分是彼此之间的信任,确保我们把决策想清楚。

We're bringing folks along.

我们带着大家一起走。

Some teams might see the exponential feel it faster than others.

有些团队可能比别的团队更早看到、更早感受到指数曲线。

So how do we have the grace to bring the organization, the growing organization and company along on that?

那我们怎么有足够的体谅,把这个组织、这个还在变大的组织和公司一起带上去?

Lenny00:17:23

So what I'm hearing here is you almost don't know what will be possible with every model release.

我听下来是这样:每次模型一发布,你们自己几乎都不知道哪些事会变成可能。

And so the important things to focus on is being adaptable as things emerge.

所以要抓住的重点,就是随着新东西冒出来保持适应力。

To your point, the product itself has to catch up to what is possible.

按你说的,产品本身得追上那些已经变成可能的事。

To your point again, just like it can do so much, but people may not understand how to do it and may not be able to do it.

还是按你说的,它能做的事情非常多,但人们可能不知道该怎么做,也可能根本做不了。

So the product making it easy and even just telling you, here's something you could do feels like an important part.

所以产品把这件事变简单,甚至直接告诉你,这儿有件事你可以做,感觉是很重要的一环。

Is that roughly what you're describing?

大致就是你在说的意思吗?

Dianne Penn00:17:53

I think so.

差不多是的。

I think there's some really interesting graphs in the original scaling law papers.

最早那几篇 scaling laws 论文里有一些很有意思的图。

And I think folks are very familiar with the scaling laws in the lens of as you add in more compute and data, what's called loss, AKA the loss from next token prediction goes down.

大家很熟悉的 scaling laws 是这个角度:当你投入更多算力和数据,所谓的 loss,也就是下一个 token 预测的 loss,会降下来。

And so it's a very smooth linear curve of the models get more intelligent as you scale them up.

所以那是一条非常平滑的线性曲线:你把模型规模拉上去,模型就变得更聪明。

What's actually also interesting in that paper is there are these very different emerging capability graphs.

那篇论文里另一个有意思的地方是,还有一批完全不同的涌现能力曲线图。

And so for example, as you add in more data and you train the models with more compute, you essentially see these actually discontinuous emerging capabilities jump.

比如说,当你加入更多数据、用更多算力去训练模型,你基本上会看到涌现能力发生非连续的跳跃。

So the models go from one plus one being a thing that it can't calculate to a thing that it could reliably calculate.

模型会从算不出一加一,变成能稳定地算出一加一。

And so these emerging capabilities, this some nature of predictability is not necessarily everyone knows the exact moment.

所以这些涌现能力,这种可预测性的性质,并不是说每个人都知道确切是哪一刻。

You need the evals to be able to assess that has actually always been a part of how this technology works.

你需要 evals 才能评估它,这一直都是这项技术运作方式的一部分。

And also what makes things like safety harder.

也让安全这类事情变得更难。

Because unless you have the evals, unless you have the systems to test, these jumps might actually happen and you don't know.

因为除非你有 evals,除非你有测试的系统,否则这些跳跃可能真就发生了,你却不知道。

Lenny00:19:16

That's so interesting that you may have developed this AI brain that can do something you're not even aware of.

太有意思了,你们可能造出了一个 AI 大脑,它能做的事情你们自己都还没意识到。

And so part of the job is just uncovering, "Wow, we just got really good at this thing.

所以工作的一部分就是去发现:“哇,我们在这件事上突然变得很厉害了。

What can we do with that?"

那我们能拿它做什么?”

Dianne Penn00:19:27

I think there's product overhang and user overhang, to maybe put it in our PM language, even on today's models.

用我们 PM 的说法,即便是在今天的模型上,也存在产品 overhang 和用户 overhang。

And I think there's a lot that we could be exploring on our current opuses and definitely with Fable, for example.

我们现在这几代 Opus 上还有很多可以挖,Fable 上肯定更是,就是一个例子。

And that discovery is actually another part of what's been in the early days of Anthropic's DNA.

这种发现其实也是 Anthropic 早期 DNA 里的另一部分。

And I think it's also continuing to be a big part of how we operate in product, in Labs and across research.

它现在也仍然是我们在产品、在 Labs、在整个研究里做事方式的一大块。

Chapter 04

Token Maxing Is Not a Solo Sport

Token maxing · 实验未必是一个人的运动
20:03 — 23:33 · Garry Tan 的 $100,000 · 公开工作 · 十次请求里长出用例
Lenny00:20:03

This makes me think about something Garry Tan's been talking about, president of YC, I don't know what his title is.

这让我想到 Garry Tan 一直在讲的一个说法——他是 YC 的总裁,具体头衔我也说不准。

He had this interesting point that if you are willing to spend $100,000 a year right now in tokens, you are living the way somebody in 2028 is going to live.

他有个很有意思的观点:如果你现在愿意一年在 token 上花 $100,000,你过的就是 2028 年的人才会过的日子。

Because by then it'll be really cheap.

因为到那时候,这些会变得非常便宜。

Everyone can work this way.

所有人都能这么干活。

But there's this alpha opportunity right now to just live in the future, go crazy on token spend.

但眼下有个 alpha 机会:直接活在未来,在 token 上放开了花。

And so there's a big opportunity for people to learn what the future's like and also just build much faster.

所以这是个很大的机会,既能提前知道未来是什么样,也能把东西做得快得多。

Thoughts on this idea and the value of token maxing, let's call it.

你怎么看这个想法,还有 token maxing——姑且这么叫——的价值?

Dianne Penn00:20:37

Yeah.

嗯。

I think I take more of a almost product lens.

我更多是从产品的视角来看。

It's almost like token spend is more the input.

可以说,token 花销更像是输入的那一端。

And really the output is what you described of experimentation.

真正的输出,是你刚说的那种实验。

And I think if we were orienting goals around experimentation, I feel like that might be the better framing of the outcomes.

如果我们把目标围绕实验来定,这大概是对结果更好的一种框定方式。

And therefore there might be different ways of achieving that outcome.

那么达成这个结果,就可能有不同的路子。

I will say internally, some of the most creative thinkers, the best prototypers do spend a lot of time with Claude, with every new version of a research model that we have.

我得说,在内部,一些最有创造力的思考者、最好的原型开发者,确实花大量时间泡在 Claude 上,泡在我们每一版新的研究模型上。

And so there is something around you have to be using the models to then come up with good, then great, then better ideas.

所以这里面确实有一点:你得一直用这些模型,才能想出好点子,再到很棒的点子,再到更好的点子。

And there's no substitute for that.

这件事没有替代品。

It's very hard to come up with a perfect strategy without touching the technology when it's moving this quickly.

技术走得这么快,不去碰它却想拿出一套完美的策略,非常难。

At the same time, I think there's other things that we could be doing.

同时,还有一些别的事我们可以做。

So one thing that we do a lot is actually working in public internally within Anthropic.

我们做得很多的一件事,其实是在 Anthropic 内部公开地工作。

And so in the early days when we had less product surfaces, there was a Slack channel where everyone, almost the entire company was testing early versions of Claude and trying different use cases.

早期我们产品形态还少的时候,有一个 Slack 频道,里面所有人——几乎全公司——都在测试 Claude 的早期版本,试各种不同的用例。

People were not calling them use cases, but you might be asking it to edit an essay or to come up with the right way to send this email.

大家当时不会管这些叫用例,但你可能是在让它改一篇文章,或者想出这封邮件该怎么发才合适。

They were all different use cases, but we all worked in public.

它们都是不同的用例,但我们都是在公开地做。

And then what you would see magically is different users or different folks on the team coming up with an idea and then other people trying different variations of that idea.

然后你会神奇地看到,不同的用户、团队里不同的人冒出一个点子,接着另一些人去试这个点子的各种变体。

And then within maybe 10 or so requests, there was something magical or potentially in a use case that emerges.

然后大概十来次请求之内,就会冒出某种神奇的东西,或者说冒出一个潜在的用例。

And I think there's a lot in not just individuals figuring out by themselves how to use this technology.

怎么用这项技术,不该只靠个人各自摸索——这里面大有文章。

I think we could be doing more to actually bring that communal discovery when we do experimentation.

在做实验的时候,我们可以做更多事,把那种共同发现真正带进来。

Experimentation is not always necessarily a individual sport.

实验未必总是一个人的运动。

Lenny00:22:55

It's so interesting.

太有意思了。

Yeah, this idea that we're not sure what this is capable of or what we could do with it.

就是这个点——我们并不确定它能做到什么、我们能拿它做什么。

And it takes all this poking around and people trying things, hearing what other people are trying to figure out what's possible.

得靠大量这样的东摸西摸,靠人去试、去听别人在试什么,才能搞清楚什么是可能的。

Such an interesting, I don't know, technology.

这真是一种,怎么说呢,特别有意思的技术。

We're just like, okay, here's what...

我们就是,好吧,这就是……

Oh, I figured out it could do this thing.

哦,我发现它能做这个。

What are you going to do with that?

那你打算拿它做什么?

Dianne Penn00:23:14

I think in a broad theme we know.

大方向上,我们是知道的。

We know that the models could write great essays or could write long form writing, but individual pain points of what can you actually solve with that and bring it to a user level that people can use, I think is something that is more exploration or experimentation based.

我们知道模型能写出很好的文章、能做长篇写作,但一个个具体的痛点——你到底能用它解决什么,又怎么把它做到用户真能上手的程度——我认为这更多要靠探索或者说实验。

Chapter 05

Is There a There There?

Anthropic Labs · 非连续赌注怎么孵
23:33 — 27:30 · 强持主题、弱持原型 · 一个工程师起步的小组
Lenny00:23:33

So following this thread, you oversee product for the Labs team, which is extremely cool.

顺着这条线往下说,你负责 Labs 团队的产品,这特别酷。

We've had Ben Mann on the podcast, Mike Krieger, whom both work on Labs now.

我们请过 Ben Mann 上节目,还有 Mike Krieger,他们俩现在都在 Labs。

Talk about Labs.

聊聊 Labs 吧。

What is Labs?

Labs 是什么?

What's come out of Labs?

Labs 做出过什么东西?

Many people have heard of these things.

很多人都听说过这些东西。

And how do they work that enables them to create such innovative ideas outside of even the core Anthropic product team?

它们是怎么运作的,才能甚至绕开 Anthropic 的核心产品团队,做出这么有创新性的想法?

Dianne Penn00:23:58

The thesis of Labs in many ways is identifying and pulling the thread on...

Labs 的命题,在很多层面上,就是找出线索、顺着这条线往下拉——

... in many ways is identifying and pulling the thread on the thread of discontinuous large bets that might not be in the core roadmap, and figuring out is there a there there?

……在很多层面上,就是找出那些可能不在核心路线图里的非连续大赌注,顺着这条线往下拉,搞清楚:里面到底有没有真东西?

And also, what is the 10X, 100X, 1000X of the there there?

还有,这个真东西的 10X、100X、1000X 又会是什么样?

And so for example, things like Claude Code, I think-

比如说,像 Claude Code 这种,我觉得——

Lenny00:24:26

I've heard of it.

我听说过。

Dianne Penn00:24:29

Things like Claude Code, things like skills, and most recently Claude Design, MCP.

像 Claude Code,像 skills,还有最近的 Claude Design、MCP。

The thing that we really try to emphasize within the teams is, especially right now, there are so many things that could be built.

我们在团队内部特别想强调的一点是:尤其是现在,能做的东西太多了。

What does it mean then to have a discontinuous bet?

那么,什么才算一个非连续的赌注?

And I think one approach that we're taking this year is you can be very strongly held opinion about the theme or the area, and then more weakly held about the exact prototype.

今年我们的一个做法是:对主题、对方向可以持非常强的观点,对具体做成哪个原型则持得松一些。

And so there's a culture of experimentation.

于是就有了一种做实验的文化。

There's a lot of the bottoms up engineers on the team are very self-enabled, self-driven to test out different ideas.

团队里很多自下而上的工程师,都很能自主推进、自我驱动地去试各种不同的想法。

And sometimes, we have a thesis, and it might not work yet, and so we then might revisit it in one to two model generations.

有时候我们有一个命题,它暂时还跑不通,那我们可能过一到两代模型再回来看。

And so this idea of these prototypes that actually end up just helping us learn, that's also valuable even if it doesn't lead to something immediately shipping.

所以,有些原型最后只是帮我们学到了东西,这本身也有价值,哪怕它没有立刻带来能上线的产品。

And so I think that allows the incubation and the charter of Labs to really accelerate and see around corners more broadly for Anthropic.

这让 Labs 的孵化和它的使命真正提速,也能替 Anthropic 在更大范围上提前看到拐角后面的东西。

Lenny00:25:44

It's so funny to think about Labs within an Anthropic, which was already so innovative and creative and just shipping like crazy, that there's value to still creating a Labs team within Anthropic.

想想挺有意思:Anthropic 本身已经这么有创新力、这么有创造力,上线速度快得离谱,在它内部再建一个 Labs 团队居然还是有价值的。

What enables Labs to work as well as it has?

是什么让 Labs 能做到今天这么好?

Because you listed all these products and it's like, what else has Anthropic shipped?

因为你列的这些产品,让人想问:Anthropic 还上线过别的什么吗?

It feels like all the biggest wins almost.

感觉最大的那些成功几乎全在这儿了。

I'm sure there are many that I'm not thinking about right now.

我相信肯定还有很多我这会儿没想到的。

What's core to creating a successful Labs org within a larger company?

在一家更大的公司里做出一个成功的 Labs 组织,核心是什么?

Dianne Penn00:26:12

I think the team culture, similar to broadly at Anthropic, I think the team culture is very valuable.

团队文化吧,跟 Anthropic 整体的情况类似,团队文化非常有价值。

I think Ben sets an incredible vision and pushes people to think about the 10X, 100X of the idea.

Ben 定的愿景非常了不起,他会推着大家去想这个想法的 10X、100X。

And the teams, the pods within Labs is small, sometimes these ideas start with one engineer.

而且 Labs 里的团队、那些小组都很小,有时候这些想法就是从一个工程师开始的。

And I think sometimes when there's almost really large teams pursuing very ambiguous large ideas, you end up actually being slowed down because of that.

而有时候一个特别大的团队去追一个非常模糊的大想法,反而正因为这样慢了下来。

So I think it's culture.

所以,是文化。

I think we actually also select for folks who actually want to do that zero to one experimentation.

我们在选人上也确实是在筛那些真心想做 0 到 1 实验的人。

And it's not easy.

这并不容易。

There's a lot of bets that we end up turning down or turning off.

有很多赌注我们最后是拒掉、或者关掉的。

And maybe we revisit them in the future, but that's hard.

也许以后我们还会回头再看,但做这种取舍很难。

That's hard when you pour your heart and soul, you're acting as a founder for a bet and it's not working yet.

你倾注心血,像创始人一样为一个赌注扛着,它却还没跑通——那种时候是真的难熬。

So I think it's like that type, it's selecting for that type of personality, folks who are really passionate and deep about the zero to one.

所以就是那一类人,是在筛这种性格的人:对 0 到 1 真正有热情、也真的钻得进去的人。

Chapter 06

What Researchers Actually Do All Day

研究员在做什么 · 把「Claude 幻觉了」翻译成可执行
27:30 — 33:49 · 愿景与迭代两条环 · 反馈拆到 tool use 层 · 顶尖研究员的画像
Lenny00:27:30

So you lead product for the research team, you work with the researchers at Anthropic.

你负责研究团队的产品,跟 Anthropic 的研究员一起工作。

A lot of people get a sense of what is research.

很多人对研究是什么大概有个感觉。

What do researchers do?

研究员到底做什么?

I think a lot of people don't totally understand these very valuable people at all the AI labs.

我觉得很多人并不真的了解各家 AI 实验室里这批非常宝贵的人。

The way I think about it, and I want to help people understand, help me understand just what are researchers doing all day.

我自己是这么理解的,我也想帮大家搞明白——帮我理解一下,研究员一整天到底在干什么。

What I imagine is they have a hypothesis for how to improve the model.

我想象的是,他们对怎么改进模型有一个假设。

They find data, they tweak some algorithms, they adjust how it's trained, and they test it, see how it did, keep iterating and keep trying to find ways to improve the model.

他们找数据,调一些算法,调整训练的方式,然后测试,看效果怎么样,不断迭代,不断找改进模型的办法。

Is that roughly right?

大致是这样吗?

Slash help us understand what researchers are doing all day.

或者说,帮我们理解一下研究员一整天在干什么。

Dianne Penn00:28:09

I think that's a lot of maybe the more day-to-day.

你说的这些,大概更多是偏日常的那部分。

I think one piece around researchers and research organizations like at Anthropic is there's also a vision of the future more broadly.

研究员和 Anthropic 这类研究组织,还有一块是对未来更宏观的愿景。

So for example, things like, I think even at the founding of the company, researchers were talking about how do we get Claude to use a computer?

比如说,早在公司刚成立那会儿,研究员就在讨论:怎么让 Claude 会用电脑?

How do we get AI to navigate a screen?

怎么让 AI 会操作屏幕?

So there's a lot of actually very founder-like energy is how I describe it within researchers are really bold and ambitious researchers, and we have a ton of those at Anthropic.

研究员身上其实有很强的创始人劲头,我是这么形容的,他们是非常大胆、有野心的研究员,Anthropic 有一大批这样的人。

So there's one layer of vision of what this technology can go.

所以有一层愿景,是关于这项技术能走到哪一步。

And then I think on this other side of the loop, there's also, now that this technology or Claude is in people's hands, how do we make it better today?

然后在这个循环的另一头,还有一块是:这项技术、或者说 Claude 现在已经到了用户手里,我们今天怎么把它做得更好?

So it's a medium and long-term, and a lot of energy thinking about that lens of the future.

所以有中期和长期这一块,很多精力花在从未来那个视角去想。

And also, in the immediate and short term, what are the improvement areas we can make?

也有眼前、短期的部分:我们能改进的地方有哪些?

And so I think you're describing a really good sense of how do we make iterative improvements on different versions of Claude?

你刚才描述的其实抓得很准,就是我们怎么在 Claude 的不同版本上做迭代改进。

The way that my team works with researchers is being very integrated and embedded in those loops, particularly areas where there's a lot of impact on users.

我的团队跟研究员合作的方式,是深度嵌进这些循环里,尤其是那些对用户影响很大的领域。

So this is things like vision, computer use, coding, agentic coding, tool use, test time compute, things where there's a direct user impact.

比如 vision、computer use、编程、agentic 编程、tool use、test time compute,这些对用户有直接影响的地方。

And then figuring out what are the ways to bring the user feedback, and ground it in a level that is understandable for researchers, and also actionable for researchers.

然后是想办法把用户反馈带进来,把它落到研究员能理解、也能据此动手的层面。

And I think that's the second piece is actually a big part of the job and sometimes a hard part of the job.

第二块其实是这份工作里很大的一部分,有时候也是很难的一部分。

So for example, we might get feedback on claude.ai.

比如说,我们可能在 claude.ai 上收到一条反馈。

"Claude hallucinated."

“Claude 出现幻觉了。”

It's very vague.

这太含糊了。

If you bring that to a researcher and you say, "Please fix Claude from being hallucinated," it's not very actionable.

你把这个拿给研究员,说“请把 Claude 的幻觉问题修掉”,这没什么可操作性。

And so part of the time of the team is understanding, okay, what's the trajectory of why that user gave that feedback?

所以团队有一部分时间花在搞清楚:这个用户给出这条反馈,背后的会话轨迹是什么?

And it's consented.

而且这是用户授权的。

And so we look at, okay, what should Claude have called tools in that moment?

然后我们看:Claude 在那个时刻本来应该调用工具吗?

Or from its current knowledge or it looked at the right document, but it looked at the wrong facts.

还是说它是凭已有的知识回答的,或者它看的文档是对的,但看错了里面的事实。

In the first case, that would have been a failure on tool use.

第一种情况,那就是 tool use 上的失败。

On the second case, it would've been a failure on, let's say, search or knowledge insertion, search synthesis, or it could be something around alignment.

第二种情况,那就是——比如说——搜索、知识注入、搜索综合上的失败,也可能是 alignment 那边的问题。

And so bring that level of detail to researchers coming up with, is this a big enough problem?

然后把这个颗粒度的细节带给研究员,一起判断:这个问题够大吗?

Figure out things like evals to then describe what we've improved it.

再想出 evals 这类东西,用来说明我们把它改进了多少。

Those are the levels of actionability, and it's the day-to-day language of their researchers.

这些就是可操作的层级,也是研究员的日常语言。

And so we try to stay very close to how to bring that in an actionable manner between users to the core model training and the research development loop.

所以我们尽量贴得很近,把用户那边的东西用可操作的方式,带到核心模型训练和研发循环里。

Lenny00:31:36

I was talking to someone the other day about how it feels like AI research is the place to be now if you want to be very successful in life.

前几天我跟人聊到,如果你想在人生里非常成功,现在感觉 AI 研究就是该去的地方。

What does it take to become a really successful researcher from which you can tell?

就你看得到的范围,成为一个真正成功的研究员需要什么?

Not everyone's brain is going to work this way, but just say people are like, "Hey, I want to explore this career path."

不是每个人的脑子都适合这条路,但假设有人说“我想试试这条职业路径”。

From what you've seen, what does it take to make it there?

从你见到的情况看,要走通这条路需要什么?

Dianne Penn00:31:59

Researchers generally, or research and product managers working with research, or both?

是泛指研究员,还是跟研究打交道的研究经理和 PM,还是两个都要?

Lenny00:32:04

Let's do both.

两个都说吧。

Dianne Penn00:32:04

Yeah.

嗯。

Lenny00:32:05

But the researchers, PMs working with researchers also going to be very successful, but it feels like everyone's trying to poach all the top researchers across every company.

但研究员,还有跟研究员合作的 PM,也会非常成功,不过感觉现在每家公司都在挖顶尖研究员。

So I know you're not an AI researcher, but just from what you've seen, just what does it take to make it in that career path?

我知道你不是 AI 研究员,但就你看到的情况,在这条职业路径上走通需要什么?

Dianne Penn00:32:22

Yeah.

嗯。

I think a lot of the most successful researchers and research leadership at Anthropic are folks who are really strong first principles thinkers about problems.

Anthropic 这边最成功的研究员和研究负责人,很多都特别擅长对问题做第一性原理思考。

They reason through problems really well, who are just passionate about their research area and have a bold description of what that could look like.

他们把问题推演得很清楚,对自己的研究领域有热情,也能大胆地描述出它可能长成什么样。

And then who are actually close to the details.

而且他们真的贴近细节。

And so our leadership, our chief scientists, our heads of fine-tuning and RL, folks are actually really close to the training runs and actually look at things like how the training run is going, evals, looking at the underlying data.

我们的管理层、首席科学家、fine-tuning 和 RL 的负责人,都非常贴近训练过程,会真的去看训练跑得怎么样、看 evals、看底层数据。

So actually staying really close and be excited to be in the details, I think have been a sign of really strong researchers and developing taste.

所以真正贴得近、乐意钻进细节里,一直是研究员够强的标志,也是 taste 在养成的标志。

And I think another piece is just their ability to think big over time and be very ambitious.

另一块是他们能长期地把事情往大了想,非常有野心。

Like the Dario, like we can transform software engineering.

就像 Dario 那样:我们可以重塑软件工程。

And I think going in that direction, you learn so much.

朝那个方向走,你会学到非常多。

You have to shoot for the stars in many ways across your ideas, I think, in order to be a successful researcher.

要成为成功的研究员,你的想法在很多方面都得往最高处打。

Chapter 07

Assume Claude 8 Already Exists

更有野心的检验法 · 以及前沿模型的安全防护
33:49 — 38:21 · 向前兼容 · 兜底系统 · 能力越强门槛越高
Lenny00:33:49

I love just this meme of just be more ambitious comes up so often now, which is so hard.

我特别喜欢「更有野心一点」这个梗,现在到处都在讲,可这事真的很难。

It's easy to say that it's hard to actually just like how big can you think, and how that's so much of what AI now unlocks, just be more ambitious.

说起来容易,真要做到很难——你到底能想多大?而这恰恰是 AI 现在解锁出来的东西,就是更有野心。

Dianne Penn00:34:01

Yeah.

嗯。

Lenny00:34:02

Yeah.

对。

Dianne Penn00:34:03

I think it's thinking through it once or twice end-to-end.

就是端到端地把它想一遍、两遍。

And then being, I think, stubborn about the area and maybe more loose around the exact approach.

然后在方向上固执一点,在具体路径上可以松一点。

It is a question we challenge ourselves with, but the technology is moving so quickly, and so how do you make sure what you're building is actually forward compatible?

这个问题我们一直在拷问自己,但技术跑得太快,你怎么确保手上正在造的东西真的向前兼容?

And so it's also actually part of, I think, the core product development loop to think bigger.

所以往大了想,其实也是核心产品开发循环的一部分。

One thing I ask the team frequently or how I think about when we're building a product is let's say Claude 8 comes around.

我常问团队一个问题,也是我们做产品时我自己的思路:假设 Claude 8 出来了。

What changes in what users do?

用户做的事情里,什么变了?

And then what does that mean for how you're building today?

那这对你今天该怎么造,意味着什么?

Is it going to be forward compatible to that experience?

它能向前兼容到那个体验吗?

So just grounding, I think being ambitious is very broad.

所以就是要落到实处,「有野心」这个说法太宽泛了。

And so trying to ground it in some ways of describing that.

所以要试着用一些具体的说法,把它落下来。

Lenny00:35:13

And also, yeah, everything heading in a direction that all is cohesive and makes sense versus just ambitious in a completely different direction.

还有就是,所有东西都朝着一个连贯、说得通的方向走,而不是光有野心、方向却完全岔开。

Speaking of ambition and Claude 8, Fable/Mythos recently feels like hit this very new tipping point with models where it used to be you have an awesome model, release it.

说到野心和 Claude 8,Fable/Mythos 最近感觉是模型上一个全新的临界点——以前是你有个很棒的模型,发布出去就行了。

Hey everyone, welcome.

大家好,欢迎。

Opus 4.5 is out.

Opus 4.5 发布了。

Everyone can use it.

所有人都能用。

Mythos went in a very different direction.

Mythos 走的完全是另一条路。

We got blocked.

我们被挡下来了。

There was a lot of scrutiny, a lot of concern about what it was capable of.

外界大量审视,很多人担心它到底能干出什么来。

All the companies had to go make sure it wasn't going to hack into all their systems.

所有公司都得去确认它不会黑进自家所有系统。

And it feels like now every model, because they continue to get better, will now have a lot more scrutiny and there'll be more restrictions on who can use them, which feels like a big deal.

感觉现在每一代模型,因为一直在变强,都会受到多得多的审视,谁能用也会有更多限制,这看着是件大事。

How do you think about that?

你怎么看这件事?

How does that change the way you operate?

这会怎么改变你们做事的方式?

Dianne Penn00:36:02

I'm going to maybe leave the policy and the export control side to folks on that and work on that.

政策和出口管制那一块,我还是留给做那块、也在推进那块的同事。

I think the product question and how we interact with these internally is, I think as you mentioned, as frontier models become more capable, the safeguards and the ways of red teaming and testing and the pre-release process also needs to evolve and adapt quickly to address that.

产品这一侧,以及我们内部怎么跟这些东西打交道——正如你提到的,随着前沿模型能力越来越强,安全防护、红队测试和测试的方式、发布前的流程,也得快速演进、快速适配,来应对这一点。

And so one example is before Fable models, we didn't have as strong of, let's say, fallback UXs and systems because our goal was to make sure that there is asymmetrical benefit for this technology and to minimize the downside or a severe risk of it.

举个例子,在 Fable 系列模型之前,我们的兜底 UX 和兜底系统没有这么强,因为我们的目标是确保这项技术带来的收益是不对称的,同时把它的下行风险、或者说严重风险,降到最低。

And so we ended up building fallback systems so that users will still get a great response from Opus 4.A immediately.

所以我们最后建了兜底系统,让用户仍然能立刻从 Opus 4.A 拿到很好的回复。

And so I think there's a piece around as we evolve and improve safety systems, how do we continue to develop and deliver great user experiences?

所以这里有一块是:在演进和改进安全系统的同时,怎么继续做出并交付很好的用户体验?

I think there's more that we can do on both sides.

两边我们都还有更多可以做的。

And so you'll see us innovating, improving on what we call now the model safeguards package more and more in the coming weeks and month.

所以接下来几周、几个月,你会看到我们在现在叫做「模型安全防护包」的东西上,越来越多地创新和改进。

Lenny00:37:29

What's really interesting and just unexpected here is it creates this really interesting advantage for Anthropic where you have access to the latest stuff.

这里真正有意思、也完全没料到的一点是,它反而给 Anthropic 造出了一个很有意思的优势——最新的东西,你们自己拿得到。

And this is going to happen at every lab.

而且每一家实验室都会这样。

Everyone's going to keep improving and it creates this unfair advantage within the Labs to have access to the best stuff that other people can't yet outside of your control.

所有人都会继续变强,于是 Labs 内部反倒有了一种不公平的优势:最好的东西你们能用,外面的人还用不上,而这还不是你们能控制的。

You'd prefer everyone use it.

你们其实更希望所有人都用上。

So it's a really interesting, this new feedback loop that's going to start where models that are so advanced are only accessible to certain companies, and that's going to be a whole new unexpected, it's like a second order effect of all these restrictions.

所以这真的很有意思,一个新的反馈循环要开始了:先进到这个程度的模型只有某些公司用得上,这会是一个全新的、意料之外的——算是这些限制带来的二阶效应。

Dianne Penn00:37:59

Our goal is to be to develop these systems in the models to be as inclusive as possible.

我们的目标是把模型里的这些系统做得尽可能包容更多人。

I think our goal is to not have that happen for the general purpose, general use technologies and to make it more accessible.

我们的目标是别让这种事发生在通用技术、日常使用的技术上,而是让它更可及。

I think this is one of our top priorities right now to reduce what we're seeing there.

减少我们现在看到的这种情况,是我们眼下最优先的事情之一。

Chapter 08

Evals Are the New PRDs

招 PM 看什么 · 抠 token 得像抠像素一样抠
38:21 — 44:16 · Mercury 口播 · 三年没改的招聘 loop · 第一性原理
Lenny00:38:21

Yeah, that makes sense.

嗯,说得通。

I would imagine you'd want as many customers if people are using this thing as possible.

可以想见,如果大家都在用这个东西,你肯定希望客户越多越好。

This episode is brought to you by Mercury, radically different banking, loved by over 300,000 entrepreneurs, and now with Command.

本期节目由 Mercury 赞助播出——彻底不一样的银行服务,超过 300,000 位创业者都爱用,现在还有了 Command。

I've been a customer of Mercury's for over six years.

我做 Mercury 的客户已经六年多了。

I have never once thought about leaving.

一次都没想过要换。

Mercury is basically what happens when banking is built by product people, not by bankers.

银行要是由做产品的人来做、而不是由银行家来做,就是 Mercury 这个样子。

They make it so easy.

他们把这些事做得特别简单。

Dare I say fun to send invoices, move money around, set up virtual cards for folks on my team.

我甚至敢说,开发票、转账、给团队里的人配虚拟卡,都变得挺好玩。

Does your bank have an API, a terminal native CLI, or an AI-ready MCP server?

你的银行有 API 吗,有终端原生的 CLI 吗,有为 AI 准备好的 MCP server 吗?

I don't think so.

我看没有。

And just recently, they launched Command, a conversational interface built directly into Mercury, which acts as your financial operator.

而且就在最近,他们上线了 Command,一个直接内建在 Mercury 里的对话式界面,相当于你的财务操作员。

I've been using Command to transfer money around, to figure out what categories I've been spending the most money in, analyze my cash flows, and just today I used it to find out how much I've made from a specific sponsor over the past year.

我一直在用 Command 转账、搞清楚我在哪些类目上花钱最多、分析我的现金流,就在今天,我还用它查了过去一年从某个赞助商那儿赚了多少钱。

I just ask, "How much have I made from X over the past year?"

我就问一句:“过去一年我从 X 那儿赚了多少?”

10 seconds later, I have an answer.

10 秒之后,答案就有了。

It is so freaking cool.

这也太酷了。

Visit mercury.com to learn more and apply online in minutes.

访问 mercury.com 了解更多,几分钟就能在线申请。

Mercury is a FinTech company, not an FDIC-insured bank.

Mercury 是一家金融科技公司,不是 FDIC 承保的银行。

Banking services provided through Choice Financial Group and Column N.A., Members FDIC.

银行服务由 Choice Financial Group 和 Column N.A. 提供,两家均为 FDIC 成员。

I want to talk a little bit about how the product role is changing and who is doing well in this new world now that AI is such a core part of our life.

我想聊一聊产品这个岗位正在怎么变,以及 AI 已经成了我们生活里这么核心的一部分之后,在这个新世界里谁做得好。

When you're hiring PMs, product people, when you're looking at people that do well in today's world, what are some things that you notice?

你招 PM、招做产品的人的时候,你看今天这个环境里做得好的那些人,你会注意到哪些东西?

What are you looking for more most?

你更看重的是——最看重的是什么?

What are you looking for more?

你更看重什么?

What's trending up in what you find is important and what's trending down?

在你觉得重要的那些东西里,哪些在往上走,哪些在往下走?

Dianne Penn00:40:04

We actually, on my team, have not changed our hiring loop for three years now.

其实我的团队,招聘流程三年来一直没变过。

So what we actually look for and the traits and how we evaluate generalists like PMs, generalists like research, product managers have actually been the same.

我们真正看什么、看哪些特质、怎么评估 PM 这类通才、research 这类通才、产品经理,其实一直是同一套。

So I think some of those traits, number one is first principles thinking.

这些特质里,第一条是第一性原理思考。

And this is really rather than pattern matching what you used to do in, let's say, consumer product or B2B SaaS, but actually figuring out in this moment for this user group with this technology, what is the user value?

它真正的意思不是去套用你过去在消费产品或者 B2B SaaS 里做过的模式,而是在此时此刻、针对这个用户群、用这项技术,去搞清楚用户价值到底是什么。

Lenny00:40:50

Is there an example of that?

能举个例子吗?

A lot of people hear first principles thinking, they're like, "Yes, I got it."

很多人听到第一性原理思考,都会说:“对,我懂。”

I'm good at this.

我很擅长这个。

What's an example of someone having really demonstrated really good first principles thinking?

有没有哪个例子,是谁真正展示出了非常好的第一性原理思考?

Dianne Penn00:41:00

I think one example is I think you think of a product manager as I own product strategy and delivering user value, but I demonstrate day-to-day by writing a PRD or writing a product vision doc.

一个例子是,你会把产品经理理解成:我负责产品战略、负责交付用户价值,但我日常的体现方式是写 PRD、写产品愿景文档。

And for my team as research product managers, the way to drive user value is to figure out the right user feedback, the evals that then can be a personification of that user need.

而在我的团队,作为 research 产品经理,推动用户价值的方式是找到对的用户反馈,找到那些能把用户需求具象化的 evals。

So we do write some product documents and PRDs, but we actually have a saying on the team of evals are the new PRDs.

所以我们也写一些产品文档和 PRD,但团队里其实有一句话:evals 就是新的 PRD。

Because in order to deliver that user value, it's not that exact artifact that people used to write in the last one to two decades.

因为要交付那份用户价值,靠的已经不是过去一二十年里大家写的那个具体产物。

It's a new way of working.

这是一种新的工作方式。

And so the first principles thinking would be, let me figure out what is the thing I should do to achieve my goals, rather than here is a set of activities that I've done and therefore I will continue to do.

所以第一性原理思考就是:让我搞清楚,为了达成目标我该做的那件事是什么,而不是“这是我一直在做的一套动作,所以我会接着这么做下去”。

Lenny00:42:10

So the idea here used to be have an idea, create a PRD, talk to people about it, align on the plan, design it, build it, ship it, see how it goes, iterate.

所以过去的路子是:有个想法,写个 PRD,跟人聊,对齐计划,做设计,把它做出来,上线,看效果,再迭代。

What I'm hearing here is it's like, okay, here's some feedback about something that's wrong or an opportunity.

我听下来是这样:好,这里有一条反馈,说明哪儿不对,或者哪儿有机会。

Step one is the eval is now how you define what the work is versus a PRD.

第一步是,现在定义这件事要做什么的是 eval,不是 PRD。

Dianne Penn00:42:32

Maybe step one would be understanding the user pain point.

第一步也许是理解用户的痛点。

And so the way to even access a user pain point is different.

而且连接触到用户痛点的方式都不一样了。

In the past, we might do a user interview.

过去我们可能会做一次用户访谈。

I think if you go deep enough, you might have the user walk you through their user flow, the pixels.

挖得够深的话,你可能会让用户带你走一遍他们的用户流程,走一遍那些像素。

Here, you have to sweat the tokens as much as you sweat the pixels, and so one activity we have on the team is reading the transcripts and understanding what was the trajectories of failed very deeply to then say, was this a hallucination?

而在这里,抠 token 得像抠像素一样抠,所以我们团队的一项日常就是读 transcript,非常深入地去理解失败的会话轨迹是什么样,然后判断:这是一次幻觉吗?

Was this Claude being overconfident?

还是 Claude 过度自信了?

So the theme of the failure actually has a lot of nuance, and then that allows you to build a description, a sustained description of that pain point.

所以失败的类型其实有很多细微差别,而这让你能写出一段描述,一段对那个痛点持续成立的描述。

So that could be essentially in a new eval.

这基本上就能变成一个新的 eval。

And is the eval on distribution?

还有,这个 eval 在分布内吗?

Is it capturing both the positive situations where this is failing and also areas when it should actually not fail?

它有没有同时覆盖到确实会失败的那些正例,以及那些本来就不该失败的地方?

And then bring that back to, let's say, research so that we can make the improvements and actually measure the quality of, okay, when we have Opus 5.5, is this area improving or not?

然后把这些带回给 research,让我们能做出改进,并且真的去衡量质量:好,等我们有了 Opus 5.5,这块到底有没有变好?

Is Claude now able to identify the right places in the document and pull the right synthesis out?

Claude 现在能不能在文档里定位到对的位置、把对的综合结论提取出来?

So it's just the actionability and shorten the distance to actionability for our stakeholders and partner teams like researchers to take action on.

所以这就是可执行性——缩短那段距离,让我们的利益相关方和 researcher 这类合作团队能更快动手。

Chapter 09

Write the Eval First

一条 eval 是怎么写出来的 · PRD 没死
44:16 — 50:19 · 那 80% 其实是 JSON 写不对 · 30-40 个例子起步 · PRD 剩下的两个位置
Lenny00:44:16

Is there an example of something like this where you found an issue or opportunity and then wrote the eval?

有没有这样的例子——你发现了一个问题或者一个机会,然后写了一条 eval?

And what is the eval looking like in most cases?

多数时候,这条 eval 会是什么样子?

When people want to picture an eval, what do they picture?

大家想象一条 eval 的时候,脑子里浮现的是什么?

Dianne Penn00:44:29

We actually pioneered this concept within Anthropic.

这个概念其实是我们在 Anthropic 内部最早做起来的。

So one of the early examples is the early Claude models were not very good at following specific schemas.

最早的例子之一是,早期的 Claude 模型不太会遵守特定的 schema。

So things like outputs in JSON.

比如按 JSON 输出。

And now that is fundamental to Claude being able to be a good agent.

而现在,这是 Claude 能当好一个 agent 的根本。

If you can't output a certain format, you don't know how to access APIs, you can call tools, et cetera.

如果你输出不了某种格式,你就不知道怎么访问 APIs、怎么调用工具,等等。

And so the initial end-to-end was I was hearing feedback around Claude 2 days.

最初那条完整的链路是这样的:Claude 2 那阵子我听到一些反馈。

Claude was not very good at following instructions.

Claude 不太会遵循指令。

So then digging in with the users, what do you mean by Claude is not good at following instructions?

于是就跟用户往下挖:你说 Claude 不擅长遵循指令,具体指什么?

Give me what situations this was happening.

告诉我这是在哪些场景下发生的。

What's the exact paragraph?

具体是哪一段?

What did you ask?

你问的是什么?

What was Claude's response?

Claude 回的是什么?

Going to that level of detail.

一直追问到这种细节。

And what I saw was something like 80% of what people meant in the early days for this failure was Claude would not write the right JSON.

我看到的结果是,早期大家说的这类失败里,大概 80% 其实是 Claude 写不出对的 JSON。

And so then, okay, let's generate maybe to start, just 30 to 40 examples of when Claude was not doing this thing correctly.

那好,一开始先攒个 30 到 40 个例子,都是 Claude 没把这件事做对的场景。

And then that actually is your eval set.

而这其实就是你的 eval set。

And you could have essentially a prompt and a response.

基本上就是一个 prompt 加一个回复。

And if that is not working in the right golden answer that you might have, then that means that the eval essentially is beneficial because it's identifying a pain point consistently.

如果它跟你手上那个正确的标准答案对不上,就说明这条 eval 有价值,因为它能稳定地指出一个痛点。

And so then we added that to our repositories for evals, and when we have versions of Claude, we actually run that eval and just check.

于是我们把它加进 evals 的仓库,之后每出一版 Claude,就真的跑一遍这条 eval,检查一下。

I think at this point, it's always 100% or 99.9 and so it's no longer a pain point.

到现在,它一直是 100% 或者 99.9,所以早就不是痛点了。

But in the early days, it was taking the user feedback, figuring out actually what they mean.

但早期就是拿用户反馈,搞清楚他们到底在说什么。

Can we reproduce it?

我们能复现吗?

Is it consistent?

它稳定出现吗?

Is it a big issue?

这是个大问题吗?

And then figuring out how to standardize it in a way that can be consumable for researchers.

然后想办法把它标准化成研究员能直接拿去用的形式。

Lenny00:46:42

It's basically test-driven development for PMs is the world we're living now where you write the test first.

这基本上就是 PM 的测试驱动开发——我们现在就活在这么个世界里:先写测试。

So is this just a core part of the product management job now at Anthropic writing evals?

所以写 evals 现在是不是已经成了 Anthropic 产品管理工作的核心部分?

Dianne Penn00:46:54

I think so.

我想是的。

I also think it's something I've talked to other PMs at other companies about, and I think it's also more and more of the skillset more broadly because a lot of the products that we're building is at the intersection of models with harnesses, with a set of contexts for a set of users.

我也跟别的公司的 PMs 聊过这件事,而且它正越来越成为一项更普遍的基本功,因为我们做的很多产品都落在几样东西的交叉处:模型、harness,还有给一群用户准备的一组上下文。

And so having things like evals actually is a way, not just for folks working on models, but generally within product to get to better user experiences because you can't improve what you can't measure.

所以 evals 这类东西不只是做模型的人要用,产品这边普遍也能靠它做出更好的用户体验,因为你没法改进你测不了的东西。

And a lot of this is very still tactile-based.

而这里面很多事仍然非常靠手感。

It's still very judgment-based.

仍然非常依赖判断力。

And so you have to stay close to the details.

所以你必须贴着细节走。

Lenny00:47:38

And also very non-deterministic, which is a big part of this just like it's not going to give you the same answer every time so you got to describe it more broadly, it's not going to be an exact match.

而且它还非常非确定性,这也是很关键的一点——它不会每次都给你同样的答案,所以你得描述得宽一些,不可能是精确匹配。

So this is a really interesting change in the way product happens and will happen is evals, writing evals versus PRDs is a big part of this.

所以产品的做法上正在发生、也还会继续发生一个很有意思的变化:写 evals 而不是写 PRDs,是其中很大一块。

Do you guys still do PRDs?

你们还写 PRDs 吗?

Is there still a one-pager describing a problem or is it replay?

还有那种描述问题的一页纸吗,还是说已经被取代了?

Okay, now you're shaking your head yes.

好,你这是在点头说有。

Dianne Penn00:48:00

We are.

还有。

We do.

还在做。

Lenny00:48:01

Wait a minute.

等一下。

Okay.

好。

Now you're shaking your head.

你这是在点头啊。

Dianne Penn00:48:03

Yes.

对。

We are.

还有。

We do.

还在做。

I think when there's a very defined problem, I think things like evals might be almost a shorthand.

问题非常明确的时候,evals 这类东西差不多就是一种简写。

I think there's other cases where PRDs are really valuable.

但另一些时候,PRDs 是真的有价值。

PRDs are great vehicles for getting a very large group of people aligned on a set of sources of truth about experience and set of goals.

PRDs 是很好的载体,能让一大群人对齐到同一批东西上——关于体验的唯一信息源,还有一组目标。

So when we do have a model, we actually, for every model, we do have a PRD, less necessarily for our researchers, but more for our growing product surfaces, for our engineering teams, for our stakeholders like legal and safety and others as just a source of truth of putting together what we're aiming to achieve so that a big group of people can row in the same direction.

所以每有一个模型,我们每个模型都确实会写一份 PRD,不太是给研究员看的,更多是给我们越来越多的产品形态、给工程团队、给法务和安全这些相关方,就当一个唯一信息源,把我们要达成的目标摆在一起,好让一大群人朝同一个方向划。

The other place where I do think PRDs are valuable are on the more ambiguous problems and opportunities.

PRDs 另一个我确实觉得有价值的地方,是那些更模糊的问题和机会。

So if we haven't shipped a thing like computer use, we don't necessarily have a set of user specific pain points always.

比如 computer use 这种我们还没上线过的东西,手上不一定总有一组具体的用户痛点。

And I think there's value in the product vision portions of a PRD to explore what could, even if a technology is not yet ready to work for everyone, how do you get it to work well for some group so you can explore the value?

而 PRD 里产品愿景那部分是有价值的,可以去探索有什么可能——哪怕一项技术还没到能服务所有人的程度,你怎么让它先在某一群人身上跑得好,好去探它的价值?

You can actually bring something that is coherent to a user group.

你确实能拿出一个完整连贯的东西交给一群用户。

So we do have PRDs.

所以我们是有 PRDs 的。

I think the application's a little different now.

只是现在用法有点不一样了。

Lenny00:49:39

Okay, this is great.

好,这太好了。

I just had Andrew, he's the head of the Codex app at OpenAI, and you guys are aligned.

我刚请过 Andrew,他是 OpenAI 那边 Codex app 的负责人,你们俩的说法是一致的。

PRD's not dead.

PRD 没死。

Still very useful for specific projects and ideas.

在具体的项目和想法上还是很有用。

Great.

很好。

Okay.

好。

We've closed the book on...

这事我们就算翻篇了……

PRD is still kicking.

PRD 还活得好好的。

Okay.

好。

So we've been talking a bit about just what kind of skills are emerging for product people.

我们刚聊了一点,产品人身上正在冒出来哪些新技能。

Is there anything else that you find has shifted in what patterns are common across people that are doing well in this new AI world in terms of product managers and folks on the product teams?

在这个新的 AI 世界里做得好的那批人身上——产品经理,还有产品团队里的人——你还发现有什么别的变了,有哪些共通的模式?

Is there anything else that you're like, okay, this is something you got to shift or something you look for more people?

还有没有别的什么,是你会说“这个你得改过来”,或者是你在人身上会更看重的?

Chapter 10

Managers Have to Ship

当管理者也得亲自发版 · 以及怎么在 AI 里找到乐趣
50:19 — 58:11 · 资深 PM 和应届同一套 onboarding · 结对着玩 · 深挖一两个
Dianne Penn00:50:19

I think maybe specifically for folks who might be mid-career or folks who have been more in a managerial product leadership seat.

可能特别是对职业中期的人,或者更多坐在管理岗、产品负责人位置上的人。

One thing that I think I feel pretty strongly about is in order to be good managers of teams and PMs working with this technology, you have to be really hands-on yourself and have spent not just time tinkering, but actually shipping with this technology.

有一件事我看法挺坚定:要带好用这项技术做事的团队和 PM,你自己必须真的亲自上手,而且不只是花时间折腾,是真的用这项技术交付过东西。

And again, being in the details and sweating the tokens along with your PMs and your engineers and your teams.

还是那句话,待在细节里,和你的 PM、工程师、团队一起抠 token。

And so even for folks that I hire who have more tenured PM experience, the onboarding plans are exactly the same as somebody who is more early career.

所以哪怕我招的人 PM 资历更老,他们的 onboarding 计划也和职业早期的人一模一样。

And it's around understanding users, reading consent and user feedback, talking to customers.

内容就是理解用户、读用户授权的数据和用户反馈、跟客户聊。

I think there's something around being able to understand what to do with this, what good looks like, and having developed that in a very hands-on manner that's important.

能不能搞清楚该拿它做什么、什么才算好,而且这套东西是非常亲自上手地练出来的——这一点很重要。

It's not necessarily easy for someone to agree or be able to see what a good or great AI product or AI feature could look like if they haven't kind of experienced building themselves.

一个人如果没自己经历过做东西的过程,要他认同、或者要他看得出一个好的、乃至出色的 AI 产品或 AI 功能该长什么样,不一定容易。

So I do feel pretty strongly that if you're a manager, you have to be hands-on; you have to spend a portion of your time actually shipping.

所以我确实看法很坚定:你如果是管理者,就得亲自上手,得拿出一部分时间真的去交付东西。

You have to kind of walk in the shoes of your teams.

你多少得站到团队的位置上。

And I always try to carve out a portion of time to actually own one to two work streams when we have models in order to keep my theory of mind, keep my sense of how the models are moving, how quickly it's improving, so I can help the team make decisions and make better decisions.

而且每当有模型在手,我总会尽量切出一部分时间,真的自己扛一到两条工作线,好维持我的心智模型,维持我对模型在往哪走、改进有多快的手感,这样我才能帮团队做决策、做出更好的决策。

Lenny00:52:35

So what I'm hearing here is no matter where you are in the ladder of hierarchy at a company, if you're not building yourself, if you're not actually talking to Claude, talking to Codex, building stuff, you're not going to make it.

所以我听下来是:不管你在公司层级的哪一层,你自己不动手做,不真的去跟 Claude 对话、跟 Codex 对话、去做东西,你就混不下去。

Dianne Penn00:52:46

And you should have fun working with this technology.

而且你用这项技术的时候,应该是开心的。

I think that's the other piece.

这是另一点。

I think the folks that would be most successful, regardless of their level, are people who love working with AI and are exploring and experimenting.

不管在什么级别,最能做成的那批人,都是喜欢跟 AI 打交道、一直在探索和试验的人。

And carving out the time, not just for the experimentation, but actually hands-on shipping end-to-end, getting the user feedback, I think has to be fundamental for everyone.

而且切出时间——不只是用来试验,而是真的亲自端到端交付、拿到用户反馈——这对每个人都得是基本功。

Lenny00:53:12

I 100% know what you mean there.

我百分之百懂你在说什么。

Just me sitting on my newsletter and this podcast just talking about stuff and like, yeah, yeah, that sounds great.

光是我守着我的 newsletter 和这档播客,在那儿聊来聊去,然后“嗯嗯,听着挺好”。

Every time I actually build something, and I tinker with all kinds of little projects, you're just like, okay, I see what's happening here.

但每次我真去做点东西——我什么样的小项目都折腾——你就会“哦,我看明白这儿在发生什么了”。

And you just get so much more.

而且你收获的东西一下子多太多了。

It's hard to exactly describe what you experience actually working with the models and building stuff, but it's a whole different world of like, okay, I see.

真正上手用模型、动手做东西是什么体验,很难准确描述,但那完全是另一个世界:哦,我懂了。

Here's what they're talking about computer use; here's what they're talking about with this limitation of this UX situation.

哦,他们说的 computer use 是这个意思;他们说的这个 UX 场景的局限,是这个意思。

And you made this really interesting point that you have to have fun with it, which is not easy for a lot of people because they're pushed to use AI or they just don't know exactly what to do with it.

你刚才那个点很有意思:你得从这里头找到乐趣。这对很多人不容易,因为他们要么是被推着去用 AI,要么根本不知道该拿它干什么。

For people that are just like, I don't know, it's just so annoying.

有些人就是那种:我也说不上来,就是烦。

I just have to do this.

我就是不得不干这个。

I hate this frigging thing.

我讨厌死这破玩意儿了。

Why do I have to work with this?

凭什么非得我用它?

Things are changing so much.

什么都在变,变得也太多了。

I'm tired.

我累了。

Advice for helping people find that joy in this work.

有什么建议,能帮这些人在这件事里找到那份乐趣。

Dianne Penn00:54:06

I think maybe I'll reemphasize something I said earlier around just that experimentation is not an individual sport.

我可能会把前面说过的再强调一遍:试验不是一个人的运动。

Some of the moments where I think I've touched practically every version of research models across 20-plus versions of Production Cloud at this point.

有那么一些时刻——到现在为止,20 多个版本的生产版 Claude 里,几乎每一版研究模型我都上手摸过。

And I think part of the joy comes from seeing other people discover use cases too.

乐趣有一部分来自看到别人也找到用例。

And so maybe one idea here would be pairing with somebody who is excited and seeing what on a use case that you care about and working together versus identifying or trying to figure out the perfect use case yourself because that might feel like work.

所以这里也许可以这么做:找一个很兴奋的人结对,围绕一个你在乎的用例一起看、一起做,而不是自己去找、自己憋出那个完美的用例,因为那种感觉像在干活。

Working with others feels like joy a lot of the time, and is there more that we could do to bring other people along?

和别人一起做,很多时候是种乐趣;那我们还能再做点什么,把更多人带上?

That's something a lot of times internally we have somebody who is very curious, and them sharing an idea of a new prototype actually brings a ton more people who are like, oh, I didn't know this could work now with Claude.

内部经常有这样的事:某个人特别好奇,他把一个新原型的想法分享出来,一下子就带动了一大批人——“哦,原来 Claude 现在能做这个了。”

And so there's just some virtuous cycles here and ways of continuing to have joy with this technology.

所以这里就有一些良性循环,也有一些办法让你一直从这项技术里得到乐趣。

Lenny00:55:22

That's such a good point.

这个点太好了。

I think that's also why Twitter is so useful for a lot of this is you see other people sharing what they've done, and it inspires you to come up with your own little ideas.

我想这也是 Twitter 在这件事上这么有用的原因:你看到别人分享自己做的东西,它会激发你想出自己的小点子。

And also it's just fun to share your own thing that you've done.

而且把自己做的东西分享出去,本身就很好玩。

So that's a really good point.

所以这个点真的很好。

Just find other people to play around with and look for use cases.

去找别人一起玩,一起找用例。

The thing I've also heard a lot is just find a problem you want to solve in your life or work and just open up Claude, Claude Code, tell it, here's what I want to do.

我还经常听到的一条是:找一个你生活里或工作里想解决的问题,打开 Claude、Claude Code,告诉它我想干什么。

And it's incredible how far you can get just with a vague idea of a problem you want to solve.

就凭一个模模糊糊的、你想解决的问题,你能走出多远,简直不可思议。

Dianne Penn00:55:53

Yeah.

对。

I think it gets hard in that there's so many different things that you could try.

难就难在,你能试的东西实在太多了。

And so you just narrowing in on either pairing with someone, working with somebody who have a lot of joy about this technology, or figuring out something that you could immediately find value.

所以你要收窄:要么和别人结对,和那些从这项技术里得到很多乐趣的人一起做;要么找一件你能立刻拿到价值的事。

Either of those things allow you to go deeper rather than more high-level about too many things.

这两条路都能让你往深里走,而不是在太多事情上浮在表面。

I find it hard to keep pace with the number of prototypes or products that are out there.

外面的原型和产品多到我自己都跟不上。

And so my lens has been: how do I go deep in one to two of them myself?

所以我的角度一直是:我自己怎么在其中一两个上钻进去?

Lenny00:56:31

That's so interesting you say that because that's exactly it.

你这么说太有意思了,因为就是这么回事。

We just had the survey that I ran with my colleague Noam asking my readers just how they're feeling about all the things going on in the tech right now and AI.

我和同事 Noam 刚做过一次调研,问我的读者们对现在科技圈和 AI 里发生的这一切是什么感受。

And one of the most interesting takeaways we had was to find that happiness is exactly what you said is go deep in a couple things versus trying to just ton of little things.

我们最有意思的结论之一,就是发现幸福感恰恰是你说的这个:在几件事上钻深,而不是去试一大堆零碎的小东西。

Find a couple things to really solve well and then go deep.

找几件事真正解决好,然后往深里钻。

And that is a source because a lot of the happiness people feel is when they finally unlocked their way for AI to actually make their lives better versus just a couple messed up, broken, half-working things.

而这就是个来源,因为人们感到的幸福,很多时候是他们终于打通了让 AI 真正改善自己生活的那条路,而不是攒了一堆乱七八糟、坏掉的、只跑了一半的东西。

Dianne Penn00:57:08

Yeah.

对。

It's, how do you go from this being a check-the-box?

就是说,怎么才能让它不再只是打勾了事?

Right?

对吧?

And so us as product people, it's then an exercise of product prioritization of your time and your energy.

所以我们这些做产品的,这就变成一次产品优先级排序的练习,排的是你的时间和你的精力。

And if the goal is to experiment with joy, then what are the inputs that you need for that?

如果目标是带着乐趣去试验,那你需要哪些输入?

But yeah, I think the secret sauce of Anthropic is the culture and the bottoms of nature of how people work and this experimenting in public.

但我确实觉得,Anthropic 的秘方就是这个文化、大家自下而上的做事方式,还有这种公开做试验。

And by doing that, it's very much about how to bring other people along.

这么做,很大程度上就是在想怎么把别人也带上。

That ends up being, I think, really valuable.

这件事最后的价值非常大。

Lenny00:57:55

Yeah, I've heard this so many times from all the labs, just like no one's exactly sure how some of this is going to be used.

是啊,这话我从各个实验室那儿听过太多次了:没人完全说得准这些东西最后会被怎么用。

And a lot of it is just putting stuff out early, seeing how people use it, seeing what's possible, and then using that information to build the actual product to lean in.

很多时候就是早点把东西放出去,看人们怎么用,看什么是可能的,然后拿这些信息去做真正的产品、往那个方向压。

Dianne Penn00:58:10

Yeah.

嗯。

Yeah.

对。

Chapter 11

Coaching, EQ, and the Fear of Brain Rot

她自己怎么用 Claude · 先有自己的观点
58:11 — 1:04:49 · 《关键对话》做成 skill · 月度业务回顾整篇交出去 · 当审稿人不当写手
Lenny00:58:11

I'm curious how, kind of on this thread of finding ways for AI to help you in your work and life, are there any interesting ways you've been using Claude lately in your work as a PM?

顺着这条线——找各种办法让 AI 帮到你的工作和生活——我很好奇,最近你自己做 PM 的时候,有没有什么有意思的 Claude 用法?

Dianne Penn00:58:22

I think there's a lot of things with Fable and things like Tag.

Fable,还有 Tag 这一类,能做的事很多。

So I think Tag is in the very early days; I think there's something around how you work in a different paradigm of allowing an agent to go off and work and them bring back product and experiences to you.

Tag 还在非常早期;换一种范式来工作——让一个 agent 自己跑出去干活,再把产品和体验带回给你——这里面有点东西。

I think one area that it's not more recent, but one that I bring up a lot with the team and I think we could do more on using AI is just how to use it to also have better conversations with each other, to be better managers.

有一个方向不算新,但我常跟团队提,也觉得我们在用 AI 上还能做更多——就是怎么用它把彼此之间的对话谈得更好,做更好的管理者。

I don't think it's necessarily just about raising the IQ of experiences we build, but also I use it a lot and actually prepping for how to have better conversations in the moment during crucial conversations.

这不一定只是把我们做出来的体验的 IQ 提上去;我自己也大量用它来做准备——关键对话真正发生的那一刻,怎么把话说得更好。

So I love that book.

那本书我很喜欢。

And so I actually have a skill that helps me figure out, am I going in the right level of detail given the situation at hand, and actually helping me be a better manager and better supporter for the team.

所以我确实有一个 skill,帮我判断:眼前这个情境下,我下探的细节颗粒度对不对;它也确实在帮我做更好的管理者、更好地支撑团队。

So for managers on the team, that's actually a thing that I've been sharing more with our managers of, okay, how do you actually use Claude to make you a better coach?

所以团队里的管理者,这也是我最近更多在跟他们分享的:你到底怎么用 Claude,让自己成为更好的教练?

Because it's hard sometimes to find the right perfect words, and the models have a lot of perfect and right words.

因为有时候很难找到那句刚刚好的话,而模型有大量刚刚好、恰当的措辞。

And I think there is something about how it can actually augment us from an EQ perspective in addition to IQ.

而且除了 IQ,它还能从 EQ 那一面增强我们——这里面是有点东西的。

Lenny01:00:04

Oh man, there's so much interesting stuff there.

天,这里面有意思的东西太多了。

So just to understand what you're doing there.

我先弄清楚你具体是怎么做的。

So you built a skill; you're just like, "Help me, Claude, build a skill."

你做了一个 skill;就是跟它说:“Claude,帮我做个 skill。”

Pulling in lessons from Crucial Conversations, the book, which it knows enough about.

把《关键对话》那本书里的心法引进来——它对这本书了解得够多。

You don't have to even give it the content.

你甚至不用把书的内容喂给它。

And then you use that skill to talk to Claude, "Hey, I have this very difficult conversation coming up with a colleague.

然后你用这个 skill 跟 Claude 说:“我马上要跟一位同事有一场很难的对话。

Give me some tips on how to approach it."

给我点建议,这事该怎么开口。”

Dianne Penn01:00:25

Yeah.

对。

And it's a great, it's almost like coaching, individualized, personalized coaching of just how to make you...

它很好,几乎就是一种教练,针对个人的、量身定制的教练,教你怎么让自己……

And there's so much context switching that we do all day.

我们一整天要做大量的上下文切换。

And having Claude help me pair and help me...

让 Claude 跟我搭档、帮我……

And maybe there are times where I end up not using suggestions from Claude, but it actually ends up being very helpful for just coming up and brainstorming.

也许有些时候我最后没用 Claude 的建议,但拿它来冒想法、做头脑风暴,确实很管用。

Am I thinking about reactions in the right way?

我预想对方的反应,方式对不对?

How do I actually go a bit deeper, faster, build trust faster, be more direct?

我怎么才能挖得更深一点、更快一点,更快建立信任,更直接?

Lenny01:01:05

Yeah.

嗯。

Man, I have so many questions here.

我这儿问题太多了。

This is so interesting.

这太有意思了。

One is just there's concern people are going to start talking the way AI writes because they're talking to AI so much, and it's going to be like, "Dianne, it's not this, but it's that."

一个是,有人担心大家会开始照 AI 写东西的腔调说话,因为跟 AI 聊得太多了,会变成:“Dianne,这不是这个,而是那个。”

I know that you're not doing that, but that's a concern people have.

我知道你没这样,但这是大家有的一个担忧。

Let me just ask about that, I guess.

那我就问问这个吧。

Do you fear there's this brain-rot atrophy stuff people talk about it where you're just so reliant on AI now and we stop learning and thinking and overlaying AI thoughts on that, being so close to it and being so integrated with AI constantly?

你会担心大家在说的那种脑腐、大脑退化吗——现在这么依赖 AI,我们不再学习、不再思考,只是把 AI 的想法叠在上面,跟它贴得这么近、时时刻刻跟 AI 融在一起?

Dianne Penn01:01:38

A lot of actually thinking process and writing process are tied together for me personally.

就我个人而言,很多思考过程和写作过程是绑在一起的。

And so I think there are ways where I use Claude to augment my thinking, but what I want to make sure, and maybe this is what you're describing, is Claude doesn't take over all of my thinking for me.

所以我会用一些方式让 Claude 增强我的思考,但我要守住的一点——可能就是你说的这个——是别让 Claude 把我的思考整个接管过去。

And so I think depending on the situation, depending on how much more personal judgment I want to have in a situation, I might come up with my own POV first and then work with Claude through that, and making sure that I maintain my sense and tone throughout.

所以要看情况,看这件事上我想保留多少自己的判断力:我可能会先出自己的 POV,再带着它跟 Claude 往下推,全程保住我自己的感觉和语气。

I think there are then other things like updates.

另外还有一些别的,比如进展汇报。

We have monthly business reviews, and then in those cases it's much more, I actually want it to be standard, and I want it to be much more like it gets a crystallized information in the right way.

我们有月度业务回顾,那种场合就更多是——我其实希望它是标准化的,希望它更像是用正确的方式把信息结晶出来。

And I have a skill, and we're augmenting and improving our skill for that.

我有一个 skill,我们也在为此增强、改进这个 skill。

I want to get to a place where the monthly business review, the writing of that, is potentially asymmetrically less valuable than the thinking.

我想走到这样一步:月度业务回顾里,写的那部分,价值可能远低于想的那部分,两头是不对称的。

And so how do I get that piece delegated to Claude fully?

那我怎么把这一块整个交给 Claude?

And I'm more of a reviewer and a verifier of that information.

我更多是这些信息的审阅者和核验者。

So I think it depends on what you're using Claude for and what you're trying to convey, and is there asymmetrical value in delegating more to Claude.

所以这要看你拿 Claude 做什么、想传达的是什么,也看多交一点出去,换来的价值是不是不对称。

Lenny01:03:11

What I'm also hearing, the first tip is really great, which was think first, have a point of view, and then use Claude as a sparring partner almost to evolve the idea, push back on the idea.

我还听出一层——第一条建议特别好——先自己想,先有一个观点,然后差不多是把 Claude 当陪练,用它把想法往前推、反过来挑这个想法。

Dianne Penn01:03:22

Yeah.

对。

Yeah.

对。

And I think this is where things like actually our alignment research and safety research is helpful because what you don't want is AI that just agrees with you.

这正是我们的 alignment 研究和安全研究派上用场的地方,因为你不想要一个只会同意你的 AI。

What you want is this technology to actually augment and grow and get to a better outcome.

你想要的是这项技术真的能增强你、让你成长、得到更好的结果。

And so sometimes it's having Claude push back makes me better.

所以有时候,是 Claude 反驳我,让我变得更好。

And so that's great.

这很好。

Like a coworker, I want somebody to push back when my ideas are not fully formed.

就像同事一样,我的想法还没完全成型时,我希望有人来反驳我。

Lenny01:03:53

I want to hear more about that.

我想多听听这个。

I've heard that when Ben Mann was on the podcast, he talked about the constitution that is built into Claude and how unintuitively the work and the focus on safety and alignment, as you said, and this constitution that describes how Claude should think and operate.

我听过 Ben Mann 上这档播客时讲到 Claude 内置的那部宪法,讲到——很反直觉地——在安全和 alignment 上下的这些功夫,就像你说的,还有这部写明 Claude 该怎么思考、怎么运作的宪法。

You would think that would limit the abilities of Claude and make it less fun and interesting.

你会以为这些会限制 Claude 的能力,让它变得没那么好玩、没那么有意思。

It's exactly the opposite.

结果恰恰相反。

Claude is the most interesting personality.

Claude 的人格是最有意思的。

I hear that constantly.

这话我一直听到。

It's just like I much prefer talking to...

就是那种,我更愿意跟……

OpenClaw famously was built on Claude, and then people were forced to switch.

OpenClaw 众所周知是基于 Claude 做的,后来大家被迫切换。

We won't get into it.

这个我们就不展开了。

Were forced to switch to ChatGPT, and they're like, "This is so bad.

被迫切到 ChatGPT,然后大家说:“这也太糟了。

This is not who I'm used to talking to."

这不是我平时聊天的那个人。”

So that is, I think, a really interesting point.

所以我认为这是一个非常有意思的点。

I just want to make sure we spend a little time on.

我想确保我们在这上面花一点时间。

Why is it?

为什么会这样?

Why is that the case?

为什么是这个结果?

Just this focus on alignment, safety, having this clear constitution?

就是这种对 alignment、对安全的专注,以及有这样一部清晰的宪法?

Why does that make Claude better and more interesting to talk to you also?

为什么这会让 Claude 变得更好,也更有意思、更让人愿意跟它聊?

Chapter 12

A Thinking Partner Doesn't Just Agree with You

会顶嘴的 Claude 才有用 · 以及 AI 为什么还写不好
1:04:49 — 1:11:41 · 让 Claude 给自己定价 · 主动性 · 谁签字比谁执笔更重要
Dianne Penn01:04:49

In order to make Claude as intelligent and as capable as possible, being able to have Claude actually push back in the right points and then add, it's like a yes or no and actually helps you come to a better conclusion.

要让 Claude 尽可能聪明、尽可能有能力,就得让它能在该反驳的地方真的反驳,然后再往上加东西——不只是一个 yes 或 no,而是真的帮你得出更好的结论。

So I've used Claude to help with things like, are we making the right pricing decision on the next version of Claude?

所以我用 Claude 来处理这样的事:下一版 Claude 的定价,我们这个决策做得对不对?

It's a little bit meta, but using a research version of Opus, asking it to figure out how it should price.

这有点自我指涉,但就是用一个研究版的 Opus,让它去想该怎么定价。

And being able to come out with better outcomes is a goal at the end of the day.

说到底,目标就是能得出更好的结果。

And so having AI not just be an assistant, not just be a doer and being delegated task, but figuring out is it doing the right thing?

所以,让 AI 不只是助理,不只是一个干活的、听人派活的角色,而是去判断:它做的事对不对?

That's actually very integrated with knowing when to push back.

这跟知道什么时候该反驳,其实是分不开的。

That's part of knowing when you should be proactive.

这是「知道什么时候该主动」的一部分。

Proactivity is not necessarily always doing a thing that you are scheduled to do.

主动性不一定总是去做那些已经排好的事。

It is knowing when to come up with a new idea.

而是知道什么时候该提出一个新想法。

And so in order for Claude to be more useful, the general approach has to be that it knows when to push back.

所以,要让 Claude 更有用,总的思路必须是:它知道什么时候该反驳。

It's a core part of the characteristics together of the models.

这是模型整体特质里的核心一环。

Lenny01:06:10

That is so interesting.

这太有意思了。

It's so interesting that that is what a big part of it being less compliant is almost what makes it better and more useful because we need that.

特别有意思的是,它没那么顺从,这一点反而很大程度上让它更好、更有用,因为我们就需要这个。

I've had so many people where they're like, "Hey, AI told me I was right."

我碰到过好多人跟我说:「嘿,AI 说我是对的。」

And they're like, "No, I wish it was to other people."

然后他们又说:「不,我倒希望它是对别人这么说。」

Dianne Penn01:06:26

Yeah.

对。

And it comes back to our earlier point around thinking.

这又回到我们前面说的思考那一点。

How do you protect your thinking?

你怎么保护自己的思考?

If you have an AI that can be a thinking partner, a thinking partner doesn't just agree with you.

如果你有一个能当思考搭档的 AI,那思考搭档不是只会点头的。

It should add to you.

它应该给你添东西。

And you should come away at the end of the day having better ideas because you worked with Claude.

而且说到底,跟 Claude 一起做完,你带走的想法应该更好。

That should be the hero goal, not just making your ideas 10% better.

这应该是最该追的目标,而不只是把你的想法提升 10%。

Lenny01:06:52

Yeah, I love this sense.

对,我很喜欢这个说法。

It used to be think 10X.

以前是「往 10X 想」。

It used to be the way founders push people like, what if we 10X this?

以前创始人推着大家往前,靠的就是这句:如果我们把这个做到 10X 会怎样?

And I love what I keep hearing is it's like, how do we go 1,000X from this idea?

而我特别喜欢现在一直听到的说法:我们怎么从这个想法做到 1,000X?

What is the most ambitious version of this?

这件事最有野心的版本是什么?

I want to come back to something that I was thinking about as we were talking about talking to Claude constantly.

我想回到刚才聊「一直跟 Claude 说话」的时候我在想的一件事。

It's very clear when AI has written something still.

现在 AI 写的东西,还是一眼就能看出来。

It's funny that it's a large language model.

有意思的是,它是个大语言模型。

You would think, of all things, it would be very good at writing, and interestingly, just no AI is very good at writing.

你会觉得,别的不说,它写东西应该特别好,但有意思的是,没有哪个 AI 写得特别好。

It's always very clear this was AI-written.

永远都很明显,这是 AI 写的。

Do you think we'll get to a place where we will not know this was AI?

你觉得我们会走到那一步吗——看不出这是 AI 写的?

Dianne Penn01:07:32

I think it depends on what's the goal that you're looking to achieve by knowing or known-

这取决于,你想靠「知道」或者「被知道」达成什么目标——

Lenny01:07:39

What's the eval?

eval 是什么?

Dianne Penn01:07:40

Yeah.

对。

What's the eval?

eval 是什么?

I actually do think there's more that we could be doing on making Claude write better.

我确实觉得,在让 Claude 写得更好这件事上,我们还能做更多。

There's actually very active efforts on my team and on the research side about making Claude write better, just generally.

我团队这边和研究那边其实都有很活跃的工作在推,就是整体上让 Claude 写得更好。

I think it should be clear where an idea is being led by you or by you, Lenny, or me, Dianne.

应该能看清楚,一个想法是由你主导,还是由你 Lenny、由我 Dianne 主导。

I think it really depends on what's the goal of that writing.

这真的取决于那次写作的目标是什么。

For something like a monthly business review, I would actually love to have that end-to-end be written by Claude.

像月度业务回顾这种,我其实特别希望它从头到尾都由 Claude 来写。

And-

而且——

Lenny01:08:22

Obviously, and not make it feel like it was written by a human.

而且显然,别把它弄得像是人写的。

That's such an interesting point you're making.

你说的这一点特别有意思。

Is it actually better for us to know that it's AI versus not?

知道它是 AI 写的,跟不知道比,是不是其实更好?

Dianne Penn01:08:29

Yeah.

对。

But it's also for maybe the lens is more around verifiability or who's verifying the output.

但也许,这事该看的更多是可验证性,或者说谁在验证输出。

Right?

对吧?

Who's signing off?

谁来签字?

Maybe less around who's writing, but who's verifying who's signing off?

也许重点不在谁执笔,而在谁验证、谁签字?

That becomes more what matters than who's writing it.

比起谁写的,这才更要紧。

Lenny01:08:52

Why do you think AI is not great at writing?

你觉得 AI 为什么写不好?

My guess is it has studied all of the best writing in all of humanity.

我猜,它把人类所有最好的写作都学过了。

It's figured out, here's the best way to write.

它琢磨出来了:最好的写法就是这样。

And there's only so many ways to write.

而写法就那么几种。

And so we've just recognized, okay, this is what AI does.

于是我们就认出来了:好,这就是 AI 的路数。

It has these tropes.

它有那么几套套路。

Is that the core of it?

这是核心原因吗?

Is there something else that's keeping it from being a great writer?

还是有别的什么东西,让它成不了一个好写手?

Ironically, being a large language model of all things, you think you'd be really great at language.

讽刺的是,它偏偏是个大语言模型,你会以为它对语言应该特别在行。

Dianne Penn01:09:21

I think part of it is also we need to invest more in training improvements to make AI continuously strong on areas like writing.

一部分原因也在于,我们需要在训练改进上投入更多,让 AI 在写作这类领域上持续变强。

I think it's also the technology's jagged-edged, like we mentioned.

也因为这项技术是能力参差的,像我们前面提到的。

So sometimes when the models were good at writing, but not agentic, our thesis is: how do we make the models more agentic or call the right tools?

所以有时候模型写作还行,但不够 agentic,我们的判断就是:怎么让模型更 agentic、或者能调用对的工具?

Now that that's improved a bit, then it's, well, now it's these other areas actually become more of the rough edges.

现在这块改善了一些,那就变成:好吧,现在换成另外那些地方成了更粗糙的边角。

And so I think we're in one of those moments worth writing where we need to actually just focus and prioritize on training the models to be great at this area.

所以在写作上,我们正处在这样一个时刻——需要真正集中精力,优先把模型在这个领域训练好。

And that is a very active area for us.

这在我们这儿是个非常活跃的方向。

So fun that you mentioned.

你提起这个,挺巧的。

Lenny01:10:11

Okay.

好。

I'm glad.

那太好了。

I'm glad.

太好了。

And also it was going to be interesting once AI is so good, we're like, "I don't know who wrote that."

还有一点也会挺有意思:等 AI 好到那个程度,我们会说:「我也不知道那是谁写的。」

But to your point, sometimes we actually want to know that it's AI.

但按你说的,有时候我们其实就是想知道那是 AI 写的。

That's really interesting.

这真的很有意思。

I never thought of it that way.

我从来没这么想过。

The other interesting part of this is that there's that comedian who was joking that we're on a plane and the WiFi's down and we're just like, "What the hell?

这件事另一个有意思的地方是,有个喜剧演员讲过这么个段子:我们坐在飞机上,WiFi 断了,我们就在那儿嚷:「搞什么啊?

The WiFi's not working on this plane.

这飞机上的 WiFi 不好使。

This sucks.

太烂了。

How dare you?"

你们怎么好意思?」

When you're in a tube in the sky flying like a bird.

可你人正坐在一根飞在天上的铁管子里,像鸟一样飞着。

And how dare you complain that the WiFi doesn't work?

你还好意思抱怨 WiFi 不好使?

Your point is there's so much advancement and so much power.

你的意思是,进步已经这么多了,能力已经这么强了。

We can't fix it all.

不可能什么都修好。

We can't make it all work the best possible and so basically writing has been not the priority, and it feels like there's more investment happening there.

也不可能让每样东西都做到最好,所以写作基本上一直不是优先项,而看起来现在那边的投入在变多。

Dianne Penn01:10:53

Yeah.

对。

I think tone and character is a priority.

语气和性格是个优先项。

I think it's this advancement of the technology is a work in progress and so we see a leap or an emergence of a jump in agentic behaviors.

这项技术的推进是个进行中的过程,所以我们会看到 agentic 行为上出现一次跃迁、一次涌现式的跳跃。

And so that is a new normal.

那就成了新的常态。

And then these other capabilities needs to continue improving.

然后另外那些能力需要继续改进。

And I think once we improve, let's say, writing and tone and character, we probably will say, how do we have Claude be even more proactive?

而等我们把写作、语气和性格改进了,大概又会问:怎么让 Claude 更主动?

Proactivity is an opportunity, and that's human nature.

主动性是一个机会,而这也是人的天性。

We want to make ourselves better.

我们想让自己变得更好。

We want to make this technology better and I think we're applying it to AI, which is the right thing.

我们想让这项技术变得更好,我们正把这一点用到 AI 上,这是对的。

We should be making it better.

我们就该让它更好。

Chapter 13

Where Human Brains Still Win

判断力、坚持,和自己内心的声音
1:11:41 — 1:17:01 · 判断力是攒出来的 · Claude Science · 给孩子留一点挣扎
Lenny01:11:41

I want to ask you a couple questions I like to ask folks working at the very center of the future that is coming.

有几个问题我喜欢问那些身处这个正在到来的未来最中心的人,也想问问你。

One is, where do you think human brains will continue to be most valuable over the years?

第一个是,未来这些年里,你觉得人脑会在哪些地方继续最有价值?

I know Anthropic's mission and vision will reach AGI, a super intelligence.

我知道 Anthropic 的使命和愿景是走到 AGI,也就是超级智能。

So in the future, maybe nowhere.

所以在未来,可能哪儿都不再有价值了。

But before we get there-

但在我们走到那一步之前——

Intelligence.

智能。

So in the future, maybe nowhere, but before we get there, where do you think human brains will continue to be most valuable, as we approach that timeline?

所以在未来可能哪儿都不再有价值了,但在走到那一步之前,随着我们逼近那个时间点,你觉得人脑会在哪些地方继续最有价值?

Dianne Penn01:12:09

We started to talk about making Claude and models better at judgment, especially in the last year or so.

我们开始讨论怎么让 Claude 和模型在判断力上做得更好,尤其是最近一年左右。

I think judgment is an area where, it's accumulation of so much nuance and so much experience.

判断力这块,是大量细微差别和大量经验累积起来的。

And these systems haven't experienced as much as humans have.

而这些系统经历过的,还没有人类多。

And so I think that hard-earned judgment is a area for product leaders and just generally will continue to be really critical.

所以这种辛苦挣来的判断力,是产品负责人的一块领域,总体上也会继续非常关键。

There are so many things AIs can build.

AI 能造的东西太多了。

Which one are the things that an org like Labs should build?

其中哪些才是像 Labs 这样的组织该造的?

A lot of that requires human judgment, persistence, so proactivity.

这里面很多都需要人的判断力、坚持,也就是主动性。

These are all traits that are beyond just general capabilities, but just behaviors and characteristics of people at that level of how do you get to the best solutions?

这些特质都超出了通用能力本身,而是人在“怎么拿到最好的解法”这个层面上的行为和特征。

How do you create the best experiences?

怎么做出最好的体验?

So I think those types of traits are actually the tactile traits that I think will continue to be important.

所以这类特质才是那种实打实的特质,会继续重要下去。

I think there is also still a lot of capabilities and subject matter expertise as well.

另外,能力和专业领域知识这块也还有很多。

I think software engineering has been really transformed by AI.

AI 已经彻底改变了软件工程。

I think there's areas like biology, life sciences.

还有生物、生命科学这些领域。

These are all things that we're just at the foot of the exponential on.

这些我们都还只是站在指数曲线的脚下。

Maybe software engineering, we're on the exponential.

软件工程可能已经在指数曲线上了。

On some of these other areas, we're not quite there yet.

另外一些领域,我们还没到那儿。

And so I think you're seeing us shift things like Claude Science investing in these areas because those are areas that I think is just bring this technology to society and having a positive benefit for society.

所以你会看到我们把 Claude Science 这类方向往这些领域上投,因为正是这些领域能把这项技术带进社会,给社会带来正向价值。

So I think there's a lot more to go there.

所以那边还有很长的路要走。

Lenny01:14:11

Another question I want to ask is, as someone with kids, how do you think about what you are encouraging them to learn?

我还想问的一个问题是,作为一个有孩子的人,你怎么考虑该鼓励他们去学什么?

Where do you think you're going to nudge them to be successful in this wild new world that we're entering?

在我们正走进的这个疯狂的新世界里,你觉得你会把他们往哪个方向推,好让他们成功?

Dianne Penn01:14:27

I actually think it's a lot of the same traits you and I probably grew up with, which is curiosity for learning, persistence, believing in your own inner voice.

其实很大程度上还是你我小时候多半也有的那些特质:对学习的好奇心、坚持、相信自己内心的声音。

Developing and then believing in your own inner voice.

先养出自己内心的声音,然后相信它。

I have a four-year-old, I have a eight-year-old.

我有一个四岁的孩子,还有一个八岁的。

It's on me to help them develop their inner voice.

帮他们养出内心的声音,是我的责任。

And whether that's being opinionated and taking a stance to me and developing that, encouraging that, I think that those types of skill sets are things that is important in the future and having their own individual voice.

不管是有主见、敢跟我亮出立场,还是去培养这一点、鼓励这一点,这类本事在未来都很重要,还有他们各自独立的声音。

Lenny01:15:10

That is so interesting.

这太有意思了。

It's so related to the answer you had when I asked about how to avoid brain rot essentially and over-relying on AI, which is just keep focused on your own point of view and your own perspective before you over-rely on AI.

这跟我问你怎么避免脑腐、怎么避免过度依赖 AI 时你给的回答太有关系了——就是在过度依赖 AI 之前,先守住自己的观点、自己的视角。

And just this idea you're describing of building that in kids is really important.

而你说的这个,从小就在孩子身上把它建起来,真的很重要。

That is interesting.

这很有意思。

And I love how all this connects judgment, persistence, a point of view of your own.

我很喜欢这些东西串起来的样子:判断力、坚持、有自己的观点。

Dianne Penn01:15:34

Yeah.

嗯。

Lenny01:15:35

Both for kids and also adults.

孩子和大人都一样。

Dianne Penn01:15:37

[inaudible 01:15:37].

[听不清 01:15:37]。

Yeah.

嗯。

Anything-

有什么——

Lenny01:15:37

Anything-

有什么——

Dianne Penn01:15:37

You think about for your ...

你为你自己的……在想什么……

Lenny01:15:39

Oh, man.

哎哟。

Well, the question I'm thinking about is just when to get them on some AI-y thing.

我在想的问题就是,什么时候让他们开始接触 AI 这类东西。

I have a three-year-old, so it's pretty early for that.

我有个三岁的孩子,所以这事还挺早的。

But how do you onboard them to this crazy thing?

但你要怎么把他们领进这个疯狂的东西?

I was at an event recently and a bunch of parents were talking about how they think about AI and their kids.

我最近参加一个活动,一群家长在聊他们怎么看 AI 和自己的孩子。

And one person had a really interesting approach, which is keep them on the very early models so that they still have to struggle a bit and not get all the answers immediately.

有一个人的做法特别有意思:让孩子一直用很早期的模型,这样他们还是得挣扎一下,不会立刻拿到所有答案。

Thought that was interesting, like an open source local model, not Fable.

我觉得这挺有意思,比如给个开源的本地模型,而不是 Fable。

Dianne Penn01:16:11

I like it.

这个我喜欢。

Lenny01:16:12

Yeah.

嗯。

Dianne Penn01:16:12

Yeah.

嗯。

Lenny01:16:12

Yeah.

嗯。

And curiosity is something.

还有好奇心也算一样。

I keep mentioning Ben Mann, but his answer, actually, to this question has always stuck with me, which is curiosity and also just he's a big fan of Montessori, which is what we're encouraging for our kids.

我老是提到 Ben Mann,但他对这个问题的回答其实一直印在我脑子里,就是好奇心,另外他非常推崇蒙台梭利,这也是我们在给自己孩子推的。

So there's something there.

所以这里面是有东西的。

Maybe a last question, just along these lines, something Fiona Fung actually suggested ask you, who was recently on the podcast.

可能是最后一个问题,还是沿着这条线,是最近上过这档播客的 Fiona Fung 建议我问你的。

How do you stay just recharged and not burnout, being in the center of this crazy storm of AI, as a mom working in ...

身处 AI 这场疯狂风暴的中心,作为一个在……工作的妈妈,你怎么让自己一直充着电、不 burnout?

We're seeing the research work at Anthropic.

我们看到的是 Anthropic 的研究工作。

We're living through the most unprecedented time, working at ...

我们正活在最史无前例的时代,在……工作……

Just being on the outside of Anthropic, it's crazy.

光是站在 Anthropic 外面看,就已经很疯狂了。

I don't even know what it's like to be on the inside.

我都不知道在里面是什么感觉。

What have you learned about avoiding burnout, staying recharged, staying sane during the middle of all this?

在这一切当中,关于怎么避免 burnout、怎么保持充电、怎么保持清醒,你学到了什么?

Chapter 14

Entering the Hive Mind

不 burnout 的办法 · 以及还需不需要 PM
1:17:01 — 1:24:36 · 一年四个模型 vs 一个季度 · 休假不欠债 · PM 该更多不是更少
Dianne Penn01:17:01

In 2024, we shipped four models in the whole year or four series of models.

2024 年一整年,我们上线了四个模型,或者说四个系列的模型。

And I think we did more than that volume in just Q2 of this year.

而今年光是 Q2 一个季度,我们出的量就超过了那一整年。

I think I've been really lucky with the team that we've grown and built, both the stakeholders on the research side and within our research product management team.

我很幸运,有我们一路带起来、搭建起来的这支团队,既包括研究侧的相关方,也包括研究产品管理团队内部的人。

I think that one of the magical parts about approaching all of this is that it's not an individual sport.

面对这一切,神奇的地方之一在于,它不是一个人的运动。

There's a sense of radical ownership and team collaboration that, I think sometimes it does feel like a high performance sport because you're in very critical decisions.

这里有一种极致的主人翁意识,也有团队协作,有时候它确实像一项高水平竞技,因为你处在非常关键的决策里。

There's new information about users, about training, and you have to make recommendations and judgments and decisions very quickly.

关于用户、关于训练,不断有新信息进来,你得很快给出建议、判断和决定。

And nobody can do that sustainably by themselves.

没人能靠一己之力长期这样撑下去。

And so I think what's really helped is having a team that is incredible, who looks out for each other, who the night before a launch, even if they're not the core DRI on that model, will stay up and help the DRI to review the blog post and make edits and come up with better demos and knowing to be each other's extra hand.

所以真正帮到我们的,是有一支了不起的团队,他们彼此照应,在发布前一晚,哪怕自己不是那个模型的核心 DRI,也会熬着帮 DRI 审博客文章、做修改、想出更好的 demo,知道要给彼此当那多出来的一双手。

I think it's very easy, if you take all of this change on your own shoulders, to feel like you're alone and to feel like you have to do everything.

要是把这些变化全扛在自己肩上,很容易觉得自己是孤身一人,觉得什么都得自己来。

But I think one of the magical parts of Anthropic is this ability for us to figure out what are those opportunities to help each other and actually then taking the next smile of mind melding.

但 Anthropic 神奇的地方之一,是我们有本事找出哪些地方可以互相帮忙,然后真的再往前走一步,做到脑波同步。

We called it entering the hive mind.

我们把这个叫做进入蜂巢思维。

There was an article about this.

有篇文章写过这事。

And I think part of that is just that allows the team to replenish.

这里面有一部分作用,就是让团队能回血。

I was just on PTO in June.

我六月刚休过 PTO。

It's not just that you can take PTO and you come back to 3X the amount of things to do.

重点不只是你能休 PTO,回来时要做的事变成 3X。

It's actually that you can take PTO and know the team can figure out the right things to do and that we individually can watch out for each other.

而是你能休 PTO,同时知道团队能想清楚该做哪些正确的事,我们每个人也会互相照应。

So I think that's a big part.

所以这是很重要的一块。

I'm really lucky, just personally.

就我个人而言,我也真的很幸运。

Also, my partner is really supportive.

另外,我的伴侣非常支持我。

This is year six of me working in AI, so Amazon and then Anthropic.

今年是我做 AI 的第六年,先是 Amazon,然后是 Anthropic。

And so he sees how much I just love the technology and what this can do.

所以他看得到我有多喜欢这项技术、多喜欢它能做到的事。

And that really helps, I think, also from a personal perspective as well.

从个人层面上说,这一点也真的帮了很大的忙。

Lenny01:19:45

I love how many of these answers connect.

我很喜欢这些回答之间有这么多呼应。

So what I'm hearing here is just working with other people, relying on other people, helping each other out when things get crazy, which is a similar answer you had for just how to find the joy and fun in this work.

所以我听到的是,就是跟别人一起干活、依靠别人、事情忙疯的时候互相搭把手,这跟你讲怎么在这份工作里找到乐趣时给的答案很像。

Just be inspired by other people, see what they're doing, work together.

从别人身上得到激励,看看他们在做什么,一起干。

Dianne Penn01:20:06

Yeah.

嗯。

Lenny01:20:08

And it's interesting, when Fiona was on the podcast recently, I was asking her just what's changed in the world of software engineering?

有意思的是,Fiona 最近上播客的时候,我问她软件工程这个领域到底变了什么?

And she pointed out it's a lot lonelier now because now we're working with agents instead of other humans.

她指出,现在孤独多了,因为我们现在是在跟 agent 一起工作,而不是跟别的人。

Teams are smaller.

团队变小了。

People are having all these fleets they're talking to constantly.

大家手上都有成队的 agent,一直在跟它们说话。

And so this is just a reminder of just the power of just actual other humans around you.

所以这正好提醒我们,身边真实的人有多大力量。

Dianne Penn01:20:26

We're asked to work and make decisions on really big things because you have more scale from the technology.

我们被要求去处理、去决策的都是很大的事情,因为技术给了你更大的规模。

And I think having individuals, having other folks more who can have some level of mind meld with what you work on, how you approach, maybe not exactly every detail, but what are their first principles?

而有一些人、有更多人能在你做的事情、你的思路上跟你有某种程度的脑波同步——不一定每个细节都对得上,但他们的第一性原理是什么?

What are the assumptions you make?

你做的假设是什么?

Then helps them back up for you or push your decision and sharpen your thinking.

这样他们就能替你顶上,或者推你的决策一把,把你的思考磨得更锋利。

So I think I really try to look for that when building the team, growing the team, hiring.

所以在搭团队、扩团队、招人的时候,我真的会去找这个。

Is this person going to care about their own ego and building out a big org or are they going to care about contributing to Anthropic and contributing to the impact of the team and orienting towards folks who are low ego, team-oriented?

这个人在意的是自己的 ego、是把摊子铺大,还是在意对 Anthropic 的贡献、对团队影响力的贡献,然后往那些不端着、以团队为先的人身上靠。

I think it's a big part of the sustainability.

这是可持续性里很重要的一块。

Lenny01:21:31

Yeah.

嗯。

A lot of it always just comes down back to culture and hiring.

很多事情最后总是回到文化和招人上。

And I know I've heard a lot, just the reason Anthropic is able to move so fast.

我也听到过很多,就是 Anthropic 为什么能跑得这么快——

I remember that moment when something shipped every day of the month.

我记得有那么一阵,一个月里天天都有东西上线。

There's a calendar of launches and people were talking about how is this possible?

有一张发布日历,大家都在讨论这怎么可能做到?

And what I heard a lot is just because everyone is so aligned around the mission and the values, it allows people to make decisions really quickly.

我听到最多的说法是,因为每个人在使命和价值观上都高度对齐,这让大家能非常快地做决定。

Before we get to our very exciting lightning round, is there anything else, Dianne, that you wanted to share?

在进入我们非常刺激的闪电问答之前,Dianne,还有什么你想分享的吗?

Anything else you wanted to touch on?

还有什么想聊到的吗?

Anything you want to maybe double down on of things we've talked about?

我们聊过的东西里,有没有哪些你想再加重强调一下的?

Dianne Penn01:22:04

This was actually really fun because I feel like your questions actually sharpened some of my thinking around how the thoughts connect.

这次其实很有意思,因为你的问题真的让我把这些想法之间怎么连起来想得更清楚了。

Lenny01:22:10

I'm your real human Claude over here.

我就是你这边的真人版 Claude。

Dianne Penn01:22:13

One thing that I really want to convey or have people take away is I think one, in the ways of working, but also just two, this is a lot of growth and change and having the joy in using this technology.

我特别想传达、想让大家带走的是:一是工作方式上的,二是这里有大量的成长和变化,以及使用这项技术本身的乐趣。

And if you're feeling like in this moment, you don't have as much of that feeling of initial joy, how do you find people who do, if this is an area that you're excited and want to work on?

如果你此刻觉得自己没有当初那种快乐了,而这又是你兴奋、想投入的领域,那你怎么去找到还有这种感觉的人?

And I think developing skillsets, replenishing skillsets in many ways of things like thinking from a first principles' manner about what you solve.

还有就是培养技能、补充技能,比如用第一性原理去思考你要解决的是什么。

I think fundamentally, you didn't ask me this, but there is this question in the community of do we still need PMs when the models are so capable, when engineers are leaning in?

从根本上说,你没问我这个,但社区里确实有这么个问题:模型这么强、工程师又这么投入的时候,我们还需要 PM 吗?

I think the role of people who are user-centric, who go into the details of understanding what users are trying to accomplish, bubbling that up in an actionable manner and doing the relentless work to do that, that, to me, is a core of a product person.

以用户为中心,钻进细节去搞清楚用户到底想完成什么,再把这些以可落地的方式提上来,并且为此持续不懈地做下去——这样的角色,在我看来就是产品人的内核。

And I actually think we need more of that.

而且我确实认为,这样的人我们需要更多。

I think we are becoming very technology layered-driven.

我们现在越来越靠技术层驱动。

And actually to make that impactful, you have to go deep, you have to be curious, you have to be super hands-on.

而要真正做出影响力,你必须钻得深,必须有好奇心,必须极度亲自上手。

And those are things that are also traits that have, I think, helped Anthropic from a product development and model development perspective and as part of the culture.

这些也都是特质,它们在产品开发和模型开发上帮到了 Anthropic,也构成了文化的一部分。

And hopefully that's valuable for others as well.

希望这些对别人也有价值。

Lenny01:23:59

Amazing.

太棒了。

What an inspiring way to end it.

这个收尾太提气了。

Oh man.

天哪。

Yeah.

嗯。

And I've been saying this too for a long time.

这话我也说了很久了。

Just now that building is easy, the hard part becomes, as you said, what should we build and is the thing we have built correct and good and worth leaning into?

既然现在做东西变容易了,难的部分就变成,像你说的,我们该做什么,以及我们做出来的东西是不是对的、好的、值得投入的?

And to me, that's what PMs do and what PMs are good at.

在我看来,这正是 PM 在做的事,也是 PM 擅长的事。

Dianne Penn01:24:18

Yeah.

嗯。

Yeah.

对。

And it's getting into the details of the user.

而且要钻进用户的细节里。

Lenny01:24:22

Yeah.

嗯。

Empathy.

同理心。

Okay, great.

好,太好了。

PMs are going to make it.

PM 会活下来的。

Okay.

好。

PRD is not dead.

PRD 没死。

All kinds of important lessons here.

这里面全是重要的经验。

Dianne, with that, we've reached a very exciting lightning round.

Dianne,那么我们就来到了非常刺激的闪电问答环节。

I've got five questions for you.

我有五个问题要问你。

Are you ready?

准备好了吗?

Chapter 15

Lightning Round

闪电问答 · 祖父那句「上面永远还有一层」
1:24:36 — 1:33:50 · How to Raise an Adult · 《辐射》· 华尔街交易台上的唯一一个
Dianne Penn01:24:36

Yep.

对。

Lenny01:24:37

First question, what are two or three books that you find yourself recommending most to other people?

第一个问题,你最常推荐给别人的是哪两三本书?

Dianne Penn01:24:43

One personal one, I really like How to Raise an Adult.

先说一本偏个人的,我很喜欢 How to Raise an Adult(《如何养出一个大人》)。

So I'm a mom.

我是个妈妈。

I think a lot about what is the things that I want to instill in my kids.

我常想,我希望在孩子身上种下些什么。

And that book is really helpful for describing, we're not trying to raise children, we're trying to raise adults.

那本书特别有帮助的一点,是它把这件事讲清楚了:我们要养的不是孩子,是大人。

So just the framing of what does that mean?

光是这个说法本身——这到底意味着什么?

And what does it mean, what are the characteristics that we want to hone and harness and foster in our kids?

它意味着什么,我们想在孩子身上打磨、调动、培养的又是哪些特质?

The other book that I was listening to on Audible recently is Incorrigible by Eric Ries.

另一本我最近在 Audible 上听的,是 Eric Ries 的 Incorrigible。

So the author-

作者他——

Lenny01:24:44

Incorruptible, I think is-

是 Incorruptible 吧,我记得——

Dianne Penn01:25:22

Incorruptible.

Incorruptible。

Yes.

对,没错。

Lenny01:25:24

Yeah.

嗯。

His recent [inaudible 01:25:26].

他最近那本[听不清 01:25:26]。

Dianne Penn01:25:27

Yeah.

对。

And I think the question of how to build great companies is important.

怎么造出伟大的公司,这个问题很重要。

I've personally just been most fascinated with how to keep great teams and great companies going further.

而我个人最着迷的,是怎么让优秀的团队、优秀的公司走得更远。

And it was very interesting to just see his framing and reframing of the question.

看他怎么给这个问题搭框架、又怎么重新搭一遍,很有意思。

I loved some of the examples around having metrics around culture.

我很喜欢里面关于给文化配指标的那些例子。

If you only measure revenue and then that's how you're going against.

如果你只考核收入,那你拿来对标的就只有收入。

But if you have other better metrics, that's actually the way to sustain the values you care about.

但如果你还有别的、更好的指标,那才是守住你在乎的那些价值观的办法。

I've been trying to think about how to actually bring that to the team level of how do we better articulate our norms, a lot of the things we talked about on the team.

我一直在想怎么把这套东西落到团队层面:我们怎么才能把自己的规范讲得更清楚,也就是我们在团队里聊过的很多事。

So I think that's also a really good read.

所以这本也很值得一读。

Lenny01:26:16

There you go.

这就对了。

That'll be your next watch, everyone, as you're listening to this, the Eric Ries episode.

各位听到这儿,下一期就看这个吧——Eric Ries 那一期。

Such a good episode.

那期特别好。

Yeah.

对。

And his book just came out, Incorruptible.

他的书也刚出,就是 Incorruptible。

Dianne Penn01:26:25

[inaudible 01:26:25].

[听不清 01:26:25]。

Lenny01:26:25

I think it was like a New York Times bestseller.

好像还上了 New York Times 畅销榜。

It's actually doing incredibly well, which I was really happy to see.

卖得特别好,我看到挺替他高兴的。

Dianne Penn01:26:31

Yeah, exactly.

对,没错。

Lenny01:26:32

Next question.

下一个问题。

Favorite recent movie or TV show you've really enjoyed.

最近最喜欢的电影或者剧集,真正看得过瘾的那种。

Most people at Anthropic do not have time to watch things, but I'm curious if you have an answer.

Anthropic 大多数人都没时间看这些,不过我很好奇你有没有答案。

Dianne Penn01:26:40

I would say during some time off last month, I did get to binge-watch Fallout on Amazon Prime.

上个月休假的时候,我确实把 Amazon Prime 上的《辐射》一口气刷完了。

So that was actually ...

那个其实……

Have you heard of it?

你听说过吗?

Lenny01:26:56

Yeah.

听过。

Yeah.

嗯。

It's based on the video game.

改编自那个电子游戏。

Dianne Penn01:26:58

Yes, it's based on the video game.

对,改编自那个电子游戏。

I think it's witty, it's humorous.

它很机灵,也幽默。

It's also super action-oriented, so highly recommend.

动作戏还特别足,强烈推荐。

Lenny01:27:08

Okay, next question.

好,下一个问题。

Do you have a favorite product you've recently discovered that you really love?

你最近有没有发现什么特别喜欢的产品?

Dianne Penn01:27:12

I really do think Claude Tag is very interesting in terms of a product experience.

从产品体验的角度,我真的觉得 Claude Tag 很有意思。

We actually have different versions of this within Anthropic, and I think it's actually been a really, really, really powerful tool.

我们在 Anthropic 内部有它的好几个版本,它确实是个非常非常强大的工具。

Lenny01:27:29

Yeah.

对。

I think some people are like, "What's the big deal?"

有些人会说:“这有什么大不了的?”

The fact that everyone at Anthropic is raving about it tells me something important is going on here.

但 Anthropic 所有人都对它赞不绝口,这说明里面确实有重要的东西。

And I'm trying to actually get it working within my Slack community that I have for paid newsletter subscribers.

我正在想办法把它接进我给付费 newsletter 订阅者开的那个 Slack 社群。

How cool would that be?

那得多酷啊?

Dianne Penn01:27:43

Yeah.

对。

Lenny01:27:43

Yeah.

对。

I'm trying to figure out how it works when it's not a company, when it's just a bunch of people that don't know each other and how that might work, but we're trying it out.

我在琢磨,如果对面不是一家公司,而是一群互不相识的人,它会怎么运转、还行不行得通,总之我们在试。

Okay.

好。

Two more questions.

还有两个问题。

Your favorite life motto that you find yourself often coming back to in work or in life.

你最喜欢的人生格言,工作或生活里你常回头想起的那句。

Dianne Penn01:27:57

So I was actually raised by my grandparents for the first 10 years of my life.

我人生的头 10 年其实是祖父母带大的。

And my parents were immigrant college and master's students in the US.

我父母那会儿在美国读本科和硕士,是移民学生。

Lenny01:28:08

Oh, wow.

哇哦。

Dianne Penn01:28:08

And my grandfather always says, no matter how far you go, there's always another level, which is, I think, a really good way, though a pretty intense way of describing his life philosophy.

我祖父总说,不管你走多远,上面永远还有一层——这句话很好地概括了他的人生哲学,尽管也挺狠。

But I go back to that whenever there's something new or unprecedented that we experience.

每次遇到什么新的、前所未有的事,我都会回到这句话。

And I think first half of this year, there was definitely a lot of that.

今年上半年,这样的事确实不少。

There was a lot of new things that we were learning, I was learning.

有很多新东西是我们在学、是我在学。

So just feeling like there's always another mountain, another opportunity to come.

那种感觉就是,前面永远还有下一座山、下一个机会。

Lenny01:28:50

Not good enough, Dianne.

还不够好,Dianne。

We need to go better.

我们得做得更好。

We need to go bigger.

我们得做得更大。

Makes me think about, actually, another Ben Mann line from his podcast episode, that this is the most normal it's ever going to be.

这让我想起 Ben Mann 在他那期播客里的另一句话:现在是往后最正常的时候了。

It's only going to get weirder and crazier.

只会越来越怪、越来越疯。

Dianne Penn01:29:04

Yeah.

对。

Lenny01:29:05

Oh my God.

天哪。

Okay.

好。

Final question.

最后一个问题。

I was poking at your LinkedIn.

我翻了翻你的 LinkedIn。

You were a high yield bond trader at JP Morgan Chase early in your career.

你职业生涯早期在 JP Morgan Chase 做高收益债交易员。

You have this redacted $100 million trading portfolio of some kind.

上面还写着一个被打码的 $100 million 交易组合什么的。

What did you learn from that time in your life that has stuck with you?

人生的那一段,你学到了什么一直留到今天的东西?

And/or is there a crazy story from that period?

或者,那段时间有没有什么疯狂的故事?

It was four years of your life.

那可是你人生里的四年。

Dianne Penn01:29:31

I think I learned, actually, a lot that I applied here at Anthropic and other jobs thereafter.

其实我学到很多,后来在 Anthropic、在别的工作里都用上了。

So when I was at JP Morgan, the trading floor, you could envision Wolf of Wall Street.

我在 JP Morgan 的时候——说到交易大厅,你脑子里浮现的可能是《华尔街之狼》。

That's very different.

实际差别很大。

Most traders, I think, are in front of a terminal.

大多数交易员就坐在一台终端前面。

They're much more doing analyses on their computers, but it's still very, I would say, male-dominated.

他们更多是在电脑上做分析,但可以说,那里仍然非常男性主导。

And so I was the only woman.

而我是那里唯一的女性。

I was the only person with my background on the trading desk.

整个交易台上,只有我是那种背景。

And I learned that there was a very good environment to building one, my sense of authentic self.

我学到的是,那其实是个很好的环境:一来,能建立起我真实的自我。

And two, that even if I was the most junior person, even if I may look different, that the best ideas and having conviction in the best ideas, irregardless of all of those other factors, is the most important thing.

二来,哪怕我是最资浅的那个,哪怕我看上去和别人都不一样,最重要的仍然是拿出最好的想法,并且对这个想法有信念——不管其他那些因素。

And so I think just bring that sense of how I show up more at work.

所以我把那股劲儿带进了我在工作里的样子。

I'm pretty vulnerable and authentic with my team.

我在团队面前挺坦露、也挺真实。

I try to really make sure that, regardless of people's levels or tenures, if they have a great idea, how to help them pursue that and to do also the same.

我很努力地做到:不管级别、不管资历,只要有好想法,我就帮他把它推下去,也希望他自己这么做。

So to put the idea out there, to actually have conviction in it, to do the follow through, to do the nitty-gritty work to make it happen.

把想法摆出来,真的对它有信念,一路跟到底,把那些琐碎的具体活儿干掉,让它成真。

So those are all things that I learned from trading.

这些都是我从做交易里学到的。

And yeah, I think applies to any job in many ways.

而且在很多方面,它们放到任何一份工作上都成立。

Lenny01:31:20

That is beautiful.

这话真好。

Where can people find you online if they want to follow you?

如果大家想关注你,在网上哪里能找到你?

And how can listeners be useful to you?

听众又能怎么帮到你?

Dianne Penn01:31:28

I don't have a large presence on social.

我在社交媒体上没什么存在感。

I think the best way to find my work, my team's work is really the Anthropic blog and when we're publishing new models, new product experiences.

想看到我和团队的工作,最好的地方就是 Anthropic 的博客,还有我们发布新模型、新产品体验的时候。

I think in terms of useful for me, I think the best thing, number one, is your feedback.

至于怎么帮到我,第一,最好的就是你们的反馈。

We actually, if you thumbs up or thumbs down on any of our product surfaces, if you contact your salesperson with feedback about the model, it will make its way to me.

你在我们任何一个产品形态上点赞或者点踩,或者把对模型的反馈告诉你的销售,这些都会传到我这里。

We actually, with every research model, I actually get pretty close into understanding favorability and feedback.

每一版研究模型,我都会贴得很近地去看好评度和反馈。

So giving us that feedback, pushing Claude, telling us where it's falling down, those help us make Claude better.

所以给我们反馈、去逼 Claude、告诉我们它在哪儿掉链子,这些都能帮我们把 Claude 做得更好。

The other thing is if you have folks in your network who seem like this type of profiled person that I just talked about, I'm hiring, the team is growing.

另一件事是,如果你身边有符合我刚才说的那类画像的人,我在招人,团队在扩。

We really will love just people who love this technology, who are deeply curious, first principles thinkers, who are fearless in questioning assumptions and who have a tinkering hackery spirit.

我们特别想要那种热爱这项技术、极度好奇、用第一性原理思考、敢于质疑假设、还带着爱折腾的 hacker 劲儿的人。

Lenny01:32:46

Wow, what a dream job.

哇,这工作也太理想了。

So basically open PM roles at Anthropic on the research team.

所以基本上就是 Anthropic 研究团队开放的 PM 岗位。

Dianne Penn01:32:52

Yes.

对。

Lenny01:32:53

And they apply, I assume, on the website, the careers page.

我猜是在官网的招聘页面上申请。

Dianne Penn01:32:56

Yes.

对。

Lenny01:32:56

Holy moly.

我的天。

All right, here we go.

好了,来吧。

Enjoy the flood of resumes you're about to receive.

好好享受马上要涌来的简历吧。

Dianne Penn01:33:01

Thank you, Lenny.

谢谢你,Lenny。

Lenny01:33:04

Dianne, thank you so much for being here.

Dianne,非常感谢你来。

Dianne Penn01:33:06

Thank you so much for having me.

非常感谢你邀请我。

Thank you for really helpful, thought-provoking questions, helping me even connect the dots on how we work, how this whole technology is coming together and being proud of people in it.

谢谢你这些很有帮助、很引人思考的问题,它们甚至帮我把一些事串了起来:我们是怎么工作的、这整套技术是怎么拼到一起的,以及为身在其中的人感到骄傲。

Lenny01:33:21

I really appreciate that.

我真的很感激。

But thank you, Dianne, for real.

但真的,谢谢你,Dianne。

Okay.

好。

Well, bye everyone.

好了,各位再见。

Thank you so much for listening.

非常感谢你的收听。

If you found this valuable, you can subscribe to the show on Apple Podcasts, Spotify, or your favorite podcast app.

如果你觉得这期有价值,可以在 Apple Podcasts、Spotify 或者你常用的播客应用上订阅本节目。

Also, please consider giving us a rating or leaving a review as that really helps other listeners find the podcast.

也请考虑给我们打个分或者留个评价,这真的能帮别的听众找到这个播客。

You can find all past episodes or learn more about the show at lennyspodcast.com.

所有往期节目和关于节目的更多信息,都可以在 lennyspodcast.com 找到。

See you in the next episode.

下期见。