Lenny's Podcast · Episode · 双语整理

From Six Months to One Day

六个月被压成一天 · Anthropic 怎么造产品

"It is very hard to be the right amount of AGI pilled. It's very easy to build the product for the super AGI strong model. The hard thing is figuring out for the current model how do you elicit the maximum capability." 要把 AGI 之药吃得"刚刚好",最难。给一个未来超级模型写产品很容易;给现在的模型把用户带到 golden path,这一关才是真功夫。

嘉宾 · Cat Wu(Head of Product, Claude Code & co-work) 主持 · Lenny Rachitsky 时长 · 1h 25m 发布 · 约 2026 年初至 5 月初(以 SRT 文件 2026-05-03 mtime 推断,无原始发布日期)
TL;DR · 速读

Anthropic PM 们最近想清楚了什么

14 个来自 Cat Wu 的判断
  1. AGI 之药要"刚刚好"才最难

    "It is very hard to be the right amount of AGI pilled. It's very easy to build the product for the super AGI strong model."

    "The hard thing is figuring out for the current model, how do you elicit the maximum capability? How do you help users go get onto the golden path?"

    照着未来超模型 spec 写产品很容易;难的是把现在的模型用到极限,把用户引到能跑通的那条路上。

  2. 六个月被压成一天

    "The timelines for a lot of our product features have gone down from 6 months to one month and sometimes to one week or even one day."

    "There should be less emphasis on aligning your multi-quarter roadmaps with your partner teams and more emphasis on okay, how can we figure out the fastest way to get something out the door?"

    一旦交付节奏从季度变成周,跨季度 roadmap 和 partner alignment 这些"老 PM 工艺"全部变成阻力,而不是助力。

  3. Research preview 是策略,不是免责声明

    "We actually ship almost all of our features in research preview. What this does is it reduces our commitment for shipping something."

    "We can just get something out in a week or two. We clearly brand this when we ship something so that users know that this is an early product."

    把"我们还在试"内嵌到 brand,公司提交压力一下变小,任何工程师都敢一周内把东西扔出门收反馈。

  4. 不是 Mythos 让他们快,是流程让他们快

    "It's not fully mythos. We do use the models internally and I think this has increased our rate of shipping a little bit, but I don't think it explains the bulk of the increase."

    "A lot of it is the process and the expectation on the team. So we're very low on process. We want to remove every single barrier to shipping things."

    别盯着 Anthropic 在用什么秘密模型 —— 真正起作用的是"任何成员都能从想法到上线 < 1 周"这条铁律。

  5. 所有角色都在融合

    "All of the roles are merging. PMs are doing some engineering work, engineers are doing PM work, designers are PMing and also landing code."

    "On our team we're pretty focused on hiring engineers with great product taste. There are many engineers on our team who are fully able to end-to-end go from see user feedback on Twitter through to ship a product at the end of the week with almost no product involvement."

    你叫自己 PM、工程师还是设计师不重要 —— 能不能 end-to-end 把一个想法落到用户手里才是要紧事。

  6. Taste 是真正稀缺的东西

    "It comes back to product taste. As code becomes much cheaper to write, the thing that becomes more valuable is deciding what to write."

    "We get tens of thousands of GitHub issues asking for every single thing under the sun and it takes a lot of care and taste to figure out which of these is worth building and what is the right way to build it."

    写代码变便宜以后,"决定写什么"才是值钱的活。Anthropic 团队基本只招两类人:有 taste 的工程师,或干过工程的 PM 和设计师。

  7. 常识和 EQ 是人类还守得住的护城河

    "Humans still provide a level of common sense that the models don't."

    "The model doesn't always have a great sense of who all the stakeholders are, how they relate to each other, what their preferences are, what are the right venues to communicate with them to keep them on board."

    一千个发布动作里,谁是 stakeholder、什么 venue 沟通最合适这种隐性知识,模型暂时还接不下来。

  8. 在龙卷风眼里要会笑

    "We try to face every challenge with a smile because there's always so much going on."

    "Maybe Sunday night there's some P0 and then by Monday there's a P0 and by Monday afternoon there's a P0000 and you're like wow, I can't believe I was so worried about that P0 from Sunday."

    不会笑的人最先 burn out。Anthropic 招人先过这一关:能不能不被 P0 节奏拖垮。

  9. 使命大于产品 KR

    "If Claude Code failed but Anthropic succeeded I would be extremely happy."

    "Mission means that teams are willing to make sacrifices that hurt their own goals and their own KRs in service of Anthropic's goals. People are very happy to make those trade-offs."

    不是嘴上说说 —— team 真的会为整体使命牺牲自己产品的优先级。OpenClaw 那次取舍就是这套逻辑的延续。

  10. PM 最难的活:画出一个月后的产品

    "The hardest skill is being able to define what the product should look like a month from now."

    "There are patterns that the best PMs can see based on how users are abusing the limits of the existing product, and the best PMs can sense that, set a direction, and steadily execute towards it."

    用户怎么"虐"现在的产品,这里头藏着下一个版本应该长成什么样。看不到这个,就只能跟着模型节奏被动反应。

  11. 写 10 个好 eval 比写 100 个垃圾 eval 重要

    "You don't need to build hundreds of evals for them to be useful. Just building 10 great evals is important."

    "Eval is this underappreciated thing that more PMs, more engineers should be working on. It's important for helping the team quantify what the goal is."

    eval 是被严重低估的工具。它本质是把"成功长什么样"写成可执行的定义,让全队朝同一个方向推。

  12. 模型变强 → harness 该减,不该加

    "We can remove a lot of prompting interventions every time the model gets smarter. We actually do this every time we launch a model."

    "We read through the entire system prompt and reflect on, okay, for each of these sections, does the model really need this reminder anymore? And if not, we'll remove it."

    "the model will eat your harness for breakfast" —— 像 to-do list 这种为旧模型加的拐杖,新模型来了就该考虑剥掉。

  13. 真问题 > 玩具 setup

    "I see a lot of people customizing their tool, adding a ton of skills and MCPs… that can even distract from your core goal of launching some product."

    "I would push people towards building apps that you're actually using every single day, because only through that usage are you actually getting the value."

    Twitter 上的"look at my setup"是反面教材;每天真用、解决了你具体问题的小 app 才是杠杆。

  14. Just do things — 工作是假的

    "Jobs are fake. If you understand the constraints, you can figure out what you can do and then just like try to do it quickly."

    "In a lot of companies, roles are very strictly defined… 'just do things' lets people feel empowered to make these decisions, empowered to operate across team boundaries just to get something done."

    Cat 在 Stripe 早期 20 人时拿到的 mental model:角色定义只是抽象,最后只看你有没有把活干掉。

Chapter 01

The AGI-pilled balance

AGI-pilled · 平衡当下与未来
00:00 — 02:15 · 冷开场 · Cat 与 Boris 的分工 · "80% mind-meld"
下面是节目开场前的"冷开场"片段(Cold open) —— 主持人 Lenny 把全集中最有冲击力的几句剪到了开头作钩子。
Cat Wu00:00:00

I think it is very hard to be the right amount of AGI pilled.

我觉得 AGI 之药要吃得"刚刚好",非常难。

It's very easy to build the product for the super AGI strong model.

照着那个未来的、超强 AGI 模型去做产品,反而是容易的。

The hard thing is figuring out for the current model, how do you elicit the maximum capability?

真正难的是面对今天这个模型 —— 你怎么把它现有的能力榨到最大?

Lenny00:00:13

I've never seen anything like the pace you folks at Anthropic are shipping at.

我从来没见过像你们 Anthropic 这种发布节奏。

Cat00:00:17

We want to remove every single barrier to shipping things.

我们想把"上线一件东西"路上的每一个障碍都拆掉。

The timelines for a lot of our product features have gone down from 6 months to 1 month and sometimes to even one day.

很多产品功能的交付周期已经从六个月压到了一个月,有时候甚至压到一天。

You're interviewing hundreds of PMs and you just keep feeling like they're approaching it very incorrectly.

你面试过几百个 PM,会一直感觉:很多人对这件事的切入方式根本就不对。

The PM role is changing a lot. It's changing really quickly.

PM 这个角色变化非常大,而且变得非常快。

The thing that is extremely important for building AI native products is iterating so quickly — figuring out a way for you to actually launch features every single week.

做 AI 原生产品最关键的一件事,就是迭代快到能让你每一周都真的把新功能扔上线。

Lenny00:00:44

What do you think are the emerging skills PMs need to develop?

你觉得现在的 PM 应该开始培养哪些新技能?

Cat00:00:48

It comes back to product taste.

归根到底,还是要看 product taste。

As code becomes much cheaper to write, the thing that becomes more valuable is deciding what to write.

代码本身越来越便宜,反而是"决定要写什么"这件事越来越值钱。

冷开场结束 · Lenny 切回 intro 自白 · 介绍嘉宾。
Lenny(intro)00:01:00

Today my guest is Cat Wu, head of product for Claude Code and co-work at Anthropic.

今天的嘉宾是 Cat Wu,Anthropic 旗下 Claude Code 与 co-work 的产品负责人。

Cat is at the center of everything that is changing in AI and product and building.

在 AI、产品、工程这三件事正发生的所有变化里,Cat 都站在最中间。

She and her team are building the product that is most changing the way that we all build our products.

她和团队做的这个产品,是当下最深刻地改变着"我们其他人怎么做产品"的那一个。

She is so full of insights and wisdom and lessons. This is an episode you cannot miss.

这一期满是真知灼见,你绝对不该错过。

With that, I bring you Cat Wu. Cat, welcome to the podcast.

事不宜迟,Cat,欢迎来到节目。

Cat00:01:35

Thanks for having me.

谢谢邀请。

Lenny00:01:37

I have so many questions. I'm so excited to have you on this podcast.

我有一肚子问题,非常激动你能来。

I want to start with giving people an understanding of your role alongside Boris.

咱们先从你和 Boris 怎么搭班子说起,让大家有个概念。

Everybody knows Boris — his episode is the number one most popular episode on this podcast. No pressure.

大家都认识 Boris —— 他那一集到现在是本播客订阅量最高的一集。希望不会给你压力。

He created Claude Code. He leads the team, ships a bazillion PRs a day from his phone. I don't even know what the number is anymore.

他是 Claude Code 的创造者,带着整个团队,每天能从手机上 ship 出无数个 PR,我已经不知道具体多少了。

I think people don't give you enough credit for the success that Claude Code has had and co-work and all the things you all are building.

我觉得 Claude Code、co-work 还有你们所有这些东西的成功,你应得的功劳被严重低估了。

Help us understand your role on the team, how you work with Boris, how you split responsibilities — what does the PM role look like on the Claude Code team?

能不能讲讲你在团队里的角色?你和 Boris 怎么协作、怎么分工?在 Claude Code 团队里,PM 这个角色实际是怎么运转的?

"It is very hard to be the right amount of AGI pilled."

AGI 之药要吃得"刚刚好",反倒是最难的事 —— 这是 Cat 整集开场抛下的那把钥匙。

Chapter 02

Six months to one day

六个月 → 一天 · PM 周期被压成一周
02:16 — 09:00 · 不再做季度 roadmap · Research Preview · 工程 / 营销 / 文档同周协同
Cat00:02:16

I feel very lucky to work with Boris. He's been an amazing thought partner.

能和 Boris 合作我觉得特别幸运,他是非常棒的 thought partner。

He's our tech lead. He's very much the product visionary.

他是我们的 tech lead,也是这个产品最主要的"看远景"的人。

He is great at setting like — this is what the product needs to be in three months, six months from now. This is what the AGI-pilled version of the product is.

他特别擅长锚定那种远景:三个月、六个月之后,这个产品应该长什么样?那个"AGI-pilled 版本"的产品是怎样的?

A lot of my role is figuring out — okay, what is the path from where we are today to that vision 3 to 6 months from now?

我的角色很大一部分,就是搞清楚:从今天的状态走到那个三到六个月后的愿景,中间这条路怎么走。

I spend more of my time on the cross-functional — making sure that our marketing team, sales team, finance, capacity, etc. are bought in on the plan and that we're all rowing the same direction.

我大部分时间花在跨职能协同上 —— 让 marketing、sales、finance、capacity 这些团队都买账这个计划,大家朝一个方向划船。

And once the feature is ready, that there aren't any blockers to shipping it.

还要保证一个功能做好了之后,上线路上没有任何卡点。

It works well because we kind of mind-meld, but it is actually remarkably blurry of a line.

能跑通是因为我们俩在很多事上"思维同步";但这条分工线其实非常模糊。

I think we're like 80% mind-meld, then there's this 20% of things that maybe I care a lot more about — so I'll drive those — and then 20% where he cares a lot more than me, and he just drives those.

大概 80% 的事我们想到一块去;剩下 20% 是我比较在意的,我就推;还有 20% 是他比较在意的,他就推。

[Sponsor break · WorkOS] —— 中段 sponsor 读稿,介绍 OpenAI / Anthropic / Cursor / Vercel / Replit 等都在用 WorkOS 把 SSO / SCIM / RBAC 等"企业准入"功能变成现成 API。略过不译。
Lenny00:04:30

Something that you shared before we started recording is the fact that you're interviewing hundreds of PMs all the time.

录之前你跟我说,你一直在面试几百个 PM。

If I had a nickel every time someone asked me for an intro to someone at Anthropic to go work as a PM, I'd have 30 billion in ARR.

每次有人找我求"介绍我去 Anthropic 当 PM"的人脉,如果我能收一毛钱,我大概能攒出 300 亿美元 ARR。

It's just like the number one place people want to go work at. So I can only imagine how many PMs you're interviewing.

这是当下大家最想去工作的地方,你面试的数量我都不敢想。

You told me that you're just seeing people doing it wrong — the way they're approaching what they think it takes to be a successful AI PM.

你说你看到很多人对"怎么做一个成功的 AI PM"这件事的切入方式都是错的。

Talk about what you're seeing, and what people need to understand about what it takes to be successful these days.

讲讲你具体看到了什么,大家需要明白哪些点,才能在今天做出成绩?

Cat00:05:03

Before AI, technology shifts were a lot slower.

AI 之前,技术变迁要慢得多。

You could plan on the 6 to 12 month time horizons.

那时候你可以按六个月到一年的尺度去规划。

Because you were shipping features at a slower rate, there was a lot more emphasis on coordinating with all the other partner teams to make sure that they're shipping features that unblock yours — because code at that time was very expensive to make.

因为发布节奏慢,很多力气都花在和兄弟团队对齐,让他们的功能去解你的卡点 —— 那时候写代码是很贵的事。

Now with AI, and with how much that has accelerated engineering, and with how quickly the model capabilities are improving, the timelines for a lot of our product features have gone down from 6 months to one month, and sometimes to one week or even one day.

现在有了 AI,工程效率被大幅拉高,加上模型能力还在飞速提升,我们很多产品功能的周期从六个月降到一个月,有时候降到一周,甚至一天。

As a PM, there should be less emphasis on aligning your multi-quarter roadmaps with your partner teams, and more emphasis on — okay, how can we figure out the fastest way to get something out the door?

所以作为 PM,你不应该再把太多精力花在和兄弟团队对齐多季度 roadmap 上,而是应该想:怎么用最快的方式把东西推出门?

How can we make like a "concept corner" of our product suite where an engineer or a PM has an idea and by the end of the week we are able to get it into our users' hands?

能不能在产品矩阵里专门划出一个"概念角",让一个工程师或 PM 有了想法,周末之前就能让用户摸到?

The PMs who do the best on AI native products are the ones who can shorten the time from having this idea to actually getting the product in the hands of users — and help define what are the most important tasks that need to work out of the box.

在 AI 原生产品上做得最好的 PM,都能把"想到"到"用户用上"这段路压到最短;同时还能定义出:产品开箱必须做对的最重要的几个任务是什么。

Lenny00:06:30

What I love about this is — people haven't grasped how fast they need to move, and how much of the job now is helping the team move fast.

我特别喜欢这个观点 —— 大家还没意识到自己需要多快,也没意识到 PM 工作现在很大一块就是帮团队加速。

What does your PM team do to help them move this fast — other than have access to the most advanced models?

你们 PM 团队具体怎么让大家这么快?除了能用上最前沿的模型之外。

Cat00:07:00

The first thing is to set clear goals — because LLMs are so general that it actually creates a lot of ambiguity in who we're building for, what problems we're trying to solve, what the top use cases are.

第一件事是把目标定得很清楚 —— LLM 太"通用"了,反而让"为谁做、解决什么、最关键的 use case 是哪几个"全都变模糊。

A great PM is able to say — okay, our key user is professional developers. The main problem we want to solve is maybe there are too many permission prompts and people are feeling fatigue. The use case is — we want professional developers at enterprises to safely get to zero permission prompts.

一个好的 PM 能直接说:我们的核心用户是专业开发者;要解决的问题是 permission prompt 太多让人疲惫;use case 就是 —— 让企业里的专业开发者在安全前提下做到零授权弹窗。

That actually sets a pretty clear goal because it rules out a lot of potential approaches for reducing permission prompts.

这样目标就很硬,因为它一下子排除了一大堆减少弹窗的备选方案。

The second thing is figuring out some repeatable process for getting these features shipped.

第二件事是搭一套可复用的发布流程。

For Claude Code, what we do is we actually ship almost all of our features in research preview.

在 Claude Code,我们几乎所有功能都是以 research preview 的形式上线的。

We clearly brand this when we ship something so that users know — this is an early product, this is just an idea, this is something we're trying to get feedback on, and that this might not be supported forever.

发布时我们会明确给它打上"research preview"的标签,让用户清楚知道:这是早期产品,只是个想法,我们在拿反馈,可能不会永远支持下去。

What this does is it reduces our commitment for shipping something. We can just get something out in a week or two.

这样做最大的作用是——把"上线"这件事的承诺压力降下来,我们一两周就能把东西扔出门。

The third thing is help create the framework for the team so that they know when to pull in cross-functional partners, and what those partners' expectations are.

第三件事是帮团队搭一个框架,让他们知道什么时候该把跨职能伙伴拉进来、对方期待什么。

We have a really tight process between engineering, marketing, and docs.

我们工程、营销、文档之间有一套非常紧的流程。

When engineers feel a feature is ready and we've dogfooded internally, they post it in our evergreen launch room.

工程师觉得一个功能可以发,而且内部已经 dogfood 过了,他们就会把它发到我们那个 evergreen launch room(常驻发布频道)里。

Then Sarah, who leads our docs, and Alex, who leads PMM, and Tara and Lydia on DevRel just jump in and can turn around the marketing announcement the very next day.

然后文档负责人 Sarah、PMM 负责人 Alex,加上 DevRel 的 Tara 和 Lydia 就会立刻接力,第二天就能把营销公告做出来。

Because we have this really tight process, it lowers the friction for any engineer to ship something. And PM is the role that should be setting this up.

因为这个流程很紧,任何工程师 ship 东西的摩擦力都很低。而搭这套流程,本来就是 PM 的职责。

"We can just get something out in a week or two."

把"我们还在试"这件事内嵌成 brand —— Research Preview 不是免责声明,是 Anthropic 把"提交压力"压低的核心机制。

Lenny00:08:59

How do PRDs fit into this?

那 PRD 在这个流程里还有位置吗?

You said that goals are a really important part — being aligned on what success looks like, who is this for, who's this not for. Are you writing PRDs? Is it just a couple bullet points? How has that evolved in the world of an AI PM?

你刚刚说目标很关键 —— 成功长啥样、给谁用、不给谁用。那你们写 PRD 吗?还是就几条 bullet?在 AI PM 的世界里这件事是怎么变的?

Chapter 03

Goals, metrics, principles — not heavy PRDs

目标 · 指标 · 原则 · 不是厚 PRD
09:12 — 10:30 · 每周 metrics readout · 团队原则文 · 仅在歧义大或基础设施类项目写 PRD
Cat00:09:12

There are two things that we do.

我们做的其实是两件事。

One is we have very rigorous metrics and we do metrics readouts with the entire team every week.

第一,我们有一套非常严的指标体系,每周都会和全员一起做 metrics readout。

The goal of this is to make sure that everyone deeply understands all the facets of our business — what our key goals are, how they're trending, and what drives them.

这件事的目的是让每个人都吃透业务的每一面 —— 关键目标是什么、走势如何、由什么驱动。

The second thing is we have this list of team principles. This includes who our key users are, why those are our key users.

第二,我们有一份"团队原则"清单,里面写清了核心用户是谁、为什么是这群人。

The reason that we articulate all of this is so that everybody on the team feels like they understand how our business works, what's important to us, and what we're willing to trade off.

把这些都说清楚,是为了让每个成员都能理解业务的运转方式、什么对我们是重要的、我们愿意为什么牺牲什么。

It lets people make decisions by themselves without feeling like they're blocked on PM or any other stakeholder.

这样大家就能自己做决定,不会觉得"卡在 PM 那里"或者卡在某个 stakeholder。

Lenny00:09:55

I love how so much of this is like, okay — we still need PMs in the future. There's so much talk of "why do we need PMs? We're just going to ship and build. We need engineers."

我特别喜欢一点 —— 你说的这些其实在告诉我们:未来仍然需要 PM。现在到处都在说"我们还要 PM 干嘛?直接造东西就完事了,只要工程师就够了。"

Cat00:10:03

Oh, we actually do PRDs sometimes.

对了,其实我们有时候还是会写 PRD 的。

For features that are particularly ambiguous, it does help to write out a one-pager on what the goals are, what the delightful use cases are, what the failure modes currently are that we need to fix.

对那些特别模糊的功能,写一页纸把目标、最让人惊喜的 use case、当前需要修的失败模式都列出来,确实有用。

There are occasionally some projects, especially things that require heavy infrastructure, that do take many months. And for those situations, we do write PRDs still.

偶尔也有一些跨多个月的项目,尤其是要动很重的基础设施的那种,这种情况我们仍然会写 PRD。

Lenny00:10:29

I want to drill a little bit further into just how you're able to move so fast.

我想再往深里挖一下:你们到底凭什么能这么快?

I've never seen anything like the pace folks at Anthropic are shipping at — someone made this calendar of launches across Anthropic and it was literally every day there was a major feature or product.

Anthropic 的发布节奏我从没见过 —— 有人做了一份你们所有发布的日历,基本上每天都有重要功能或新产品。

One question people had online is — you guys built this incredible model Mythos that's still in preview because it's so powerful people are a little afraid of what it can do.

网上有人问:你们最近做出的 Mythos 这个模型,因为太强了还在 preview 阶段,大家有点怕它能干嘛。

Have you guys been using this? Is this part of the reason you've been able to move so fast?

你们内部在用它吗?这是不是你们能跑这么快的原因之一?

Chapter 04

Process beats the magic model

流程胜过秘密模型 · 顺手回应 leak / OpenClaw
11:00 — 14:30 · "不是 Mythos 让我们快" · 源代码泄漏的 process failure · OpenClaw 取舍逻辑
Cat00:11:03

We've been moving pretty fast for several quarters now.

我们其实已经快了好几个季度了。

So I think it's not fully Mythos. Mythos is an incredibly powerful model. We do use the models internally and I think this has increased our rate of shipping a little bit, but I don't think it explains the bulk of the increase.

所以原因不全在 Mythos。Mythos 是一个非常强的模型,我们内部确实在用,这让发布速度多少提升了一点,但提升的大头不是它。

I think a lot of it is the process and the expectation on the team.

我觉得真正起作用的,是流程和团队的预期设置。

We're very low on process. We want to remove every single barrier to shipping things.

我们流程极简,目标是把"上线"路上的每一个障碍都拆掉。

We want to make sure every single person on the team feels empowered to take their idea from just an idea to out in the world in less than a week — sometimes even in a day.

我们要让每个成员都觉得自己有权力,把一个想法从"只是想到"推到"已经发出去",而且一周内 —— 有时甚至一天内 —— 就完成。

"It's not fully Mythos. A lot of it is the process and the expectation on the team."

别再追问 Anthropic 是不是偷偷用了什么秘密模型 —— 真正让他们快的是"任何成员都能在 7 天内把想法推上线"这条预设规则。

Lenny00:11:41

Cool. Oh man, what an advantage to have the best model and also be building product. That's so cool.

酷。说真的,既能用最强的模型又能做产品,这种优势太爽了。

Cat00:11:46

We are very lucky to be able to work with the Frontier models.

能直接接触前沿模型,我们确实很幸运。

Lenny00:11:49

There are a couple of side things I want to go on side quests on.

我想顺着这个话题岔出去问几件事。

One is — a week ago or so, the whole source code of Claude Code leaked. Somebody got it out there. I think it was a mistake someone made.

第一件事是,大概一周前 Claude Code 的全部源代码泄漏了,被人传出来,看起来是某人犯了错。

Is there anything you can comment there — what happened, what went wrong, what should people know?

这件事你方便讲讲吗 —— 发生了什么、哪里出了问题、有什么是大家该知道的?

Cat00:12:15

We immediately looked into this when we saw it.

我们看到的时候第一时间就去查了。

We realized that this was the result of human error. There was a human working with Claude to write a PR. This was just an update to how we release our packages, and it actually went through two layers of human review.

我们后来确认这是人为失误。当时有一位同事在和 Claude 一起写 PR,改的只是我们怎么 release 包的流程,而且实际上经过了两层人工 review。

So this was a result of human error. We've hardened our processes to make sure that it doesn't happen in the future.

所以根因是人为错误。之后我们已经加固了流程,确保未来不会再发生。

Lenny00:12:42

Is this person still at Anthropic? Are they doing alright?

这位同事还在 Anthropic 吗?他还好吗?

Cat00:12:42

Yes. Yes. It's a process failure, and the most important thing is to just learn from it and to add more safeguards so that doesn't happen again.

在的、在的。这本质上是流程失效;最重要的是从中学到东西,然后加更多的安全护栏,让它不再重演。

That's what we've been focused on, and most of those have shipped.

这就是我们一直在做的事,大部分加固已经上线了。

Lenny00:12:54

Another question I had is OpenClaw.

另一个问题是 OpenClaw。

Recently there's been this move to keep people from using their Claude subscription with their open Claudes. People got really upset — they're confused why this is happening. It feels like there's harm caused to the open source community.

最近你们出了一个动作:不让大家把 Claude 订阅接到他们的 open Claude 工具上。很多人很不满,搞不清你们为什么这么做,感觉像是对开源社区造成了伤害。

What do people need to understand about what went into this decision?

你能讲讲这个决定背后是怎么想的吗?

Cat00:13:18

We've been seeing a lot of demand for Claude, and we've been working very hard to both scale our infrastructure and also to make our harness more token efficient so that you can get more usage out of it.

Claude 的需求一直非常猛,我们一边在拼命扩基础设施,一边在让 harness 在 token 上更省,让大家用同样的额度能跑更多事。

It wasn't designed for third party products, which have different usage patterns than our first party ones.

订阅最开始就不是为第三方产品设计的 —— 它们的使用模式和我们自家产品很不一样。

We spent a bunch of time trying to figure out what is the most seamless transition that we can offer.

我们花了不少时间在想,能给出的最平滑的过渡是什么样。

I was very happy to be able to say that everyone gets some credits alongside their subscription.

最后我们能让每个人都拿到一笔附带的 credits,这一点我个人挺开心的。

But yeah — we did have to make the hard decision that we needed to prioritize our first party products and our API.

但确实 —— 我们必须做这个艰难的决定:把第一方产品和 API 放在优先位。

So this is a decision that resulted from that.

这件事就是这条优先级带出来的结果。

Lenny00:14:00

To me it makes so much sense. You guys are subsidizing this usage at like 200 bucks a month and there's basically unlimited use of this.

在我看来完全说得通。一个月 200 美元,等于你们在拿钱补贴这个用量,而且基本是不限量的。

Businesses are trying to make money. You can't just give away compute when it's so in demand. So I get it.

公司终归要赚钱,算力这么紧的时候,你不能再白送出去 —— 我很理解。

Chapter 05

Inside the PM org at Anthropic

Anthropic 的 PM 组怎么搭(30-40 人)
14:30 — 17:30 · Research / CDP / Cloud Code / Enterprise / Growth · Diane 的 Research PM 组 · Managed Agents
Lenny00:14:26

Coming back to the PM team — what does the PM team look like at Anthropic? How many PMs are there? How are they organized?

回到 PM 团队 —— Anthropic 的 PM 组怎么搭?有多少 PM?怎么组织的?

Cat00:14:30

We have a few PM teams. I think we're maybe around 30 or 40 PMs right now.

我们有几个 PM 团队,目前总人数大概在 30-40 个 PM。

We have the research PM team, who Diane leads — this team is responsible for understanding all of the feedback from our customers for our models, feeding that to the research team to act on, and they also shepherd the model launch.

第一支是 Diane 带的 research PM 团队 —— 负责把客户对模型的反馈吃透、传给研究团队推进,同时也负责模型发布的执行。

There's the Cloud Developer Platform team that maintains the APIs that Claude Code is built on top of, and they also release things like Managed Agents — which is a way for you to build your agents and we can host them on your behalf.

第二支是 Cloud Developer Platform 团队,他们维护 Claude Code 所依赖的 API,也负责发布像 Managed Agents 这样的产品 —— 让你来定义 agent,由我们替你托管运行。

Then there's Claude Code, that works on both Claude Code and the co-work core products.

第三支是 Claude Code 团队,负责 Claude Code 和 co-work 这两个核心产品。

There's Enterprise that helps make Claude Code and co-work easier to adopt for all of our enterprise customers — so this is everything from cost controls, RBAC, security controls, just making sure that these enterprises feel very confident and comfortable using our tools.

第四支是 Enterprise,负责让企业客户更容易上 Claude Code 和 co-work —— 从成本管控、RBAC、安全管控,一直到让客户用得放心。

Then we also have our growth team that is responsible for growing across our entire product suite.

第五支是 growth 团队,负责整条产品线的增长。

We work very closely with them on Claude Code and co-work growth, and I know they also work with our other teams on CDP growth — growth of people who use the Cloud API.

我们和他们在 Claude Code、co-work 的增长上配合非常紧;他们和别的团队也一起做 CDP 的增长 —— 也就是 Claude API 用户的增长。

Lenny00:16:15

Speaking of growth — Amol was just on the podcast. He had this interesting insight that most people haven't been sharing.

说到 growth —— Amol 刚来过这个播客。他抛了一个观点,我觉得很少人讲。

There's always this sense that we need fewer PMs in the future. Why do we need PMs? Engineers can just ship. His take is that because engineers are moving so fast, PMs and designers are squeezed.

现在到处都觉得未来需要的 PM 更少了 —— 工程师自己就能发,要 PM 干嘛?Amol 的看法是反过来的:正因为工程师跑得太快,PM 和设计师反而被挤压了。

There's less time to stay on top of everything that is happening — there's a feature shipping every day. So his take is he needs more PMs because it's hard to keep up.

每天都有功能上线,根本没时间跟上每件事。所以他的结论是他需要更多 PM,光是"跟上"都不容易。

Do you feel like there will be an increase in hiring of PMs? What do you think is going on with the PM profession long term?

你觉得 PM 招聘会增加吗?长期看 PM 这个行业会怎么走?

Chapter 06

Roles merge — taste is the moat

角色融合 · taste 才是稀缺品
17:30 — 21:00 · 工程师做 PM · PM 写代码 · 设计师 land PR · "engineer with taste" 才是关键招聘画像
Cat00:16:30

I think all of the roles are merging.

我觉得所有这些角色都在融合。

PMs are doing some engineering work, engineers are doing PM work, designers are PMing and also landing code.

PM 在做一些工程的活,工程师在做 PM 的活,设计师既做 PM 又能 land 代码。

You can either hire a lot more engineers who have great product taste, or you can keep your engineering hiring the same and hire a lot more PMs to help guide some of their work.

这种情况下,要么多招一批有很强 product taste 的工程师,要么工程招聘不变、多招 PM 来引导工程的工作。

On our team we're pretty focused on hiring engineers with great product taste. This way we can reduce the amount of overhead for shipping any product.

我们团队的选择是后者反过来 —— 我们重点招"有 product taste 的工程师",这样把每个产品上线的协调成本压到最低。

There are many engineers on our team who are fully able to end-to-end go from "see user feedback on Twitter" through to "ship a product at the end of the week" with almost no product involvement.

我们团队里很多工程师可以从"在 Twitter 上看到用户反馈"端到端走到"周末发出一个产品",中间几乎不用 PM 介入。

This I think is actually the most efficient way to ship something.

在我看来,这才是发布东西最高效的方式。

So engineer and PM are kind of overlapping, and you will get a lot of benefit from having more of either.

所以工程师和 PM 已经互相重叠了,任何一边多招点人都有用。

I think product taste is still a very rare skill to have, and we'll pretty much hire anyone who we feel has demonstrated this strongly.

而 product taste 仍然是一种非常稀缺的能力 —— 只要我们认为某个人在这上面证明得足够强,我们基本就会招。

Lenny00:17:25

And your background was in engineering, right?

你自己原本是做工程的,对吗?

Cat00:17:27

Yeah, I was an engineer for many years. I was then a VC very briefly before joining Anthropic.

对,我做过很多年工程师。之后短暂地做过一阵 VC,然后才加入 Anthropic。

Almost all the PMs on our team have either been engineers or ship code here on Claude Code.

我们团队几乎所有 PM,要么以前是工程师,要么现在就在 Claude Code 上 ship 代码。

That's one of the things that I think helps build trust with the team, and also enables us to move a lot faster.

这件事是我们和工程团队建立信任的关键,也是我们能跑这么快的原因之一。

Actually our designers also have been front-end engineers before.

连我们的设计师之前也都是前端工程师。

Lenny00:17:54

There's definitely this merging happening, the Venn diagrams you're combining.

确实在发生融合,你们在把好几个圈合到一起。

The big question for a lot of people is — if you're coming from engineering or product or design, which of those core skills is going to be most valuable?

很多人的疑问是 —— 如果你的底子是工程、产品或者设计,这三种里哪一种核心能力以后最值钱?

I could see at Anthropic and on Claude Code, engineering is very valuable. I'm curious if at other companies, having a design background and becoming a PM is more valuable.

在 Anthropic、在 Claude Code 这边,工程显然很值钱。我好奇在别的公司是不是设计背景转 PM 反而更值钱。

Cat00:18:16

I still think it comes back to product taste.

我还是觉得最后还是回到 product taste。

As code becomes much cheaper to write, the thing that becomes more valuable is deciding what to write.

代码越来越便宜,反过来"决定写什么"就越来越值钱。

What is the right UX for this feature? What is the most delightful way that a user can experience it?

这个功能的正确 UX 是什么?用户体验它最让人惊喜的方式是什么?

We get tens of thousands of GitHub issues asking for every single thing under the sun, and it takes a lot of care and taste to figure out which of these is worth building and what is the right way to build it.

我们 GitHub 上有几万个 issue,各种五花八门的需求都有 —— 在这里面挑出"哪些值得做、要怎么做",非常吃 taste。

That skill set can come from any background, but I think that's the most important thing.

这种能力可以来自任何背景 —— 但它本身是最重要的。

The reason why an engineering background is particularly useful, at least for the next few months, is — if you have an engineering background, you have a better sense for how hard something should be.

至少在接下来几个月,工程背景特别有用的原因是 —— 你对"一件事应该多难"有更准的直觉。

That's often a factor in what you choose to build. If something is very easy to build, then maybe instead of debating it, you just spend an hour doing it.

这会直接影响你选什么去做。一件事如果很容易做,那不用讨论,直接花一小时做掉就行。

But if something is harder to build and you know that upfront — you know that this will cost a lot more for our team to get out the door — so it helps a bit with the prioritization.

如果一件事很难做、而且你提前就知道,那意味着团队要付出更大代价才能上线,这种判断对优先级帮助很大。

Lenny00:19:27

You said "for the next few months." Is that just because the models will get so good potentially in the next few months — you may not even need to know that as much?

你说"接下来几个月" —— 是因为模型几个月后可能强到这种判断都不太需要了?

Cat00:19:30

I think the valued skill set does change quite frequently, and so it's really hard to predict more than a few months out.

我觉得"值钱的技能"这件事一直在变,变得很频繁,所以超过几个月就很难预测。

It's less a commentary on what shift I think will happen, and more a commentary that I think large shifts will happen.

所以这句话其实不是在说"我预判会有什么变化",而是在说"我相信会有大变化"。

"As code becomes much cheaper to write, the thing that becomes more valuable is deciding what to write."

代码变成便宜的那一刻,真正稀缺的能力一夜之间从"会写"切到了"知道写什么" —— Anthropic 招聘画像背后就是这一句话。

Chapter 07

Where humans still matter

AI 时代,人类还做什么
21:00 — 28:30 · 常识 / EQ · 龙卷风眼里的冷静 · brutally prioritize · 牺牲产品一致性 · /powerup
Cat00:20:03

There's a large increase in coding capability which then changes what other roles are valuable.

编码能力上来一大截,会反过来改变"哪些其他角色值钱"这件事。

The most important thing is to be able to have first-principles thinking — where you can figure out how the tech landscape is changing, what the team really needs from you, and to jump in and fix that hole.

最重要的能力是 first-principles thinking —— 能看清技术格局怎么变、团队真正缺你哪一块,然后自己跳进去补那个洞。

The work is becoming more amorphous, which means that a great PM is able to understand what all the gaps are, figure out what the highest priority ones are, and then figure out — okay, how do I learn that skill set, or what is the skill set that I have that I can apply to this challenge?

现在工作的边界越来越模糊。一个好的 PM 要能看到所有缺口,挑出最重要的那些,然后想:我怎么学那块能力,或者我现有哪块能力能拿过来打这场仗?

The current environment values people who are able to wear a lot of hats, are able to swap them, and are very low-ego about what work they do to help the team move faster.

现在的环境最值钱的是这种人:身上能戴好几顶帽子、能在帽子之间随时切换,对"自己做什么活才让团队跑得更快"这件事完全没有 ego。

Lenny00:21:06

I love this answer. There's this question I've been asking people in your shoes — folks at the bleeding edge of what AI is capable of — which is just like, where will human brains continue to be useful and necessary for a while, until we get to super intelligence?

我喜欢这个答案。我最近一直在问像你这样在 AI 最前沿的人一个问题 —— 在我们走到超级智能之前,人类的大脑还会在哪些地方继续有用、继续必要?

What I'm hearing is — picking the things to work on, knowing where the market's going, figuring out what to prioritize. And then knowing if the thing you've built is good and right, and getting it out there in some early version. Does that sound right?

我从你这听到的是 —— 挑该做什么、看清市场往哪走、定优先级;然后判断做出来的东西好不好、对不对,把它的早期版本扔到外面去。这样讲对吗?

Cat00:21:43

I think humans still provide a level of common sense that the models don't.

我觉得人类仍然在提供模型给不出的那种常识。

There's a thousand moving pieces to any product launch. Some of them are very small, but there's always a lot that could potentially go wrong.

任何一次产品发布都有成千上万个动来动去的零件,很多看起来很小,但任何一个都可能出岔子。

The model doesn't always have a great sense of who all the stakeholders are, how they relate to each other, what their preferences are, what are the right venues to communicate with them to keep them on board.

模型对"谁是 stakeholder、他们之间什么关系、各自有什么偏好、用什么场合沟通能让他们一直 on board",这些常常拿捏不准。

A lot of this more tacit, common sense, EQ-kind-of knowledge is still very valuable.

这种偏隐性的、常识层面的、EQ 性质的知识,目前还非常值钱。

Of course, we want the models to get better at this, and I think they will be. But right now I think there are still gaps.

当然我们希望模型在这上面能进步,我相信它们会;但当下确实还有缺口。

Lenny00:22:30

How do you deal as a human going through so much constant change — being on the inside of the tornado? Maybe it's calm there, but how do you stay on top of what's going on, how do you stay sane?

作为一个普通人,身处这种持续巨变,你怎么扛?在龙卷风眼里也许反而是静的,但你怎么跟得上每件事、怎么保持头脑清醒?

Cat00:22:39

Our team is full of people who lean into the chaos.

我们团队里都是那种"愿意扎进混乱"的人。

We try to face every challenge with a smile because there's always so much going on. There are always so many risks and tricky situations that you know if you get too stressed about anything, you'll burn out.

每一个挑战我们都努力笑着面对 —— 因为事情永远那么多,风险和棘手情况永远那么多,任何一件压你太久,你就会 burn out。

We really look for people who can look at a challenge and be like — that's going to be hard, but I'm excited to tackle it, and I'm going to do the best that I possibly can. I know I won't be perfect but I'll be able to sleep at night knowing that I did my best.

我们招的是这样一种人:看到挑战会说"这会很难,但我很期待把它干掉,我会尽力,可能做不到完美,但晚上睡得着,因为我尽力了"。

Lenny00:23:15

That's an interesting answer to "what skills will be important in this future" — because, I forget who said this, maybe Ben Mann — that this is the most normal the world will ever be.

这对"未来什么技能重要"是个有意思的答案 —— 我忘了是谁说的,可能是 Ben Mann:今天这个世界,会是未来回看时最"正常"的一天。

Cat00:23:23

It definitely gets harder.

确实会越来越难。

There are a lot of weeks where maybe Sunday night there's some P0, and then by Monday there's a P0, and by Monday afternoon there's a P0000, and you're like — wow, I can't believe I was so worried about that P0 from Sunday.

很多个礼拜是这样:周日晚上来一个 P0,周一又来一个 P0,周一下午来一个 P0000,然后你回头想 —— 哇,我居然为周日那个 P0 那么紧张。

You just have to acknowledge that there's only so much that you can do, that you need to sleep well so that you can make good decisions next day, and just brutally prioritize where you spend your time.

你必须承认自己能做的就那么多,必须睡好,第二天才能做出好决定 —— 然后心狠手辣地把时间花在最重要的地方。

What's the most important thing to get right? And be okay letting things go.

最该做对的是哪件?然后接受其他事被放掉。

There are products that we ship that aren't as polished as I wish they were. But our top goal is to help empower professional developers, and if a product isn't successful — as long as it's not blocking the core use case — it's okay because we'll hear the feedback and we'll fix in the next release.

我们 ship 出去的有些产品没我希望的那么打磨。但我们的最高目标是让专业开发者更强,只要这个产品的失败没有挡住核心 use case,那就 OK —— 反馈会进来,下个版本就修。

Launching a feature that is buggy is the kind of thing that would have kept me up at night. But it is something that I am now able to live with, knowing that we're going to get that quick feedback and fix it in the next release.

放一个有 bug 的功能上线,以前是会让我整夜睡不着的事;现在我能接受了 —— 我知道反馈会很快回来,下个版本就改。

Lenny00:24:30

There's that GIF, I think it's from Pirates of the Caribbean, of this guy walking down a pair of stairs on a ship, and the whole ship is just being demolished around him, and he's so chill, just strolling down the staircases, everything's falling apart.

有个 GIF,我记得是《加勒比海盗》里的:一个家伙在船上慢慢踱下楼梯,整艘船在他周围塌成废墟,他一脸淡定,就这么走下去。

Everyone I've met from Anthropic is just so chill and so optimistic.

我接触过的 Anthropic 人都是这种调调:特别淡定,特别乐观。

Cat00:24:51

If you don't have it, you'll get pretty burnt out.

如果你没有这种"淡定+乐观",你很快会被烧穿。

We tend to hire people who have been in the industry for a while and have experienced lots of ups and downs, and have a good sense for what gives them energy and how to maintain their energy over time.

我们偏向招那种"已经在这行待了一阵、经历过起起伏伏"的人 —— 他们清楚什么事让自己有能量、怎么把能量在长期里维持住。

Lenny00:25:20

Something I wanted to ask — there's these roles blurring, engineers becoming PMs, everyone's everyone. What do we lose in that world?

我还想问 —— 角色一路融合,工程师变 PM,大家变成一团模糊,这种情况下我们会失去什么?

Do we lose career ladders and clear career paths? Do we lose design consistency, code quality? What are some things we're sacrificing for the greater good?

职级阶梯、清晰的成长路径会不会丢?设计一致性、代码质量会不会丢?为了"更大的好",我们具体在牺牲什么?

Cat00:25:42

We're sacrificing product consistency.

我们正在牺牲的是产品一致性。

Historically, when code was expensive to write, you would carefully plan out everything in your product suite, how every product relates to each other, what the use case for every single one is, how they integrate — you would pretty much have one product for each use case.

以前代码贵,你会非常仔细地规划整条产品线 —— 产品之间什么关系、每个的 use case 是什么、怎么集成,基本上一个 use case 对应一个产品。

Now with AI moving so quickly and with so many ideas to test out, we do sometimes have features that overlap with each other.

现在 AI 跑得这么快、要试的想法这么多,有时候我们的功能之间确实会有重叠。

A lot of times it's because there are two form factors that we love internally and we want the external audience to tell us which one is better.

很多时候是因为内部对两种 form factor 都很喜欢,我们想让外部用户告诉我们哪一种更好。

What that means for someone who's a new user is — a new user might not know what is the best path to accomplish X. There is more education we need to do.

对新用户来说,代价是 —— 他们不知道做 X 应该走哪条路。我们要做更多的教育工作。

I think users also feel like it's hard to keep up with the latest. Usually in traditional PM, you ship a feature every month or quarter, so it's really easy for a user to understand — okay, I just need to check in once a month.

用户也觉得跟不上最新。在传统 PM 节奏里,你每月或每季度发一个功能,用户很容易做"我每月看一次就行了"这种安排。

With these agentic tools — not just Claude Code and co-work, but across the whole ecosystem — people feel this need to check Twitter every single day to see what the absolute latest thing is.

而现在这一波 agentic 工具 —— 不只是 Claude Code 和 co-work,整条生态都一样 —— 大家觉得自己得每天刷 Twitter 才能跟上"最新的最新"。

There's more we can do to help people feel less like they're on this ever-increasingly-fast treadmill.

我们还能在产品上做更多事,让大家不觉得自己被绑在一台越来越快的跑步机上。

I would love people to feel like they can just open these tools, the tools will educate them or teach them what they want to know, and they can feel more brought along.

我希望大家打开工具的时候,工具会主动教他们想知道的事,让他们感觉是被工具带着往前走的。

Lenny00:27:48

I saw you launched this really interesting feature the other day — I think it's /powerup — where it walks you through all the cool ways and basically the best practices to use Claude Code. Is that along these lines?

前几天看到你们上的一个挺有意思的功能 —— 好像叫 /powerup —— 它会带你走一遍 Claude Code 的最佳用法。是不是这条思路上的产物?

Cat00:27:57

Yeah, exactly.

对,正是。

In the past, we didn't actually want to do something like /powerup because we felt like the product should be intuitive enough that you don't actually need to go through any tutorial.

以前我们其实不想做 /powerup 这样的东西,觉得产品应该直观到不用任何教程。

Over time, we've realized that there's just so many features and there's so much demand for a built-in onboarding experience that we diverged a bit from our original principle of "no onboarding flow" and added this — because there's just so many users who wanted to know there's 100 features, what are the 10 that I absolutely need to use?

但功能越来越多,大家又强烈想要内建的 onboarding,我们后来稍微背离了原来"不做 onboarding 流"的原则,加上了这个 —— 因为太多用户想知道:你这有一百个功能,有哪 10 个是我非用不可的?

Chapter 08

Enterprise with launch culture

B2B 也用 C 端的发布节奏
28:30 — 34:55 · Anthropic 反共识地用消费节奏服务企业 · 使命 vs Focus · OpenClaw 决定的逻辑根
Lenny00:28:32

Anthropic has been really successful with B2B enterprises, where traditionally you don't launch a bunch of stuff. You just kind of have a quarterly release maybe — it's the opposite of "every day we got something new."

Anthropic 在 B2B 企业市场做得非常成功 —— 而传统上 B2B 是不会狂上线东西的,可能就一个季度发一次,跟"每天都来点新东西"完全相反。

The run Anthropic has been on is just otherworldly. Anthropic was way behind when it started — one of the least funded companies, didn't have distribution, OpenAI was way ahead. It was like "no way Anthropic has any chance to compete long term."

Anthropic 这条增长曲线简直不像人间产物。一开始 Anthropic 被远远甩在后面 —— 融资最少之一、没有分发渠道、OpenAI 遥遥领先,大家都觉得"Anthropic 长期不可能竞争"。

Now it's just killing it — beating the biggest companies. The growth is uh like 11 billion in ARR with one-month % growth that by the time this comes out is probably even higher.

现在却完全在屠榜 —— 击败那些最大的公司。增长大概是 110 亿美元 ARR,上个月环比增长幅度,等这一期播出的时候大概只会更高。

Just being on the inside — what are some ingredients that have allowed Anthropic to be this successful and come from behind?

作为内部人,你能讲讲 Anthropic 能从落后到追上、做得这么好,有哪些关键配方吗?

Cat00:29:29

The two most important things are — one, this unifying mission. It's hard to state how important this is.

最重要的两件事 —— 第一,统一的使命。它有多重要,怎么强调都不为过。

We hire people who care most about bringing safe AGI to all of humanity.

我们招的人,最在乎的就是把安全的 AGI 带给全人类。

This is something that we reference frequently in our decisions about what our entire product or should focus on shipping.

在决定整个产品该上线什么时,我们会频繁地拿这个使命来对齐。

Because we put this mission above any individual product line, we're able to make very fast decisions that cut across the entire org and execute on them in a unified way.

因为我们把这个使命摆在任何单一产品线之上,我们能做出跨整个组织的快速决策,并且统一执行。

This is something that I've never seen at a company of our scale.

我在我们这个体量的公司里从来没见过这种状态。

Lenny00:30:12

Just to make sure that's clear — having the number-one mission be safety alignment, making sure AI is good for the world, having that as a clear mission makes decisions a lot easier?

我确认一下 —— 你们把"安全对齐、让 AI 对世界有益"放在首位,把它定成一个清晰的使命,这件事会让所有决策都变容易?

Cat00:30:24

If there are two competing priorities, we'll talk about which one is more important for Anthropic's mission, and it makes it a lot easier to decide which we prioritize.

两件事撞在一起的时候,我们会讨论哪一个对 Anthropic 的使命更重要,这样就很容易决定优先级。

Then everyone will stand behind the one that we decide.

然后大家都会站在我们决定的那一边。

Sometimes that means — hey, we want to ship something on Claude Code, but this other thing is more important. So we deprioritize shipping this and we just wait until later.

有时候这意味着:嘿,我们想在 Claude Code 上发个东西,但另一件更重要,那 Claude Code 这件就先押后,延后再发。

Lenny00:30:50

What's interesting about that is — it explains, versus another company that maybe rhymes with bopen-bi, did a lot of different things. What I'm hearing is — we're not going to launch a social network, we're not going to launch a feed of interesting information, because it's not aligned to this mission. That has kept Anthropic focused.

有意思的是,这就解释了 —— 跟某个名字押韵 bopen-bi 的公司比起来,你们没去做一堆杂事。我听到的是:我们不会去做社交网络、不会去做一个"有趣信息流",因为这些跟使命对不齐。这件事让 Anthropic 一直保持聚焦。

Cat00:31:10

When I think about mission, I think about putting Anthropic's goals ahead of any individual or any individual product.

我对 mission 的理解是 —— 把 Anthropic 的目标放在任何一个个体、或任何一条产品线之上。

For me, the second thing we're very good at is focus. Mission to me is slightly different.

对我来说,第二件我们做得很好的事是 focus。mission 和 focus 在我心里其实有区别。

Mission means that teams are willing to make sacrifices that hurt their own goals and their own KRs in service of Anthropic's goals and KRs. People are very happy to make those trade-offs.

mission 的意思是:团队愿意为 Anthropic 的目标和 KR 牺牲自己的目标和 KR,而且大家很乐意做这种取舍。

An extreme example — if Claude Code failed but Anthropic succeeded, I would be extremely happy. The whole team is very willing to make decisions that follow that chain of thought.

一个极端的例子 —— 如果 Claude Code 失败了但 Anthropic 成功了,我会非常开心。整个团队都很愿意按这条逻辑做决定。

"If Claude Code failed but Anthropic succeeded I would be extremely happy."

这一句最能解释 OpenClaw 取舍 —— 不是产品线的内斗,而是整支团队真的愿意为 Anthropic 整体目标牺牲自己的产品。

Lenny00:31:58

Do you feel like the OpenClaw decision is part of this — like, this is not furthering Anthropic's mission, we need to stop this because it's not working in the way we want?

你觉得 OpenClaw 那个决定是不是这条逻辑的一部分 —— 这件事没在推进 Anthropic 的使命,所以要刹车?

Cat00:32:10

One of the most important things for Anthropic is to grow the number of users that we're able to reach.

Anthropic 最重要的事之一是扩大能服务到的用户规模。

One of the ways we're able to do this is with the Claude subscriptions on our first party products, and so we very much want to double down on that — but that does come at the expense of third party products sometimes.

实现这件事的一种方式是通过我们第一方产品的 Claude 订阅,所以我们非常想在这一边重押 —— 而这有时候就要以牺牲第三方产品为代价。

Lenny00:32:28

We've been talking about Claude, co-work, all these things. There's Claude Code, Claude Desktop, co-work. What's the best way to understand when to use which?

咱们一直在聊 Claude、co-work 这些。Claude Code、Claude Desktop、co-work 三个产品 —— 怎么理解什么时候用哪个最合适?

Cat00:32:44

I tend to use Claude Code in the terminal when I'm just kicking off a one-off coding task and I want all of the latest features.

我自己在做一次性的编码任务、又想用上最新功能时,会用 terminal 里的 Claude Code。

The CLI is our initial product surface and it's also the one where our features often land first, so it's the most powerful of all the tools.

CLI 是我们最初的产品形态,新功能也常常先落在 CLI 上,所以它是这三者里最强的。

Desktop really shines when you're doing something that requires front-end work.

Desktop 在涉及前端工作时最闪光。

One thing I love to do is use our preview feature — if I'm building a web app, I'll use Claude Code and Desktop with the preview pane open on the right hand side, so I can actually see the web app I'm making in real time as I'm chatting with Claude.

我特别喜欢用 preview 这个功能 —— 做 web app 时,我会同时开 Claude Code 和 Desktop,右边挂着 preview pane,一边和 Claude 对话,一边实时看着 web app 在变。

Desktop is also great for people who want something more graphical. A terminal can feel very unfamiliar to someone non-technical.

Desktop 对于喜欢更图形化界面的人也更友好。terminal 对很多非技术用户来说很陌生。

Desktop is also great for getting an at-a-glance view of everything that's happening — you can see your CLI terminal sessions, your other desktop sessions, your sessions kicked off on web and mobile. It's a one-stop control plane.

Desktop 还能让你一眼看见所有任务 —— CLI terminal 会话、其他 desktop 会话、网页和手机端开起来的会话,全部集中。它是个 one-stop 的 control plane。

Web and mobile is really great for kicking things off on the go. CLI and desktop both require you to be on your local laptop.

Web 和 mobile 最适合"在路上随手起任务"。CLI 和 desktop 都得在你自己的 laptop 上。

Sometimes you're out and about, you're touching grass, going on a walk, and you don't have your laptop open — I can't count the number of people I've seen holding their laptop open like tethered to their phone while they're outside.

有时候你在外面散步、touch grass,laptop 不在手上 —— 我数不清看到多少人在户外抱着 laptop 用手机做热点的样子。

For me, what mobile lets you do is kick off these tasks on the go, so that you don't need to bring your laptop everywhere.

对我来说,mobile 的意义就是 —— 在路上也能起任务,不必随时背着 laptop。

Chapter 09

Mission, kill, demo Friday

使命第一 · 砍功能 · 周五 demo
34:55 — 43:00 · CLI / Desktop / Web 怎么各司其职 · co-work 让 PM 一夜出 20 页 deck
Lenny00:34:57

I love that. I've seen people on planes — it's just like such a meme now — "I need to finish, let this agent finish. I can't shut this down. I need Wi-Fi."

我喜欢这个。我在飞机上见过的场面已经成了梗 —— "我得让这个 agent 跑完,别断我!我要 Wi-Fi!"

Cat00:35:04

For co-work, the role this fills is — there's a lot of work that everyone does where the output isn't code.

co-work 解决的是 —— 我们每个人每天有大量工作,输出物根本不是代码。

Whether that's getting to Slack zero or inbox zero, creating a slide deck for some customer meeting, writing a quick doc on what the goals of a feature are, or what the launch plan is — all these tasks produce non-code outputs, and co-work is best positioned for that.

不管是把 Slack 清干净、inbox zero,还是做一份给客户会议用的 deck,或者写一篇"这个功能的目标 / 发布计划"的小文档 —— 这些任务产出都不是代码,co-work 是干这些事的最佳形态。

The way I split the products in my mind is — if I'm building something where the output is code, I use Claude Code or Desktop or Claude Code on mobile. If the output is anything that's not code, I use co-work.

我心里的划分方式是:输出是代码,用 Claude Code / Desktop / Claude Code on mobile;输出不是代码,用 co-work。

Lenny00:35:48

People are sleeping on the success that co-work is having. It's growing incredibly fast, and people still don't understand maybe what it's for.

大家都没意识到 co-work 火到什么程度了。它增长得非常快,但很多人还搞不清它到底是干嘛的。

Give us a couple of use cases — in your work as a PM, what are some really interesting, maybe unexpected ways you use co-work to save time, get more work done?

能不能讲两个 use case —— 你做 PM 时,用 co-work 救你时间、提升产出最有意思甚至最意外的几种用法是什么?

Cat00:36:08

If you're getting started on co-work, the first thing you really need to do is connect all the data sources that are relevant to your role.

用 co-work 第一件事:把所有跟你角色相关的数据源都接上。

Co-work can only do a great job if it has access to all the context it needs to curate the output for you.

co-work 必须能拿到这些 context,才有办法为你定制产出。

For me, that means I connect it to my Google Calendar, my Slack, my Gmail, my Google Drive — so it can find relevant context, ask questions, pull in threads. This substantially improves the quality of the result.

我自己接的是 Google Calendar、Slack、Gmail、Google Drive,这样它能去找相关上下文、提问、把对话拽过来 —— 结果质量会显著提升。

Last night I was working on the Code with Claude conference — there are a few talks I'm giving there. One talks about the transition of Claude Code from an assistant to a full-on agent.

就拿昨晚为例,我在准备 Code with Claude 会议的内容 —— 那场会上我有几场演讲,其中一场讲的是 Claude Code 从助手升级成完整 agent 的转变。

One thing I wanted to do in this talk was to showcase all of the products we've been shipping that enable this transition, and figure out what are the success stories internally that we can use as demos.

我想在讲座里展示所有为这次转变铺路的产品发布,以及找出内部哪些成功故事可以拿来当 demo。

I have my Google Drive connected, I have Slack connected. Alex, our product marketer, put together a draft of the points he thinks we should cover.

我接了 Google Drive 和 Slack。我们的产品市场经理 Alex 写了一份草稿,列了他觉得该讲的点。

I fed this all into co-work. I told it the narrative I want to tell. And it just worked for an hour.

我把这些全喂给了 co-work,告诉它我想讲的叙事是什么。然后它就跑了一个小时。

It walked through Twitter to see what we launched. It looked through our evergreen launch room. It looked in our Claude Code announce channel where our team posts demos of how they've been getting the most value out of Claude Code.

它去 Twitter 看了我们都发了什么,翻了我们的 evergreen launch room,去了我们的 Claude Code announce 频道(团队成员在那里贴 Claude Code 用得最爽的 demo)。

It synthesized all this together into a 20-page deck that I woke up to this morning. I read through it and it was pretty good.

然后它把这一切整合成了一份 20 页的 deck,等我早上一醒就看到了。我读了一遍,挺不错。

There were a few tweaks — I like my slides to have extremely minimal words and it was a little too wordy. But it was far faster than what I would be able to produce.

有几处需要调 —— 我的 slide 偏向极简文字,它做得稍微啰嗦了点。但比我自己产出的速度快太多。

Because co-work has access to our whole design system, it actually looks like an Anthropic designer put it together. When you visually see it, you're like "oh, this is incredibly polished."

因为 co-work 接了我们整套 design system,它做出来的视觉上像是 Anthropic 自己的设计师做的 —— 你一眼看到会觉得"哇,这非常 polished"。

Making this slide deck would have taken me hours, but instead it turns out a draft that is actually quite good — so I can focus on making sure the demos are amazing.

这份 deck 我自己做要好几个小时,而它直接给我一份相当能用的草稿 —— 我就能把精力放在把 demo 做漂亮上。

"Co-work just like went off for a few hours and built the whole slide deck."

这是 Cat 的标配工作流:连数据源 → 写叙事和参考材料 → 让 co-work 自己跑几个小时 → 醒来读改 → 把人留在最值钱的事上(demo 本身)。

Lenny00:38:45

This sounds like a dream come true to PMs — putting decks together is so annoying.

PM 们听到这个简直是美梦成真 —— 拼 deck 实在太烦了。

Cat00:38:49

It's so slow.

慢得要死。

Lenny00:39:00

To help people try this for themselves — step one is connect what? Slack, what else do you suggest?

想让别人也跑一遍 —— 第一步是接什么?Slack 之外你还推荐接什么?

Cat00:39:07

Slack, Google Calendar, Gmail, G Drive. You should connect your communications tools and where you store your source-of-truth data for what your team cares about, what you care about, and what you're working on.

Slack、Google Calendar、Gmail、G Drive。你该接的就是沟通工具,以及你团队和你自己工作 source of truth 数据的存放点。

Lenny00:39:21

And what was the prompt roughly that you put in there to generate this deck?

你 prompt 大概怎么写的,让它生成这份 deck?

Cat00:39:26

I just wrote — make me a slide deck for the Code with Claude conference. This is what our PMM suggested it should cover. This is the current draft I made that I don't like. This is one I made manually that I don't like, but I linked it.

我大概是这么写的:给 Code with Claude 大会做一份 slide deck。我们 PMM 建议的内容大纲是这个。我之前做的草稿我不喜欢,但我把链接挂上了。

"Can you start by creating a proposed outline with details? Also, make sure it doesn't overlap too much with a keynote talk, which is more important."

"先帮我做个带细节的拟议大纲,另外要保证别和那场 keynote 重得太厉害,那一场更重要。"

Then Claude read a bunch of the links I sent and created a proposed outline. I read through its proposal and all the different ideas, and just made a decision on what I wanted to actually be in the final deck.

然后 Claude 读了我发它的一堆链接,搞了一个拟议大纲。我把它的提案和各种想法读了一遍,自己拍板决定 deck 里实际要放什么。

This is an example of what the role of the PM still is today. Claude is a great brainstorming partner. It can synthesize a massive amount of information and present all the possibilities to you. But the role of the PM is still to make the end decision of what should belong in the final product.

这就是今天 PM 这个角色的样子 —— Claude 是非常好的 brainstorming partner,能把海量信息整合起来摆在你面前。但最终决定"什么应该进最后版本",仍然是 PM 的工作。

For this, I decided I wanted the talk to cover the progression from making local tasks successful, to making every PR green, to helping engineers land more PRs — and for each, which demo would be the most compelling.

这次我定的是 —— 这场讲要讲一条递进:让本地任务跑通 → 让每个 PR 都绿 → 让工程师 land 更多 PR;每一段对应哪个 demo 最有说服力。

After that decision about the outline, co-work just went off for a few hours and built the whole slide deck.

大纲拍板之后,co-work 就自己跑了几个小时,把整份 deck 拼出来了。

Lenny00:40:50

How does it know the design system of Anthropic?

它怎么知道 Anthropic 的 design system?

Cat00:41:00

For this, we already have a standardized deck we use across all of our external engagements. I gave Claude access to that — it can see what colors we use, what fonts, the different slide formats that are possible. It has like 20 of these example slides.

这次我用的是我们对外已经统一的 deck 模板。我把它给 Claude 看 —— 它能看到我们用什么颜色、什么字体、可选的 slide 板式。模板里大概有 20 张样例 slide。

You can also connect to your Figma MCP if you have your slide format saved there, and it can pull that in.

如果你的 slide 板式存在 Figma 里,也可以接 Figma MCP,让它从那边拉。

Lenny00:41:48

Along those lines — what's in your stack of tools as a PM at Anthropic, other than the Anthropic tools? You mentioned Slack — anything else?

顺着这条问 —— 你作为 Anthropic 的 PM,除了 Anthropic 自家工具之外,你的 stack 里还有什么?你刚提了 Slack —— 还有别的吗?

Cat00:42:02

My stack is pretty heavily Claude Code, co-work. Anthropic largely runs on Slack — it's like the core OS of our company.

我的 stack 里 Claude Code 和 co-work 占大头。Anthropic 基本运行在 Slack 上 —— Slack 像是我们公司的 core OS。

Day-to-day, maybe 30% of my time is pushing the boundaries of what co-work can do, so I have a very strong sense of what we're not good at. And I spend a lot of time talking with the model to understand why it makes the mistakes that it does.

日常,大概 30% 时间我在测 co-work 的边界,这样我能非常清楚我们在哪些地方做得还不够;另外大量时间花在和模型对话,搞清它为什么犯它犯的那些错。

We actually have a lot of internal tools that we make. One of the things that Claude Code has really unlocked for our entire company is — it really lowers the barrier to making any custom app you want.

我们内部还做了大量自研工具。Claude Code 给整家公司解锁的一件事是 —— 做任何一个定制 app 的门槛被显著降低了。

We've seen this surge in personalized work software that people are building for custom use cases instead of using tools that don't perfectly fit.

我们看到一波"私人化办公软件"的浪潮 —— 人们为自己很具体的 use case 自己写 app,而不是去凑合用那些不完全合身的现成工具。

Chapter 10

Co-work in the wild

Co-work 在内部的真实用法
43:00 — 48:30 · 销售自动化定制 deck · Slack 的不可替代 · Applied AI 团队是第二大 token 消费者
Lenny00:43:06

I got to hear more. What are some examples? What are things you've built, other people built, that are really popular and useful?

再多说点,有什么例子?你或别人做出来的、特别受欢迎、特别好用的东西?

Cat00:43:12

One of the sales folks on Claude Code realized he was making these repetitive decks over and over. He has this web app he built with examples of the core Claude Code decks we know work well — like a 101, a 201, and Mastering Claude Code.

Claude Code 团队有一个销售意识到,他在反复做差不多的 deck。他做了一个 web app,把已经验证有效的 Claude Code 核心 deck 模板放了进去 —— 101、201、Mastering Claude Code 这种。

He has a way to input specific customer context that pulls from Salesforce, from Gong, from other notes, so we can customize the decks for specific customers.

他做了一个入口,把特定客户的 context 喂进去 —— 从 Salesforce 拉、从 Gong 拉、从其他笔记拉,这样就能为这一家客户定制 deck。

It'll pull out things like — okay, this customer is using Bedrock, or Claude Code for Enterprise, or Console — which affects what features are available.

它会自动抓取关键事实 —— 比如这客户用的是 Bedrock、Claude Code for Enterprise 还是 Console,这决定了哪些功能对他们开放。

It'll pull things like — this customer is concerned about the code review stage of the SDLC — and add a slide about our code review features there.

它会抓:这位客户最在意 SDLC 里的 code review 环节 —— 然后在 deck 里加一张介绍我们 code review 功能的 slide。

It'll pull — this customer needs to be HIPAA compliant or needs XYZ security controls — and we'll add a slide or two about that.

它还会抓:这家客户要做 HIPAA 合规、或者要某几项安全管控 —— 然后加一两张对应的 slide。

If this is a customer on Vertex or Bedrock and doesn't want to use Claude for Enterprise, we'll just take out some of the slides that are Claude-for-Enterprise-only features.

如果客户在 Vertex 或 Bedrock 上、不想用 Claude for Enterprise,它会直接把那些只在 Claude for Enterprise 上才有的功能页删掉。

Normally this is manual work that could take 20-30 minutes — so people either spend that time, or just decide not to do it and use the general deck.

这件事手动做要 20-30 分钟,所以人们要么硬花这个时间,要么干脆放弃定制、直接用通用 deck。

With this it takes a few seconds and you get a tailored deck.

现在几秒钟就能拿到一份量身定做的 deck。

Lenny00:44:42

What's interesting is — Slack is the tool that nobody's trying to create their own version of. Slack just continues to win and is kind of the OS of so many companies.

有意思的是 —— Slack 是那种"几乎没人想自己重做一个"的工具。它一直在赢,几乎是很多公司的 OS。

It's so interesting — people talk about Salesforce, "we don't need SaaS software anymore, we're going to build our own." Slack is a durable tool that nobody wants to compete with and build a better version of.

很有意思 —— 大家说 Salesforce 时是"以后不需要 SaaS 了,自己造",但 Slack 是那种顽强的工具,几乎没人想去做更好的对手。

Cat00:45:00

It's pretty important communications infrastructure, and they do the core task of helping everyone get real-time updates incredibly well.

Slack 是一种很关键的沟通基础设施,他们把"让所有人拿到实时更新"这件事做得特别好。

Lenny00:45:13

People hate on Slack, but it's really great at what it's trying to do, and the most cutting-edge teams are hooked on it.

大家爱吐槽 Slack,但它做它该做的事真的非常好 —— 最前沿的团队都离不开它。

Cat00:45:21

I also love how easy they've made it to customize it. We love making Slack bots — this kind of hackability means we're able to integrate with Slack the way we want to. Really appreciate Slack's work on that.

我也很喜欢它的可定制性。我们特别爱做 Slack bot,这种 hackability 让我们可以按自己想要的方式接进去。这一点非常感谢 Slack。

[Sponsor break · Vanta] —— 中段 sponsor 读稿,介绍 Vanta 的合规自动化(SOC 2 / ISO 27001 / HIPAA 等 35+ 框架)。略过不译。
Lenny00:46:30

You talked about all these different teams and how they use Claude Code and co-work. Other than engineering — I imagine engineering is the biggest token spender — what's the second-place function right now for tokens?

你刚提到很多团队怎么用 Claude Code 和 co-work。除了工程团队 —— 我猜工程是 token 消耗第一 —— 第二大消耗 token 的职能是哪个?

Cat00:47:04

Applied AI is amazing at pushing the boundaries of what Claude Code and co-work can do.

Applied AI 团队特别擅长把 Claude Code 和 co-work 的能力推到极限。

A lot of our applied AI team spends time with our customers helping them adopt our API. Our applied team will, for example, make prototypes on behalf of these customers, which Claude Code makes much faster than it used to be.

他们大量时间在和客户一起,帮客户落地我们的 API。比如他们会替客户做 prototype —— Claude Code 让这件事比以前快得多。

They also have the dual goal of needing to manage a lot of customer comms, customer inbound, historical context, call notes — so they're heavy on co-work and on Claude Code.

他们同时还要管大量客户沟通、inbound、历史 context、电话记录 —— 所以他们既重度用 co-work,又重度用 Claude Code。

Lenny00:47:42

Just to understand applied AI — is that like a forward-deployed engineering sort of role? How would most people describe what the applied AI team is doing?

想确认一下 Applied AI —— 这是不是类似 forward-deployed engineering 的角色?常人会怎么描述他们的工作?

Cat00:47:50

Yeah, it's helping our customers adopt the latest API and model features across their company — both for powering their products and for internal acceleration.

对,他们的活就是帮客户在公司内部把最新的 API 和模型特性落地 —— 既用来驱动客户自己的产品,也用来加速客户内部工作。

Lenny00:48:05

Got it — so it's like customer success / go-to-market-y, kind of like a deploy-engineering sort of role.

明白了 —— 有点像 customer success / GTM,但又是偏 deploy engineering 的形态。

Cat00:48:10

Exactly. It's a very technical go-to-market person.

没错,本质上是非常技术的 GTM 人。

Cat00:48:19

We also see them pushing the boundaries of what co-work can do.

我们也看到他们在把 co-work 的能力推到极限。

A lot of these folks cover multiple customers, and on a high day they can have 5 to 10 customer engagements. What they often use co-work to do is — the night before, they'll ask it to summarize all my customer meetings coming up the next day.

他们一个人要 cover 很多客户,忙的一天可能要开 5-10 个客户 engagement。他们常用的 co-work 工作流是 —— 头一天晚上让 co-work 总结"明天我所有客户会议"。

What are all the things that this customer has asked me for? What's top of mind for them? What are the action items from past meetings?

这位客户问过我什么事?他们最关心的点是什么?过去会议上的 action item 是什么?

Co-work will put together this dossier — this brief — of what they should be aware of going into the next meeting.

co-work 就会把这些拼成一份 dossier、一份 brief,告诉他们去下一场会前该知道哪些事。

Co-work can also research answers — if a customer asked "okay, when is feature X going to launch?" — co-work can help the AI person research through Slack to get the latest ETA and add it to the notes, so during the customer call they have the absolute latest.

co-work 还能帮他们查答案 —— 客户问"功能 X 什么时候发?"co-work 就替这位同事去 Slack 里翻最新的 ETA,加到笔记里,等客户电话进来,他们手上就是最新信息。

These are workflows people are building for themselves and sharing with others on their team.

这些都是个人为自己搭、再分享给团队的 workflow。

Chapter 11

FDE + Claude's character

FDE 工程师 · Claude 的性格
48:30 — 58:45 · token 花费 / 工资比 · 一个月后产品想象力 · 让模型自我反思 · 5 人 taste 圈 · 写 10 个好 eval · Amanda 与 Claude 性格
Lenny00:49:25

Something that comes up a lot recently — token spend exceeding people's salary. Are there numbers floating around Anthropic of how much engineers spend on tokens — a month, a day?

最近一个被反复讨论的话题是 —— token 花费超过工资。Anthropic 内部有没有"工程师每月 / 每天 token 开销"这种数字流传?

Cat00:49:50

It's clear to us that as the models get better, people delegate far more tasks to it and they spend a lot more hours in tools like Claude Code and co-work.

我们看得很清楚:模型越强,人们就越愿意把更多事丢给它,在 Claude Code 和 co-work 里待的时间也就越长。

We do see the token cost per engineer or per any knowledge worker increase every time there's a model jump or a substantial product improvement.

每次有一次模型跃迁、或者有一次重要产品升级,我们都能看到工程师 / 知识工作者的人均 token 成本上升。

I think it's still much lower than what the average engineer salary is, but we see the percentage increasing over time.

这个金额比工程师的平均工资仍然低不少,但占工资的比例确实一直在涨。

Lenny00:50:21

I believe you guys have basically unlimited tokens. You can use as much as you want. Is that right?

我以为你们内部基本上 token 不限,想用多少用多少 —— 是这样吗?

Cat00:50:33

We can use a lot of tokens. Some people do run into limits.

我们能用很多 token。但确实有人会撞到上限。

Lenny00:50:36

Okay, there's a limit. "Boris, shut it down." It's so interesting how many advantages come from having the most advanced model — it's such an interesting flywheel.

原来还是有上限的。"Boris,关掉啊。"用上最强模型带来的优势之多,真挺有意思的 —— 一个很特别的飞轮。

Cat00:50:55

We also believe a lot in empowering our internal teams to build as fast as possible.

我们非常相信"让内部团队尽可能快地造东西"。

We trust that everyone understands how much capacity that serving these models truly costs, and we trust our team to use the tokens responsibly.

我们相信每个人都清楚服务这些模型实际要烧多少容量,也相信团队会负责任地用 token。

It's very frowned upon to waste tokens, but we do trust individuals to make that judgment call.

浪费 token 是被很 frown 的,但我们相信个人能做出判断。

Lenny00:51:14

Coming back to the PM role — what do you think are the emerging skills that PMs need to develop, and that you / AI companies most look for when hiring PMs these days?

回到 PM 这个角色 —— 你觉得当下 PM 应该开始培养、AI 公司最看重的新兴技能是哪些?

Cat00:51:35

I think the hardest skill is being able to define what the product should look like a month from now.

最难的能力,是能画出"产品在一个月后应该长什么样"。

There's a lot of ambiguity in what models are capable of in that timeline and how user behavior will change.

模型一个月后能做什么、用户行为会怎么变,模糊度都极高。

But there are patterns the best PMs can see based on how users are abusing the limits of the existing product. The best PMs can sense that, set a direction, and steadily execute towards it — and change the path if the model capabilities are much better or worse than expected.

但用户怎么"虐"现有产品的边界,这里面有最好的 PM 能看出的模式 —— 他们能凭直觉定一个方向,稳步往那走;如果模型能力比预期强或弱,就修路径。

It is very hard to be the right amount of AGI pilled. Everyone can see this future where the models are extremely smart and can do almost everything, in which case you actually don't need that complicated a product. You can just have a text box where you tell the model what you want — it's so smart that it can add any tool or integration it needs to get the job done. It knows when it's uncertain, and it can ask clarifying questions.

把 AGI 之药吃得"刚刚好"非常难。每个人都能看到那个未来:模型超级聪明,几乎能干所有事 —— 那时候你不需要复杂的产品,只要一个文本框,你告诉它你想要什么,它聪明到能自己加工具、加集成,把活干完;它知道自己不确定的时候,会反问澄清问题。

It's pretty easy to build the product for the super AGI strong model. The hard thing is figuring out for the current model — how do you elicit the maximum capability? How do you help users get onto the golden path? How do you guide users to interact with the model's strengths and patch its weaknesses? This skill is pretty rare.

给超级 AGI 模型设计产品很简单。难的是面对今天这个模型 —— 怎么把它能力榨到最大?怎么把用户引到那条 golden path?怎么引导用户去碰模型的长处、避开它的短板?这种能力非常稀缺。

"It is very hard to be the right amount of AGI pilled."

PM 最难的活,从来不是"为未来超模型规划",而是"为今天这个模型把用户带到能用的那条路上"。

Lenny00:53:19

How do you build that skill? Is it just using each model — basically understanding the limits of each model, having taste into what the model maybe is capable of, what it's great and not great at, where it's changed?

这种能力怎么练?是不是就是把每一代模型都用一遍,搞清楚它的边界,养出对"模型大概能干嘛、强在哪、弱在哪、什么变了"的 taste?

Cat00:53:32

It's spending a ton of time talking and using the model.

就是花大量时间和模型对话、和模型一起干活。

One of the things I really like to do is ask the model to introspect on its own behaviors.

我特别喜欢的一招是 —— 让模型对自己的行为做反思。

Sometimes when I notice the model does something unexpected — for example there are situations where the model will make a front-end change and run tests but not actually use the UI — it's actually pretty useful to ask the model to reflect on why it did this.

有时候我发现模型做了一件出乎意料的事 —— 比如它改了前端、跑了测试、但没真的去点 UI —— 让它反思"自己为什么这么做"非常有用。

Sometimes they'll say — hey, there was something confusing in the system prompt, or I didn't realize that the front-end verification was part of this task, or I delegated the verification to this sub-agent and the sub-agent didn't do the test and I didn't check its work.

有时候它会说:嘿,system prompt 里有一处让我搞糊涂了;或者我没意识到前端验证是这个任务的一部分;或者我把验证交给了一个 sub-agent,而那个 sub-agent 没跑测试,我也没去检查它的工作。

A lot of times just being very curious about why the model made the decision it did will show you what misled it, so you can fix the harness to close this gap.

很多时候,只要你对"模型为什么做了那个决定"保持极度好奇,就会发现是什么误导了它 —— 然后你就知道该怎么改 harness 把这个洞补上。

The other thing that helps is to figure out who the taste users are who you trust the most to give you accurate feedback about the model.

另一件有用的事是 —— 找出你最信任的、能对模型给出准确反馈的"taste 用户"。

Usually there's a handful of people who are much better than others at articulating what makes a specific model or model-harness combination good.

通常会有那么几个人,在表达"某个特定模型 / 模型+harness 组合好在哪里"上比别人强很多。

A lot of people will give you feedback, but not everyone's feedback is as qualified. Finding a group of those five people you trust is really important for getting very fast feedback.

能给你反馈的人很多,但反馈质量参差。找到一组你信任的"五人小队",对拿到快速且高质量的反馈非常关键。

The third thing that is useful but not everyone loves doing is building evals. You don't need to build hundreds of evals for them to be useful. Just building 10 great evals is important for helping the team quantify what the goal is, what their progress towards it is, and what they're missing.

第三件事是很多人不爱做但很有用的事 —— 写 eval。不用写几百个,写 10 个足够好的 eval 就很重要 —— 它能帮团队把"目标是什么、当前距离目标多远、还差什么"量化下来。

Eval is this underappreciated thing that more PMs, more engineers should be working on.

eval 是被严重低估的工作,更多 PM、更多工程师应该花时间做这件事。

Lenny00:55:33

There's this trend — that the future of product management is writing evals — because essentially it's "what does success look like, let me concretely define it, and then we'll know."

有个趋势是 —— 产品经理的未来就是写 eval。本质上 eval 就是把"成功长啥样"具体写下来,然后大家都知道朝哪走。

How much of your time are you spending writing evals?

你自己花多少时间写 eval?

Cat00:55:46

The importance of evals varies based on the feature you're working on or the problem you're trying to solve.

eval 的重要性取决于你在做什么功能、解什么问题。

There are a lot of folks on our team who do spend a lot of time on eval. We have a small pod that collaborates very closely with research to more precisely understand our Claude Code behaviors and what the largest areas of improvement are, and to measure those concretely.

我们团队有不少人在 eval 上花很多时间 —— 有个小 pod 和 research 配合得很紧,专门去精确理解 Claude Code 的行为,定位最有改进空间的几个地方并把它量化。

I personally jump into evals when there's a feature I think needs more product definition. Often the output is — okay, here are five evals I made, this is how you run them, these are the ones that succeed, these are the ones that don't, and this is the prompt I've used to increase the success rate.

我自己会在某个功能"产品定义还不够清"的时候跳进去写 eval。通常产出是:这是我写的 5 个 eval、跑法是这样、哪几个能跑通、哪几个跑不通、我用什么 prompt 把成功率拉了上来。

It varies based on the feature — not every feature needs it, but features such as memory benefit a lot from this.

看具体功能,不是每个功能都需要,但像 memory 这类功能受益就非常大。

Lenny00:56:50

The point you made about people being very good at evaluating models is so interesting. Is there anyone specific you want to shout out who's very good at this?

你说"有些人特别擅长评估模型",这个点很有意思。有具体想点名表扬的人吗?

Cat00:57:00

Two people who I think are incredible at this — one is Amanda, who molds Claude's character. It's such a hard role because the task is so ambiguous. Even coding is easier because you can verify the success, whereas crafting the character requires a very strong sense of conviction in who Claude should be.

两个人我觉得特别强 —— 一个是塑造 Claude 性格的 Amanda。这个角色非常难,因为任务本身极度模糊。写代码反而容易,因为成功能被验证;而塑造一个 character 需要你对"Claude 应该是谁"有非常强烈的信念。

She has an incredible ability to not only mold the character, but to articulate what the goals are, what's successful and what's not.

她不仅能把这个 character 雕出来,还能把"目标是什么、成功是什么、不成功是什么"清晰地讲出来。

The other group of people I really trust is the Claude Code team. We often have team lunches, and whenever there's a new model we're testing, one of the fastest ways for us to get feedback is to go to every single person at the team lunch and just be like — "Hey, what's your vibe on the model?"

另一群我特别信任的人是 Claude Code 团队。我们经常一起吃午饭。只要在测一个新模型,我们最快的反馈方式就是 —— 在午饭桌上挨个问:"你对这个模型的 vibe 是什么?"

Often we'll get feedback like — "Okay, this model is not fully explaining its thinking, it's too abrupt." Or — "this model loves writing a ton of memories but we're not sure if the memories are high quality." Or — "this model loves to test itself, which is great." Or — "this model isn't testing itself enough."

经常拿到的反馈是 —— "这个模型不会把自己的思路解释清楚,太突兀";或者 —— "这个模型很爱写一堆 memory,但我们不太确定 memory 质量";或者 —— "这个模型特别爱给自己加测试,挺好";或者 —— "这个模型给自己加的测试不够多"。

That informs what data we look at to verify if it's a larger pattern. We have a ton of data, but it is very hard to extract insights — so the feedback from this group helps us inform what hypotheses to test, and we can then extract data to verify.

这些反馈会指导我们去看哪些数据,验证是否是大范围的规律。我们的数据非常多,但从中抽出 insight 非常难 —— 这群人的反馈帮我们决定"要去验证哪个假设",然后我们再到数据里去验。

Lenny00:58:45

This point about the character of Claude — I had Ben Mann on the podcast (co-founder), and he talked about how the character / the constitution of Claude is such an important part of Claude.

关于 Claude 的性格 —— 我请过 Ben Mann(联合创始人)上播客,他讲过 Claude 的 character / constitution 是 Claude 极其重要的一部分。

I didn't realize until afterwards just like — with OpenClaw, one of the reasons people are sad is the personality of Claude. Because Claude's personality is so good and fun and interesting, unlike other models. The way he put it is — "the personality is what makes Claude so good at so many things."

直到后来我才意识到 —— OpenClaw 那件事让大家难过的原因之一就是 Claude 的人格。因为 Claude 的人格特别好、特别有趣,跟别的模型不一样。Ben 当时的说法是:"正是这个性格让 Claude 在很多事情上做得特别好。"

It feels like a trivial side thing — okay, it's going to be funny and interesting and talk in a fun way — but it's so core to the success of Claude.

表面上像是个无足轻重的"附带特性" —— 反正它幽默、有趣、说话好玩 —— 但其实是 Claude 成功的核心。

Is there anything you can share about why the character / personality is so key?

关于"为什么这个性格 / 人格如此关键",你有什么可以补充的吗?

Chapter 12

Model eats harness · the vision

模型吃掉 harness · 未来的 Claude Code
58:45 — 1:09:00 · 性格让模型好用 · 砍 to-do list · 每个任务先 task,然后 6 个并行,再到 50/100 个 · "AI 是杠杆"
Cat00:59:34

When you reflect on everyone you've worked with, there are some people where you're like — I really like their energy, I really like their vibe.

回想你共事过的所有人,总有一些会让你觉得 —— 我喜欢他的能量,我喜欢他的 vibe。

When people think about Claude and Claude Code, this is one of the things they bring up the most — they really love that Claude is lighthearted and fun, but also extremely competent at your task.

大家提到 Claude 和 Claude Code 时,最常说的就是 —— 他们特别喜欢 Claude 既轻松好玩,又对你的任务做得极其到位。

People really like that Claude's low-ego. If you tell it "hey, you did this thing wrong," it's like truly sorry — "oh shoot, thanks for telling me, let me fix it, let's work together."

大家也很喜欢 Claude low-ego。你告诉它"嘿,这件事你做错了",它会真心地说"哎呀,谢谢你告诉我,我来修,我们一起搞定"。

It's also very positive. If you're feeling like — oh, this is an insurmountable task, I don't know how to get started — Claude is like, "Okay, it's okay. These are the steps that I think we should take. Do you want me to get started on it for you?"

它也非常正向。当你觉得"这事我搞不定,完全不知道从哪下手"的时候,Claude 会说:"没事,这是我建议我们走的几步。要我先帮你起头吗?"

Part of what makes a great co-worker is this positivity, this bias towards action, this ability to give you earnest feedback — not just agreeing with every single thing you say.

一个好同事所需的特质:这种正向、bias towards action,以及"能给你真诚反馈,而不是无脑同意你说的每一句"。

We try to imbue this into Claude because we think it makes it a lot more enjoyable to work with.

我们尝试把这种特质灌进 Claude,因为这样和它一起干活会愉快得多。

Lenny01:00:45

There's something I want to come back to — you talked about how when new models come out, you often have to revisit things you've built. That's so interesting and so frustrating maybe — "oh god damn it we shipped this thing, now we have to rethink it."

我想回到一个点 —— 你刚说每次新模型出来,你们常常要回头重新审视已有的东西。这件事既有意思又有点崩溃 —— "我擦,这功能刚发了,现在又得重新想。"

Talk about how often you have to come back with a new model and be like — okay, we have to redo this product we launched a few months ago.

讲讲新模型出来后,你们多久会一次回头说 —— 好,这个几个月前发的产品,得重做。

Cat01:01:03

A lot of the changes we make with a new model is removing features that are no longer needed.

新模型来了,我们大部分改动其实是 —— 删掉那些已经用不上的功能。

A lot of times we add features to the product as a crutch for the model because it's not naturally doing it itself.

很多功能其实是当时给模型加的拐杖,因为模型自己做不到。

The classic example is a to-do list. When we first launched Claude Code, people would ask it to do these large refactors, and Claude Code would say "okay cool, I need to change these 20 call sites" and would go and change five of them and then stop.

最经典的例子是 to-do list。Claude Code 刚发的时候,有人让它做大型 refactor,它会说"好,这 20 个调用点都要改"—— 然后改了 5 个就停了。

We were like — okay, how do we force it to remember to get every single one of these 20? Sid on our team was like — okay, what would a human do? A human would make a list of everything they need to change. Similar to how in VS Code you would look up all the call sites and it would be a list on the left side, and you'd go through them one by one. How do we give Claude a tool like this?

我们当时想:怎么逼它记得这 20 个一个不落?团队里的 Sid 想:人会怎么做?人会把所有要改的地方列一张清单 —— 像 VS Code 里查所有 call site 那样,左边是个列表,挨个改下去。怎么给 Claude 也搭一个这种工具?

So he added a to-do list, and we found that with that, Claude was actually able to fix all 20 call sites.

于是他加了一个 to-do list,我们发现:加上之后 Claude 真的能把 20 个调用点都改掉。

But with Opus 4 and later models, we realized we didn't need to force it to use this to-do list. It would naturally use it itself.

但到了 Opus 4 以及之后的模型,我们发现根本不需要逼它用 to-do list,它自己就会用。

For the earlier models, we had to keep reminding it — "Hey, did you finish everything on the to-do list? You can't finish until you're done with everything." For the later models, without prompting, it just naturally thinks to do everything on the to-do list.

早期模型时我们要不停提醒:"嘿,to-do list 上的事都做完了吗?没全完成不能算结束。"到后来的模型,不用任何提示,它自己就会把列表上的事都做完。

These days, the to-do list is still nice as a user — you can more clearly see what Claude is working on. But honestly, it's such a deemphasized part of the product right now that the model may use it, the model may not. It's really not necessary for it to make thorough changes anymore.

今天,to-do list 作为用户体验仍然挺好 —— 你能更清楚地看到 Claude 在干什么。但说实话,它在产品里已经被严重弱化了,模型用不用都行,完全不再是它做完整改动的必要条件。

Lenny01:02:44

I forget who said this on the podcast — "the model will eat your harness for breakfast." What I'm hearing is — you remove things over time that you've added on top of the model where it was not operating the way you wanted. And as the models get smarter, it just becomes simpler and simpler for it to do the thing you want.

我忘了哪一期里有人说过 —— "the model will eat your harness for breakfast"。我现在听到的是 —— 你们在时间里逐步删掉那些当初为了"模型行为不达标"而加在它头上的东西。模型越聪明,它做你想要的事就越简单。

"The model will eat your harness for breakfast."

每次新模型上,Anthropic 都会把整个 system prompt 翻一遍,问每一段:这条还需要吗?能砍就砍。to-do list 是这种"加完该剥掉"的最清晰例子。

Cat01:03:04

We can remove a lot of prompting interventions every time the model gets smarter. We actually do this every time we launch a model.

每次模型变聪明,我们都能砍掉一大批 prompting 干预。每次模型发布,我们都会做这件事。

We read through the entire system prompt and we reflect on, okay — for each of these sections, does the model really need this reminder anymore? And if not, we'll remove it.

我们会把整个 system prompt 通读一遍,逐段反思:这一段提醒,模型还需要吗?不需要,就删掉。

The most exciting thing that new models unlock though is entirely new features.

不过新模型最让人兴奋的,是它解锁全新的功能。

There are a lot of features that we've been testing with prior models and the accuracy wasn't high enough for us to want to launch them.

有很多功能我们用之前的模型试过,但准确率不够高,所以没上。

One example is code review. We tried to build a code review product a few times, and we launched simpler versions in the past — the /code-review command. It was only with the most recent models that we felt — okay, this code review is so good that our engineering team relies on this code review to pass before we merge PRs.

一个例子是 code review。我们试过几次做 code review 产品,过去也发过更简版本 —— /code-review 命令。直到最新一代模型,我们才觉得 —— OK,这个 code review 已经好到我们工程团队可以把"过了 review 才 merge PR"作为规则了。

We've always dreamed of Claude being able to be a reliable code reviewer that can confidently catch the majority of bugs. It was only with Opus 4.5 / 4.6 and Sonnet 4.6 that we felt — okay, we are now able to run multiple code-review agents simultaneously to traverse the entirety of the codebase, and to synthesize a set of real issues that an engineer needs to address before merge.

我们一直梦想 Claude 能做一个可靠的 code reviewer,能放心地抓出大多数 bug。直到 Opus 4.5 / 4.6 加 Sonnet 4.6 这一波,我们才觉得 —— 我们现在可以并行跑多个 code-review agent 通读整个 codebase,然后把"工程师 merge 前必须处理的真实问题"整合输出。

Lenny01:04:39

This is another trend very common on this podcast — build something that will possibly be possible in the next six months. Be at the edge of what's working, and then it'll catch up and become an amazing product, and you'll be ahead of everyone.

这又是一个本播客常出现的规律 —— 提前做一个"未来六个月可能 work"的东西。停在能跑通的边缘,等模型追上,它就变成一个出色的产品,而你已经领先所有人。

Cat01:04:52

Yeah, exactly. It's pretty important to build products that don't necessarily work yet, so that you know — okay, what is missing for this product to work? Then with the newest model, you can swap it into the prototype you've already made and see — okay, does this new model close that gap?

没错。做"目前还跑不通"的产品本身就很重要 —— 它能让你看清"要让它跑通还差什么"。新模型来了之后,你直接把它换进已经写好的 prototype 里,看看这次能不能补上那个缺口。

Lenny01:05:12

How much can you speak to where things are going with Claude and co-work — the vision?

关于 Claude 和 co-work 的未来 —— 你能聊多少?这个 vision 大概是什么样?

I imagine you don't want to give away too much, but it feels like there's all these awesome features being added — dispatch control from phone, mobile app, all these things. What's a way to understand the vision long term?

我猜你也不想剧透太多,但你们一直在加各种酷功能 —— 在手机上调度、mobile app 等等。怎么从长期理解你们的 vision?

Cat01:05:32

We think about this in terms of building blocks.

我们用 building blocks 来想这件事。

For both Claude Code and co-work, the core building block is making individual tasks successful — you produce some output, you give it a clear prompt description, is it able to consistently produce acceptable output that you can either merge or share with your colleagues or external audience? The task is the core building block.

对 Claude Code 和 co-work 来说,核心积木都是"让一个任务跑成功" —— 你想要某种输出,你写清楚 prompt,它能不能稳定地给你一个可以 merge、可以发给同事或外部受众的产出?任务,是最底层的积木。

As the models get smarter, the task success rate gets a lot higher. Then we see people moving towards doing multiple tasks at the same time. So multi-coding was a big thing towards the end of 2025 and it's only increased since then.

模型越聪明,任务成功率越高。接着我们看到人们开始并行做多个任务 —— multi-coding 在 2025 年底成了一个大趋势,从那之后只增不减。

We see this as — okay, great, one task works, and now you can do six tasks at a time. As the models get even smarter, the way we are extrapolating this is — okay, next maybe you're going to run 50 Claudes at a time or hundreds of Claudes at a time.

我们的看法是:好,一个任务跑通了,现在能并行 6 个;模型再变聪明,我们外推:下一步可能是同时跑 50 个、甚至上百个 Claude。

What is the infrastructure we need to build to enable that? At that point you're probably not going to run everything locally on your machine anymore — there's just not enough RAM. So we're thinking — how do we make it easier for you to manage all these? These will probably run remotely.

为了支撑那种规模,我们需要什么基础设施?到那时,你大概率不会在本机上跑所有东西 —— 内存根本不够。所以我们在想:怎么让你更容易管这么多任务?它们大概率会远程跑。

How do we build the interface so that you, as a human, know which tasks you need to look into? How do we make sure the agent is fully verifying work — so that when you look at a task and it says it's done, you can very quickly verify and fully trust that it is done to your spec?

界面怎么设计,才能让你这个人类一眼知道"我现在要看哪几个任务"?怎么保证 agent 是真的在自我验证 —— 这样当任务标"完成"的时候,你能快速核对、完全信得过它真的按你的要求做完了?

How do we make sure this process is self-improving? When you do see a task that isn't done to your liking, you can give it feedback and the model will know for every future run to incorporate that feedback so it never makes that mistake again.

怎么让整个流程能自我提升?当你看到一个任务做得不合你意,你给它反馈,以后每次跑它都能把这条反馈吃进去,再也不犯同样的错。

This is the progression we're bringing our users along for.

这是我们正在带着用户走的递进路线。

Lenny01:07:23

There's a lot of people listening — product managers, founders, cross-functional folks — there's a lot of worry about their roles and the future of their careers.

听众里很多 PM、创始人、跨职能岗位的人 —— 大家都很焦虑自己的角色和职业未来。

What advice would you have for people to not just survive this transition to a very AI-driven world, but to thrive? What do people need to hear, need to be doing?

你会给大家什么建议,不只是熬过这次转型,还要在这个高度 AI 驱动的世界里活得好?有什么是大家需要听到、需要去做的?

Cat01:07:49

I think AI gives everybody a ton more leverage than they used to.

AI 给每个人的杠杆,比以前大得多。

Anytime you realize that you're doing some manual task multiple times, think about how you can use Claude Code, co-work, or other AI tools to automate that for you.

只要你发现自己在重复做某件手工活,就想想怎么用 Claude Code、co-work 或者别的 AI 工具替你自动化掉。

Most people have creative parts of their job that they absolutely love, and tedious parts that they really hate doing. The beauty of AI is that it can do those tedious parts for you — it can learn from every time you've done that manual task and generalize, then run it automatically — so you can focus on the creative parts.

大多数人的工作里既有他们超爱的创造性部分,也有他们讨厌的琐碎部分。AI 美在哪里 —— 它可以替你做琐碎部分;它可以从你每次手工做这件事的方式里学,泛化,然后自动跑 —— 让你可以专注于创造性部分。

My push for people is — figure out the repetitive parts that you can pass to Claude. Iterate on those automations until the success rate is very high. Then focus on what more you can be doing for your team, your product, your company that people haven't had the bandwidth to pick up so far.

我对大家的建议是 —— 找出可以丢给 Claude 的那些重复活,把这些自动化迭代到成功率很高;然后把腾出来的精力,投到你团队 / 产品 / 公司里那些"一直没人有时间做"的事上。

What's that pet project that you always thought the company should do but you've never had bandwidth to do? If AI can take care of the grunt work, then you have this extra 20% time you didn't have before.

那个你一直觉得公司应该做、但你从没腾出过手做的小项目是什么?如果 AI 能接走 grunt work,你就能多出 20% 的时间。

My push is to lean into these tools, hand off the work that you're not excited to do, figure out how it can accelerate you, and as a result, you'll be able to do so much more.

所以我的建议是 —— 拥抱这些工具,把不让你兴奋的工作交出去,搞清它怎么帮你提速;结果就是你能做的事会比以前多得多。

Chapter 13

Find real problems, not setups

找真问题 · 别再炫配置
1:09:00 — 1:15:30 · 95% 不是自动化 · "your setup is so awesome — but what are you building?" · Karpathy 二分法 · Action-based 时代
Lenny01:09:19

Something core to what you just shared — and I fully agree with — is "find problems to solve with AI."

你刚说的有一个核心点我特别同意 —— 找用 AI 来解决的真问题。

There's all this potential in what these tools can do. For a lot of people, the hardest part is "what should I actually do?"

这些工具潜力很大。对很多人来说最难的就是"我到底该做什么?"

What you're saying is — pay attention to things that you are doing constantly, that you can automate. Pay attention to ideas that have been floating around that you haven't had time to do. It's basically — solve a problem for yourself.

你的意思是 —— 留心你一直在做、可以自动化的事;留心一直飘在你脑子里、你没时间做的想法。归根到底就是:解决一个你自己的问题。

Cat01:09:45

Exactly. I would also push listeners towards focusing on bringing your automations from "okay, this is a cool concept" to "hey, this actually works 100% of the time."

没错。我还想推大家一把 —— 把自己的自动化从"这是个不错的 concept"推到"它真的 100% 能跑通"。

Sometimes I see users trying to automate something, getting it to 90 / 95% accuracy, and then giving up on it. If an automation doesn't work 100% of the time, it's not really an automation.

有时候我看到用户把某个自动化做到 90 / 95% 的准确率,然后就放弃了。如果一个自动化不能 100% 跑通,那其实算不上是自动化。

That last 5 to 10% does take more time. Building the automation is often a lot slower than you doing it yourself.

最后那 5% 到 10% 确实费时间。把自动化做出来,常常比你自己手动做一次要慢得多。

I would encourage listeners to put in that time — scope some automation that you really want to get to 100%. Put in the elbow grease to teach Claude your preferences, give it feedback so it can improve its skill, so that it can get to that 100%. Then really, you'll be able to rely on it.

我鼓励大家舍得花这个时间 —— 锁定一个你真的想做到 100% 的自动化;下功夫教 Claude 你的偏好、给它反馈、让它的技能改进到 100%。到那时,你才真的能依赖它。

There's just not much value in a 95%-there automation.

95% 那种"差一点"的自动化,价值真的有限。

Lenny01:10:44

I am super guilty of that. This is really good advice for me.

我自己就深陷这一条。这建议对我太对症。

Cat01:10:48

I am guilty of this too. I've been teaching co-work to try to get me to inbox zero for Gmail, and it has been very time consuming, and definitely not there yet.

我自己也犯。我一直在教 co-work 帮我把 Gmail 清到 inbox zero,极其费时,目前也远远没到位。

Lenny01:11:02

Funny enough, that's exactly where my mind goes. I have this workflow where every email I get, it looks for things that are spammy — like "hey, can I come on your podcast?" — and categorizes it into a folder called "spammy."

巧了,我也想到这件事。我搭了一个 workflow,每封邮件来都会被检查是不是 spammy —— 比如那种"嘿能不能上你的播客?"—— 然后被自动归到一个叫 spammy 的文件夹。

It's 95% great, but then there's like — "oh wow, I missed an email because it went in there." This is a good push for me to get it to perfect.

95% 工作得很好,但偶尔会"哎呀,我漏掉一封因为它进了那个文件夹"。你这话正好推我一把,去把它做到 100%。

Cat01:11:29

We're also working on making the flow for customizing these commands a lot easier.

我们也在让自定义这种命令的流程变得更顺。

Right now I think you have to know too many concepts. You have to know to define a skill. You have to know to use the skill and give it feedback. Then you have to know to tell co-work to update the skill based on all the feedback. Then you have to know where to read the skill to make sure the feedback was incorporated the way you want.

现在你要懂的概念太多 —— 你得知道怎么定义一个 skill、怎么用这个 skill 并给它反馈、怎么告诉 co-work 把这些反馈合进 skill,还得知道去哪里读这个 skill 确认反馈的确按你想要的方式合进去了。

It's our job to make this flow really seamless so that it doesn't feel painful to do.

把这条 flow 变得真的丝滑、让人做起来不痛苦,是我们的活。

Lenny01:11:57

Is there anything else, Cat, you wanted to share — anything you wanted to double down on that we haven't already touched on, before we get to our very exciting lightning round?

Cat,正式进入闪电问答之前,还有什么你想说的、你想重申的、我们没聊到的?

Cat01:12:10

I see a lot of people playing around with AI, building prototype apps and tinkering with workflows.

我看到很多人在玩 AI,做 prototype app、捣鼓 workflow。

I would really push people towards building apps that you're actually using every single day — because only through that usage are you actually getting the value.

我特别想推大家去做"你真的每天都在用的 app" —— 只有真用,才真正在获得价值。

If you build a prototype app that isn't helping you get more done, then the AI isn't really adding value to your day.

如果你做的只是一个不能让你产出更多的 prototype,那 AI 其实没真的给你的一天加什么。

Lenny01:12:38

There's only so much you learn from "okay, I just one-shotted something — oh, that's cool" — and then you never come back to it.

那种"好,我一发就成 —— 噢,挺酷的"然后你再也不回头看的东西,你能学到的就那么多。

Cat01:12:45

And you're not getting much leverage from it.

而且你从中拿到的杠杆也很有限。

Cat01:12:49

I also think there's a lot of people who spend a lot of time customizing their workflow.

我还看到很多人在自定义 workflow 上花了海量时间。

There are two ends of the spectrum — one is people who never customize or never build automations. But there's the polar opposite end of people who obsess about customizing their tool, adding a ton of skills and MCPs and these workflow improvements — and I think sometimes that can even distract from your core goal of launching some product or building some feature.

这件事有两个极端 —— 一头是从不定制、从不写自动化的人;另一头是对工具定制极度上头、加一堆 skill / MCP / workflow 微优化的人 —— 我觉得有时候这种沉迷甚至会让你偏离"发布产品 / 做出功能"的真正目标。

There's a lot of fun in customizing — and we definitely want to make our products very hackable so it can work really well for you. But there is a limit to how much it's useful.

定制是真的有乐趣 —— 我们也确实想让产品非常 hackable,让它能为你专属地工作得好。但有一个边际效用的边界。

There's a camp of people who maybe spend so much time customizing that they're not sleeping and not doing the core task they originally set out to do.

有一群人在定制上花到不睡觉,反而把当初定下要做的核心任务给丢了。

Lenny01:13:41

I see a lot of that on Twitter — "look at my setup, it's out of control, it's so optimized." Then — what are you actually building?

我在 Twitter 上看到很多这种 —— "看我这套 setup,炸裂、超优化"。然后 —— 你到底在 build 什么?

"No, but my setup is so awesome, it gets so much done."

"不不,我的 setup 真的太牛,产出很多。"

Cat01:13:52

I think the simple setups actually work better.

我觉得简单 setup 反而更管用。

"The simple setups actually work better."

"看我配置多牛"是反面教材。Cat 的判断:简单 setup + 真的每天在解决某件事的小 app,是 AI 时代最被低估的杠杆。

Lenny01:13:59

There's this Karpathy tweet that came out yesterday, where he talked about this divide that's interesting between people who tried ChatGPT / Claude back in the day, were like "nah, this is terrible" and gave up — they're so cynical, "no way, it's not actually that big of a deal" — and then there's people who are using it to code, who see the full intense power of it.

昨天 Karpathy 发了一条 tweet,讲一个挺有意思的分化 —— 一群人当年试过 ChatGPT / Claude,觉得"啊这玩意儿不行",就放弃了,然后变得很犬儒"AI 没那么大不了";另一群人在用它写代码,亲眼看见它真正炸裂的力量。

People on both sides don't understand the other side and how they see the world. So your advice is really good — actually use it for real things and see how good it actually has gotten.

两边都不理解对方眼中的世界。所以你的建议正好对症 —— 用它去做你真实的事,你才会看到它到底变得有多好。

Cat01:14:38

The big shift is that the 2024 generation of products were chat-based, and the Claude Code generation of products is action-based.

最大的转变是 —— 2024 那一代产品是基于 chat 的;Claude Code 这一代产品是基于 action 的。

The big "aha" moment people have is when Claude can just do things on your behalf. It is an amazing feeling to know that the agent is capable of doing so much more than telling you what to do — the agent can actually just do it itself. When people feel that, that's the eye-opening moment.

人们的"aha 时刻"是当 Claude 能直接替你把事干掉。意识到 agent 不只是告诉你"该怎么做",而是能直接去做 —— 这种感觉太牛。当一个人亲身体会到这一刻,他的眼睛就打开了。

Lenny01:15:10

Shout out to the Claude Code Chrome extension — you can just watch it doing stuff and you'd be like — "Fill out this form for me." "Alright, here I go."

必须给 Claude Code 的 Chrome extension 一个 shout out —— 你可以坐在那看着它干活,跟它说"帮我填这个表",它就"好,出发"。

Cat01:15:18

Exactly.

没错。

Lenny01:15:24

Anything else before we get to our very exciting lightning round?

进入闪电问答之前,还有什么要说的?

Cat01:15:26

No, let's do it.

没啦,开始吧。

Chapter 14

Lightning round + AGI life

闪电问答 · AGI 之后做什么
1:15:30 — 1:25:30 · 三本书 · Free Solo · Waymo · "Just do things" · AGI 之后的攀岩与 BD
Lenny01:15:24

Let's do it. Cat, I've got five questions for you. Welcome to the lightning round. Are you ready?

来吧。Cat,我有五个问题要问你 —— 欢迎来到闪电问答。准备好了吗?

Cat01:15:32

I'm ready.

准备好了。

Lenny01:15:34

First question — what are two or three books that you find yourself recommending most to other people?

第一题 —— 有哪两三本书是你最常推荐给别人的?

Cat01:15:38

I really like How Asia Works. It's a story about economic development and the policies and governments that make long-lasting successful economies.

我很喜欢 How Asia Works。这本书讲经济发展,以及哪些政策和政府制度能造就长期成功的经济体。

The other book I'm really into is The Technology Trap. It's about the past few technology revolutions — the industrial revolution, the computer revolution — and how this has affected workers. The reason I really like it is because there's a lot we can learn from history to make sure that this transition goes well.

另一本我很喜欢的是 The Technology Trap。讲的是过去几次技术革命 —— 工业革命、计算机革命 —— 对劳动者的影响。我喜欢它的原因是,我们能从历史里学到很多,让这一次的转型走得更好。

Maybe on a fun note — I really like The Paper Menagerie. It's a book of short stories about coming of age and AI and self-discovery.

轻松一点的一本是 The Paper Menagerie(《纸animals园》),一本短篇集,讲成长、AI 和自我发现。

Lenny01:16:30

Favorite recent movie or TV show you have really enjoyed?

最近你看得很爽的电影或剧?

Cat01:16:34

I really like Drive to Survive. There's no deeper meaning — there's just something very satisfying about people being so obsessed with a singular engineering goal, and the purity of their pursuit.

我很喜欢 Drive to Survive(《极速求生》)。没什么深刻含义 —— 就是看一群人为单一的工程目标着魔、那种纯粹的追逐感,特别让人满足。

I also really love Free Solo, about Alex Honnold climbing El Capitan without a harness. Similarly — it's such a pure achievement to be able to climb this extremely challenging, dangerous route, and to have the mental focus to do it knowing that if you make a single mistake, you die.

我还非常喜欢 Free Solo —— Alex Honnold 不带保护绳爬 El Capitan。同样的纯粹 —— 能完成这条极难极危险的路线,同时还能保持那种"一次犯错就会死"的心智专注,这本身就是一种纯粹的成就。

Lenny01:17:17

It's insane. Yeah, that movie is out of control. And it's interesting how these relate in some way to the work you do.

太疯狂了。那部片真是炸裂。有意思的是,它们某种程度上和你做的工作其实是相通的。

Cat01:17:22

I actually am a rock climber. I first watched Free Solo before I climbed rocks, and I thought it was impressive — I didn't understand how impressive it was.

我自己也攀岩。我第一次看 Free Solo 是在我开始攀岩之前 —— 当时只觉得"挺厉害",其实根本不知道有多厉害。

It's one of the rare movies where the more you know about it, the more you're blown away by how insane this is. The kinds of moves he's doing on the wall are things I don't think I will ever be able to do in my lifetime, even if it were set in a gym one foot off the ground — with a rope.

这是少见的"你越懂越觉得疯"的电影。他在墙上做的那些动作,放进室内攀岩馆、离地一英尺、还系着保护绳,我这辈子都不一定做得到。

Lenny01:17:50

Did you see the documentary on that other guy, the younger one that went on like ice mountain?

你看过另一位年轻攀冰者那部纪录片吗?

Cat01:17:54

I did. That one was very sad.

看了。那一部很难过。

Lenny01:17:56

But that was wild. Okay — favorite product you recently discovered that you really love?

那部确实疯狂。下一题 —— 最近你发现并爱上的产品是什么?

Cat01:18:00

The product that has most changed my life outside of Claude products is probably Waymo. I'm a diehard Waymo user — use it twice a day, get to and from work.

除 Claude 自家产品之外,最改变我生活的产品大概是 Waymo。我是死忠用户 —— 每天用两次,通勤来回都靠它。

Two things I really like — one, I don't feel bad if a Waymo is waiting for me. So I feel less pressure to be right at the curbside the moment it arrives.

我喜欢它两点。第一,Waymo 等我我不会内疚 —— 所以"它一到我必须立刻在路边"这种压力没有了。

The second thing is it lets me be a bit more productive. When I'm in the car with another human, I typically try not to do work calls — I feel a little rude if I'm on my laptop the whole time.

第二,它让我更能产出。如果车里有别的人(司机),我一般不愿意接工作电话,觉得"我一直挂在 laptop 上"有点失礼。

But one thing I really appreciate about Waymo is — I can call into a work call, I'm not worried about someone overhearing me, not worried about "is this rude, am I talking too loud, do I need to ask someone to change the music?" This has given me back like 30 minutes every day.

而 Waymo 让我可以放心地接工作电话 —— 不用担心别人听见,不用想"我是不是失礼了、声音是不是太大、要不要请别人调小音乐"。这件事每天给我多挤出了大概 30 分钟。

Lenny01:18:55

All these second-order effects of technology — it's so interesting.

技术的这些二阶效应太有意思了。

Cat01:18:59

I always thought Waymo needed to be priced lower than Uber and Lyft to succeed, but actually I'm very happy to pay a 2x premium for it.

我以前一直以为 Waymo 必须比 Uber 和 Lyft 便宜才有戏,实际上我很愿意为它多付一倍。

Lenny01:19:30

Two more questions. Do you have a favorite life motto that you often come back to in work or in life?

还有两题。你工作或生活里反复回去的座右铭是哪一句?

Cat01:19:35

Just do things.

Just do things —— 就把事干掉。

Cat01:19:38

There's a lot of value in first-principles thinking — if you know what you're optimizing for and have strong first principles, then you can normally deduce what the right course of action is, and clearly articulate it to all the stakeholders. Then you should just do it.

first-principles thinking 真的很值钱 —— 你只要知道自己在优化什么、把核心 first principles 抓得很稳,通常就能推出正确的行动路径,把它清楚地讲给所有 stakeholder,然后就直接干。

I think jobs are fake. If you understand the constraints, you can figure out what you can do, and then just try to do it quickly, learn from the mistakes, and apologize or fix them if you did something wrong.

我觉得"工作"这件事是假的。你只要搞懂约束条件,就能想出自己能做什么,然后快点去做、从错误里学,做错了道歉或者修掉。

"Jobs are fake. If you understand the constraints, you can figure out what you can do."

这是 Cat 在 Stripe 早期 20 人时被打通的 mental model。她把它一直带到了 Anthropic —— "Just do things"是这套 mental model 的执行版口号。

Lenny01:20:08

"You could just do things" — whoever said that.

"你其实可以直接去做"—— 不知道是谁先说的,但这话太对了。

Cat01:20:10

It's liberating actually to tell people this.

把这句话说给别人听,其实是很解放的。

In a lot of companies, roles are very strictly defined — okay, this is what the PM does, this is what the designer does, this is what the engineer does. Even team scopes are very rigidly defined — like, this corner of the codebase we touch, and this corner we're not allowed to touch.

很多公司里角色定义非常死 —— PM 做什么、设计师做什么、工程师做什么。连团队 scope 也僵硬 —— 这块 codebase 我们能碰、那块不能碰。

What "just do things" lets people do is they feel empowered to make these decisions, empowered to operate across team boundaries — just to get something done.

"Just do things"让大家感到自己有权做决定、有权跨越团队边界做事 —— 只为把事干掉。

Lenny01:20:38

That feels like a big important skill. People call it agency. Just do the things.

这听起来是一个非常重要的能力,大家把它叫 agency —— 把事干掉。

Cat01:20:46

Bias towards action. All these ways of describing — just don't wait for permission.

bias towards action。本质上是一句话:别等许可。

This is my favorite reason to work at a startup at some point in your life. One thing that was very life-changing for me was working at Stripe when we were 20 people. There was no process, and we had really big problems we needed to solve.

这也是我"建议你这辈子至少在创业公司里待一段时间"的最大理由。我当年加入 Stripe 时只有 20 个人,那是非常改变我的一段经历 —— 没有流程,但要解决的问题非常大。

I really appreciate Alex and the rest of the team for empowering me and the rest of the team to figure things out without any boundaries for what sales is supposed to do, what ops is supposed to do, what engineering is supposed to do.

我特别感谢 Alex 和当时的团队,他们把我们都赋能了 —— 没人规定 sales 应该干嘛、ops 应该干嘛、工程应该干嘛 —— 我们就自己摸索。

You have all the tools at your disposal. You have some ambitious, hairy problem statement, and you can do whatever you need to get to a good solution.

你有所有工具可用;你面对的是一个野心很大、很乱的问题陈述;然后你可以做任何你需要做的事,直到把它推到一个好答案。

Lenny01:21:28

You almost need that experience to build that skill. A lot of people go through school or college — "do the thing we tell you to do and you'll get a good grade" — and you have to unlearn that, and just do the thing that needs to be done, even if people think it's dumb, because you think it's the right thing.

几乎要有这种经历你才能 build 出这个能力。很多人在学校 / 大学里被训练成"做老师让你做的事,就能拿好分"。你得把这套 unlearn 掉,做你认为该做的事 —— 哪怕别人觉得蠢。

Cat01:21:46

Yeah. Exactly.

对,完全是。

Lenny01:21:47

Two more final questions. One — when Claude thinks, there are all these I-don't-know-if-you-call-them-verbs. What's the term for these things?

最后两个问题。一个是 —— Claude 思考的时候,会冒出一堆"动词"之类的东西,这些有官方名字吗?

Cat01:21:55

Thinking words.

叫 thinking words。

Lenny01:21:56

Thinking words. Interestingly, these all leaked in the source code. Do you have a favorite thinking word?

thinking words。有意思的是,源代码泄漏的时候这些词都出来了。你有最喜欢的一个吗?

Cat01:22:03

I really like "manifesting." It's also the sticker that I have on my favorite [laptop / item].

我超爱 "manifesting"(显化)。这也是我最爱那一台机器上的贴纸。

Lenny01:22:10

Clearly the winner. Okay, final question — asked Boris this too. With AGI potentially arriving in our lifetime, when you potentially don't have to work, what are you going to do? What are you going to do with all your time?

"manifesting" 完胜。最后一题 —— 这道我也问过 Boris。如果 AGI 真的在我们这一代到来,你可能不再需要工作 —— 你打算干嘛?你会怎么用你的时间?

Cat01:22:23

I think it will take a long time for AGI to diffuse across society. So the immediate thing is actually just helping bring the world along.

我觉得 AGI 在整个社会上扩散会是一个漫长过程。所以眼下的事其实是 —— 帮整个世界跟上。

My non-serious answer for after this happens is — I'll probably just do a lot of rock climbing. I'll probably move to Fontainebleau and just live amongst 10,000 boulders and climb for a bit.

不太正经的回答是 —— 我大概会狂去攀岩。可能会搬到 Fontainebleau,在一万块抱石之间住着,爬上一阵子。

There are also so many books I want to read — my goal is to be able to read one or two books a week, and I'm currently at probably 0.5. The backlog is pretty big.

还有一堆书想读 —— 我的目标是一周读一到两本,现在大概 0.5,积了很多。

There's so much we can learn from history, and so much I don't understand as well as I would love to. I don't know anything about physics or robotics or any hardware or aerospace. There are just so many interesting topics. I'm excited to learn — even knowing that the AI will already know it.

历史里能学的太多,我想搞懂的事我现在都还差得远。物理、机器人、硬件、航空航天我什么都不会。这么多有意思的题目 —— 我很想学,即便知道 AI 早就比我懂。

Lenny01:23:26

Cat, this was amazing. You're awesome. Two follow questions. Where can folks find you online if they want to reach out and follow what you're up to? And how can listeners be useful to you?

Cat,这一期太精彩了。你太厉害。两个收尾问题 —— 想关注你或联系你的人,可以在哪里找到你?听众怎么样能帮上你的忙?

Cat01:23:35

The best way to reach out is — I'm @CatWu on Twitter. Feel free to tag me, feel free to DM me. I read all my DMs. I don't always respond to every single one, but I will read them all.

联系我最好的方式是 Twitter,@CatWu。欢迎 tag 我、DM 我。我所有的 DM 都会读 —— 不一定每条都回,但都会读。

The thing that is most helpful is — tell us where Claude Code and co-work aren't working well for you.

对我们最有用的事是 —— 告诉我们 Claude Code 和 co-work 在哪里没做好。

We're very grateful for the amount of positive feedback. But the things we thrive on are edge cases, errors, specific tasks we can reproduce where Claude Code or co-work fail. If you share that with us and we can reproduce it, this is something we can actively improve for our next generations of models and harnesses.

正面反馈我们非常感激,但真正让我们成长的是 edge case、错误、能复现的失败任务。你只要把这种情况给我们,我们能复现,我们就能在下一代模型和 harness 里把这块改好。

Lenny01:24:25

Extremely cool. People on Twitter are not shy with sharing this feedback — so keep it coming.

太棒了。Twitter 上大家都很敢于反馈 —— 请继续。

Cat01:24:30

Please, please share the problems that you're having with us.

真心拜托 —— 把你遇到的问题分享给我们。

Lenny01:24:34

It's really cool to see all your team being so active on Twitter, responding to people. So what I'm hearing is — this is actually stuff you guys see and react to.

看你们整个团队在 Twitter 上这么活跃、给人回复,真挺酷。我听到的是 —— 你们真的在看、真的在回应。

Cat01:24:44

We appreciate everyone being so engaged with us. It gives the team a ton of energy. We have this channel called "user love," and whenever you guys share a success story, we post it there. Whenever you share issues with our product, we put it into our feedback channel — that way our broader team is able to act on it.

大家这么积极地和我们互动,我们非常感激,真的给团队带来很多能量。我们内部有一个频道叫 "user love",你们分享的成功故事会被贴在那里;你们反馈的问题会进我们的 feedback 频道,这样更多团队成员都能去 follow up。

Lenny01:25:02

That is so cool to know. Thanks for sharing that. Well, Cat, thank you so much for being here.

这太棒了,谢谢你告诉我们。Cat,真心感谢你来。

Cat01:25:07

Thanks for having me.

谢谢邀请。

[Outro · Lenny] —— 收尾感谢、订阅 / 评分提示。略过不译。