BuildSpeak每日 builder 文摘
今日归档生词本关于
🎙 播客Training Data· 2026 年 7 月 21 日· 10,850 词 · 约 54 分钟

Factory's Matan Grinberg: The Coming ‘Dark Factory’ Where Software Builds Itself

SPACE 播放 / 暂停·←→ 上一句 / 下一句
Speaker 100:00 - 00:13
Bezos at Amazon, it's customer obsession. But in our mind, that's an input metric. Like, you don't wanna measure input metrics. It doesn't matter if you're customer obsessed. Like, you could be customer obsessed, and they file a restraining order against you because they don't like what it is that you're doing.
Speaker 100:00 - 00:13
Bezos 在 Amazon 讲的是 customer obsession(客户至上式执着)。但在我们看来,那是一个输入指标。你其实不该去衡量输入指标。你是否“痴迷客户”并不重要。比如说,你完全可能很痴迷客户,结果他们因为不喜欢你的做法,反而去申请 restraining order(限制令)来阻止你。
Speaker 100:13 - 00:28
Like, our job is to build something so good that our customers themselves become obsessed with us. That is our job. It's like, you know, the analogy is if you're a coach of a basketball team, you don't wanna tell your players before they come out there like, hey, guys. Make sure to sweat. It's like, what?
Speaker 100:13 - 00:28
我们的工作,是把东西做得足够好,好到让客户自己对我们产生执着。这才是我们的工作。就像,打个比方,如果你是一个篮球队教练,你不会在球员上场前对他们说,嘿,伙计们,记得多出汗。就像,啊?
Speaker 100:28 - 00:41
Like, no. Like, score points. Like, we need to score points. And in doing so, yeah, you're probably gonna sweat. And I think similarly, to create obsessed customers, you probably need to be really obsessed yourself with the customers, but the output is what matters.
Speaker 100:28 - 00:41
不,重点不是这个。重点是得分。我们得得分。而在这个过程中,是的,你大概率会出汗。我觉得类似地,要创造出对你高度着迷的客户,你自己大概也需要非常执着地关注客户,但真正重要的是输出结果。
Speaker 200:58 - 01:03
We're here in the studio with Matane from Factory. This is our second time
Speaker 200:58 - 01:03
我们现在在录音棚,和来自 Factory 的 Matane 一起。这是我们第二次
Speaker 101:03 - 01:04
with Matane. Thanks for having me.
Speaker 101:03 - 01:04
和 Matane 一起。感谢邀请我来。
Speaker 201:04 - 01:09
Here in the small and elite group of second time training data attendees. So thank you.
Speaker 201:04 - 01:09
进入这个规模很小但很精英的“第二次参加 training data 节目嘉宾”群体。所以,谢谢。
Speaker 101:09 - 01:09
Oh, yeah.
Speaker 101:09 - 01:09
哦,对。
Speaker 201:09 - 01:26
Matane is the cofounder and CEO of Factory, which makes droids, which are autonomous agents for the art of software development. Yes. Indeed. And, Matane, we're gonna jump right in because I think you guys are a little bit of a dark horse candidate in this world of software development. It is a market that has absolutely taken off.
Speaker 201:09 - 01:26
Matane 是 Factory 的 cofounder 和 CEO,Factory 做的是 droids,也就是面向软件开发这门技艺的 autonomous agents(自主 agent)。对,没错。Matane,我们直接切入正题吧,因为我觉得你们在软件开发这个世界里有点像一匹 dark horse(黑马)。这绝对是一个已经彻底爆发的市场。
Speaker 201:26 - 01:36
There are folks like Cloud Code and Cognition and others who who have a lead, but you guys are coming up strong. Talk about the competitive dynamics and what makes Factory special.
Speaker 201:26 - 01:36
像 Cloud Code、Cognition 以及其他一些团队,目前确实占有领先优势,但你们的势头也非常强。谈谈这里面的竞争态势,以及 Factory 的特别之处是什么。
Speaker 101:36 - 02:05
It's been a wild ride. We started Factory three three and a half years ago now. So in April 2023, when the world and the enterprise in particular was barely ready for GitHub Copilot, let alone fully autonomous agents. And so, I think the first two years, it was kind of our journey in the desert is is how I like to refer to it because we were focused on fully autonomous agents, but engineers weren't ready. Procurement teams at the enterprise weren't ready.
Speaker 101:36 - 02:05
这一路可以说非常疯狂。我们是在大约三年半前创办 Factory 的。也就是说,到了 2023 年 4 月,整个世界,尤其是企业界,连 GitHub Copilot 都还只是勉强准备好接受,更不用说完全自主的 agent(智能体)了。所以,我觉得最初那两年,某种程度上就像我们“在沙漠中的旅程”——我喜欢这么形容——因为我们专注于完全自主的 agent,但工程师们还没准备好,企业里的采购团队也没准备好。
Speaker 102:05 - 02:42
And so I think retrospectively, we really like honed our craft and learned a lot about how to build for developers in the enterprise. But, you know, it it it took a lot of time to actually come around to when they were ready to receive it. And so we're kind of now emerging much more. Some of these other players like Anthropic or OpenAI who have a ton of distribution are going in and, you know, bringing their incredible tools like Cloud Code or Codex. The thing that enterprises are really caring about that we have learned through those two years is they do not want anyone to kind of be their single point of failure.
Speaker 102:05 - 02:42
所以回过头看,我认为我们确实打磨了自己的能力,也学到了很多关于如何为企业中的开发者构建产品的经验。不过,说到底,还是花了很长时间,才等到他们真正准备好接受这类东西。所以我们现在算是开始更明显地走出来了。像 Anthropic 或 OpenAI 这样的其他参与者,拥有非常强大的分发能力,正在进入这个市场,并带来像 Cloud Code 或 Codex 这样的惊人工具。而企业真正关心的一点——这是我们在那两年里学到的——是他们不希望任何一方成为自己的单点故障。
Speaker 102:42 - 03:01
They do not want anyone to kind of control their fate. And so something that really matters is model independence. Everyone learned from cloud where, you know, back in the cloud days, it was like AWS or, or, or Azure being like, Hey, you know, come on in. Sign this three year contract. It's gonna be so cheap.
Speaker 102:42 - 03:01
他们不想让任何人掌控自己的命运。所以,真正重要的一点是 model independence(模型独立性)。大家都从 cloud(云计算)时代吸取了教训:那时候 AWS 或 Azure 会说,嘿,快来吧,签个三年合同,会非常便宜。
Speaker 103:01 - 03:08
We're gonna subsidize it. It'll be great. And then a couple years later, when it came time to renewal, they would 10 x the the the contract.
Speaker 103:01 - 03:08
我们会给补贴,一切都会很棒。然后几年之后,等到续约的时候,他们就会把合同价格直接提高到 10 倍。
Speaker 203:08 - 03:10
Dave Gravity. We got you now.
Speaker 203:08 - 03:10
Dave Gravity。现在我们套牢你了。
Speaker 103:10 - 03:14
Yeah. We got you. What? Are you gonna do a two year migration to go to someone else? Like, no way.
Speaker 103:10 - 03:14
对,我们拿住你了。怎么,你还打算花两年时间迁移到别人那里去吗?根本不可能。
Speaker 103:14 - 03:31
Everyone has scars from that now. And so everyone knows, look, Claude Code is fantastic. Codex from OpenAI is fantastic. We cannot put our fate in any one of these model providers' hands. Also, you just look at the risk profiles of the model labs versus the cloud providers.
Speaker 103:14 - 03:31
现在每个人身上都有这种伤疤。所以大家都明白,Claude Code 很棒,OpenAI 的 Codex 也很棒。但我们不能把自己的命运交到任何一家 model provider(模型提供商)手里。另外,你只要看看这些 model lab(模型实验室)和 cloud provider(云服务提供商)的风险画像就知道了。
Speaker 103:32 - 04:03
What's the last piece of drama that came out of a, one of the cloud providers versus like the model labs? It seems like there's kind of always some sort of chaos of, you know, internal fighting or getting in spats with the government or, you know, any other entities. And so, if you're gonna, you know, build this very important part of your business, you wanna make sure that you're robust to any of these changes. And that's something that we've learned over those kind of initial two years is like developers really care about things being modular. They want to know that they can customize it to what they want.
Speaker 103:32 - 04:03
最近一次从某家 cloud provider 和 model labs 之间闹出来的风波是什么?感觉好像总是会有某种混乱,比如内部争斗、和政府起冲突,或者和其他什么机构发生摩擦。所以,如果你要把这个非常重要的业务环节搭建起来,你就会想确保它能经受住这些变化。这也是我们在最初那两年里学到的一点:developers 真的很在意 modular(模块化)。他们想知道自己可以按需求去定制。
Speaker 104:03 - 04:26
They want to know that if there's a new model that comes out that's faster or cheaper or more performant, they can kind of hot swap it in. And that's, I think one of the biggest reasons why a lot of the largest enterprises are taking the momentum that they've had from a Codex or a Cloud Code, and then are carrying that into factory because they get that performance from these fantastic models, but they do it without the vendor lock in that, you know, the model labs directly.
Speaker 104:03 - 04:26
他们还想知道,如果有新 model 出来,更快、更便宜或者性能更强,他们可以直接把它 hot swap(热切换)进去。我觉得,这也是为什么很多大型企业会把他们从 Codex 或 Cloud Code 上积累起来的势头延续到 factory:他们既能获得这些优秀 models 带来的性能,又不用承受直接绑定 model labs 所带来的 vendor lock-in(供应商锁定)。
Speaker 204:26 - 04:31
And if I'm the enterprise, I'm gonna be like, wait a minute, am I now just getting locked into factory? What's the answer to that?
Speaker 204:26 - 04:31
但如果我是企业客户,我会说,等一下,那我现在是不是又被锁进 factory 里了?这个问题怎么回答?
Speaker 104:31 - 04:55
So the good, it's a really good question because that is something that you might think of like, okay, wait, so we're just switching the, the lock end point. All of the modularity that we build is such that if at some point you wanted to say, hey, you know what? Factory's not staying at the frontier anymore. Whether it's like the automations that you build or the skills registry that we help you create, the work that we've done stays in your code base. And any of the automations that we've created, the artifacts also live in your code base.
Speaker 104:31 - 04:55
这个问题特别好,因为你确实可能会这么想:好吧,等等,我们是不是只是把锁定的终点换了个地方。我们构建的所有 modularity(模块化能力)都意味着,如果未来某个时候你想说,嘿,你知道吗?Factory 已经不再处于最前沿了。无论是你构建的 automations(自动化流程),还是我们帮助你建立的 skills registry,这些工作成果都会留在你的 code base 里。而且我们创建的任何 automations,它们的 artifacts(产物)也都存在你的 code base 里。
Speaker 104:55 - 05:24
In other words, there aren't really things that we're saying, like our tribal knowledge about your org that we're keeping on our side and not giving to you. And that's part of the relationship that we have with customers is like, we similarly want to make sure we're providing the best experience possible. If we help you arbitrage between different models to get cost optimization, we're giving you that optimization. We're not taking that away from you. And I think that's a really important part of the trust that we're building with with these enterprises.
Speaker 104:55 - 05:24
换句话说,并不存在这种情况:比如关于你们 org 的某些 tribal knowledge(部落式经验知识)被我们留在自己这边,不交给你。这也是我们和客户关系中的一部分:我们同样希望确保自己提供的是尽可能好的体验。如果我们帮你在不同 models 之间做权衡以实现 cost optimization(成本优化),那这个优化结果是交给你的,我们不会再把它拿走。我认为,这是我们和这些企业建立信任中非常重要的一部分。
Speaker 205:24 - 05:54
You and I were talking probably a couple months ago at this point, and I was trying to give you credit for having the right vision for this market two, three years ago. And you responded with something along the lines of, thank you, but being two, three two or three years early is the same as being wrong. Yes. Which I thought was a wonderful response in so many ways. Can you talk about, like that, those two years in the desert, how did it feel to have this vision that turned out to be right that nobody appreciated for a year or two?
Speaker 205:24 - 05:54
你和我大概是几个月前聊过一次,当时我想称赞你,说你在两三年前就对这个市场有了正确的判断。而你的回应大概是:谢谢,但早了两三年,和看错了其实没区别。对。 我觉得这句回答在很多层面上都非常精彩。你能不能讲讲那种“在沙漠里的两年”是什么感觉?当你拥有一个后来被证明是对的愿景,但在一两年里却没有人认可,它是什么感受?
Speaker 205:54 - 05:58
Can you just talk about like that journey and what it has done to the DNA of, of your company?
Speaker 205:54 - 05:58
你能不能谈谈这段历程,以及它对你们公司 DNA 产生了什么影响?
Speaker 105:58 - 06:21
Yeah. I mean, in the moment it's really, really difficult because, you know, I hadn't had a job before. I dropped out of my PhD to start this company. And, you know, over the course of those two years, convinced, you know, 20 of the smartest people that I've ever met to quit what it was that they were doing and, you know, join factory and join us on this mission. And these are people with families.
Speaker 105:58 - 06:21
对,我的意思是,在那个当下,这真的、真的非常困难,因为我以前从来没有上过班。我是从 PhD 退学出来创办这家公司的。而且在那两年里,我说服了 20 个我见过最聪明的人,放下他们原本在做的事情,加入 factory,和我们一起投入这项使命。而这些人很多都是有家庭的。
Speaker 106:21 - 06:47
These are people with kids who are, like, dedicating years of their lives to this problem and going, you know, customer after customer, and they, like, they weren't ready for agents. They didn't get it. Also, the models weren't as performant, but I think a lot of it was behavioral. And, I mean, even just a fun anecdote of, like, giving developers an NPS survey. If you ever are giving a developer an NPS survey, they do not like whatever it is that you're giving it to them.
Speaker 106:21 - 06:47
这些人很多都有孩子,差不多是把自己人生中的好几年都投入到这个问题上,一个客户接一个客户地去推进;但他们当时还没准备好接受 agent。他们不理解。还有一个原因是,模型那时的表现也没那么强,但我觉得很大一部分其实是行为层面的。再讲个有趣的小 anecdote(轶事):如果你给开发者做 NPS survey,他们基本不会喜欢你拿来让他们打分的那个东西。
Speaker 106:47 - 07:03
Because, like, developers, they vote with their feet. They are very clear what they like and what they don't like. And if you're like, I wonder if they like it, they definitely don't. And but during that time, I think there were there was a lot that we were learning. There was a lot that I myself was like, I'd never had a job before.
Speaker 106:47 - 07:03
因为开发者会用脚投票。他们喜欢什么、不喜欢什么,态度都非常明确。要是你还在想“他们到底喜不喜欢呢”,那答案基本就是:他们肯定不喜欢。不过在那段时间里,我觉得我们也学到了很多。我自己也有很多事情是在摸索中学的——毕竟我之前从来没有工作过。
Speaker 107:03 - 07:24
Enterprise sales is not something that comes obvious to to a physicist. But at the end of the day, it doesn't it doesn't matter. There's no you don't get any, you know, bonus points for being early because, like, who who cares? Like, there's no consolation prize. It's either you do the thing or you don't do the thing, and that's all that matters.
Speaker 107:03 - 07:24
Enterprise sales(企业销售)这种事,对一个 physicist(物理学家)来说并不天然容易上手。但说到底,这其实不重要。你不会因为起步早就得到什么额外加分,因为谁在乎呢?没有安慰奖。事情只有两种结果:要么你做成了,要么你没做成;真正重要的只有这一点。
Speaker 107:24 - 07:47
And for the team, it was really tough. There were points where we ended up getting good at enterprise sales, but the product still wasn't good. And that's a very tricky position to be in because we ended up, you know, getting to a point where we were like just under 2,000,000 in revenue and the product was not good. And there was a point in time where we realized this, because if you're really good at sales, you can sign contracts. That's like, you can definitely do that.
Speaker 107:24 - 07:47
对团队来说,那段时间也特别艰难。有一阵子我们确实变得很擅长 enterprise sales,但产品本身还是不够好。这会把你置于一个非常棘手的位置,因为最后我们做到的收入规模差不多接近 2,000,000,但产品并不好。后来有个时点我们意识到了这一点:如果你销售能力真的很强,你是可以把合同签下来的,这点完全做得到。
Speaker 107:48 - 08:27
But if you're doing that and the developers don't like your product, it's like a ticking time bomb because it's eventually they're gonna churn and it's gonna be really, really bad. We realized this and we proactively gave all of those customers their money back. And I remember that was one of the most difficult decisions to make because not only is there a, you know, a group of, you know, 20 people who are getting ridiculous offers from all the labs, they have these huge, you know, financial incentives to go elsewhere. There are all these other companies that are doing well and they decided to do this. And then we're gonna say, oh yeah, Hey, by the way, that, you know, little bit of revenue we managed to get, we're actually gonna give it back because we don't think product is making their developers happy.
Speaker 107:48 - 08:27
但如果你在这么做,而开发者又不喜欢你的产品,那这就像一颗滴答作响的定时炸弹,因为他们最终一定会 churn(流失),到时候局面会非常、非常糟。我们意识到这一点后,主动把所有这些客户的钱都退了回去。我记得那是最难做出的决定之一,因为不只是说团队里那二十来个人都拿到了各个 labs 给出的夸张 offer,他们有非常强的财务激励去别的地方;外面还有很多其他公司发展得很好,而他们还是选择留下来做这件事。结果我们却要说,哦对了,顺便一提,我们好不容易拿到的那一点收入,其实要退回去,因为我们觉得这个产品并没有让他们的开发者满意。
Speaker 108:27 - 08:29
We also had
Speaker 108:27 - 08:29
我们还曾经有过
Speaker 208:29 - 08:30
Why did tell you make that decision?
Speaker 208:29 - 08:30
你们为什么会做出那个决定?
Speaker 108:30 - 08:54
You know, we sold them on a good vision and convinced them that, you know, this is the right team to work with and that we were going to deliver the solution for them. But we realized that the way that we had sold them on it and the product that we were delivering was not up to snuff in a way that I don't think it would hold true to one of our operating principles. And one of our operating principles that I really like is create obsessed customers.
Speaker 108:30 - 08:54
我们当时是用一个很好的愿景去说服他们的,让他们相信这是一支值得合作的团队,也相信我们会为他们交付解决方案。但后来我们意识到,我们当初向他们承诺的方式,以及我们实际交付的产品,都没有达到应有的水准;而且在我看来,这不符合我们的一条 operating principle(运营原则)。我非常喜欢我们的一条 operating principle,就是:create obsessed customers。
Speaker 208:54 - 08:54
Yeah.
Speaker 208:54 - 08:54
对。
Speaker 108:54 - 09:12
This kind of like flips over Bezos' thing, where Bezos at Amazon, it's customer obsession. But in our mind, that's an input metric. And, look, input metrics are like, you don't wanna measure input metrics. It doesn't matter if you're customer obsessed. Like, you could be customer obsessed, and they file a restraining order against you because they don't like what it is that you're doing.
Speaker 108:54 - 09:12
这有点像是把 Bezos 那套理念翻过来了。Bezos 在 Amazon 讲的是 customer obsession(客户至上式痴迷)。但在我们看来,那是一个输入指标。说到底,输入指标这种东西,你其实不该拿来衡量结果。你是不是痴迷客户并不重要。比如,你完全可能对客户过度痴迷,结果他们因为不喜欢你做事的方式,反而去申请限制令。
Speaker 109:12 - 09:28
Like, our job is to build something so good that our customers themselves become obsessed with us. That is our job. It's like, you know, the analogy is if you're a coach of a basketball team, you don't wanna tell your players before they come out there like, hey, guys, make sure to sweat. It's like, what? Like, no.
Speaker 109:12 - 09:28
我们的工作,是把产品做得好到让客户自己对我们上瘾、变得痴迷。这才是我们的工作。这就像一个类比:如果你是篮球队教练,你不会在球员上场前跟他们说,嘿,伙计们,记得多出汗。什么?当然不是。
Speaker 109:28 - 09:46
Like, score points. Like, we need to score points. And in doing so, yeah, you're probably gonna sweat. And I think similarly, to create obsessed customers, you probably need to be really obsessed yourself with the customers, but the output is what matters. And I think that coming back to this, the product that we were delivering was not creating obsessed customers.
Speaker 109:28 - 09:46
你得得分。我们需要得分。在这个过程中,是的,你大概率会出汗。我觉得同样地,要创造出痴迷我们的客户,你自己大概也必须非常痴迷客户,但真正重要的是输出结果。我想回到这一点:我们当时交付的产品,并没有创造出那种痴迷型客户。
Speaker 109:46 - 10:02
And we wanted to make sure, like, this was a group of the smartest people I've ever met. We were getting there. Like, we were getting a lot of intuition. Things were starting to come together internally. Like we could see internally, we were starting to become a lot more agent native in how we were doing things.
Speaker 109:46 - 10:02
而且我们想确认的是,这是一群我见过最聪明的人。我们当时正在接近那个目标。我们获得了很多直觉判断,很多事情在内部开始逐渐成形。我们在内部已经能看到,我们做事的方式正开始变得越来越 agent native(以 agent 为原生核心)。
Speaker 110:02 - 10:25
And the product was kind of scratching that itch, but we were kind of ahead of our customers. And we wanted to maintain trust with our customers so that when it does hit, we can come back to them and say, hey, guys, this is the real deal. I promise. To build that credibility, we had to say, Hey, look, you know, even though you were maybe happy to continue, we're gonna give you this back and say three months from now, I think it'll be ready. Give us some time.
Speaker 110:02 - 10:25
那个产品算是在某种程度上挠到了痒处,但我们有点走在客户前面了。我们希望维护与客户之间的信任,这样等时机真正到来时,我们可以回去对他们说,嘿,伙计们,这次是真的。我保证。为了建立这种可信度,我们必须坦诚地说,嘿,你看,哪怕你们也许愿意继续用下去,我们还是要把这个退回给你们,并且说,三个月后,我觉得它就准备好了。给我们一点时间。
Speaker 110:25 - 10:27
And I promise we will knock your socks off.
Speaker 110:25 - 10:27
我保证,到时候一定会让你们惊艳。
Speaker 210:27 - 10:29
How did your customers react when you had that conversation?
Speaker 210:27 - 10:29
当你和客户进行那番谈话时,他们是什么反应?
Speaker 110:30 - 10:42
Some of them were like, oh, great. Sounds good. Because I think it wasn't something that they, you know, were obsessed with. Some of them were a little bit confused. But I think generally, it's especially enterprises, they're not used to these things.
Speaker 110:30 - 10:42
他们中有些人会说,哦,太好了,听起来不错。因为我觉得那并不是让他们特别痴迷的东西。有些人则有点困惑。但总体来说,尤其是 enterprises(企业客户),他们并不习惯这种做法。
Speaker 110:42 - 10:45
A lot of times enterprise budget, once it's gone, it's gone and no one really cares.
Speaker 110:42 - 10:45
很多时候,enterprise budget(企业预算)一旦花没了,也就没了,基本上没人真的会在意。
Speaker 210:45 - 10:46
Yeah.
Speaker 210:45 - 10:46
对。
Speaker 110:46 - 11:05
And so some of them didn't even know if they had a mechanism by which to take back money. But, you know, it's a difficult thing to tell also, like, investors who believe in you. Like, you know, remember having the conversation with Sean. And Sean, obviously, he's stayed really close with the company, so he was very, like, on the same page. But it's kind of a scary thing to be like, hey.
Speaker 110:46 - 11:05
所以他们当中有些人甚至都不知道自己有没有一种机制可以把钱收回来。但你也知道,这件事也很难开口告诉那些相信你的投资人。我记得当时和 Sean 聊过这件事。Sean 很显然一直和公司走得很近,所以他的想法和我们非常一致。但要突然去说,嘿,这其实还是挺吓人的。
Speaker 111:05 - 11:13
By the way, know, remember all those updates and you're saying, hey. Look. The, you know, revenue's going up. It's about to go down to zero. It was a scary thing.
Speaker 111:05 - 11:13
顺便说一句,你还记得之前那些更新吧,你一直在说,嘿,你看,revenue(营收)在上涨。可它马上就要掉到零了。这真的是件很可怕的事。
Speaker 111:13 - 11:35
And I think it was kind of a leap of faith of, like, we see the signal internally early of, like, this is the direction we need to go. We need to kind of pivot the approach on the product. But I remember that all hands where we told the whole team, it it was like, oh my that was, like, one of the worst months of my life. Like, I was just because no like, not everyone was gonna say, like, what the hell is this? What's going on?
Speaker 111:13 - 11:35
我觉得那有点像一次 faith leap(信念上的冒险):我们在内部很早就看到了某种信号,知道这才是我们必须前进的方向。我们需要调整产品的方法,算是做一次 pivot(战略转向)。但我记得那次 all hands(全员会),我们把这件事告诉整个团队时,天啊,那大概是我人生中最糟糕的一个月之一。我当时整个人都……因为不可能每个人都会直接说“这到底是什么?到底发生了什么?”
Speaker 111:35 - 11:41
But it's kind of the looks on their faces where they kinda go a little bit pale, and they're like, oh boy, like, is this just the early signs and we're about
Speaker 111:35 - 11:41
但更明显的是他们脸上的表情:脸色有点发白,然后像是在想,哦天哪,这会不会只是某种早期信号,说明我们马上就要——
Speaker 211:41 - 11:45
to sink completely? How'd you keep the team together through that?
Speaker 211:41 - 11:45
——彻底沉下去了?你们是怎么在那种情况下把团队稳住的?
Speaker 111:46 - 12:22
I think honestly, the only reason the team stayed together is we were so ruthless about hiring early on where it was like people that are genuinely really, really obsessed with the mission, which our mission is to bring autonomy to software engineering. And, like, really, really caring about that, making sure everyone was also, like, very clear feedback loops as to, like, this the fate is in our hands. It's not like this is, like, oh, something that I go do. It's like we all have a part to play in, you know, making this work. And I think embracing how much it sucked was also, I think, something that was very valuable.
Speaker 111:46 - 12:22
说实话,我觉得团队之所以能坚持在一起,唯一的原因是我们在早期招人时极其 ruthless(严格、毫不妥协):我们找的都是那种真正、真正对使命非常着迷的人。我们的使命是把 autonomy(自主性)带到 software engineering(软件工程)里。大家必须真的非常在乎这件事。同时也要确保每个人都非常清楚 feedback loops(反馈闭环):公司的命运掌握在我们自己手里。这不是那种“哦,这是某个人去做的事”,而是我们每个人都在让这件事成功这件事上扮演一部分角色。我还觉得,坦然接受那段时间真的很难熬,这一点本身也非常有价值。
Speaker 212:22 - 12:23
Just being honest about it.
Speaker 212:22 - 12:23
就是要对这件事保持诚实。
Speaker 112:23 - 12:31
You're being super honest about, like, yeah, this sucks. Like, you overlook the look at those competitors. The revenue's going up like crazy. Like, this is not good. Like, we are in a very bad position.
Speaker 112:23 - 12:31
你当时是非常坦诚的,会直接说,没错,这很糟。你再看看那些竞争对手,营收增长快得惊人。这情况不妙。我们当时的处境非常糟糕。
Speaker 112:31 - 12:59
Like, we just had to give back all of our revenue. Like, we need to really get our shit together. And in the moment, I think retrospectively, those are the moments where really the deepest bonds are made. Like, if you talk to people who are, like, athletes or even, like, academics or whatever, whenever you're in the, like, stressful period, whether it's, like, cramming before finals or, you know, in intense, like you know, we have some some rowers on our team, and I think that's an example we always go to. Like, that's a pure pain.
Speaker 112:31 - 12:59
我们当时几乎不得不把所有营收都吐回去。我们真的必须好好整顿自己。事后回看,我觉得正是在那种时刻,最深的纽带才会真正形成。如果你去问那些运动员,或者学者之类的人,你会发现,每当你处在那种高压时期——不管是期末前突击复习,还是那种高强度的状态,我们团队里也有一些 rowers,我们总会拿这个来举例——那就是纯粹的痛苦。
Speaker 212:59 - 13:00
Pure pain sport.
Speaker 212:59 - 13:00
纯粹痛苦的运动。
Speaker 113:00 - 13:16
It's pain. It's literally just there is one number that quantifies your performance. It's just what is your time on your two k or your time in that? But, like, embracing that is what creates those enduring bonds such that afterwards, like, we know what it's like to be at rock bottom. We know what it's like to lose.
Speaker 113:00 - 13:16
就是痛苦。真的就是这样:有一个数字可以量化你的表现,比如你划 2k 的时间,或者你在那项比赛里的成绩。但正是拥抱这种东西,才会建立起那种持久的纽带,让你在事后会觉得,我们知道跌到谷底是什么感觉。我们知道失败是什么感觉。
Speaker 113:16 - 13:35
We know what it's like. I mean, when we first started the company, our valuation was 5,000,000. Like a lot of our competitors, a lot of the companies out there these days, they don't know what it's like to not be a unicorn. That's like manifestly, that is what they are day one. Whereas like we have been there kind of in those dark moments and not a single person left.
Speaker 113:16 - 13:35
我们知道那是什么感觉。我的意思是,我们刚创办公司时,估值只有 5,000,000。如今很多竞争对手、很多公司,根本不知道“不是什么 unicorn(独角兽公司)”是什么感觉。对他们来说,这几乎是显而易见的——他们从第一天起就是那样。而我们经历过那些黑暗时刻,却没有一个人离开。
Speaker 113:35 - 13:48
Yeah. That makes us so resilient and so strong that, you know, going forward, things are going a lot better now, But there are going to be really bad times, but we have that resiliency in our DNA that I'm not sure some of these other companies do.
Speaker 113:35 - 13:48
对,这让我们变得非常有韧性,也非常强大。所以,往后看,虽然现在情况已经好了很多,但未来一定还会有非常糟糕的时候。不过我们的 DNA 里已经有了这种韧性,我不确定其他一些公司是否也具备这一点。
Speaker 313:48 - 13:57
I love that. So talk us talk to us about what changed. And I'm curious your comment from earlier that the models getting better is not the most important thing that happens.
Speaker 313:48 - 13:57
我喜欢这一点。所以跟我们讲讲,到底是什么发生了变化。我也很好奇你前面那句评论:model(模型)变得更好,并不是最重要的事情。
Speaker 113:57 - 13:58
Because at least in my mind, the model's
Speaker 113:57 - 13:58
因为至少在我看来,model(模型)的
Speaker 313:58 - 14:01
getting better is the most important thing that happens. Just help me understand.
Speaker 313:58 - 14:01
变得更好是最重要的事情。你就帮我理解这一点。
Speaker 114:01 - 14:19
Yeah. So so a couple of things. So one is the interaction pattern that we were building for before was too ambitious. Like to your point, we were right in that what we were building for was fully autonomous agents, but it was two years too early, which makes it wrong. And fully autonomous agents require a complete change in behavior from the developer.
Speaker 114:01 - 14:19
对。所以有几件事。其一是,我们之前一直在构建的交互模式太激进了。就像你说的,我们瞄准的确实是 fully autonomous agents(全自主 agent),这一点没错;但它早了两年,于是就等于错了。而 fully autonomous agents 需要开发者彻底改变自己的行为方式。
Speaker 114:19 - 14:35
And we were trying to do that out of the box before they were even using tools like Copilot. It was just too much of a leap. It was too much of a step function jump. So it's an important day. 09/26/2025 was when we first put out, basically, the the droid CLI.
Speaker 114:19 - 14:35
而我们当时甚至是在开发者还没开始使用像 Copilot 这样的工具之前,就想把这套东西开箱即用地推出来。这步子迈得太大了,是一种阶跃式的巨大跳变。所以那是个重要的日子。09/26/2025 是我们第一次推出,基本上,就是 droid CLI 的时间。
Speaker 114:35 - 14:56
And the droid CLI met developers where they were in a manner that previously these fully autonomous agents did not. And also its performance was like completely state of the art and it was model agnostic. So it could use every model that was out there. September 26 was also two years after we initially started. So the world had gotten much more used to using things like auto complete.
Speaker 114:35 - 14:56
而 droid CLI 以一种此前那些 fully autonomous agents 做不到的方式,贴近了开发者当下的工作状态。另外,它的性能也完全是 state of the art(业界最先进水平),而且它是 model agnostic(模型无关)的。所以它可以使用当时市面上的每一个模型。9 月 26 日也正好是我们最初启动这件事两年之后。所以那时整个世界已经更习惯使用像 auto complete(自动补全)这样的东西了。
Speaker 114:56 - 15:33
Like by by late twenty twenty five, most engineers were using an auto complete tool and many were starting to, at the time, use like a chat interface to ask an agent to go do changes like wholesale. So like the more agentic interaction. However, what we see is that, like, if you go back now and use in this, like, agentic interaction, some of these older models, they're still good. So the biggest thing that changed was developers and, in particular, in the enterprise, like, being open minded to this new way of working. In particular, you know, developers, they've established their workflows over the last thirty years.
Speaker 114:56 - 15:33
到 2025 年底时,大多数工程师都已经在使用某种 auto complete tool(自动补全工具),而且很多人当时也开始通过类似 chat interface(聊天界面)的方式,让一个 agent 去大范围地直接修改代码,也就是更 agentic(更具 agent 特征)的交互方式。不过,我们现在看到的是,如果你回过头,在这种 agentic interaction(agent 式交互)里使用一些较老的模型,它们其实仍然很好用。所以变化最大的是开发者,尤其是企业里的开发者,他们开始更愿意接受这种新的工作方式。尤其是,开发者在过去三十年里已经建立起了自己的工作流。
Speaker 115:33 - 15:39
They can be stubborn. A lot of them were like, no. No. No. Like, my craft could never be done by an, you know, an AI tool.
Speaker 115:33 - 15:39
他们可能很固执。很多人当时都会说,不,不,不,我这门手艺不可能交给什么 AI 工具来做。
Speaker 115:40 - 15:49
So a lot of it was just, like, understanding how to work with these tools and having the willingness to go in and try and also the intuition about what are the guardrails that you need to provide in order for it to succeed.
Speaker 115:40 - 15:49
所以很大一部分其实只是去理解该怎么和这些工具协作,并且愿意亲自进去试一试;同时还要有那种直觉,知道为了让它成功,你需要提供哪些 guardrails(护栏 / 约束)。
Speaker 315:49 - 15:50
Yeah.
Speaker 315:49 - 15:50
对。
Speaker 115:51 - 16:01
And so I think it was a combination of both of these things. The model's getting better, so you need to do less in the way of providing guardrails. But, also, developers lowering their guard and being like, okay. You know what? Let me go try and do these things.
Speaker 115:51 - 16:01
所以我认为这是这两方面共同作用的结果。模型在变得更好,因此你在提供 guardrails 这件事上需要做得更少了。但与此同时,开发者也放下了戒备,开始觉得:好吧,你知道吗?让我去试着用这些东西做点事情。
Speaker 116:01 - 16:18
It's gonna go do things I don't like. And then, also, there's a certain degree to which when Andre Karpathy tweets about something, then every engineer suddenly is like, okay. You know, maybe this is true. And Andre started to tweet about these agentic work. Early on, he wasn't as open to it.
Speaker 116:01 - 16:18
它会去做一些我不喜欢的事。然后还有一点是,当 Andre Karpathy 发推谈到某件事时,突然每个工程师都会觉得,哦,行。你知道,也许这是真的。And Andre 开始发推讨论这种 agentic work。早期他对此并没有那么开放。
Speaker 116:18 - 16:36
And then him being more open to it genuinely just changed some people's minds, which is funny, but that's some of the things that go into behavior changes. Like, you hear it from people you trust. You start seeing it, you know, from people within your organization who are maybe a little bit more agent native, but that's that's kind of these things together is what what changed that.
Speaker 116:18 - 16:36
然后他变得更开放这件事,确实改变了一些人的想法,这很有意思,但行为变化本来就有一部分是这么发生的。比如,你会从你信任的人那里听到这件事。你也开始在组织内部看到,有些人可能会更 agent native,但这些因素加在一起,才是促成这种变化的原因。
Speaker 316:36 - 16:38
And now we're all gonna be on Slack. We
Speaker 316:36 - 16:38
现在我们都会上 Slack。我们
Speaker 116:39 - 16:44
might we might be on we might be pushing the limits of Slack, which I think is gonna be another interesting thing. That's cool. But yeah.
Speaker 116:39 - 16:44
可能——我们可能已经在——我们可能正在逼近 Slack 的极限,我觉得这也会是另一件很有意思的事。很酷。不过,是的。
Speaker 316:44 - 16:50
Okay. So September 2025, you launched the Droid CLI. You said Frontier Performance Soda. What does that mean for you?
Speaker 316:44 - 16:50
好。那是 2025 年 9 月,你们发布了 Droid CLI。你当时说过 Frontier Performance Soda。这对你来说是什么意思?
Speaker 116:50 - 17:15
There's like the benchmarks, which have a very short half life. Like anytime there's a good benchmark, it gets bench maxed within like three to six months. Yeah. At the time, I think the one that we kind of championed when we launched and kind of, it ended up becoming a pretty good benchmark was terminal bench. So prior to that, the one that was kind of leading was SWE bench, which was kind of took some open source projects and some examples of issues that were then solved.
Speaker 116:50 - 17:15
benchmark(基准测试)这类东西的半衰期很短。就像,任何一个好的 benchmark,基本都会在三到六个月内被刷到极限。对。当时我觉得我们发布时主推的那个,而且后来也确实成了一个相当不错的 benchmark,就是 terminal bench。在那之前,比较领先的是 SWE bench,它基本上是选取一些 open source project(开源项目)以及一些后来被解决的问题案例。
Speaker 117:15 - 17:46
The problem with that was it was very focused on like Python and like scripting or like individual file changes. Whereas terminal bench was more one, it was in the terminal setting. So it was things like scheduling runs and things that were not just, like, changing the code file, but general software development tasks. And that was something that we ended up, you know, having really frontier performance on. Now it's, like, bench maxed to the extreme to where it's, like, I think, you know, models that come out now are, like, 90% on it.
Speaker 117:15 - 17:46
它的问题在于,它非常聚焦于 Python,以及脚本类任务,或者单个文件的修改。而 terminal bench 更偏向于,第一,它是在 terminal(终端)环境里。所以它包含像调度运行这样的事情,以及那些不只是改代码文件,而是更一般的软件开发任务。那也是我们最终在上面取得非常 frontier performance(前沿性能)的地方。现在它基本已经被刷到极致了,到了——我想,你知道,现在出的模型在它上面大概都能到 90%。
Speaker 117:46 - 17:53
And I think there's a very short time horizon from putting out a good benchmark to then it being kind of in the training data.
Speaker 117:46 - 17:53
而且我觉得,从发布一个好的 benchmark,到它随后某种程度上进入 training data(训练数据),这个时间窗口是非常短的。
Speaker 317:53 - 18:00
What goes into building a great and is it that is it a great harness? And it seems like there's almost a lot of FUD in the ecosystem of my harness is better than your harness.
Speaker 317:53 - 18:00
要打造一个优秀的 harness,到底需要哪些要素?以及,什么样的 harness 才算优秀?感觉生态里似乎有很多 FUD(恐惧、不确定性、怀疑),都在说“我的 harness 比你的 harness 更好”。
Speaker 118:00 - 18:00
And Yeah.
Speaker 118:00 - 18:00
嗯,对。
Speaker 318:00 - 18:11
Yeah. You know, you need to own the model to have a good harness or actually you're you have a better harness if you don't own the model. Like Yeah. What's your mental model for for Yeah. You know, maxing aside, what keeps you at the frontier?
Speaker 318:00 - 18:11
对。你知道,常见的一种说法是,你得自己拥有 model(模型)才能做出好的 harness,或者反过来说,如果你不拥有 model,你反而会有更好的 harness。对吧。所以你的 mental model(思维框架)是什么?先把营销话术放一边,是什么让你始终站在前沿?
Speaker 118:11 - 18:32
Yeah. So a couple so some general things that matter are the way you do caching. So, you know, cash tokens end up being, like, a tenth as expensive. And so one big piece of performance for a given harness is what what is your, like, rate of of token caching? Another example would be how do you perform while in compression or compaction?
Speaker 118:11 - 18:32
对。所以有几个比较通用的重要点。比如你做 caching(缓存)的方式。你知道,cache tokens(缓存 token)最后的成本可能只有十分之一左右。所以,对于一个既定的 harness 来说,一个很大的性能因素就是:你的 token caching(token 缓存)命中率到底有多高?再比如,你在 compression 或 compaction(压缩 / 紧缩)过程中表现如何?
Speaker 118:33 - 19:01
So typically, when you're dealing with a a long session, you're gonna exceed the context limit of the model itself. And so the harness will do some sort of, you know, summarization, compression, compaction, whatever you wanna call it. And the way that you perform during that compaction is a big determining factor of how good your harness is. And tests that they do for that are like, they call it needle in the haystack where you have some long thread and maybe there's one piece of information that's really important. How often will your harness preserve that through compaction?
Speaker 118:33 - 19:01
通常来说,当你处理一个很长的 session(会话)时,你会超出 model 本身的 context limit(上下文长度限制)。所以 harness 会做某种 summarization、compression、compaction,不管你怎么称呼它。而你在这种 compaction 期间的表现,是决定 harness 好坏的一个关键因素。他们为此做的一类测试,叫 needle in the haystack(草堆里找针):你有一段很长的线程,也许其中有一条信息特别重要。你的 harness 在 compaction 之后,能有多高概率把那条信息保留下来?
Speaker 119:01 - 19:26
Other examples are like tool use or how does it use the environment to validate whatever work that it's doing? These are things that you can kind of have individual metrics on and that we kind of have our own internal benchmarks to measure how do the out of the box agents do versus how does factory perform. I think one thing that naively everyone believed initially was if you train the model and you build the harness, you're gonna make them better together.
Speaker 119:01 - 19:26
其他例子还包括 tool use(工具使用),或者它会怎样利用 environment(环境)来验证自己正在做的工作。这些东西某种程度上都可以有各自独立的指标,我们内部也有自己的 benchmark(基准测试),来衡量开箱即用的 agents(智能体)表现如何,以及 factory 的表现如何。我觉得,最初大家一个很天真的共同看法是:如果你同时训练 model、同时构建 harness,它们就会一起变得更好。
Speaker 319:26 - 19:27
Yeah.
Speaker 319:26 - 19:27
对。
Speaker 119:27 - 19:38
And much to the chagrin of many of my friends at OpenAI and Anthropic, this is not true. If you build a harness that supports different models, that harness will be better.
Speaker 119:27 - 19:38
但很遗憾——我的很多在 OpenAI 和 Anthropic 的朋友可能不太愿意听到——事实并非如此。如果你构建的是一个支持不同 models(模型)的 harness,这个 harness 反而会更好。
Speaker 319:38 - 19:41
What's the like, my intuition would be model harness co design makes you better.
Speaker 319:38 - 19:41
我直觉上会觉得,model 和 harness 的协同设计会让你做得更好,这一点到底是怎么回事?
Speaker 119:41 - 19:42
Yes.
Speaker 119:41 - 19:42
对。
Speaker 319:42 - 19:44
What's the intuition for why it's actually not?
Speaker 319:42 - 19:44
为什么从直觉上看,它实际上并不是这样?
Speaker 119:44 - 20:13
It's very analogous to the idea maybe, like, I don't know, ten years ago of if you were to be like, Hey, I want to train my personal AI back in like ML days before like GPT-three. I want to train my personal AI. I'm going to give it all of my data because I want it to know me. Turns out the answer was train it on the whole Internet, and it'll be so much better for you than if it were just trained on your data. So there's a sort of analog that emerges where it's what data is to a model, models are to a harness.
Speaker 119:44 - 20:13
这很类似于一个想法——大概十年前吧,如果你当时会说,嘿,我想训练我自己的 personal AI,那还是 ML 时代、还在 GPT-three 之前。我想训练我的 personal AI,所以我要把我所有的数据都喂给它,因为我希望它了解我。结果答案是:把整个 Internet 都拿来训练它,它对你的帮助反而会比只用你的数据训练更大。所以这里会出现一种类比:对于 model 来说,data 是什么;对于 harness 来说,models 就是什么。
Speaker 120:15 - 21:06
Where the more models you expose to a harness, you avoid overfitting that harness to the nuances of that model in particular. And there are certain intricacies about different models that you can learn from and then improve different models performance in your own harness. And this was why, for example, we kind of stopped doing it because TerminalBench got so bench maxed. But initially, when, like, every new Opus or GPT model would come out, it would perform better on TerminalBench in Droid than it would in Claude Coder Codex, which is why like, and this is something that, you know, I think was somewhat frustrating too because I deal like, from a lab perspective, you ideally want it so that it's better together because then that means you have to use their harness and you can't use a different one. But I think the reality is it it it's, you know, having that multimodal harness ends up getting kind of frontier on on all of those assets.
Speaker 120:15 - 21:06
也就是说,你让一个 harness 接触的 models 越多,就越能避免让这个 harness 过度拟合某一个特定 model 的细微特性。而且,不同 models 身上有一些特定的复杂之处,你可以从中学习,然后在你自己的 harness 里改进不同 models 的表现。这也是为什么,比如说,我们后来某种程度上不再继续这么做了,因为 TerminalBench 已经被刷到 benchmark(基准测试)上限了。但最开始时,每次有新的 Opus 或 GPT model 出来,在 Droid 里的 TerminalBench 表现都会比在 Claude Coder Codex 里更好。所以这也是为什么——而且这件事我觉得某种程度上也挺让人沮丧的,因为如果从 lab 的视角看,你理想上会希望“绑在一起时效果更好”,因为那样别人就必须用他们的 harness,不能换别的。但我觉得现实是,拥有那种 multimodal harness(多模型/多模态的调度框架),最后往往会在所有这些资产上都逼近前沿水平。
Speaker 221:06 - 21:11
A good, like, example or illustration of that? Conceptually, makes sense. Is there, like, an easy way to illustrate it?
Speaker 221:06 - 21:11
有没有一个比较好的例子或说明?概念上我能理解。有什么容易说明它的方式吗?
Speaker 121:11 - 21:19
Maybe maybe a good example of it is, like, if you're familiar with the different behaviors of Opus and GPT 5.6 right now.
Speaker 121:11 - 21:19
也许一个不错的例子是,如果你现在熟悉 Opus 和 GPT 5.6 的不同行为的话。
Speaker 321:19 - 21:20
I am. He's not.
Speaker 321:19 - 21:20
我熟悉,他不熟悉。
Speaker 121:20 - 21:36
Opus tends to be I mean, loosely loosely. I mean, to be fair, honestly, these days, I'm not doing it as much either. But I will say this. Loosely, Opus is kind of like that super friendly colleague where you're like, hey. I wanna go do these 20 tasks, and they're like, okay.
Speaker 121:20 - 21:36
Opus 往往是——我是说,大体上、大体上吧。公平地说,老实讲,这些天我自己也没怎么做这类比较了。不过我可以这么说:粗略来看,Opus 有点像那种超级友好的同事。你对他说,嘿,我想去做这 20 个任务;他会说,好啊。
Speaker 121:36 - 21:42
Cool. Hey. By the way, five of those tasks, I realized we didn't need to do it. Don't worry about it. I got other of these done.
Speaker 121:36 - 21:42
很好。对了,顺便说一下,其中有 5 个任务我发现其实没必要做。别担心。我已经把另外一些做完了。
Speaker 321:42 - 21:44
Did it this way. Tonight's not a good time. Let's pick it up
Speaker 321:42 - 21:44
我是这么做的。今晚不是个合适的时间。我们明早再接着做
Speaker 121:44 - 21:58
in the morning. Yeah. Like, let's go let's go get a beer afterwards and hang out, whatever. Meanwhile, like, GPT 5.6 is like, absolutely, I will do every single one of those, and nothing will stop me. I'm not gonna sleep until this it's, like, kind of very OCD and, you know, meticulous.
Speaker 121:44 - 21:58
吧。对。比如说,我们之后去喝杯啤酒,放松一下,随便聊聊。与此同时,GPT 5.6 则像是,会完全照做,我会把你说的每一件事都完成,什么都拦不住我。在这事做完之前我都不睡。它有点像,非常 OCD(强迫症倾向),你知道的,非常一丝不苟。
Speaker 121:58 - 22:31
But sometimes, you know, you want one where it's like, it actually realizes, hey, that list of 20 that you gave me, actually, here's a better way of doing it anyway. You know, 5.6 is more methodical. If you build a harness for each of those, there are actually different things that that harness will then be good or bad at. So for example, one thing that, you know, typically agents will do is they'll they'll have a to do list of, like, if you have a task, it'll go and generate a to do list. And the Claude code harness can, in some cases or and this is maybe less relevant now, but I think earlier, this is a just a more illustrative example.
Speaker 121:58 - 22:31
但有时候,你会想要这样一种模型:它真的能意识到,嘿,你给我的那 20 项清单,其实不管怎样,我这里有个更好的做法。你知道,5.6 更偏方法论、更按部就班。如果你为每一种情况都构建一个 harness(测试/约束框架),那么这个 harness 实际上也会各有擅长和不擅长的地方。比如说,agent(智能体)通常会做的一件事,就是它们会有一个 to do list(待办清单);如果你给它一个任务,它就会去生成一个待办清单。而 Claude code harness 在某些情况下可以——或者说,这一点现在也许没那么相关了,但我觉得在更早的时候,这是一个更有说明性的例子。
Speaker 122:32 - 23:02
Earlier, it was really strict to make sure it would stick to the to do list because the model itself would typically wander. Meanwhile, Codex wouldn't do that because the model itself was really, really OCD about that. But if you're a user, you wanna have the same experience regardless. Like, you wanna make sure if you switch to a different model, you're not gonna suddenly lose track of whatever things that you are working on. And so there are certain things where like, maybe in some cases you really want robust tool use and there are tools that you use to do these to do lists.
Speaker 122:32 - 23:02
更早的时候,它会非常严格地确保自己坚持那个 to do list,因为模型本身通常会跑偏。与此同时,Codex 不会这样,因为模型本身在这方面真的非常非常 OCD。但如果你是用户,你会希望无论如何都有一致的体验。比如,你会想确保如果切换到另一个模型,不会突然就丢掉自己正在处理的那些事情的脉络。所以在某些情况下,你可能确实非常需要稳健的 tool use(工具使用),而且你会用一些工具来处理这些 to do list。
Speaker 123:02 - 23:26
You want really robust tool use and you wanna make sure that no matter what, if I'm a user, I wanna see my to do list there. Like there were some cases where it would just like not have the to do list. And so that these are things that kind of improve the general performance and that the to do list matters because you're doing some crazy migration and you don't have the to do list. And then you're in this long session where there's compaction that might get lost in the summarization. And then now you forgot what your seventh step was, and that could be one of the failure modes.
Speaker 123:02 - 23:26
你会希望 tool use 非常稳健,而且无论如何,如果我是用户,我都想在那里看到我的 to do list。因为确实有些情况下,它就是会莫名其妙没有那个 to do list。所以这些东西其实是在提升整体表现。而 to do list 之所以重要,是因为你可能正在做某种很疯狂的 migration(迁移),结果你没有 to do list。然后你又处在一个很长的 session(会话)里,中间还有 compaction(压缩),这些内容可能会在 summarization(总结)时丢失。然后你就忘了自己的第七步是什么,而这就可能成为一种 failure mode(失败模式)。
Speaker 123:26 - 23:28
That's kind of an example of of how Good example.
Speaker 123:26 - 23:28
这算是一个关于“怎么会这样”的例子。很好的例子。
Speaker 223:28 - 23:29
Yeah. Yeah. Yeah.
Speaker 223:28 - 23:29
对。对。对。
Speaker 323:29 - 23:34
That's a great example. Okay. So we talked about one type of maxing, benchmark maxing. Let's talk about token maxing. Yes.
Speaker 323:29 - 23:34
这是个很好的例子。好,我们刚才谈了一种 maxing,也就是 benchmark maxing(基准测试最大化)。我们来谈谈 token maxing。好。
Speaker 323:34 - 23:40
Because it feels like the world has changed a lot. We've gone from token maxing to now cost rationalization. What does that mean for factory?
Speaker 323:34 - 23:40
因为感觉世界已经变了很多。我们已经从 token maxing(token 使用最大化)走到了现在的 cost rationalization(成本理性化)。这对 factory 意味着什么?
Speaker 123:41 - 24:00
Yeah. So maybe I'll I'll lay this out just to so we're all on the same page of, like, the way that we see what what's led us to this token maxing. So loosely, there was, like, this phase one where maybe phase zero was, like, no one believed in AI. Then phase one, everyone believes in AI. And then boards were like, mister CEO, what are you doing about AI?
Speaker 123:41 - 24:00
对。所以也许我先把这个脉络讲清楚,这样大家对我们如何看待、以及是什么把我们带到这种 token maxing 上,能有一致理解。大致来说,有这么一个 phase one,或者说 phase zero 可能是:当时没人相信 AI。然后到 phase one,所有人都相信 AI。接着董事会就会问:mister CEO,你在 AI 这件事上到底做了什么?
Speaker 124:00 - 24:07
What's your AI strategy? And mister CEO is like, shit. I don't know. Like, what's our AI strategy? CTO, like, make sure everyone goes and uses AI.
Speaker 124:00 - 24:07
你的 AI strategy(AI 战略)是什么?然后 mister CEO 就会想,妈的,我也不知道。比如,我们的 AI strategy 到底是什么?于是 CTO 就会说,确保所有人都去使用 AI。
Speaker 124:07 - 24:20
And so then phase two is, you know, CTO is like, okay. Shit. We gotta make sure everyone uses AI. Let's start putting it in performance reviews. Let's make public, like, bench or public, like, rankings of who's using tokens the most, because everyone's stubborn.
Speaker 124:07 - 24:20
然后到了 phase two,就是 CTO 会说,好吧,妈的,我们得确保每个人都用 AI。那就开始把它写进 performance reviews(绩效评估)吧。再公开一些类似 benchmark(基准)或公开排名,看看谁用的 tokens 最多,因为大家都很固执。
Speaker 124:20 - 24:30
No one wants to use this stuff. They're all skeptical. And then we enter phase three, which is everyone sees these ratings. They see that it's part of their perf reviews, and they're like, okay. I'm gonna use AI for everything.
Speaker 124:20 - 24:30
没人想用这些东西。大家都持怀疑态度。然后我们进入 phase three,这时所有人都看到了这些评分,也看到这已经成了他们 perf reviews(绩效考核)的一部分,于是他们会想,好吧,那我什么都用 AI 来做。
Speaker 124:30 - 24:40
And that's kind of phase three. It's this token maxing where people are using, like, Opus for literally everything. Like, what's the weather in SF? Opus. Tell me.
Speaker 124:30 - 24:40
这差不多就是 phase three。也就是这种 token maxing:人们真的把 Opus 用在一切事情上。比如,SF 今天天气怎么样?Opus,告诉我。
Speaker 124:40 - 25:03
I don't know. Like, there are banks that we are working with where they are spending literally hundreds of thousands of dollars a month on people asking things like, literally, what is the weather? Or, like, tell me about Python. Like, trivial questions that you could Google, people are asking Opus. And the reality is this happened because we were so worried about adoption that we overcorrected and we're like, adoption by any means necessary.
Speaker 124:40 - 25:03
我也不知道。比如有些我们正在合作的银行,他们每个月真的会花几十万美元,让员工去问一些字面意义上像“今天天气怎么样”这样的问题。或者“给我讲讲 Python”。这种你 Google 一下就能知道的琐碎问题,人们却在问 Opus。现实是,会出现这种情况,是因为我们太担心 adoption(采用率)了,于是矫枉过正,变成了“不惜一切代价推动 adoption”。
Speaker 125:04 - 25:38
And I think that's actually it's like a decent approach. Like, it's probably faster to do that and then curb usage or or make usage more responsible than it is to start limited and be like, you know, you can only use it for this thing. Because when you have people that are stubborn, first, you wanna just prove that it works, and then you can get kinda more mature about it. Where factory fits in, I think one of the most important things that we do is that we have the factory router, which allows you to dynamically route to different models based on the task that you're doing. So, if you you're asking what the weather is, you probably don't need the very frontier of human intelligence to answer that for you.
Speaker 125:04 - 25:38
而我觉得,这其实算是一种还不错的方法。也就是说,先这么做、然后再去抑制使用,或者让使用变得更负责任,可能比一开始就设很多限制、说“你只能把它用在这件事上”要更快。因为面对那些很固执的人时,你首先想做的是证明这东西确实有效,然后你才能在这件事上变得更成熟。至于 factory 在这里如何发挥作用,我认为我们做得最重要的一件事之一,就是我们有 factory router,它可以根据你正在处理的任务,动态地路由到不同的模型。所以,如果你只是在问天气如何,那你大概不需要动用人类智能最前沿的模型来回答这个问题。
Speaker 325:38 - 25:39
Or or you really do.
Speaker 325:38 - 25:39
或者你其实真的会这么做。
Speaker 125:39 - 26:13
Mean, it depend I don't know. Depends on what kind of answer you're looking for, giving you a full, like, down to the molecular level of what's happening, But, you know, allowing that. But also more importantly for every enterprise, something that no one's dealing with yet, but twelve months from now is going to be the case is, not everyone needs the same tokens. Having a blanket kind of token cap for every individual in some large bank, let's say, makes no sense. So every CIO is gonna need to answer for every incremental token, where do we put it?
Speaker 125:39 - 26:13
我的意思是,这要看情况,我也不知道。取决于你想要什么样的答案——是想要一个完整的、比如一直细到分子层面到底发生了什么的解释;当然,也可以那样讲。但另外更重要的是,对每一家 enterprise(企业)来说,有件事现在还没人真正处理,但十二个月后就会变成现实:不是每个人都需要同样数量的 token。比如说,在某家大型银行里,给每个人都设一个一刀切的 token 上限,这根本说不通。所以每个 CIO 都得回答:每新增一个 token,我们该把它投到哪里?
Speaker 126:14 - 26:42
And right now, it is super not obvious how you would do that. Like, right now we're saying, oh, you know, the PMs who are, like, vibe coding dashboards get the same token limits as, like, the engineers who are building, like, critical infrastructure. That's probably not the best thing to do. Or similarly, you might be dealing with COBOL code bases where Opus is not the best model to use, but instead maybe some fine tuned model on that code base in particular. The point of the router is that we can kind of accommodate these different constraints where maybe you say, you know what?
Speaker 126:14 - 26:42
而现在,怎么做这件事其实一点都不明显。比如现在我们会说,哦,你知道,那些在随手 vibe coding 做 dashboard 的 PM,拿到的 token 限额,和那些在搭建关键基础设施的工程师是一样的。这大概率不是最佳做法。再比如,你可能在处理 COBOL code base,这时 Opus 未必是最适合的 model(模型);相反,也许某个专门针对那套代码库 fine-tuned(微调)过的模型更合适。router(路由器)的意义就在于,我们可以某种程度上适配这些不同的约束条件——比如你可以说,知道吗?
Speaker 126:42 - 26:53
This part of the org, they're just vibe coding. They can use Gemini Flash. This part of the org, they're doing COBOL. We fine tuned this great model to work on COBOL. Let's route to that when we're working on that part of the code base.
Speaker 126:42 - 26:53
组织里的这一部分人,他们就是在 vibe coding。那他们可以用 Gemini Flash。组织里的这一部分人,他们在做 COBOL。我们专门 fine-tuned 了一个很棒的模型来处理 COBOL。那当我们在处理那部分代码库时,就把请求路由到那个模型。
Speaker 126:53 - 27:14
Maybe this other part, we really care about reliability, so let's generate the code with OpenAI, test it with Anthropic, review it with Gemini, things like that. And we can actually take in your routing procedure instructions in natural language. So you could even say things like it's not purely deterministic. It can even be like, hey. You know, Pat, I don't know.
Speaker 126:53 - 27:14
也许还有另一部分,我们非常看重 reliability(可靠性),那就用 OpenAI 生成代码,用 Anthropic 做测试,用 Gemini 做审查,诸如此类。而且我们实际上可以用 natural language(自然语言)接收你的路由流程指令。所以你甚至可以说,这不一定是纯粹 deterministic(确定性)的。它甚至可以像是,嘿,你知道,Pat,我也不清楚。
Speaker 127:14 - 27:15
Like, I don't know what he's doing.
Speaker 127:14 - 27:15
比如,我不知道他在做什么。
Speaker 327:15 - 27:17
Like Give tag Gemini Flash.
Speaker 327:15 - 27:17
比如,给他打上 Gemini Flash 这个 tag(标签)。
Speaker 227:17 - 27:17
Give him
Speaker 227:17 - 27:17
给他。
Speaker 127:17 - 27:17
Flash. I
Speaker 127:17 - 27:17
突然想到。我
Speaker 327:17 - 27:18
don't know.
Speaker 327:17 - 27:18
不知道。
Speaker 127:18 - 27:32
Or, you know, I think we really need to avoid having them use open models because, you know, whatever reason, we don't like the way open models perform here. And we'll do internal benchmarking to know which models are better at which of these tasks.
Speaker 127:18 - 27:32
或者,你知道,我觉得我们确实需要避免让他们使用 open models,因为,怎么说呢,出于各种原因,我们不喜欢 open models 在这里的表现。我们会做内部 benchmarking(基准测试),来弄清楚哪些模型更擅长这些任务中的哪些部分。
Speaker 327:32 - 27:35
How close are the open models at this point? Which one's the best?
Speaker 327:32 - 27:35
现阶段 open models 差得还有多远?哪个是最好的?
Speaker 127:35 - 27:48
GLM 5.2 is incredible. It's at the point where internally we have no token limits for our engineers, and, like, half of our tokens are open to open models. Wow. Yeah. Because they're just faster, and they're cheaper.
Speaker 127:35 - 27:48
GLM 5.2 非常惊人。现在已经到了这样一个程度:我们内部对工程师基本没有 token 限制,而且,大概有一半的 token 都开放给 open models。哇。对。因为它们就是更快,而且更便宜。
Speaker 127:48 - 28:03
They're just as performant. And I think the thing that everyone gets wrong is everyone is comparing, like, GLM 5.2 to the latest model, like OPUS 4.8 or GPT 5.6. But, really, they should be compared to OPUS 4.7 or GPT 5.5.
Speaker 127:48 - 28:03
它们的性能也一样强。我觉得大家普遍搞错的一点是,所有人都在拿 GLM 5.2 和最新模型相比,比如 OPUS 4.8 或 GPT 5.6。但实际上,它更应该和 OPUS 4.7 或 GPT 5.5 来比较。
Speaker 328:04 - 28:04
Why?
Speaker 328:04 - 28:04
为什么?
Speaker 128:05 - 28:23
Because, generally, the open models come later, and they're they're kind of a generation behind. And that's kind of the the the frontier models will be frontier. The question is, are the open models getting as good as, like, frontier minus one? And the answer is unequivocally yes, which I think is a really, really interesting outcome. It's great for consumers.
Speaker 128:05 - 28:23
因为一般来说,open models 都会晚一些出来,它们某种意义上会落后一代。这也算是常态——frontier models(前沿模型)终究会是前沿模型。问题在于,open models 是否已经好到可以达到“frontier 减一代”的水平?答案是毫无疑问的:是。我觉得这是一个非常、非常有意思的结果。这对消费者来说是好事。
Speaker 128:23 - 28:57
And by consumers, I don't mean like individuals. I mean, the consumers of the APIs, because if you're a, you know, a business that is doing in our software engineering, your job is at a very high level to solve problems. And if we can allow you to solve those problems faster and with cheaper models that are just as performant, that means you can solve more problems. Like, that is a good thing. And it is a very good world where there is not, like, a monopoly on intelligence, but instead kind of a a a garden of intelligence that you can pick and choose, you know, when you'd like.
Speaker 128:23 - 28:57
我这里说的 consumers 不是指个人。我指的是 API 的消费者,因为如果你是一个从事软件工程的企业,你的工作在很高层面上就是解决问题。如果我们能让你用更快、且同样有性能表现但更便宜的模型来解决这些问题,那就意味着你可以解决更多问题。这是好事。一个没有 intelligence(智能)垄断、而是更像一个 intelligence(智能)花园、你可以按需挑选的世界,是一个非常好的世界。
Speaker 128:57 - 29:25
It's something that we joke about is, like, you know, on this intelligence allocation thing, if you're if you're trying to get a a tutor for your daughter in algebra, you can probably find someone cheaper than Albert Einstein to be that tutor. Now it might be that she eventually goes and becomes, like, a leading, you know, physicist or something, in which case, yeah, maybe let's let's get Albert Einstein in there. But most likely, you can get, you know, a high school student or something like that. And it's probably much more cost effective for you as well-to-do so. So
Speaker 128:57 - 29:25
我们会拿这个开玩笑:说到这种 intelligence 分配,如果你想给你女儿找一个 algebra(代数)家教,你大概率能找到一个比 Albert Einstein 更便宜的人来做这件事。她以后当然也可能真的成为某种顶尖 physicist(物理学家),那样的话,嗯,也许我们确实该把 Albert Einstein 请来。但更可能的是,你找个高中生之类的人就够了。而且对你来说,这样做的成本效益很可能也高得多。所以——
Speaker 229:26 - 29:38
since you guys do the model routing, like if you look at the, you know, if there's a pie chart that shows the complexion of models being used by your customer base today, what did it look like a few months ago? What does it look like today? What do you think it'll look like in a year?
Speaker 229:26 - 29:38
既然你们会做 model routing(模型路由),如果你看一张饼图,显示你们客户群体如今所使用模型的构成,那几个月前它是什么样?今天又是什么样?你觉得一年后会是什么样?
Speaker 129:38 - 29:51
Yeah. I will caveat this with saying that right now enterprises haven't gone too opinionated yet into the routing procedures. Okay. This is something that will happen over the next six to twelve months. But right now they're just going from no router to router.
Speaker 129:38 - 29:51
对。不过我要先补充一句:目前 enterprises(企业)在 routing procedures(路由流程)这件事上还没有形成太强的主观策略。好吧。这会是在接下来六到十二个月里发生的事。但现在,他们只是从“没有 router(路由器)”变成“有 router”。
Speaker 129:51 - 30:09
That's kind of the first change. Then it's going to be like the exact nature of the, of the routing. At the beginning of the year, there's less than 1% of tokens went to open models. In the first quarter, it became a single digit percent. It is now crossed into being a double digit percent of tokens.
Speaker 129:51 - 30:09
这是第一层变化。接下来变化的会是 routing(路由)本身的具体方式。年初的时候,流向 open models(开放模型)的 token(词元)还不到 1%。到了第一季度,它变成了个位数百分比。现在已经跨过了两位数百分比这个门槛。
Speaker 130:10 - 30:17
Percent Now of tokens is not always the same as percent of cost because the open tokens are cheaper, but it is it's pretty crazy to see the the growth there.
Speaker 130:10 - 30:17
不过,token(词元)占比不一定总是等于成本占比,因为 open token 更便宜,但看到那边的增长还是相当惊人。
Speaker 330:17 - 30:18
Yeah. What's your forecast?
Speaker 330:17 - 30:18
嗯。你的预测是什么?
Speaker 130:20 - 30:52
My sense is that we will asymptote towards vast majority being open just because it provides you more optionality and it's cheaper. But that doesn't mean they're gonna be like, that's of token share, not necessarily of leverage share. Because maybe there are 1% of tokens that are incredibly, incredibly valuable and are, like, very key decision making, and then the rest are more, like, implementation tokens or kind of lower stakes, if you will. I don't think there's gonna be a world in which, like, it's ever gonna be a 100%.
Speaker 130:20 - 30:52
我的感觉是,我们会渐近到“绝大多数都是 open”的状态,仅仅因为它给你更多 optionality(可选性),而且更便宜。但那并不意味着它们会——这里说的是 token share(词元份额),不一定是 leverage share(杠杆份额/关键价值份额)。因为也许有 1% 的 token 极其、极其有价值,承担的是非常关键的决策,而其余的更像是 implementation tokens(执行类词元),或者说是风险更低的部分。我不认为会出现一个世界,让它最终变成 100%。
Speaker 330:52 - 30:52
Yep.
Speaker 330:52 - 30:52
对。
Speaker 130:52 - 31:42
Think the frontier of intelligence will inherently always be valuable for every business just because the stakes are going to get higher and the kind of intricacy with which you think is going to be more important, but we'll be better at offloading certain tasks. And this is like, you can loosely think of this already with the way orgs are structured, where, you know, in general, engineering leaders are more tenured engineers who in theory have, like, more wisdom and each kind of minute of their brain power is higher leverage in theory. And even you know, you can also imagine, like, consider a human engineer and try mapping over the course of their day, like, how much brain power they're using. And, like, you know, it's probably gonna be really low for a lot of it, but then there are gonna be some moments where they're, like, going pretty high. Like, they're deeply concentrating and thinking about some, you know, systems design problem or whatever.
Speaker 130:52 - 31:42
我认为,智能的前沿能力对每一家企业来说天然都会始终有价值,因为利害关系会越来越高,而你思考问题时所需的那种复杂程度也会变得更重要;但与此同时,我们也会更擅长把某些任务卸载出去。这个其实已经可以从组织结构中粗略理解:一般来说,engineering leaders 往往是资历更深的工程师,理论上拥有更多智慧,而且他们每一分钟的脑力从理论上说都有更高的杠杆效应。你也可以设想一下,一个 human engineer 在一天当中,脑力使用强度是如何分布的。很可能其中大部分时间都比较低,但也会有一些时刻非常高强度——比如他们在高度专注地思考某个 systems design 问题之类的事情。
Speaker 131:42 - 32:08
All of those low leverage moments, we wanna automate away. And, like, we wanna like, those, like, very high leverage moments, sometimes, like, you know, we're referring to them as, like, the eureka moments or the moments where they're, like, doing something that's very high leverage. What if those aren't just moments, but what if those are, like, hours at a time? Because you don't have to deal with all the other stuff. And I think that's kind of the the way to think about intelligence allocation is if you're an engineer and you're writing docs, that is such a low leverage use of your time.
Speaker 131:42 - 32:08
所有那些低杠杆的时刻,我们都希望把它们自动化掉。然后,那些高杠杆时刻——有时我们会把它们称为 eureka moments(灵光一现的时刻),或者说他们在做某种极高杠杆事情的时刻——如果这些不再只是片刻,而是能持续几个小时呢?因为你不需要再处理其他那些杂事了。我认为,这就是理解 intelligence allocation(智能分配)的一种方式:如果你是工程师,却在写 docs,那么这对你的时间来说就是一种杠杆极低的使用方式。
Speaker 132:08 - 32:32
Like, you've become an expert in your craft and you used to spend hours writing docs. Like, I remember it was actually valuable. Like, I remember Stripe had so much alpha for just having incredible docs. But imagine all the other stuff those incredible engineers could do if it wasn't writing documentation. Like, we should live in a world where everyone can have docs as good as Stripe, and that is, like, strictly beneficial for everyone.
Speaker 132:08 - 32:32
你已经成了自己专业领域里的专家,却还要花上好几个小时写 docs。就像我记得,这件事过去其实是有价值的。我记得 Stripe 当年仅仅因为拥有极其出色的 docs,就获得了很大的 alpha(超额优势)。但试想一下,如果不用写 documentation,那些出色的工程师本来还能去做多少别的事情。我们应该生活在这样一个世界里:每个人都能拥有和 Stripe 一样好的 docs,而这对所有人来说都只会是纯粹的好事。
Speaker 132:32 - 32:36
And then the question is, okay, what do those really smart engineers do with their time once they don't have to do that?
Speaker 132:32 - 32:36
接下来的问题就是,好,那么这些真正聪明的工程师,在不用再做这些事之后,要把时间花在什么地方?
Speaker 332:37 - 32:53
Maybe this is good time to talk about business model, given that, you know, especially the rise of open weight models, the cost differential, I imagine that means very different things for your cost structure, but very similar value delivered to customers. How do you think about business model and pricing?
Speaker 332:37 - 32:53
也许现在正是谈 business model(商业模式)的好时机,因为你知道,尤其随着 open weight models 的兴起,成本差异我想会对你们的成本结构产生非常不同的影响,但交付给客户的价值却可能非常相似。你们是如何思考 business model 和 pricing(定价)的?
Speaker 132:53 - 33:10
Yeah, this is more what our customers want and need as opposed to what we want and need. So for example, I think right now usage based is clearly the way to go. We wanna be aligned with like what they are doing and what we are doing. I think seat based doesn't make sense at least for what we are doing. My sense is that eventually we will change to outcome based.
Speaker 132:53 - 33:10
对,这更多取决于我们的客户想要什么、需要什么,而不是我们想要什么、需要什么。比如说,我认为目前 usage based(按使用量计费)显然是正确的方式。我们希望和他们在做的事、以及我们在做的事保持一致。我觉得 seat based(按席位计费)至少对我们现在做的事情来说并不合理。我的感觉是,最终我们会转向 outcome based(按结果计费)。
Speaker 133:11 - 33:25
Now I don't think the enterprise is ready for that. And we've learned our lesson from those first two years. We are not going to impose things. Right? But my suspicion is that, you know, in the twenty thirties, things will probably look more like outcome based.
Speaker 133:11 - 33:25
但我现在并不认为 enterprise(企业客户)已经为此做好准备。我们也从最初那两年的经历中吸取了教训。我们不会强行推行某些东西,对吧?不过我的判断是,到了 2030 年代,情况大概率会更像是 outcome based。
Speaker 233:25 - 33:29
Yeah. What does outcome based mean for your market? What would be the definition of an outcome? So
Speaker 233:25 - 33:29
对。那么对你们这个市场来说,outcome based 意味着什么?outcome 的定义会是什么?
Speaker 133:30 - 34:04
maybe here's a way to put it. So right now we chart we we are usage based. Like, more tokens you use, you know, the more you pay, the more we get. Now, since we are model independent, we kind of, with our router, we are kind of pointing a token cannon at either OpenAI, Anthropic, AWS, GCP, you know, any one of these people. To a certain degree, this is like a really dumbed down version of a marketplace where right now there is a a buy the buy side is an engineer who wants a task done.
Speaker 133:30 - 34:04
也许可以这样来表达。我们现在的收费方式是 usage based(按使用量计费)。比如,你用的 token 越多,你付的钱就越多,我们赚到的也越多。由于我们是 model independent(模型无关)的,所以通过我们的 router,我们有点像是在把一门 token cannon 指向 OpenAI、Anthropic、AWS、GCP,或者这些供应方中的任何一家。在某种程度上,这很像一个被大幅简化了的 marketplace(市场):现在的 buy side(买方)是一个想把某项任务完成的工程师。
Speaker 134:04 - 34:15
And then you have the model providers who are saying like, either in benchmarks right now, they're like, we perform at this cost and this performance, and then we determine who we go to for that given task.
Speaker 134:04 - 34:15
然后另一边是 model providers(模型提供方),他们会说,比如从当前的 benchmarks(基准测试)来看,我们能以这样的成本提供这样的性能,接着我们再根据这个具体任务来决定把任务交给谁。
Speaker 234:15 - 34:15
Yeah.
Speaker 234:15 - 34:15
对。
Speaker 134:15 - 34:39
There's a world in which, you know, if it's so important to get these tokens, they might kind of like, bid in a certain way of saying like, look, here is our cost for this task. We will get this task done at this cost no matter what, but they're pricing it such that, you know, they hope that they can make a margin there. Yeah. They price it wrong, they're at a negative margin. If they price it right and win the bid, then they get the positive margin.
Speaker 134:15 - 34:39
有一种可能的世界是,如果拿到这些 token 真的如此重要,他们可能会以某种方式进行 bidding(竞价),相当于说:看,这是我们完成这个任务的报价。无论如何,我们都会以这个成本把任务做完;但他们的定价会使他们希望自己还能从中赚到利润。对。如果他们定价错了,利润率就是负的;如果他们定价对了并赢得竞标,那他们就能拿到正利润。
Speaker 134:39 - 34:51
And the way you determine if the task was successful is by some validation loops. Because no one is using these tools anymore where it's just like, write me code. Great. Thank you. It's generally write me code, and here's how I know it was done well.
Speaker 134:39 - 34:51
而判断任务是否成功的方法,是通过一些 validation loops(验证循环)。因为现在已经没人再用那种只是“给我写段代码”“好,谢了”这种工具了。通常的情况是“给我写段代码,而且这是我判断它是否写得好的标准”。
Speaker 134:51 - 35:22
And similarly, if you are like a model lab and you are given, here's a task, here's the validation criteria, you'll be able to say roughly how much you think you would be willing to pay to get those tokens. And, you know, you wanna have some some margin on that. And then in that world, that's basically that's a way that that you kind of dynamically shift from usage based to outcome based. I think that there are so many questions with this, and this is very much forward looking. But I think there's a lot of questions about how do you subdivide tasks, you know, divvying that up, I think is something that's not obvious.
Speaker 134:51 - 35:22
类似地,如果你是一个 model lab(模型实验室),有人给你一个任务和对应的 validation criteria(验证标准),你就能大致判断,为了拿到这些 token,你愿意付出多大成本。当然,你也会希望里面留出一些利润空间。这样一来,这基本上就提供了一种从 usage based(按使用量计费)动态转向 outcome based(按结果计费)的方式。我认为这里面还有非常多的问题,这也非常偏向未来视角。但我觉得其中一个很大的问题是,如何把任务拆分开来,你知道,怎么去分配这些部分,我认为这并不是显而易见的。
Speaker 135:22 - 35:25
Yeah. But as these tools get better, doing things like that actually become way easier.
Speaker 135:22 - 35:25
对。不过随着这些工具变得更强,做这类事情实际上会容易得多。
Speaker 235:25 - 35:32
Yeah. That's fascinating. Yeah. Yeah. If you can scope a task and then create a competitive marketplace, that'd be a fascinating version of the future.
Speaker 235:25 - 35:32
对。这很有意思。是的,是的。如果你能把任务界定清楚,然后建立一个有竞争性的 marketplace,那会是一种很有意思的未来形态。
Speaker 135:32 - 35:42
Yes. And as a user, it then creates an incentive to be very thorough in your validation criteria. Yeah. Because like, you know, there are stories of like, you you ask an agent to like fix my code and it deletes your code. Yeah.
Speaker 135:32 - 35:42
是的。而且对于用户来说,这也会形成一种激励,让你把 validation criteria(验证标准)写得非常详尽。对。因为你知道,确实有这种故事:你让一个 agent 去“修复我的代码”,结果它把你的代码删掉了。对。
Speaker 135:42 - 35:44
It's like, you know, the solution is just get rid of it all together. It's like
Speaker 135:42 - 35:44
这就像,你知道,解决办法就是把它整个都去掉。就像——
Speaker 335:44 - 35:45
that Silicon Valley episode.
Speaker 335:44 - 35:45
那一集 Silicon Valley。
Speaker 135:45 - 35:46
You can
Speaker 135:45 - 35:46
你可以
Speaker 335:46 - 35:47
see how precious it was.
Speaker 335:46 - 35:47
看出它当时有多珍贵。
Speaker 135:47 - 35:51
Yeah. But, like, so you need to make sure your tests are very thorough because technically, it could hit all of
Speaker 135:47 - 35:51
对。但话说回来,你得确保你的测试非常彻底,因为从技术上讲,它可能会命中你所有的——
Speaker 335:51 - 35:53
your maladies. Will go rogue.
Speaker 335:51 - 35:53
各种病症。会失控。
Speaker 135:53 - 35:54
Yeah. Exactly. Yeah. Yeah. Yeah.
Speaker 135:53 - 35:54
对。没错。对。对。对。
Speaker 135:54 - 35:56
Yeah. Good reference.
Speaker 135:54 - 35:56
对。这个引用很好。
Speaker 335:57 - 36:03
Maybe zooming out a little bit. You named the company Factory. Actually, you named it Droid before before Factory.
Speaker 335:57 - 36:03
也许稍微把视角拉远一点。你把公司命名为 Factory。其实,在 Factory 之前,你先把它命名为 Droid。
Speaker 136:03 - 36:03
That's
Speaker 136:03 - 36:03
那就是
Speaker 336:03 - 36:16
right. You But named it Factory before this concept took off, and now it feels like everybody wants to build a software factory. Where do you think we are today in in terms of the building of software factories, and how close are we to the ultimate vision of a software factory?
Speaker 336:03 - 36:16
对。你在这个概念真正火起来之前就把它命名为 Factory,而现在感觉好像每个人都想打造一个 software factory(软件工厂)。你觉得我们今天在构建 software factory 这件事上处于什么阶段?我们距离 software factory 的终极愿景还有多远?
Speaker 136:16 - 36:32
Yeah. Everyone has a software factory whether they know it or not. It's just a very inefficient one. So it's kind of like it feels like, you know, pre industrialization where, like, you know, people were manually, like, you know, sewing things together or, like, woodworking or whatever it might be. And these things are very inefficient.
Speaker 136:16 - 36:32
是的。每个人其实都有一个 software factory,只是他们自己未必意识到,而且那通常是个效率很低的工厂。这有点像工业化之前的时代,人们靠手工把东西缝起来,或者做木工之类的。不管具体是什么,这些方式都非常低效。
Speaker 136:32 - 37:04
Like, right now, if you go to an organization that has more than 10,000 people and you're to ask about the process by which they decide and release a feature, there is, like, hundreds or maybe thousands of people in that process. And most likely, they couldn't even draw it for you. Like, there's very low likelihood that they would know what what that process looks like. That is not because they think that is the right way of doing things. That is just kind of the nature of building large software as it is kind of today.
Speaker 136:32 - 37:04
比如现在,如果你去一家拥有超过一万名员工的组织,问他们决定并发布一个功能的流程是什么,那么这个流程里会涉及几百人,甚至上千人。而且很可能,他们甚至没法把这个流程画给你看。也就是说,他们其实很大概率并不真正清楚这个流程到底长什么样。这不是因为他们觉得这就是正确的做事方式,而只是因为在当下,构建大型软件本身往往就是这样一种状态。
Speaker 137:04 - 37:22
But with these systems, so much tribal knowledge can be codified. So much of this stuff that typically would require, oh, we need to ask this guru who's been here for thirty years who has the wisdom. Oh, we then need this approval and that approval. Oh, and I forgot there was some doc that said we always have to do this checklist. And it relies so much on kind of human behavior and redundancy.
Speaker 137:04 - 37:22
但有了这些系统,很多 tribal knowledge(部落知识、隐性经验)都可以被 codify(编码化、明确化)。很多原本通常需要“哦,我们得去问那个在这里待了三十年的 guru,他最有经验”“哦,然后还需要这个审批、那个审批”“哦,我忘了还有份文档说我们总是得走这份 checklist(检查清单)”之类的事情,其实都高度依赖人的行为以及大量重复环节。
Speaker 137:24 - 37:49
So much of that can be automated and refocused on like what actually moves the needle for our business. And I think this move towards software factories is a move towards how do we figure out what are the actual inputs that determine what features we need to build? And that might be inputs from the customers, inputs from the market, inputs from, like, you know, product leaders at the company. And let's be very clear. These are the signals, the inputs that we are taking in here.
Speaker 137:24 - 37:49
其中很大一部分都可以被自动化,并重新聚焦到那些真正能推动我们业务增长的事情上。我认为,向 software factory 的转变,本质上是在思考:到底哪些真实输入决定了我们需要构建什么功能?这些输入可能来自客户,来自市场,也可能来自公司内部的产品负责人。而且我们要非常明确:这些就是我们在这里接收和使用的 signals(信号)与 inputs(输入)。
Speaker 137:49 - 38:04
Okay. Great. We have those signals. Then what is the process by which we build this? And really like mapping out the assembly lines of how you are building software is really important because then you get to close the loop and say, did this actually deliver outcome for our business?
Speaker 137:49 - 38:04
好,很好。我们已经有了这些信号。接下来,构建它的流程到底是什么?真正把你构建软件的 assembly lines(装配线)映射清楚,是非常重要的,因为这样你才能闭环,并判断:这件事到底有没有为我们的业务带来结果?
Speaker 138:04 - 38:37
Talking before about the tokenomics, if you're that CIO and you're faced with that question of where do you put every incremental token, really the question two years from now is gonna become, where do you put every incremental dollar? And so you're gonna have to be asked, do you put that incremental dollar towards head count or towards tokens? And if tokens to where in the org? And these are things that you can only really know when you have these kind of feedback loops that give you examples of like, Hey, by the way, we made those decisions based on this data and it did not matter at all. We added these new features and no one cared.
Speaker 138:04 - 38:37
前面谈到 tokenomics(token 经济模型)时,如果你是那个 CIO,并且面临“每多出来一个 token(代币/计量单位)该投到哪里”的问题,那么两年后,这个问题真正会变成“每多出来一美元该投到哪里”。到时候你必须回答:这笔增量预算是投向 head count(人员编制),还是投向 tokens?如果投向 tokens,那又应该投到组织里的哪个地方?而这些问题,只有在你拥有这类 feedback loops(反馈回路)时,你才真正有可能知道答案——它们会给你具体例子,比如:“顺便说一句,我们当时基于这些数据做了那些决策,结果根本没有影响。我们增加了这些新功能,但没有人在乎。”
Speaker 138:37 - 38:59
It didn't create more retention. It didn't create more usage or whatever metrics that business is looking to optimize. And the only way to do this is like, you need kind of more rigor and more process. It almost feels like, like ten years from now, we're gonna look back at this previous era of software and it's gonna feel like businesses in like ancient times where they didn't do accounting. It is like It's gonna be
Speaker 138:37 - 38:59
它并没有带来更高的 retention(留存)。也没有带来更多使用量,或这类企业想要优化的其他 metrics(指标)。而要做到这一点,唯一的办法就是你需要更多的严谨性和更多流程。感觉几乎像是,十年后我们回头看这个上一代软件时代时,会觉得那时的企业就像古代那些不做 accounting(会计)的商号一样。那会变得像是
Speaker 238:59 - 39:04
like, it's gonna be like marketing in the day of mad men. Yes. Where it's like all creative and you have no idea what's actually working.
Speaker 238:59 - 39:04
就像 Mad Men 时代的 marketing(营销)一样。对。那种一切都靠创意驱动,而你根本不知道到底什么真正有效。
Speaker 139:04 - 39:10
It makes no like, it's like, oh, yeah. Let's ship that feature. Oh, I think it went well. Like, yeah, we had I got some metrics on that. Yeah.
Speaker 139:04 - 39:10
这完全说不通。就像,哦,对,我们把那个功能上线了。哦,我觉得效果不错。对,我们确实——我拿到了一些相关 metrics(指标)。对。
Speaker 139:10 - 39:17
It's like, no. If you guys read the the blog post that Jack Dorsey put out Yeah. About how every company is like an AGI
Speaker 139:10 - 39:17
应该是,不。如果你们读了 Jack Dorsey 发的那篇 blog post(博客文章),对,里面讲到每家公司都像一个 AGI
Speaker 239:17 - 39:17
Yeah.
Speaker 239:17 - 39:17
对。
Speaker 139:18 - 39:34
There's also this degree to which if your company is an AGI, you wanna optimize the weights. Yeah. You wanna figure out what nodes are doing what things, which are load bearing, which are not, which need more tokens. Where do you need more nodes? And in order to do like, you don't train a model by vibes.
Speaker 139:18 - 39:34
还有一点是,如果你的公司本身就是一个 AGI,你就会想去优化它的 weights(权重)。对。你会想弄清楚哪些 nodes(节点)在做哪些事情,哪些是 load bearing(承重/关键支撑)的,哪些不是,哪些需要更多 tokens(资源配额),哪里需要更多 nodes。因为要做这些——你不会靠感觉去训练一个模型。
Speaker 139:35 - 39:53
I mean okay. Actually, you kind of do. But you don't I guess, more importantly, you don't do back prop in a model by vibes. Like, you are running those actual, like, calculations, and you are seeing when we change this node, what happens. Now you might be making bets on how to change the model by vibes, but you're like, you're it's pretty, like, mathematical in what you were doing.
Speaker 139:35 - 39:53
我的意思是,好吧,其实某种程度上你确实会。但我想更重要的是,你不会靠感觉去对一个模型做 back prop(反向传播)。你是在运行那些真实的计算,而且你会看到:当我们改动这个 node 时,会发生什么。现在,你也许会凭感觉去下注,决定该如何改这个模型,但你实际在做的事情本身是相当数学化的。
Speaker 139:53 - 40:20
Meanwhile, at companies, you know, people are determining token budgets just by shooting from the hip. People are laying people off by shooting from the hip and just being like, oh yeah, like 20,000. There is no way there is science to laying off 20,000 people. That is just like, here's a chunk and let's just see what happens. Instead, I think in these organizations, the way they can do things is much more mathematical of like this part of the business matters a lot and does better if we give it more tokens.
Speaker 139:53 - 40:20
与此同时,在公司里,你知道,人们却只是凭拍脑袋来决定 token budgets(token 预算)。人们也是凭拍脑袋裁员,只是说,哦对,就 20,000 吧。裁掉 20,000 人这件事不可能有什么科学依据。那纯粹就是,先砍这么一大块,再看看会发生什么。相反,我觉得这些组织做事的方式本来可以更数学化一些,比如:业务的这一部分非常重要,而且如果我们给它更多 tokens,它就会做得更好。
Speaker 140:20 - 40:39
It doesn't actually matter if we give it more humans. So let's give them more tokens. There might be other parts of the business where actually giving them more tokens doesn't matter, but more people matter. Because if we build more relationships with our customers and deeper relationships with our customers, that matters. But these are things that we're gonna need, like, quantitative insight on, and you need a software factory to do that.
Speaker 140:20 - 40:39
其实给它更多 humans(人)并不重要。那我们就给他们更多 tokens。也可能业务里还有别的部分,实际上给更多 tokens 并不重要,但更多人却很重要。因为如果我们和客户建立更多关系、建立更深的关系,那是重要的。但这些事情都需要某种 quantitative insight(量化洞察),而要做到这一点,你需要一个 software factory(软件工厂)。
Speaker 140:39 - 40:44
Otherwise, you're just, like, shooting from the hip and just guessing, which won't work as well.
Speaker 140:39 - 40:44
不然的话,你基本上就是在凭感觉开枪、瞎猜,而那样效果不会太好。
Speaker 340:44 - 40:50
In the limit, how much do you think people will spend on tokens versus on engineering headcount?
Speaker 340:44 - 40:50
从长期极限来看,你觉得人们在 token 上的花费,相比在 engineering headcount(工程团队人数)上的花费,会是多少?
Speaker 140:50 - 41:25
It'll depend on the business. I think every business will have a balance and it just depends on like, like, they're just gonna be like, an easy example is generally salespeople, they probably don't need that many tokens if they're good salespeople. Because, generally, the where they provide the most alpha is, like, when they're in the seat face to face with their customers, talking about the customer's problems, understanding, you know, how they build software in our case, and how we can make that, you know, more efficient, more productive. They can use tokens a little bit of, like, oh, whatever, generate them some, you know, AI debrief, take some notes, like, help them with a follow-up. But, like, that's so minimal, the number of tokens.
Speaker 140:50 - 41:25
这要看具体业务。我觉得每个业务都会有一个平衡点,这取决于——举个简单的例子,通常 salespeople(销售人员)如果本来就很优秀,他们大概不需要那么多 token。因为一般来说,他们提供最多 alpha(超额价值)的地方,是他们坐下来和客户面对面交流时,讨论客户的问题,理解——以我们这里为例——他们是怎么构建软件的,以及我们怎样能让这件事更高效、更有生产力。他们可以稍微用一点 token,比如随便生成个 AI debrief(AI 复盘摘要)、记点笔记、帮忙做后续跟进之类。但总的来说,token 的数量其实非常少。
Speaker 141:25 - 41:51
It basically doesn't matter. Like, if you add more tokens to the sales team, it probably won't change their output. If you add more humans to the sales team, it probably will. Meanwhile, engineering teams are pretty different, where engineering teams, generally, it seems like the you want people to own an outcome end to end, but then if you give them more tokens, they can produce a lot more. And so it seems like there and then there's a lot of kind of places in between of like operations, finance, marketing.
Speaker 141:25 - 41:51
基本上没那么重要。比如说,如果你给销售团队增加更多 token,他们的产出大概不会有什么变化;但如果你给销售团队增加更多人,产出大概率会变。与此同时,engineering teams(工程团队)就很不一样了:通常看起来,你会希望人对某个结果端到端负责,但如果你给他们更多 token,他们就能产出多得多。所以看起来这里面还有很多介于两者之间的岗位,比如 operations(运营)、finance(财务)、marketing(市场)。
Speaker 141:51 - 42:15
These are places where are neither here nor there, where I think they're they're somewhere in between and it kind of depends on your business. But I think every business is going to have to ask, like, what is our core competency? It's something that we see a lot in the market or we used to see, and now they finally kind of hit reality. But what we used to see is, oh, like, we're gonna build our own like software development agents. And we're like, okay, like you're a, like a consumer, like logistics company.
Speaker 141:51 - 42:15
这些职能既不完全属于这一边,也不完全属于那一边,我觉得它们处在中间地带,具体还是要看你的业务。但我认为每家公司都必须问自己:我们的 core competency(核心能力)是什么?这是我们在市场上经常看到的情况,或者说以前经常看到,而现在他们终于多少碰到现实了。我们以前常见的是,有人会说,哦,我们要自己构建软件开发 agent。然后我们会说,好吧,可你们是一家 consumer logistics company(消费物流公司)。
Speaker 142:16 - 42:24
Like, are you sure you wanna do that? They're like, yeah, yeah. We're, this is a core, we have to do this. It's like, okay. And then six months later, it's like, wait, actually, this is not a core competency for our business.
Speaker 142:16 - 42:24
你们确定真的想这么做吗?他们会说,对对对,这就是核心,我们必须做。这时我们会说,好吧。结果六个月后,他们又会说,等等,其实这并不是我们业务的核心能力。
Speaker 142:24 - 42:48
We don't wanna hire, you know, AI engineers to be doing this. Our core competency is, you know, consumer logistic. That's what we wanna focus on. And I think this is an opportunity for every business to double down on their core competency and what matters for them and then procure externally whatever it is that doesn't matter for them. Like a trivial example of this is like, I don't know, in the days of the early Internet, you probably had to be a programmer to build a website.
Speaker 142:24 - 42:48
我们不想为了这个去雇 AI engineers(AI 工程师)。我们的核心能力是 consumer logistics,这才是我们想专注的东西。我认为,这是一个让每家公司都能进一步押注自身核心能力、聚焦真正重要事项的机会,然后把那些对自己并不重要的东西向外部采购。一个很简单的例子是,我不知道,在早期 Internet 时代,你大概得会编程才能做一个网站。
Speaker 142:48 - 43:05
And, like, websites generally help if you're a pizza shop because you wanna have, you know, people come to your pizza shop. They wanna be able to like, whatever. At that time, would you say it was a core competency of, like, a pizza shop to have engineers? Like, certainly not. Like, that is kind of a byproduct of, like, a brief moment in time.
Speaker 142:48 - 43:05
而网站通常确实有帮助——如果你开的是一家 pizza shop,因为你想让人们来你的 pizza shop,他们需要能够,怎么说,完成相应操作。但在那个时候,你会说拥有工程师是 pizza shop 的核心能力吗?当然不是。这在某种程度上只是某个短暂历史时刻的副产品。
Speaker 143:05 - 43:29
But then there were companies out there that help you build a website. You don't need to be technical. And then this is why we live in a world where, like, most pizza shops don't have an engineering department, which I think is probably a good thing. And I think similarly, a lot of businesses have dealt with the reality of if you want to do x, y, z other thing, you have to bring in people of this type of role. But I think that's been, like, something you had to do, not because it's a core competency of the business.
Speaker 143:05 - 43:29
但后来出现了一些公司,能帮你搭建网站。你不需要懂技术。也正因如此,我们才生活在这样一个世界里:大多数 pizza 店都没有 engineering department,而我觉得这大概是件好事。类似地,我认为很多企业一直都在面对这样一个现实:如果你想做 x、y、z 这些别的事情,你就得引入某一类岗位的人。但我觉得,这更多是一种“你不得不做”的事,而不是因为它是企业的核心能力。
Speaker 143:30 - 43:43
And allowing businesses to focus and double down on the things that they're best at, I think is gonna be good for the consumers of their business. And so I think we're just gonna see, like, a lot, like, ruthless refocusing on what actually matters, which is gonna be cool to see.
Speaker 143:30 - 43:43
让企业把注意力集中起来,并加倍投入到自己最擅长的事情上,我认为这对它们的消费者会是件好事。所以我觉得,我们将会看到大量非常坚决的重新聚焦,回到那些真正重要的事情上,而这会很值得一看。
Speaker 243:43 - 44:05
On that, so, you know, every company kind of has to go through this process of reinvention. You know, ten or twenty years ago, people talked about digital transformation. And I don't know if anybody's given it a buzzword now, but AI transformation, something of that sort. Couple years ago, you ran into a bunch of organizations that just weren't ready to deal with autonomous agents. Things you've seen your customers start to change.
Speaker 243:43 - 44:05
说到这个,你知道,每家公司某种程度上都得经历这样一个自我重塑的过程。十年或二十年前,人们谈的是 digital transformation。现在我不知道有没有人给它起一个新的 buzzword,但大概可以叫 AI transformation,类似这种。几年前,你会碰到一大批根本还没准备好应对 autonomous agents 的组织。你已经看到客户开始发生的一些变化。
Speaker 244:05 - 44:19
And so the question is, when you look at your customers as they kind of go up this maturity curve and sort of reinvent themselves for the future, Any good, like, tricks or techniques that you've seen them use to repot themselves a bit?
Speaker 244:05 - 44:19
所以问题是,当你看着你的客户沿着这条 maturity curve 逐步向上、并为未来重新塑造自己时,你有没有看到他们会用哪些不错的小技巧或方法,来稍微给自己“换盆”一下?
Speaker 144:19 - 44:39
Yeah. I mean, I think, surprisingly, like, the companies that have been doing, like, company wide hackathons really end up doing well. It seems, like, relatively trivial, but, like, just setting aside a day for everyone in the workforce is just, like, build shit with AI. It really sets the tone and sets the pace.
Speaker 144:19 - 44:39
对。我是说,挺令人意外的是,那些一直在做 company wide hackathons 的公司,最后往往表现得很好。这听起来似乎相对琐碎,但只要专门留出一天时间,让整个 workforce 的所有人都去“用 AI build shit”,它真的会定下基调,也会把节奏带起来。
Speaker 244:39 - 44:40
So just give me a look. It's
Speaker 244:39 - 44:40
那就给我看一下。它——
Speaker 344:40 - 44:45
not I I tried to force him to build some of the coding agents. It didn't go so well.
Speaker 344:40 - 44:45
不是,我之前试着逼他去做一些 coding agents,结果不太顺利。
Speaker 144:45 - 44:45
We'll work on it.
Speaker 144:45 - 44:45
我们会继续努力。
Speaker 244:45 - 44:48
We'll do it after this one. You know, we gave it a great effort.
Speaker 244:45 - 44:48
我们会在这件事之后再做。你知道,我们已经为此付出了很大的努力。
Speaker 144:49 - 44:57
But that's it. Like, it literally just setting aside the time to, like, do it. And, like, even if it fails miserably, like, it's fine. And also, like, the orgs that are okay with failing. Yeah.
Speaker 144:49 - 44:57
但也就是这样了。说白了,真的就是要专门腾出时间去做。而且,就算最后惨败了,也没关系。还有就是,那些能够接受失败的 orgs(组织)也很关键。对。
Speaker 144:57 - 45:12
Like, it feels like it feels like there are some who are like, we need to do it exactly right. We need to make the right decision from day one. No air like, you're gonna make mistakes. Everyone is going to. And the orgs who are kind of leaning into it and embracing it to a certain degree, I think, are succeeding.
Speaker 144:57 - 45:12
感觉好像有些人会觉得,我们必须把这件事做得绝对正确。我们必须从第一天起就做出正确决定。不是这样的——你一定会犯错。每个人都会。至于那些某种程度上愿意投入其中、去接纳这件事的 orgs(组织),我觉得他们正在取得成功。
Speaker 145:12 - 45:27
Like, one of our largest customers is is not necessarily known to be, like, at the absolute frontier of AI, but I think for them, they were just like, look. This matters. We were kind of there have been other transit transformations that we relate to. We're not gonna be late to this. Like, we're just gonna go in.
Speaker 145:12 - 45:27
比如说,我们最大的客户之一,并不一定以站在 AI 最前沿而闻名,但我觉得对他们来说,他们就是在想,听着,这件事很重要。我们以前也经历过其他类似的转型。我们不会在这件事上落后。我们就是要直接进去做。
Speaker 145:27 - 45:51
We might mess up, but, like, obviously, respecting, like, secure the things that you're not allowed to mess up Sure. Put those aside. But, like, let's go and get our engineers to mess around and build this stuff and see where it breaks and understand what they like and what they don't like. I think that really matters a lot in the ones that we're seeing succeed. And also the ones who are like pretty bold in reinventing the processes that they've put in place.
Speaker 145:27 - 45:51
我们可能会搞砸,但当然,要注意把那些你绝对不能出错的东西保护好。没错,那些先放到一边。但除此之外,就去做吧,让你们的工程师去折腾、去构建这些东西,看看它会在哪里出问题,理解他们喜欢什么、不喜欢什么。我觉得,在我们看到那些成功的案例里,这一点真的非常重要。还有,那些敢于大胆重塑自己既有流程的人,也更容易成功。
Speaker 145:51 - 46:04
And just saying like, hey, it's there's no sacred cows. Like, let's let's put this aside, try something out. If it doesn't work, put that sacred cow right back. And I think that's that's been kind of a determining factor there. And when it comes from within.
Speaker 145:51 - 46:04
也就是要说,嘿,这里没有什么 sacred cows(不可触碰的旧规)。来吧,先把它放到一边,试点新东西。如果不行,再把那个 sacred cow 原样放回去。我觉得这一直是一个相当关键的决定性因素。尤其是当这种推动来自内部的时候。
Speaker 146:04 - 46:12
If it comes from the board, probably not gonna go well. Yeah. If it comes from within, like, the tech team or the the ICs or the leadership, that's when we see it go better.
Speaker 146:04 - 46:12
如果这事是从 board(董事会)那里压下来的,多半不会有什么好结果。对。如果它是从内部来的,比如 tech team(技术团队)、ICs(独立贡献者)或者领导层发起的,那就是我们看到它进展更顺利的时候。
Speaker 346:12 - 46:19
Do you have any predictions for the most important changes that are gonna happen in your space over the next, call it, twelve months?
Speaker 346:12 - 46:19
你有没有什么预测:在接下来的,姑且说,十二个月里,你所在这个领域会发生哪些最重要的变化?
Speaker 146:19 - 46:37
A lot of AI consumption is going up like crazy, and everyone's super, super excited because our revenue is going wild. Like, lot of this is synchronous usage. In other words, like, if everyone woke up sick tomorrow, like, a lot of Claude code usage would be zero. Because it's all just, hey, Claude code or, hey, Codex or, hey, Droid. Right?
Speaker 146:19 - 46:37
现在很多 AI 消费量都在疯狂上涨,大家都超级、超级兴奋,因为我们的营收也在狂飙。比如说,这里面很大一部分其实是同步使用(synchronous usage)。换句话说,如果明天所有人一觉醒来都病倒了,那很多 Claude code 的使用量就会变成零。因为这一切本质上都只是:嘿,Claude code,或者,嘿,Codex,或者,嘿,Droid。对吧?
Speaker 146:37 - 47:00
I think in twelve to twenty four months, like, 90% of tokens will be asynchronous tokens. So these are gonna be, you know, droids on their own autonomously being like, hey, here's some signal that I found from a customer. Let's go fix it or let's go create a first pass solution to this. And I think that is gonna be where the real, like, agent native stuff begins. Because right now we're still kind of in, like, copilot mode.
Speaker 146:37 - 47:00
我觉得在未来 12 到 24 个月里,90% 的 token 都会变成异步 token(asynchronous tokens)。也就是说,这些会是自主运行的 droids,自己在那里说:嘿,我从客户那里发现了某个信号。我们去把它修掉,或者先为这个问题做一个第一版方案。我认为真正 agent native 的东西会从这里开始。因为现在我们其实还多少处在 copilot 模式里。
Speaker 147:00 - 47:32
Like, if you're going to an agent and say, hey, go do this for me, it is more agentic because it's not gonna come back and ask you a ton, but it's still like you are kicking it off. Like, yeah, if you guys have ever been to Tesla's factories, which is one of the sources of inspiration for the name is, like, it's just robotic arms everywhere going and doing stuff. Like, it's not like there are people there, like, going and, you know, attaching the widget to the thing. And this idea of, like, a dark factory where, like, the lights are off and things are just happening, that is where software development's going. That's kind of where the the name came from is, like you know, Elon was always talking about the factory is the machine that builds the machine.
Speaker 147:00 - 47:32
比如,如果你去对一个 agent 说,嘿,替我做这个,它确实更有 agentic 的特征,因为它不会回来问你一大堆问题,但本质上还是你在启动它。像,如果你们有人去过 Tesla 的工厂——这也是这个名字灵感的来源之一——你会看到到处都是机械臂在运转、在做事。不是说还有人在那里手动把某个小部件装到另一个东西上。那种所谓 dark factory(黑灯工厂)的概念——灯关着,但事情自己在发生——这就是软件开发要去的方向。这个名字某种程度上也是这么来的:Elon 一直在说,工厂才是制造机器的机器。
Speaker 147:32 - 47:47
Yeah. And that's been something that we took to heart. And I guess also that combined with his whole thing about how you're destined to become the opposite of your name. And in our case, you know, factory becomes artisanal. It was kind of a good a good flip there.
Speaker 147:32 - 47:47
对。那也是我们非常认同的一点。我想,再加上他那套“你最终会变成自己名字的反面”的说法。放到我们这里,就是 factory 最后变成了 artisanal。这种反转还挺妙的。
Speaker 247:47 - 47:53
So What's your most optimistic version of the future, both for factory and for the world at large? So I
Speaker 247:47 - 47:53
那么,你对未来最乐观的版本是什么?无论是对 factory,还是对整个世界?所以我——
Speaker 147:53 - 48:24
think short term, there's gonna be a lot of turbulence because I think a lot of companies have misallocated resources pretty poorly. There's been a lot of bloat. And I think the correction that's gonna happen there is gonna be really painful for a lot of people. And I think that's something that I think every AI CEO should really bear much more responsibility than they currently are for. And also figuring out ways to address and kind of ameliorate in some way because this is something that's gonna be very painful for a lot of people.
Speaker 147:53 - 48:24
我觉得短期内会有很多动荡,因为我认为很多公司在资源配置上做得非常糟,资源错配很严重。也积累了很多臃肿问题。而我认为接下来发生的那种纠偏,对很多人来说会非常痛苦。我也认为,这是每一个 AI CEO 都应该比现在承担更多责任的事情。同时也应该去想办法应对它,并在某种程度上缓解它,因为这会让很多人非常难受。
Speaker 148:24 - 48:46
Now I I have optimism that we can actually address that faster than we think. We just need to start now in terms of addressing that. Now the longer term and why I think this is a good thing is why I don't believe at all, like, you know, the BS that people are saying of, oh, engineers are going away. Generally, there is a huge number of problems in the world. A large subset of those problems can be solved with software.
Speaker 148:24 - 48:46
不过我依然乐观地认为,我们实际上能比自己想象得更快地解决这个问题。只是我们现在就得开始着手处理。再说长期,以及为什么我认为这总体上是件好事:我完全不相信那些人在说的鬼话,比如“工程师要消失了”。总体而言,这个世界上有大量问题,而其中很大一部分都可以靠软件来解决。
Speaker 148:47 - 49:10
A small subset of those problems are currently being solved with software. And so in the short term, this means that, okay, first, there's a given problem that was over allocated engineering resources. So, okay, we need to reallocate those. Reallocate those is a very kind of cold way of saying some people are going to lose their jobs. But I think the, the thing that's going to happen in the longer term is we need engineers.
Speaker 148:47 - 49:10
但目前真正正在用软件解决的,只是这些问题中的一小部分。所以短期来看,这意味着:好,首先,有些既定问题曾经被过度配置了 engineering resources(工程资源)。那好,我们需要重新配置这些资源。所谓“重新配置”,其实是一种很冷冰冰的说法,说白了就是有些人会失去工作。但我认为,从更长期来看,真正会发生的事情是:我们仍然需要工程师。
Speaker 149:10 - 49:28
Engineers are some of the best systems thinkers and the best problem solvers. And there are so many problems that can be solved with software that are not being solved with software. And so that means that we are going to take those engineers and have them go and solve problems that previously were not being solved. That is such a net good for the world. Because, again, there are so many of these problems that we are not solving.
Speaker 149:10 - 49:28
工程师是最擅长系统性思考的人群之一,也是最优秀的问题解决者之一。而且,有太多本可以用 software(软件)解决的问题,还没有被 software 解决。所以这意味着,我们会让这些工程师去解决那些过去没有被解决的问题。这对世界来说绝对是巨大的净利好。因为再次强调,这类我们尚未解决的问题实在太多了。
Speaker 149:28 - 49:58
And, also, there's so many problems that we are maybe solving but with really shitty software. And, like, this is gonna enable people to solve it with incredible software. And, you know, the vision for factories that we are kind of the the factory that allows them to go and build this incredible software to solve these different problems. And these problems range from, like, things that are trivial to, you know, like, government software typically is not very good, whether it's, like, DMV or, like, IRS web like, all that stuff is generally a pretty poor experience. We don't need to live like that.
Speaker 149:28 - 49:58
而且,还有很多问题也许正在被解决,但用的是非常糟糕的 software。这个趋势将使人们能够用极其出色的 software 去解决这些问题。你知道,我们对 Factory 的愿景有点像是一座工厂,这座工厂让他们能够去构建这种了不起的 software,用来解决各种不同的问题。这些问题的范围很广,从一些看似琐碎的事,到比如政府 software——通常都不太好,不管是 DMV,还是 IRS 的网页之类,总体体验都相当差。我们没必要一直这样生活下去。
Speaker 149:58 - 50:26
Like, we can live in we can live in a world where all software is really fantastic, but also things like, you know, pharmaceutical research. Like, so much that goes into solving diseases is not just, like, a biology problem. A lot of it requires the best software engineers in the world. And previously, those problems haven't allocated the right dollars to attract the best engineers. But now because of what's happening, I think we will be much more closely allocated to, like, these are the biggest problems.
Speaker 149:58 - 50:26
我们完全可以生活在一个所有 software 都非常出色的世界里;同时,还有像 pharmaceutical research(制药研究)这样的领域。比如说,治疗疾病这件事里有太多工作,并不只是 biology(生物学)问题。其中很大一部分需要世界上最优秀的 software engineers(软件工程师)。而过去,这些问题并没有分配到足够的资金去吸引最优秀的工程师。但现在,随着正在发生的一切,我认为资源配置会更紧密地对准这些真正最大的问题。
Speaker 150:26 - 50:40
Let's get the best minds and the best problem solvers to solve that. I think it's kind of our job as an industry to do that relocation reallocation as quickly as possible. So it's not ten years, but maybe like six months or a year.
Speaker 150:26 - 50:40
让最优秀的头脑和最强的问题解决者去解决这些问题。我认为,作为一个行业,我们的职责某种程度上就是尽快完成这种 relocation、reallocation(重新迁移与重新配置)。这样就不是等上十年,而可能只需要六个月或一年。
Speaker 350:40 - 50:57
Wonderful. Matan, I think the clarity and consistency of your vision over time has just always been very inspiring. Then just seeing how much you've grown as a leader and how much Factory has grown as a company, even since the last time we did this Training Data episode, it's truly awe inspiring. So thank you for joining us again to share what you're up to.
Speaker 350:40 - 50:57
太棒了。Matan,我觉得,你的愿景一直以来都非常清晰且前后一致,这始终都很鼓舞人心。再看到你作为一名领导者成长了这么多,以及 Factory 作为一家公司成长了这么多——哪怕只是和我们上一次录制这期 Training Data 节目相比——也真的令人惊叹。所以,非常感谢你再次加入我们,分享你最近在做的事情。
Speaker 150:57 - 50:59
I appreciate it a lot. Thank you. Thank you.
Speaker 150:57 - 50:59
我非常感激。谢谢。谢谢。
原文 ↗https://www.youtube.com/watch?v=ZesOukBjPmI
BuildSpeak — 关于本项目BUILT IN PUBLIC · 跟随 builders 而非 influencers