BuildSpeak每日 builder 文摘
今日归档生词本关于
🎙 播客No Priors· 2026 年 6 月 26 日· 7,523 词 · 约 38 分钟

Why Traditional Benchmarks Fail Modern AI Models with OpenAI Research Scientist Noam Brown

SPACE 播放 / 暂停·←→ 上一句 / 下一句
Speaker 100:00 - 00:13
With GBT three, you couldn't scale test time compute. Like, if you gave it a budget of $10,000,000 and said, okay. Well, let's see what GBT three can do. It really can't do that much. The precarious frameworks and responsible scaling policies, they don't really account for the amount of test time compute.
Speaker 100:00 - 00:13
用 GBT three 时,你没法扩展 test time compute(测试时计算)。比如说,如果你给它一个 $10,000,000 的预算,然后说,好吧,我们来看看 GBT three 能做什么。它其实也做不了太多。那些脆弱的 framework(框架)和 responsible scaling(负责任扩展)政策,并没有真正把 test time compute 的规模考虑进去。
Speaker 100:13 - 00:32
They just say, okay. Well, what's the capability of the model? The problem is we're in a world now where the capability of the model is a function of how much money you put into it, basically. If you give it a budget of $10,000, it can do a lot more than what it can do with a budget of $10. Give it a budget of $10,000,000, it can do even At what budget should you evaluate these models?
Speaker 100:13 - 00:32
它们只是说,好吧,那这个 model(模型)的 capability(能力)是什么?问题在于,我们现在所处的世界里,model 的 capability 基本上取决于你往里面投了多少钱。如果你给它 $10,000 的预算,它能做的事会比只有 $10 预算时多得多。给它 $10,000,000 的预算,它还能做得更多。那么,你究竟应该在什么预算水平上评估这些 model 呢?
Speaker 100:32 - 00:36
The policies that exist today don't really address that question.
Speaker 100:32 - 00:36
现有的政策其实并没有真正回答这个问题。
Speaker 200:43 - 01:03
Hi, listeners. I'm Sarah Gore, and welcome back to No Priors. Today, I'm here with Noam Brown, one of our godfathers of AI reasoning. We talk about the broken state of evaluations, very large scale test time compute, how he thinks about recursive self improvement, and what's next on the horizon for competition at the frontier. Welcome.
Speaker 200:43 - 01:03
大家好,听众朋友们。我是 Sarah Gore,欢迎回到 No Priors。今天和我一起的是 Noam Brown,他是 AI reasoning(AI 推理)领域的奠基者之一。我们聊到了 evaluation(评估)体系如今的失灵状态、超大规模的 test time compute、他如何看待 recursive self improvement(递归式自我改进),以及 frontier(前沿)竞争的下一阶段会是什么。欢迎你。
Speaker 201:03 - 01:05
Noam, I'm so excited to have you back.
Speaker 201:03 - 01:05
Noam,我真的特别高兴你能再次来节目。
Speaker 101:05 - 01:06
That's great to be back. Yeah.
Speaker 101:05 - 01:06
很高兴再次回来。是的。
Speaker 201:06 - 01:21
You are our first guest. I'm very proud of my taste in friends and researchers for the pod, given, you know, how important, you know, inference time scaling has become to the industry. You should be proud too, having actually pioneered
Speaker 201:06 - 01:21
你是我们的第一位嘉宾。考虑到 inference time scaling(推理时扩展)如今对整个行业已经变得多么重要,我对自己为这个播客挑选朋友和研究者的眼光感到非常自豪。你也应该为此感到自豪,毕竟你确实开创了
Speaker 101:21 - 01:23
it. Played apart, yeah, among many others.
Speaker 101:21 - 01:23
它。算是参与其中了,对,和很多其他人一起。
Speaker 201:24 - 01:37
You just wrote this essay that really resonated about large scale test time compute and why the industry is not evaluating these models as robustly as it should be. What was the motivation for it?
Speaker 201:24 - 01:37
你刚写了这篇文章,关于大规模 test-time compute(测试时算力)以及为什么行业并没有像应有的那样稳健地评估这些模型,引发了很多共鸣。写这篇文章的动机是什么?
Speaker 101:37 - 02:08
Yeah. The motivation was we released 5.5, and the initial reaction was kinda skepticism that it was a substantially better model. This to be fair, that only lasted for a few hours before people had some time to play around with it and and try it out themselves, and they saw that it was actually substantially better. But I think a lot of the skepticism came from the benchmark grid that was published. Basically, whenever a new model is released, there's this benchmark grid where they show all these different benchmarks on on the x axis and then the performance of different models on the y axis, and you can just, like, compare different models.
Speaker 101:37 - 02:08
对。动机是我们发布了 5.5,而最初的反应有点怀疑,觉得它并没有实质性地变得更好。公平地说,这种怀疑只持续了几个小时,后来大家有时间自己上手玩一玩、试一试,就发现它实际上确实好了很多。但我觉得,很多怀疑都来自当时发布的 benchmark grid(基准测试表)。基本上,每当有新模型发布时,都会有这么一张 benchmark grid:x 轴列出各种不同的 benchmark(基准测试),y 轴展示不同模型的表现,这样你就可以直接比较不同模型。
Speaker 102:08 - 02:28
It's like a single number for a model on a single benchmark. And if you look on paper at the difference between, like, five point five and five point four or or other models, it wasn't it was an improvement, but it wasn't a huge improvement. It was only a few percentage points on some benchmarks. So people looked at that, and they were skeptical that it was actually a better model. Once they played around with it, the story changed.
Speaker 102:08 - 02:28
它本质上就是“一个模型在一个 benchmark 上的一个单一分数”。如果你从纸面上看,比如 five point five 和 five point four,或者和其他模型之间的差别,确实是有提升,但并不是那种特别巨大的提升。在一些 benchmark 上也就只高了几个百分点。所以人们看了之后,会怀疑它是否真的更好。但一旦他们亲自上手体验,结论就变了。
Speaker 102:28 - 02:53
I think the reason why it doesn't show up as so much better on the benchmarks is because the benchmarks are being presented the benchmark results are being presented in the wrong way. They're not controlling for the amount of test on compute that is being used on that benchmark question. It turned out that 5.5 is just much more efficient with its thinking. If you run it at max settings, 5.4 is thinking for a lot longer. It takes longer to get back a response than 5.5.
Speaker 102:28 - 02:53
我认为它在 benchmark 上没有显得好那么多,原因在于 benchmark results(基准测试结果)的呈现方式是错的。它们没有控制每一道 benchmark 题目所使用的 test-time compute(测试时算力)数量。结果发现,5.5 只是“思考”效率高得多。如果你把设置拉到最大,5.4 实际上会思考更久,返回响应的时间也比 5.5 更长。
Speaker 102:54 - 03:12
And once you control for the amount of thinking time, actually, you can see that 5.5 is a substantial jump over 5.4. That is, I think, people's day to day experience with it. And then when I mention this to people, the reaction the typical question I get is like, okay. Well, why not just have 5.5 think for as long as 5.4? And the question is like, well, how long should they think for?
Speaker 102:54 - 03:12
一旦你控制住思考时间的长短,其实就会发现,5.5 相比 5.4 是一次相当大的跃升。我觉得,这也符合大家日常使用它时的体验。然后当我把这点告诉别人时,典型的反应通常是:好吧,那为什么不让 5.5 也像 5.4 一样思考那么久?而问题在于:那它们到底应该思考多久?
Speaker 103:12 - 03:31
Typically, the response I get is, well, until the performance plateaus. Right? This is at some point where the performance on the benchmark is gonna plateau, and you just evaluate to that point. The thing is, the point at which it plateaus is actually really far out these days. I mean, if it's true in g p d three land back in 2022, models couldn't really think productively for that long, and so you could just run them until they plateau.
Speaker 103:12 - 03:31
我通常得到的回答是,那就一直到性能进入 plateau(平台期)为止,对吧?也就是说,benchmark 上的表现总会在某个点趋于平稳,那你就评估到那个点。问题是,如今这个进入 plateau 的点其实已经非常靠后了。我的意思是,如果放在 2022 年的 g p d three 那个时代,模型其实还不能持续进行那么长时间的有效思考,所以你确实可以直接让它跑到 plateau 为止。
Speaker 103:31 - 03:52
It's not that far away. But what we're seeing today with the modern models is that 5.5 and other models can think for if you scaffold them reasonably well, can think for weeks even, before having performance plateau on some of these benchmarks. And so the point at which they plateau is simply too far out to reasonably test.
Speaker 103:31 - 03:52
那个点并不算太远。但我们今天在现代模型上看到的是,5.5 和其他模型如果给它们设计合理的 scaffold(脚手架式提示/流程),甚至可以连续思考几周,在某些 benchmark 上性能才会进入 plateau。所以,它们达到 plateau 的那个点,已经远到无法被合理地测试了。
Speaker 203:52 - 04:00
We all need to actually reinforce either, like, a patience limit or a budget limit from a token perspective now, and that wasn't true a few years ago.
Speaker 203:52 - 04:00
现在,我们实际上都需要从 token 的角度引入某种 patience limit(耐心上限)或者 budget limit(预算上限),而这在几年前还不是个问题。
Speaker 104:00 - 04:19
Exactly. And so I think the proper way to and so my claim is the proper way to evaluate the models now is you either have some kind of budget for the benchmark, whether it's tokens or cost or time or whatever, or you plot the performance as a function of the amount of test time compute that's going into the model. And then it becomes much more clear how to compare the performance between these different models.
Speaker 104:00 - 04:19
没错。所以我认为——我的主张是——现在评估这些模型的正确方式,是你要么给 benchmark 设定某种预算,无论是 tokens、成本、时间还是什么;要么把性能画成一个函数,横轴是投入到模型中的 test-time compute(测试时计算)量。这样一来,比较这些不同模型之间的性能就会清楚得多。
Speaker 204:20 - 04:40
Given the model evaluation cycle and the fact that performance does not asymptote for many tasks over quite a long period of time, what do you do about that issue, the fact that some of the evals that you would want to run are both beyond the scope of budget or time that's reasonable given the current model release cycle?
Speaker 204:20 - 04:40
考虑到模型评估周期,以及这样一个事实:对许多任务来说,性能在相当长一段时间内都不会渐近到上限,那么你怎么处理这个问题?也就是说,你想运行的一些 eval(评测)已经超出了在当前模型发布周期下合理的预算或时间范围。
Speaker 104:40 - 05:04
I mean, I think for things like cyber, we've seen, and actually the AISI in their evaluations has shown that the models continue to improve at a 100,000,000 tokens. You know, if you run them for a 100,000,000 tokens, they're still improving at beyond that point. Mhmm. And that can take a very long time to run. But you also do see that, like, the performance is is it's not just, like, a discontinuous jump.
Speaker 104:40 - 05:04
我的意思是,我觉得像 cyber 这类事情上,我们已经看到——实际上 AISI 在他们的评估中也表明——这些模型在 100,000,000 tokens 的规模上仍在持续提升。也就是说,如果你让它们跑到 100,000,000 tokens,甚至超过那个点之后,它们仍然在进步。嗯。而且这可能需要很长时间才能跑完。但你也会看到,性能提升并不只是那种突发性的、不连续的跳跃。
Speaker 105:04 - 05:25
It's actually, like, you can see the slope of improvement over those 100,000,000 tokens. And so you you could probably do some kind of evaluation up to a certain budget and then just say, okay. Well, this is what we project the performance to look like. And I think this this there hasn't been a lot of research on this yet. I actually think this would be a great paper to publish if there's any academics out there looking for something to research.
Speaker 105:04 - 05:25
实际上,你可以看到在这 100,000,000 tokens 过程中的提升斜率。所以你大概可以先在某个预算上限内做评估,然后再说,好吧,这就是我们预测的性能表现会是什么样。我认为目前这方面还没有很多研究。实际上,我觉得如果有学者正在寻找研究题目,这会是一篇很值得发表的论文。
Speaker 105:25 - 05:34
Can you predict what the performance looks like at an inference budget of, let's say, $10,000, only using inference budgets up to 10 or 100 or 10 or $100?
Speaker 105:25 - 05:34
你能否预测,在 inference budget(推理预算)比如说 $10,000 的情况下,性能会是什么样子,而你手头只使用了不超过 10、100,或者 10 到 $100 的 inference budget 数据?
Speaker 205:34 - 05:42
So maybe an orthogonal question for you. Do you think users are systematically, like, not thinking long enough with their models about problems?
Speaker 205:34 - 05:42
那我换个相对正交的问题问你。你是否认为,用户在使用模型解决问题时,系统性地没有让模型“思考”足够久?
Speaker 105:42 - 05:44
What do you mean by not thinking long enough?
Speaker 105:42 - 05:44
你说的“没有思考足够久”是什么意思?
Speaker 205:44 - 06:08
If you can build an agent or control the amount of test time compute being used, like, there's what is done by the model itself, and there's what the user can do. Do you think that, you know, the industry is using test time compute an optimal amount, way undershooting it, or it's, you know, it's a problem in the models where they just need to be able to do that thinking faster?
Speaker 205:44 - 06:08
如果你可以构建一个 agent,或者控制所使用的 test-time compute 量,那么这里既有模型自身在做的部分,也有用户可以做的部分。你觉得,行业现在对 test-time compute 的使用量是最优的、严重偏少的,还是说问题出在模型本身——它们只是需要能够更快地完成这种思考?
Speaker 106:08 - 06:32
I I think it depends on the problem. I think this idea that the models, you just let them think for a week or whatever, then they respond. It's it sounds nice, and, yes, the benchmarks look great, but it's not very practical when working because, like, okay, you ask the model a question, then you sit there for a week waiting for it to come back to you. Mhmm. I think what people have have found most effective is to kind of, like, iterate quickly with the models.
Speaker 106:08 - 06:32
我觉得这取决于具体问题。我觉得那种想法——模型你就让它思考一周之类的,然后它再回复——听起来不错,而且,是的,benchmark(基准测试)上的结果看起来很亮眼,但在实际工作中并不太实用,因为比如说,你问模型一个问题,然后你就得坐在那里等一周,等它回来答复你。嗯。我觉得人们发现最有效的方式,还是和模型进行快速迭代。
Speaker 106:32 - 06:46
And so the thinking time, I think, needs to be flexible. When it makes sense to respond quickly to the user, it should respond quickly. And then when it makes sense to think for a long time and the user wants it to think for a long time, then it makes sense to think for a long time. I think people have been striking the right balance given what they have to deal with right now.
Speaker 106:32 - 06:46
所以我觉得思考时间需要是灵活的。什么时候适合快速回复用户,它就应该快速回复;而当适合长时间思考,并且用户也希望它长时间思考时,那长时间思考就是合理的。我觉得考虑到当下人们要处理的现实约束,他们基本上已经找到了合适的平衡。
Speaker 206:47 - 07:02
How would you characterize know, there's a lot of talk about benchmark maxing and the ability to gain different benchmarks. What would you characterize the, like, landscape of benchmarks as today? And then do you have, like, favorites that you think are more indicative of capability than others?
Speaker 206:47 - 07:02
你会怎么描述现在的 benchmark(基准测试)格局?现在大家谈很多 benchmark maxing(把基准分数刷到极致)以及在不同 benchmark 上提分的能力。那你会如何概括当下的 benchmark 生态?另外,你有没有一些自己偏爱的 benchmark,觉得它们比其他 benchmark 更能反映真实能力?
Speaker 107:03 - 07:41
So the benchmark maxing thing is also motivation for for writing an essay that I think it's really easy to show you can do much better than previous benchmarks or or previous previous models on benchmarks by just, for example, scaffolding a bunch of models together. So if you say, okay. Well, we're going to instead of just running this model once, we're gonna run it five times and take the best of the five responses or, like, ask a judge which one it it thinks is best, then you can get much higher scores than that model. And so it's really easy to make something that looks a lot better on paper but is actually not better once you control for the amount of test time compute. That is one thing that I'm worried about when it comes to benchmark maxing.
Speaker 107:03 - 07:41
所谓 benchmark maxing 这件事,也是促使我去写那篇文章的原因之一。我认为,要展示“你在 benchmark 上比之前的 benchmark 结果,或者比之前的模型做得好很多”,其实是很容易的。举个例子,你只要把一堆模型用 scaffold(脚手架式组合)方式串在一起就行。比如你说,好,我们不只是把这个模型跑一次,而是跑五次,然后从五个回答里选最好的一个,或者让一个 judge(裁判模型)来判断哪个最好,那你就能拿到比这个模型本身高得多的分数。所以,做出一个在纸面上看起来好很多、但实际上在控制了 test-time compute(测试时算力)之后并没有更好的东西,是非常容易的。这是我对 benchmark maxing 的一个担忧。
Speaker 107:41 - 08:10
I mean, it, like it's it's a little misleading is the only concern that I have. And then as far as, like, the benchmarks themselves, I think there is always a risk of, like, just optimizing for the benchmark. And I've I've certainly encouraged my team, and I think at OpenAI, we're pretty good about not trying to optimize for for specific benchmarks. But once you put out a benchmark, it's there's it's always at risk of just being optimized for. And I think one way to one way to address that is to keep held out private sets that isn't publicly available.
Speaker 107:41 - 08:10
我的意思是,我唯一的担心在于,这多少有点误导。至于 benchmark 本身,我觉得始终存在一种风险:大家只是去针对 benchmark 做优化。我当然一直在鼓励我的团队这样做,而且我觉得在 OpenAI,我们在“不去专门针对某些特定 benchmark 优化”这件事上做得还不错。但一旦你把一个 benchmark 发布出去,它就总是有被专门优化的风险。我觉得一个应对办法,是保留一些 held-out private sets(保留的私有测试集),不对外公开。
Speaker 208:10 - 08:30
The most popular fallback advice for, you know, figure out if a model is significantly better or not is to just play with it for a while. Do you have anything more sophisticated than that that you suggest people do? Like, do you create your own set of new evaluations each time besides private holdback at OpenAI?
Speaker 208:10 - 08:30
现在最常见的一条保底建议是:如果你想判断一个模型是不是显著更好了,那就直接上手玩一阵子。除了这个之外,你有没有什么更复杂、更系统的方法会建议大家去做?比如,在 OpenAI,除了 private holdback(私有保留集)之外,你们会不会每次都自己再设计一套新的 evaluation(评估)?
Speaker 108:30 - 08:34
I think everybody has their own set of questions that they like to ask the model whenever it comes out.
Speaker 108:30 - 08:34
我觉得每个人都会有自己的一套问题,每次有新模型出来时,就会拿这些问题去问它。
Speaker 208:34 - 08:34
Mhmm.
Speaker 208:34 - 08:34
嗯。
Speaker 108:34 - 09:04
For me lately, it's been I I use them to make poker bots and see how good they can make a poker bot. I think it's a nice eval because there is very little open source code for making poker bots. And there's a lot of published essay there's a lot of published papers on it, but you really have to reason through everything. And it's, like it requires a lot of just reasoning and iteration and, like, a lot of small gotchas that I can kind of I've already worked through myself, so I can see where the models fail along the way. They've gotten really good at it now.
Speaker 108:34 - 09:04
对我来说,最近我会用它们来做 poker bot(扑克机器人),看看它们能把一个扑克机器人做到多好。我觉得这是个很好的 eval(评估)方式,因为关于如何做扑克机器人,现成的 open source code(开源代码)非常少。相关的论文倒是发表了很多,但你真的得把所有东西自己推理清楚。而且这里面很需要推理和迭代,也有很多细小的坑;这些坑我自己基本都已经踩过了,所以我能看出模型在过程中会在哪里失败。不过它们现在在这方面已经变得非常厉害了。
Speaker 209:04 - 09:15
Can you describe perhaps, like, with your poker bot creation, like, how reasoning might have progressed in model releases for you guys over a few releases?
Speaker 209:04 - 09:15
你能不能大致讲讲,比如结合你们做 poker bot 的经历,在连续几次 model 发布中,你们看到的 reasoning 能力是怎么逐步进展的?
Speaker 109:15 - 09:30
Yeah. When the early models were really bad at it. Like, they could not basically do anything. And then 5.2, I was able to work with it to make a reverse solver. So that's, like, the final stage Mhmm.
Speaker 109:15 - 09:30
可以。早期的模型在这方面真的很差,基本上什么都做不了。到了 5.2,我已经能和它配合做出一个 reverse solver 了。所以那算是最后一个阶段。嗯。
Speaker 109:30 - 09:48
Of poker. And that itself was, I thought, really impressive. I I had to work with it a little bit, but I was actually really impressed because I was able to make the reverse solver probably about five times faster than I would have alone. There were a couple things that I got tripped up on. Blockers was always a big, big issue.
Speaker 109:30 - 09:48
是在 poker 里。而这件事本身,我觉得已经非常令人印象深刻了。我确实还得和它磨合一下,但我真的很惊讶,因为我做出这个 reverse solver 的速度,大概比我自己单独来快了五倍。中间也有几件事把我卡住了。blockers 一直都是个很大很大的问题。
Speaker 109:48 - 10:08
But overall, like, you know, with a with a bit of gentle steering, it just kind of like I I kinda felt like a grad student where, okay, they would run into issues, but at least, like, I would know what those issues were and know how to fix it. And I could just make suggestions, and it would go off and then do it. And then pretty quickly, would actually come back with something really good. Mhmm. And then especially the optimization, thought was very impressive.
Speaker 109:48 - 10:08
但总体来说,你知道的,只要稍微温和地引导一下,它就会自己往前推进。我当时的感觉有点像是在带一个 grad student:好,虽然它会遇到问题,但至少我能知道那些问题是什么,也知道该怎么修。我只要给出一些建议,它就会自己去做。然后很快,它实际上就会带回来相当不错的结果。嗯。尤其是在 optimization 上,我觉得非常厉害。
Speaker 110:08 - 10:40
It was able to make it, like, 10 times faster than what I was able to do because it was just able to optimize the code so well. The downsides with 5.2 is I felt like it was gaslighting me a lot, and I always had to be very careful checking it and making sure, like, okay. Is it actually doing what it said it did? Are there any things that are, like, glaring issues that it's not recognizing or it's just pretending aren't issues? I remember there was, like, one point where for one of the models I was playing around with it, not 5.2, I kind of, like, as a as a unit test, I told it, okay.
Speaker 110:08 - 10:40
它能把速度做到大概比我自己做的快 10 倍,因为它确实非常擅长优化代码。5.2 的缺点是,我感觉它经常会 gaslight 我,所以我总得非常小心地检查它,确认,好,它到底有没有真的做到它声称自己做了的事?有没有什么很明显的问题,它自己没意识到,或者干脆假装那不是问题?我记得有一次,我在玩其中一个模型,不是 5.2,我有点像是在做一个 unit test,我跟它说,好。
Speaker 110:40 - 10:51
Well, let's say I have a $100 in the pot, and I fold. How much am I losing? And the model said $92. And I was like, that's crazy. I I have a $100 in the pot, I just fold it.
Speaker 110:40 - 10:51
那假设底池里有 100 美元,而我弃牌了。我会亏多少?那个模型回答说是 92 美元。我当时就想,这也太离谱了。我有 100 美元在底池里,然后我直接弃牌。
Speaker 110:51 - 10:55
How do I not lose a $100? And it said, oh, you know, it's 92. It's close to a 100. It's fine. It's no big deal.
Speaker 110:51 - 10:55
我怎么可能不是亏掉 100 美元?然后它说,哦,你知道,是 92。已经很接近 100 了。没关系,不是什么大问题。
Speaker 110:55 - 11:06
And I was like, clearly, this is a problem. Right? So the models did have this problem where they would gaslight you a lot. Mhmm. But once we got to 5.5, I actually thought it was way better.
Speaker 110:55 - 11:06
我当时就想,这显然是个问题。对吧?所以这些模型确实有这种问题,它们经常会 gaslight 你。嗯。但到了 5.5,我确实觉得它好多了。
Speaker 111:06 - 11:25
It was able to basically do a zero shot. And in fact, I've been working on just doing a full scale poker solver, and it it's basically able to do the whole thing with some gentle steering from me. And I wouldn't be surprised if, six months or a year from now, the model is able to do zero shot an entire poker solver, basically my entire PhD thesis, in one go.
Speaker 111:06 - 11:25
它基本上已经能做到 zero-shot(零样本)了。事实上,我最近一直在做一个完整规模的 poker solver,而它基本已经能在我稍微引导一下的情况下把整件事做出来。六个月或一年之后,如果这个 model 能够 zero-shot(零样本)一次性完成整个 poker solver,基本相当于把我的整篇 PhD thesis 一口气做完,我一点也不会意外。
Speaker 211:26 - 11:49
Let's talk about the larger implications of needing to evaluate these models relative to, let's say, speed of their reasoning or efficiency versus token volume, right, or dollar budget or whatever your scaler is. Can you describe some of the larger implications in your essay, including around safety evaluations?
Speaker 211:26 - 11:49
我们来谈谈一个更大的影响:在评估这些 model 时,需要结合它们的推理速度,或者效率与 token 数量、美元预算,或者不管你用什么 scaler(衡量尺度)之间的关系来看,对吧?你能不能讲讲你文章里提到的一些更广泛的影响,包括 safety evaluations(安全评估)方面的内容?
Speaker 111:49 - 12:08
Yeah. The safety evaluations thing, it's it's a bit of an inconvenient truth thing where okay. So I guess for background, a lot of the all of the labs have these things called either responsible scaling policies, preparedness frameworks. They go by various names. But the idea is that whenever a model's released, they go through a series of evaluations to measure, are there dangerous capabilities?
Speaker 111:49 - 12:08
对。关于 safety evaluations(安全评估)这件事,它有点像一个 inconvenient truth(令人不舒服的事实)。我先交代一下背景:现在很多,或者说所有 lab 都有一些东西,叫 responsible scaling policies,或者 preparedness frameworks,它们名字各不相同。但核心思路是,每当一个 model 要发布时,都会经过一系列评估,来衡量它是否具备危险能力。
Speaker 112:09 - 12:23
Could these models do things that we're we we wouldn't want, a bad actor to do? And if the model isn't very capable, then it's no big deal. But if it is very capable, if it could be used, for example, to make bioweapons, then you want to put in mitigations against that.
Speaker 112:09 - 12:23
这些 model 会不会做到一些我们不希望 bad actor(恶意行为者)去做的事情?如果 model 能力不强,那问题不大。但如果它非常强,比如说它可能被用来制造 bioweapons(生物武器),那你就会希望针对这种风险加入一些 mitigations(缓解措施)。
Speaker 212:23 - 12:24
Mhmm.
Speaker 212:23 - 12:24
嗯。
Speaker 112:24 - 12:41
But the question is, okay. Well, how do you evaluate whether the model is capable of that? And they have, like, various protocols about, like, how they do these evaluations. But a lot of these frameworks were developed around the era of ChatGPT, either before or after, when test time compute scaling was not really as much of a thing. Mhmm.
Speaker 112:24 - 12:41
但问题在于,好,那你要怎么评估一个 model 是否具备这种能力?他们确实有各种 protocol(流程规范)来说明如何做这些评估。但很多这类 framework 都是在 ChatGPT 那个时代前后形成的,当时 test-time compute scaling(测试时算力扩展)还并不是一个特别重要的东西。嗯。
Speaker 112:41 - 12:59
And it made sense. Like, with GPT three, you couldn't scale test time compute. Like, if you gave it a budget of $10,000,000 and said, okay. Well, let's see what g p t three can do, it really can't do that much more than what you could do with, like, $10 or $1. The preparedness frameworks and responsible scaling policies, they don't really account for the amount of test time compute.
Speaker 112:41 - 12:59
而这在当时是说得通的。比如 GPT-3,你没法对它做 test-time compute scaling(测试时算力扩展)。如果你给它 $10,000,000 的预算,然后说,好,我们来看看 GPT-3 能做到什么,它其实并不会比你只花 $10 或 $1 时强出太多。现在这些 preparedness frameworks 和 responsible scaling policies,其实都没有真正把 test-time compute 的规模考虑进去。
Speaker 112:59 - 13:08
They just say, okay. Well, what's the capability of the model? The problem is we're in a world now where the capability of the model is a function of how much money you put into it, basically.
Speaker 112:59 - 13:08
它们只是说,好,这个 model 的能力是什么。问题在于,我们现在所处的世界里,model 的能力基本上是你往里面投入多少钱的函数。
Speaker 213:08 - 13:08
Mhmm.
Speaker 213:08 - 13:08
嗯。
Speaker 113:08 - 13:25
If you give it a budget of $10,000, it can do a lot more than what it can do with a budget of $10. If you give it a budget of $10,000,000, it could do even more. And so at what budget should you evaluate these models? The policies that exist today don't really address that question.
Speaker 113:08 - 13:25
如果你给它 $10,000 的预算,它能做的事会比只有 $10 预算时多得多。如果你给它 $10,000,000 的预算,它还能做得更多。所以问题是,你到底应该在什么预算水平下评估这些 model?今天现有的这些 policy,其实并没有真正回答这个问题。
Speaker 213:25 - 13:25
Mhmm.
Speaker 213:25 - 13:25
嗯。
Speaker 113:25 - 13:50
Some do some do better than others, but for the most part, this is not really a factor that's being heavily considered. Now whether it should be released anyway, I don't I don't wanna wade into this question. I think there's, you know, there's arguments on both sides. But I think the important thing to recognize is that this is a question that is not being we're we're just kind of, like, you know, pretending that this issue doesn't exist, and I think it's important to just, you know, one way or the other account for it.
Speaker 113:25 - 13:50
有些做得比另一些更好,但总体而言,这其实并不是一个被认真纳入考量的因素。至于是否无论如何都应该发布,我不想卷入这个问题。我认为,怎么说呢,双方都有各自的论点。但我觉得重要的是要认识到,这是一个并没有被真正讨论的问题——我们某种程度上只是,像是在假装这个问题不存在。我认为重要的是,无论怎样,至少要把它算进去、正面处理。
Speaker 213:50 - 14:16
Yeah. It was the mirror image of the capability question of if the models can continue to do more and more without asymptoting on some tasks at very large budgets, then they should also be able to do so for tasks we don't want them to do as a society. Right? And so testing for that and what budget is allocated, it also seems out of sync from the model release cycle itself. Right?
Speaker 213:50 - 14:16
对,这其实是能力问题的镜像:如果这些 model(模型)在非常大的 budget(预算)下,能在某些任务上持续做得越来越多,而不是趋于平缓,那么它们也应该同样能够在那些作为社会我们并不希望它们去做的任务上表现出这种能力。对吧?所以,对这类能力进行测试,以及为此分配多少 budget,这看起来也和 model 的发布周期本身并不同步。对吧?
Speaker 214:16 - 14:43
There's been this acceleration of, you know, you get a new model every sometimes few days and weeks at this point versus six months. And you have a line in the essay where you say, like, the the only way to truly evaluate an agent on some very long running task might be to run it for a year, and that's gonna be true of both, like, useful and negative tasks. Right? And so how do you think about that versus the model release cycle?
Speaker 214:16 - 14:43
现在已经出现了这种加速:你会得到一个新 model,有时按如今的节奏是每隔几天、几周,而不是过去的六个月。你在那篇文章里有一句话,大意是,要真正评估一个 agent(智能体)在某种超长周期任务上的表现,唯一的方法可能就是让它跑上一年,而且无论是有用的任务还是负面的任务,这一点都成立。对吧?那么你怎么看待这一点与 model 发布周期之间的关系?
Speaker 114:44 - 14:53
Yeah. This this is also an interesting dynamic where, basically, as the models have become stronger, they've they're more they're better able to operate over longer horizons.
Speaker 114:44 - 14:53
对。这也是一种很有意思的动态变化:基本上,随着这些 model 变得更强,它们也越来越能够在更长的时间跨度上运作。
Speaker 214:53 - 14:54
Mhmm.
Speaker 214:53 - 14:54
嗯。
Speaker 114:54 - 15:12
So, again, with g p d three, if you wanted to run it for, you know, a week, there's really not much you could do to scaffold it into something useful that could actually run for a week. But we're seeing now with the most recent models that you can actually scaffold, for example, 5.5 into doing a series of experiments that can run for weeks, for months.
Speaker 114:54 - 15:12
所以,还是拿 g p d three 来说,如果你想让它运行一周,实际上你几乎没法通过搭 scaffold(脚手架式封装/工作流)把它变成一个真正有用、并且能连续跑上一周的东西。但我们现在看到,借助最新的 model,你实际上已经可以搭出这样的 scaffold,比如让 5.5 去执行一系列实验,而且这些实验可以持续跑上几周、几个月。
Speaker 215:12 - 15:19
Have you given your poker solver task, like, infinite budget yet?
Speaker 215:12 - 15:19
你有没有给你的 poker solver 任务分配过那种近乎无限的 budget?
Speaker 115:19 - 15:27
I haven't really scaffolded something together where I just tell it, like, okay, just run this for for weeks. I think I could it I I Until it has
Speaker 115:19 - 15:27
我还没有真正搭出一个系统,然后直接告诉它,像是,好,就让这个连续跑上几周。我觉得我可以,我……直到它具备
Speaker 215:27 - 15:28
some types.
Speaker 215:27 - 15:28
某些类型。
Speaker 115:28 - 16:04
Could probably give it slash goal and just, like, yeah, tell it to go nuts. But I I think at this point, it could 100% do the reverse solver if I just give it slash goal. I don't think it's at the level yet where it could do, like, the full poker solver if I gave it just slash goal and told it, yeah, go go run for a month. But we're going to pretty soon be at that point where I I probably could just tell it, like, yeah, go work on this for a month, and then come back to me with a a full, complete PokerSolver that's state of the art. And the problem is if you want to evaluate the capabilities of a model, what it can do after running for a month, the only way to be fully sure is to actually run it for a month.
Speaker 115:28 - 16:04
可能只要给它 /goal,然后,差不多就是,嗯,告诉它放手去做就行了。但我觉得以现在这个阶段,如果我只是给它 /goal,它已经 100% 能做 reverse solver 了。我觉得它还没到那种程度:如果我只给它 /goal,然后告诉它,嗯,去跑一个月,它就能做出完整的 poker solver。不过很快我们应该就会到那个点:我大概可以直接告诉它,比如,嗯,去把这个做一个月,然后回来给我一个完整、全面、state of the art(最先进)的 PokerSolver。问题在于,如果你想评估一个模型的能力,想知道它在连续运行一个月之后能做到什么,那么唯一能完全确定的方法,就是真的让它跑一个月。
Speaker 116:04 - 16:33
And if you wanna know after six months, the only way to know fully is to run it for six months. Now there I'll I'll get to, like, things we could do to address that a little bit later, but, like, it's important to recognize that the model release cycle is look. We're releasing new models, like, every two or three months at this point. And so a model comes out, it takes two or three months to push it to its limits, and then you have another model come out. And so nobody actually knows what the ceiling of capabilities are for these models because nobody's actually run them for long enough to really tell.
Speaker 116:04 - 16:33
如果你想知道六个月之后会怎样,唯一能完全知道的方法,就是让它跑六个月。这个问题我之后会讲到一些可以稍微缓解的办法,但重要的是要认识到,目前的模型发布周期是这样的:我们现在基本上每两三个月就会发布一个新模型。所以一个模型出来后,要花两三个月才能把它逼到能力极限,然后下一个模型又出来了。因此,实际上没有人真正知道这些模型的能力上限在哪里,因为根本没有人让它们持续运行足够久,来真正看清这一点。
Speaker 116:33 - 17:02
When slash goal came out, for example, I mean, people started running things that it took over a week for it to finish, and so people actually didn't realize that this was a big deal until after a week, until a week after it was released. Mhmm. I think that's gonna be more and more true. You know, the implications of that are, I think, pretty interesting because what do the labs do to, like, fully evaluate their models before they're released? It's actually very difficult because, yeah, you would have to the only way to to really do the evaluations is then delay the model release cycle.
Speaker 116:33 - 17:02
比如说,当 /goal 刚出来的时候,人们开始让它跑一些任务,而这些任务要一个多星期它才会完成,所以大家其实直到一周后、也就是它发布一周之后,才意识到这是一件大事。嗯。我觉得这种情况会越来越普遍。我认为这里面的含义很有意思,因为这就引出了一个问题:这些 labs(实验室)在发布模型之前,到底要怎么才能充分评估它们的模型?这其实非常困难,因为,没错,如果你真的要做这种评估,那么唯一真正可行的办法就是推迟模型发布周期。
Speaker 117:03 - 17:06
And, you know, there's a lot of competitive pressure right now to not do that.
Speaker 117:03 - 17:06
而你也知道,现在有很大的竞争压力,促使大家不要这么做。
Speaker 217:06 - 17:14
Do you think there's, like, exciting latent capability in the models that are already released that people have not fully explored given timeline?
Speaker 217:06 - 17:14
你觉得,在已经发布的这些模型里,是否存在一些令人兴奋的 latent capability(潜在能力),只是因为时间线的限制,人们还没有把它们充分探索出来?
Speaker 117:14 - 17:44
I think absolutely. I think actually a really great example is the Erdos unit distance problem. So for the viewers that don't know, like, we used an internal model at OpenAI a few weeks ago to disprove the unit Erdos unit distance conjecture. Now I'm not a mathematician, but this seems like it was a a pretty big deal. The in the math community, it was, like, the first first problem that a lot of mathematicians had really spent a lot of time on, and the model was able to do something that they weren't able to do and do it in a way that was actually interesting and useful for mathematicians.
Speaker 117:14 - 17:44
我觉得绝对有。其实一个非常好的例子就是 Erdos unit distance problem。给不了解的观众解释一下:几周前,我们在 OpenAI 用了一个内部模型,推翻了 unit Erdos unit distance conjecture。虽然我不是数学家,但这看起来似乎是一件相当大的事。在数学界,这好像是第一个有很多数学家真正花了大量时间研究的问题,而这个模型做到了他们没能做到的事,而且它完成这件事的方式,对数学家来说也确实是有趣且有用的。
Speaker 117:44 - 17:59
Honestly, it did it at a budget that was dirt cheap. I mean, we didn't put a lot of effort into this. We just we trained a new model, and we were just curious what it could do, and we ran it through some problems. And this one, at a pretty low budget, it was like, oh, yeah. I think I have a disproof, and then we were able to verify that, yeah, that disproof is correct.
Speaker 117:44 - 17:59
说实话,它做到这件事的 budget(预算)低得离谱。我的意思是,我们并没有为此投入很多精力。我们只是训练了一个新模型,然后单纯想看看它能做什么,于是就让它去跑一些问题。这个问题在相当低的预算下,它就像是,哦,对,我觉得我有一个 disproof(反证)了,然后我们也验证了,没错,这个反证是正确的。
Speaker 118:00 - 18:16
After we announced the results, a bunch of people found that you could get the answer out of 5.5 as well. If now it's not as simple as just asking 5.5, hey. Here's the neural unit distance conjecture. What's the disprove? You had to scaffold it a bit.
Speaker 118:00 - 18:16
在我们公布结果之后,很多人发现,其实用 5.5 也能得到这个答案。当然,这并不是说你只要直接问 5.5:“嘿,这里有个 neural unit distance conjecture。它的 disprove 是什么?”就行了。你还是得稍微给它搭个 scaffold(脚手架式流程)。
Speaker 118:16 - 18:29
You had to, like, steer it a bit. And so somebody found, okay. You asked 5.5. List a bunch of ways that you could tackle this problem. And then for it it lists one of the paths that are actually promising to get to the disprove.
Speaker 118:16 - 18:29
你确实得稍微“引导”它一下。于是有人发现,哦,明白了。你去问 5.5:列出一堆你可以如何处理这个问题的方法。然后它在这些方法里,会列出其中一条实际上很有希望走到 disproving(证伪)的路径。
Speaker 118:29 - 18:54
And then it you tell it, like, okay. Explore this some more. And then if you do this enough times, it actually ends up arriving at the disprove. Now what this means is you could, in principle, ask 5.5 to you know, as as a general purpose scaffold, list a bunch of different strategies, and then for each strategy, tell it to investigate that strategy. And then it would probably be able to arrive at the disprove with a general purpose scaffold.
Speaker 118:29 - 18:54
然后你再告诉它,比如说,好,继续更深入地探索这个方向。要是你这样做足够多次,它实际上最后确实会走到 disprove(证伪)。这意味着,原则上你可以让 5.5 作为一个通用 scaffold(脚手架),先列出一堆不同策略,然后针对每个策略,再让它去调查那个策略。这样的话,它大概率就能靠一个通用 scaffold 走到 disprove(证伪)。
Speaker 118:54 - 19:18
Now that scaffold would be very expensive. I mean, it would probably cost I I just ballpark, like, a thousand to 10 to a $100,000. But it would be possible, and it would have been possible for somebody to disprove the Erdosia to disc distance conjecture before we did using a general purpose model. And nobody had explored sufficiently, well, happens if I put a $100,000 worth of compute into 5.5? What could it do?
Speaker 118:54 - 19:18
当然,那个 scaffold 会非常昂贵。我是说,我粗略估计一下,成本大概会在一千到十万、甚至十万美元左右。但这是可行的,而且在我们之前,理论上就可能有人用一个通用模型去证伪 Erdosia to disc distance conjecture。只是一直没有人充分探索过:如果我往 5.5 里投入价值 10 万美元的 compute(算力),它到底能做到什么?
Speaker 119:18 - 19:21
And the answer is, yeah, you probably could get stuff like that out of it.
Speaker 119:18 - 19:21
而答案是,对,你大概确实能从它那里得到这类结果。
Speaker 219:21 - 19:24
So people should be experimenting more with the current generation in terms of
Speaker 219:21 - 19:24
所以,人们应该更多地拿当前这一代模型来做实验,看看它们在……方面到底能做到什么。
Speaker 119:24 - 19:42
Well, this is, I think, an interesting question of, is it worth it to experiment with because, again, the model release cycle is every every couple months, we put out a new model that's even more powerful. And so the cost of disproving the air to ocean at distance congestion drops by, like, 10 or a 100 x with every model release cycle, probably, in some cases, more.
Speaker 119:24 - 19:42
不过,我觉得这里有个很有意思的问题:这样做值不值得去实验?因为,还是那句话,模型发布周期基本上每隔几个月就是一轮,我们会推出一个更强的新模型。所以,去证伪那个 air to ocean at distance congestion 的成本,几乎会随着每一轮模型发布下降 10 倍或 100 倍,在某些情况下甚至更多。
Speaker 219:42 - 19:48
So You've seen the meme that's like, oh, alright. Like, why why bother doing any engineering work when I should just wait for the next
Speaker 219:42 - 19:48
所以你肯定见过那个 meme(梗图):哦,好吧,那我为什么还要做任何 engineering(工程)工作?我还不如直接等下一次——
Speaker 119:48 - 19:55
model release? On vacation and come back two months later, and then it's, you know, a thousand times cheaper. So agree with that? I Is that
Speaker 119:48 - 19:55
model release(模型发布)呢?去度个假,两个月后再回来,到时候你知道的,成本就便宜一千倍了。所以你同意这种说法吗?我是说,这是吗?
Speaker 219:55 - 19:58
what you're doing right now at OpenAI, just waiting for the next model release?
Speaker 219:55 - 19:58
你现在在 OpenAI 做的事情,就是等着下一个 model 发布吗?
Speaker 119:58 - 20:12
I think I I mean, I will say we're in we're in a period where progress is very fast, and, like, yeah, the models are becoming more capable. I I can say, like, at OpenAI, one of the things that we're we're actively not doing look. We have a lot of mathematicians. We have a lot of physicists. People are very excited about what these models can do right now Mhmm.
Speaker 119:58 - 20:12
我想,我是说,我们现在正处在一个进展非常快的时期,而且,没错,models 正在变得越来越有能力。我可以说,在 OpenAI,我们正在主动避免做的一件事是——你看,我们有很多数学家,也有很多物理学家。大家都对这些 models 目前能做到什么感到非常兴奋,嗯哼。
Speaker 120:12 - 20:58
Especially, you know, the internal models. We are trying to encourage people to not spend all their time just, like, going through all the mathematical open problems, physics problems, and just seeing pushing the models to their limits to see what they can prove or disprove Because we really think the focus should be on how do we make even more capable models, how can we get them get them out safely to the world as quickly as possible so that all the scientists in the world can use these models to solve the problems themselves. So, yeah, in some sense, we are thinking about this that, yes, it's really tempting to just put all of our efforts into scaling up these models and see what they can do with their limits right now, but really the focus should be on how do we use these models to make even more powerful models, even more capable models that can do everything much more cost effectively.
Speaker 120:12 - 20:58
尤其是,你知道,内部 models。我们在努力鼓励大家不要把所有时间都花在,比如,把各种数学未解问题、物理问题过一遍,然后不断把 models 推到它们的极限,看看它们能证明什么、证伪什么。因为我们真的认为,重点应该放在:我们如何做出能力更强的 models,如何尽可能快且安全地把它们交付给全世界,这样全世界所有的科学家都能用这些 models 自己去解决问题。所以,对,在某种意义上,我们确实在想这件事:没错,把我们所有精力都投入到扩展这些 models、看看它们在当前极限下能做什么,这确实非常诱人;但真正的重点应该是,我们如何用这些 models 去造出更强大的 models、能力更强的 models,并且让它们以高得多的成本效益去完成各种事情。
Speaker 220:59 - 21:23
What is changing about the direction or allocation of resources for research in your mind given your beliefs about this very large scale the impact of very large scale test time compute? How does this interact with the idea of recursive self improvement, for example, where, you know, it's a dominant idea for how, you know, any lab gets to the best capability model.
Speaker 220:59 - 21:23
鉴于你对这种超大规模 test time compute(测试时计算)的影响的看法,在你看来,研究方向或资源分配正在发生什么变化?这又如何与 recursive self improvement(递归式自我改进)这样的想法相互作用?比如说,你知道,这一直被认为是任何实验室获得最强 capability model(能力最强的模型)的主导性思路。
Speaker 121:23 - 21:32
So one thing I should clarify, I don't think we're at the point where, okay, you just give it an arbitrary an an extremely high inference budget, and it's just it's just super intelligent across the board.
Speaker 121:23 - 21:32
所以,有一件事我应该先澄清:我不认为我们已经到了这样一个阶段——就是,好吧,你只要给它任意的、极高的 inference budget(推理预算),它就会在各个方面都变得超级智能。
Speaker 221:32 - 21:33
Slash goal
Speaker 221:32 - 21:33
斜杠目标
Speaker 121:33 - 21:37
Yeah. Baseline. Okay. G p d seven or whatever and then, like, yeah, just go nuts.
Speaker 121:33 - 21:37
对。基线。好吧。G p d seven 之类的,然后,就是,没错,直接放开干。
Speaker 221:37 - 21:39
What's between us and there then?
Speaker 221:37 - 21:39
那么,我们距离那种状态之间还差什么?
Speaker 121:39 - 22:08
I think having played around with the model so okay. So first of all, there are some benchmarks where the models will just not improve if they have more inference budget. So I think a lot of factual retrieval kind of questions fall into this category of if you ask a person when was Abraham Lincoln born and they don't know the date, they could sit there. They could think about it for a week. If they if they don't have access to a computer or something, they're not gonna be able to do better answering that question if they thought about it for a week compared to five seconds.
Speaker 121:39 - 22:08
我觉得,在实际上手玩过这个 model 之后,情况大概是这样的。首先,有一些 benchmark,即使给 model 更多 inference budget(推理预算),它们也不会变好。所以我认为,很多事实检索类问题都属于这一类:如果你问一个人 Abraham Lincoln 是哪年出生的,而他不知道这个日期,那他可以坐在那里想。他就算想一周,如果没有电脑之类的工具可查,也不可能比想五秒钟时更会回答这个问题。
Speaker 122:09 - 22:38
Same with the model. If you actually, interestingly enough, if you give the model these kinds of, like, factual retrieval questions and you give them a little bit of time to think, they do actually do But if you give them a week, they're not suddenly gonna do better at remembering dates. There are so there's some benchmarks where they clearly improve with more test time compute, and there's some where they don't. I think on the other extreme, there are benchmarks where they kind of obviously will keep improving limit without limit with more test time compute. So the example I like to point to is sudoku.
Speaker 122:09 - 22:38
model 也是一样。实际上,很有意思的是,如果你给 model 这种事实检索类问题,再给它一点思考时间,它确实会做得更好一些;但如果你给它一周时间,它也不会突然更擅长记住日期。所以,有些 benchmark 会随着更多 test-time compute(测试时算力)而明显提升,也有些不会。另一方面,在另一个极端,也有一些 benchmark 很显然会随着更多 test-time compute 持续提升,几乎没有上限。我最喜欢举的例子是 sudoku。
Speaker 122:39 - 23:07
If you it's there's a really simple strategy to solving sudoku, which is just try a bunch of different random numbers and then see if it fits the criteria, if it if it matches all the constraints. And if it doesn't, just try a different random combination of numbers. And, clearly, with enough time, you will be able to solve any Sudoku puzzle with this strategy. You can kind of trivially see, like, okay, any model could keep doing better and better if it was just given more test compute. So you have and and all the benchmarks kind of exist somewhere between these two extremes.
Speaker 122:39 - 23:07
解 sudoku 其实有一个非常简单的策略:就是尝试很多组不同的随机数字,然后看它是否符合要求、是否满足所有约束。如果不满足,那就换一组随机数字继续试。很显然,只要时间足够长,用这个策略你最终能解出任何一个 Sudoku 谜题。所以你几乎可以很直观地看出来:只要给更多 test compute,任何 model 都可以不断做得更好。所有这些 benchmark,基本都分布在这两个极端之间的某个位置。
Speaker 123:07 - 23:39
The models are not at the level where if you just give them enough test time compute, they will be able to do all of our jobs just because, yeah, there's some benchmarks where they will not improve. There are some things where they they will not improve. One thing I see for research in particular is they don't have very good research taste right now, and so I think they're actually a very good complement to researchers, especially, you know, I've found, like, I've found that I've been much more effective by using these models, but they're not able to fully replace the whole research cycle. Now does that change with time? Probably.
Speaker 123:07 - 23:39
这些 model 还没有达到这样一种程度:只要给足够多的 test-time compute,它们就能完成我们所有人的工作。因为,没错,确实有一些 benchmark 它们不会提升;有些事情它们就是不会变得更好。尤其在 research 上,我看到的一个问题是,它们现在还没有很好的 research taste(研究品味/判断力)。所以我认为,它们实际上是 researchers 很好的补充工具。特别是,就我自己的体验来说,使用这些 model 之后,我的效率高了很多;但它们还不能完整替代整个 research cycle(研究周期)。以后这会不会改变?大概会。
Speaker 123:39 - 23:51
I mean, I think the models are getting better across the board. Some some things are getting better faster than others, but they're not at the point where they're fully replacing researchers with just enough test time compute.
Speaker 123:39 - 23:51
我的意思是,我认为这些 model 在各个方面都在进步。有些方面进步得比另一些更快,但它们还没到这样一个点:只靠足够多的 test-time compute,就能完全取代 researchers。
Speaker 223:51 - 23:59
Can you give an example or two of, like, asking the model to do a research task for just like this? Is a terrible idea.
Speaker 223:51 - 23:59
你能不能举一两个例子,说明让 model 就这样直接去做 research task(研究任务)为什么会是个糟糕主意?
Speaker 123:59 - 24:36
I mean, I I think going back to my poker solver example, I was really impressed with the model's ability to optimize the algorithms that I had developed in my in my PhD. It was honestly it it was it was shocking to see how inefficient I was, in retrospect, and they were able to make it, like, you know, 1,000 x faster. And then I was like, okay. Can you come up with an algorithm that is better than the algorithms that I came up with or that anybody else came up with? You and go ahead and, like, look at all the published work and synthesize that and then try to come up with something novel, and it it's not able to do it.
Speaker 123:59 - 24:36
我的意思是,回到我之前那个 poker solver 的例子,我对 model 优化我在 PhD 期间开发出的那些算法的能力印象非常深刻。说实话,事后回头看,我当时的低效程度简直让人震惊,而它们居然能把速度提升到大概 1,000x。然后我就想,好,那你能不能提出一种比我想出来的、或者比其他任何人想出来的都更好的算法?你去看看所有已发表的工作,把它们综合起来,然后试着提出一些新的东西。但它做不到。
Speaker 124:36 - 24:52
And I can give it a lot of time, it's it's still not able to do it. Now it's possible that if I scaffolded something and, like, kind of constrained it a bit more, that maybe it could eventually come up with something better, but it would take a lot of it's not just as simple as saying, okay. Please come up with a better algorithm.
Speaker 124:36 - 24:52
而且我就算给它很多时间,它还是做不到。现在也不是说完全不可能;如果我给它搭一个 scaffold(脚手架式流程),再多加一些约束,也许它最终可能会提出更好的东西。但这需要做很多额外工作,并不是简单地说一句“好,请想出一个更好的算法”就可以。
Speaker 224:52 - 24:54
And how do you think that that gets improved?
Speaker 224:52 - 24:54
你觉得这要怎么改进?
Speaker 124:55 - 25:17
What I've seen is with every model release cycle, it does get better at this sort of thing. It's still it's still bad in my opinion, but it it's not as bad as it used to be. And I wouldn't be surprised if at some point same thing with coding, same thing with math, where there's just, like, this inflection point where suddenly it's actually good enough to be useful. I wouldn't be surprised if we encounter that point for research tastes as well.
Speaker 124:55 - 25:17
我看到的是,每一轮 model 发布周期,它在这类事情上确实都会变得更好。按我的看法,它现在依然还是做得不好,但已经没有以前那么糟了。而且我不会惊讶于某个时刻会出现类似 coding、类似 math 的情况——就像出现一个拐点,突然之间它已经好到足够有用了。如果我们在 research taste(研究品味)上也遇到那个点,我也不会意外。
Speaker 225:18 - 25:23
Even that, what do you like, what is your framing of RSI today? Like, how should we think about it?
Speaker 225:18 - 25:23
即便如此,你会怎么看,你对今天的 RSI 的框架是什么?也就是说,我们应该怎样理解它?
Speaker 125:24 - 25:51
The models are definitely accelerating what researchers can do inside the labs. But I think they are accelerating some things and not other things. And currently, we're at the point where, okay, if something goes a 100 x faster, you get bottlenecked by the things that don't go a 100 x faster. Over time, the things that we're getting bottlenecked on are going to shrink, and and there will be, I think, a kind of a gradual takeoff in that respect. But it's more about transforming right now, it's more about transforming what researchers do rather than fully replacing the researchers.
Speaker 125:24 - 25:51
这些 model 的确在加速研究人员在 labs 内部能做的事情。但我认为,它们是在加速某些事情,而不是所有事情。现在我们所处的阶段是:如果某件事快了 100 倍,你就会被那些没有快 100 倍的事情卡住。随着时间推移,我们当前受限的那些瓶颈会逐渐缩小,我认为,从这个角度看,会出现一种渐进式的起飞。但就目前而言,这更像是在改变研究人员做事的方式,而不是完全替代研究人员。
Speaker 225:52 - 25:57
So that actually implies that you don't think we're close to a very fast takeoff right now.
Speaker 225:52 - 25:57
所以这其实意味着,你认为我们现在并不接近一次非常快速的起飞。
Speaker 125:58 - 26:38
I think fast takeoff is relative. Things are moving very fast, but I think there is this hypothesis that you could have basically an overnight intelligence explosion where the models discover some kind of breakthrough to make themselves smarter, and then that leads to more breakthroughs that make themselves even smarter immediately. And you have, basically, in an instance, the model's just, you know, becoming very superhuman across the board in in moments. And I don't think we're headed to that world largely because of the fact that the models rely so much on large scale test on compute in order to achieve their greatest intelligence. If you if it requires so much test on compute to unlock the full capabilities of the model, then that means you're bottlenecked by time.
Speaker 125:58 - 26:38
我觉得 fast takeoff(快速起飞)是相对的。事情的发展确实非常快,但有一种假说认为,基本上可能会出现一场一夜之间的 intelligence explosion(智能爆炸):model 发现某种能让自己变得更聪明的突破,然后这又立刻带来更多突破,让自己变得更聪明。于是,基本上就在一瞬间,model 会在几乎所有方面都变得极度超人。我不认为我们正走向那样的世界,主要原因在于,model 在达到最高智能水平时,非常依赖大规模的 test-time compute(测试时算力)。如果要释放 model 的全部能力需要如此大量的 test-time compute,那就意味着你会受制于时间。
Speaker 126:40 - 27:05
Things can only go so fast because the models need to run for long enough to actually do something really, really powerful. Time itself becomes a bottleneck to what we can do. And I I think that is the case right now for a lot of the labs that, ultimately, I think the biggest bottleneck for all of us is time. And that's why all the researchers are working so intensely right now. It's it's just so many so many hours per week are being put into this because we all see what the overhang is.
Speaker 126:40 - 27:05
事情只能快到一定程度,因为 model 需要运行足够长的时间,才能真正做出非常非常强大的事情。时间本身就成了我们所能做到之事的瓶颈。我认为,现在很多 labs 的情况就是如此:归根结底,我觉得我们所有人最大的瓶颈都是时间。这也是为什么现在所有研究人员都在如此高强度地工作。大家每周都在投入非常非常多的工时,因为我们都看得到那个 overhang(潜在积压能力)。
Speaker 127:05 - 27:09
We see what the capabilities are, and we're just bottlenecked by how quickly can we do things.
Speaker 127:05 - 27:09
我们看得到这些能力是什么,而我们真正受限的,只是我们能多快把事情做出来。
Speaker 227:10 - 27:17
What do you think is on the frontier that is less explored now? Like, we've talked about multi agent before.
Speaker 227:10 - 27:17
你觉得现在哪些前沿方向还比较少被探索?比如,我们之前聊过 multi agent。
Speaker 127:17 - 27:41
I think multi agent is quite explored. I think there's sufficient scale? I think there's a lot more that could be done, but it's also one of the things that's that's hard to do at small a lot of research is hard to do at small scale. I think multi agent in particular, it really requires in order to fully unlock the capabilities, think requires, like, frontier models. I think we've seen some pretty interesting multi agent scaffolds.
Speaker 127:17 - 27:41
我觉得 multi agent 已经被探索得相当多了。我觉得规模上已经足够?我认为还有很多事可以做,但这也是那种很难在小规模上做的方向之一,很多研究都很难在小规模上推进。我觉得 multi agent 尤其如此:如果真要把它的能力完全释放出来,我认为需要 frontier models(前沿模型)。我觉得我们已经看到一些相当有意思的 multi agent scaffolds(多 agent 脚手架)。
Speaker 127:41 - 28:11
I think, they're able to do a lot, but I think it's really just scratching the surface of what it will be able to do. I mean, one way that I think about it is if if you look at human civilization, it's not that humans have become smarter over it's not that they evolved to become smarter over, you know, the past fifty thousand years. It's that humans are able to do a lot more today than they were back in caveman times because there have been billions of humans thinking for a long time and building off of each other's accumulated knowledge.
Speaker 127:41 - 28:11
我觉得,它们已经能做很多事了,但我认为这真的还只是触及了它未来能力的表面。我的一个思考方式是:如果你看人类文明,并不是说人类在过去五万年里变得更聪明了,也不是说他们进化得更聪明了。人类今天之所以能做比穴居人时代多得多的事,是因为有数十亿人长期思考,并在彼此积累的知识之上不断构建。
Speaker 228:11 - 28:15
We have, like, very good retrieval and scaffolding versus fifty thousand years ago.
Speaker 228:11 - 28:15
相比五万年前,我们现在拥有非常好的 retrieval(检索)和 scaffolding(脚手架式组织)。
Speaker 128:15 - 28:36
It's not even I wouldn't even call it a scaffold. This is, like this is a very, like, organic emergent property of just, like, humans being able to accumulate knowledge, share it, and build off of it. We're not seeing that with AI models today. They kind of they're they're born into a world for and they exist for a very short context window, and then they just, like, disappear. Mhmm.
Speaker 128:15 - 28:36
这甚至——我甚至都不想把它叫作 scaffold。这更像是一种非常有机的涌现属性:人类能够积累知识、分享知识,并在此基础上继续构建。今天的 AI models(AI 模型)还没有呈现出这一点。它们有点像是被生到一个世界里,只存在于一个非常短的 context window(上下文窗口)中,然后就那样消失了。嗯。
Speaker 128:36 - 29:00
And, yeah, there are things that you can kinda do to, like, continue them, but it's very limited. I do think eventually we will and we're starting to see, like, signs that we're entering a world where they can coordinate on a large scale. I think MoltBook and Openclaw, when they first came out, I think it was obviously a bit overhyped, but they were an indication of where things could go in the future. And I do think that eventually we get to that kind of world.
Speaker 128:36 - 29:00
是的,当然也有一些办法可以某种程度上让它们延续下去,但非常有限。我确实认为,最终我们会——而且我们已经开始看到一些迹象——进入这样一个世界:它们可以在大规模上进行协同。我觉得 MoltBook 和 Openclaw 刚出来的时候,显然有点被过度炒作了,但它们确实指示了未来事情可能发展的方向。我也确实认为,我们最终会到达那样一种世界。
Speaker 229:01 - 29:04
Of some sort of coordinated compounding state.
Speaker 229:01 - 29:04
某种协调的、复利式累积状态。
Speaker 129:04 - 29:11
Yeah. The ability of the models to to share knowledge on a more global level and be able to build on that knowledge productively.
Speaker 129:04 - 29:11
对,就是模型能够在更全局的层面共享知识,并且能够富有成效地在这些知识之上继续构建的能力。
Speaker 229:11 - 29:47
Given this set of beliefs and your work, like, how would you characterize just competition at the frontier between the three kingdoms, if there is no overnight takeoff. It's just researchers grinding away, trying to make good high taste algorithmic and investment decisions about where to go, and then compute allocation, and then policy decisions, and eval decisions. It feels like slightly more grounded than I support I suppose, like racing towards some immediate hard takeoff that nobody can catch you on.
Speaker 229:11 - 29:47
考虑到这一整套信念以及你所做的工作,比如说,如果不存在一夜之间的 takeoff(突飞猛进),那么你会如何描述三大王国在前沿上的竞争?如果只是研究人员埋头苦干,努力在去哪里投入上做出高质量、有品味的 algorithmic(算法)和 investment(投资)决策,然后再做 compute allocation(算力分配)、policy(政策)决策,以及 eval(评估)决策。这种图景我想比那种“朝着某个立刻发生、别人根本追不上的 hard takeoff(硬性爆发)狂奔”的说法,要更 grounded(更贴近现实)一些。
Speaker 129:47 - 30:14
I think the competition is very intense right now. I do think the models that exist today are accelerating what researchers at the Frontier Labs can do. There's obviously, like I said, limits to that right now, but the ability to use the models to improve the the model research is a real thing, and it's it it is like an amplifying force. I think that will continue to be true. I think they'll become more true over time.
Speaker 129:47 - 30:14
我认为现在的竞争非常激烈。我确实认为,今天已经存在的这些模型,正在加速 Frontier Labs 的研究人员所能做的事情。显然,正如我刚才说的,这在当下还是有限度的,但利用模型去改进模型研究这件事是真实存在的,而且它确实像是一种放大力量。我认为这会继续成立,而且我觉得随着时间推移,这一点会变得越来越真实。
Speaker 130:14 - 30:42
One thing that I am comforted by is I think all the researchers at the frontier labs I all the frontier labs, I think, recognize what is at stake and what these models like, what what the what the risks are. And that's something that I I find comforting that I think everybody really understands. Like, okay. This is a pretty serious thing, and it can lead to really great things or it can lead to really bad things. And, yes, there's a competitive dynamic between the labs, but, like, we can also try to figure out how we all get to the positive outcomes rather than the very negative outcomes.
Speaker 130:14 - 30:42
有一件事让我感到些许安心,那就是我认为所有 frontier labs 的研究人员,我想所有 frontier labs 都一样,都认识到了利害所在,也认识到了这些模型的风险是什么。我觉得这一点很令人宽慰,因为我认为大家都真的明白:好吧,这是一件相当严肃的事情,它可能带来非常好的结果,也可能带来非常糟糕的结果。是的,labs 之间存在竞争态势,但我们也可以努力去弄清楚,怎样才能让大家都走向积极的结果,而不是非常负面的结果。
Speaker 230:43 - 31:03
I I think, you know, I'd be remiss to ask just because you have been right very early for a long time on the importance of test time compute and reasoning as a framework. Like, are there ways in which you use the models that you should you would encourage others to? Right? Is it just goal everything? I think for a
Speaker 230:43 - 31:03
我觉得,你知道,我如果不问这个问题就有点失职了,因为你在 test time compute(测试时算力)以及 reasoning(推理)作为一种框架的重要性上,很早就判断对了,而且对了很久。比如,在你使用这些模型的方式里,有没有哪些做法是你会鼓励别人也去采用的?对吧?是不是凡事都把目标先交给它?我觉得对于很多
Speaker 131:03 - 31:33
lot of people, they worked I mean, this is probably not even true for your audience necessarily, but there's a lot of people that experimented with AI back in, like, 2023 and felt like they couldn't trust the outputs and then don't use it for really high stakes decisions. And, actually, I think the models have progressed to a point where they are very good for these kinds of things. I mean, I asked them tax advice, or I bought a condo recently, and I was asking it for advice. I'm like, okay. Well, what's all the paperwork that I have to fill out, and, like, how do how do I what does it all mean?
Speaker 131:03 - 31:33
人来说,他们曾经——我的意思是,这对你的受众来说可能甚至未必成立——但确实有很多人在 2023 年那会儿试过 AI,结果觉得自己没法信任它的输出,所以后来就不会把它用于真正 high stakes(高风险、高重要性)的决策。可实际上,我认为模型已经进步到了一个程度,在这类事情上它们已经非常好用了。我的意思是,我会问它 tax advice(税务建议),或者我最近买了一套 condo,我也会向它征求建议。我会想,好吧,那我到底需要填写哪些 paperwork(文书材料),以及我该怎么处理,这一切到底是什么意思?
Speaker 131:33 - 31:35
It's actually really good for these kinds of questions.
Speaker 131:33 - 31:35
它处理这类问题其实真的非常好。
Speaker 231:35 - 31:36
Mhmm.
Speaker 231:35 - 31:36
嗯。
Speaker 131:36 - 31:50
So I use it day to day for for a lot of this kind of stuff, and I think they're at a point now where they've actually been at a point for a while now where I feel like I can just trust the outputs, arguably more than I could trust the output from from a human person. An expert
Speaker 131:36 - 31:50
所以我在日常生活中会用它来处理很多这类事情,而且我觉得它们现在已经到了这样一个阶段——实际上到这个阶段已经有一段时间了——我感觉自己就是可以信任它们的输出,甚至可以说,可能比我信任一个 human person(真人)给出的输出还要更多。一个 expert(专家)
Speaker 231:50 - 32:04
human. Yeah. Okay. I have two two final questions for you. One is, is there something you think that the rest of the research community doesn't agree with you on or doesn't understand the importance of quite yet?
Speaker 231:50 - 32:04
人类。对,好。最后我还有两个问题想问你。其中一个是:你是否觉得,有什么事情是研究界其他人并不同意你的,或者说他们还没有真正理解其重要性的?
Speaker 132:05 - 32:07
Oh, this is such a good question. I wish I had time to think about this ahead of time.
Speaker 132:05 - 32:07
哦,这真是个特别好的问题。真希望我之前有时间先想一想这个问题。
Speaker 232:07 - 32:15
You can you can just hang out with me and think about it. Yeah. Is it weird to be, like, consensus now? You're a bit salty three years ago when you're like, why don't people understand how important this is?
Speaker 232:07 - 32:15
你可以就跟我一起待着,边聊边想。对。现在这算不算已经成了某种 consensus(共识)?三年前你还有点不爽,会想,为什么大家就是不明白这件事有多重要?
Speaker 132:15 - 32:21
I still I still feel like it's not consensus, though, because, like, you know, people still don't publish the benchmarks this way.
Speaker 132:15 - 32:21
不过我还是觉得这还不算 consensus(共识),因为,你知道,大家还是没有用这种方式来发布 benchmark(基准测试)结果。
Speaker 232:21 - 32:22
Oh, that's true. Yeah.
Speaker 232:21 - 32:22
哦,这倒是真的。对。
Speaker 132:23 - 32:27
Yeah. That's actually why I wrote like inertia. That's kinda yeah. But that's kinda why I wrote the essay. I was just like, look.
Speaker 132:23 - 32:27
对。这其实就是我写那篇文章的原因,有点像是 inertia(惯性)吧。对,差不多就是这样。但这也正是我写那篇文章的原因。我当时就是想说,听着。
Speaker 132:27 - 32:42
I mean, we can talk about this, but, like, yeah, the part of the motivation is, like, I would talk to researchers about we it makes sense to show the benchmarks with an x axis. Whether it's tokens or cost or time, there there should be an x axis. And everybody would say, like, yeah. That makes sense. We should do that.
Speaker 132:27 - 32:42
我的意思是,这个我们可以展开聊,但,对,动机的一部分就是,我会和研究人员讨论:把 benchmark(基准测试)结果用带有 x 轴的方式展示是有意义的。不管这个 x 轴是 tokens(token 数)、cost(成本)还是 time(时间),都应该有一个 x 轴。然后每个人都会说,对,这很有道理。我们应该这么做。
Speaker 132:42 - 32:43
But
Speaker 132:42 - 32:43
但是
Speaker 232:43 - 32:47
But they're not acting with the importance of, like, a good heart. Like, this is we have to measure the correct thing.
Speaker 232:43 - 32:47
但他们并没有带着那种真正重视这件事的态度去行动。就好像,这件事关乎我们必须衡量正确的东西。
Speaker 132:47 - 32:58
Well, really, their response is people expect us to to publish the Grid. Mhmm. And then well, okay. Well, why do people expect the Grid to be published? Because everybody publishes the Grid.
Speaker 132:47 - 32:58
嗯,说到底,他们的回应是:人们期待我们发布 Grid。嗯哼。然后,好吧。那为什么人们会期待 Grid 被发布呢?因为所有人都会发布 Grid。
Speaker 132:58 - 33:28
And so you kind of end up in this this bad equilibrium where everybody kind of knows that it's a bad equilibrium, but, like, nobody wants to break out. And I I felt like, okay. Well, if I just hopefully come out and say, like, look, guys, let's all recognize that we're in a bad equilibrium, and let's move to this different equilibrium where we're we're plotting things with an x axis, that hopefully that can no. Next time there's a model release, a company can feel comfortable not publishing the grid, at least not at the very front, the top line, and, we can have a more productive evaluation of these models.
Speaker 132:58 - 33:28
所以你最后会落入这种糟糕的均衡:每个人其实都知道这是个坏均衡,但就是没人愿意跳出来打破它。而我当时觉得,好吧,如果我直接站出来说,大家看,让我们都承认我们现在处在一个坏均衡里,然后转到另一个不同的均衡——在那里我们用 x 轴来绘图——也许这样就可以,不,应该说,下次再有 model release(模型发布)时,公司就能更自在地选择不发布那个 grid,至少不用把它放在最前面、最显眼的 headline 位置。这样我们也就能对这些模型做出更有成效的评估。
Speaker 233:29 - 34:13
Then a last question for you. How do you think about companies across all of these specialized domains who feel the value that they have is essentially like the routing layer, the choice layer of, you know, my goal is composed of a bunch of discrete tasks. Some require more intelligence and less. And within my job as a vendor is to solve that problem or achieve the optimal outcome with taking into account the budget constraints. And so I will manage, like, the parallelization and how how much inference do you spend on it from what model.
Speaker 233:29 - 34:13
最后一个问题。你怎么看那些分布在各种专业领域中的公司:它们认为自己的价值本质上在于 routing layer(路由层)、choice layer(选择层)。也就是说,我的目标是由一系列离散任务组成的,其中有些任务需要更多 intelligence(智能),有些则不需要那么多。而作为 vendor(供应商),我的工作就是解决这个问题,或者在考虑预算约束的前提下实现最优结果。所以我会去管理,比如并行化,以及针对不同模型该投入多少 inference(推理)计算。
Speaker 234:14 - 34:30
Because I I think the the FrontierLab point of view is that that routing happens both within the you know, behind the API, behind the application, and then some of it in the model itself. And that's pieces of that are clearly being externalized in all these applications.
Speaker 234:14 - 34:30
因为我觉得,FrontierLab 的视角是,这种 routing 一部分发生在 API 背后、application(应用)背后,另一部分则发生在模型本身内部。而其中有些部分显然正在这些应用里被 externalize(外部化)。
Speaker 134:30 - 34:46
Yeah. I do I do think this is related to the fact that, like, benchmarks should be evaluated with an x axis of tokens or cost. I I have seen some evals recently that show, like, okay. Well, with with a routing layer, you can achieve much better performance Mhmm. By basically doing consensus among the models.
Speaker 134:30 - 34:46
对。我确实觉得这和这样一个事实有关:benchmark(基准测试)应该用 tokens 或 cost 作为 x 轴来评估。我最近确实看到一些 evals(评测)显示,好吧,有了 routing layer,你可以获得好得多的性能,嗯哼,基本方法就是让多个模型之间形成 consensus(共识)。
Speaker 134:46 - 35:02
Yeah. And, like, I definitely believe that if you do consensus among the models, that you're gonna achieve better performance than any individual model. But it's important to ask, like, are you gonna do better than having that model basically think for longer? Like, once you control for the amount of test time compute, is it is it actually still doing better? Mhmm.
Speaker 134:46 - 35:02
对。而且,我当然相信,如果你让多个模型做 consensus,那么你得到的性能会比任何单个模型都更好。但重要的是要问:这会不会比让那个模型本质上思考更久还要更好?也就是说,一旦你控制了 test time compute(测试时计算量),它是不是仍然真的更好?嗯哼。
Speaker 135:02 - 35:04
That's that's the question that you want to figure out.
Speaker 135:02 - 35:04
这才是你真正想搞清楚的问题。
Speaker 235:04 - 35:15
Okay. That's very principle of you, which is like, yes, routing is fine, but it's all subject to the same budget question. Yep. Right. If you put it on the same scaler, then you can make an optimal decision, and I think maybe I win.
Speaker 235:04 - 35:15
好的,这很符合你的原则性。也就是说,没错,routing 可以,但它同样要受制于同一个预算问题。对。只要把它放到同一个标尺上,你就可以做出最优决策,而我觉得也许那样我就赢了。
Speaker 135:15 - 35:33
Mhmm. I I I don't even know necessarily that I I would believe that the routing does better, but then there's still a question of, is it gonna do significantly better? Is it very fragile? Is it reflective of real world use cases compared to benchmarks? Because, like, one issue you could run into is that you could optimize for certain benchmarks with the routing, and then show, like, oh, yeah.
Speaker 135:15 - 35:33
嗯。我我我甚至不一定知道我我会相信 routing 的效果更好,但接下来仍然有个问题:它会不会显著更好?它是不是非常脆弱?和 benchmark(基准测试)相比,它是否反映了真实世界的 use case(使用场景)?因为,比如,你可能会遇到的一个问题是,你可以针对某些 benchmark 用 routing 做优化,然后展示说,哦,对。
Speaker 135:33 - 35:49
We see there's big improvement on these benchmarks. But in real world use cases, it actually ends up not being a significant improvement. So I I would say, at the very least, like, I would say, you wanna control for test on compute, and then you also wanna have all the same skepticism about benchmarks that you would normally have.
Speaker 135:33 - 35:49
我们看到这些 benchmark 上有很大的提升。但在真实世界的 use case 里,最终它实际上可能并不是显著的改进。所以我我至少会这么说:你会想要控制测试时的 compute(计算资源),然后你也会想要像平常一样,对 benchmark 保持同样的怀疑态度。
Speaker 235:50 - 35:57
Awesome. Noam, thanks so much, and and for being on the mission, for, breaking us out of this false equilibrium.
Speaker 235:50 - 35:57
太好了。Noam,非常感谢你,也感谢你参与这项 mission(使命),帮助我们打破这种虚假的 equilibrium(均衡)。
Speaker 135:57 - 35:58
Yeah. It's great to be back.
Speaker 135:57 - 35:58
是啊,很高兴再次回来。
Speaker 236:00 - 36:16
Find us on Twitter nopriorspod. Subscribe to our YouTube channel if you wanna see our faces. Follow the show on Apple Podcasts, Spotify, or wherever you listen. That way you get a new episode every week. And sign up for emails or find transcripts for every episode at nopriors.com.
Speaker 236:00 - 36:16
欢迎在 Twitter 上关注我们:nopriorspod。如果你想看看我们的真人出镜,也欢迎订阅我们的 YouTube 频道。也请在 Apple Podcasts、Spotify 或你平时收听的平台上关注这档节目。这样你每周都会收到新一期节目。你还可以在 nopriors.com 订阅邮件,或查看每一期节目的文字稿。
原文 ↗https://www.youtube.com/watch?v=AZrU6y3pUcU
BuildSpeak — 关于本项目BUILT IN PUBLIC · 跟随 builders 而非 influencers