BuildSpeak每日 builder 文摘
今日归档生词本关于
🎙 播客Training Data· 2026 年 6 月 24 日· 8,059 词 · 约 40 分钟

Memory and Continual Learning: Engram's Dan Biderman and Jessy Lin

SPACE 播放 / 暂停·←→ 上一句 / 下一句
Speaker 100:00 - 00:24
What about pre training or even post training makes it possible for the models to generalize in these magical emergent ways and controlling that process so that a company has a set of private data? How do we make the models learn that just as well as the models know, like, the capital of France or, you know, like, how to write Python? So I think it's a really fun problem to think about.
Speaker 100:00 - 00:24
那么,究竟是 pre training(预训练)甚至 post training(后训练)中的什么,让模型能够以这种近乎神奇的 emergent(涌现)方式实现泛化?以及,如何控制这个过程,使得当一家公司拥有一组私有数据时,我们能让模型把这些内容学得和它知道 France 的首都、或者知道怎么写 Python 一样好?所以我觉得这是一个非常有趣、值得思考的问题。
Speaker 200:41 - 00:57
Welcome to training data. We are delighted to have Don Biederman and Jesse Lynn, cofounders of Ngram today. Ngram is a NeoLab focused on memory and continual learning and two of the hottest topics in all of AI research today. And Sean and I are delighted to dig in on those topics with you today.
Speaker 200:41 - 00:57
欢迎来到 Training Data。今天我们非常高兴请到 Ngram 的联合创始人 Don Biederman 和 Jesse Lynn。Ngram 是一家 NeoLab,专注于 memory(记忆)和 continual learning(持续学习)——这也是当今整个 AI 研究中最热门的两个话题。Sean 和我也非常高兴,今天能和你们一起深入探讨这些话题。
Speaker 300:57 - 00:59
Awesome. Happy Great. To be
Speaker 300:57 - 00:59
太棒了。很高兴。很棒。能够来到这里
Speaker 200:59 - 01:09
So maybe to kick off, the Engram website says, we don't see the world through the lens of pretraining or post training. Our models are always training. What does that mean?
Speaker 200:59 - 01:09
那我们也许先从这里开始,Engram 网站上写着:我们不是通过 pretraining(预训练)或 post training(后训练)的视角来看世界。我们的模型始终都在训练。这是什么意思?
Speaker 101:09 - 01:51
I think models today obviously know a lot of things. They're incredibly smart, but we think the bottleneck for making these models more useful these days is not really raw intelligence, but understanding new and evolving context. Whether it's a new task that you're doing or a particular context for a job or something like this, how do you bake that into the model weights the same way that pre training and post training bakes that into the model weights very deeply? This is why we think of ourselves as working on these fundamental problems of memory and continual learning, which are really two sides of the same coin. How do you make the models learn new things and bake them deeply into the weights of the model?
Speaker 101:09 - 01:51
我觉得,如今的模型显然已经知道很多东西了。它们非常聪明,但我们认为,要让这些模型在当下变得更有用,瓶颈其实并不在于原始智力,而在于理解新的、不断演变的 context(上下文)。无论是你正在做的一项新任务,还是某个工作的特定 context,或者类似的东西,关键在于:你如何把这些内容像 pre training 和 post training 那样,深深地烘焙进 model weights(模型权重)里?这也是为什么我们认为自己是在解决 memory 和 continual learning 这些基础问题——它们其实是一枚硬币的两面。你如何让模型学会新的东西,并把它们深深写入模型的 weights 之中?
Speaker 201:51 - 02:01
And is your premise then that memory as a separate database or a separate thing that you've shoved into the context window is not true memory and is not true continual learning?
Speaker 201:51 - 02:01
那么,你们的前提是否是:把 memory 作为一个单独的数据库,或者作为一个被塞进 context window(上下文窗口)的独立东西,这并不是真正的 memory,也不是真正的 continual learning?
Speaker 102:01 - 02:39
I think all of these tools will kind of come together. So these days, the way that people are solving these problems is with context engineering. So you take a huge prompt, maybe you keep talking to the model over many, many turns and hours and reorganize the context to better understand what you're trying to do. We think these kinds of things, like tool use, context engineering will play a part, but I think an under leveraged tool these days is using the same training pipeline or framework or kind of workflow that the frontier labs are using to make these models really good at frontier math or code, but applying that to every kind of domain, every kind of context that you have, let's say in a company.
Speaker 102:01 - 02:39
我认为,这些工具最终都会某种程度上融合在一起。如今,人们解决这些问题的方式主要是 context engineering(上下文工程)。也就是你给模型一个巨大的 prompt(提示词),也许你会在很多很多轮、持续数小时的对话中不断和模型交流,并重组 context,以便更好地理解你想完成的事情。我们认为,这类东西——比如 tool use(工具使用)、context engineering——都会发挥作用;但我觉得,如今一个尚未被充分利用的工具,是使用 frontier labs 正在采用的同一套 training pipeline(训练流水线)或 framework(框架)或 workflow(工作流),这些方法已经被用来让模型在 frontier math(前沿数学)或 code(代码)上表现得非常出色,而我们可以把它应用到每一种 domain(领域)、每一种 context 上,比如说,一家公司内部拥有的各种 context。
Speaker 302:40 - 03:00
Yeah. To me, it's like as an individual, taking notes and having sticky notes is a very valuable thing. We should never discard this. But whenever we get back to business the next day, we always have some sort of trace of memory in our brain, some new intuition about how things should be and where should we look. So these two things should come together.
Speaker 302:40 - 03:00
对我来说,这有点像是:作为个人,记笔记、贴便利贴都是非常有价值的事情,我们绝不该抛弃这些。但当我们第二天回到工作中时,我们的大脑里总会留下一些 memory(记忆)的痕迹,一些关于事情应该如何推进、以及我们应该去哪里寻找的新直觉。所以,这两者应该结合起来。
Speaker 303:00 - 03:23
And current solutions are more kind of externalized memory. And this has two issues. One is that the amount of tokens we will all collectively individually generate is going be in the tens of millions of tokens per day soon. So just keeping it and searching through it is going to be and rereading it, it's going to be pretty expensive, but it's going to also be pretty hard, pretty confusing for the models, unless we have major, major breakthroughs in
Speaker 303:00 - 03:23
而当前的解决方案更像是一种外置化的记忆(externalized memory)。这有两个问题。其一是,我们所有人作为整体、也作为个体,很快每天都会生成数以千万计的 token。所以如果只是把这些内容保存下来、再去搜索、再去重读,成本会相当高;而且除非我们在……方面取得非常非常重大的突破,否则这对模型来说也会相当困难、相当容易造成混淆。
Speaker 203:23 - 03:25
Depends how on billions of tokens for Sean.
Speaker 203:23 - 03:25
这要看 Sean 那边是不是已经到几十亿 token 的规模了。
Speaker 303:26 - 03:27
That's good.
Speaker 303:26 - 03:27
那不错。
Speaker 403:27 - 03:28
Depends on the day.
Speaker 403:27 - 03:28
得看是哪一天。
Speaker 203:29 - 03:35
Could you maybe tell us a little bit about the Ngram architecture or the Ngram product and how it works?
Speaker 203:29 - 03:35
你能不能跟我们稍微讲讲 Ngram architecture 或者 Ngram product,以及它是如何工作的?
Speaker 103:35 - 04:15
Yeah, mean, at a high level, what we're trying to do is take any context. There's all these different workspaces, let's say. We're working with partners like Notion, Microsoft, and Harvey that have these places where people are doing a lot of work over a long period of time. There's all this context, both in terms of documents that you've already written as a team, as well as now people are interacting with these agents more and more in these products, or having conversations, giving them feedback, and figuring out how to have a model that deeply understands that context. So not just reading the files at test time, but really understanding it the way that an employee that's worked at your company for years has.
Speaker 103:35 - 04:15
可以。我的意思是,从高层来看,我们想做的是获取任何上下文(context)。可以说,现在存在各种不同的工作空间(workspaces)。我们正与 Notion、Microsoft 和 Harvey 这样的合作伙伴合作,它们都有一些场景:人们会在这些地方长期进行大量工作。这里面包含着大量上下文,既包括你们团队已经写过的文档,也包括现在人们越来越多地在这些产品里与这些 agent 互动,或者与它们对话、给它们反馈。我们要做的是,弄清楚如何构建一个能够深度理解这些上下文的模型。所以不只是让它在 test time 读取文件,而是真正像一个已经在你公司工作多年的员工那样去理解这些内容。
Speaker 104:15 - 04:49
So you understand at a high level, Oh, these are the initiatives across the company. This is the way that we do things. You've studied how to run the hiring pipeline or how to do this kind of thing within the company and can operate just as well as anybody else can in the company. What we're doing is training per team models within these workspaces that deeply understand those contexts and can improve with time on the things that people care about. The way that we do this at a technical level maybe is training these into weights.
Speaker 104:15 - 04:49
这样它在高层上就会明白:哦,这些是整个公司的各项计划;这是我们做事的方式。它已经学会了如何运作招聘流程,或者如何在公司内部处理这类事情,并且能够像公司里其他任何人一样好地开展工作。我们正在做的是,在这些工作空间里为每个团队训练专属模型,让它们深度理解这些上下文,并且随着时间推移,在人们关心的事情上不断提升。从技术层面来说,我们实现这一点的一种方式,可能就是把这些内容训练进 weights(权重)里。
Speaker 104:49 - 05:07
We do a lot of adapter fine tuning. Adapters of many types. I think people have looked into this for decades at this point, whether it's LORAs or prefixes or sparse architectures. I think all of these tools are at our disposal. Then figuring out what the right data is.
Speaker 104:49 - 05:07
我们会做大量 adapter fine-tuning(适配器微调)。adapter 有很多类型。我觉得到现在为止,人们研究这些东西已经有几十年了,无论是 LORAs、prefixes,还是 sparse architectures(稀疏架构)。我认为这些工具都可以为我们所用。然后接下来就是弄清楚什么样的数据才是合适的数据。
Speaker 105:08 - 05:31
How do you turn any kind of raw document or interaction into useful training signal for the model? Again, we have a variety of tools now like supervised fine tuning, RL, on policy distillation, all of these things that the field has developed and trying to fit these pieces together into a model that learns continuously on the things that people care about.
Speaker 105:08 - 05:31
你如何把任何类型的原始文档或交互,转化为对模型有用的训练信号?再说一次,现在我们已经有各种工具,比如 supervised fine tuning(监督微调)、RL(强化学习)、on-policy distillation(同策略蒸馏)等等,这些都是这个领域发展出来的方法;现在要做的,是尝试把这些部分拼接起来,构建一个能够围绕人们关心的事情持续学习的模型。
Speaker 305:31 - 06:03
Yeah, and it's not a bet that tools are not there. Like our models always work under the assumption that some knowledge is externalized, some tools are always there. But what you need to do is you need to figure out, and that's the hard task is what needs to be internalized and what can be externalized. And even for stuff that's externalized, many individuals and companies have their own bespoke tools and ways of doing things. Not everyone has the same, you know, bash CLI tools that, you know, the frontier models are training on and how to get the models to better understand your bespoke setup, I think is its own interesting thing.
Speaker 305:31 - 06:03
是的,而且这并不是在赌工具不存在。我们的模型一直都建立在这样一个假设之上:有些知识是 externalized(外部化)的,有些工具总是存在的。但你需要做的是弄清楚——而这正是困难所在——哪些东西需要 internalized(内化),哪些东西可以 externalized(外部化)。即便是那些已经 externalized 的内容,很多个人和公司也都有自己定制的 bespoke tools(专用工具)和做事方式。并不是每个人都拥有那些 frontier models(前沿模型)训练时所接触的同样的 bash CLI tools;而如何让模型更好理解你自己的 bespoke setup(定制环境),我认为这本身就是一件很有意思的事。
Speaker 206:04 - 06:19
And so is the premise then that my Notion agent will be a custom agent that is LoRa fine tuned or some way with an adapter tuned so that it's constantly learning on new content that's added into my Notion workspace. Is that Yeah, the
Speaker 206:04 - 06:19
所以你的前提是,我的 Notion agent 会是一个定制 agent,会经过 LoRa fine tuned,或者以某种方式做 adapter tuning(适配器调优),从而能够对不断添加到我的 Notion workspace 里的新内容持续学习。是这个意思吗?对,那个——
Speaker 306:20 - 06:26
and they're working with many models, they're the early users of all the frontier models, and they're probably going to keep doing that.
Speaker 306:20 - 06:26
而且他们在与很多模型合作,他们是所有 frontier models 的早期用户,而且他们很可能还会继续这么做。
Speaker 206:26 - 06:30
Does this approach work on the frontier models, are the closed frontier models or not?
Speaker 206:26 - 06:30
这种方法适用于 frontier models 吗?那些 closed 的 frontier models 也可以吗?
Speaker 306:30 - 06:49
We need white box access to the weights, right? So, we can partner with companies that have closed source weights and do this with them. But it's easiest for us to do it with open source models. But any model that's a transformer model, we can do our thing to it.
Speaker 306:30 - 06:49
我们需要对权重有 white-box access(白盒访问权限),对吧?所以,我们可以和那些拥有 closed-source weights(闭源权重)的公司合作,并和他们一起做这件事。但对我们来说,用 open-source models(开源模型)来做是最容易的。不过,只要是 transformer model(Transformer 模型),我们都可以对它施加我们这套方法。
Speaker 206:49 - 07:10
And what's the trade off then when people are comparing the before and after using you? Is it that they're no longer sending so much context? And so the trade off is like you burn more compute upfront to learn your company's way of doing things into the weights, and then you're sending less context to the model on every inference pass. Is that the rough trade Yeah,
Speaker 206:49 - 07:10
那么,当人们比较使用你们前后的差异时,权衡点是什么?是不是他们不再需要发送那么多 context(上下文)了?也就是说,这种 trade-off(权衡)大致是:你先额外消耗更多 compute(算力),把你公司做事的方式学进权重里;然后在每一次 inference pass(推理过程)中,你发送给模型的 context 就会更少。这个理解大致对吗?对,
Speaker 307:10 - 07:53
that's one thing. The fact that you don't have to research things and reread things and the fact that you don't have to write like monstrous system prompts, can give you, you know, two orders of magnitude reduction in token inference consumption. It's not like, you know, 50% or it can be 100x fewer tokens because many things, especially things that relate to people and teams and organization and priorities, these are things that you can't really find in one document unless like you really have it really regimented and document everything. And these kinds of things, the model can kind of implicitly learn by training on some of the data and answer within 100 tokens, the best frontier models would consume 100,000 tokens doing. So these kinds of examples are interesting.
Speaker 307:10 - 07:53
这是其中一点。还有一点是,你不必再去研究各种东西、反复重读材料,也不必再写那种臃肿得吓人的 system prompts(系统提示词);这些都能让你的 token inference consumption(推理 token 消耗)降低两个数量级。不是说只减少 50%,而是可能少 100 倍,因为很多事情——尤其是和人、团队、组织以及优先级相关的事情——你很难只在一份文档里找到,除非你真的把一切都高度规范化并完整记录下来。而这类东西,模型可以通过在部分数据上训练,某种程度上隐式地学会,然后用 100 个 token 之内给出回答;而最好的 frontier models 做同样的事,可能要消耗 100,000 个 token。所以这类例子很有意思。
Speaker 307:54 - 08:19
And also the quality, there are tasks that are not supernatural for the current generation of the models. And we kind of think there's going to be consistently this gap of like three to six months ahead, where there's certain things that are bespoke that people are just exploring, the models are not fully great for them, the models will at some point be great for them, but if you can autonomously learn in a very lightweight way, it will give value in that time and cameras capabilities.
Speaker 307:54 - 08:19
还有质量方面的问题:有些任务对当前这一代模型来说并不是“超人级”的。我们基本认为,这里会持续存在一个大约提前三到六个月的空档期:有些事情是高度定制化的,人们还只是在探索,模型还没有真正把它们做好;模型终究会在某个时点把它们做好,但如果你能以一种非常轻量的方式进行自主学习,那么在这段时间里以及能力演进的过程中,它就能创造价值。
Speaker 208:19 - 08:22
Why train on the workspace level versus the individual level, for example?
Speaker 208:19 - 08:22
比如说,为什么要在 workspace 层级而不是个人层级进行训练?
Speaker 308:22 - 08:45
Either is fine for us. It's just easier to start with, you know, teams of people have, you know, are more, you know, disciplined in how they collect context and in the amount of context they have over years. And it's easy for us to start there. But every person's computer and every person's phone one day is a useful target for our technologies. And in fact, it will be very interesting to go there.
Speaker 308:22 - 08:45
对我们来说两种都可以。只是更容易从团队开始——你知道,团队里的人通常在收集 context(上下文)方面更有纪律性,而且他们多年积累下来的 context 数量也更多。对我们而言,从这里切入更容易。但未来,每个人的电脑和每个人的手机,都会成为我们技术的一个有价值的目标。事实上,走到那一步会非常有意思。
Speaker 308:45 - 08:51
We just think the big deposits of information are now in teams of people collaborating in knowledge work.
Speaker 308:45 - 08:51
我们只是认为,如今信息的大型沉淀库主要存在于从事知识工作的协作团队之中。
Speaker 208:51 - 09:24
Is it a feature or a bug that there is so much fact memorization built into large language models. And there's a school of thoughts that the models just wrote memorizing the fact that the capital of France is Paris is actually a bad thing. And what we would prefer for the models to do is abstractly learn the concepts of countries and capital cities, but not to memorize all these facts in the weights. And so I'm curious what you think about disentangling memorization versus learning, how it's done in the models today, and then how you're thinking of approaching
Speaker 208:51 - 09:24
大语言模型里内置了如此多的事实记忆,这到底是 feature 还是 bug?有一种观点认为,模型仅仅去记住“France 的首都是 Paris”这样的事实,其实是一件坏事。更理想的情况是,模型能够抽象地学会国家和首都城市这些概念,而不是把所有这些事实都硬记在 weights(权重)里。所以我很好奇,你怎么看待 memorization(记忆)和 learning(学习)的解耦:今天的模型里它们是怎么实现的,以及你们打算如何处理这个问题。
Speaker 109:24 - 09:52
Yeah, I think it's a really interesting question. To some extent, you need to remember stuff in order to compose them into more complex concepts. I think the thing that's missing is figuring out what's important to remember. I think even now when you think about learning new knowledge, if you look at a lot of these academic benchmarks, it's like, how can we learn very specific facts, like the length of a bridge in this African country? That's not something that you really want the models to devote capacity for.
Speaker 109:24 - 09:52
是的,我觉得这是一个非常有意思的问题。从某种程度上说,你需要记住一些东西,才能把它们组合成更复杂的概念。我认为真正缺失的是:弄清楚什么是值得记住的。即使现在,当你思考学习新知识这件事时,如果你去看很多这类学术 benchmark(基准测试),它们关注的往往是:我们怎样才能学到非常具体的事实,比如某个 African 国家里一座桥的长度?这并不是你真正希望模型为之投入 capacity(容量)的东西。
Speaker 109:52 - 10:32
It's not something that we devote capacity to. I think if you look at human memory, you can say a lot more about this, but it's lossy because part of the feature of intelligence is compressing what's important and separating that from what's not important. I think you can't really separate fact learning from non fact learning or skill learning as some people would like to think. If you take a model and some people have done this with models where you strip out all the facts and just have it the pure core or something like this, it's very unnatural as a model. It doesn't know basic things and you need that.
Speaker 109:52 - 10:32
这也不是我们想投入 capacity 的东西。我认为,如果你去看人类记忆——关于这一点你其实可以讲很多——它是有损的,因为智能的一个特征,本来就是压缩那些重要的内容,并把它们和不重要的内容区分开来。我认为,你并不能像有些人想的那样,真正把事实学习和非事实学习、或者 skill learning(技能学习)分离开来。如果你拿一个模型——有些人确实这样做过——把所有事实都剥离掉,只留下某种纯粹的核心之类的东西,那对模型来说其实是非常不自然的。它会连一些基本的东西都不知道,而这些东西你是需要的。
Speaker 210:33 - 10:36
Why do you need that? Why can't you look up facts and then just have
Speaker 210:33 - 10:36
你为什么需要那些?为什么不能把事实查出来,然后只保留
Speaker 110:37 - 11:00
I think if those you look at facts? How the models think, if you need to recall basic facts in order to take the next step in your thinking, you can't get very far. Maybe that's a high level intuition, but it's part of the reason why we think training is really important. In order to think more and more complex and deep thoughts about things, you kind of need to internalize something so that you can compose them into more abstract concepts.
Speaker 110:37 - 11:00
我认为,如果你看事实的话——模型是怎么思考的——如果你在思考中为了迈出下一步,需要先回忆起一些基本事实,那你其实走不了太远。也许这只是一个高层次的直觉,但这也是我们为什么认为训练真的很重要的一部分原因。为了对事物进行越来越复杂、越来越深入的思考,你某种程度上需要把一些东西内化,这样你才能把它们组合成更抽象的概念。
Speaker 411:00 - 11:01
Yeah, and
Speaker 411:00 - 11:01
对,而且
Speaker 311:01 - 11:52
there have been efforts before that were hard to scale to try and disentangle the two and pre train the models in a way that allows it to retrieve and search for things and not internalize them. The recipe we know to hill climb on collectively right now is this fact pre training step. And I think the magic of, or the mystery of this approach is that, you know, traditionally in CS, we would have, you know, databases as its own curriculum, and we would have algorithms and the databases is like facts about the world and capitals of whatever, store them or query them. There's also algorithms of how do you efficiently manipulate information and get some answers in a sample efficient way. And I think the magic of deep learning is that these two things are now mushed together, and we need all these smart people and philanthropic interpretability to try and break them apart.
Speaker 311:01 - 11:52
以前也有人做过一些尝试,想把这两者拆开,但很难扩展;他们试图用一种方式来 pre-train(预训练)模型,让它能够去检索、去搜索信息,而不是把这些信息内化进去。就我们目前共同不断优化、逐步爬坡的配方来说,就是这个基于事实的 pre-training(预训练)步骤。我觉得这种方法神奇、或者说神秘的地方在于,传统上在 CS(计算机科学)里,我们会把 databases 当成一个独立课程;我们也会有 algorithms,而 databases 更像是关于世界的事实、各种国家的首都之类——你把它们存起来,或者去查询它们。与此同时,还有 algorithms 这部分:你如何高效地操作信息,并以 sample-efficient(样本高效)的方式得到答案。我认为 deep learning(深度学习)的神奇之处就在于,这两件事现在被糅在一起了,而我们需要所有这些聪明的人,以及带有公益性质的 interpretability(可解释性)研究,来试着把它们重新拆开。
Speaker 311:52 - 12:31
And I think a lot of what we're seeing now in the adoption of AI into the economy is that these things are gradually separating again, where companies have their own context, and they really handle them with care and engineer them with care. There's a generic model that's completely a stranger to these contexts. And the model is operating on them. But for us, it's clear that there needs to be a certain convergence, at least with some cadence, where the facts and the stories and the details are getting mixed into the model. It has disadvantages as well, because if you have to, you know, capitals of countries are, you know, they can change, but it's not very frequent, but there's many other facts that are changing all the time.
Speaker 311:52 - 12:31
我认为,我们现在在 AI 被采纳进经济体系的过程中看到的很多现象是,这些东西正在逐渐再次分离:公司有它们自己的 context(上下文/语境),而且它们确实会非常谨慎地处理这些 context,也会非常谨慎地对其进行工程化。与此同时,还有一个完全不熟悉这些 context 的通用模型。这个模型在这些 context 之上运行。但对我们来说,很清楚的一点是,至少需要以某种节奏实现一定程度的收敛,也就是说,这些事实、叙事和细节需要被混合进模型里。这也有缺点,因为比如说,国家首都这样的事实当然也可能变化,只是变化不那么频繁,但还有很多其他事实是在不断变化的。
Speaker 312:31 - 12:35
Just imprinting them into weights is a challenging thing to do.
Speaker 312:31 - 12:35
仅仅把这些东西压进 weights(权重)里,本身就是一件很有挑战的事。
Speaker 212:35 - 12:46
I see. So you're saying it's a false dichotomy that's trying to separate algorithms from databases here. What really matters is how to distinguish what's important to remember versus what's not important.
Speaker 212:35 - 12:46
我明白了。所以你的意思是,试图在这里把 algorithms 和 databases 分开,是一种错误的二分法。真正重要的是,如何区分什么是值得记住的,什么是不重要、不必记住的。
Speaker 312:46 - 12:48
Exactly. It's an it's
Speaker 312:46 - 12:48
没错。这是一个——
Speaker 212:47 - 12:53
open part of how we dream. Are you guys taking any inspiration from that in terms of ranking?
Speaker 212:47 - 12:53
这是我们如何进行 dream(“做梦”式生成/联想)的一个开放部分。你们在 ranking(排序)这方面,有没有从中获得什么启发?
Speaker 112:53 - 13:15
Very, very loosely, I think. Just the idea that that's kind of a phase that's missing maybe where you take a context and you deeply internalize it. Right now it's like everything happens at test time. You look at the context that the user gives you and you do some thinking on the fly. But again, you can't get very far or you can get so far maybe, and you make mistakes along the way.
Speaker 112:53 - 13:15
我觉得,这是一个非常、非常宽泛的想法。大意只是:这里可能缺少了一个阶段——你拿到一个 context(上下文),然后把它深度内化。现在更像是一切都发生在 test time(测试时)。你查看用户给你的 context,然后临场做一些思考。但再说一次,这样你走不了太远,或者说也许能走一段,但过程中会不断犯错。
Speaker 113:15 - 13:21
How do you digest that back into the model so that next time you do it, you do it the right way and make even more progress?
Speaker 113:15 - 13:21
你要怎样把这些再消化回模型里,这样下次再做时,它就能用正确的方式去做,并且取得更大的进展?
Speaker 313:21 - 13:46
Yeah. And what are dreams? Dreams are pretty crazy things to say we want to build an AI that's like our dreams sounds a little bit like a nut thing to do. There's not a lot of coherence there. But what's interesting there is like, what happens in our dreams, we see things, we talk to ourselves, and we experiment with the affordances of what can we do and can't we do in the world and social situations and in, you know, any, it's heavily biased towards social stuff, right?
Speaker 313:21 - 13:46
对。那 dreams(梦)又是什么呢?梦是很疯狂的东西。说我们想造一个像我们的 dreams 一样的 AI,听起来多少有点疯。但那里其实没有太多连贯性。不过有意思的是,在梦里会发生什么?我们会看到事物,会自言自语,还会去试验各种 affordances(可供性)——在这个世界里、在社交场景里,以及你知道的各种情境中,我们能做什么、不能做什么。而且它明显非常偏向 social(社交)相关的内容,对吧?
Speaker 313:46 - 14:11
So for us too, with things we're building is, you know, we give the models the time to then go back, retreat from the actual interaction and experiment with its affordances. What can it do in an environment? What does it know? How fast can it handle these kind of tail extreme things, the same ones that we dream about at night? You guys come from academic backgrounds.
Speaker 313:46 - 14:11
所以对我们来说也是一样:对于我们正在构建的东西,我们会给模型时间,让它从真实互动中退出来,回头去试验它的 affordances。它在一个环境中能做什么?它知道什么?面对这类长尾极端情况时,它处理得有多快——也就是那些我们夜里也会梦到的情况。你们几位都是学术背景出身。
Speaker 314:11 - 14:16
What's a canonical example that motivates this problem or
Speaker 314:11 - 14:16
有什么典型的例子能说明这个问题,或者
Speaker 414:17 - 14:19
that's a win so far?
Speaker 414:17 - 14:19
到目前为止,有什么成功案例吗?
Speaker 314:19 - 14:48
Yeah. I have one example, and Jesse can give another one. A hypothetical one, for example, imagine one of the AI labs, say OpenAI has to win some math Olympiad in a week time from now. Would they construct a catalog of all the math textbooks and really have people annotate which chapters to get and which graphs to see? Or will they actually collect this, synthesize some training data, launch a training job, see where it lands in five, six days, start evaluating it and stuff like that.
Speaker 314:19 - 14:48
对。我有一个例子,Jesse 也可以再讲一个。先说一个假设性的例子:比如某家 AI lab,假设 OpenAI,必须在一周后赢下一项数学奥林匹克竞赛。他们会不会去整理一整套数学教材目录,再让人逐章标注该学哪些章节、该看哪些图表?还是说,他们其实会把这些资料收集起来,合成一些 training data(训练数据),启动一个 training job(训练任务),看看五六天后训到什么程度,然后开始评估之类的。
Speaker 314:48 - 15:18
So it's obvious for anyone who's trained models that there's a superior way to integrate across the ideas and capabilities and involves this kind of magic of training. And we are clear that this has to happen in those high stake domains of math and coding and cyber and stuff. We just think much of this magic can actually end up in the hands of many more people in interesting ways. Why isn't it just the Foundation Model Labs that own the end product here? How do you go between giants?
Speaker 314:48 - 15:18
所以,对任何训练过模型的人来说都很明显:要把各种想法和能力更优地整合起来,是有一种更好的方式的,而这涉及 training(训练)这种近乎魔法的过程。我们也很清楚,这件事必须发生在数学、coding(编程)、cyber(网络安全)这些高风险领域里。我们只是认为,这种“魔法”其实可以以很有意思的方式落到更多人的手中。为什么最终产品一定只能由 Foundation Model Labs 掌控?你要怎样在这些巨头之间找到空间?
Speaker 115:18 - 16:09
Yeah. I think the worldview that we have is a bit different from the Frontier Lab worldview, where it's like we want one model that's bigger and bigger, that's more and more intelligent across a variety of domains. Instead, how we see it, we imagine this world where everybody has their own model. A lot of the things that people want to learn are either private, things that will never see the light of day in a post training dataset, or even conflicting, like, Oh, the way that I want to do the task is different from how another company or another individual wants to. I think a lot of these things we're already seeing are hard to train into the models with the same tools that we have used for decades in machine learning, which is you have really clean supervision, you have ground truth reward signals, and you create a nice environment and you train the model to use the tools to better accomplish this coding task.
Speaker 115:18 - 16:09
是的。我觉得我们的世界观和 Frontier Lab 的世界观有点不同。后者更像是想要一个单一模型,而且这个模型要越来越大、在各种领域都越来越智能。相反,我们设想的是一个每个人都有自己模型的世界。很多人们想让模型学习的东西,要么是私密的、永远不会出现在 post-training(后训练)数据集里的内容,要么甚至是彼此冲突的,比如说,哦,我希望任务按我的方式来做,而另一家公司或另一个人希望按他们的方式来做。我认为,我们现在已经看到,很多这类东西很难用过去几十年机器学习里一直在使用的同一套工具训练进模型:你有非常干净的 supervision(监督),有 ground truth reward signals(真实奖励信号),再构造一个良好的环境,然后训练模型使用工具去更好地完成这个 coding(编程)任务。
Speaker 116:09 - 16:33
Instead, a lot of the things that actually happen out in the world are very ambiguous or it's hard to say what makes something good. I think a lot of these things are very specific to individuals and I think very misaligned or not very aligned with how the Frontier Labs think about the whole training pipeline and what kind of models will exist in the longer term.
Speaker 116:09 - 16:33
相反,世界上真实发生的很多事情都非常模糊,很难说到底什么才算“好”。我认为,这里面很多东西都高度因人而异,而且我觉得它们与 Frontier Labs 对整个训练 pipeline(流程)的理解,以及对长期会存在什么样模型的设想,存在很强的不匹配,或者说并不怎么一致。
Speaker 316:33 - 17:04
Yeah, and to add to it, I think, what is the P0 for the Frontier Labs? Some of you here are pretty close with them. It's getting to AGI, getting this one generic model that's extremely capable in coding and math, and then using it to automate the economy or to solve really hard, you know, long term problems in cryptography and defense or whatever. And it's pretty clear what needs to happen to push this, you know, more pre training, bigger models, more data, more RL, more inference time compute, that kind of stuff. That's P0.
Speaker 316:33 - 17:04
对,再补充一点,我觉得 Frontier Labs 的 P0 是什么?你们这里有些人与他们走得很近。那就是实现 AGI,做出这一个通用模型,让它在 coding 和 math 上都极其强大,然后用它去自动化经济,或者解决 cryptography、defense 之类那些非常困难的长期问题。不难看出,要推动这件事需要什么:更多 pre-training(预训练)、更大的模型、更多数据、更多 RL(强化学习)、更多 inference-time compute(推理时算力)之类的东西。这就是他们的 P0。
Speaker 317:04 - 17:23
That's where the majority of expenditure and talent goes. And definitely all of them are thinking about memory and all of them are thinking about continual learning. It's just more of a product kind of effort right now. We think it deserves its own attention. We think breakthroughs need to happen there.
Speaker 317:04 - 17:23
这也是绝大多数资金投入和人才流向的地方。而且他们当然都在思考 memory(记忆),也都在思考 continual learning(持续学习)。只是目前这更多还是一种 product(产品)层面的工作。我们认为这值得被单独拿出来重点关注。我们认为那里需要真正的突破。
Speaker 317:23 - 17:49
And Demis and the Sequoia event about a month ago said pretty clearly that we need new breakthroughs around these topics, and obviously they're thinking about them. We're just focusing exclusively on this. And we think certain things around incentives of where the data is and who owns the model are pretty interesting. So if you could learn from many humans or organizations at scale, without necessarily sending someone to work with them shoulder to shoulder, that would be a pretty big unlock.
Speaker 317:23 - 17:49
大约一个月前,Demis 在 Sequoia 的活动上也说得很清楚:我们需要围绕这些话题取得新的突破。显然,他们也在思考这些问题。只是我们是把全部注意力都集中在这上面。我们还认为,在激励机制层面,数据在哪里、模型归谁所有,这些问题都很有意思。所以,如果你能在不必派人去和他们肩并肩一起工作的情况下,仍然从大量的人类或组织那里规模化地学习,那将会是一个相当大的 unlock(关键突破)。
Speaker 117:49 - 18:12
And maybe another point on that is like, I think a lot of things need to look different in the world. So one is there needs to be new research breakthroughs. Two is new infrastructure for training small models for everybody rather than one big model, one big run. Then the third, I think, is a different way of combining research and product. Right now, I think there's researchers in these frontier labs.
Speaker 117:49 - 18:12
也许关于这一点,另一个要说的是,我认为这个世界里很多东西都需要换一种样子。第一,需要新的研究突破。第二,需要新的基础设施,用来为每个人训练 small models(小模型),而不是只训练一个大模型、只做一次大规模 run(训练运行)。第三,我认为还需要一种结合 research(研究)和 product 的不同方式。现在,我觉得这些 frontier labs 里是有研究人员的。
Speaker 118:12 - 18:47
They train the model. They throw it over the fence to the product team who then prompts or context engineers new product surfaces on top of the core models. But in this world where the models are always training, I think the inputs that users provide are very intricately tied to what the models learn from, what the training signal is. There needs to be a lot more of a integrated loop between research and product. While we're focused on tackling a lot of the core research challenges, and that's our background, I think we're also very focused on how to deploy this as quickly as possible to learn from actual feedback in the real world.
Speaker 118:12 - 18:47
他们训练模型,然后把它“扔过墙”交给 product team(产品团队),再由后者通过 prompting(提示词设计)或 context engineering(上下文工程)在核心模型之上做出新的产品界面。但在一个模型始终处于训练中的世界里,我认为用户提供的输入,与模型会从什么中学习、训练信号是什么,是极其紧密地绑在一起的。research 和 product 之间需要一个更加一体化的闭环。虽然我们现在专注于解决很多核心研究挑战,这也是我们的背景所在,但我觉得我们也同样非常关注如何尽快把这些东西部署出去,从现实世界中的真实反馈里学习。
Speaker 118:47 - 19:17
What motivated you to work on this problem? I think it's obviously one of the grand challenges in I think everybody's talking about it these days because the models are so smart, what else is left? I think learning at the edges, learning the remainders of what makes these models useful. It's not just about raw intelligence anymore, it's about learning new things. I think it also feels very fundamental because it goes back to really understanding what makes the model so good.
Speaker 118:47 - 19:17
是什么驱动你去做这个问题的?我觉得这显然是一个重大挑战之一。我想现在几乎每个人都在谈这件事,因为模型已经这么聪明了,那还剩下什么?我认为,关键是在边缘处学习,去学习那些让这些模型变得有用的“剩余部分”。现在已经不只是原始智能高不高的问题了,而是要学会新的东西。我也觉得这件事非常基础、非常根本,因为它最终回到了一个问题:到底是什么让模型如此出色。
Speaker 119:17 - 20:03
Right now the models incidentally know a lot of things from pre training and we don't really understand why. It's like the internet was just this gift granted to us where there's a diverse set of data that contains all of these different examples of coding and writing and all of these other things. It just happened that way. Now to figure out how to crack this problem of continual learning, it's about figuring out what about pre training or even post training makes it possible for the models to generalize in these magical emergent ways and controlling that process so that a company has a set of private data. How do we make the models learn that just as well as the models know the capital of France or how to write Python?
Speaker 119:17 - 20:03
现在,这些模型会因为 pre-training(预训练)而“顺带”知道很多东西,而我们其实并不真正理解原因。这有点像互联网只是被当作一份礼物赐给了我们:那里有一组多样化的数据,包含了各种不同的 coding(编程)示例、写作示例,以及所有这些其他东西。事情就是这样碰巧发生了。现在,要想弄清楚如何攻克 continual learning(持续学习)这个问题,关键在于弄明白:pre-training,甚至 post-training(后训练)中的哪些因素,使模型能够以这种近乎神奇的涌现式方式进行泛化;以及如何控制这个过程,让一家公司拥有一组 private data(私有数据)时,我们也能让模型像知道法国首都是什么、或者知道如何写 Python 那样,把这些数据学得一样好。
Speaker 120:04 - 20:07
So I think it's a really fun problem to think about.
Speaker 120:04 - 20:07
所以我觉得,这是一个非常有意思、很值得思考的问题。
Speaker 220:07 - 20:10
And Don, you came from the neuroscience world, is that right?
Speaker 220:07 - 20:10
Don,你是 neuroscience(神经科学)领域出身的,对吧?
Speaker 320:10 - 20:17
Yes. So I was initially interested in questions around consciousness and the human condition and things like that.
Speaker 320:10 - 20:17
是的。我最初感兴趣的是关于 consciousness(意识)、human condition(人类处境)之类的问题。
Speaker 220:17 - 20:18
Are the models conscious?
Speaker 220:17 - 20:18
这些模型有意识吗?
Speaker 320:18 - 20:48
I don't have any advanced thoughts on this more than you would read. Don't think so, but it's important that smart people are thinking about it. I would say like, I was interested in how humans think, how humans perceive, and as Amos Forsky, the Israeli psychologist used to say, like, he's not interested in artificial intelligence, he's interested in natural stupidity. So, would say like, I started kind of similarly trying to see how people and animals experience the world. Gradually, you know, my inclinations took me to the stats and AI domains.
Speaker 320:18 - 20:48
除了你能读到的那些内容之外,我对此没有什么更深入的想法。我觉得没有,但重要的是,要有聪明的人认真思考这个问题。我想说的是,我当时感兴趣的是 humans(人类)如何思考、如何感知;就像 Israeli psychologist(以色列心理学家)Amos Forsky 过去常说的那样,他对 artificial intelligence(人工智能)不感兴趣,他感兴趣的是 natural stupidity(自然的愚蠢)。所以我会说,我一开始也有点类似,试图去看人和动物是如何体验这个世界的。后来,渐渐地,你知道,我的兴趣倾向把我带到了 stats(统计)和 AI(人工智能)这些领域。
Speaker 320:48 - 21:08
And there I figured that so many of the same problems of memory and continual learning are really, really urgent. And the kind of solutions we have in the current systems are pretty far from what we have in biology. And I'm not one of these people who would say that the machine should be like, you know, like the animal or the human brain. I don't think so. There's many things computers can do better than us.
Speaker 320:48 - 21:08
到了那里,我发现很多相同的 memory(记忆)和 continual learning(持续学习)问题其实都非常、非常紧迫。而且我们目前系统中的这类解决方案,和 biology(生物学)中的做法相比,差距还相当大。我并不是那种会说机器应该像动物、或者像人脑那样工作的人。我不这么认为。计算机有很多事情确实能比我们做得更好。
Speaker 321:08 - 21:42
But human memory has these like very different things in it. It's, you know, if you want to store a whole code base, or you can use computer, you don't even need AI on the computer to store everything losslessly and just get it. But the human brain evolved to work in these constraints of, know, information capacity and to have these fuzzy representations that can then, you know, be abstracted and form connections and form the next day. Current systems don't really have that beyond the generic pre training step. And I was really interested in, you know, what are ways to build that in?
Speaker 321:08 - 21:42
但 human memory(人类记忆)里面有一些非常不同的东西。比如说,如果你想存储整个 code base(代码库),你完全可以用计算机;你甚至都不需要在计算机上用 AI,就能把所有东西无损地存下来并直接取用。但 human brain(人脑)是在这些约束条件下演化出来的,比如 information capacity(信息容量)的限制,以及拥有这些模糊表征;而这些表征随后可以被抽象化、建立连接,并在第二天继续形成新的东西。当前系统其实并没有这种能力,除了那个通用的 pre-training(预训练)步骤之外。我当时真正感兴趣的是:有哪些方法可以把这种东西构建进去?
Speaker 321:42 - 21:43
What are ways to learn from that?
Speaker 321:42 - 21:43
我们可以从中学到什么?
Speaker 421:44 - 22:20
This is more of a philosophical question. You mentioned in the brain, there's a bunch of different real estate, different co processing units, whatever. Modern computer architecture, there's CPUs, GPUs, memory, there's different co processors. With the bitter lesson, do you think that what's happening is that LLMs are converged to, say, one coprocessor that's just totally dominant? It's like everything, all compute is gonna happen in the GPU equivalent of a language model?
Speaker 421:44 - 22:20
这更像是一个哲学问题。你提到在大脑里,有很多不同的“地盘”、不同的协处理单元(co-processing units)之类的东西。现代计算机架构里,有 CPU、GPU、memory(内存),还有各种不同的协处理器。结合 bitter lesson,你觉得正在发生的事是不是:LLM 最终收敛到某一个完全占主导地位的协处理器?就好像一切、所有计算,都会发生在某种相当于 language model 的 GPU 上?
Speaker 422:20 - 22:47
Or do you think that these models are kind of building a bunch of coprocessors emergently inside the model? Take with memory, do you think that the models themselves will just build whatever part of the brain equivalent would be that's good at memory? Or do you think there needs to be another standalone architecture with it?
Speaker 422:20 - 22:47
还是说,你觉得这些模型是在模型内部以涌现(emergently)的方式构建出一堆协处理器?以 memory 为例,你觉得模型本身会自行长出某种相当于大脑中负责记忆、并且擅长记忆的那一部分吗?还是说,它需要配套另一种独立的架构?
Speaker 222:47 - 22:50
Yeah, is memory an emergent property almost versus distinct
Speaker 222:47 - 22:50
对,所以 memory 几乎算是一种涌现属性,而不是一个独立的
Speaker 422:50 - 22:58
process? Yeah, and almost everything. Is everything that we need in intelligence will just be emergent with better training data and more scaled compute?
Speaker 422:50 - 22:58
过程?对,而且几乎一切都是这样。我们在 intelligence(智能)中所需要的一切,是否都会随着更好的训练数据和更大规模的 compute(算力)而自然涌现出来?
Speaker 322:58 - 23:12
Yeah. I would say just on a more superficial perspective on the current deployment of AI, it's way more than just GPUs, and we're seeing all these sandboxes exploding and models operating on other computers trying things. I more mean on the
Speaker 322:58 - 23:12
对。我会说,仅仅从当前 AI 部署的一个比较表层的视角来看,它远不只是 GPU,我们也看到各种 sandbox(沙箱)在快速涌现,模型在其他计算机上运行、尝试各种事情。我更想表达的是在
Speaker 423:12 - 23:15
model architecture level rather than on Yeah, the
Speaker 423:12 - 23:15
model architecture(模型架构)层面,而不是在——对,在这个
Speaker 323:15 - 23:44
so other experiment, either there have been many previous experiments on different architectures that we contributed to, like the state space family and others to try and handle very, very long context more efficiently. The thing with all these methods, it ends up being a trade off, usually a trade off between memory and accuracy and memory not in the behavioral cognitive sense, memory in the computer sense, right? Instead of having the memory footprint of the transformer attention, which is quadratic in the sequence length, these models are-
Speaker 323:15 - 23:44
所以,另一类实验,或者说其实此前已经有很多关于不同架构的实验,我们也参与过其中一些,比如 state space 这一类架构,以及其他一些方法,试图更高效地处理非常非常长的 context(上下文)。这些方法的问题在于,最后都会变成一种权衡,通常是 memory 和 accuracy(准确性)之间的权衡——这里的 memory 不是行为或认知意义上的记忆,而是计算机意义上的内存,对吧?也就是说,不再采用 transformer attention 那种随序列长度呈二次增长的 memory footprint(内存占用),这些模型则是——
Speaker 423:44 - 23:47
Some are claiming they have sub quadratic Yeah,
Speaker 423:44 - 23:47
有些人声称他们做到了 sub-quadratic(次二次复杂度),对,
Speaker 323:48 - 24:11
some are claiming, some do have it, right? And some of the best Chinese model have layers that are inspired by those state space architectures and are not quadratic in cost. Thing is that in our hands, we find that you always compromise accuracy for this memory. There's no free lunch. And what we're saying is like, look, if you're really bitter less and pilled, do you want to do is you want to think, how can I burn more compute?
Speaker 323:48 - 24:11
有些人是在声称,有些人确实做到了,对吧?而且一些最好的 Chinese model 的层受到了那些 state space architectures(状态空间架构)的启发,在成本上不是 quadratic(二次)的。问题是,以我们的实测来看,你总是在这种 memory(记忆)上拿 accuracy(准确率)做交换。天下没有免费的午餐。我们的意思是,听着,如果你真的是那种极度“bitter-less and pilled”的人,你想做的其实是去想:我怎么才能烧掉更多 compute(算力)?
Speaker 324:11 - 24:35
And how can I burn it on new context that I have not seen before? So we're as bitter less and pilled as anyone else. And we are not betting that the overall direction of AGI is gonna end anywhere soon. Just think there's more compute to scale. If I truly want to understand Sean and Sean's work and Sean's context, just rereading files is not going to make it, especially for a special person like you.
Speaker 324:11 - 24:35
以及我怎么把这些 compute 烧在我之前没见过的新 context(上下文)上?所以我们和任何人一样“bitter-less and pilled”。而且我们并不押注 AGI 的整体方向会很快走到尽头。我们只是认为,还有更多 compute 可以继续 scale(扩展)。如果我真的想理解 Sean、Sean 的工作以及 Sean 所处的 context,光靠反复读文件是远远不够的,尤其是像你这样一个特殊的人。
Speaker 324:36 - 24:40
Gotta train a Special derogatory. We gotta train a 100,000,000,000,000
Speaker 324:36 - 24:40
得训练一个特别——带点贬义地说——特殊的东西。我们得为这家伙训练一个 100,000,000,000,000
Speaker 424:40 - 24:41
parameters for this guy.
Speaker 424:40 - 24:41
parameter(参数)规模的模型。
Speaker 224:42 - 24:52
Cosine. Cosine. Yep. What are you finding that people care most about their models learning? Like, is it memorizing facts about the organization?
Speaker 224:42 - 24:52
Cosine。Cosine。对。你们发现人们最在意自己的模型学会什么?比如说,是去记住关于这个组织的事实吗?
Speaker 224:52 - 25:02
Is it remembering like, Ah, no, we do CI this way? What are people actually hoping to And then maybe this feeds into how you do the ranking of memory slots and all that.
Speaker 224:52 - 25:02
还是说,它要记住类似“啊,不,我们的 CI(持续集成)是这么做的”这种东西?人们实际上希望它学会什么?然后也许这就会进一步影响你们怎么给 memory slots(记忆槽位)排序之类的。
Speaker 125:02 - 25:34
Yeah, well, think if you look at what people are spending their time in the app layer doing these days, it's a lot of just trying to make the model work well for your use case. Oh, I want the model to, let's say, design my website with my brand style. That's a very common example these days. But there's many kinds of different tasks that people do with agents, like learning how to run a workflow or your particular way of writing, let's say. So there's many kinds of things.
Speaker 125:02 - 25:34
对,不过我觉得,如果你看现在人们在 app layer(应用层)花时间做的事,很大一部分其实只是想办法让模型在你的 use case(使用场景)里表现得更好。哦,我希望模型,比如说,能按照我的品牌风格来设计网站。这是现在非常常见的一个例子。不过,人们用 agents(智能体)做的任务种类很多,比如学习怎么运行一个 workflow(工作流),或者学习你特定的写作方式,诸如此类。所以它可以学很多不同类型的东西。
Speaker 125:34 - 25:45
And honestly, I think when we think about these methods, kind of going back to this distinction between facts and skills, there really is none. I think the methods are kind of agnostic to that.
Speaker 125:34 - 25:45
说实话,我认为当我们思考这些方法时,回到“事实”和“技能”的区分上,其实并没有真正的区别。我觉得这些方法对此基本是 agnostic(无偏向、无关)的。
Speaker 325:45 - 26:20
Yeah, to me, it's like the natural thing. Almost all the app layers are basically a frontier model wrapped in a loop with search tools and stuff. And what they're all interested in doing with us is finding ways to kind of interface with their data in a way that's, you know, faster, more efficient, and also is more contextual. So almost all of them, it's like, we want to have our, you know, our firm knowledge, you know, be encoded in something that's more efficient that I don't have to research. We want to have the model know in a targeted way, who's the person I should triage a thing to.
Speaker 325:45 - 26:20
对,我觉得这几乎是很自然的方向。几乎所有 app 层,本质上都是一个 frontier model(前沿模型)外面套了一个带 search tools(搜索工具)之类的循环。而他们几乎都希望和我们一起做的,是找到某种方式,以更快、更高效、同时也更具上下文性的方法去和他们的数据对接。所以对他们中的几乎所有人来说,想法都类似于:我们希望把自己的机构知识编码进某种更高效的东西里,这样我就不用每次都去研究;我们希望模型能以一种有针对性的方式知道,这件事应该分流给谁来处理。
Speaker 326:21 - 26:46
And we're just showing them that with pretty lightweight training, these things can be instinctual to the models. They don't have to have these very involved long REPL loops to solve them. So it's, in a sense, it's like, you know, it's a rag killer kind of thing. Again, we can always do a rag and we can always retrieve, but that's the thing that people are interested in interfacing with very large data planes and automating very repetitive things this way.
Speaker 326:21 - 26:46
而我们正在向他们展示的是,通过相当轻量的训练,这些能力可以变成模型的“本能”。它们不需要靠那种非常复杂、很长的 REPL 循环来解决问题。所以从某种意义上说,这有点像是个“RAG killer”的东西。再说一次,我们当然始终可以做 RAG,也始终可以 retrieve(检索),但人们真正感兴趣的是:如何以这种方式去对接超大的 data planes(数据平面),并把那些高度重复的事情自动化。
Speaker 226:46 - 27:07
Yeah, and I want to double click on this RAG killer thing, and I'm sorry to beat a dead horse, I just don't fully grok it yet. Is the premise that there's some trade off between doing RAG versus updating your model weights? Is it the idea that you should be doing both? Like what types of things should be done in the weights versus what types of things should be external ized to RAG?
Speaker 226:46 - 27:07
对,我想再深入追问一下这个“RAG killer”的说法,也抱歉我一直揪着这个问题不放,我就是还没有完全 grok(真正理解)它。它的前提是不是:做 RAG 和更新模型 weights(权重)之间存在某种权衡?还是说,应该两者都做?比如,什么类型的东西应该写进 weights,什么类型的东西又应该 externalize(外置)到 RAG 里?
Speaker 327:07 - 27:29
I think it's an unsolved problem. I don't think anyone has answer to it. We're all working on it. It's also the fundamental question of like biological memory, what should be internalized versus what not. I do think that things that are like, you know, do you need to internalize the room number in a hotel that you were in like a year ago?
Speaker 327:07 - 27:29
我觉得这还是个未解决的问题。我不认为现在有人有答案。我们都还在研究它。这其实也对应着一个关于 biological memory(生物记忆)的根本问题:什么应该被内化,什么不应该。我的确觉得,有些东西比如说——你是否需要把一年前住过的某个酒店房间号内化下来?
Speaker 327:29 - 27:50
Probably no, not in your neural tissue. Probably that's good to write down, but do you need to internalize maybe the password to your home right now? Probably it's useful for the next few years to have that imprinted somewhere. So yeah, how does this translate into like knowledge work and products? This is still something we figure out and we try to take the approach that we try to use as few heuristics as possible.
Speaker 327:29 - 27:50
大概率不需要,至少不需要存在你的 neural tissue(神经组织)里。那种东西大概写下来就好了。但比如你现在家里的密码,也许就需要内化;至少在接下来几年里,把它印刻在某个地方大概是有用的。所以,是的,这该如何映射到 knowledge work(知识工作)和产品上?这仍然是我们正在摸索的事情,而我们的做法是尽量少依赖 heuristics(启发式规则)。
Speaker 327:50 - 28:05
It's easy to run filters on the data and say like, I'm going to keep this, discard that, train on this, train on that. But as humans, you know, we watch TikTok and we, get exposed to a lot of garbage and still the brain is able to learn and not completely go off the rails. We think models should be the same as well.
Speaker 327:50 - 28:05
对数据做过滤很容易,然后说,我要保留这个、丢掉那个、用这个训练、不用那个训练。但作为人类,你知道,我们会刷 TikTok,也会接触到很多垃圾内容,但大脑依然能够学习,而不会彻底失控。我们认为模型也应该如此。
Speaker 128:05 - 28:15
Yeah, maybe concretely in the short term, I think a lot of what people are worried about these days is the huge inference costs of running these agents for days on end.
Speaker 128:05 - 28:15
对,也许从短期更具体地说,我觉得最近很多人担心的一点是:让这些 agent 连续跑上好几天,会带来极其高昂的 inference(推理)成本。
Speaker 228:16 - 28:17
High inference costs a good thing.
Speaker 228:16 - 28:17
高昂的 inference(推理)成本是件好事。
Speaker 128:18 - 28:19
Consuming tokens for Sonia
Speaker 128:18 - 28:19
为 Sonia 消耗 token(令牌)。
Speaker 428:21 - 28:22
works with fireworks.
Speaker 428:21 - 28:22
与 fireworks 配合使用。
Speaker 328:22 - 28:24
She really loves XIA.
Speaker 328:22 - 28:24
她真的很喜欢 XIA。
Speaker 228:24 - 28:25
Love inference.
Speaker 228:24 - 28:25
热爱 inference(推理)。
Speaker 328:27 - 28:28
We love inference too.
Speaker 328:27 - 28:28
我们也热爱 inference(推理)。
Speaker 128:28 - 28:44
Yeah. In the short term, I think that's the immediate pain point. Why are you reading the same files over and over again, even in the same query? But definitely across people in the same company, they're running the same queries on the same documents over and over again. That should be something the model just knows.
Speaker 128:28 - 28:44
对。短期来看,我觉得这就是眼下最直接的痛点。为什么你会一遍又一遍地读取同样的文件,甚至在同一次 query(查询)里也是如此?而且更不用说,同一家公司里的不同人,会一遍又一遍地在相同文档上运行相同的 query(查询)。这本来就应该是模型直接知道的东西。
Speaker 128:44 - 28:51
In the same way you ask an employee, they don't type into the search box, What was I working on yesterday? They just know.
Speaker 128:44 - 28:51
就像你问一名员工时,他们不会在搜索框里输入“我昨天在做什么?”他们就是知道。
Speaker 228:51 - 28:53
But doesn't caching solve that?
Speaker 228:51 - 28:53
但缓存不能解决这个问题吗?
Speaker 128:53 - 29:24
I think to some extent, yeah. But I think going back to this question of what should be internalized versus what's something you retrieve at test time, I think, again, a lot of it is about building on your knowledge. If you are always doing RAG, you can't make associations like, Oh, I see somebody on the team is doing this research. I recall at an abstract level, Oh, there's this related thing that you might want to know about. You didn't even ask about it.
Speaker 128:53 - 29:24
我觉得在某种程度上,可以。但回到这个问题:什么应该被 internalized(内化进模型),什么又应该是在 test time(测试时)再去 retrieve(检索)的东西,我觉得归根结底,很大一部分都和“基于已有知识继续构建”有关。如果你总是在做 RAG,你就无法形成这样的联想:哦,我看到团队里有人在做这项研究;我在一个抽象层面上想起,哦,还有一个相关的东西你可能会想了解。哪怕你根本没有问到它。
Speaker 129:24 - 29:31
But I think these kinds of associations can only happen in weights because they're not really about, you asked me to search for this, I'm going to search for this.
Speaker 129:24 - 29:31
但我认为,这类联想只能发生在 weights(权重)里,因为它并不是那种“你让我搜这个,我就去搜这个”的过程。
Speaker 329:31 - 30:03
And also, I think the main limitation with retrieval systems in general and in AI specifically is the problem is not so much what to store and where to put it. It's the problem is like how to address it, like how to query the thing. Do you know what to look for even? And this involves some sort of intuition that sometimes the models don't have interestingly enough, they don't know where to look. And especially if you're limited to the current way of doing things, which is keyword search, it is just easier to scale in RL and least involved in terms of like infra for embeddings and stuff.
Speaker 329:31 - 30:03
另外,我觉得 retrieval system(检索系统)总体上、尤其在 AI 里,主要限制并不那么在于“存什么、存在哪里”。真正的问题更像是:要怎么寻址它,怎么 query(查询)那个东西。你甚至知道该找什么吗?这需要某种直觉,而有意思的是,模型有时候恰恰缺少这种直觉——它们不知道该去哪里找。尤其如果你受限于当前这套做法,也就是 keyword search(关键词搜索),那它只是更容易在 RL(强化学习)里扩展,而且在 embeddings(嵌入)之类的 infra(基础设施)方面投入也更少。
Speaker 330:03 - 30:43
So yeah, knowing what to search is something that's intuitive and can happen in the way it's and also about caching and inference. Like, of this company started with us taking it like a deep dive into like KV caches and caching. And this is a fascinating thing, right? KV cache is a monstrosity of the current way of doing things that, you know, think about it, a KV cache for a single like Wikipedia article for some, you know, Taylor Swift or something like this, it will be like 80 gigabytes of HBM memory on the GPU and an entire Lama, it's for say a 70B Lama model. And the entire weights of the model would be about 100 gigabytes.
Speaker 330:03 - 30:43
所以,知道该搜什么这件事,本身就是一种直觉,也可能发生在模型内部。还有关于 caching(缓存)和 inference(推理)这点——这家公司一开始其实是让我们深入研究 KV cache 和 caching。这个事情很有意思,对吧?KV cache 可以说是当前这套做法下一个相当怪异的产物。你想想看,比如一篇 Wikipedia 文章,像 Taylor Swift 之类的单篇词条,它对应的 KV cache 在 GPU 上可能就要占到 80GB 的 HBM memory(高带宽内存);而如果拿一个 70B 的 Llama model 来说,整个模型的 weights 大概也就 100GB。
Speaker 330:43 - 31:15
And with some distortion, they remember the entire internet. And how come this thing is so, one thing is so bit efficient. And we have this proof of existence that gradient descent can pack a lot of information in very few numbers. Whereas this KV cache thing, take a few tens of kilobytes of article and it becomes those 80 gigabytes of brain state. So sure, you can cache this, you can load this, you'll have issues with disk to HBM stuff, people are working on it, it's pretty interesting.
Speaker 330:43 - 31:15
而模型的 weights 带着一定失真,却能记住几乎整个互联网。那为什么这个东西会这么……一种方式的 bit efficiency(比特效率)这么高?我们已经有了一个“存在性证明”:gradient descent(梯度下降)可以把大量信息压进很少的数字里。而相比之下,KV cache 这个东西,把几十 KB 的文章,变成了那 80GB 的“脑状态”。所以当然,你可以缓存它,也可以把它加载进来;disk 到 HBM 之间的数据搬运会带来一些问题,人们也正在解决这个问题,这很有意思。
Speaker 331:15 - 31:37
But what if we can take those 80 gigabytes, spend some compute offline, maybe also in Fireworks and file, but then compress it and make it really, really small so that the thing we load in cache is like a thousand X smaller. That would have tremendous implications for how we load things, how fast we can do things and what the fidelity of their representation is.
Speaker 331:15 - 31:37
但如果我们能把那 80GB 拿出来,在线下花一些 compute(算力),也许也会在 Fireworks 里做一些处理,然后把它压缩得非常非常小,让我们加载进 cache(缓存)的东西缩小到原来的千分之一,会怎样?那将对我们如何加载内容、我们能多快完成这些事情,以及它们表征的 fidelity(保真度)产生极其巨大的影响。
Speaker 231:37 - 31:38
Super interesting.
Speaker 231:37 - 31:38
特别有意思。
Speaker 431:38 - 31:49
Yeah. What are some of the things that could happen in the next year or two that would be like the chat GBT moment of memory? Or do you think that that's not how things will play out?
Speaker 431:38 - 31:49
是啊。接下来一两年里,可能会发生哪些事情,会像“memory 的 ChatGPT 时刻”那样?还是说你觉得事情不会那样发展?
Speaker 131:49 - 32:08
It's a good question. I don't know. I think the first proof of concept of the thing that people keep talking about with continual learning, which is you have an intern that you can teach things over time and it actually gets better. I think everybody's waiting to see that. No matter how sophisticated the context engineering approaches are these days, they're not getting there.
Speaker 131:49 - 32:08
这是个好问题。我不知道。我觉得,大家一直在谈的 continual learning(持续学习)如果第一次出现真正的 proof of concept(概念验证),那会很重要:也就是你有一个 intern(实习生)一样的东西,你可以随着时间不断教它,它也确实会越来越好。我觉得所有人都在等着看到这个。无论现在的 context engineering(上下文工程)方法看起来多么复杂,它们都还做不到这一点。
Speaker 132:08 - 32:20
I think you need all of these tools at your disposal to make that happen. But I think it will be something like that, where it's like the model's actually getting smarter. Woah, it's different from yesterday.
Speaker 132:08 - 32:20
我觉得,要实现那件事,你需要把所有这些工具都用起来。但我想,它会是类似这样的时刻:模型真的在变得更聪明。哇,它和昨天不一样了。
Speaker 332:20 - 32:54
Yeah. And it's important to say that the ChatGPT model was not anticipated. We've just read about all the That's different for us products to do, not for- The product directions that certain people had before ChatGPT was different. I feel like to me, the example is like, look, if you, you know, resign from your job today and your sole mission was to make a model that's better for you, and you would use OpenAnthropic and all these frontier models, and you just 20 fourseven engineer the context right skills, your way to move the needle is very limited as an individual. You'll just be better off waiting for the next version of the model and you'll take it from there.
Speaker 332:20 - 32:54
是啊。而且很重要的一点是,要说明 ChatGPT 这个模型其实并不是事先被预料到的。我们之前读到的更多是那些“我们产品能做这个、做那个”的说法,而不是——在 ChatGPT 出现之前,某些人设想的产品方向是不同的。对我来说,一个例子是:如果你今天辞职,然后把唯一目标设定为做一个对你自己更有用的模型,而且你会用 OpenAnthropic 和所有这些 frontier models(前沿模型),然后你 24/7 地把 context(上下文)和相关能力工程化到位,作为个人,你真正能推动的空间其实非常有限。你最好还是等下一个版本的模型出来,再在那个基础上继续做。
Speaker 332:54 - 33:22
And we would like to see a future where actually the more time you spend on the thing actually translates to the quality of performance, at least in the things and domains you care about. And this is pretty hard to achieve. And the only reason we think it could be achieved is if you start scaling compute and training on these data without destroying them all, importantly, which is pretty hard. This is just for fun, rapid fire questions going off
Speaker 332:54 - 33:22
而我们希望看到这样一种未来:你在某件事上投入的时间越多,最后就真的能转化为更高的表现质量,至少在那些你在意的事情和领域里是这样。这其实很难实现。而我们之所以认为它有可能实现,唯一的原因是:如果你开始扩大 compute(算力)规模,并用这些数据去训练,同时又不会把这些数据全都“毁掉”——这一点很关键,而这本身也相当难。这只是随便玩玩的 rapid fire questions(快问快答),话题刚好转到
Speaker 433:22 - 33:24
just memory. When's the last
Speaker 433:22 - 33:24
memory(记忆)。最近一次
Speaker 333:24 - 33:38
time you reached surprised about something in AI in any area? When reading about fundraising. A lot of surprises every day. I would say all of us felt a little bit of a change around the capabilities of the coding agents.
Speaker 333:24 - 33:38
你在 AI 的任何领域里,因为什么事情而感到惊讶,是什么时候?如果说到 fundraising(融资)的新闻,那每天都有很多让人惊讶的事。我会说,我们所有人都多少感觉到,coding agents(编码 agent)的能力出现了一点变化。
Speaker 133:38 - 33:38
That's
Speaker 133:38 - 33:38
那就是
Speaker 333:38 - 34:06
true. But we've been, you know, dabbling with these things and trying to make them work in more effortful ways before. So it didn't come as a complete surprise. But yeah, I think to me, the main events were GetUp Copilot, that for me was just the main event and ChatGPT and then seeing the agentic stuff. We all anticipated, I think, and different people had different expectations on how far it can go and how long horizon it can go.
Speaker 333:38 - 34:06
这倒是真的。不过,你知道,我们之前就一直在浅尝这些东西,也试着用更费力的方式让它们真正发挥作用。所以这并不算是完全出乎意料。但对我来说,主要事件还是 GitHub Copilot,那对我来说就是最重要的事件,然后是 ChatGPT,再然后是看到这些 agentic(具备 agent 特性的)东西。我想我们都预料到了,只是不同的人对它到底能走多远、能延伸到多长的时间跨度,有不同的预期。
Speaker 334:06 - 34:25
But I feel, yeah, we're yet to see something fundamentally different and people are working on completely new ways of doing things now. But yeah, to me, models actually changing in a way that's not harmful and learning new things on the fly that are, you know, personally and economically viable. That's interesting.
Speaker 334:06 - 34:25
但我的感觉是,是的,我们还没有看到某种从根本上不同的东西,而现在人们确实正在研究全新的做事方式。不过对我来说,真正有意思的是:model(模型)以一种无害的方式发生改变,并且能够动态学习新东西,而且这些东西在个人层面和经济层面都是可行的。
Speaker 234:25 - 34:43
Right now, there's this idea of like, we're each gonna have a token wallet that we're going to bring around to companies or to different apps, different workspaces. Do you think that we're gonna end up with, like, a memory bank, a memory wallet that we're gonna move around to across the digital world as we go?
Speaker 234:25 - 34:43
现在有这样一种想法:我们每个人都会有一个 token wallet(token 钱包),然后把它带到不同公司、不同 app、不同工作空间里去。你觉得最终我们会不会也拥有某种 memory bank、memory wallet(记忆钱包),让我们在数字世界中穿梭时也能随身带着走?
Speaker 134:43 - 35:06
I think it's an interesting question. I don't know if we've fully figured out what the right product form factor is in this sense. In a way, even with ChatGPT memory, let's say, I don't want it to remember across my personal and work context. Oh, It's like, you might like these sheets because you trained a model on a GPU last week. It's like, that's totally irrelevant.
Speaker 134:43 - 35:06
我觉得这是个很有意思的问题。我不确定我们是否已经完全想清楚了,在这个意义上正确的产品 form factor(形态)到底是什么。某种程度上说,即便以 ChatGPT memory 为例,我也不希望它在我的个人场景和工作场景之间跨上下文记忆。比如它说:“哦,你可能会喜欢这些床单,因为你上周在 GPU 上训练过一个模型。”这就完全不相关。
Speaker 135:06 - 35:24
To some extent it's because the memory is flawed, but also I think you do want memory in your, I guess, tools and the products that you use to be separated to have control over that. So I personally think there needs to be some separation there, but I guess to be determined what that might look like.
Speaker 135:06 - 35:24
某种程度上,这是因为这种 memory(记忆)本身并不完美;但另一方面,我也觉得,你确实会希望自己使用的工具和产品里的 memory 是彼此分开的,这样你才能对它有控制权。所以我个人认为这里需要某种分隔,不过具体会是什么样子,还有待观察。
Speaker 335:24 - 35:56
Yeah. And I think a holy grail is you go to work and you just burn through all these tokens and you create all this value. And somehow, you know, all the IP and stuff stays with the company, but somehow the skills you learned, the things you invented, your ways of doing things, some of them you can take with you as well to your next job in a way that's, you know, sanitized and not, you know, harmful to any other company's IP. So I do think like carrying a set of skills will be interesting. We do it in our biology right now, and we just, you know, sign NDAs and have like ethical rules around it.
Speaker 335:24 - 35:56
对,我觉得一个 holy grail(终极目标)是:你去上班,然后大量消耗这些 token,创造出很多价值。与此同时,所有的 IP(知识产权)之类的东西都留在公司;但你学到的技能、你发明出来的东西、你的做事方式,其中有一部分也能以某种“净化过的”、不会伤害其他公司 IP 的方式,被你带去下一份工作。所以我确实觉得,能够携带一整套技能会很有意思。我们现在在生物意义上其实就是这么做的,只不过会签 NDA(保密协议),并遵守一些伦理规则。
Speaker 335:56 - 36:09
But I think doing it in a digital world would be pretty interesting and pretty rewarding because it will force each of us to push the frontier and implement AI more deeply in our companies and our individual life and then be rewarded for it. I started a PhD in
Speaker 335:56 - 36:09
但我觉得,如果这件事发生在数字世界里,会非常有意思,也会很有回报,因为这会迫使我们每个人去推动前沿,在公司和个人生活中更深地落实 AI,然后也因此获得回报。我是在读 PhD 时开始接触这个的——
Speaker 436:09 - 36:26
the SAS firm in 2007 Stanford. AI was boring as hell at the time. It was all statistical learning. And there's basically two areas, computer vision and NLP. So vision and language were kind of the two areas, and I think that's still true.
Speaker 436:09 - 36:26
2007 年我在 Stanford 的 SAS firm。那时候 AI 无聊得要命,基本全是 statistical learning(统计学习)。而且大体上只有两个领域:computer vision 和 NLP。所以 vision 和 language 算是两大方向,我觉得现在某种程度上也还是这样。
Speaker 436:27 - 36:48
In 2012, Alex, that happened. Vision was dominating for six years or whatever. Are you guys surprised that language seems to be the language approach seems to be dominating over vision in progress? Question two, do you think vision has any chance of coming back? How do you think about this?
Speaker 436:27 - 36:48
到了 2012 年,Alex,那件事发生了。之后 vision 主导了六年左右之类的时间。你们会不会惊讶于,现在看起来似乎是 language,这种 language 路线,在进展上压过了 vision?第二个问题,你觉得 vision 还有机会重新回来吗?你们怎么看这件事?
Speaker 136:48 - 37:21
Yeah, I think it is pretty surprising to me. I mean, some people maybe saw it coming, but I think I've always kind of been interested in language as, I don't know, I guess a medium for communication and so many complex abstract things can be done in language. I do think, I imagine in the longer term, language and vision will combine in this more unified system where we take in inputs from all of these different modalities and understand them in this abstract way.
Speaker 136:48 - 37:21
是的,我觉得这件事挺让我惊讶的。我的意思是,也许有些人早就预见到了,但我一直都对语言很感兴趣——我不知道,我想它是一种 communication(交流)的媒介,而语言里能承载很多复杂而抽象的东西。我确实认为,从更长期来看,语言和视觉会结合成一个更统一的系统,我们会接收来自各种不同 modality(模态)的输入,并以这种抽象的方式去理解它们。
Speaker 337:23 - 38:02
Yeah. To me, I've never been interested in language. It seemed to me such an advanced capability that, you know, very, the entire animal kingdom has very different forms of speech and language than what we, you know, and how we communicate with ourselves and writing. And I was always, as many other leaders in AI had this thought that, you know, the natural thing is you have to experience the world, act in it, and vision, and action that will be the key. But then I've, you know, like anyone else, who's seen the chat GPT moment and went to do some work at Mosaic and stuff like that to learn how the sausage is made on NLP side.
Speaker 337:23 - 38:02
对我来说,我其实从来没有对语言感兴趣过。它在我看来是一种过于高级的能力——你知道,整个 animal kingdom(动物界)都有各种与我们完全不同的 speech(发声)和 language(语言)形式,也不同于我们彼此交流和书写的方式。而且我一直和许多 AI 领域的领导者一样有一种想法:很自然的路径应该是,你必须去体验这个世界、在其中行动,而 vision(视觉)和 action(行动)才会是关键。但后来我也和其他人一样,经历了 chat GPT 的那个时刻,然后去 Mosaic 那边做了一些工作之类的,想看看 NLP(自然语言处理)这边的门道到底是怎么回事。
Speaker 338:02 - 38:36
And the thing that's striking is that like the language should be pretty hard, like each word has this one hot embedding vector that's as dissimilar to any other word than it is, you know, to, you know, it's a completely high dimensional space, and it's really artificial in a sense. And we learn it with models that are order of magnitude bigger than the best vision models. And still, know, things work pretty well. I do think there's a lot of juice to be squeezed in an image and video. I think you guys doing good investments in this space.
Speaker 338:02 - 38:36
而真正引人注目的是,语言本来似乎应该是很难的——每个词都有一个 one-hot embedding vector(独热嵌入向量),它与任何其他词的差异都几乎一样大;你知道,这是一个完全高维的空间,从某种意义上说也非常人工。我们还用比最好的 vision models(视觉模型)大一个数量级的模型来学习它。即便如此,事情依然运作得相当不错。我确实觉得,image(图像)和 video(视频)里还有很多潜力可挖。我觉得你们在这个领域的投资做得不错。
Speaker 338:37 - 38:40
But I think the two would keep being interesting
Speaker 338:37 - 38:40
但我认为这两者还是会继续以不同的方式保持有趣。
Speaker 438:40 - 38:49
in different ways. I mean, now I'll tell you my That was my lead up. Now I'm gonna tell you the crackpot theory. Oh, wow. And this podcast is not for me to pontificate.
Speaker 438:40 - 38:49
以不同的方式有趣。我的意思是,现在我来告诉你们——前面那些只是铺垫。现在我要讲我的 crackpot theory(离经叛道的理论)了。哦,哇。而且这个 podcast(播客)本来也不是让我在这里高谈阔论的。
Speaker 438:49 - 39:42
It's for you guys, but this is something I've been thinking a lot about, and I just you're the right people to share this with. I I was pretty shocked that language kinda surpassed vision, and I underestimated what was happening with LLMs in like 2018, 2019, 2020, because I just had this bias towards vision. And when I look back on it now, like I think what's basically happening is that in biology, like vision has a massive fundamental advantage over language in biology, and maybe I'm wrong, but basically like the bit rate that your brain can process optical data through the eye is and this is my I'm not a biologist. This is just kinda my dumb assessment. It seems many orders of magnitude greater.
Speaker 438:49 - 39:42
它是给你们的,但这是我一直想了很多的一个东西,我只是觉得你们是最适合分享这个想法的人。我对语言某种程度上超过了视觉这件事感到非常震惊,而且在 2018、2019、2020 年那会儿,我低估了 LLMs(大语言模型)正在发生的事,因为我就是对视觉有这种偏见。现在回头看,我觉得本质上发生的是:在生物学里,视觉相对于语言有一种巨大的、根本性的优势。也许我错了,但大致来说,你的大脑通过眼睛处理 optical data(光学数据)的 bit rate(比特率)——先说明一下,我不是生物学家,这只是我有点笨拙的判断——看起来要高出许多个数量级。
Speaker 439:42 - 40:28
And there's a lot of, like, optical processing that happens, like, even before you reach, you know, like, electrons. And so it's just, like, the total bit rate that is of training data that's kind of being processed and then making it to your brain seems many or semantic greater than the audio data where, you know, it's sound waves, where sound waves are fundamentally, like, much slower bit rate than light. Yeah. And then there's almost an upscaling from the acoustics to electronics, which make it into your brain, whereas there's a downscaling from photons to electrons with vision. Whereas in computers today, everything is electronic.
Speaker 439:42 - 40:28
而且在到达——你知道,就像到达 electrons(电子)之前,已经发生了大量 optical processing(光学处理)。所以,总的 bit rate,也就是那种被处理并最终传到你大脑里的 training data(训练数据)总量,似乎比 audio data(音频数据)大出许多——或者说语义量级上更大;音频毕竟是 sound waves(声波),而声波在本质上就是比光具有低得多的 bit rate。是的。然后,声音这边几乎像是从 acoustics(声学信号)向 electronics(电子信号)做了一次 upscaling(上采样),才进入你的大脑;而视觉这边则是从 photons(光子)到 electrons(电子)的 downscaling(下采样)。而在今天的计算机里,一切都是电子的。
Speaker 440:28 - 41:11
So it's kinda like you nerfed vision and you promoted language where the it's like all processing is on the same playing field. It's all electronic. And I just I think this might this is like my crazy ass dumb non technical crackpot theory, but I think this might be part of why just from an information theory perspective, that maybe language and vision are on a similar playing field by the time you get to LLMs. And then LLMs were just a really, really smart architecture that's better suited for language than for vision. How dumb does this sound, especially to you, Don, the neuroscientist?
Speaker 440:28 - 41:11
所以这有点像是你削弱了 vision(视觉),同时抬高了 language(语言),因为所有处理都处在同一个竞技场上,全部都是电子的。我只是觉得,这也许——这是我这个疯狂、愚蠢、非技术性的 crackpot theory(离经叛道的理论)——但我觉得,从 information theory(信息论)的角度看,这可能部分解释了为什么到了 LLMs 这里,语言和视觉也许站到了一个相近的起跑线上。然后 LLMs 只是一个非常、非常聪明的 architecture(架构),它更适合语言,而不是视觉。这个说法听起来有多蠢,尤其是在你这位 neuroscientist(神经科学家)Don 面前?
Speaker 341:11 - 41:28
Jessie also has some background in cognitive computational science, right? So I would say my point here is like, look, much of what we're doing in knowledge work, we haven't evolved to do, right? We're sitting on these computers reading these things, writing these memos, whatever. We are not evolved to do this. It's new to us.
Speaker 341:11 - 41:28
Jessie 也有一些 cognitive computational science 的背景,对吧?所以我想说的是,你看,我们现在在知识工作里做的很多事情,其实都不是我们进化出来就擅长做的,对吧?我们坐在电脑前读这些东西、写这些备忘录之类的。我们并不是为了做这些而进化的。这对我们来说是新的。
Speaker 341:28 - 41:56
Our brains are not wired for this. Still, nevertheless, it's useful to have LLMs to do this for us. And you know, as humans, we're heavily vision biased, you know, other rodents are more olfactory biased, and I've worked on these things myself before. So what's the real estate in the brain that's allocated to vision and, you know, occipital lobes versus like language areas, temporal lobe, probably more vision, I'll have to check with Chad GPT, I think that's the situation.
Speaker 341:28 - 41:56
我们的大脑并不是为此而“布线”的。尽管如此,让 LLMs(大语言模型)替我们做这些事仍然是有用的。而且你知道,作为人类,我们在很大程度上是 vision biased(视觉偏向)的,其他啮齿类动物则更 olfactory biased(嗅觉偏向),这些东西我自己以前也研究过。所以,大脑里分配给视觉的“地盘”到底有多少,比如 occipital lobes(枕叶)对比 language areas(语言区域)、temporal lobe(颞叶)这些,可能还是视觉更多吧,我得去问问 Chad GPT,我觉得情况大概是这样。
Speaker 441:56 - 41:57
You don't know from memory?
Speaker 441:56 - 41:57
你自己不记得吗?
Speaker 341:57 - 42:10
No, man. I'm externalizing. I'm a big rag believer in my personal lifestyle. But I think that In the limit, we're all it's all rag. I internalize just, you know, important things like my emotions to you.
Speaker 341:57 - 42:10
不记得,老兄。我在做 externalizing(外部化)。在我个人生活方式里,我是个坚定信奉 rag 的人。但我觉得,到了极限情况下,我们所有人——一切都会是 rag。我只把一些重要的东西内化,比如你带给我的情绪。
Speaker 342:10 - 42:36
No, just kidding. Sorry. Anyways, yeah, and vision is dominating when people are training vision language models, end up the language ends up dominating the vision content there. But yeah, it's hard to say that because a certain brain is more biased towards a certain modality doesn't mean necessarily that we're going to more efficiently do it. I do think that efforts on like brain computer interfaces should take this into account.
Speaker 342:10 - 42:36
不,开玩笑的。抱歉。总之,是的,视觉是占主导的;人们在训练 vision language models(视觉语言模型)时,最后往往是语言内容反而压过了视觉内容。不过,很难简单地下结论说,因为某个大脑对某种 modality(模态)偏向更多,就一定意味着我们做这件事会更高效。我确实认为,像 brain computer interfaces(脑机接口)这样的努力,应该把这一点考虑进去。
Speaker 342:36 - 42:45
How do you then relay it back to the brain? That's where I think it's really important to think what real estate do we have there right now. But for knowledge work, it's equally fine if it's text, I think.
Speaker 342:36 - 42:45
那你之后要怎么把它再传回大脑呢?这就是为什么我觉得,去思考我们现在那里面到底还有什么“地盘”可用,非常重要。但对于知识工作来说,如果用 text(文本)来承载,我觉得也完全没问题。
Speaker 242:46 - 42:52
Last question. If everything goes right, what does the world look like in five, ten years? And then what does Ngram's role in it?
Speaker 242:46 - 42:52
最后一个问题。如果一切顺利,五年、十年后的世界会是什么样子?然后 Ngram 在其中会扮演什么角色?
Speaker 142:52 - 43:23
I think I'm imagining a world where everyone has their own model that is really different from the other person's model and from the frontier model, and all of these serve different purposes. To have a model that really I think people often talk about knowing you, but also helping you in the ways that make sense to you personally, whether it's an individual or a team. I think there's an element of having different kinds of intelligence everywhere.
Speaker 142:52 - 43:23
我想象的是这样一个世界:每个人都会有属于自己的 model(模型),而且它会和别人的模型不同,也会和 frontier model(前沿模型)不同;这些模型都会服务于不同的目的。我认为,拥有一个真正——人们常常谈论“了解你”的模型——但更重要的是,它还能以对你个人而言真正有意义的方式帮助你,无论这个对象是个人还是团队。我觉得,这里面有一种让不同类型的 intelligence(智能)无处不在的意味。
Speaker 343:24 - 44:07
Yeah. And to me, actually, it's a variant of the story where in neuroscience, we know that memory and navigation are pretty closely related, same circuits in the brain that, you know, represent landmarks in space are in charge of some, you know, elements of episodic memory and things like this. And for me, I think the company can be, you know, the actual LLM interface to the data plane for everyone. So sharing some similarities to great companies like, you know, Databricks and Oracle, where, you know, we form these memories that happen to be neural memories with models that happen to be personalized and happens to be there's hundreds of millions of them, but they're basically a neural interface to the data plane in a way that's very different from what we know. And it's more efficient, it's more associative.
Speaker 343:24 - 44:07
是的。对我来说,这其实是同一个故事的一个变体:在神经科学里,我们知道,记忆和导航之间的关系相当紧密;大脑中那些表征空间地标的回路,也负责情景记忆等内容中的某些要素。而对我来说,我认为这家公司可以成为面向所有人的、连接 data plane(数据平面)的实际 LLM interface。它和 Databricks、Oracle 这样的伟大公司有一些相似之处:我们形成这些“记忆”,只不过它们恰好是 neural memories(神经记忆);使用的模型恰好是个性化的,而且这样的模型会有数亿个;但从根本上说,它们是通向 data plane 的一种 neural interface(神经接口),其方式与我们已知的完全不同。它更高效,也更具联想性。
Speaker 344:07 - 44:14
It's not representing the file system as it is, it's representing a brain state of that file system. So that's for me a vision.
Speaker 344:07 - 44:14
它并不是按文件系统原本的样子去表征文件系统,而是在表征这个文件系统的一种 brain state(脑状态)。所以对我来说,这就是一个愿景。
Speaker 244:14 - 44:19
Beautiful vision to end on. Thank you guys so much for coming by to share what you're building.
Speaker 244:14 - 44:19
这是个很美的愿景,也很适合作为结尾。非常感谢你们来这里分享你们正在构建的东西。
Speaker 344:19 - 44:20
Awesome. Love it. Thank Thank you, guys.
Speaker 344:19 - 44:20
太棒了。我很喜欢。谢 谢谢你们,各位。
原文 ↗https://www.youtube.com/watch?v=aiR7F4jqjXY
BuildSpeak — 关于本项目BUILT IN PUBLIC · 跟随 builders 而非 influencers