Speaker 100:00 - 00:21
If you accept as the truth that we're gonna be running at the limit, then what that means is that the way to get more intelligence is to be more efficient. We can't get more intelligence by applying more force if we're already at the limit. We have to be more thoughtful about how we use what we have. We build tools. We build external organs that help us solve problems. Speaker 100:00 - 00:21
如果你接受这样一个事实:我们将会在极限状态下运行,那么这意味着,获得更多 intelligence(智能)的方式就是提高效率。如果我们已经到了极限,就不可能靠施加更大力量来获得更多 intelligence。我们必须更认真地思考,如何使用我们已经拥有的东西。我们会构建工具。我们会构建外部器官,来帮助我们解决问题。
Speaker 100:21 - 00:30
You know, we we have an external stomach. We call it kitchen. Now we're creating an external brain. What is the implications of an external brain? Pretty profound. Speaker 100:21 - 00:30
你知道,我们有一个外部胃,我们叫它 kitchen。现在我们正在创造一个外部大脑。外部大脑会带来什么影响?相当深远。
Speaker 100:30 - 00:31
Nobody actually really knows. Speaker 100:30 - 00:31
实际上没有人真的知道。
Speaker 200:32 - 00:48
Hi. I'm Matt Turk. Welcome back to the mad podcast. Open source AI is having yet another moment with powerful new models arriving almost weekly, and my guest today is one of the very best people to unpack it all. Brian Catantaro leads Nemotron, NVIDIA's family of open foundation models. Speaker 200:32 - 00:48
大家好,我是 Matt Turk。欢迎回到 mad podcast。开源 AI 又一次迎来了高光时刻,功能强大的新模型几乎每周都在出现,而今天的嘉宾正是最适合帮我们拆解这一切的人之一。Brian Catantaro 负责 Nemotron,也就是 NVIDIA 的开放 foundation models(基础模型)系列。
Speaker 200:48 - 01:31
Now not everyone realizes NVIDIA has a massive effort to build a frontier AI models, but it employs hundreds of AI researchers and Neutron three Ultra immediately became the number one US open weights model when it was released just a couple weeks ago. We begin this conversation with the state of open source AI and the race between The US and China, and then we go deep inside Nemotron. Four bit training, hybrid member transform architecture, mixture of experts, multi token prediction, and multi teacher distillation all in plain language. Finally, get a rare look at how a modern AI research organization actually runs, how you get many brilliant minds to build one model instead of 100 papers. Please enjoy this awesome conversation with Brian Catantaro. Speaker 200:48 - 01:31
现在,并不是所有人都意识到 NVIDIA 正在大规模投入构建前沿 AI 模型,但它确实雇用了数百名 AI 研究人员,而 Neutron three Ultra 在几周前发布时,立刻成为美国排名第一的 open weights(开放权重)模型。我们的对话将从开源 AI 的现状,以及 The US 和 China 之间的竞赛谈起,然后再深入 Nemotron 内部。four bit training(4 比特训练)、hybrid member transform architecture、mixture of experts(专家混合)、multi token prediction(多 token 预测)以及 multi teacher distillation(多教师蒸馏),都会用通俗语言来讲解。最后,你还将难得一窥一个现代 AI 研究组织究竟是如何运作的:怎样让许多聪明绝顶的人共同构建一个模型,而不是产出 100 篇论文。请享受这场与 Brian Catantaro 的精彩对话。
Speaker 201:33 - 01:56
Alright, Brian, excited to do this. It seems that open source is having a banner year. So you guys at NVIDIA just released Nemotron three Ultra, which is an important moment and the best open source, open weight model in The US. And that was just a few days ago. And then even more recently, GLM 5.2 came out, and that was another moment. Speaker 201:33 - 01:56
好,Brian,很高兴来做这期节目。看起来开源正迎来一个辉煌之年。所以你们 NVIDIA 刚刚发布了 Nemotron three Ultra,这是一个重要时刻,也是 The US 最好的开源、开放权重模型。那还是几天前的事。而就在更近一点的时候,GLM 5.2 也发布了,那又是另一个重要时刻。
Speaker 201:56 - 02:10
So it seems that things are accelerating in open source AI. It feels like a great place to start. What's your assessment, about where we are and how wide the gap between closed source and open source currently is? Speaker 201:56 - 02:10
所以看起来开源 AI 的发展正在加速。这感觉是个很好的切入点。你如何判断我们现在所处的位置?以及 closed source(闭源)和 open source(开源)之间的差距目前到底有多大?
Speaker 102:10 - 02:36
Well, it's really exciting to see all of the energy going into open technologies for AI because we know that, open technologies make it possible for people to innovate. You know, the Internet is such a great example of that. We actually did have closed Internets. I don't know if you remember things like America Online and Prodigy back in the day, And they were great. And open Internet has also, been amazing. Speaker 102:10 - 02:36
嗯,看到这么多力量投入到 AI 的开放技术中,确实非常令人兴奋,因为我们知道,开放技术能让人们进行创新。Internet 就是一个绝佳的例子。实际上,我们确实曾经有过封闭的 Internet。我不知道你是否还记得早年的 America Online 和 Prodigy 之类的东西,它们都很不错。而开放的 Internet 同样也一直非常了不起。
Speaker 102:36 - 03:08
Right? Like, so many different companies have been able to figure out how to transform their work, thanks to, an open technology. The application of the Internet to retail is very different from the application of the Internet to health care or manufacturing, but all of them have been totally transformed, by the Internet. AI, I believe, is also a very transformational technology and also a technology that needs to be applied in very diverse ways. And because of that, I believe that open technologies for AI are really fundamental. Speaker 102:36 - 03:08
对吧?很多不同的公司都已经找到了如何借助一种开放技术来改造自身工作的方式。Internet 在零售业中的应用,与 Internet 在医疗保健或制造业中的应用非常不同,但它们全都被 Internet 彻底改变了。我相信,AI 也是一种具有强大变革性的技术,而且也是一种需要以非常多样化方式加以应用的技术。正因为如此,我认为 AI 的开放技术确实是非常基础、非常根本的。
Speaker 103:09 - 03:23
And it's very exciting to see continued investment and development of open technologies from, for AI from so many different organizations around the world. And, you know, I I hope that that continues. Speaker 103:09 - 03:23
而且,看到世界各地这么多不同的组织持续为 AI 的开放技术进行投资和开发,真的非常令人振奋。并且,你知道,我希望这种势头能够持续下去。
Speaker 203:23 - 03:43
And what's your sense for how far behind open source is compared to closed source? This has been the the big trend over the last few years has been this sort of narrowing gap. Do do you think that open source is almost there or the bar keeps getting raised by the close source models? Speaker 203:23 - 03:43
你觉得 open source 相比 closed source 还落后多少?过去这几年一个很大的趋势,就是这种差距在不断缩小。你认为 open source 已经快追上了吗,还是说 closed source models 的门槛一直在被抬高?
Speaker 103:43 - 04:17
Well, I I feel like this question, it's maybe a tempting question because, you know, it's fun to set up kind of competition, but but I actually feel like the whole AI community is moving very fast. And if you look, for example, at the progress in AI, whether it's closed or open just over the past three months, it's been incredible. And so if you're in a field that's moving really, really fast, I think that's more important than any particular gaps that might exist between different models because the most important thing is, you know, how is AI developing as a field? Speaker 103:43 - 04:17
嗯,我觉得这个问题之所以常被提起,可能是因为它很容易让人想把事情包装成一种竞争,这样也挺有意思;但实际上,我感觉整个 AI 社区都在以非常快的速度前进。比如说,如果你看看 AI 在过去短短三个月里的进展,无论是 closed 还是 open,都非常惊人。所以,如果你身处一个发展速度真的、真的非常快的领域,我认为这比不同模型之间可能存在的任何具体差距都更重要,因为最重要的是,AI 作为一个领域究竟在怎样发展。
Speaker 204:17 - 04:34
What do you think the drivers are to continue to progress in open source AI? Is that the communities and big companies like NVIDIA being behind it? Is that the global competition with China? What propels, open source AI forward? Speaker 204:17 - 04:34
你认为推动 open source AI 持续进步的驱动力是什么?是社区本身,以及像 NVIDIA 这样的大公司在背后支持吗?是和 China 的全球竞争吗?到底是什么在推动 open source AI 向前发展?
Speaker 104:34 - 04:58
You know, I think there's a number of things that that are pushing open technologies for AI forward. One is just the demand. You know, there's so many organizations that want to customize AI and wanna integrate it deeply into their work in a way that really requires open technologies for AI. And so so I think, the demand is certainly there. I think also it's just the best way to develop technology. Speaker 104:34 - 04:58
你知道,我认为有很多因素在推动 AI 的开放技术向前发展。其中一个就是需求本身。你知道,有太多组织希望能够定制 AI,并把它深度整合进自己的工作流程里,而这确实需要 AI 的开放技术。所以我认为,需求肯定是存在的。我还认为,这本身也是发展技术的最佳方式。
Speaker 104:59 - 05:29
And we've seen this, you know, for for many decades that technologies developed in the open move quicker because we can all learn from each other. And, in an era where we're undergoing the most exciting thing to happen in technology in our lifetimes with the development and the deployment of AI, what else do do computer scientists want to work on other than making AI awesome? And if working together as a community is the best way to do that, then that's also a driver that pushes the community towards openly developing technology. Speaker 104:59 - 05:29
而且我们已经看到,几十年来一直如此:以开放方式开发的技术推进得更快,因为我们都能彼此学习。在这样一个时代里,随着 AI 的开发和部署,我们正经历着有生以来技术领域最令人兴奋的事情。那么,computer scientists 除了把 AI 做得非常出色之外,还会想做什么呢?如果以社区协作的方式共同推进,是实现这一点的最佳路径,那么这也会成为一种驱动力,推动整个社区朝着公开开发技术的方向前进。
Speaker 205:29 - 06:09
To ask maybe a slightly cynical question, there is at least a part of the community that's wondering whether open source as an ecosystem, not Nvidia, but in general, been progressing in part based on the ability to distill closed source models and in a world where we seeing the Anthropics and Fable 5s of the world starting to discourage distillation. Do you think there is a chance that open source AI progress may slow down in that context or as a result? Speaker 205:29 - 06:09
我想提一个也许稍微有点犬儒的问题。至少社区里有一部分人在想,open source 作为一个生态系统——不是单指 Nvidia,而是整体而言——其进展在一定程度上是否依赖于对 closed source models 进行 distill(蒸馏)的能力;而在一个像 Anthropics 和 Fable 5 这样的参与者开始不鼓励 distillation(蒸馏)的世界里,你觉得在这种背景下,或者因此,open source AI 的进展有可能放缓吗?
Speaker 106:09 - 06:51
You know, in my mind, there's no question that when the, technology community decides to make huge investments in the most transformational technology of our time, that there's gonna be rapid progress. And also that that technology is not gonna be controlled by a small group of people, because that's just not the way that, the industry works. You know, we we, do our best work. We have the most impact with our work when we're able to, each think about it in our own way and apply it in our own way. So, you know, I love, the closed AI APIs, whether from Anthropic or other people. Speaker 106:09 - 06:51
你知道,在我看来,这毫无疑问:当 technology community 决定对我们这个时代最具变革性的技术进行巨额投资时,就一定会出现快速进展。而且,这项技术也不会被一小群人所控制,因为这根本不是这个行业的运作方式。你知道,只有当我们每个人都能用自己的方式去思考它、用自己的方式去应用它时,我们才能做出最好的工作,才能让我们的工作产生最大的影响。所以,你知道,我很喜欢 closed AI APIs,不管是来自 Anthropic 还是其他人。
Speaker 106:51 - 07:06
I think they're amazing. You know, I'm really, really impressed with the work that those labs are doing. But they're not the only labs in the world. There's lots of labs around the world, and lots of people have a good idea. It's not the case that there's only a few labs that have the monopoly on all good ideas. Speaker 106:51 - 07:06
我觉得它们非常了不起。你知道,我真的、真的对那些实验室正在做的工作印象深刻。但它们并不是世界上仅有的实验室。全球有很多实验室,也有很多人有好想法。并不是说,只有少数几个实验室垄断了所有好的想法。
Speaker 107:06 - 07:26
That's just not true. That's not how humanity operates. There's a there's a lot of bright people on this planet. And, you know, the community, of course, cares deeply about this technology. It's obviously so transformational, has such profound impacts on so many things, that that, of course, many people, wanna be involved in that. Speaker 107:06 - 07:26
这根本不是真的。人类不是这样运作的。这个星球上有很多聪明人。而且,你知道,这个 community 当然非常关心这项技术。它显然具有如此强的变革性,会对这么多事情产生如此深远的影响,所以当然会有很多人想参与其中。
Speaker 107:26 - 07:42
And, so I think over time, we're gonna see that, community oriented approaches to developing and deploying AI are gonna continue to strengthen and be widely adopted because that's really the history of how we build things as a as a human, species. Speaker 107:26 - 07:42
所以,我认为随着时间推移,我们会看到,以 community 为导向的 AI 开发与部署方式会继续增强,并被广泛采用,因为这其实就是我们作为人类这个物种构建事物的历史方式。
Speaker 207:42 - 08:21
Do you think that is globally true as well? So, you know, in particular, with respect to China, this perception that, yes, a lot of people have great ideas around the world. However, a lot of progress from Chinese models were directly inspired or perhaps generated through distillation from the closed source models. Is that just kind of like press rage bait or from the perspective of a leading AI researcher, you're very impressed by the novel ideas that come out of China as well? Speaker 207:42 - 08:21
你认为这一点在全球范围内也成立吗?特别是关于 China,现在有一种看法是:没错,世界各地很多人都有很好的想法;但 Chinese models 的很多进展,都是直接受 closed source models 启发,或者甚至可能是通过 distillation(蒸馏)产生的。这种说法只是媒体为了博眼球、激起争议,还是说,从一位顶尖 AI researcher 的视角来看,你同样对 China 产出的原创想法印象非常深刻?
Speaker 108:23 - 08:52
Perhaps unusually, I, actually did work at a Chinese company for about two and a half years. I worked at Baidu. I worked in the Silicon Valley AI Lab, along with Andrew Ng and as well as Dario Amadei. And, we all worked, for a Chinese company and saw how smart, hardworking, creative, inventive our colleagues were, at the rest of Baidu. And, you know, that experience has has stuck with me. Speaker 108:23 - 08:52
也许有点少见,但我实际上曾在一家 Chinese company 工作了大约两年半。我当时在 Baidu,供职于 Silicon Valley AI Lab,和 Andrew Ng 以及 Dario Amadei 一起工作。我们都在为一家 Chinese company 工作,也看到了 Baidu 其他同事有多么聪明、勤奋、富有创造力和发明精神。你知道,那段经历一直让我印象深刻。
Speaker 108:53 - 09:13
I think it's absolutely false to say that, you know, the achievements of of some other country are all being, created by sort of, you know, copycat mentality. It's just not it's just not true. Now do we all learn from each other, in the technology community? Of course. You know? Speaker 108:53 - 09:13
我认为,说其他某个国家的成就全都是由某种“模仿者心态”造就的,这种说法绝对是错误的。事实根本不是这样。那么,我们在 technology community 里会不会彼此学习?当然会。你知道?
Speaker 109:13 - 09:49
Of course of course, we learn from each other. But, you know, I I would say, you know, it's been a really good thing for the world that the Chinese, AI community has been so, open with what they've been building. I think it's enabled a tremendous number of companies to build things that they couldn't have done without, that community. And I think it's also spurred, technological progress throughout the AI, ecosystem. So, you know, I'm really grateful for, the contributions that our, colleagues in China have made over the years. Speaker 109:13 - 09:49
当然,当然,我们会彼此学习。但你知道,我会说,Chinese AI community 对他们正在构建的东西一直如此开放,这对世界来说是一件非常好的事。我认为,这让大量公司得以构建原本离开那个 community 就做不出来的东西。我也认为,这推动了整个 AI ecosystem 的技术进步。所以,你知道,我真的非常感谢我们在 China 的同行这些年来所作出的贡献。
Speaker 109:49 - 10:30
And, you know, I I would love to encourage a spirit of openness amongst AI labs around the world outside of China as well. You know, I was really excited when, OpenAI released the GPT OSS models, a a while back, and then, of course, Google's been doing great work with Gemma. Absolutely thrilling to see that. And, you know, we're pushing Nemotron along here at NVIDIA as well. So I I think there's a there's a chance for, the rest of the world to catch up to China, in the sense that, you know, we can understand the benefits of working together, as a community to build technologies for AI in a way that I think China has frankly been leading. Speaker 109:49 - 10:30
而且,你知道,我也很希望鼓励一种开放精神,不仅是在 China,也是在全球其他地区的 AI labs 之间。你知道,前一阵子 OpenAI 发布 GPT OSS models 的时候,我真的很兴奋,当然,Google 用 Gemma 做的工作也非常出色。看到这些真的令人非常振奋。而且,你知道,在 NVIDIA,我们也在推进 Nemotron。所以我认为,世界其他地区是有机会赶上 China 的;也就是说,我们可以理解协作的好处,把大家作为一个 community(社区)联合起来,以共同构建 AI 技术,而坦率地说,在这方面 China 一直走在前面。
Speaker 210:30 - 10:42
Great. What is the case for a customer to be using open source models, these days? What is your fundamental advantage? Speaker 210:30 - 10:42
很好。如今,客户为什么要使用 open source models(开源模型)?你们最根本的优势是什么?
Speaker 110:43 - 11:17
Every company is built around a secret. This is a secret that has to do with not just their intellectual property, but also their platform, which has to do with how do they interact with problems and customers. How do they think about solutions, to what what their customers need? And it is always the case that the value of AI is greater when it can be more tightly connected with those secrets because, you know, AI depends on data critically. So the more valuable the data that goes in, the more valuable the solution becomes. Speaker 110:43 - 11:17
每家公司都是围绕某种秘密建立起来的。这种秘密不仅关系到他们的 intellectual property(知识产权),也关系到他们的平台,也就是他们如何与问题和客户互动。他们如何思考解决方案,去满足客户真正的需求?而且,一直以来都是这样:当 AI 能更紧密地连接这些秘密时,它的价值就会更大,因为,你知道,AI 在关键层面上依赖 data(数据)。所以,输入的数据越有价值,最终形成的解决方案也就越有价值。
Speaker 111:17 - 12:04
Now, every company, when it's thinking about how to deploy AI, has to think through what are the implications for the core secrets of our company. And, there's a lot of circumstances where, due to trade secrets or or, you know, trying to think of think through the business model or or even regulatory requirements that, you know, there's data that you really have to treat very carefully by law. And it is much better to do that when you are able to think that through and implement it yourself. Thinking about the integration of AI, the way that AI interacts with customers, the guardrails that are put in place. You know, every company, has a specific understanding of its customers and and therefore what what the customer needs. Speaker 111:17 - 12:04
现在,每家公司在考虑如何部署 AI 时,都必须想清楚:这对我们公司的核心秘密意味着什么。有很多情况下,出于 trade secrets(商业机密),或者,你知道,为了理清 business model(商业模式),甚至是出于 regulatory requirements(监管要求),有些数据依法就必须被非常谨慎地处理。而当你能够自己把这些问题想透并亲自实施时,这件事会好得多。包括思考 AI 的集成方式、AI 如何与客户互动、以及需要设置哪些 guardrails(防护约束)。你知道,每家公司对自己的客户都有特定的理解,因此也更清楚客户真正需要什么。
Speaker 112:04 - 12:25
And, the amazing thing about open technologies for AI is that they allow custom, customization. Right? So companies can think this through. They can build things that that really matter for them. And, you know, I started out this conversation talking about the Internet and about how the Internet the deployment of the Internet has been done in very different ways for very different industries. Speaker 112:04 - 12:25
而 AI 的开放技术最了不起的一点在于,它们允许定制,也就是 customization(定制化)。对吧?所以公司可以把这些问题真正想透,可以构建那些对自己确实重要的东西。而且,你知道,我在这段对话一开始谈到过 Internet,也谈到过 Internet 的部署方式在不同行业中其实非常不同。
Speaker 112:26 - 12:38
And there's a lot of desire to do that, as we see AI change the way that we work and play throughout the entire economy. This is really spurring a lot of demand for open technologies for AI. Speaker 112:26 - 12:38
而且,随着我们看到 AI 正在改变整个经济中我们的工作和娱乐方式,大家也越来越希望能以这种方式来做。这确实正在激发对 AI 开放技术的大量需求。
Speaker 212:39 - 12:57
Great. I'd love to go into a bit of a deep dive into Nematron. But before we do that, maybe a few minutes on your story, your background. What was your past to where you are today, including the the Baidu detour? Speaker 212:39 - 12:57
很好。我很想更深入聊聊 Nematron。不过在那之前,也许先花几分钟讲讲你的经历、你的背景。你是如何一步步走到今天的,包括中间去 Baidu 的那段经历?
Speaker 112:57 - 13:13
So, I started work at, NVIDIA in 2008. At the time, I was a graduate student trying to figure out parallel computing for artificial intelligence, and I thought, NVIDIA had a chance of changing the way computers work with AI. Speaker 112:57 - 13:13
所以,我是在 2008 年开始在 NVIDIA 工作的。那时候,我还是一名 graduate student(研究生),正在尝试搞清楚如何把 parallel computing(并行计算)用于 artificial intelligence(人工智能),而我当时觉得,NVIDIA 有机会改变计算机利用 AI 的方式。
Speaker 213:13 - 13:16
Which was a lonely presumably a lonely quest, right, in 2008? Speaker 213:13 - 13:16
那大概是一段很孤独的探索吧,对吗,尤其是在 2008 年?
Speaker 113:17 - 13:32
Oh, it was it was very chaotic. Back then, people thought I was crazy. And, you know, I remember going to ICML in 2008. I published my first paper training models on the GPU, and people asked me why I was there. People said, this is not a good paper for ICML. Speaker 113:17 - 13:32
哦,那时非常混乱。那会儿大家都觉得我疯了。你知道,我记得自己在 2008 年去 ICML。我发表了第一篇关于在 GPU 上训练模型的论文,结果别人还问我为什么会在那里。有人说,这不是一篇适合发在 ICML 的好论文。
Speaker 113:32 - 13:48
We're we're we just do fancy math here. And I was like, well, but I think computing actually matters a lot for AI. If we could train bigger models that had more capacity to learn, we could probably solve more problems. And, they kinda nodded their heads in nose. So, like, well, why I'm not really sure why you're here. Speaker 113:32 - 13:48
他们的意思是,我们——我们这里只做高深的数学。我当时想的是,可我认为计算实际上对 AI 非常重要。如果我们能训练更大的模型,让它们拥有更强的学习容量,也许就能解决更多问题。而他们大概只是点点头,但其实并不买账。就像在说,好吧,我还是不太明白你为什么会在这儿。
Speaker 213:48 - 13:51
Isn't a GPU a thing for gaming as well presumably? Speaker 213:48 - 13:51
GPU 不也是拿来打游戏的东西吗,大概?
Speaker 113:51 - 13:57
Right. Yeah. There's there's also that. Right? Which we we continue to, run into that, that idea. Speaker 113:51 - 13:57
对,没错。确实也有这一层意思,对吧?而且我们一直到现在还会碰到这种看法。
Speaker 113:57 - 14:15
Actually, a GPU is whatever NVIDIA says it is. You know, we make them. So a GPU is is a is a thing that we make in order to accelerate the world's most important computations, which in 1995 was graphics. And, you know, for a long time now, it's been AI. So, anyway, I started at NVIDIA. Speaker 113:57 - 14:15
其实,GPU 是 NVIDIA 说它是什么,它就是什么。毕竟是我们制造它们的。所以,GPU 是我们为了加速这个世界上最重要的计算任务而打造的东西——在 1995 年,这个任务是图形处理。而到后来很长一段时间里,这个任务一直都是 AI。总之,我就是这样加入 NVIDIA 的。
Speaker 114:15 - 15:07
I was in the research group, doing strange things about trying to make, compilers, libraries, for AI on the GPU. That led to, the creation of, first Copperhead, which was a it was a a Python, embedded language that, compiled to the GPU, which I think foreshadowed a lot of things in TensorFlow and and PyTorch. And then then that led to the creation of cooDNN, was NVIDIA's first product for for deep learning on the GPU. And I I really enjoyed working on that, but I was always wanting to see more, firsthand about the applications of AI. And at NVIDIA, I was mostly working on, you know, libraries and compilers for AI. Speaker 114:15 - 15:07
我当时在 research group,做一些很特别的事情,比如尝试为 GPU 上的 AI 开发 compiler(编译器)和 library(库)。这后来促成了 Copperhead 的诞生——它是一种嵌入式 Python 语言,可以编译到 GPU 上运行。我觉得它在很多方面都预示了后来 TensorFlow 和 PyTorch 里的不少东西。再后来,这又推动了 cuDNN 的诞生,它是 NVIDIA 第一个面向 GPU 上 deep learning(深度学习)的产品。我非常喜欢做这项工作,但我一直想更直接地看到 AI 的应用层面。在 NVIDIA,我大多还是在做 AI 的 library 和 compiler。
Speaker 115:07 - 15:33
So I thought, well, you know, when Andrew Ng asked me to go, build the Silicon Valley AI Lab with him, at Baidu, I thought, oh, this is a great opportunity because even back then, Baidu, was very advanced in its application of AI to its core business. And so so that was a a fantastic opportunity for me. The Baidu Silicon Valley AI Lab was an amazing place, full of brilliant people that were working really hard. Speaker 115:07 - 15:33
所以我想,嗯,当 Andrew Ng 邀请我和他一起在 Baidu 建立 Silicon Valley AI Lab 时,我觉得,哦,这是一个绝佳的机会,因为即使在那时候,Baidu 在把 AI 应用到其核心业务这件事上就已经非常领先了。所以那对我来说确实是一个非常棒的机会。Baidu Silicon Valley AI Lab 是个了不起的地方,里面都是才华横溢、而且工作极其投入的人。
Speaker 215:33 - 15:42
What was it like working with a a young Dario? Was there, like, any signs that he could become, you know, who he has become? Speaker 215:33 - 15:42
和年轻时的 Dario 一起工作是什么感觉?当时有没有什么迹象表明,他会成为如今这样的他?
Speaker 115:42 - 16:05
Dario, was brilliant from the beginning. I remember, I interviewed him. I was on the panel, and, at the time he, had been working in bioinformatics. So he he hadn't been working on deep learning or or the things that we call AI these days. But it was very clear that he learned extremely quickly and also that he thought extremely deeply. Speaker 115:42 - 16:05
Dario 从一开始就非常出色。我记得我面试过他。我当时是面试小组成员,而那时候他一直在做 bioinformatics。所以他当时并没有在做 deep learning,或者我们如今所说的 AI 这些东西。但很明显,他学东西特别快,而且思考也非常深入。
Speaker 116:05 - 16:37
I think, you know, the thing I admire most about Dario, is the strength of his conviction. You know, I've been working in this field, for a long time, and I've believed also that AI is gonna transform the world. But I don't think that I believed in it as completely as Dario did. And perhaps that was because, you know, my academic training during my PhD was full of a lot of caution. I don't know if you remember, but AI was old and bad in 2005 Speaker 116:05 - 16:37
我想,我最佩服 Dario 的一点,是他信念的力量。你知道,我在这个领域工作很久了,我也一直相信 AI 会改变世界。但我觉得,我并没有像 Dario 那样如此彻底地相信这件事。也许那是因为,我 PhD 期间接受的学术训练里充满了很多谨慎。我不知道你还记不记得,但在 2005 年,AI 还是一种又老又不被看好的东西。
Speaker 216:38 - 16:39
It will never work. Speaker 216:38 - 16:39
它永远都行不通。
Speaker 116:40 - 16:52
That people did with computers. They started doing it in 1945. Right? And and so there had been so many grandiose promises that failed to deliver over the years. And so I came to AI with a lot of caution. Speaker 116:40 - 16:52
那是人们拿计算机做的一件事。他们从 1945 年就开始做这个了,对吧?所以,这些年里已经有过太多夸夸其谈的宏大承诺,最后都没能兑现。所以我进入 AI 这个领域时,是带着很强的谨慎态度的。
Speaker 116:53 - 17:11
In fact, back then we used to call it machine learning, which was basically a dodge. Like, we just didn't want people to know that we were we were working on AI because then they would be like, oh, we've heard about that. It never works. Right? So, I came to AI with a little bit of this, like, academic caution like, oh, we should hedge a little bit. Speaker 116:53 - 17:11
其实,在那个时候我们通常把它叫作 machine learning,这基本上是一种回避说法。就像是,我们只是不想让别人知道我们在做 AI,因为那样他们就会说,哦,这个我们听说过,从来都不管用。对吧?所以,我进入 AI 时,心里多少带着一点这种学术上的谨慎,觉得,哦,我们应该稍微保守一点。
Speaker 117:11 - 17:41
Like, I don't know if now's the time. And Dario, his strength of conviction and his understanding of the moment of how, the technology was developing this time it was actually going to work. And then the implications of that on, you know, how the technology should be developed, what kind of institutions, to build. I think he's done a a spectacular job. And so, yeah, working with him, it was always always a fun experience. Speaker 117:11 - 17:41
就像,我也不知道现在是不是那个时机。而 Dario,他那种坚定的信念,以及他对当下这个时刻的理解——理解这一次技术是如何发展的,理解它这次真的会成功。然后还有这件事所带来的影响,比如技术应该如何被开发,应该建立什么样的 institutions(机构)。我认为他做得极其出色。所以,是的,和他一起工作始终都是一种非常愉快的经历。
Speaker 217:41 - 17:45
So then you went back to NVIDIA and walk us through the journey? Speaker 217:41 - 17:45
所以后来你又回到了 NVIDIA,给我们讲讲这段历程吧?
Speaker 117:45 - 18:04
Yeah. So, ten years ago, actually, in 2016, Jensen called me up and said, hey, would you like to come back and build an applied research lab? And I thought that would be a fantastic, opportunity. I you know, I've always loved NVIDIA. I've loved the way the company works, the convictions the company holds. Speaker 117:45 - 18:04
是的。所以,十年前,准确说是在 2016 年,Jensen 给我打电话,说,嘿,你愿意回来创建一个 applied research lab(应用研究实验室)吗?我觉得那会是一个非常棒的机会。你知道,我一直都很喜欢 NVIDIA。我喜欢这家公司的运作方式,也喜欢它所坚持的那些信念。
Speaker 118:04 - 18:25
You know, NVIDIA is a very unique company. It follows through over long time periods, you know, and I've seen that with CUDA. I've seen it with our deep learning technologies. I've seen it with our ray tracing graphics technologies, our AI for graphics. You know, over and over again, NVIDIA is not afraid to put in five or ten years worth of research in order to change the world. Speaker 118:04 - 18:25
你知道,NVIDIA 是一家非常独特的公司。它会在很长的时间跨度上真正把事情做到底,我在 CUDA 上看到了这一点,也在我们的 deep learning(深度学习)技术上看到了这一点,还在我们的 ray tracing graphics(光线追踪图形)技术、我们的 AI for graphics 上看到了这一点。一次又一次,NVIDIA 都不惧怕投入五年或十年的研究,只为改变世界。
Speaker 118:25 - 18:55
You know? And, working at a company that has that strength of conviction and the ability to follow through is kind of an ideal thing for me. I just really I just really love the support that the company gives, gives its researchers, to invent the future. And so, I thought I'd come back. The first project that, that I worked on, actually became DLSS, which, some of your, audience may know about, but DLSS is our real time, AI for graphics. Speaker 118:25 - 18:55
你知道吗?而且,对我来说,能在一家拥有这种信念强度、也有能力长期贯彻到底的公司工作,几乎是理想状态。我真的非常喜欢公司给予研究人员的支持,让他们去发明未来。所以我就决定回来。实际上,我参与的第一个项目后来变成了 DLSS,你们的一些观众可能知道它;DLSS 就是我们面向图形的实时 AI。
Speaker 118:55 - 19:35
And it makes a small GPU run like a big GPU. It's about 10 times more efficient because rather than computing, the color of every pixel for every frame, we use AI to infer the color. And, you know, these days, 23 out of every 24 pixels, is being generated by our AI model when you're using DLSS, to play games, and gamers love it. It's become the standard way of playing games because it's just so much more responsive and it's more beautiful. Our AI, we train it offline on huge datasets, and it's able to render graphics in real time, more beautifully than, traditional methods do. Speaker 118:55 - 19:35
它能让小 GPU 跑出大 GPU 的效果。它的效率大约高 10 倍,因为我们不是为每一帧计算每一个像素的颜色,而是用 AI 去推断颜色。你知道,如今在你使用 DLSS 玩游戏时,每 24 个像素里有 23 个都是由我们的 AI model(模型)生成的,而玩家们非常喜欢它。它已经成为玩游戏的标准方式,因为它的响应更快,画面也更漂亮。我们的 AI 是在离线状态下用海量数据集训练出来的,因此它能够以实时方式渲染图形,而且效果比传统方法更漂亮。
Speaker 119:36 - 20:22
We recently actually announced DLSS five which is a fully generative version of DLSS, and, I, am so excited about it. It represents culmination of ten years worth of research on how to make, real time graphics much more beautiful. And so, so that's part of, part of the journey here for me was real time AI for graphics. But then at the same time, we also started a language modeling project, and this was back in 2017, you know, before transformers, were big and before language modeling, started taking over the world. But, you know, I just had this intuition maybe built on, you know, some of the things that that I had seen, while working, at Baidu. Speaker 119:36 - 20:22
我们最近实际上公布了 DLSS five,这是 DLSS 的一个 fully generative(完全生成式)版本,我对此非常兴奋。它代表了十年研究的集大成,核心是在于如何让实时图形变得更加漂亮。所以,对我来说,这段历程的一部分就是面向图形的实时 AI。但与此同时,我们也启动了一个 language modeling(语言建模)项目,那是在 2017 年,还早于 transformers 真正爆发、也早于 language modeling 开始席卷世界的时候。不过,你知道,我当时就是有一种直觉,可能是建立在我在 Baidu 工作时看到的一些事情之上。
Speaker 120:22 - 20:46
I I just had this intuition that, you know, working with text and understanding text was gonna lead to better reasoning, which was gonna lead to better application of AI in all sorts of domains. And so, so we started this project, called Megatron. Megatron stands for the biggest baddest transformer. That's why we named it that. And it was really a systems project to show the world how to train the largest transformer models on NVIDIA's hardware. Speaker 120:22 - 20:46
我当时就是有一种直觉:处理文本、理解文本,将会带来更好的 reasoning(推理)能力,而这又会推动 AI 在各种领域中得到更好的应用。所以我们启动了这个叫作 Megatron 的项目。Megatron 的意思是“最大、最强的 transformer”,所以我们才给它起了这个名字。而它本质上是一个 systems project(系统项目),目的是向世界展示,如何在 NVIDIA 的硬件上训练最大规模的 transformer models(模型)。
Speaker 120:46 - 21:28
Back at the time, some of your, audience may or may not remember this, but there was there were being claims made that the only way to train big transformer models was on the TPU, because after all, the transformer had been invented at Google. And, so, you know, we looked we looked at, you know, we loved the the transformer paper. We thought, wow, this has amazing potential. We tried it out on our own language modeling tasks and it worked so much better than the RNNs that we had been using before. And also we saw immediately that there was an enormous systems opportunity to co optimize the GPU, the networking, all of the compilers and software that would enable people to scale, transformer based language models, really dramatically. Speaker 120:46 - 21:28
在当时,你们的一些观众可能记得,也可能不记得,那时候有人声称,训练大型 transformer models 的唯一方式就是用 TPU,毕竟 transformer 本来就是在 Google 发明的。所以,我们当时看了那篇 transformer 论文,非常喜欢,觉得,哇,这东西潜力惊人。我们把它用在自己的 language modeling 任务上试了一下,结果比我们之前一直在用的 RNNs 好得多。同时,我们也立刻看到,这里面存在巨大的系统机会,可以对 GPU、网络、以及所有编译器和软件进行协同优化,从而让人们能够把基于 transformer 的 language models 规模大幅扩展起来。
Speaker 121:28 - 21:54
And we we thought, you know, this is this is something that that, could really have an impact. So we started the Megatron project, which then led to, I think, basically helping the whole industry figure out how to train, extremely large, LLMs, and also led to the foundations of today's Nemotron project where, you know, NVIDIA trains, its own LLMs, for its own purposes. So that's kind of the the history. Speaker 121:28 - 21:54
我们当时觉得,这确实可能产生真正的影响。于是我们启动了 Megatron 项目,而我认为,这后来基本上帮助整个行业摸索出了如何训练超大规模 LLMs,也奠定了今天 Nemotron 项目的基础——也就是 NVIDIA 为了自身用途训练自己的 LLMs。所以这大概就是这段历史。
Speaker 221:54 - 22:19
Great journey. Okay. So let's go into all things, Neemotron. And before we get into the specifics, that's the obvious question that I'm sure you've been asked many times, which is why does NVIDIA care in the first place to be building model and investing very significant efforts into creating its own family of frontier models? Speaker 221:54 - 22:19
很精彩的经历。好的。那么我们来深入聊聊 Neemotron 的一切。在进入具体细节之前,先问一个显而易见、我相信你已经被问过很多次的问题:为什么 NVIDIA 一开始会在意亲自构建 model(模型),并投入非常大的精力去打造属于自己的 frontier models(前沿模型)家族?
Speaker 122:19 - 22:57
You know, Neutron has two jobs. The first job is to help us understand how to build the systems of the future. NVIDIA is an accelerated computing company, and that means thinking through the world's most important computational challenges from first principles and designing systems, which includes a lot of software, in order to make it possible for people to invent and deploy things that never could have been done with standard computing. But in order to do that, NVIDIA has to deeply understand everything about how AI works. That's how we codesign all of the systems and software, for our main product line. Speaker 122:19 - 22:57
你知道,Neutron 有两项任务。第一项任务,是帮助我们理解如何构建未来的系统。NVIDIA 是一家 accelerated computing(加速计算)公司,这意味着我们要从第一性原理出发,思考世界上最重要的计算挑战,并设计系统——其中也包括大量软件——从而让人们能够发明并部署那些用标准计算根本无法完成的东西。但要做到这一点,NVIDIA 就必须非常深入地理解 AI 运作的一切原理。这也是我们如何为自己的主力产品线协同设计全部系统和软件的方式。
Speaker 122:57 - 23:24
So the first job of Neutron is to make sure that NVIDIA continues to exist so that we can continue delivering meaningful acceleration in an era where Moore's law has died. And the the acceleration that we get these days comes through specialization. But, again, specialization comes through understanding. So that's Neutron's first job is to help NVIDIA understand how to build its core products. NVIDIA's second or Neutron's second job is to support the ecosystem. Speaker 122:57 - 23:24
所以,Neutron 的第一项任务,就是确保 NVIDIA 能持续存在下去,这样我们才能在一个 Moore's law 已经终结的时代,继续提供有意义的加速。如今我们获得的加速,来自 specialization(专用化、专业化)。但归根结底,specialization 来自理解。因此,Neutron 的第一项任务,就是帮助 NVIDIA 理解如何构建它的核心产品。NVIDIA 的第二项任务,或者说 Neutron 的第二项任务,则是支持整个 ecosystem(生态系统)。
Speaker 123:25 - 24:14
One of the most valuable things that NVIDIA has built over the years is, all of the people around the world who build and deploy amazing AI, using NVIDIA's technologies. And, we think that it's necessary for open technology for AI to continue to exist, from NVIDIA to help, support that. Neutron's not trying to be the only open technology for AI. We love all technology for AI for the very straightforward reason that whenever AI, is further developed and further deployed, it's an opportunity for our business. So so this is this is, you know, we're we're very explicitly trying to develop our ecosystem because that's good business for us, But we're not trying to be the only provider of technologies for this ecosystem. Speaker 123:25 - 24:14
多年来,NVIDIA 打造出的最有价值的东西之一,就是遍布全球、使用 NVIDIA 技术来构建并部署惊人 AI 的那些人。我们认为,要让 AI 的 open technology(开放技术)继续存在,NVIDIA 有必要出手支持。Neutron 并不想成为 AI 唯一的 open technology。我们热爱所有 AI 技术,原因其实很直接:只要 AI 得到进一步发展并进一步部署,对我们的业务来说就是机会。所以,这一点我们说得非常明确——我们确实是在积极发展自己的 ecosystem,因为这对我们是门好生意;但我们并不想成为这个 ecosystem 中技术的唯一提供者。
Speaker 124:14 - 24:28
We love seeing, other companies contribute as well. The the most important thing for Neutron's second job is just making sure that it continues to be possible for companies of all shapes and sizes to build and deploy their own AI. Speaker 124:14 - 24:28
我们也很高兴看到其他公司作出贡献。对 Neutron 的第二项任务来说,最重要的就是确保各种规模、各种类型的公司,都仍然能够构建并部署属于自己的 AI。
Speaker 224:28 - 24:31
By the way, Moore's Law is dead. Is that is that is that official? Speaker 224:28 - 24:31
顺便说一句,Moore's Law 已经死了。这算是官方定论了吗?
Speaker 124:31 - 24:33
It's been dead for years. Speaker 124:31 - 24:33
它已经死了很多年了。
Speaker 224:33 - 24:35
It's been dead for years? Like, why why is that? Speaker 224:33 - 24:35
已经死了很多年?为什么会这样?
Speaker 124:35 - 24:57
Well, you just look at the the progress, in semiconductor manufacturing. You know, the the original statement of Moore's law was economic. Right? It was about we can afford to put twice as many transistors on the same chip in every, whatever, twenty four months, whatever the the time period is. And, these days, that is absolutely not the case, and it hasn't been for probably five or ten years. Speaker 124:35 - 24:57
嗯,你只要看看半导体制造领域的进展就知道了。你知道,Moore's law 最初的表述其实是经济层面的,对吧?它说的是:每隔——不管是 24 个月还是多久——我们都能负担得起在同一块 chip 上放入两倍数量的 transistor(晶体管)。而如今,情况绝对不是这样,而且大概五年到十年前就已经不是这样了。
Speaker 124:57 - 25:14
Right? Now we are still scaling our systems, right, through a number of ways. One is just applying a lot more silicon to it. Right? We are also getting transistors are continuing to get smaller and and more efficient, although at a slower pace, but they're also getting quite a bit more expensive at the same time. Speaker 124:57 - 25:14
对吧?不过我们仍然在通过多种方式扩展系统,对吧?一种方式就是直接给它投入更多 silicon(硅)。对吧?另外,transistor 也仍在继续变得更小、更高效,虽然速度变慢了,但与此同时它们也确实变得贵了不少。
Speaker 125:17 - 26:07
So the, you know, in an era where where Moore's law was alive, the best way to make the system of the future was to take the system of the present and then just shrink it and and maybe double it at the same time. Right? But in an era where where we've been living for a while now where you don't get economic benefits from taking your existing design and shrinking it, you really have to be more clever about how you use every part of the system. That that's an era where accelerated computing is is much more valuable than ever because the the work of thinking through the prop problem from first principles and codesigning absolutely everything from transistors to algorithms and applications, in order to reduce waste and and deliver meaningful acceleration, that's more valuable than ever. Speaker 125:17 - 26:07
所以,你知道,在 Moore's law 还有效的时代,打造未来系统的最佳方式,就是拿当下的系统,把它缩小,然后也许同时再翻一倍,对吧?但在我们已经生活了一段时间的这个时代里,把现有设计缩小并不会带来经济收益,所以你真的必须更聪明地使用系统中的每一个部分。在这样的时代,accelerated computing(加速计算)比以往任何时候都更有价值,因为从第一性原理出发重新思考这个问题,并对从 transistor 到 algorithm(算法)再到 application(应用)的所有环节进行 codesign(协同设计),以减少浪费并实现有意义的加速,这件事比以往任何时候都更有价值。
Speaker 226:07 - 26:32
Fantastic. To playback, what you were saying a minute earlier, It makes good business sense for NVIDIA to be in the model business because one, it helps, design better chips and two, whatever is good for AI is ultimately good for NVIDIA, which makes a lot of sense. That Neutron effort is reasonably recent. Right? It started in 2023, I believe, maybe. Speaker 226:07 - 26:32
太好了。回过头来复述一下你刚才说的意思:NVIDIA 进入 model 业务在商业上是说得通的,因为第一,它有助于设计更好的 chip;第二,任何对 AI 有利的事,最终也都会对 NVIDIA 有利,这很有道理。那个 Neutron 项目是相当近期的,对吧?我记得它是从 2023 年开始的,也许吧。
Speaker 226:32 - 26:41
Walk us quickly through the key releases. I believe in 2023, there was Neutron three eight b as a key release, or am I missing a step? Speaker 226:32 - 26:41
你快速带我们过一下几个关键发布吧。我记得 2023 年的一个关键发布是 Neutron three eight b,还是说我漏掉了中间某一步?
Speaker 126:41 - 26:55
Yes. Yes. Yes. So, you know, the the, you know, the the numbering is somewhat lost to time. It I almost feel like we're in the Lord of the Rings, and it's like, you know, there's, like, some ancient, like, relics that we're digging up out of an old mine. Speaker 126:41 - 26:55
对,对,对。所以,你知道,这个编号多少已经湮没在时间里了。我几乎感觉我们像是在 Lord of the Rings 里,仿佛在某座旧矿井里挖掘出一些古老的 relic(遗物)。
Speaker 126:56 - 27:15
You know, this is a long time ago. You know, the original what what, we originally called Nemotron One was actually a project that we did with Microsoft. We jointly trained a 530,000,000,000 parameter model. I believe that was released in 2021. And so this is GPT three era. Speaker 126:56 - 27:15
你知道,那已经是很久以前了。最初那个——我们最早称为 Nemotron One 的东西——其实是我们和 Microsoft 一起做的一个项目。我们联合训练了一个 530,000,000,000 参数的 model。我想那是在 2021 年发布的。所以那还是 GPT three 时代。
Speaker 127:15 - 27:35
And that's what at the time we called it Megatron Turing NLG. Turing was the what Microsoft was calling their their language model efforts, at the time. But, that, in retrospect, we called Nemotron one. Then along the way, we built a few more. We got up to Nemotron three. Speaker 127:15 - 27:35
而那在当时我们把它叫作 Megatron Turing NLG。Turing 是 Microsoft 当时对他们 language model 工作所使用的名称。不过,事后回看,我们把那算作 Nemotron one。后来一路上,我们又做了几个,最后到了 Nemotron three。
Speaker 127:36 - 28:05
And then, Lama came along, and we were really excited about that. We were very, happy that Meta was supporting the open AI technology space. And so we we started, you know, taking our language model technology and adding it to LAMA models, which then resulted in LAMA Nematron one. And, you know, that was the first, reasoning model, built on LAMA. We were really proud of that. Speaker 127:36 - 28:05
然后,Lama 出现了,我们对此真的非常兴奋。我们非常高兴 Meta 在支持 open AI 技术领域。所以我们开始把自己的 language model(语言模型)技术加入到 LAMA models 中,随后就产生了 LAMA Nematron one。你知道,那是第一个基于 LAMA 构建的 reasoning model(推理模型)。我们对此非常自豪。
Speaker 228:05 - 28:07
And that was 2025? Speaker 228:05 - 28:07
那是 2025 年吗?
Speaker 128:08 - 28:27
Might have been '24. I believe, I can't remember. Somewhere around there. And then, yes, and then we we continued, to to develop that. And, you know, last we we so we the numbers kinda started over again. Speaker 128:08 - 28:27
可能是 2024 年。我想是,但我记不清了。大概就是那个时间。然后,是的,我们继续开发它。然后,你知道,后来我们的编号某种程度上又重新开始了。
Speaker 128:27 - 28:52
We released a a Nemotron two. I believe it was last year. And then we quickly followed that up with Neutron three because we we needed to put MOE support in. Neutron two didn't have MOE support and that made it uncompetitive against, other models like GPT OSS 20 b was just, like, so fast because of MoE. And so we were like, okay. Speaker 128:27 - 28:52
我们发布了一个 Nematron two。我记得应该是去年。接着我们很快又推出了 Neutron three,因为我们需要加入 MOE support。Neutron two 没有 MOE support,这让它相比其他模型缺乏竞争力,比如 GPT OSS 20 b,靠着 MoE 简直快得惊人。所以我们当时想,好吧。
Speaker 128:52 - 29:14
We've got we've got to put put the MoE in, so that became Nematron three. Now we're in a a slightly difficult state because we're working on Nematron four. Right? But we already released the Nematron four, which was, in 2024, we released a three forty b, model called Nematron four. And so I'm not exactly sure how we're gonna, solve this marketing problem. Speaker 128:52 - 29:14
我们必须把 MoE 加进去,于是就有了 Nematron three。现在我们的处境有点尴尬,因为我们正在做 Nematron four,对吧?但我们其实已经发布过 Nematron four 了——那是在 2024 年,我们发布了一个名为 Nematron four 的 three forty b 模型。所以我也不太确定我们要怎么解决这个 marketing 问题。
Speaker 129:14 - 29:41
I didn't create this marketing problem. So, I'll I'll I'll do my best to to make it clear that Neutron four of of of whatever the neck whenever we release that is different from the the 2024 Neutron four. But in any case, we've been working on this for a long time. I I think more important to us than any particular generation is just the sustained commitment that NVIDIA has to developing these models. We've been doing it for a while. Speaker 129:14 - 29:41
这个 marketing 问题不是我造成的。所以,我会尽力说明清楚:我们之后无论什么时候发布的那个 Neutron four,和 2024 年的那个 Neutron four 是不一样的。不过无论如何,我们做这件事已经很久了。我觉得相比某一个具体代际,对我们来说更重要的是 NVIDIA 在开发这些模型上的持续投入与承诺。我们已经做了一段时间了。
Speaker 129:42 - 30:08
I think our models have gotten dramatically more useful in the past year, which is a reflection of two things. Primarily, one is that, the whole company has come together. So there are many different teams around NVIDIA that now understand how important this is to NVIDIA's future. And so there's dramatically more people and better ideas that are going into Neematron. And then, number two, along with that, we've been able to scale the compute resources that go into it. Speaker 129:42 - 30:08
我认为在过去一年里,我们的模型已经变得有用得多,这反映了两点。首先,最主要的一点是,整个公司都协同起来了。所以现在 NVIDIA 内部有很多不同团队都明白,这件事对 NVIDIA 的未来有多重要。因此,投入到 Neematron 中的人大幅增加了,想法也更好了。第二,与此同时,我们也得以扩大投入其中的 compute resources(计算资源)规模。
Speaker 130:08 - 30:32
Obviously, it's very important to have good computing infrastructure to build AI. We've, recently increased our investment substantially because, we believe that that this is really, really key to our company's future. Fascinating. But I just to continue the thought, I think it's really important that everybody knows that we've been doing this for a long time. We are increasing our investments substantially, and NVIDIA is a company that follows through. Speaker 130:08 - 30:32
显然,要构建 AI,良好的 computing infrastructure(计算基础设施)非常重要。最近,我们大幅增加了投资,因为我们相信这对公司的未来真的、真的非常关键。很有意思。不过我想顺着这个思路继续说下去:我认为让所有人都知道这一点非常重要——我们做这件事已经很久了。我们正在大幅增加投资,而且 NVIDIA 是一家说到做到的公司。
Speaker 130:32 - 30:37
You know, we followed through over ten plus years with CUDA, and we're doing that with Neutron now. Speaker 130:32 - 30:37
你知道,十多年里我们一直持续推进 CUDA,现在我们也在对 Neutron 做同样的事。
Speaker 230:37 - 31:10
That's very helpful because, I think the the the broader world is is just starting to catch up to the fact that there is a very substantial open source frontier AI research effort that's been happening. So it's really interesting to hear that there's been this progression. And now there's this family of models that we're going to talk about in a second. Another important moment seems to be, the creation, just in March, three months ago of the Nemotron coalition. Do you wanna explain briefly what that is? Speaker 230:37 - 31:10
这很有帮助,因为我觉得,更广泛的世界才刚开始意识到:一直以来都有一股非常强大的 open source frontier AI research(开源前沿 AI 研究)力量在推进。所以听到这一路演进真的很有意思。现在又有了这个我们马上要谈到的 model(模型)家族。另一个重要时刻似乎是,就在三个月前的 3 月,Nemotron coalition 的成立。你愿意简要解释一下那是什么吗?
Speaker 131:11 - 31:37
So Nemotron exists to help support the ecosystem, and we were thinking, well, this is a different kind of AI project than other projects around the industry. Right? Because, we're not actually trying to dominate in any way. We're just trying to support. We don't we're not trying to control, we're, the way that AI is, being integrated into all these companies. Speaker 131:11 - 31:37
Nemotron 的存在是为了帮助支持整个 ecosystem(生态系统),而我们当时在想,嗯,这和行业里的其他项目是不一样的一类 AI 项目。对吧?因为我们其实并不是想以任何方式去主导。我们只是想提供支持。我们不是想控制 AI 被整合进这些公司的方式。
Speaker 131:37 - 31:59
We're just trying to make sure there's good AI. But we thought, well, maybe if we worked with people while we develop it, then it's gonna be more useful for them. It'll be easier to integrate because we will consider what they need from the beginning. And, you know, Neutron has always been collaborative. I was telling you that, you know, long long time ago, our first big model that we trained, we we did with Microsoft. Speaker 131:37 - 31:59
我们只是想确保有好的 AI。但我们想到,也许如果我们在开发过程中就和大家一起合作,那它对他们来说会更有用。它也会更容易被集成,因为我们从一开始就会考虑他们需要什么。而且,你知道,Neutron 一直都是协作式的。我刚才跟你说过,很久很久以前,我们训练的第一个大型 model(模型),是和 Microsoft 一起做的。
Speaker 131:59 - 32:12
Microsoft. Right? It was a joint effort where NVIDIA and Microsoft researchers worked side by side to build that. And that, that ended up, I think, helping both NVIDIA and Microsoft. I think we both learned a lot from that, experience. Speaker 131:59 - 32:12
Microsoft。对吧?那是一次联合努力,NVIDIA 和 Microsoft 的研究人员并肩合作来构建它。而我觉得,那最终对 NVIDIA 和 Microsoft 都有帮助。我想我们双方都从那次经验中学到了很多。
Speaker 132:12 - 32:54
And so because Neutron is not trying to compete with, other companies but rather support, because we're gonna be putting it out there openly anyway, why not collaborate before the thing is built rather than Neumatron being a project that NVIDIA does all on its own and then posts on the Internet and says, hey. Why don't you try this? We think it might be good. Why don't we make sure that it's good for the partners that that are interested by working with them before Neumatron is even created? And incorporating, you know, any sort of feedback, evaluations, environments, benchmarks, or any other, kinds of technology that other people want to bring. Speaker 132:12 - 32:54
所以,正因为 Neutron 不是要和其他公司竞争,而是要提供支持,而且反正我们本来就会把它公开发布出来,那为什么不在它建成之前就先开展合作呢?而不是让 Neumatron 成为一个完全由 NVIDIA 单独完成、然后再发到 Internet 上说,嘿,你们来试试看吧,我们觉得它可能不错的项目。为什么不在 Neumatron 甚至还没被创建出来之前,就通过和感兴趣的伙伴合作,确保它对他们确实有用呢?并且把各种 feedback(反馈)、evaluation(评估)、environment(环境)、benchmark(基准测试),以及其他任何别人想带进来的技术类型都整合进去。
Speaker 132:54 - 33:17
It turns out that the entire ecosystem, there's a lot of companies that really want open models to succeed. And so they have a self interest. They have their own vested self interest in making sure that open technologies are excellent. And so why not, work with them and and let them contribute, however they'd like to making Neutron better. So that's the the idea of the Neutron coalition. Speaker 132:54 - 33:17
结果发现,整个 ecosystem(生态系统)里有很多公司都真的希望 open model(开放模型)取得成功。所以他们本身就有明确的 self-interest(自身利益),要确保 open technologies(开放技术)足够优秀。既然如此,为什么不和他们合作,并让他们以任何他们愿意的方式出力,来把 Neutron 做得更好呢?这就是 Neutron coalition 的理念。
Speaker 133:17 - 33:37
It is not an exclusive coalition. We're not trying to be the only model out there. All the companies that we work with are free to to continue doing the work however makes sense to them. And yet, you know, these companies want to work with us because they wanna make sure that, open technologies for AI keep, developing quickly and that they have a chance to influence how that happens. Speaker 133:17 - 33:37
这不是一个排他的 coalition(联盟)。我们并不是想成为唯一存在的模型。所有与我们合作的公司,都可以自由地继续按照对他们来说合理的方式推进自己的工作。与此同时,这些公司之所以想和我们合作,是因为他们希望 AI 的 open technologies(开放技术)能继续快速发展,并且他们也能有机会影响这一发展是如何发生的。
Speaker 233:37 - 33:49
Great. What's the current state of the Neumontron family? You got Nano, you got Super, you got Ultra. What do those models do and what are the use cases for them? Speaker 233:37 - 33:49
很好。Neumontron 家族目前的情况怎么样?你们有 Nano、Super 和 Ultra。这些模型分别做什么,适用场景是什么?
Speaker 133:49 - 34:30
So Nano is a 30,000,000,000 total, 3,000,000,000 active parameter model. Supers one twenty and twelve, and Ultra is five fifty and fifty five. They're designed really to fit, you know, it's kind of small, medium, and large deployment scenarios. You know, nano can be really capable for things that, you know, don't require nearly as much knowledge or reasoning but obviously for the for the most capable model, you go for ultra. Super in a lot of ways is our most popular model because it represents kind of a great balance between, cost and and intelligence. Speaker 133:49 - 34:30
Nano 的总参数量是 30,000,000,000,活跃参数量是 3,000,000,000。Super 是一百二十和十二,Ultra 是五百五十和五十五。它们的设计本质上是为了适配不同部署场景,可以理解为小、中、大三档。Nano 对那些不太需要大量知识或复杂推理的任务来说已经非常有能力;当然,如果你需要能力最强的模型,就会选择 Ultra。很多时候,Super 是我们最受欢迎的模型,因为它在成本和智能之间实现了非常好的平衡。
Speaker 134:30 - 35:02
So we we kinda like, having this small, medium, and large, approach to building a family just because our customers, seem to respond to that, pretty well. But, you know, the most important thing from NVIDIA's point of view that people are doing with LLMs is agents. Right? Is, building agentic workflows. Having it having an agent working on your behalf, solving problems for you night and day, is such an exciting way of approaching the problems that we have to solve. Speaker 134:30 - 35:02
所以,我们确实很喜欢这种按小、中、大来构建模型家族的方式,因为客户似乎对此反馈很好。不过,从 NVIDIA 的视角来看,人们使用 LLM(大语言模型)最重要的方向是 agents。对吧?也就是构建 agentic workflows(智能体工作流)。让一个 agent 代表你工作,日夜不停地为你解决问题,这是应对我们必须解决的问题时一种非常令人兴奋的方式。
Speaker 135:03 - 35:08
And, it's our dream to make Nemotron amazing for that purpose. That's that's our goal. Speaker 135:03 - 35:08
而且,我们的梦想就是让 Nemotron 在这个用途上表现得非常出色。这就是我们的目标。
Speaker 235:09 - 35:18
To double click on this, at a hell of a Nemotron family is focused on agentic reasoning with a particular focus on making it efficient. Is that is that the right headline? Speaker 235:09 - 35:18
为了进一步确认这一点,可以这样概括吗:整个 Nemotron 家族都聚焦于 agentic reasoning(智能体推理),并且特别强调效率?这是不是一个准确的标题式总结?
Speaker 135:19 - 35:41
That's right. Yeah. Nemotron has always been, speed first approach to building models because NVIDIA is an accelerated computing company. As I was saying, we're trying to think through what is the problem here computationally from first principles. And, you know, Nemotron, three family has a lot of things in it that are, we're really proud of. Speaker 135:19 - 35:41
没错。是的。Nemotron 一直以来都是以速度优先的方式来构建模型,因为 NVIDIA 是一家 accelerated computing(加速计算)公司。就像我刚才说的,我们会试着从第一性原理出发思考:这里在计算层面上的问题到底是什么。并且,Nemotron 3 家族里有很多内容是我们非常自豪的。
Speaker 135:41 - 36:02
For example, Nemotron Ultra and Super, were pretrained using four bit arithmetic. We pretrained those in MVFP four, which, you know, is a not trivial thing to do to invent the algorithm so that your model can converge to an excellent result using such coarse arithmetic. It required a lot of invention. Really proud of that. Speaker 135:41 - 36:02
例如,Nemotron Ultra 和 Super 在预训练时使用了 four-bit arithmetic(四位运算)。我们是在 MVFP4 下对它们进行预训练的;你知道,这并不是一件简单的事。要发明出这样的算法,使模型在如此粗粒度的运算下依然能够收敛到非常优秀的结果,这需要大量创新。我们对此非常自豪。
Speaker 236:02 - 36:08
Do you wanna explain maybe for for people what, four bit is versus 16 bit, for example? Speaker 236:02 - 36:08
你愿意不愿意顺便给大家解释一下,比如说,four-bit 和 16-bit 相比到底是什么概念?
Speaker 136:08 - 36:32
You know, actually, there was a fantastic post I saw in Hacker News yesterday where somebody let you, upload a picture, and then it would basically posterize it, basically reduce the colors to fit different number formats including NVFP four and MXFP eight and some of the other formats that are out there. And so you could kind of swipe around and look what it does to the colors of the picture. And, you know, it's it's really quite dramatic. Four bits is not a lot of bits. Right? Speaker 136:08 - 36:32
你知道吗,实际上,我昨天在 Hacker News 上看到一篇很棒的帖子:有人做了个东西,让你上传一张图片,然后它会基本上把图片做成 posterize 效果,也就是减少颜色数量,以适配不同的数字格式,包括 NVFP four、MXFP eight,以及其他一些现有格式。这样你就可以来回滑动,看看这些格式会怎样改变图片的颜色。然后,你知道,这个效果其实相当明显。四 bit 的确不算多,对吧?
Speaker 136:32 - 37:04
That's only 16 value. Now, of course, these are all, what are called, block scaled formats. So, groups of numbers also come with an eight bit, scaling factor. And the the specifics of this can get rather complicated, so maybe they're not quite as important. But the the reason why we want to do this is because, first of all, we have dramatically higher, throughput for these formats in our GPUs, specifically on, Blackwell Ultra. Speaker 136:32 - 37:04
那也就只有 16 个取值。当然,这些都属于所谓的 block scaled formats,也就是说,一组数字还会配一个 8 bit 的 scaling factor(缩放因子)。这里面的具体细节会变得相当复杂,所以也许没那么重要。但我们之所以想这么做,首先是因为,在我们的 GPUs 上,这些格式的 throughput(吞吐量)要高得多,尤其是在 Blackwell Ultra 上。
Speaker 137:04 - 37:28
And secondly, we know that it's gonna save an enormous amount of energy. One one way to think about, the computational problem of AI is that we are going to be running at the limit. Whatever the limit is, it could be, an economic limit. Like, we only have so many billion dollars to to buy servers with. It could be a power limit. Speaker 137:04 - 37:28
第二,我们知道这样会节省大量能源。理解 AI 计算问题的一种方式是:我们终究会在某个极限上运行。不管那个极限是什么,可能是经济上的极限。比如说,我们只有这么几十亿美元可以拿来买服务器。也可能是电力上的极限。
Speaker 137:28 - 37:49
We only have so many gigawatts that we can afford to to train a model with. Whatever the limit is, we're gonna be running at that limit. The every organization is is because why? Because the value of intelligence is so high, you know, that that people are gonna they're gonna invest because they they know that they're gonna get return. The the value of intelligence is is enormous. Speaker 137:28 - 37:49
我们只能负担得起用这么多 gigawatt(吉瓦)的电力来训练模型。不管这个极限是什么,我们都会在那个极限上运行。每个组织都会这样,为什么?因为 intelligence(智能)的价值太高了,你知道,人们会愿意投入,因为他们知道自己会获得回报。智能的价值是巨大的。
Speaker 137:50 - 38:14
So if you if you accept as the truth that we're gonna be running at the limit, then what that means is that the way to get more intelligence is to be more efficient. We can't get more intelligence by applying more force if we're already at the limit. We have to be more thoughtful about how we use what we have. And, you know, four bit number formats are dramatically cheaper to move around. They take up less space in memory. Speaker 137:50 - 38:14
所以,如果你接受“我们会一直在极限上运行”这个事实,那么它意味着,想获得更多智能,方法就是提高效率。如果我们已经处在极限上,就不能再靠施加更多资源来获得更多智能。我们必须更用心地使用手头已有的资源。而且,你知道,四 bit 数字格式在搬运时成本要低得多,占用的 memory(内存)空间也更小。
Speaker 138:15 - 38:50
They take up less picojoules when you move them from the memory in even on the chip, around the chip. Much less, energy when you compute on them. And so, so that's really, driving, you know, the investment in four bit formats. And I think these days, four bit formats for deployment are very well established. It's it's pretty pretty straightforward these days to make a a good quantized four bit checkpoint that you can deploy and that gets you a lot of inference cost and speed advantages. Speaker 138:15 - 38:50
把它们从 memory 里移动出来时——甚至只是在 chip(芯片)内、绕着芯片传输——消耗的 picojoules(皮焦耳)也更少。对它们做计算时,能耗也低得多。所以,这才是真正推动我们投资四 bit 格式的原因。我认为现在,四 bit 格式用于 deployment(部署)已经非常成熟了。如今要做出一个效果不错、可部署的 quantized(量化)四 bit checkpoint(检查点)已经相当直接,而且它能在 inference(推理)成本和速度上带来很大的优势。
Speaker 138:51 - 39:25
But using four bit formats for pretraining, that's quite a bit more challenging because you have this numeric solver that's, you know, optimizing the weights and, you know, it it can be quite sensitive. So if you if you don't, treat the numbers right, your model can diverge. And instead of actually getting a model done through pre training, you end up with, you know, basically just that run diverged, which is, you know, always always scary. So it it took a lot of invention for us, to be able to pretrain Neutron, in four bit. We're really proud of that. Speaker 138:51 - 39:25
但把四 bit 格式用于 pretraining(预训练)就更具挑战性得多,因为这里有一个 numerical solver(数值求解器)在优化 weights(权重),而它可能相当敏感。所以如果你不能正确处理这些数字,模型就可能 diverge(发散)。这样一来,你不但无法真正完成预训练得到一个模型,最终还可能只是得到“这次训练跑发散了”这样的结果,而这总是很吓人的。所以,为了能够用四 bit 来预训练 Neutron,我们做了很多创新。对此我们真的很自豪。
Speaker 239:25 - 39:37
Okay. Great. Alright. So as we, get into, slightly more technical things, the architecture of NeMOtron is hybrid. Is that right? Speaker 239:25 - 39:37
好的,很棒。好,那么当我们进入稍微更技术性一点的话题时,NeMOtron 的架构是 hybrid(混合式)的,对吗?
Speaker 239:37 - 39:48
So it's a combination of transformer and member state space, which is a slightly more exotic form of of of architecture. Walk us through that. Speaker 239:37 - 39:48
所以它是 transformer 和 member state space 的组合,这是一种稍微更“异域”一点的架构形式。给我们详细讲讲这个吧。
Speaker 139:48 - 40:37
Yeah. You know, we published a paper in 2024 that showed that you actually get a smarter model by combining state space models with transformers. And we we actually did a sweep of, you know, how much of the model should be full attention and how much of it should be, a state space model in order to get the the lowest perplexity, basically the best the best language model that you could get. And we found that you actually want it to be mostly a state space model with a little bit of attention. And it's kind of the intuition behind that is that the state space model seem to be better at kind of this intuitive intuitive kind of impressionistic understanding of a sequence, because they're, you know, they're kind of summarizing the entire sequence into a constant space. Speaker 139:48 - 40:37
对。你知道,我们在 2024 年发表了一篇论文,表明把 state space model 和 transformers 结合起来,实际上能得到一个更聪明的模型。我们当时还系统扫了一遍参数空间:模型里到底应该有多少部分是 full attention(全注意力),又有多少部分应该是 state space model,才能得到最低的 perplexity(困惑度),也就是基本上你能做出的最好的 language model。然后我们发现,理想情况其实是:大部分应该是 state space model,再加上一小部分 attention。其背后的直觉大概是,state space model 似乎更擅长对一个序列形成那种更直观、印象式的理解,因为它们本质上是在把整个序列总结进一个常量大小的空间里。
Speaker 140:37 - 41:04
That's how they work. Right? So instead of having the ability to look at the entire sequence randomly, they summarize everything at every step into a constant, cache or a little scratch pad that they're they're working on. And, that constraint seems to actually make them smarter at some tasks that involve, like, global understanding. On the other hand, the advantage of full attention is that it can pick out very specific, bits of information and look at those exactly. Speaker 140:37 - 41:04
它们就是这么工作的,对吧?也就是说,它们不是能够随机地查看整个序列,而是在每一步都把所有内容总结进一个恒定大小的 cache(缓存),或者说一个它们正在操作的小 scratch pad(草稿板)里。而这种约束似乎反而让它们在某些任务上变得更聪明,尤其是那些需要全局理解的任务。另一方面,full attention 的优势在于,它可以挑出非常具体的信息片段,并且精确地去查看它们。
Speaker 141:04 - 41:22
It doesn't lose anything. There's no lossy compression going on. You can actually see the whole thing. And so we found that, you know, using both of these together was actually better than using either one on their own. And that is independent of the speed benefit. Speaker 141:04 - 41:22
它不会丢掉任何东西。这里不存在有损压缩。你确实可以看到完整内容。所以我们发现,把这两者结合起来,实际上比单独使用任意一种都更好。而这一点和速度收益是无关的。
Speaker 141:22 - 41:44
That is just the model is smarter. And since we've published that, I think a lot of other labs have also found this to be true. A lot of models these days are being built with hybrid SSM approaches. For example, QN has done that. Kimi is using what they call Kimi linear attention these days. Speaker 141:22 - 41:44
这纯粹是因为模型更聪明了。自从我们发表那篇论文之后,我觉得很多其他实验室也都发现这点是成立的。如今很多模型都在采用 hybrid SSM(混合式状态空间模型)方案来构建。比如,QN 就这么做了。Kimi 现在用的是他们所说的 Kimi linear attention。
Speaker 141:44 - 42:31
So it's become, I think, quite widely adopted to use some sort of state space model in conjunction with, full attention for, for the the base architecture. Now it also has some speed benefits because, the amount of memory that you need to hold, that state space, cache is actually constant with respect to your sequence length, which then means that, generally, you can fit much higher batches on the GPU when you're training and doing inference, because the memory, requirement is lower. And it keeps the GPU, fuller and busier, and therefore, you know, provide some some pretty important pretty important efficiency benefits as well. Speaker 141:44 - 42:31
所以我觉得,现在在基础架构里把某种 state space model 和 full attention 结合使用,已经变得相当普遍了。同时它还有一些速度上的好处,因为你用来保存那个 state space cache 的内存量,相对于序列长度其实是常量,这就意味着无论在训练还是推理时,通常都能在 GPU 上塞进更大的 batch,因为内存需求更低。这样也会让 GPU 保持更满、更忙,因此还能带来一些相当重要的效率收益。
Speaker 242:31 - 42:42
So the models are also based on an MOE, a mixture of experts architecture. Walk us through that and maybe remind people what MOE is in the first place. Speaker 242:31 - 42:42
所以这些模型也基于 MOE,也就是 mixture of experts 架构。给我们讲讲这个,也顺便提醒一下大家,MOE 到底是什么。
Speaker 142:42 - 43:02
So mixture of experts is a form of sparsity. The idea is, wow, you wanna train a model on the entire Internet. You want it to remember absolutely everything about the history of everything. But when you're answering a particular question, does it seem reasonable that it needs to actually think about the entire universe in order to answer that question? Actually, no. Speaker 142:42 - 43:02
mixture of experts 是一种 sparsity(稀疏性)形式。它的想法是这样的:你想在整个互联网的数据上训练一个模型。你希望它记住关于万事万物历史的一切。但当它在回答某一个具体问题时,它真的有必要为了回答这个问题而去“思考”整个宇宙吗?其实并不需要。
Speaker 143:02 - 43:20
It seems like it's quite sparse. Right? It seems like, we're we're we're using a language model to explore a very tiny space of ideas in order to answer a question or solve a problem. We want the model to be able to draw from the entire universe. We wanna train it so that it understands everything that it possibly can. Speaker 143:02 - 43:20
这看起来好像相当稀疏,对吧?感觉就像,我们在用一个 language model(语言模型)探索一个非常微小的想法空间,以便回答某个问题或解决某个问题。我们希望这个模型能够从整个宇宙中汲取信息。我们想训练它,让它尽可能理解它所能理解的一切。
Speaker 143:20 - 43:59
But when it's actually running, it doesn't really need to see all of that information. There's been a variety of approaches to sparsity that try to take advantage of this property, but mixture of experts has been the most successful. And the way that it works is that the neural network has what's called a router that is learned that, is gonna decide to send activations to a subset of the experts, for every token that's flowing through every layer of the model. It's gonna be making choices about, which fraction of the model is gonna actually get to interact with this token as we try to understand it, build up representations of the problem, and then generate the next token that we're gonna output. Speaker 143:20 - 43:59
但当它实际运行时,它其实并不需要看到所有这些信息。围绕这种稀疏性,已经有各种不同的方法试图利用这一特性,不过 mixture of experts 最成功。它的工作方式是:神经网络里有一个经过学习得到的 router,它会决定,对于流经模型每一层的每一个 token,要把 activations 发送给一部分 experts。它会不断做出选择:在我们试图理解这个 token、构建对问题的表征、然后生成我们将要输出的下一个 token 时,模型中到底有哪一部分能够真正与这个 token 发生交互。
Speaker 243:59 - 44:15
So it's a little bit like if I have a company with five fifty employees, but 55 of them are in engineering, I want the 55 employees who are specialists to come to my meeting about engineering and not the rest of the company. Speaker 243:59 - 44:15
这有点像,如果我有一家公司,有 550 名员工,但其中 55 人在 engineering,那么我希望来参加我的 engineering 会议的是这 55 名专业员工,而不是公司的其他人。
Speaker 144:15 - 44:28
That's right. Yeah. Or you can think about it as a library. Like, if you go into a library to do research, you don't read all of the books in the library. Like, your first job is to figure out which books you need to look at in order to find the answer to your question. Speaker 144:15 - 44:28
没错。对。你也可以把它想成一座图书馆。比如,你去图书馆做研究时,你不会把图书馆里的所有书都读一遍。你的第一项工作,是先弄清楚你需要看哪些书,才能找到你问题的答案。
Speaker 144:28 - 44:51
And so, so that's kind of the the idea behind MOEs. Now MOEs have fascinating implications for the systems that we built. So with Blackwell, for example, NVIDIA went all in on MOEs. That's why we built NVL 72, which allows up to 72 of our GPUs to read and write each other's memory, at very high speeds, very low latency. And why is that important? Speaker 144:28 - 44:51
所以,这就是 MOEs 背后的基本思路。现在,MOEs 对我们构建的系统有非常引人入胜的影响。比如说,对于 Blackwell,NVIDIA 在 MOEs 上是全力投入的。这就是为什么我们构建了 NVL 72,它可以让多达 72 个我们的 GPUs 以非常高的速度、非常低的延迟相互读写彼此的内存。那为什么这很重要呢?
Speaker 144:51 - 45:27
It's because as you put a token through the stack of layers, at every layer, you have a router that's routing that token somewhere else. Why don't you partition your experts so that, you know, the experts are not sitting every expert on every GPU, but you have a subset of the experts assigned to each GPU. And then you're routing the tokens between the GPUs very dynamically as you push, the token through the network. Now, this is impossible to predict in advance where the tokens need to go because it's very specific to that particular token for that particular model. And so that's why we built NVL 72. Speaker 144:51 - 45:27
因为当你把一个 token 送过整叠层时,在每一层里,你都有一个 router,会把这个 token 路由到别的地方。为什么不把 experts 做分区呢?也就是说,不是让每个 GPU 上都放着每一个 expert,而是给每个 GPU 分配一部分 experts。然后,当你把 token 推过这个网络时,就可以在 GPUs 之间非常动态地路由这些 token。现在,token 需要去哪里,是不可能提前预测的,因为这对那个特定模型里的那个特定 token 来说是高度具体的。这也就是为什么我们构建了 NVL 72。
Speaker 145:27 - 45:54
And that's why, Blackwell is so amazing for inference for, you know, today's AI models is because we thought deeply about mixture of experts when we were building it. And this is speaking to Neematron's first job. You know, if if we hadn't been working on understanding AI, we wouldn't have been able to build Blackwell properly. And that, you know, has has translated directly into, you know, increased deployment, of Blackwell, which, you know, we're we're we're very excited about. Speaker 145:27 - 45:54
这也就是为什么 Blackwell 对于 inference 来说、对于当今的 AI models 来说,如此惊艳——因为我们在构建它时,深入思考了 mixture of experts。这也呼应了 Neematron 的第一项工作。你知道,如果我们当时没有致力于理解 AI,我们就不可能把 Blackwell 正确地构建出来。而这也直接转化成了 Blackwell 部署量的增长,这一点我们非常兴奋。
Speaker 245:55 - 46:00
Is what you just described, called latent MOE, or is that a different concept? Speaker 245:55 - 46:00
你刚才描述的这个,叫 latent MOE 吗,还是说那是一个不同的概念?
Speaker 146:00 - 46:35
Latent MOE is a specific, innovation that we have in Nemotron three family. And, what it does is actually, reduces the amount of communication that has to be sent through NVLink during MOE computations by basically down projecting it. So, you know, every token is it produces a vector. And the idea is, like, we're gonna take that vector and learn a way to compress it and then send that compressed thing through the network, and then we're gonna uncompress it at the other end. And as a result, we save on network bandwidth, and we also get four times the number of experts for the same inference cost. Speaker 146:00 - 46:35
Latent MOE 是我们在 Nemotron three family 中的一项特定创新。它的作用实际上是:在进行 MOE 计算时,通过基本上把信息做降维投影,来减少必须通过 NVLink 发送的通信量。也就是说,每个 token 都会产生一个向量。我们的思路是,把这个向量以某种可学习的方式压缩,然后把压缩后的内容通过网络发送,再在另一端解压。这样一来,我们既节省了网络带宽,也能在相同的 inference cost(推理成本)下,把 experts 的数量提升到四倍。
Speaker 146:35 - 46:48
So you could think about it as like, you know, our library of books got four times bigger, and we get to, you know, read four times more books, at the same inference cost because of because of this particular innovation. Speaker 146:35 - 46:48
所以你可以把它理解为:我们的“藏书馆”规模变成了原来的四倍,而因为这项创新,我们可以在相同的 inference cost(推理成本)下,“阅读”四倍数量的书。
Speaker 246:48 - 46:53
Is MOE in general, becoming the default architecture for Frontier AI? Speaker 246:48 - 46:53
总的来说,MOE 是否正在成为 Frontier AI 的默认架构?
Speaker 146:54 - 47:03
Yeah. I believe MOEs have been the default, in Frontier AI for a long time. They're just a really good combination of inference cost and intelligence. Speaker 146:54 - 47:03
是的。我认为 MOE 很长时间以来就一直是 Frontier AI 中的默认选择了。它在 inference cost(推理成本)和 intelligence(智能水平)之间,确实提供了非常好的组合。
Speaker 247:03 - 47:04
Great. Great. Speaker 247:03 - 47:04
很好。很好。
Speaker 147:04 - 47:26
But they have drawbacks as well. You know, they they take a lot more memory. If you have a a very small amount of memory, a dense model is gonna be smarter. And they also, they tend to work best either if you're running at batch size one, so you're running basically a single job, or you're running a huge data center with, like, infinite queries coming in. In the middle, they can be a little bit tricky. Speaker 147:04 - 47:26
但它们也有缺点。你知道,MOE 会占用更多内存。如果你的内存非常有限,dense model(稠密模型)会更聪明一些。而且,它们通常在两种情况下效果最好:要么你是在 batch size 为 1 的情况下运行,也就是基本上只跑单个任务;要么你是在一个超大数据中心里运行,像是有无限多的查询不断涌入。介于这两者之间的场景,就会稍微棘手一些。
Speaker 247:26 - 47:40
Another important characteristic of Neumatron three Ultra is a 1,000,000 token context, the the long context window. How important is that in the overall mix, and what does it enable the model to do? Speaker 247:26 - 47:40
Neumatron three Ultra 的另一个重要特性是 1,000,000 token context,也就是超长上下文窗口。这在整体能力组合中有多重要?它又让模型能够做到什么?
Speaker 147:40 - 48:14
The longer the context length, the more challenging problems we can solve with a language model. That allows us to do things like append all sorts of information to a query, which could be a code base. It could be instructions. You know, in in the long term, I'm hoping that I have my own personal LLM that's able to read all of my emails, you know, and help me answer questions about that. You know, the more information that we can attach to a particular query, the the more useful the model can be. Speaker 147:40 - 48:14
context length(上下文长度)越长,我们就越能用语言模型解决更有挑战性的问题。这让我们可以把各种各样的信息附加到一个 query(查询)上,比如一个 code base,或者一组 instructions(指令)。从长期来看,我希望自己能拥有一个个人 LLM,能够读取我所有的 emails,并帮我回答与之相关的问题。我们能够附加到某个特定 query 上的信息越多,模型就会越有用。
Speaker 148:14 - 48:35
Now, it can get more and more expensive, right, to reason over large amounts of of input data. And, and so that's one of the the reasons why there's usually a limit on how big the context length can be. But with Neemotron three, we we tried to push it as far as we could go. We think a million tokens is a lot of tokens, and you can do a lot of things with that. Speaker 148:14 - 48:35
现在,对大量输入数据进行推理的成本会越来越高,对吧。所以这也是为什么 context length(上下文长度)通常会有上限的原因之一。但在 Neemotron three 里,我们尽可能把这个上限往前推进。我们认为一百万个 token 已经是非常多的 token 了,你可以用它做很多事情。
Speaker 248:36 - 48:51
Prisma is particularly helpful in, sort of multistep, agentic workflows, and there's, this whole separate discussion around context compaction to make sure that the model doesn't get lost in too many tokens. So, like, how do you how do you all think about this? Speaker 248:36 - 48:51
Prisma 在某种多步骤的 agentic workflow(智能体式工作流)中特别有帮助;另外还有一个独立的话题,就是 context compaction(上下文压缩),以确保模型不会因为 token 太多而“迷失”。所以,你们大家是怎么考虑这个问题的?
Speaker 148:51 - 49:23
A 100%. I mean, you know, compaction is that's a thing if you're using an agentic workflow you deal with all the time. And compaction tends to work pretty well, you know, because language models, are pretty good at at identifying the most relevant things and summarizing, and you're basically trying to summarize your context when you compact it. So, compaction, is is not a bad approach. I think having models that can just natively, reason about larger amounts of data is just inherently more useful. Speaker 148:51 - 49:23
绝对是这样。我的意思是,如果你在使用 agentic workflow,那么 compaction 就是你一直都要面对的事情。而且 compaction 往往效果相当不错,因为 language model(语言模型)其实很擅长识别最相关的信息并进行总结,而你在做 compaction 时,本质上也就是在总结你的上下文。所以,compaction 并不是一个坏办法。但我认为,拥有能够原生地对更大量数据进行推理的模型,本身会更有用。
Speaker 149:23 - 49:26
So, of course, we wanna push the the boundary on that as well. Speaker 149:23 - 49:26
所以,当然,我们也想继续推动这方面的边界。
Speaker 249:26 - 49:31
Great. Can you talk about the multi token prediction, which is also very interesting? Speaker 249:26 - 49:31
很好。你能谈谈 multi token prediction(多 token 预测)吗?这个也非常有意思。
Speaker 149:32 - 50:10
If you're running at a low batch size, which is when you are trying to get the most interactivity if you're in a data center, so you want you want the model to respond as quickly as possible, and it's okay for it to be more expensive. Your token your cost per token might might be higher, but you want the result as quickly as possible. Or if you're running, locally, so you might be running at batch size one just because you're the only person using it. It turns out that, the GPU has extra execution capabilities that are just lying there unused. The bulk of the work when you're running in these scenarios is actually fetching the weights from memory. Speaker 149:32 - 50:10
如果你是在较低 batch size(批大小)下运行——这通常是你在 data center(数据中心)里追求最高交互性的时候——那么你希望模型尽可能快地响应,即使这样成本更高也没关系。你的每个 token 的成本可能会更高,但你希望尽快得到结果。或者,如果你是在本地运行,那么你可能就是以 batch size one 运行,因为只有你一个人在用它。事实证明,在这些情况下,GPU 其实还有额外的执行能力处于闲置状态。因为在这些场景里,运行时的大部分工作其实是从内存中取回 weights(权重)。
Speaker 150:10 - 50:40
And then you push the token past those weights, and then you fetch more more, weights from memory. But it turns out if you if you push two tokens or even five tokens through those same, weights, it would cost basically the same amount of time. Because the the expensive thing is not doing the math to push the token through the weights. The expensive thing is just reading all of those weights from memory, all those parameters they have to come in. And so the idea with multi token prediction is to take advantage of this by having the model predict multiple tokens at once. Speaker 150:10 - 50:40
然后你让 token 通过这些 weights,接着再从内存中取回更多的 weights。但事实证明,如果你让两个 token,甚至五个 token,通过同样的这些 weights,所花的时间基本是一样的。因为昂贵的部分并不是做那些数学计算,让 token 通过这些 weights;真正昂贵的是从内存中读取所有这些 weights,把那些 parameters(参数)都调进来。所以 multi token prediction 的思路,就是利用这一点,让模型一次预测多个 token。
Speaker 150:40 - 50:59
Let's say that the model predicts five tokens. We know the first token is correct. The next four tokens may or may not be correct. So then what we do is on the next pass, we take those four tokens and we stick them into the model, and then, run it through. And at the end, we check, you know, the model then predicts another set of tokens. Speaker 150:40 - 50:59
假设模型预测了五个 token。我们知道第一个 token 是正确的,接下来的四个 token 可能正确,也可能不正确。于是我们下一轮会把这四个 token 放进模型里,再跑一遍。然后在结束时,我们再检查——这时模型又会预测出另一组 token。
Speaker 150:59 - 51:26
Right? Then we check, were the extra tokens we predict last time correct? If so, then we just accept them, and then we get, like, a four x speed up. And if they were incorrect, then we only accept the ones that were correct and then, you know, proceed from there. So the benefit of this is it it doesn't degrade accuracy at all because you're using the model to double check. Speaker 150:59 - 51:26
对吧?然后我们会检查:我们上一次预测的那些额外 token 是否正确?如果正确,那我们就直接接受它们,这样大概就能获得 4 倍加速。如果不正确,那我们就只接受其中正确的那些,然后再从那里继续往下走。所以它的好处是完全不会降低准确率,因为你是在用模型本身做二次检查。
Speaker 151:26 - 51:47
Right? So all the speculation is gonna get checked during the next token that you you run through the model. So it doesn't degrade your accuracy at all to turn on multi token prediction, but it can give you a speed up and it's probabilistic depending on the acceptance rate of your, predictor. You know? So if your predictor's more accurate, the acceptance rate goes higher, you get a higher speed up. Speaker 151:26 - 51:47
对吧?所以所有这些 speculative(推测式)内容,都会在你下一次让模型跑下一个 token 时被检查一遍。因此,开启 multi token prediction 并不会丝毫降低准确率,但它可以带来加速,而且这是概率性的,取决于你的 predictor(预测器)的 acceptance rate(接受率)。对吧?所以如果你的 predictor 更准确,acceptance rate 就会更高,你得到的加速也就更大。
Speaker 151:47 - 52:16
So with, you know, with, our recent Neutron models, you know, we're pretty proud of our acceptance rates, but we're always trying to make them better, you know, always trying to improve that acceptance rate. This is a really good example of accelerated computing. You know, with multi token prediction, the speed that you get is a function of the accuracy of your model. The more accurate your model is, the faster the inference is, the cheaper the inference is, the more accurate it is. That's not usually how it works, but in this case, that's how it works. Speaker 151:47 - 52:16
所以,拿我们最近的 Nemotron 模型来说,我们对自己的 acceptance rate 还是相当自豪的,但我们也一直在努力把它做得更好,一直在想办法提升这个 acceptance rate。这是 accelerated computing(加速计算)的一个非常好的例子。用 multi token prediction 时,你能获得的速度,本质上是模型准确率的函数。你的模型越准确,inference(推理)就越快,inference 也越便宜,而且还更准确。通常事情不是这么运作的,但在这里,它就是这样运作的。
Speaker 152:16 - 52:47
And what that implies is that, you know, if we're trying as NVIDIA, as a company to provide meaningful acceleration to the world's most important computational workloads, this has to be an important part of how we, think about it. You know? If there's a three x cost reduction or speed improvement, for inference, which is the most important computational workload of 2026, If that's on the table and it depends on the accuracy of the multi token prediction network, then that's something that NVIDIA needs to understand very deeply because it's gonna affect, our business directly. Speaker 152:16 - 52:47
而这意味着,如果我们站在 NVIDIA 这家公司角度,想要为全球最重要的计算工作负载提供有意义的加速,那这就必须成为我们思考问题时非常重要的一部分。对吧?如果 inference——也就是 2026 年最重要的计算工作负载——存在 3 倍的成本下降或速度提升空间,而这又取决于 multi token prediction network 的准确率,那么这就是 NVIDIA 必须非常深入理解的东西,因为它会直接影响我们的业务。
Speaker 252:47 - 52:58
Fascinating. And just to to continue on the on on the tour, multi teacher distillation. We talked about distillation a little bit upfront. What does that mean in the context of, Nemotron three? Speaker 252:47 - 52:58
很有意思。那就顺着这个话题继续往下聊,multi teacher distillation。我们前面已经稍微谈到过 distillation 了。在 Nemotron three 的语境里,这具体是什么意思?
Speaker 152:58 - 53:30
So with Nemotron three Ultra, we did post training using something called multi domain on policy distillation. And what that, entails is that, you know, we have many different aspects of the model we want to improve. For example, science understanding is different from math theorem proving, which is different from coding, which is different from agent harness, interactions. Right? There's there's, with Neutron three, I think we had about, ten ten or 15 of these teachers. Speaker 152:58 - 53:30
所以在 Nemotron three Ultra 中,我们做了 post training(后训练),用的是一种叫 multi domain on policy distillation 的方法。这具体意味着什么呢?就是我们有很多不同方面想要提升这个模型。比如说,science understanding 和 math theorem proving 不一样,和 coding 不一样,也和 agent harness 交互不一样。对吧?在 Neutron three 里,我记得我们大概有 10 到 15 个这样的 teacher(教师模型)。
Speaker 153:31 - 54:06
And, so the idea is that you take these teacher models and you push them as far as you can go on some specific domain. So you just don't worry about making good at everything, just make it really, really smart at this one domain. Then you have a collection of these models and you want to create one model that learns, to be good at everything. And we do that using a specific reinforcement learning technique that a lot of labs these days use called, MOPD. And the the good thing about this is that because the teachers are supervising, they can give really dense rewards to the the student model. Speaker 153:31 - 54:06
所以这个思路是:你拿这些 teacher model,把它们在某个特定 domain(领域)上尽可能往前推。也就是说,你先别管它是不是样样都强,只要让它在这一个领域里变得非常、非常聪明。然后你就会得到这样一组模型,接着你想创建一个单一模型,让它学会样样都擅长。我们是通过一种特定的 reinforcement learning(强化学习)技术来做到这点的,现在很多实验室都在用,叫 MOPD。它的好处在于,因为这些 teacher 在监督,所以它们可以给 student model 提供非常密集的 reward(奖励)。
Speaker 154:06 - 54:39
Basically, every token is getting supervised, and so the student can learn really quickly, and then become, you know, almost as good as all of the teachers at all of the things. So, one benefit of this is that it really helps the team work together better. You know, if if, you don't have a technique like this and you have, let's say, 500 people working to try to make a model better and one team's like, well, I'm trying to make it better at this thing. And then another team's like, I'm trying to make it better at that thing. There can be a tug of war where it's like, well, who wins? Speaker 154:06 - 54:39
基本上,每一个 token 都在被监督,所以 student 可以学得非常快,然后最终在所有这些事情上都变得几乎和所有 teacher 一样好。所以这样做的一个好处是,它确实能帮助团队更好地协同工作。你想,如果没有这种技术,而你有,比如说,500 个人都在努力让一个模型变得更好,一个团队说“我在努力让它在这件事上变得更强”,另一个团队说“我在努力让它在那件事上变得更强”,那就可能出现一种拔河局面:到底谁说了算、谁赢?
Speaker 154:39 - 54:57
You know? With and and if you have to make a choice like, oh, I'm gonna make I'm gonna choose to prioritize this one over that one, then you make the other team feel like their work doesn't matter. You know? It's just really hard. One of the challenges of of building AI in in 2026 is that you have to figure out how to get the people to work together even though you're only building one thing at the end of the day. Speaker 154:39 - 54:57
你知道吗?如果你必须做出这样的取舍,比如说,哦,我要决定优先做这个而不是那个,那你就会让另一个团队觉得他们的工作不重要。你知道吗?这真的很难。到了 2026 年,构建 AI 的挑战之一在于:尽管说到底你最终只是在构建一件东西,但你必须想办法让大家协同工作。
Speaker 154:58 - 55:05
And so this particular technology, has been really instrumental in helping more people work together to make NeMOtron stronger. Speaker 154:58 - 55:05
所以,这项特定技术在帮助更多人一起协作、让 NeMOtron 变得更强这件事上,确实起到了非常关键的作用。
Speaker 255:05 - 55:34
Fascinating. So is that just as much a technology question as a human organization question? Okay, fantastic. Let's put a pin in this and get back to this in a second because it's a fascinating topic. In terms of the post training that you just alluded to, one of the exciting things that you all did in the context of Nemotron is also to publish the data, the training data. Speaker 255:05 - 55:34
很有意思。所以这既是一个技术问题,也是一个人的组织管理问题,对吗?好,非常好。我们先把这个话题放一放,等会儿再回来,因为这确实很有意思。说到你刚才提到的 post training(后训练),你们在 Nemotron 相关工作中做的一件很令人兴奋的事,也是把数据、也就是训练数据,公开发布了。
Speaker 255:34 - 56:05
Does that include per industry data for specific reinforcement learning tasks? Yes. That's the beauty of a conversation like this today where you guys can actually talk about those things. So where does one get the data from for post training reinforcement learning focused efforts? One of the key questions in the world today is that LLMs or AI systems have become great at coding and great at math. Speaker 255:34 - 56:05
这里面是否包括按行业划分、面向特定 reinforcement learning(强化学习)任务的数据?是的。这也是今天这种对话的妙处——你们其实可以直接谈这些事。所以,对于以 post training reinforcement learning 为重点的工作,人们到底该从哪里获取数据?当今世界里的一个关键问题是,LLM(大语言模型)或 AI 系统已经变得非常擅长 coding,也非常擅长 math。
Speaker 256:05 - 56:24
So like the next big question is can they become great at law and consulting and you know, all sorts of different domains? And part of the the black box of closed models is, like, how people go about doing all all of this, where do they get the data from? To the extent that you can talk about all of this, I'd be very curious about how you guys have have gone about it. Speaker 256:05 - 56:24
所以下一个大问题就像是:它们能不能也变得擅长 law、consulting,以及你知道的,各种不同领域?而 closed models(闭源模型)的黑箱部分之一就在于,人们究竟是怎么做这一切的,他们的数据到底从哪里来?在你们方便谈的范围内,我会非常想了解你们是如何处理这件事的。
Speaker 156:25 - 57:02
It's not an easy question to answer because it is quite complex, but I would say, we rely on a number of things. One is that we do purchase data from, companies that that, you know, are are building datasets that you can purchase. And to the extent that, you know, we have the rights to redistribute, or to to to open up that data, we do as part of, our our, NeMotron data effort. You know, with with Neutron, we are trying to be maximally open with the data that we release because our goal is to support the ecosystem. Right? Speaker 156:25 - 57:02
这个问题不容易回答,因为它确实相当复杂,但我会说,我们依赖很多种方式。其一是,我们会从一些公司购买数据,这些公司本身就在构建可供购买的 datasets(数据集)。而在我们拥有再分发权,或者说有权公开这些数据的情况下,我们就会把它们作为我们 NeMotron data effort 的一部分开放出来。你知道,借助 Neutron,我们在发布数据这件事上正尽可能做到最大程度的开放,因为我们的目标是支持整个生态系统。对吧?
Speaker 157:02 - 57:44
Our goal is is not to be the only model out there, and we love it when we hear of other models around the industry that are using our datasets, to to make their AI stronger because that means we're succeeding in our job to keep the ecosystem thriving and growing. Now we also are big believers in synthetic data generation. We use an enormous amount of compute, running language models on our own systems to create synthetic data that then helps our models be better at, solving problems in specific domains. And we release a lot of that data as well. Now it's, of course, not very straightforward to do this. Speaker 157:02 - 57:44
我们的目标并不是成为唯一存在的 model(模型),而且当我们听说行业里的其他模型正在使用我们的 datasets 来增强它们的 AI 时,我们会非常高兴,因为这意味着我们在维持生态系统繁荣和增长这件事上做对了。与此同时,我们也非常相信 synthetic data generation(合成数据生成)。我们投入了海量 compute(算力),在我们自己的系统上运行 language models(语言模型)来生成 synthetic data,这些数据随后帮助我们的模型更好地解决特定领域中的问题。我们也公开发布了其中很多数据。当然,这件事做起来并不是那么直接。
Speaker 157:44 - 58:01
Like, you know, AI is always garbage in, garbage out. So, you have to work really hard to make sure that any synthetic data that you create is actually adding value that's actually helping the model generalize and solve problems, more intelligently. But those are the the primary ways that we go about, building our datasets. Speaker 157:44 - 58:01
比如说,你知道的,AI 一直都是 garbage in, garbage out(输入垃圾,输出垃圾)。所以,你必须非常努力地确保你创建出来的任何 synthetic data 都是在真正增加价值,真正帮助模型更好地 generalize(泛化),并且更智能地解决问题。不过,这些就是我们构建 datasets 的主要方式。
Speaker 258:01 - 58:36
Since we're talking about post training and RL in different domains, just curious to get your thoughts on where we go from here in terms of generalization. So just to, build on what I was saying a second ago, the industry seems to be marching from coding and math, which are domains with verifiable rewards, to different industries. Do you think that this is where things are going and that the AI industry as a whole is going to be able to cover those next few domains as efficiently as coding or math? Speaker 258:01 - 58:36
既然我们在谈 post training 和 RL(强化学习)在不同领域中的应用,我很好奇你对接下来在 generalization(泛化)方面的发展怎么看。接着我刚才的话说,行业似乎正从 coding 和 math 这些具有可验证奖励的领域,迈向其他行业。你觉得这就是未来的方向吗?整个 AI 行业是否也能像在 coding 或 math 中那样高效地覆盖接下来的几个领域?
Speaker 158:36 - 59:41
Coding is really special because it's a very intellectual exercise that created a lot of economic value, which then meant that we had an enormous amount of tokens that we could, learn from, as well as, tooling that allows us to verify whether, you know, our our models are actually solving problems. So coding is always gonna have a special place in our heart, and and something that I think AI is gonna continue to get much better at, because we we have this, special relationship with it. You know, with regards to other domains, I think what I'm excited about has to do with significantly more diverse environments for AI to learn in during reinforcement learning. I believe that, you know, I mean, reinforcement learning is is such a general form of, teaching an AI how to solve problems. We're just getting started at figuring out how to apply that. Speaker 158:36 - 59:41
Coding 的确很特别,因为它是一种非常智力密集的活动,而且创造了大量经济价值,这意味着我们拥有海量的 token(文本片段)可以用来学习,同时也有工具能够验证,我们的 models 是否真的在解决问题。所以 coding 一直都会在我们心中占据特殊地位,而且我认为 AI 在这方面还会持续变得更强,因为我们和它之间有这种特殊关系。至于其他领域,我真正感到兴奋的是,AI 在 reinforcement learning 过程中可以进入显著更多样化的环境来学习。我相信,reinforcement learning 本身是一种非常通用的、教 AI 如何解决问题的方式。我们现在才刚开始摸索如何去应用它。
Speaker 1 | 59:41 - 1:00:16 And I think as our environments get more sophisticated, the AI then learns more understanding of the problems that it's trying to solve as well as the implications of the actions that it can take, then it becomes much better at, at actually, solving those problems. When I look at the the, environments that we're using today, they're still, fairly simple, all things considered. And I think, that's gonna become significantly more complex and diverse over the next few years.
我认为,随着我们的环境变得更复杂,AI 也会更理解它试图解决的问题,以及它所采取行动会带来什么影响;那样一来,它在真正解决这些问题时就会强得多。现在回头看我们今天使用的这些环境,综合来看,它们仍然相当简单。我认为,未来几年这类环境会变得显著更复杂、也更多样。
Speaker 2 | 1:00:16 - 1:00:32 Alright. So you you mentioned, you know, making 500 people work together, and I said, that we would get back to it because it it's so interesting. So just, taking a step back, like, tell us about the research organization at, NVIDIA. Like, how is it structured? How how does it all work?
好的。你刚才提到了“让 500 个人一起协作”,我当时说我们会再回到这个话题,因为它非常有意思。所以先退一步讲,跟我们说说 NVIDIA 的研究组织吧。它是怎么构建的?整体是如何运作的?
Speaker 1 | 1:00:32 - 1:00:56 Well, NVIDIA is not structured according to an org chart. We have one, but it's not actually the best way of understanding how we work. My team, for example, is not part of the official NVIDIA research team. My team is actually part of the organization that builds the GPU. And my team is not the only team building Neutron.
NVIDIA 并不是按照 org chart(组织架构图)来运作的。我们当然有 org chart,但它其实不是理解我们如何工作的最佳方式。比如说,我的团队就不属于官方的 NVIDIA research team。我的团队实际上属于负责构建 GPU 的组织。而且,构建 Neutron 的也不只是我的团队。
Speaker 1 | 1:00:56 - 1:01:53 There's probably 10 teams around the company that have significant involvement in building Neutron, in different parts of the company, in in enterprise software, in our AI software division, the part of NVIDIA that actually designs the GPU also significantly is involved in in building Neemotron. So there's there's so many different teams that that have to work together. We always like to say that the mission is the boss, rather than, the organization. But, what that implies is that people have to figure out how to work together, which is challenging, in the sense that humans are naturally tribal creatures. And it's, not natural for us to, be friendly with people we don't know very well or trust coworkers that we don't have, you know, success working with in the past.
公司里大概有 10 个团队深度参与了 Neutron 的构建,分布在公司的不同部门里,包括 enterprise software、我们的 AI software division,以及 NVIDIA 中实际负责设计 GPU 的那部分团队,也都在 Neutron 的构建中有很深的参与。所以有非常多团队必须协同工作。我们总喜欢说,真正的老板是 mission(使命),而不是组织本身。但这也意味着,人们必须想办法彼此协作,而这很有挑战性,因为 humans 天生就有部落属性。对我们来说,天然地去和不太熟悉的人建立友好关系,或者去信任那些过去没有一起成功合作过的同事,并不是一件自然的事。
Speaker 1 | 1:01:54 - 1:02:27 And, you know, actually, the name Nemotron, reflects that. We had, the Nemo team, which and the Megatron team, which was building, primarily focused on systems research for, for building large language models. And, you know, we decided to work together and then start calling our projects Nemo Tron reflecting, you know, sort of the the collaboration between these teams. Since then, Neumatron has has dramatically expanded. There's so many more teams that are part of the effort.
实际上,Nemotron 这个名字也反映了这一点。我们当时有 Nemo team,还有 Megatron team;后者主要专注于构建 large language models 的 systems research。后来我们决定一起合作,于是开始把项目叫作 Nemo Tron,以体现这些团队之间的协作关系。从那以后,Neumatron 已经大幅扩展了,参与这项工作的团队多了很多。
Speaker 1 | 1:02:28 - 1:02:50 And it's really important, that we have structured it, in this open way inside of NVIDIA. You know, we are inviting, volunteers from around the company to come help build NVIDIA's AI. We think it's very important to the future of the company. And, you know, as that vision continues to develop, more and more people wanna join. That's fantastic.
而且,我们以这种开放的方式在 NVIDIA 内部来组织这件事,真的非常重要。我们在邀请公司各处自愿加入的人来帮助构建 NVIDIA 的 AI。我们认为这对公司的未来至关重要。随着这一愿景不断发展,会有越来越多人想加入,这非常棒。
Speaker 1 | 1:02:50 - 1:03:24 We're really excited about that. And it means that we then have to figure out, how to organize the work so that everybody has a chance to contribute and feel heard and feel like their ideas are are, you know, fairly evaluated, on the on the path towards impact. We we have a a formal process for doing that. We have an internal website where people, share ideas and then those ideas, are assigned to one of 25 different, leads that are, you know, over various parts of building Neutron. They interact with those ideas.
我们对此真的很兴奋。这也意味着,我们必须想清楚如何组织工作,好让每个人都有机会贡献、都能感到自己的声音被听见,也能觉得自己的想法在通往实际影响的过程中得到了相对公平的评估。我们有一套正式流程来做这件事。我们有一个内部网站,人们会在上面分享想法,然后这些想法会被分配给 25 位不同的负责人之一,他们分别负责构建 Neutron 的不同部分,并会与这些想法进行互动。
Speaker 1 | 1:03:24 - 1:03:53 Some of those ideas get further developed. Some of those ideas get deferred until, you know, the next the next time we we go around building a new model. But we're trying to build Neutron in an open and inclusive way, so that, you know, we can really come together as a company to build it. I think, you know, organizations that figure out how to collaborate to build AI succeed. Organizations that struggle with control over who owns the AI tend to, waste a lot of effort.
其中一些想法会被进一步发展;另一些想法则会被暂缓,等到下一轮我们构建新 model(模型)时再处理。但我们正努力以一种开放且包容的方式来构建 Neutron,这样我们整个公司就真的能一起把它做出来。我认为,能够找到协作方式来构建 AI 的组织会成功;而那些在“AI 到底归谁控制”这个问题上挣扎的组织,往往会浪费很多精力。
Speaker 1 | 1:03:54 - 1:04:03 And so NVIDIA's success and Neutron's success, I think, is directly proportional to our ability to collaborate. It's something that I care deeply about.
所以我认为,NVIDIA 的成功以及 Neutron 的成功,都与我们的协作能力成正比。这是我非常在意的一件事。
Speaker 2 | 1:04:03 - 1:04:32 Fantastic. But you you mentioned earlier that despite the fact that you work at the number one undisputed leader in GPUs, you all as a research organization don't have all the GPUs that you would want in the world. So like how does the allocation of GPUs and compute happen? Is that based on how promising an idea is or early success? Do you, give GPUs, withdraw GPUs, based on success?
太棒了。不过你刚才提到,尽管你在 GPUs 领域公认、毫无争议的头号领导者这里工作,但作为一个研究组织,你们手上也并没有世界上你们想要的那么多 GPUs。所以,GPUs 和 compute(算力)的分配到底是怎么发生的?这是基于一个想法看起来有多有前景,还是基于早期成果?你们会根据成效来发放 GPUs,或者收回 GPUs 吗?
Speaker 1 | 1:04:33 - 1:05:03 It's a really complicated question, and it's it's obviously a a difficult problem for everyone in the industry to figure out how to allocate their their compute. Inside Nematron, you know, so we we have a budget, for Nematron. And inside Nematron, we allocate compute based on what we think the needs of the project are. We have a hierarchy. So we had a we have a set of programs and and in inside of each program, we had a have a set of projects.
这是个非常复杂的问题,而且很显然,如何分配 compute(算力)也是整个行业里每家公司都很难处理的问题。在 Nematron 内部,我们有属于 Nematron 的预算。而在 Nematron 内部,我们会根据我们认为各个项目的需求来分配算力。我们有一个层级结构。也就是说,我们有一组 program(项目方向),而在每个 program 里面,我们又有一组 project(具体项目)。
Speaker 1 | 1:05:04 - 1:05:39 And each of them put forward their requests. And then, you know, we have a two week cycle where we review requests and we review the budget and then we make decisions, in kind of a hierarchical way and then, you know, compute gets decided that way. Now having said that, this is something that I think we can still do better at. It's hard when we're making decisions about compute allocation because, every researcher is convinced that their idea, could change the world if it just got a thousand times more GPUs attached to it. Right?
每个项目都会提交自己的申请。然后我们会以两周为一个周期,审查这些申请、审查预算,再以一种层级化的方式做出决策,算力也就这样被确定下来。话虽如此,我认为这件事我们依然可以做得更好。在做算力分配决策时,这很难,因为每个研究人员都坚信:只要再给自己的想法配上一千倍的 GPUs,它就可能改变世界。对吧?
Speaker 1 | 1:05:39 - 1:05:50 And they they might be right. It might actually be true. And yet, we're running at the limit. We don't have a thousand x more GPUs for every idea that that we have. We we have to operate within, the limits that that we have.
而且他们也可能是对的。这可能确实是真的。但即便如此,我们也是在极限条件下运行。对于我们手里的每一个想法,我们都没有多出一千倍的 GPUs。我们必须在现有的限制之内运作。
Speaker 1 | 1:05:51 - 1:06:38 And so it is, a challenging process. We try to incorporate as many people's perspectives into that as possible so that it's as much as possible a shared sense of understanding, maybe not agreement. So there may be times when one project feels like it really deserved more GPUs because the impact of that would have been so high, but it didn't get it. We hope in that circumstance that they have an understanding of why some other project did get more GPUs and why that was considered more of a priority during this particular allocation round, for the company, so that people can at least, understand, you know, that there's a reason, for the allocations that we have. Having said that, you know, this process is always improving.
所以,这确实是一个很有挑战的过程。我们会尽可能把更多人的视角纳入其中,好让大家尽可能形成一种共同的理解感——也许不一定是共同认同。所以,有时某个项目会觉得自己理应获得更多 GPUs,因为那样带来的影响会非常大,但它最终没有拿到。我们希望在这种情况下,他们能够理解为什么另一个项目得到了更多 GPUs,以及为什么在这一轮面向整个公司的资源分配中,那件事被认为优先级更高。这样一来,人们至少能够明白,我们当前的这些分配背后是有理由的。话虽如此,这个流程也一直在持续改进。
Speaker 1 | 1:06:38 - 1:06:52 There's always more work to be done to make this more transparent and more fair. And and then, of course, my my number one is just to get more GPUs so that, you know, we we can also fund more things because I would like to do that too.
总还有更多工作要做,才能让这件事更透明、更公平。当然,我的头号目标就是弄到更多 GPU,这样你知道的,我们也能资助更多项目,因为我也很想这么做。
Speaker 2 | 1:06:52 - 1:06:58 How do you balance useful research with great exploratory research?
你怎么在有用的研究和优秀的探索性研究之间取得平衡?
Speaker 1 | 1:06:59 - 1:07:21 My belief is that research needs to be bootstrapped. Research is a chicken and egg problem. So, it is always the case that every researcher believes if I just had a lot more resources, my idea would change the world. Actually, it's important that researchers feel that way because if you didn't feel that way, you wouldn't have the conviction that's required to go do something crazy and new. Right?
我的看法是,研究需要靠 bootstrapping(自举)来启动。研究本身就是一个鸡生蛋、蛋生鸡的问题。所以情况总是这样:每个研究者都会相信,如果我只是有更多得多的资源,我的想法就会改变世界。其实,研究者有这种感觉很重要,因为如果你没有这种感觉,你就不会有那种做出疯狂而全新事情所必需的信念。对吧?
Speaker 1 | 1:07:21 - 1:07:38 So you have to believe. And and so, of course, you start with that belief. But then how do you translate that belief into something that other people can understand, right, that other people are willing to invest in? This is what I call the the chicken and egg problem. Right?
所以你必须相信。而且当然,一开始就是从这种信念出发。但接下来,你怎么把这种信念转化成其他人能够理解、也愿意为之投入的东西呢?这就是我所说的鸡生蛋、蛋生鸡问题。对吧?
Speaker 1 | 1:07:38 - 1:08:00 Because, like, once your research idea is is obviously good and impactful, it's easy to get resources, but how do you get it to be obviously good and impactful without those resources? Right? So so the way you solve chicken and egg problems is by bootstrapping. This is an iterative, problem solving approach where you do something small. You get some sort of signal about this is a good idea, and you tell people about that.
因为比如说,一旦你的研究想法显然是好的、而且有影响力的,拿到资源就很容易;但如果没有这些资源,你又怎么让它变得显然是好的、显然有影响力的呢?对吧?所以,解决这种鸡生蛋问题的方法就是 bootstrapping(自举)。这是一种迭代式的问题解决方法:你先做一点小的事情,获得某种信号,表明这是个好主意,然后你把这件事告诉别人。
Speaker 1 | 1:08:00 - 1:08:11 And then you ask for just a little bit more. And, if people saw like, oh, yeah. That, you know, that experiment turned out pretty well. That's pretty intriguing. We should probably do a little bit more there.
然后你再去争取多一点点资源。如果人们看到,哦,是啊,那个实验结果相当不错,挺有意思的。我们也许应该在这上面再多做一点。
Speaker 1 | 1:08:11 - 1:08:33 Then you're on track. Right? And, that that, over time, you know, iterate a lot, iterate quickly, iterate many times, you can bootstrap to, you know, finding significant resources for your idea and also usually attracting more people, to come along with it on the way because they have a chance to see that this idea is gonna change the world, and then they wanna be part of it.
那你就上路了,对吧?而且,随着时间推移,你会反复迭代、快速迭代、迭代很多次,你就能通过自举为自己的想法争取到可观的资源,同时通常也会吸引更多人一路加入,因为他们有机会看到这个想法将会改变世界,于是他们也想成为其中的一部分。
Speaker 2 | 1:08:33 - 1:08:48 Is that how the moonshots at NVIDIA got started as well over the years, whether that's in in AI or otherwise? So it it it was bottoms up, somebody coming up with a good idea versus Jensen saying this is what we need to do?
这些年来,NVIDIA 里的那些 moonshots 也是这样启动的吗,无论是在 AI 还是其他领域?也就是说,它是自下而上的——有人提出一个好想法——而不是 Jensen 说这就是我们需要做的事?
Speaker 1 | 1:08:48 - 1:09:08 Well, you know, Jensen has lots of good ideas too. And so the company is very responsive to to his ideas and that's, that's important as well. But Jensen very explicitly says all the time, this is a company of volunteers. You know, each of us is here because we choose to. We could we could be doing something else in our lives, but we we choose to be here.
嗯,你知道,Jensen 也有很多好想法。所以公司对他的想法反应非常迅速,这一点也很重要。但 Jensen 也一直说得非常明确:这是一家由 volunteers(自愿者)组成的公司。你知道,我们每个人在这里,都是因为我们自己选择来。我们本来可以去做人生里的其他事情,但我们选择待在这里。
Speaker 1 | 1:09:09 - 1:09:35 And so, you know, we we tend to make decisions, especially for early stage research. It it tends to be very bottoms up, because, you know, it's it's sort of an invitation. Like, bring bring your best ideas. Let's let's figure out, you know, what are all of our best ideas and then we'll take a step from there. Now do we sometimes have top down, ideas, that are important for the company strategy?
所以,你知道,我们做决策时,尤其是在早期研究阶段,往往是非常 bottoms up(自下而上)的,因为这某种程度上像是一种邀请:把你最好的想法带来。让我们一起弄清楚,我们所有人最好的想法分别是什么,然后再从那里往前走。当然,我们有时候也会有 top down(自上而下)的想法,而且这些想法对公司战略很重要,对吗?
Speaker 1 | 1:09:35 - 1:09:47 Of course. You know? Of course. NVFP four pre training is one of those, you know. So we decided as leadership of the company, we're gonna really invest in in FP four hardware.
当然会。你知道,当然会。NVFP four pre training 就是其中之一。所以公司管理层决定,我们要真正大力投资 FP four hardware。
Speaker 1 | 1:09:47 - 1:10:03 Now it's time to go invent some optimization algorithms that succeed in using it. And, so so we told the team. We didn't say to the team, you have to work on NVMe FP four pre training. What we said is, there's an opportunity. We're making a big investment.
现在就到了去发明一些 optimization algorithms(优化算法)的时候了,要让它们真正能够成功利用这个东西。所以,我们是这样告诉团队的。我们并没有对团队说,你们必须去做 NVMe FP four pre training。我们说的是,这里有一个机会。我们正在做一笔重大投资。
Speaker 1 | 1:10:03 - 1:10:15 And if we can figure this out, it will be significant for our company. And then we let the people who are interested in that work on that. And as a result, we succeeded. You know? So, so it is a balance of, like, bottoms up and tops down.
如果我们能把这件事搞定,它对我们公司会有重大意义。然后,我们就让那些对这项工作感兴趣的人去做。结果我们成功了。所以,这其实就是 bottoms up 和 top down 之间的一种平衡。
Speaker 1 | 1:10:17 - 1:10:42 But, it always has this boot strapping feeling even even with something like NVFP four where there's a significant, like, strategic top down component. The actual technical solution, which is very intricate and complex and has a lot of moving parts, that came, from the researchers themselves. And, you know, that's my belief is that research always comes from the researchers themselves. You can't tell research exactly to how to go solve a problem because then it wouldn't be research. It would be engineering.
但是,它始终都有一种 bootstrapping(自举)式的感觉,哪怕像 NVFP four 这样带有明显战略性 top down 成分的事情也是如此。真正的技术解决方案——而这件事本身非常精细、复杂,涉及很多相互联动的部分——是研究人员自己提出来的。而且,你知道,我一直相信,research(研究)永远来自研究者自身。你没法精确地告诉研究该怎样去解决一个问题,因为如果能那样做,它就不是 research 了,而会变成 engineering(工程)。
Speaker 1 | 1:10:42 - 1:10:53 But, in a world of, of AI where the most important problems we have to solve all have this research component, there needs to be freedom for researchers to innovate if we're gonna make progress.
但是,在 AI 的世界里,我们必须解决的那些最重要的问题,全都带有这种 research 成分;如果我们想取得进展,就必须给研究人员留出创新的自由。
Speaker 2 | 1:10:53 - 1:11:26 Listening to everything you're saying, I'm struck by how entrepreneurial the culture at NVIDIA still seems to be. So like, oh, look, I'm it's a very large company, I'm I'm sure there's all sorts of politics and you mentioned the tribal instincts. So like, I'm sure all of this is happening, but especially given you know, how long the company has been around a phenomenal success, the fact that people have been making a lot of money internally, it still seems to be very entrepreneurial, bottoms up driven, maybe meritocratic. Is that the right takeaway?
听你刚才这一切,我很明显地感受到,NVIDIA 的文化似乎依然非常 entrepreneurial(创业导向)。所以就像,当然,你看,这是一家非常大的公司,我也相信内部肯定有各种各样的 politics(办公室政治),而且你也提到了那种 tribal instincts(部落式本能)。所以,这些事情我肯定都在发生。但尤其是考虑到,你知道,这家公司已经存在这么久了,取得了惊人的成功,内部很多人也赚到了很多钱,它看起来仍然非常有创业气质、由 bottoms up 驱动,也许还带有 meritocratic(绩效导向/任人唯贤)的特点。这样的理解对吗?
Speaker 1 | 1:11:26 - 1:11:53 Yeah. I mean, one thing that's very unusual about NVIDIA is the tenure of its leadership. Jensen Huang has been running the company for thirty three years, but he's not alone. There are a lot of other very senior leaders in the company who have been there for three decades or longer, including my boss. And, these people remember what it feels like to work at a very small NVIDIA, and they know what it feels like to work at a very large NVIDIA.
Speaker 1 | 1:11:26 - 1:11:53 对。我是说,NVIDIA 有一件很不寻常的事,就是其领导层的任期之长。Jensen Huang 掌管这家公司已经三十三年了,但并不只有他一个。公司里还有很多非常资深的领导者已经在那里待了三十年甚至更久,包括我的老板。而且,这些人既记得在一个非常小的 NVIDIA 工作是什么感觉,也知道在一个非常大的 NVIDIA 工作是什么感觉。
Speaker 1 | 1:11:53 - 1:12:10 They have a shared sense of ownership for the company. You know, NVIDIA is a place we often say, no one fails alone. And the the the point of that, that's just a statement of fact. Right? You work at a company.
Speaker 1 | 1:11:53 - 1:12:10 他们对公司有一种共同的主人翁意识。你知道,在 NVIDIA,我们常说一句话:没有人会独自失败。这里的重点是,这其实只是一个事实陈述。对吧?你是在一家公司里工作。
Speaker 1 | 1:12:10 - 1:12:19 It's a one company. You all succeed together. You all fail together. You work in accelerated computing. Accelerated computing is the composition of thousands of technologies.
Speaker 1 | 1:12:10 - 1:12:19 这是一家公司。你们所有人一起成功,也一起失败。你做的是 accelerated computing(加速计算)。而 accelerated computing 是由成千上万项技术组合而成的。
Speaker 1 | 1:12:19 - 1:12:51 If any of them fail to deliver acceleration, the value is destroyed. It doesn't matter whether the chip is great. If the compiler sucks, at the end of the day, the thing that you're selling is time and capability to researchers that are trying to build the future of AI. And if they don't get that, it doesn't matter whether it was the, you know, the transistor or the math unit or the compiler or the library or the networking or anything else along the way that that failed to to live up to its, expectations. The whole thing in composition fails, the whole value is destroyed.
Speaker 1 | 1:12:19 - 1:12:51 只要其中任何一项没能实现加速,价值就会被摧毁。chip 再好也没有意义。如果 compiler 很烂,那么说到底,你卖给那些试图构建 AI 未来的研究人员的,是时间和能力。如果他们得不到这些,那么究竟是 transistor、math unit、compiler、library、networking,还是流程中其他任何环节没能达到预期,其实都不重要。只要整个组合失效,整体价值就全都被摧毁了。
Speaker 1 | 1:12:51 - 1:12:58 And so we have a deep understanding of that culturally at NVIDIA, and it is something that motivates the way that we work together.
Speaker 1 | 1:12:51 - 1:12:58 所以,在 NVIDIA,这一点已经深深植入我们的文化理解之中,也正是它推动了我们彼此协作的方式。
Speaker 2 | 1:12:58 - 1:13:25 Maybe to close the conversation, I'd I'd love to zoom out, get your take from, you know, the perspective of somebody who's, like, as deep into all of this as it as it gets about where things may be going. So, like, who knows in a few years, but I don't know. In the next year or two, maybe there's some visibility. I read somewhere that you're not necessarily a big, singularity, kind of a kind of person. Is that is that fair?
Speaker 2 | 1:12:58 - 1:13:25 也许作为这段对话的收尾,我想把视角拉高一点,从像你这样深度参与这一切的人角度,听听你对未来走向的看法。几年之后会怎样,当然谁也说不好,但我想在接下来一两年里,也许还是能看到一些趋势。我在某处读到过,你不一定算是那种很相信 singularity(奇点)的人。这样说公平吗?
Speaker 2 | 1:13:25 - 1:13:28 True. And why is that?
Speaker 2 | 1:13:25 - 1:13:28 是这样。为什么呢?
Speaker 1 | 1:13:29 - 1:13:58 Well, I think that intelligence is just so incredibly multifaceted. You know, I always think, about this question, like, if a company were to be looking for its next CEO, would it find the next CEO by looking for somebody who won the International Math Olympiad? Probably not. Right? Even though, like, it's incredible for people like, I could never even compete in any way at the International Math Olympiad.
Speaker 1 | 1:13:29 - 1:13:58 嗯,我认为 intelligence(智能)实在是一个极其多面向的东西。你知道,我总会想到这样一个问题:如果一家公司要寻找它的下一任 CEO,它会通过寻找一个赢过 International Math Olympiad 的人来找到下一任 CEO 吗?大概不会,对吧?尽管那确实非常了不起——像我这样的人,根本不可能以任何方式在 International Math Olympiad 里竞争。
Speaker 1 | 1:13:58 - 1:14:17 Those people are amazing. Right? They have just incredible brilliance. That's not the right kind of brilliance to run a company. If we look, for example, at, other aspects of our culture that that are really important, for example, musicians, What kind of intelligence does it take to become a hit musician?
那些人太了不起了,对吧?他们有着惊人的才华。但那并不是经营一家公司的那种合适的才华。比如说,如果我们看看文化中其他一些非常重要的领域,像是音乐家,要成为一个大热的音乐人,需要什么样的 intelligence(智力/智慧)?
Speaker 1 | 1:14:19 - 1:14:30 Don't assume that it's all luck. It's not. These people are working hard and they're very smart in ways that I might not understand with my PhD. Right? I might not have that kind of intelligence.
别以为那全靠运气。并不是。这些人非常努力,而且他们聪明的方式,可能连拥有 PhD 的我都未必理解。对吧?我自己可能就不具备那种 intelligence(智力/智慧)。
Speaker 1 | 1:14:31 - 1:14:47 And so when I think about intelligence, I think it's just so multifaceted and so contextual. You know, it really depends on the situation. It's not just about raw intelligence. Raw intelligence is kinda like the horsepower of an engine, but an engine running without wheels doesn't go anywhere. Right?
所以当我思考 intelligence(智力/智慧)时,我会觉得它其实是非常多面的,而且非常依赖语境。你知道,真的要看具体情境。它不只是 raw intelligence(原始智力)的问题。raw intelligence 有点像发动机的马力,但一台没有轮子的发动机,哪儿也去不了。对吧?
Speaker 1 | 1:14:47 - 1:15:31 So so intelligence, the impact of intelligence has a lot to do with, the context that the intelligence has put in the harness, the platform. And so when I think about that, I think, you know, the singularity is although it's an attractive idea, I think that it's it's a really a wrongheaded idea because it it doesn't really, take into account, these other factors. So I believe that artificial intelligence is gonna continue to develop at a rapid pace. It's gonna unlock significant capabilities, for people in every aspect of our world economy, people doing every kind of work. I'm very excited about the opportunities that it's that it's going to bring.
所以,intelligence(智力/智慧)的作用和影响,很大程度上取决于它所处的语境,取决于这种 intelligence 被套上了什么样的 harness(约束/框架)、被放在什么样的平台上。所以想到这里,我会觉得,singularity(奇点)虽然是个很有吸引力的想法,但我认为它其实是一个相当错误的想法,因为它并没有真正把这些其他因素考虑进去。所以我相信 artificial intelligence(人工智能)会继续快速发展。它会为世界经济中各个层面的人释放出重大的能力,帮助从事各种工作的每一个人。我对它将带来的机会非常兴奋。
Speaker 1 | 1:15:32 - 1:15:51 I am also a little bit, concerned with how we're gonna manage the transition. So I do think that transitions are hard for humans in general. Like, we're we're conservative generally. And, you know, there there is gonna be a lot of change. This is a profound change in the way that we think and the way that we work, the way that we learn.
但与此同时,我也有一点担心,我们要如何管理这场转变。所以我的确认为,转变对人类来说通常都很难。比如说,我们总体上是偏保守的。而且,你知道,接下来会有很多变化。这将深刻改变我们的思考方式、工作方式,以及学习方式。
Speaker 1 | 1:15:53 - 1:16:07 Ultimately, I have faith in our ability as humans to figure it out. You know, we've done it in the past. This is how this is who we are. We we build tools. We build external organs that help us solve problems.
但归根结底,我对我们人类解决问题的能力是有信心的。你知道,我们过去就做到过。这就是我们,这就是人类的本性。我们会制造工具。我们会制造 external organs(外部器官),来帮助我们解决问题。
Speaker 1 | 1:16:07 - 1:16:16 You know, we we have an external stomach. We call it kitchen. It creates enormous value for us. We can eat things that we couldn't eat without a kitchen. Right?
你知道,我们有一个 external stomach(外部胃),我们叫它 kitchen。它为我们创造了巨大的价值。有了 kitchen,我们就能吃下原本没有 kitchen 时吃不了的东西。对吧?
Speaker 1 | 1:16:16 - 1:16:31 Now we're creating an external brain. You know, the the implications of the external stomach were pretty profound for us as a species. They led to agriculture, which led to organized societies the way our cities are built. So we think about what is the implications of an external brain. Pretty profound.
现在我们正在创造一个 external brain(外部大脑)。你知道,external stomach 对我们这个物种的影响已经是非常深远的了。它带来了 agriculture,而 agriculture 又带来了有组织的社会,进而塑造了我们今天城市的建造方式。所以我们再想一想,external brain 会带来什么影响。那将会非常深远。
Speaker 1 | 1:16:32 - 1:17:06 Nobody actually really knows. But what I do believe in is the power of humanity to solve problems and to learn and to incorporate new technologies in ways that benefit us. I also believe that the problems we face as a planet all require more intelligence. Every single one of them, whether that's inequality, or climate, change or, you know, any of the other, structural, problems that that I think are very worrisome that we face. The solutions to those are gonna require invention and intelligence.
事实上,没有人真正知道答案。但我所相信的是,人类具备解决问题、学习,并以对我们有益的方式吸收新技术的能力。我也相信,作为一个星球,我们面临的所有问题都需要更多的 intelligence(智能)。每一个都是如此,无论是不平等、climate change(气候变化),还是你知道的其他那些我认为非常令人担忧的结构性问题。这些问题的解决方案都将需要发明与智能。
Speaker 1 | 1:17:07 - 1:17:50 And what that means for me is that the only kinds of tools that we can really create moving forward are going to be AI, because the problems that we face are all about intelligence. And regardless of the technological approach to solving those problems, the solutions will always be called AI. And so that makes me hopeful for the future, but also, you know, somewhat, you know, respectful of the challenge that it is going to bring to us as we we try to figure out how to live in a new way with this new external brain. But I believe in our our ability to to learn and to change, and and I I think, ultimately, this is gonna make our lives better.
这对我来说意味着,未来我们真正能够创造的工具类型只会是 AI,因为我们面临的问题本质上都与 intelligence 有关。而且,无论用什么技术路径去解决这些问题,最终这些解决方案都会被称为 AI。所以这让我对未来抱有希望,但同时,也会对它将给我们带来的挑战保持某种敬畏——因为我们要努力弄清楚,如何与这个新的外部大脑一起,以一种新的方式生活。但我相信我们学习和改变的能力,而且我认为,归根结底,这会让我们的生活变得更好。
Speaker 2 | 1:17:50 - 1:18:09 Do you guys feel the AI backlash that seems to be forming internally? Is that something that you all perceive, think about? And if so, do you think it's a communication problem that our industry may have, you know, in particular given what you just said about, all the obvious potential of AI?
你们是否感受到一种似乎正在内部形成的 AI 反弹情绪?这是你们都会察觉到、会去思考的事情吗?如果是的话,你们觉得这是不是一个沟通问题,尤其是考虑到你刚才所说的 AI 那些显而易见的潜力?
Speaker 1 | 1:18:09 - 1:18:41 You know, I'm always worried about the way that the public thinks about technology and interacts with it. It matters a lot. And it is definitely the case that societies that want technological advancement have more technological advancement than societies that that don't want change. So I think it is, actually important to think about it. One thing that's interesting about AI is that I believe it tends to be, much more accepted when it is, part of everyday life.
你知道,我一直都很在意公众如何看待技术、如何与技术互动。这一点非常重要。而且情况也确实是:希望技术进步的社会,会比那些不愿意改变的社会拥有更多技术进步。所以我认为,这件事实际上很值得认真思考。AI 有一个有趣之处在于,我相信,当它成为日常生活的一部分时,人们往往会更容易接受它。
Speaker 1 | 1:18:41 - 1:18:58 And then at that point, people stop thinking about it as AI. It's just, oh, this is the tool that I use. Like, do you care whether it's AI that's helping you, route your car when you ask the map application to help you drive somewhere? Like, I mean, it is. There is actually sophisticated AI that's going into that.
而到了那个时候,人们就不会再把它当作 AI 来看了。它只是“哦,这是我在用的工具”。比如说,当你让地图应用帮你规划开车路线时,你会在意帮助你导航的是不是 AI 吗?我的意思是,确实是。这里面实际上用到了相当复杂的 AI。
Speaker 1 | 1:19:00 - 1:19:09 But you're not really thinking about that. Right? You're just using a tool. And, so I feel like people's acceptance of AI, you know, comes with experience. Right?
但你其实并不会去想这个,对吧?你只是在使用一个工具。所以我觉得,人们对 AI 的接受,某种程度上是随着经验而来的。对吧?
Speaker 1 | 1:19:10 - 1:19:18 The more experience we have working with it, the more we learn how to work with it productively. I think the more comfortable we we become with it.
我认为,我们与它协作的经验越多,就越能学会如何高效地与它一起工作,也就会对它越感到自在。
Speaker 2 | 1:19:18 - 1:19:39 Great, Brian. So it's it's been a fascinating conversation. Maybe as a as a very last question to make sure we cover it, I wanna make sure that we talk about safety. What is the state of safety currently? And where does open source and closed source sort of fit in the safety conversation today?
很好,Brian。这一直是一场非常精彩的对话。也许作为最后一个问题,为了确保我们覆盖到这一点,我想确认我们谈谈安全。当前安全领域的现状是什么?以及 open source 和 closed source 如今在安全讨论中大致处于什么位置?
Speaker 1 | 1:19:39 - 1:20:37 Safety is on everybody's minds right now. You know, watching the Fable release and the, the way that the government interacted with that, I think, is a consequence of, concerns about safety, about these models, you know, they get stronger and stronger, and then they could be, misused. And, you know, there's different approaches to thinking about safety and and trying to define safety. I have maybe a a a slightly unorthodox opinion about this, which is that I think open technologies are generally safer because there's more sunlight. You know, when more people are thinking about the safety of a technology and evaluating it and then contributing to making it safer, I think that's inherently, safer than, having a small group of people, being in charge of safety for everyone else.
现在每个人都在关注安全。你知道,看着 Fable 的发布,以及我认为政府与之互动的方式,那是对安全担忧的结果:担心这些 model(模型)会变得越来越强,然后可能被滥用。而且,你知道,人们对安全有不同的思考方式,也会尝试以不同方式定义安全。对此我可能有一个稍微有点非主流的看法:我认为开放技术通常更安全,因为有更多“阳光”照进来。你知道,当更多人都在思考一项技术的安全性、评估它,并共同让它变得更安全时,我认为这从根本上说,比由一小群人替其他所有人负责安全,要更安全。
Speaker 1 | 1:20:37 - 1:21:06 I also think with, artificial intelligence, because it is really about ideas, it's really about, exploring ideas in different ways that diversity is more safe than monoculture. And what that means is that there's gonna be different beliefs. Like, diversity isn't just about, like, the easy stuff. Diversity is about the hard stuff. Like, when people have deeply felt disagreements, they really, really totally disagree with each other.
我也认为,就 artificial intelligence(人工智能)而言,因为它本质上真的关乎思想,关乎以不同方式探索思想,所以多样性比单一文化更安全。而这意味着,会存在不同的信念。比如,多样性不只是那些容易处理的东西。多样性关乎的是那些困难的东西。比如,当人们之间存在深刻而强烈的分歧时,他们彼此真的、真的会完全不同意对方。
Speaker 1 | 1:21:07 - 1:21:52 Making it possible for people to explore, their ideas in a diverse way, I think is more safe than, trying to create a walled garden where, you know, certain ideas are considered safe and certain ideas are are considered unsafe. And, you know, this is, controversial in today's AI, environment, which I think is interesting because we've had hundreds of years of tradition, that speak directly to this. You know, in in, United States, for example, we have, laws about freedom of conscious conscience and, freedom of speech. And, you know, it's not because we didn't consider for thousands of years, would it have been safer if we didn't have those. Right?
我认为,让人们能够以多样化的方式探索自己的想法,比试图建立一个“围墙花园”更安全;在那种环境里,你知道,某些想法被视为安全,某些想法则被视为不安全。而且,你知道,在今天的 AI 环境里,这个观点是有争议的;我觉得这很有意思,因为几百年来我们一直有直接讨论这个问题的传统。你知道,在 United States,举例来说,我们有关于良心自由和言论自由的法律。而且,你知道,这并不是因为几千年来我们没有思考过:如果没有这些自由,会不会更安全。对吧?
Speaker 1 | 1:21:52 - 1:22:19 We we tried that. We we tried actually having a monoculture about, like, these ideas are safe to talk about. These eyes ideas are safe to believe. And we found that to be much less safe than a pluralism where, we officially don't take a position about what ideas are safe. We actually found that as much safer as a society to, to support diversity than it is, to try to keep everybody safe, top down.
我们试过。我们确实试过建立一种单一文化,比如规定这些想法可以安全地讨论,这些想法可以安全地相信。结果我们发现,这比一种多元主义要不安全得多;在多元主义中,我们官方上并不对哪些想法是安全的采取立场。实际上,我们发现,对一个社会来说,支持多样性要安全得多,而不是自上而下地试图让每个人都“安全”。
Speaker 1 | 1:22:19 - 1:22:26 And so, I believe that open technologies for AI are inherently, the safest way of building AI.
所以,我相信,对 AI 而言,开放技术从根本上就是构建 AI 最安全的方式。
Speaker 2 | 1:22:26 - 1:22:33 Alright. Love it. Controversial take to close the conversation. Brian, it's been fabulous. Thank you so much.
好的,喜欢这个观点。用一个有争议的看法来结束这场对话。Brian,太精彩了。非常感谢你。
Speaker 2 | 1:22:33 - 1:22:35 Really appreciate your spending time with us today.
非常感谢你今天花时间和我们交流。
Speaker 1 | 1:22:35 - 1:22:36 Thanks for inviting me.
谢谢邀请我。
Speaker 2 | 1:22:37 - 1:22:55 Hi. It's Matt Turk again. Thanks for listening to this episode of the MAD podcast. If you enjoyed it, we'd be very grateful if you would consider subscribing if you haven't already or leaving a positive review or comment on whichever platform you're watching this or listening to this episode from. This really helps us build a podcast and get great guests.
你好,又是 Matt Turk。感谢你收听这一期 MAD podcast。如果你喜欢这期内容,而你还没有订阅的话,我们会非常感激你考虑订阅,或者在你观看或收听这一期节目的平台上留下积极的评价或评论。这对我们持续做好这个 podcast、邀请到优秀嘉宾真的很有帮助。
Speaker 2 | 1:22:55 - 1:22:57 Thanks and see you on the next episode.
谢谢,我们下期节目见。