投稿 视频

2026 年 AI 现状:大模型、编程、缩放律与 AGI

State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI | Lex Fridman Podcast #490

原始信息 · SOURCE State of AI in 2026: LLMs, Coding, Scaling Laws, China, Agents, GPUs, AGI | Lex Fridman Podcast #490

视频 作者 / 主持:Lex Fridman 来源:YouTube · Lex Fridman Podcast 发布: 时长:4 小时 39 分钟(4:39:02) 原文语言:英文 youtube.com

  • Nathan Lambert — Ai2 后训练负责人 / RLHF Book 作者 · 主页
  • Sebastian Raschka — 机器学习研究者 / Build a Large Language Model From Scratch 作者 · 主页
  • Lex Fridman — 主持人 · 主页
摘要 · SUMMARY

Lex Fridman 与 Ai2 后训练负责人 Nathan Lambert、从零实现大模型的作者 Sebastian Raschka 盘点 2026 年初的 AI 现状。2025 年 1 月 DeepSeek R1 之后,中国开权实验室从 DeepSeek 一家变成 GLM、MiniMax、Kimi K2 Thinking 等一群;Nathan 估计开权还会维持几年,因为美国大厂不愿订阅中国 API。消费端他赌 Gemini 继续追 ChatGPT,企业与代码仍是 Anthropic;OpenAI 擅长落地新范式(Deep Research、Sora、o1)。架构从 GPT-2 到 gpt-oss 仍是同一族:MoE、GQA、RMSNorm 都是微调。缩放律三条轴——预训练、可验证奖励强化学习(RLVR,词是 Tulu 3 团队先起的)、推理时缩放——他都看好;预训练集群 OLMo 3 约 200 万美元,服务才是十亿级。2026 年吉瓦级 Blackwell 集群是 2022–23 签的电和数据中心。 他们讲预训练 / 中训练 / 后训练、工具使用、持续学习与上下文学习、百万 token 上下文今年或到 200–500 万。996 约 72 小时。Nathan 说「一个模型统治一切」的梦有点在死,能力是锯齿状的;Claude 几乎能从零重建 Slack。收购:Groq 约 200 亿美元、Scale 近 300 亿、Manus 八个月 20 亿;达里奥不会卖。Llama 5 不会是开权。十年后记住的更可能是深度学习,不一定是 Transformer。

English summary

Lex Fridman sits with Ai2 post-training lead Nathan Lambert and Sebastian Raschka, author of Build a Large Language Model From Scratch, for a state-of-AI conversation at the start of 2026. After DeepSeek R1 in January 2025, China’s open-weight labs multiplied — GLM, MiniMax, Kimi K2 Thinking — and Nathan expects open weights to last a few more years because US tech will not pay Chinese APIs. On consumer chat he bets Gemini keeps closing on ChatGPT; enterprise and code stay Anthropic; OpenAI still lands new paradigms (Deep Research, Sora, o1). Architectures from GPT-2 to gpt-oss are the same family: MoE, GQA, RMSNorm are tweaks. He is bullish on three scaling axes — pre-training, RL with verifiable rewards (RLVR, a term from the Tulu 3 team), and inference-time scaling. OLMo 3’s cluster was about $2 million; serving is billions. 2026’s gigawatt Blackwell clusters were power and data-center contracts from 2022–23. They walk through pre- / mid- / post-training, tool use, continual vs in-context learning, and context windows at a million tokens now, maybe 2–5 million this year. 996 is about 72-hour weeks. Nathan says the one-model-to-rule-everything dream is kind of dying; capabilities are jagged; Claude could almost rebuild Slack from scratch. Deals: Groq about $20 billion, Scale almost $30 billion, Manus $2 billion in eight months; Dario will never sell. Llama 5 will not be open-weight. In a hundred years, deep learning is more likely to be remembered than the transformer.

时间轴 · 27 个章节
  1. 00:00 开场
  2. 01:39 赞助、评论与感想
  3. 16:29 中美谁赢 AI 竞赛
  4. 25:11 ChatGPT vs Claude vs Gemini vs Grok:谁在赢
  5. 36:11 最好的编程 AI
  6. 43:02 开源 vs 闭源大模型
  7. 54:41 Transformer:2019 年以来大模型怎么演化
  8. 1:02:38 AI 缩放律:死了还是还在撑
  9. 1:18:45 AI 怎么训:预训练、中训练、后训练
  10. 1:51:51 后训练:LLM 令人兴奋的新研究方向
  11. 2:12:43 给新手的建议
  12. 2:35:36 AI 工作文化(每周 72+ 小时)
  13. 2:39:22 硅谷泡沫
  14. 2:43:19 文本扩散模型
  15. 2:49:01 工具使用
  16. 2:53:17 持续学习
  17. 2:58:39 长上下文
  18. 3:04:54 机器人
  19. 3:14:04 通往 AGI 的时间线
  20. 3:21:20 AI 会取代程序员吗?
  21. 3:39:51 AGI 的梦在死吗?
  22. 3:46:40 AI 怎么赚钱?
  23. 3:51:02 2026 年的大收购
  24. 3:55:34 OpenAI、Anthropic、Google DeepMind、xAI、Meta 的未来
  25. 4:08:08 AI 曼哈顿计划
  26. 4:14:42 英伟达、GPU 和 AI 计算集群的未来
  27. 4:22:48 人类文明的未来

官方章节来自 Lex 节目页 / YouTube 描述(含赞助段)。对话文字稿时间戳不含片头赞助,约有 14 分半偏移;章节标题用 YouTube 时间。时长由最后一章 4:22:48 加上文字稿末段估得 4:39:02。yt-dlp 因 YouTube 429/机器人校验未能跑通。英语为清理过的口述。

00:00开场Introduction

Lex Fridman

下面这场对话讲人工智能的现状:过去一年令人兴奋的技术突破,以及我们觉得接下来一年可能发生的事。有时会非常技术,但我们尽量让圈外的人也能听进去,同时绝不把内容变浅。能和 AI 社区里我最喜欢的两个人做这期,是很大的荣幸:Sebastian Raschka 和 Nathan Lambert。他们都是受尊敬的机器学习研究者和工程师,同时也是很好的传播者、教育者、写作者和 X 发帖人。

The following is a conversation all about the state of the art in artificial intelligence, including some of the exciting technical breakthroughs and developments in AI that happened over the past year, and some of the interesting things we think might happen this upcoming year. At times it does get super technical, but we do try to make sure that it remains accessible to folks outside the field without ever dumbing it down. It is a great honor and pleasure to do this kind of episode with two of my favorite people in the AI community, Sebastian Raschka and Nathan Lambert. They are both widely respected machine learning researchers and engineers who also happen to be great communicators, educators, writers, and X posters.

Sebastian 写了两本我强烈推荐的书:《从零构建大语言模型》和《从零构建推理模型》。我真心相信,在机器学习和计算机科学里,最好的学习方式就是自己从零做出来。Nathan 是艾伦人工智能研究所的后训练负责人,也写了关于人类反馈强化学习的权威书。两人都有很好的 X 账号和 Substack。这是 Lex Fridman Podcast。亲爱的朋友们,Sebastian Raschka 和 Nathan Lambert。

Sebastian is the author of two books I highly recommend: Build a Large Language Model from Scratch, and Build a Reasoning Model from Scratch. I truly believe in machine learning and computer science, the best way to learn and understand something is to build it yourself from scratch. Nathan is the post-training lead at the Allen Institute for AI, and author of the definitive book on reinforcement learning from human feedback. Both of them have great X accounts, great Substacks. This is the Lex Fridman Podcast. And now, dear friends, here’s Sebastian Raschka and Nathan Lambert.

01:39赞助、评论与感想Sponsors, Comments, and Reflections

YouTube 此段为赞助朗读与主持评论,官方文字稿从对话正片开始,此处不逐句翻译广告。

16:29中美谁赢 AI 竞赛China vs US: Who wins the AI race?

Lex Fridman

一个有用的镜头是所谓 DeepSeek 时刻。大约一年前,2025 年 1 月,开权的中国公司 DeepSeek 放出 DeepSeek R1,可以说让所有人吃惊:接近或达到前沿表现,据称算力少得多、便宜得多。从那时到今天,研究和产品两边的 AI 竞争都疯了,一直在加速。国家层面谁在赢?是中国那一组公司,还是美国那一组?Sebastian?

One useful lens is the so-called DeepSeek moment. About a year ago in January 2025, the open-weight Chinese company DeepSeek released DeepSeek R1 that I think it’s fair to say surprised everyone with near or at state-of-the-art performance, with allegedly much less compute for much cheaper. From then to today, the AI competition has gotten insane, both on the research level and the product level. Who’s winning at the international level? The set of companies in China or the set of companies in the United States? Sebastian?

Sebastian Raschka

「赢」很宽。DeepSeek 肯定赢了做开权模型的人心。赢有多个时间尺度:今天、明年、十年。我确定的一件事是:2026 年,我不认为会有哪家公司独占别人没有的技术。研究者经常换工作、换实验室、轮转。差异化因素会是预算和硬件约束。想法不会专有,实施想法的资源会。我目前看不到赢家通吃。

Winning is a very broad term. DeepSeek is definitely winning the hearts of the people who work on open-weight models. Winning has multiple timescales: today, next year, ten years. One thing I know for sure is that I don’t think in 2026 there will be any company having access to a technology that no other company has. Researchers are frequently changing jobs, changing labs. They rotate. The differentiating factor will be budget and hardware constraints. I don’t think the ideas will be proprietary, but rather the resources needed to implement them. I don’t currently see a winner-takes-all scenario.

Nathan Lambert

录制此刻,Anthropic Claude Opus 4.5 的炒作已经完全疯了。我这几周用它做东西,几乎到了梗的程度。几个月前 Google 的 Gemini 3 放出,营销和 wow 很强。然后 11 月底 Opus 4.5 出来,炒作还在涨,Gemini 3 大家反而不怎么谈了——尽管它出来时人人说这是 Gemini 夺回 Google 结构性优势的时刻。Gemini 3 是很好的模型,我还在用,只是差异化变低了。想法空间很流动,但 Anthropic 文化上以狠赌代码出名,Claude Code 现在对他们管用。即使想法自由流动,很多东西卡在人力和组织文化。Anthropic 至少看起来最不混乱。

To demarcate the point in time when we’re recording this, the hype over Anthropic’s Claude Opus 4.5 has been absolutely insane. I’ve used it and built stuff in the last few weeks, and it’s almost gotten to the point where it feels like a bit of a meme. A few months ago Gemini 3 from Google got released, and the marketing and wow factor was super high. Then at the end of November Claude Opus 4.5 was released and the hype has been growing, while people don’t really talk about Gemini 3 as much, even though when it came out everybody was like, this is Gemini’s moment to retake Google’s structural advantages. Gemini 3 is a fantastic model, and I still use it. Differentiation is lower. The idea space is very fluid, but culturally Anthropic is known for betting very hard on code, and this Claude Code thing is working out for them. Even if the ideas flow pretty freely, so much of this is bottlenecked by human effort and the culture of organizations. Anthropic seems to be presenting as the least chaotic.

另一边,中国有很多不祥的技术,实验室远不止 DeepSeek。DeepSeek 在中国踢开了一场运动,类似 ChatGPT 在美国踢开聊天机器人运动。现在大量中国科技公司在放非常强的前沿开权模型。我会说 DeepSeek 正在失去中国开源模型第一的王冠。Z.ai 的 GLM、MiniMax、Moonshot 的 Kimi K2 Thinking,尤其近几个月更亮。新 DeepSeek 仍然很强,但 2025 年 DeepSeek 提供的平台,让更多中国公司用一种新的运作方式放出极好的模型。这些模型是开权的。顺着这条轨迹,美国公司正在做的商业模式可能有风险。目前很多人在美国为 AI 软件付钱;历史上在中国和世界其他地方,人不大为软件付钱。

On the other side, there’s a lot of ominous technology from China where there are way more labs than DeepSeek. DeepSeek kicked off a movement within China similar to how ChatGPT kicked off a movement in the US. There are now tons of tech companies in China releasing very strong frontier open-weight models, to the point where I would say DeepSeek is kind of losing its crown as the preeminent open model maker in China. Z.ai with their GLM models, MiniMax’s models, and Kimi K2 Thinking from Moonshot, especially in the last few months, have shone more brightly. The new DeepSeek models are still very strong, but that could be looked back on as a big narrative point: in 2025 DeepSeek provided this platform for way more Chinese companies releasing these fantastic models. These models are open weight, and depending on this trajectory, the business models American companies are doing could be at risk. Currently a lot of people are paying for AI software in the US, and historically in China and other parts of the world, people don’t pay a lot for software.

Lex Fridman

DeepSeek 这类模型因为开权而有人心。你觉得中国公司还会开权多久?

Some of these models like DeepSeek have the love of the people because they are open weight. How long do you think the Chinese companies keep releasing open-weight models?

Nathan Lambert

我会说几年。和美国一样,没有清晰的商业模式。这些中国公司很聪明,意识到同样的约束:很多美国顶级科技和 IT 公司出于安全顾虑不会给中国公司付 API。开权模型是他们影响、参与美国巨大且增长中的 AI 支出市场的方式。政府会看到这在国际上建立很多影响力,所以有很多激励继续。但建模型和做研究非常贵,某时点我会预期整合。我不预期那是 2026 年的故事;2026 年开源模型建造者会比 2025 年更多,很多重要的会在中国。

I would say for a few years. Like in the US, there’s not a clear business model. These Chinese companies are smart and realize the same constraints: a lot of top US tech companies and other IT companies won’t pay for an API subscription to Chinese companies for security concerns. Open-weight models are an ability to influence and take part in a huge growing AI expenditure market in the US. The government will see that is building a lot of influence internationally, so there’s going to be a lot of incentives to keep it going. Building these models is very expensive, so at some point I expect consolidation. I don’t expect that to be a story of 2026; there will be more open model builders throughout 2026 than in 2025. A lot of the notable ones will be in China.

Sebastian Raschka

DeepSeek 失去王冠,某种程度上是,但他们仍略微领先。不是 DeepSeek 变差了,是别人在用 DeepSeek 的想法。比如 Kimi,同一架构,他们在训。然后又是那种蛙跳:因为模型更新,某一刻可能更好一点。这回到没有明确赢家:一个人放出东西,另一个进来,最新的大概总是最好的。

DeepSeek losing its crown — to some extent yes, but they’re still slightly ahead. It’s not that DeepSeek got worse; the other ones are using the ideas from DeepSeek. Kimi, same architecture, they’re training it. Then this leapfrogging where they might be a bit better because they have the more recent model. There won’t be a clear winner. One person releases something, the other one comes in, and the most recent model is probably always the best model.

Nathan Lambert

中国公司激励也不同。DeepSeek 很保密。MiniMax、Z.ai 这类创业公司已经交了 IPO 材料,在争西方心智份额。DeepSeek 出名是对冲基金幻方量化建的,我们不完全知道他们用模型做什么、在不在乎这个。

Chinese companies have different incentives. DeepSeek is very secretive, whereas MiniMax and Z.ai have literally filed IPO paperwork and they’re trying to get Western mindshare. DeepSeek famously is built by a hedge fund, Highflyer Capital, and we don’t know exactly what they use the models for or if they care about this.

25:11ChatGPT vs Claude vs Gemini vs Grok:谁在赢ChatGPT vs Claude vs Gemini vs Grok: Who is winning?

Lex Fridman

你觉得哪个模型赢了 2025?哪个会赢 2026?

What model do you think won 2025, and what model do you think is going to win ’26?

Nathan Lambert

消费聊天机器人语境下,问题是:你愿意赌 Gemini 超过 ChatGPT 吗?直觉上有点险,因为 OpenAI 是在位者,科技里在位有很多好处。2025 年动量在 Gemini 这边,但他们起点很低。Bard 那些早期尝试 RIP。他们能顶着组织混乱把它做成,要给很大credit。但也很难赌输 OpenAI:他们看起来总是很乱,但很擅长把东西落地。

In the context of consumer chatbots, the question is: are you willing to bet on Gemini over ChatGPT? In my gut it feels like a bit of a risky bet because OpenAI has been the incumbent and there are so many benefits to that in tech. The momentum in 2025 was on Gemini’s side, but they were starting from such a low point. RIP Bard and those earlier attempts. Huge credit to them for powering through the organizational chaos. But it’s hard to bet against OpenAI because they always come off as so chaotic, but they’re very good at landing things.

我个人对 GPT-5 评价很杂,但它一定帮他们省了很多钱:高光功能是路由器,大多数用户不再那么烧 GPU。很难把我喜欢模型的哪些点和真正能区分大众的点拆开。2026 我会说一句有风险的话:Gemini 会继续追 ChatGPT。两边都在极端规模上运转时,Google 有规模,也能把研究和产品分得更好一点。OpenAI 你听到的是运营混乱、追逐高影响的东西,很创业公司。软件和企业这边,我觉得 Anthropic 会继续成功,他们一次次为此设置自己。

Personally I have very mixed reviews of GPT-5, but it must have saved them so much money with the high-line feature being a router where most users are no longer charging their GPU costs as much. It’s very hard to dissociate the things I like out of models versus the things that are actually going to be a general public differentiator. For 2026 I’ll say something even though it’s risky. I think Gemini will continue to make progress on ChatGPT. Google has the scale when both of these are operating at such extreme scales, and Google has the ability to separate research and product a bit better, whereas you hear so much about OpenAI being chaotic operationally and chasing the high-impact thing, which is a very startup culture. On the software and enterprise side, I think Anthropic will have continued success as they’ve again and again been set up for that.

Lex Fridman

基础设施上,你觉得 TPU 给他们优势吗?

In infrastructure, you think TPUs give them an advantage?

Nathan Lambert

很大程度上因为英伟达芯片的利润率疯狂,Google 可以从上到下开发整栈,不用付这笔利润,而且他们建数据中心有先手。高交付周期、高成本上很难的利润率,Google 有历史优势。如果会有新范式,最可能来自 OpenAI。他们的研究部门一次次展示能落地新研究想法或产品:Deep Research、Sora、o1 思考模型——这些定义性的东西都来自 OpenAI。这必须是他们作为组织的顶级特质之一。很难赌输这个。但我觉得今年很多会是规模,以及优化模型里可称为低垂果实的东西。

Largely because the margin on NVIDIA chips is insane and Google can develop everything from top to bottom to fit their stack and not have to pay this margin, and they’ve had a head start in building data centers. All of these things that have both high lead times and very hard margins on high costs, Google has a kind of historical advantage. If there’s going to be a new paradigm, it’s most likely to come from OpenAI. Their research division again and again has shown this ability to land a new research idea or a product. Deep Research, Sora, o1 thinking models — all these definitional things have come from OpenAI, and that’s got to be one of their top traits as an organization. It’s kind of hard to bet against that, but I think a lot of this year will be about scale and optimizing what could be described as low-hanging fruit in models.

Lex Fridman

智能和速度显然有权衡。GPT-5 在幕后想解决的就是这个:大众到底要智能还是要速度?

Clearly there’s a trade-off between intelligence and speed. This is what GPT-5 was trying to solve behind the scenes. Do people actually want intelligence, the broad public, or do they want speed?

Sebastian Raschka

我觉得有个开关其实很好。个人使用里,大多数时候查东西我用 ChatGPT 问一个快问题,要信息快。日常任务用快模型。现在自动模式已经挺好,不必专门说思考或不思考。有时我也要 pro:写完东西丢进去,彻底检查引用、想法、格式、图号。我不必马上要,可以吃晚饭让它跑。如果每个查询都等 30 分钟甚至 10 分钟,我会疯。

I think it’s a nice variety, the option to have a toggle. For my personal usage, most of the time when I look something up, I use ChatGPT to ask a quick question and get the information fast. For most daily tasks I use the quick model. Nowadays the auto mode is pretty good where you don’t have to specifically say thinking or non-thinking. Then again I also sometimes want the pro mode. When I have something written, I put it into ChatGPT and say do a very thorough check: are all my references correct, my thoughts, formatting, figure numbers. I don’t need that right away. I can finish my stuff, maybe have dinner, let it run. I would go crazy if for each query I had to wait 30 minutes, or even 10 minutes.

Nathan Lambert

我坐在这边快疯了,你居然用路由器和非思考模型。我怎么活。我重度用 ChatGPT 很久了,从没碰过 GPT-5 非思考。语气,还有更容易出错。有些是从 OpenAI 放出 o3 开始的,那是第一个做 Deep Research、找很多来源再整合的模型,我习惯了。为工作找任何信息——论文或代码引用——我只用 GPT-5.2 thinking 或 pro。我经常同时跑五个 pro 查询。

I’m sitting over here losing my mind that you use the router and the non-thinking model. How do you live with that? I’ve been heavily on ChatGPT for a while. I never touched GPT-5 non-thinking. Its tone and then its propensity for errors. Some of this is from when OpenAI released o3, the first model to do Deep Research and find many sources and integrate them. I became habituated with that. I will only use GPT-5.2 thinking or pro when I’m finding any sort of information query for work, whether that’s a paper or some code reference. I will regularly have five pro queries going simultaneously.

快的东西我用 Gemini,或者有时能 Google 的东西。它解释得好,我信它有那层背景知识。Gemini app 好了很多。代码和任何哲学讨论,我用 Claude Opus 4.5,也总是开扩展思考。扩展思考和推理时缩放只是让模型边际更聪明。进展很快时我总站那边,因为你不知道何时会解锁新用例。实时信息或在 AI Twitter 上挖我见过的东西,有时用 Grok。Grok 4 出来时,Grok 4 Heavy——他们的 pro 变体——其实很好,我印象很深,然后肌肉记忆还是打开 ChatGPT,就跟丢了。

I use Gemini for fast things or stuff that I could sometimes Google. It’s good at explaining things and I trust that it has this background of knowledge. The Gemini app has gotten a lot better. For code and any sort of philosophical discussion, I use Claude Opus 4.5, also always with extended thinking. Extended thinking and inference-time scaling is just a way to make the models marginally smarter. I will always edge on that side when the progress is very high because you don’t know when that’ll unlock a new use case. I sometimes use Grok for real-time information or finding something on AI Twitter that I knew I saw. When Grok 4 came out, Grok 4 Heavy — their pro variant — was actually very good and I was pretty impressed, and then I just kind of lost track of it with muscle memory from having the ChatGPT app open.

Lex Fridman

我确实用 Grok 4 Heavy 做调试。别的解决不了的硬核调试,我觉得它最好。你说 ChatGPT 是最好的界面;同样原因,对我来说 Gemini 更好——也许只是动量。我爱上了他们大海捞针的能力。塞进很多上下文、要找非常具体的信息、确保它全程跟上,我觉得 Gemini 最好。

I actually do use Grok 4 Heavy for debugging. For hardcore debugging that the other ones can’t solve, I find that it’s the best. You say ChatGPT is the best interface. For me, for that same reason — this could be just momentum — Gemini is the better interface. I fell in love with their needle-in-the-haystack capabilities. If I ever put in something that has a lot of context but I’m looking for very specific information to make sure it tracks all of it, I find Gemini has been the best.

36:11最好的编程 AIBest AI for coding

Lex Fridman

我们还没怎么提编程。很多人非常在乎。我基本上一半 Cursor、一半 Claude Code,因为体验根本不同,都有用。你们现在用什么?

We didn’t really mention programming. That’s another use case a lot of people deeply care about. I use basically half-and-half Cursor and Claude Code, because I find them to be fundamentally different experiences and both useful. What do you use? What’s the current vibe?

Sebastian Raschka

我用 VS Code 的 Codeium 插件。方便,就是插件,聊天界面能进仓库。我知道 Claude Code 不太一样,更智能体,碰更多东西,整个项目给你做。我还没舒服到那一步,也许我是控制狂,我仍想看见发生了什么。Codeium 现在是甜点:帮我,但不完全接管。

I use the Codeium plugin for VS Code. Very convenient. It’s just a plugin, then a chat interface that has access to your repository. Claude Code is a bit different. It is a bit more agentic. It touches more things; it does the whole project for you. I’m not quite there yet where I’m comfortable with that because maybe I’m a control freak, but I still like to see what’s going on. Codeium is the sweet spot for me right now where it is helping me, but it is not taking over completely.

Lex Fridman

我用 Claude Code 的一个原因是练用英语编程这个技能。体验根本不同。不是微观管理生成细节、看 diff——Cursor 里你可以——而是在设计空间里思考、在宏观上引导。那是另一种想编程过程的方式。而且 Claude Code 似乎更能吃满 Claude Opus 4.5。

One of the reasons I do use Claude Code is to build the skill of programming with English. The experience is fundamentally different. As opposed to micromanaging the details of the generation and looking at the diff — which you can in Cursor — you are understanding the code deeply as you progress, versus just thinking in this design space and guiding it at a macro level. That’s another way of thinking about the programming process. Also, Claude Code just seems to be a better utilization of Claude Opus 4.5.

Nathan Lambert

并排很好。你可以开 Claude Code、Cursor、VS Code,选同样的模型,问同样的问题。Claude Code 在那个领域好太多,惊人。Claude Code 让从零建东西变得有趣,你信它会做出东西。网站、刷新工具、数据分析。我博客上我们爬 Hugging Face,留每个数据集和模型随时间的下载数。Claude 就说「我用了那些数据,没问题」。那本来要我几天。我有足够情境意识去看趋势是否说得通、去核对。但那是一种美妙的界面:有个中介,你不必做维护不同网页项目那种可怕的底层活。

It’s a good side-by-side. You can have Claude Code open, Cursor open, VS Code open, select the same models, ask questions. Claude Code is way better in that domain. It’s remarkable. Claude Code makes it fun to build things, particularly from scratch where you trust that it’ll make something. Websites and refreshing tooling, which I use it for, or data analysis. On my blog we scrape Hugging Face so we keep the download numbers for every dataset and model over time. Claude was just like, “yeah, I’ve made use of that data, no problem.” That would’ve taken me days. I have enough situational awareness to be like, okay, these trends obviously make sense, and you can check things. But that’s just a wonderful interface where you can have an intermediary and not have to do the awful low-level work to maintain different web projects.

Sebastian Raschka

从零建 LLM 很好玩,也学很多。看图,图可能有错。看概念解释,你可能误解。但如果有代码而且代码能跑,你知道它对。没有误解,它精确,否则跑不起来。这就是编码的美。它不撒谎。它基本上是数学。数学书里你也可能有错永远发现不了,因为你读的时候不跑数学。代码可以验证。

Building an LLM from scratch is a lot of fun and a lot to learn. You can look at figures, but figures can have mistakes. You can look at conceptual explanations, but you might misunderstand them. But if there is code and the code works, you know it’s correct. There’s no misunderstanding; it’s precise. Otherwise it wouldn’t work. That’s the beauty behind coding. It doesn’t lie. It’s math, basically. Even with math you can have mistakes in a book you would never notice because you aren’t running the math while reading. With code you can verify it.

43:02开源 vs 闭源大模型Open Source vs Closed Source LLMs

Lex Fridman

刚谈了一堆闭权模型。开源的图景呢?哪些有意思、为什么?我们已经提了 DeepSeek。

We just talked about a bunch of the closed-weight models. Let’s talk about the open ones. Tell me about the landscape of open LLM models. Which are interesting? Which stand out and why? We already mentioned DeepSeek.

Nathan Lambert

看我们能不看笔记报出多少?DeepSeek、Kimi、MiniMax、Z.ai、Antlang。我们在报中国的。

Do you wanna see how many we can name off the top of our head? DeepSeek, Kimi, MiniMax, Z.ai, Antlang. We’re just going Chinese.

Sebastian Raschka

再加 Mistral AI、Gemma、OpenAI 的开源 gpt-oss。NVIDIA 有一个很酷的,Nemotron 3。年底东西特别多。Qwen 可能是那个——

Let’s throw in Mistral AI, Gemma, gpt-oss, the open source model by OpenAI. NVIDIA had a really cool one, Nemotron 3. There’s a lot of stuff, especially at the end of the year. Qwen might be the one—

Nathan Lambert

Qwen 是我正要说的那个明显名字。至少 10 个中国、至少 10 个西方。OpenAI 放出了他们 GPT-2 以来第一个开源模型。gpt-oss 其实很强,做一些别的模型不太会做的事。自私地推一下西方:美国和欧洲都有完全开放的模型。我在艾伦人工智能研究所,我们在建 OLMo,放数据、代码、所有这些。现在真正有人竞争,试图把一切都放出来,好让别人能训这些模型。Institute for Foundation Models / LM360 的 K2,瑞士研究联盟 Apertus,Hugging Face 的 SmolLM 很受欢迎。NVIDIA 的 Nemotron 也开始放数据。斯坦福 Marine Community Project,让人可以开 GitHub issue 实现新想法,然后在稳定的语言建模栈上跑。这份名单 2024 年小得多,当时大概就 AI2。

Qwen was the obvious name I was gonna say. You can get at least 10 Chinese and at least 10 Western. OpenAI released their first open model since GPT-2. gpt-oss is actually a very strong model and does some things that the other models don’t do very well. Selfishly I’ll promote Western companies; both the US and Europe have these fully open models. I work at the Allen Institute for AI where we’ve been building OLMo, which releases data and code and all of this. Now we have actual competition for people trying to release everything so other people can train these models. Institute for Foundation Models / LM360 has had their K2 models. Apertus is a Swiss research consortium. Hugging Face has SmolLM, which is very popular. NVIDIA’s Nemotron has started releasing data as well. Stanford’s Marine Community Project, a pipeline for people to open a GitHub issue and implement a new idea and run it in a stable language modeling stack. That list was way smaller in 2024. I think it was just AI2.

中国开源语言模型往往大得多,作为 MoE 峰值更高。我们很喜欢的很多——Gemma 或 Nemotron——美国这边往往更小,这开始在变。Mistral Large 3 出来了,巨大 MoE,12 月,架构很像 DeepSeek。创业公司 Reka AI,以及 Nemotron,都在逗远大于 1000 亿参数的 MoE,4000 亿参数量级,时间线是 2026 年一季度。今年人们用中国开源模型和美国开源模型分别做什么,这个平衡会变。我个人会很兴奋地看。

Chinese open language models tend to be much bigger and that gives them this higher peak performance as MoEs, whereas a lot of these things we like a lot, whether Gemma or Nemotron, have tended to be smaller models from the US, which is starting to change. Mistral Large 3 came out, a giant MoE, very similar to DeepSeek architecture, in December. A startup, Reka AI, and Nemotron have teased MoE models way bigger than 100 billion parameters, in the 400 billion parameter range, coming in this Q1 2026 timeline. This kind of balance is set to change this year in terms of what people are using the Chinese versus US open models for, which I’m personally going to be very excited to watch.

Lex Fridman

巨大的 props,能报这么多。你报 LLaMA 了吗?

Huge props for being able to name so many of these. Did you actually name LLaMA?

Nathan Lambert

没有。

No.

54:41Transformer:2019 年以来大模型怎么演化Transformers: Evolution of LLMs since 2019

Lex Fridman

也许该退一步讲讲 Transformer 架构本身。

It may be useful to step back and talk about transformer architecture in general.

Sebastian Raschka

也许从 GPT-2 架构开始,那个从「Attention Is All You Need」衍生的 Transformer。那篇论文的 Transformer 有两部分:编码器和解码器。GPT 只抓住解码器。它本质上仍是神经网络,里面有注意力机制。一次预测一个 token。过 embedding 层,过 Transformer block。block 里有注意力模块和全连接层,中间有些归一化。本质上就是带注意力机制的神经网络层。

Maybe we should start with GPT-2 architecture, the transformer derived from the “Attention Is All You Need” paper. That paper had a transformer with two parts: an encoder and a decoder. GPT went with just the decoder. It is essentially still a neural network and it has this attention mechanism inside. You predict one token at a time. You pass it through an embedding layer. There’s the transformer block. The transformer block has attention modules and a fully connected layer, and some normalization layers in between. Essentially neural network layers with this attention mechanism.

从 GPT-2 走到 gpt-oss-120b,比如有 Mixture of Experts 层。不是 GPT-OSS 发明的,几年前就有。本质是把模型做大,但每次前向不消耗更多算力。全连接层非常贵:一千输入一千输出就是一百万连接。想法是把它扩成多个前馈网络。比如 256 个,但不同时用。有一个路由器:基于这个输入 token,用这个全连接网络会有用。在那个语境里叫专家。Mixture of Experts 就是多个专家。输入偏数学,会用不同专家;英译西也许咨询不同专家。不是那么斩钉截铁「这只做数学、那只做西班牙语」,更模糊。想法是往网络里装更多知识,但不是所有知识每次都用。生成 token 时更有选择。路由器选择哪些 token 去哪个专家。更复杂,更难训,很多会崩、collapse。所以 OLMo 3 仍用 dense。行话里 dense 和 sparse 有区分:MoE 是 sparse,很多专家只有几个激活;dense 是只有一个全连接模块,总在用。

Coming from GPT-2 when we move on to gpt-oss-120b, there is, for example, the Mixture of Experts layer. It’s not invented by GPT-OSS; it’s a few years old. It’s a tweak to make the model larger without consuming more compute in each forward pass. The fully connected layer is very expensive. A thousand inputs and a thousand outputs is a million connections. The idea is to expand that into multiple feedforward networks. Instead of having one, let’s say you have 256, but you don’t use all of them at the same time. You now have a router that says, based on this input token, it would be useful to use this fully connected network. In that context it’s called an expert. Mixture of Experts means you have multiple experts. Depending on what your input is — more math-heavy — it would use different experts compared to translating English to Spanish. It’s not as clear-cut as this is only an expert for math and this for Spanish. It’s a bit more fuzzy. The idea is you pack more knowledge into the network, but not all the knowledge is used all the time. During token generation you are more selective. A router selects which tokens should go to which expert. It adds more complexity. It’s harder to train. A lot can go wrong, like collapse. That’s why OLMo 3 still uses dense. There’s jargon: dense versus sparse. Mixture of Experts is considered sparse because we have a lot of experts but only a few are active. Dense is the opposite: you only have one fully connected module, and it’s always utilized.

Lex Fridman

从 GPT-2 到今天,到底落实了多少新想法?这些架构真的有多不同?

Fundamentally, how many new ideas have been implemented from GPT-2 to today? How different really are these architectures?

Sebastian Raschka

比如 MoE。gpt-oss-120b 的注意力是 Group Query Attention,从 multi-head 到 GQA 是轻微调整。他们把 LayerNorm 换成 RMSNorm,只是另一种归一化,不是大变。非线性激活函数——熟深度网络的人知道,就像把 sigmoid 换成 ReLU。不从根本上改变网络,只是微调。差不多就这些。没有真正根本不同。仍是同一架构。你可以从一个加这些改动变成另一个。

Picture Mixture of Experts. The attention in gpt-oss-120b would be Group Query Attention. A slight tweak from multi-head attention to Group Query Attention. I think they replaced LayerNorm by RMSNorm, but it’s just a different normalization, not a big change. It’s just a tweak. The nonlinear activation function — for people familiar with deep neural networks, it’s the same as changing sigmoid with ReLU. It’s not changing the network fundamentally. That’s about it. It’s not really fundamentally that different. It’s still the same architecture. You can go from one into the other by just adding these changes.

我书里是 GPT-2 模型,因为简单、很小,大约 1.24 亿参数。奖励材料里我有从零做的 OLMo、从零做的 Gemini 3、其他从零模型。我总是从 GPT-2 模型开始,加不同组件,就从这到那。某种谱系。

In my book that’s a GPT-2 model because it’s simple and very small, 124 million parameters approximately. In the bonus materials I do have OLMo from scratch, Gemini 3 from scratch, and other from-scratch models. I always start with my GPT-2 model and just add different components and you get from one to the other. It’s kind of like a lineage.

Lex Fridman

能不能给人一个直觉:往外看 AI 世界突飞猛进,同时架构根本上没变。湍流、进展、收益到底发生在哪?

Can you build up an intuition? When you zoom out there’s so much rapid advancement in the AI world, and at the same time fundamentally the architectures have not changed. Where is all the turbulence, the turmoil of the advancement happening? Where are the gains to be had?

Sebastian Raschka

训练网络有不同阶段。预训练。当年 GPT-2 就只有预训练。现在有预训练、中训练、后训练。我觉得现在处于后训练焦点阶段。预训练如果你把更高质量数据放大,仍有优势。但有 GPT-2 没有的能力解锁。ChatGPT 基本上是 GPT-3 模型,GPT-3 架构上和 GPT-2 一样。新的是加了监督微调和人类反馈强化学习。更多在算法侧,而不是架构。

There are different stages where you develop or train the network. Pre-training. Back in the day it was just pre-training with GPT-2. Now you have pre-training, mid-training, and post-training. Right now we are in the post-training focus stage. Pre-training still gives you advantages if you scale it up to better, higher-quality data. Then we have capability unlocks that were not there with GPT-2. ChatGPT is basically a GPT-3 model, and GPT-3 is the same as GPT-2 in terms of architecture. What was new was adding supervised fine-tuning and reinforcement learning with human feedback. It’s more on the algorithmic side rather than the architecture.

Nathan Lambert

系统也变很多。听英伟达宣布,他们会说你现在做 FP8、现在可以做 FP4。实验室在搞清楚如何把更多算力塞进一个模型,训得更快、塞更多数据,更快找到更好配置。大规模训练时看的指标是每 GPU 每秒 token。打开 FP8 训练可以从比如 10K 到 13K,意味着每个参数用更少内存。少存信息、少通信,训得更快。

The systems also change a lot. If you listen to NVIDIA’s announcements, they talk about things like, you now do FP8, you can now do FP4. Labs are figuring out how to utilize more compute to put into one model, which lets them train faster and put more data in. You can find better configurations faster. You can look at tokens per second per GPU as a metric when you’re doing large-scale training. You can go from like 10K to 13K by turning on FP8 training, which means you’re using less memory per parameter. By saving less information, you do less communication and you can train faster.

1:02:38AI 缩放律:死了还是还在撑AI Scaling Laws: Are they dead or still holding?

Lex Fridman

大问题:我们讲了预训练背后的架构。缩放律在预训练、后训练、推理、上下文长度、数据、合成数据上还撑得住吗?

The big question: we talked about the architecture behind pre-training. Are the scaling laws holding strong across pre-training, post-training, inference, context size, data, and synthetic data?

Nathan Lambert

我想先从缩放律的技术定义开始。缩放律是幂律关系。x 轴——你在缩放的东西——可以想成算力和数据的组合,它们有点像;y 轴是对未见过的下一个 token 的 held-out 预测准确率。模型是自回归的:你留一套模型没见过的文本,训练后它会多准。人们发现那是非常可预测的关系,缩放律就来了。那个技术含义仍在继续,然后问题是用户从中得到什么。还有更多种缩放。OpenAI 的 o1 出名在引入推理时缩放。不那么出名的是:你也可以缩放强化学习训练,得到 log 的 x 轴、y 轴性能近似线性增加。

I’d like to start with the technical definition of a scaling law, which informs all of this. The scaling law is the power-law relationship between — you can think of the x-axis, what you are scaling, as a combination of compute and data, which are kind of similar — and then the y-axis is held-out prediction accuracy over next tokens. Models being autoregressive: if you keep a set of text the model has not seen, how accurate will it get when you train? Scaling laws came when people figured out that was a very predictable relationship. That technical term is continuing, and then the question is what do users get out of it. There are more types of scaling. OpenAI’s o1 was famous for introducing inference-time scaling. Less famously, also showing you can scale reinforcement learning training and get this log x-axis and then a linear increase in performance on the y-axis.

现在大概三条轴。传统缩放律讲预训练:模型多大、数据集多大。然后缩放强化学习:这种试错学习能做多久。然后推理时算力:让模型在具体问题上生成更多 token。我偏看好:它们都还真的在工作。但低垂果实大多被摘了,尤其过去一年在可验证奖励强化学习 RLVR,以及推理时缩放。所以这些模型用起来感觉那么不同:以前第一个 token 立刻来,现在它们会跑几秒、几分钟甚至几小时,生成这些隐藏思考,才给你答案第一个字。那全是推理时缩放,能力上几乎是阶跃。它启用了工具使用,启用了我们刚谈的好得多的软件工程。

There are kind of these three axes now. Traditional scaling laws are talked about for pre-training — how big your model is and how big your dataset is — then scaling reinforcement learning, how long can you do this trial-and-error learning, then inference-time compute, just letting the model generate more tokens on a specific problem. I’m kind of bullish; they’re all really still working, but the low-hanging fruit has mostly been taken, especially in the last year on Reinforcement Learning with Verifiable Rewards, RLVR, and then inference-time scaling. That’s why these models feel so different to use. Previously you would get that first token immediately. Now they’ll go off for seconds, minutes, or even hours generating these hidden thoughts before giving you the first word of your answer. That’s all about inference-time scaling, a wonderful kind of step function in how the models change abilities. They enabled this tool use stuff and enabled this much better software engineering.

我们说「启用」,几乎完全是因为 RLVR 训练让模型很容易捡起这些技能。你看推理过程里模型生成很多 token 时,它经常在做:试一个工具,看返回,再试另一个 API,看返回,看问题解决没。训练时模型很快学会这个。最终给了这种通用基础:模型能在你的仓库里很好地用 CLI、帮你处理 Git、挪东西、整理,或搜索找更多信息。一年前坐在这些椅子上,我们不太会想到模型做这些。今年发生了,彻底改变我们怎么想用 AI。非常神奇。但不清楚下一条解锁这种东西的路径是什么。持续学习后面会谈。AI 某些领域 buzz 很多,但没人知道下一次真正的阶跃何时来。

When we say enabled, almost entirely downstream of the fact that this Reinforcement Learning with Verifiable Rewards training just let the models pick up these skills very easily. If you look at the reasoning process when the models are generating a lot of tokens, what it’ll often be doing is: it tries a tool, looks at what it gets back, tries another API, sees what it gets back and if it solves the problem. The models, when you’re training them, very quickly learn to do this. That gives this kind of general foundation where the model can use CLI commands very nicely in your repo, handle Git, move things around, organize things, or search to find more information — which, if we were sitting in these chairs a year ago, is something we didn’t really think of the models doing. This has happened this year and has totally transformed how we think of using AI. Very magical. But it’s not clear what the next avenue will be in terms of unlocking stuff like this. We’ll get to continual learning later. There’s a lot of buzz around certain areas of AI, but no one knows when the next step function will really come.

Lex Fridman

你对每种缩放都看好。先从预训练:是不是在暗示预训练缩放的低垂果实已经被摘了?预训练触到平台了,还是你连预训练都看好?

You’re bullish on every version of scaling. Can we start at the beginning? Pre-training: are we implying the low-hanging fruit on pre-training scaling has been picked? Has pre-training hit a plateau, or are you still bullish even on pre-training?

Nathan Lambert

预训练已经极端昂贵。放大预训练也意味着你要给用户服务一个非常大的模型。松散地,GPT-4 和类似模型最大大约一万亿参数。有很多传闻说训练更高效之后它们其实变小了。你想让模型更小,因为服务成本按比例下降。训这些模型的成本相对把它们服务给几亿用户来说很低。DeepSeek 有个出名的数字:按云市场价格预训练大约五百万美元。OLMo 3 论文 2.4 节我们写了 GPU 集群为训练坐了多久——包括工程问题、多种子——租集群处理所有麻烦大约两百万美元。很多人能拿到一两千万去训一个模型,但服务百万用户的经常性成本真的是十亿级算力。一千张 GPU 租用你可以一天付十万。这些公司可能有几百万张 GPU。

Pre-training has gotten extremely expensive. To scale up pre-training is also implying you’re going to serve a very large model to the users. It’s been loosely established the likes of GPT-4 and similar models were around one trillion parameters at the biggest size. There’s a lot of rumors that they’ve actually gotten smaller as training has gotten more efficient. You want to make the model smaller because then your costs of serving go down proportionately. The cost of training these models is really low relative to the cost of serving them to hundreds of millions of users. DeepSeek had this famous number of about five million dollars for pre-training at cloud market rates. In the OLMo 3 paper, section 2.4, we detailed how long we had the GPU clusters sitting around for training — engineering issues, multiple seeds — about two million dollars to rent the cluster to deal with all the problems and headaches. A lot of people could get one to 10 million dollars to train a model, but the recurring costs of serving millions of users is really billions of dollars of compute. A thousand GPU rental you can pay 100 grand a day for. These companies could have millions of GPUs.

如果缩放真的给你更好的模型,财务上值不值?我觉得会慢慢往外推,随着 AI 解决更有说服力的任务——比如 Claude Opus 4.5 让 Claude Code 就是能干活。我 7 月启动了一个叫 ATOM 的项目,American Truly Open Models,那是一个真正 vibe-coded 的网站。过去几周回来刷新,Claude Opus 4.5 对比当时能用的模型,直接碾碎了 6、7 月建的时候那些问题。可能是更大的模型。里面变量很多,但仍有进展。

If scaling is actually giving you a better model, is it going to be financially worth it? We’ll slowly push it out as AI solves more compelling tasks — like Claude Opus 4.5 making Claude Code just work. I launched this project called the ATOM project, American Truly Open Models, in July, and that was like a true vibe-coded website. I came back to refresh it in the last few weeks and Claude Opus 4.5, versus whatever model was available at the time, just crushed all the issues it had from building in June and July. It might be a bigger model. A lot of things go into this, but there’s still progress coming.

Lex Fridman

你在讲缩放律 y 轴的细微处:体验到的和基准上的,实际智能可能不同。但关于预训练的直觉:如果把算力规模放大,模型会更好吗?不是财务可不可行,只问定律本身,模型会更聪明吗?

You’re speaking to the nuance of the y-axis of the scaling laws — the way it’s experienced versus on a benchmark, the actual intelligence might be different. Still, your intuition about pre-training: if you scale the size of compute, will the models get better? Not whether it’s financially viable, but from the law aspect, do you think the models will get smarter?

Nathan Lambert

会。这有时听起来像 AI 公司领导层几乎幻灭地说:「已经撑了 13 个数量级的算力,为什么会停?」根本上相当不可能停。只是最终我们甚至测不了更大尺度,因为更多算力带来的所有问题。有很多讨论说 2026 年是非常大的英伟达 Blackwell 计算集群——吉瓦级设施——上线的一年。这些全是 2022 和 2023 年、ChatGPT 之前或刚出来时签的电和数据中心合同。建这些更大集群训模型有两到三年交付周期,同时显然有巨大兴趣建比那更多的数据中心。关键是:新集群要来了。实验室会有更多训练算力。他们会用。但这不是给定的。我见过那么多进展,我预期它,也预期模型稍大一点。我更会说我们今年会看到 2000 美元订阅;我们已经看到 200 美元订阅。那可以再 10 倍。这些都下游于一个更大的模型,只多提供一点点刀刃。

Yeah. This sometimes comes off as almost disillusioned from leadership at AI companies saying this, but they’re like, “it’s held for 13 orders of magnitude of compute; why would it ever end?” Fundamentally it is pretty unlikely to stop. Eventually we’re not even going to be able to test the bigger scales because of all the problems that come with more compute. There’s a lot of talk on how 2026 is a year when very large NVIDIA Blackwell compute clusters — gigawatt-scale facilities — are coming online. These were all contracts for power and data centers signed and sought out in ’22 and 2023, before or right after ChatGPT. It took this two-to-three-year lead time to build these bigger clusters to train the models, while there’s obviously immense interest in building even more data centers than that. That’s the crux: these new clusters are coming. The labs are going to have more compute for training. They’re going to utilize this, but it’s not a given. I’ve seen so much progress that I expect it, and I expect a little bit bigger models. I’d say it’s more like we’ll see a $2,000 subscription this year; we’ve already seen $200 subscriptions. That could 10x again. These are the kind of things that could come — all downstream of a bigger model that offers just a little bit more of a cutting edge.

1:18:45AI 怎么训:预训练、中训练、后训练How AI is trained: Pre-training, Mid-training, and Post-training

Lex Fridman

也许这里适合定义预训练、中训练、后训练。

This might be a good place to define pre-training, mid-training, and post-training.

Sebastian Raschka

预训练是经典的一次预测下一个 token。一大堆语料。Nathan 因为 OLMo 3 大概也有很有意思的洞察,论文很大一部分聚焦正确的数据配比。预训练本质上就是在海量互联网数据、书、论文等上做交叉熵、下一个 token 预测。这些年变了一点:以前人能扔什么扔什么。现在不只是原始数据,也有合成数据,人会改写某些东西。合成数据不一定意味着纯 AI 编造。也可以拿维基百科文章,改写成问答,或总结、奖励、做成更好的数据。

Pre-training is the classic training, one next-token prediction at a time. You have a big corpus of data. Nathan probably also has very interesting insights because of OLMo 3. A big portion of the paper focuses on the right data mix. Pre-training is essentially just training across entropy loss, next-token prediction on a vast corpus of internet data, books, papers and so forth. It has changed a little over the years. People used to throw in everything they can. Now it’s not just raw data. It’s also synthetic data where people rephrase certain things. Synthetic data doesn’t necessarily mean purely AI-made-up data. It’s also taking something from a Wikipedia article and then rephrasing it as a Q&A, or summarizing it, rewarding it, making better data that way.

中训练……以前也叫预训练。我觉得叫中训练是因为只有预训练和后训练、中间没有,有点怪。中训练通常和预训练类似,但更专门。同一算法,但你聚焦比如长上下文文档。预训练不做是因为没有那么多长上下文文档。我们有一个专门阶段。LLM 仍是神经网络,有灾难性遗忘问题。你教它一些东西,它忘别的。不是 100% 忘,但没有免费午餐。人也一样:十年前的数学你问我,我不知道,得再看。

Mid-training used to be called pre-training. I think it’s called mid-training because it was awkward to have pre-training and post-training but nothing in the middle. What’s the actual training? Mid-training is usually similar to pre-training, but a bit more specialized. Same algorithm, but you focus, for example, on long-context documents. The reason you don’t do that during pre-training is you don’t have that many long-context documents. We have a specific phase. One problem of LLMs is still that it’s a neural network; it has catastrophic forgetting. You teach it something, it forgets other things. It’s not 100% forgetting, but there’s no free lunch. Same with humans. If you ask me some math I learned 10 years ago, I wouldn’t know; I would have to look at it again.

我不想拟人化 LLM,但数量不总是更好,要有选择。中训练是在末端对高质量内容有选择,所以 LLM 最后看到的是好东西。后训练是所有微调:监督微调、DPO、带人类反馈的 RLVR 等等。精炼阶段。成本也有意思:预训练现在花很多钱。RL 少一点。RL 其实不太教知识,更像解锁知识。更像技能学习:怎么用预训练里已有的知识解决问题。2025 年其实有三篇 RL 做预训练的论文。但我不觉得生产里有人那么干。

I don’t want to anthropomorphize LLMs, but quantity is not always better because it’s about being selective. Mid-training is being selective in terms of quality content at the end, so the last thing the LLM has seen is the quality stuff. Post-training is all the fine-tuning: supervised fine-tuning, DPO, RLVR with human feedback and so forth. The refinement stages. The cost thing: pre-training, you spend a lot of money on that right now. RL a bit less. RL, you don’t really teach it knowledge; it’s more like unlocking the knowledge. More like skill learning, how to solve problems with the knowledge it has from pre-training. There are actually three papers this year, or last year, 2025, on RL for pre-training. I don’t think anyone does that in production.

Nathan Lambert

目前是玩具例子。很多人觉得合成数据对训模型不好。你提到 DeepSeek 有一篇 OCR 论文。很多实验室都有;AI2 有一篇,别人有多篇。原因是网上有海量 PDF 和其他数字文档,格式不容易编码成文本。所以用 DeepSeek OCR 或我们叫 OLMo OCR 的东西,抽出可能数万亿 token 的候选数据。预训练数据集规模是万亿量级,按万亿 token 计。

Toy examples for now. A lot of people think of synthetic data as being bad for training models. You mentioned DeepSeek got an OCR paper. A lot of labs did; AI2 had one, others had multiple. The reason each of these labs has these is there’s vast amounts of PDFs and other digital documents on the web in formats that aren’t encoded with text easily. You use these, like DeepSeek OCR or what we called OLMo OCR, to extract what can be trillions of tokens of candidate data. Pre-training dataset size is on the order of trillions; it’s measured in trillions of tokens.

1:51:51后训练:LLM 令人兴奋的新研究方向Post-training explained

Nathan Lambert

2025 年最大的一件是学会这种可验证奖励的强化学习,RLVR。你可以放大那里的训练,意味着大量这种迭代的生成-打分循环,让模型在工具使用和软件侧学到有意思的行为:自己搜索、跑命令、看输出。那种训练也很好地启用推理时缩放。结果是这个范式连得很好:这种 RL 训练启用推理时缩放。但推理时缩放本可以用别的方式被发现。完美风暴:模型变了很多,训练方式是主因。这大幅改变了人做后训练的方法。

The biggest one from 2025 is learning this reinforcement learning with verifiable rewards, RLVR. You can scale up the training there, which means doing a lot of this iterative generate-grade loop, and that lets the models learn both interesting behaviors on the tool-use and software side. Searching, running commands on their own and seeing the outputs. That training enables this inference-time scaling very nicely. It just turned out this paradigm was very nicely linked, where this kind of RL training enables inference-time scaling. But inference-time scaling could have been found in different ways. Kind of this perfect storm where the models change a lot, and the way they’re trained is a major factor. This has changed how people approach post-training dramatically.

Lex Fridman

能描述一下 DeepSeek R1 带火的 RLVR 怎么工作吗?

Can you describe RLVR, popularized by DeepSeek R1? Can you describe how it works?

Nathan Lambert

有趣的事实:我当时在起 RLVR 这个词的团队上,来自 DeepSeek 之前我们的 Tulu 3 工作。我们不太居功说是把缩放 RL 做火的人,但学者能有的乐趣之一是命名和影响话语,因为闭源实验室只能说那么多。你可以在没有算力训模型的时候,把事情框成一个社区能围着 RLVR 这个词聚起来。然后 DeepSeek 是做出训练突破的人:他们缩放了强化学习。让模型生成答案,再给完成打分对不对,那个准确率就是强化学习的奖励。

Fun fact, I was on the team that came up with the term RLVR, which is from our Tulu 3 work before DeepSeek. We don’t take a lot of credit for being the people to popularize the scaling RL, but as much fun as academics get is the ability to name and influence the discourse, because the closed labs can only say so much. One of the things you can do as an academic is, while you might not have the compute to train the model, you can frame things in a way that a community can come together around this RLVR term. Then DeepSeek are the people that did the training breakthrough: they scaled the reinforcement learning. They have the model generate answers and then grade the completion if it was right, and that accuracy is your reward for reinforcement learning.

强化学习经典上是一个在环境里行动的智能体,环境给它状态和奖励,你试图最大化这个奖励。语言模型这里,奖励通常是一套可验证任务上的准确率,数学题或编码任务。到事实领域开始模糊,某种意义上也可验证,或对指令的约束,比如「只用以 A 开头的词回答」。这些东西某种方式都可验证。核心想法是找更多可验证的问题,让模型试很多次,同时做这些 RL 梯度更新。基础设施从人类反馈强化学习 RLHF 演化而来,那个时代他们优化的分数是聚合人类偏好的学得奖励模型。你换了问题域,让优化能走到大得多的尺度,这启动了模型能做什么、人怎么用它们的一次重大变化。

Reinforcement learning is classically an agent that acts in an environment, and the environment gives it a state and a reward back, and you try to maximize this reward. In the case of language models, the reward is normally accuracy on a set of verifiable tasks, whether math problems or coding tasks. It starts to get blurry with things like factual domains. That is also in some ways verifiable, or constraints on your instruction, like respond only with words that start with A. All of these things are verifiable in some way. The core idea is you find a lot more of these problems that are verifiable and you let the model try it many times while taking these RL gradient updates. The infrastructure evolved from reinforcement learning from human feedback, RLHF, where in that era the score they were trying to optimize was a learned reward model of aggregate human preferences. You kind of changed the problem domains and that let the optimization go on to much bigger scales, which kind of kickstarted a major change in what the models can do and how people use them.

Lex Fridman

RLVR 适合哪些领域?

What kind of domains is RLVR amenable to?

Nathan Lambert

数学和代码是出名的。然后很多工作在所谓 rubrics 上,跟人听过的 LLM-as-a-judge 相关。对训练集里每个问题,我会再找另一个语言模型问:「这个问题的好答案会是什么样?」然后你可以一遍遍试,按这个 rubric 打分。那不一定像数学和代码那样可验证,但这个 rubrics 想法,以及其他更含糊的科学问题,是很多注意力所在。他们想把这套方法推进更开放的领域,让模型能学更多。

Math and code are the famous ones, and then there’s a lot of work on what is called the rubrics, related to a word people might have heard, LLM-as-a-judge. For each problem in my training dataset, I will then have another language model and ask it, what would a good answer to this problem look like? Then you could try the problem a bunch of times over and over and assign a score based on this rubric. That’s not necessarily verifiable like a math and code domain, but this rubrics idea and other scientific problems where it might be a little more vague is where a lot of the attention is. They’re trying to push this set of methods into these more open-ended domains so the models can learn a lot more.

Sebastian Raschka

RLVR 漂亮的地方是:你问 LLM 一道数学题,你知道正确答案,让 LLM 自己想办法,你不怎么约束它。可以加一些约束,比如用同一种语言、别西英来回切。但基本上放手。你只给问题和答案,LLM 的任务是到达正确答案。实践里漂亮的是:LLM 会一步步描述,像学生或数学家推导解。它用那些步骤,帮自己提高准确率。然后你说的推理缩放。推理缩放松散意味着使用 LLM 时花更多算力,这里就是模型用更多 token。DeepSeek R1 论文里,他们训得越久,回复越长。

The interesting, beautiful thing here is you ask the LLM a math question, you know the correct answer, and you let the LLM figure it out, but how it does it — you don’t really constrain it much. Some constraints you can add, like use the same language or don’t switch between Spanish and English. Let’s say you’re pretty much hands-off. You only give the question and the answer, and then the LLM has the task to arrive at the right answer. The beautiful thing in practice: the LLM will do a step-by-step description, like how a student or a mathematician would derive the solution. It will use those steps and that helps the model improve its own accuracy. Then inference scaling. Inference scaling loosely means spending more compute while using the LLM during inference, and here the inference scaling is that the model would use more tokens. In the DeepSeek R1 paper, they showed the longer they train the model, the longer the responses are.

2:12:43给新手的建议Advice for beginners

Lex Fridman

想岔开一点谈教育和学习。如果有聪明人在听,对编程和 AI 感兴趣,我猜从零做是好的开始。你会建议人做什么?

I was wondering if we could take a bit of a tangent and talk about education and learning. If you’re somebody listening who’s a smart person interested in programming and interested in AI, I presume building something from scratch is a good beginning. What would you recommend people do?

Sebastian Raschka

我会亲自从实现一个能在自己电脑上跑的简单模型开始。从零建模型的目标不是做出你每天个人项目用的东西。它不会成为取代现有开权模型或 ChatGPT 的私人助手。是为了精确看见什么进 LLM、什么出 LLM、预训练在你自己电脑上怎么工作。然后你学预训练、监督微调、注意力机制。

I would personally start by implementing a simple model from scratch that you can run on your computer. The goal of building a model from scratch is not to have something you use every day for your personal projects. It’s not going to be your personal assistant replacing an existing open-weight model or ChatGPT. It’s to see exactly what goes into the LLM, what exactly comes out, and how pre-training works on your own computer. Then you learn about pre-training, supervised fine-tuning, and the attention mechanism.

到某时你会碰到极限,因为小模型只能做那么多。学大规模 LLM 的问题是,做更大模型指数级更复杂:不只是模型变大。你得想参数怎么在多 GPU 上分片。KV cache 都有多种实现。一种是理解它怎么工作,像逐步 concatenate 列表增长的缓存,但在 GPU 上不最优。你会预分配 tensor 再填。那又加 20 或 30 行代码。每件事你都加这么多代码。书的诀窍基本上是理解 LLM 怎么工作。它不会是生产级 LLM,但一旦有了,你能理解生产级的。

At some point you will reach a limit because smaller models can only do so much. The problem with learning about LLMs at scale is it’s exponentially more complex to make a larger model because it’s not just that the model becomes larger. You have to think about sharding your parameters across multiple GPUs. Even for the KV cache, there are multiple ways you can implement it. One is just to understand how it works, like a cache you grow step-by-step by concatenating lists, but that wouldn’t be optimal on GPUs. You would pre-allocate a tensor and then fill it in. That adds another 20 or 30 lines of code. For each thing, you add so much code. The trick with the book is to understand how the LLM works. It’s not going to be your production-level LLM, but once you have that, you can understand the production-level LLM.

大多数例子放一张 GPU。有些 MoE 奖励材料可能要多 GPU,但目标是一张 GPU。漂亮的是你可以自我验证。几乎像 RLVR。从零写这些时,你可以拿 Hugging Face Transformers 库里的现有模型。Transformers 库很棒,但要学 LLM,我觉得那不是最好的起点,因为代码太复杂。它要适配太多用例,有人在生产用,必须非常精巧,所以缠在一起,不好线性读。所有有开权模型的前沿实验室都有 Transformers 版本,从 DeepSeek 到 gpt-oss。那是加载它们的规范方式。但 Transformers 库也不是生产推理用的。人用 SGLang 或 vLLM,又加一层复杂度。

Most of the examples I have fit on one GPU. I have some bonus materials on some MoE models; one or two may require multiple GPUs, but the goal is to have it on one GPU. The beautiful thing is you can self-verify. It’s almost like RLVR. When you code these from scratch, you can take an existing model from the Hugging Face Transformers library. The Hugging Face Transformers library is great, but if you want to learn about LLMs, that’s not the best place to start because the code is so complex. It has to fit so many use cases and some people use it in production. It’s intertwined and hard; it’s not linear to read. All frontier labs that have open-weight models have a Hugging Face Transformers version, from DeepSeek to gpt-oss. That’s the canonical way you can load them. But even the Transformers library is not used in production for inference. People use SGLang or vLLM, and it adds another layer of complexity.

我会建议:如果想理解比如 OLMo 3 怎么实现,看模型 hub 里的权重和 config 文件。你看见他们用了多少层、用 group query attention。然后在大约 100 行人类可读的 config 里看见所有组件。从你的 GPT-2 模型开始加这些。酷的是你可以加载预训练权重,看它们在你的模型里是否工作。你想匹配 Transformers 模型的同样输出,然后基本上当可验证奖励,让架构正确。有时要我一天。OLMo 3 的挑战是位置编码的 RoPE;他们有 YaRN 扩展和一些自定义缩放。一开始我匹配不上,但在这种挣扎里你理解东西。最后你知道它对,因为你可以对照参考实现做单元测试。我觉得那是最好的学习方式之一。基本上逆向工程。

What I would recommend: if I want to understand how OLMo 3 is implemented, I would look at the weights in the model hub and the config file. You can see they used so many layers, they use group query attention. Then you see all the components in a human-readable 100-line config file. Then you start with your GPT-2 model and add these things. The cool thing is you can then load the pre-trained weights and see if they work in your model. You want to match the same output you get with a Transformers model, and then you can use that as a verifiable reward to make your architecture correct. Sometimes it takes me a day. With OLMo 3, the challenge was RoPE for the position embeddings; they had a YaRN extension and some custom scaling. I couldn’t quite match it at first, but in this struggle you kind of understand things. At the end you know you have it correct because you can unit test it against the reference implementation. That’s one of the best ways to learn. Basically you reverse-engineer something.

Nathan Lambert

我觉得今天每个想进 AI 的人都该做这个,所以我喜欢你的书。我从 RL 和机器人来到语言模型,从来没花时间把基础都学一遍。Transformer 架构今天就像过去深度学习那么根本,人需要学它。很多人被压垮的是:怎么把这个应用到产生影响或找到职业路径。

That is something everyone interested in getting into AI today should do, and that’s why I liked your book. I came to language models from the RL and robotics field, so I never had taken the time to just learn all the fundamentals. The Transformer architecture is as fundamental today as deep learning was in the past, and people need to learn it. Where a lot of people get overwhelmed is how to apply this to have an impact or find a career path.

2:35:36AI 工作文化(每周 72+ 小时)Work culture in AI (72+ hour weeks)

Lex Fridman

能描述一下 996 这种文化吗?我相信可以说它发明在中国、被硅谷采纳。996 是什么?早上 9 点到晚上 9 点——

Can you describe 9/9/6 as a culture? I believe you could say it was invented in China and adopted in Silicon Valley. What’s 9/9/6? It’s 9:00 AM to 9:00 PM—

Sebastian Raschka

一周六天。

six days a week.

Lex Fridman

一周六天。那是 72 小时?这基本上是硅谷 AI 公司的标准吗?越来越这种 grind 心态。

Six days a week. What is that, 72 hours? Is this basically the standard in AI companies in Silicon Valley? More and more this kind of grind mindset.

Sebastian Raschka

也许不是精确那样,但有往那边的趋势。有意思的是几乎翻过来了:我在学术界时觉得那样,因为当教授你要写基金、教书、做研究,三份工作合一,想成功就超过全职。现在像 Nathan 刚说的,教授相比实验室,压力或工作量甚至更少,因为——

Maybe not exactly like that, but I think there is a trend towards it. It’s interesting — it almost flipped, because when I was in academia I felt like that, because as a professor you had to write grants, teach, and do your research. Three jobs in one, more than a full-time job if you want to be successful. Now, like Nathan just said, the professors in comparison to a lab have even less pressure or workload than at a frontier lab because—

Nathan Lambert

我觉得他们干活很多。他们只是非常满足。跟学生一起工作,有持续的导师跑道,使命非常以人为本。事情走得很快、很混乱的时代,这对人很有回报。

I think they work a lot. They’re just so fulfilled. Working with students, having a constant runway of mentorship and a mission that is very people-oriented. In an era when things are moving very fast and are very chaotic, it’s very rewarding to people.

Sebastian Raschka

创业公司有那种压力。你必须成。人投入时间真的重要,但真的很难,因为你必须不断交付。我待过创业公司。我过得不错,但不知道能不能永远这样。节奏有意思,正是我们开头谈的:这些模型在互相蛙跳,不断试图相对对手再走下一步。现在就是无情。

At a startup there’s this pressure. You have to make it. It is really important that people put in the time, but it is really hard because you have to deliver constantly. I’ve been at a startup. I had a good time, but I don’t know if I could do it forever. It’s an interesting pace and exactly like we talked about in the beginning. These models are leapfrogging each other, and they are just constantly trying to take the next step compared to their competitors. It’s just ruthless right now.

Nathan Lambert

蛙跳和多个玩家其实是语言模型进展一个被低估的驱动:竞争深深嵌进去。这些公司有意创造很强的文化。比如 Anthropic 以文化上深度投入、有组织出名。我们很少听到他们,Anthropic 每个人似乎很对齐。处在超紧的文化里,再有这种竞争动态,会让你拼命干、做出更好的东西。但代价是人力资本。你只能这样一段时间,人肯定在 burnout。我写过一篇 burnout,自己进进出出,尤其一边全模式训练一边当管理者。疯了的工作。《Apple in China》里 Patrick McGee 讲苹果工程师为在中国建供应链有多拼。他提到他们有「拯救婚姻」项目,播客里说人因为这种强度死过。这是基于人的消耗创造进展的完美环境。人的消耗就是我们开头的 996,人真的 grind。

This leapfrogging nature and having multiple players is actually an underrated driver of language modeling progress where competition is so deeply ingrained. These companies have intentionally created very strong cultures. Anthropic is known to be culturally deeply committed and organized. We hear so little from them, and everybody at Anthropic seems very aligned. Being in a culture that is super tight and having this competitive dynamic is a thing that’s going to make you work hard and create things that are better. But that comes at the cost of human capital. You can only do this for so long, and people are definitely burning out. I wrote a post on burnout as I’ve tread in and out of this myself, especially trying to be a manager while doing full-mode training. It’s a crazy job. In the book Apple in China, Patrick McGee talked about how hard the Apple engineers worked to set up the supply chains in China. He mentioned they had “saving marriage” programs, and he said in a podcast that people died from this level of working hard. It’s a perfect environment for creating progress based on human expense. The human expense is the 996 we started this with, where people do really grind.

Sebastian Raschka

我也读了这本书。我觉得他们有个暗号:如果有人必须回家陪家人拯救婚姻。同事说「这是这种情况的红色警报,这周末得让那人回家」。但同时我不觉得他们是被强迫干活。他们对产品那么有激情,你会进入那种心态。我当学者、当独立的人时也有过。我过劳,不健康。我有过背部、颈部问题,因为没有该有的休息。但不是有人逼我;是因为我想干,因为东西令人兴奋。

I also read this book. I think they had a code word for if someone had to go home to spend time with their family to save the marriage. Then the colleagues said, okay, this is red alert for this situation. We have to let that person go home this weekend. At the same time I don’t think they were forced to work. They were so passionate about the product that you get into that mindset. I had that sometimes as an academic, and as an independent person. I overwork, and it’s unhealthy. I had back issues and neck issues because I did not take the breaks I should have. But it’s not because anyone forced me; it’s because I wanted to work because it’s exciting stuff.

Nathan Lambert

OpenAI 和 Anthropic 就是那样。他们想做这份工作。

That’s what OpenAI and Anthropic are like. They want to do this work.

2:39:22硅谷泡沫Silicon Valley bubble

Lex Fridman

也有一种热度在积,尤其在硅谷,跟缩放律想法对齐。有一种炒作:世界会在几周尺度上被改造,你想站在中心。我有幸和各种各样的人谈话,看见世界各地这些泡沫和回音室。人怎么形成它们很迷人。公平地说硅谷是一种回音室、一种筒仓和泡沫。我觉得泡沫其实很有用、很有效。不一定负面,因为你可以极度高产。可能是乔布斯那种现实扭曲场:你们互相说服突破就在眼前,通过互相说服,让突破真的就在眼前。

There’s also a feeling of fervor that’s building, especially in Silicon Valley, aligned with the scaling laws idea. There’s this hype where the world will be transformed in a scale of weeks and you want to be at the center of it. I have the great fortune of having conversations with a wide variety of human beings, and I get to see all these bubbles and echo chambers across the world. It’s fascinating to see how we humans form them. I think it’s fair to say Silicon Valley is a kind of echo chamber, a kind of silo and bubble. I think bubbles are actually really useful and effective. It’s not necessarily a negative thing because you can be ultra-productive. It could be the Steve Jobs reality distortion field, because you just convince each other the breakthroughs are imminent, and by convincing each other of that, you make the breakthroughs imminent.

Nathan Lambert

Byrne Hobart 写过一本书给泡沫分类。一种是金融泡沫,涉及投机,是坏的;另一种有效地是为了建设,因为它推人去建。我觉得 AI 处在这个里,但我担心它转成金融泡沫。

Byrne Hobart wrote a book classifying bubbles. One of them is financial bubbles, which involve speculation and are bad, and the other is effectively for build-outs, because it pushes people to build. I do think AI is in this, but I worry about it transitioning to a financial bubble.

Lex Fridman

但在想法空间里,那个泡沫创造现实扭曲场。意味着你在偏离现实。如果走太远,同时还 996,你可能错过人类经验的一些根本方面。这是硅谷的常见问题。非常特定的地理区域。你可能不理解中西部视角,或不理解美国和世界其他不同人的经验。你们用某种方式和彼此说话、说服彼此某件事,那会让你真正陷入麻烦。无论 AI 大成功变成强大技术,还是不是,两条轨迹你都可能给自己找麻烦。

Also in the space of ideas, that bubble creates a reality distortion field. That means you are deviating from reality, and if you go too far while also working 996, you might miss some fundamental aspects of the human experience. This is a common problem in Silicon Valley. It’s a very specific geographic area. You might not understand the Midwest perspective or the experience of all the other different humans in the United States and across the world. You speak a certain way to each other and convince each other of a certain thing, and that can get you into real trouble. Whether AI is a big success and becomes a powerful technology or it’s not, in either trajectory you can get yourself into trouble.

Nathan Lambert

我甚至不太理解这个,但旧金山 AI 梗已经到了「permanent underclass」那种。想法是 2025 年最后六个月是在 AI 创业公司或模型里建持久价值的唯一时间。否则所有价值会被现有公司收走,你因此会穷。那是旧金山走得太远的例子。我仍觉得对真正想在 AI 里产生影响的年轻人,人在旧金山是最可能做成的地方。但有权衡。

The SF AI memes have gotten to the point where the “permanent underclass” was one of them. This was the idea that the last six months of 2025 was the only time to build durable value in an AI startup or model. Otherwise all the value will be captured by existing companies and you will therefore be poor. That’s an example of the SF thing that goes so far. I still think for young people who are really passionate about having an impact in AI, being physically in SF is the most likely place where you’re going to do this. But it has trade-offs.

Lex Fridman

我觉得旧金山是不可思议的地方,但有一点泡沫。如果你走进那个泡沫——它极其有价值——也要走出来。读历史书、读文学,去世界其他地方。Twitter 和 Substack 不是整个世界。

I think SF is an incredible place, but there is a bit of a bubble. And if you go into that bubble, which is extremely valuable, just get out also. Read history books, read literature, and visit other places in the world. Twitter and Substack are not the entire world.

Nathan Lambert

我一起干活的一个人要搬去旧金山,我得给他一本《Season of the Witch》。旧金山 1960 到 1985 的历史,嬉皮革命、城市里浮现的文化、HIV/AIDS 危机,以及其他。那么近,那么多动荡和伤,也有爱。没人知道这个。好书,推荐。一帮会走出去的旧金山朋友推荐给我。我在那儿住过,没体会到这个上下文,而且那么近。

One of the people I worked with is moving to SF, and I need to get him a copy of Season of the Witch. It’s a history of SF from 1960 to 1985 that goes through the hippie revolution, the culture emerging in the city, the HIV/AIDS crisis, and other things. That is so recent, with so much turmoil and hurt, but also love in SF. No one knows about this. It’s a great book; I recommend it. A bunch of my SF friends who do get out recommended it to me. I lived there and I didn’t appreciate this context, and it’s just so recent.

2:43:19文本扩散模型Text diffusion models

Lex Fridman

今年你们提到令人兴奋的一件是文本扩散模型的缩放,以及对文本扩散不同的探索。那是什么,有什么可能?跟现在的 LM 不同的路?

One of the things you guys mentioned that’s exciting this year is the scaling of text diffusion models and a different exploration of text diffusion. Can you talk about what that is and what possibilities it holds? Different kinds of approaches than the current LMs?

Sebastian Raschka

我们谈了很多 Transformer,特别是自回归 Transformer,像 GPT。不意味着没人做别的。人一直在找下一件大事,不找几乎是傻的。现在 Transformer 最好用,但不该把蛋全放一个篮子。人在开发自回归 Transformer 的替代。其中一个是文本扩散模型。

We talked a lot about the transformer architecture and the autoregressive transformer specifically, like GPT. It doesn’t mean no one else is working on anything else. People are always on the lookout for the next big thing, because it would be almost stupid not to. Right now the transformer architecture is the thing and it works best, but it’s always a good idea to not put all your eggs into one basket. People are developing alternatives to the autoregressive transformer. One of them would be text diffusion models.

听众可能从图像生成知道扩散模型,Stable Diffusion 把它做火。那时人用 GAN。然后有这种迭代去噪图像的扩散过程,时间一长图像质量很好。现在人说:文本也能试吗?直觉上还不太通,因为感觉不像像素那样连续可微。是离散文本,去噪过程怎么实现?有点像 Google 的 BERT。原始 Transformer 有编码器和解码器。解码器是我们现在 GPT 在用的。编码器更像并行技术:多个 token 并行填。GPT 一次一个 token 自回归补全。BERT 是句子有缺口——mask 掉——一次迭代填那些缺口。

Listeners may know diffusion models from image generation; Stable Diffusion popularized it. Back then people used GANs. Then there was this diffusion process where you iteratively de-noise an image, and that resulted in really good quality images over time. Now people are like, okay, can we try this also for text? It doesn’t make intuitive sense yet because it feels like it’s not something continuous like a pixel that we can differentiate. It’s discrete text, so how do we implement that de-noising process? It’s kind of similar to the BERT models by Google. The original transformer had encoder and decoder. The decoder is what we are using right now in GPT. The encoder is more like a parallel technique where you have multiple tokens that you fill in in parallel. GPT models do autoregressive completion one token at a time. In BERT models you have a sentence that has gaps — you mask them out — and then one iteration is filling in those gaps.

文本扩散有点像那样:从一些随机文本开始,然后多次迭代填缺失部分或精炼。酷的是可以同时做多个 token,有更高效的承诺。权衡当然是质量。可能更快,但去噪步越多文本越好。人在看:作为自回归模型的替代,同样质量更少算力,是否成立。现在有论文暗示:要同样质量,你得把去噪步拧上去,最后花的算力和自回归一样。另一个缺点:虽然并行,有些任务不并行。推理任务或工具使用——你得问代码解释器要中间结果——扩散模型就有点棘手。有一些混合。主要想法是怎么并行化。有意思的路。现在大多是研究模型,像 LaMDA 和其他。我看到创业公司有一些部署模型,但还没有 Gemini 或 ChatGPT 那个尺度的大扩散模型。Google 有个宣布说要推 Gemini Diffusion,放在他们 Nano 2 模型的语境里。他们说大多数基准同样质量,可以生成得快得多。我不觉得文本扩散会取代自回归 LLM,但会是给快、便宜、大规模任务的东西。也许未来免费档会是那种。

Text diffusion is kind of like that, where you start with some random text, then fill in the missing parts or refine them iteratively over multiple iterations. The cool thing is this can do multiple tokens at the same time, so it has the promise of being more efficient. The trade-off is quality. It might be faster, but the more de-noising steps you do, the better the text becomes. People are trying to see if that is a valid alternative to the autoregressive model in terms of giving you the same quality for less compute. Right now there are papers that suggest if you want the same quality, you have to crank up the de-noising steps and then you end up spending the same compute you would spend on an autoregressive model. The other downside is that while it’s parallel, some tasks are not. For reasoning tasks or tool use where you have to ask a code interpreter to give you an intermediate result, it is kind of tricky with diffusion models. There are some hybrids. The main idea is how can we parallelize it. An interesting avenue. Right now there are mostly research models, like LaMDA and some other ones. I saw some by startups, some deployed models, but there is no big diffusion model at scale yet on the level of Gemini or ChatGPT. There was an announcement by Google where they said they are launching Gemini Diffusion, and they put it into context of their Nano 2 model. They said for the same quality on most benchmarks, we can generate things much faster. I don’t think the text diffusion model is going to replace autoregressive LLMs, but it will be something for quick, cheap, at-scale tasks. Maybe the free tier in the future will be something like that.

Nathan Lambert

有几个例子已经开始用了。为什么好这么多:GPT-5 这种模型响应要时间,是一次生成一个 token。扩散这个想法本质上是把补全里所有那些 token 一批生成,所以可以快得多。我听到的创业公司是代码创业公司:你有代码库,有人在 vibe coding。「做这个改动」,一个 code diff 本质上是模型一个巨大回复。不必有那么多外部上下文,用这些扩散模型可以很快拿到。他们用文本扩散生成很长的 diff,因为自回归会要几分钟,那段时间对面向用户的产品造成很多 churn。每一秒你都在丢用户。

There are a couple of examples where it’s actually started to be used. When a model like GPT-5 takes time to respond, it’s generating one token at a time. This diffusion idea is essentially generating all of those tokens in the completion in one batch, which is why it could be way faster. The startups I’m hearing are code startups where you have a codebase and somebody is effectively vibe coding. They say, make this change, and a code diff is essentially a huge reply from the model. It doesn’t have to have that much external context, and you can get it really fast by using these diffusion models. They use text diffusion to generate really long diffs because doing it with an autoregressive model would take minutes, and that time causes a lot of churn for a user-facing product. Every second, you lose users.

2:49:01工具使用Tool use

Lex Fridman

今年和接下来几年工具使用的未来?会有很多发展吗,怎么集成进整栈?

What’s the future of tool use this year and in the coming years? Do you think there’s going to be a lot of developments there, and how that’s integrated into the entire stack?

Sebastian Raschka

我觉得现在主要在专有 LLM 这边,但开源工具里会看到更多。这是巨大解锁,因为你真的可以把某些任务从死记硬背外包给实际计算——不必让 LLM 记住 23 加 5 是多少,用计算器。

I do think right now it’s mostly on the proprietary LLM side, but we will see more of that in open-source tooling. It is a huge unlock because then you can really outsource certain tasks from just memorization to actual computation — instead of having the LLM memorize what is 23 plus 5, just use a calculator.

Lex Fridman

你觉得这能帮忙解决幻觉吗?

Do you think that can help solve hallucinations?

Sebastian Raschka

不解决,但减少。LLM 仍需要知道何时请求工具调用。第二,不意味着互联网总是对的。你可以搜 1998 年世界杯谁赢,但仍需找到对的网站、拿到对的信息。你仍可能去错网站拿错信息。我不觉得会完全解决,但在改善。今年早些时候还有一篇很酷的论文——我觉得是 12 月 31 日,所以严格说不是 2026,但很近——关于递归语言模型。把这个再推远一点。Nathan 你之前说学术界因为算力预算更难做酷研究。我记得他们全用 GPT-5,甚至没用本地模型。想法是:长上下文任务,不要让 LLM 一次做完或一条链做完,拆成子任务。让 LLM 决定什么是好的子任务,然后递归调用 LLM 去解。再加工具——每个子任务也许上网收集信息,最后再合。会有很多解锁:不一定改进 LLM 本身,改进怎么用它、它能用什么。工具使用现在一个缺点是你必须给 LLM 用工具的许可。那需要信任,尤其你想解锁比如让 LLM 回邮件,或只是分类。我今天不知道会不会给 LLM 进我邮箱。这是巨大风险。

Not solve it, but reduce it. Still, the LLM needs to know when to ask for a tool call. Second, it doesn’t mean the internet is always correct. You can do a web search for who won the World Cup in 1998, but it still needs to find the right website and get the right information. You can still go to the incorrect website and get incorrect information. I don’t think it will fully solve it, but it is improving. There was another cool paper earlier this year — I think December 31st, so not technically 2026, but close — on the recursive language model. That’s a cool idea to take this even a bit further. Nathan, you mentioned earlier it’s harder to do cool research in academia because of the compute budget. If I recall correctly, they did everything with GPT-5, so they didn’t even use local models. The idea is, for a long-context task, instead of having the LLM solve all of it in one shot or in a chain, you break it down into sub-tasks. You have the LLM decide what is a good sub-task and then recursively call an LLM to solve that. Then adding tools — each sub-task maybe goes to the web and gathers information, and then you pull it all together at the end. There’s going to be a lot of unlock using things like that where you don’t necessarily improve the LLM itself, you improve how the LLM is used and what it can use. One downside right now with tool use is you have to give the LLM permission to use tools. That will take some trust, especially if you want to unlock things like having an LLM answer emails for you, or just sort them. I don’t know if I would today give an LLM access to my emails. This is a huge risk.

Nathan Lambert

工具使用最后一点:开源和闭源模型用工具的方式非常不同。开源模型,人去 Hugging Face 下载,然后那人会想「我要什么工具?」也许 X.ai 是我偏好的搜索提供商,别人可能在乎另一个搜索创业公司。你放模型时,它需要对多种工具有用,这真的很难,因为你在做通用推理引擎,这其实是 gpt-oss-120b 擅长的。闭源模型这边,你把特定工具深深集成进体验。我觉得开源模型很难复制我喜欢用闭源模型做的一些事,那种你可以引用公开和私有信息的混合。我每三到六个月试一次网上的 Codex,就是 prompt 一个模型去更新我的某个 GitHub 仓库。

Open versus closed models use tools in very different ways. With open models, people go to Hugging Face and download the model, and then the person’s going to be like, what tool do I want? Maybe X.ai is my preferred search provider, but someone else might care for a different search startup. When you release a model, it needs to be useful for multiple tools, which is really hard because you’re making a general reasoning engine, which is actually what gpt-oss-120b is good for. On the closed models, you’re deeply integrating the specific tool into your experience. I think open models will struggle to replicate some of the things I like to do with closed models, where you can reference a mix of public and private information. Something I keep trying every three to six months is Codex on the web, which is just prompting a model to make an update to some GitHub repository that I have.

2:53:17持续学习Continual learning

Lex Fridman

持续学习——长期题目,也是重要问题。训模型成本上升,它越来越重要。能解释一下持续学习是什么,今年和接下来几年推进它有多重要吗?

Continual learning — a longstanding topic and an important problem. I think that increases in importance as the cost of training models goes up. Can you explain what continual learning is and how important it might be this year and in the coming years to make progress?

Nathan Lambert

这很相关旧金山时代精神里:什么是 AGI,什么是 ASI。我们今天的语言模型能做什么?我觉得语言模型能解很多任务,但 AI 社区一个关键里程碑是 AI 能替代任何远程工作者,接收信息、解数字任务。限制是语言模型不会像员工那样从反馈学习。你雇一个编辑,他们可能搞砸,但你会告诉他们,他们不再做。语言模型没有这种快速修改自己、学习的能力。想法是:如果我们要到真正通用、可适应、能进任何远程工作场景的智能,它需要能从反馈和在职学习里快速学习。我个人更看好语言模型只要能提供很好的上下文。你可以写很长的文档:「我有所有这些信息。这是我写过的所有博客。我喜欢这种写作;我的声音基于这个。」但很多人并不把这些提供给模型。智能体模型才刚开始。权衡是:我们需要用持续学习这件事去更新这个模型的权重,让它们学得快?还是反方:我们只要给它们更多上下文和信息,它们会因为有很多上下文而且很聪明,而显得学得快。

This relates a lot to this SF zeitgeist of: what is AGI, Artificial General Intelligence, and what is ASI, Artificial Superintelligence? What are the language models we have today capable of doing? Language models can solve a lot of tasks, but a key milestone for the AI community is when AI can replace any remote worker, taking in information and solving digital tasks. The limitation is that a language model will not learn from feedback the same way an employee does. If you hire an editor, they might mess up, but you will tell them, and they don’t do it again. Language models don’t have this ability to modify themselves and learn very quickly. The idea is, if we are going to get to something that is a true, general adaptable intelligence that can go into any remote work scenario, it needs to be able to learn quickly from feedback and on-the-job learning. I’m personally more bullish on language models being able to just provide very good context. You can write extensive documents: I have all this information. Here are all the blog posts I’ve ever written. I like this type of writing; my voice is based on this. But a lot of people don’t provide this to models. The agentic models are just starting. It’s this kind of trade-off: do we need to update the weights of this model with this continual learning thing to make them learn fast? Or, the counterargument is we just need to provide them with more context and information, and they will have the appearance of learning fast by just having a lot of context and being very smart.

Lex Fridman

术语:持续学习指持续改权重,让模型基于新进来的信息适应、调整,而且持续、迅速、频繁。你另一边提的通常叫上下文学习。你学东西,有巨大上下文窗口,每次 prompt 系统都可以再塞额外信息。两者都可以正当地看成学习。只是学习发生的地方不同。

Continual learning refers to changing the weights continuously so that the model adapts and adjusts based on the new incoming information, and does so continually, rapidly, and frequently. The thing you mentioned on the other side is generally referred to as in-context learning. As you learn stuff, there’s a huge context window. You can just keep loading it with extra information every time you prompt the system. Both can legitimately be seen as learning. It’s just a different place where you’re doing the learning.

Sebastian Raschka

说实话,持续学习——更新权重——我们已经有不同风味。区分是:你对每个人做个性化定制模型,还是在全局模型尺度做?我们已经有从 GPT-5 到 5.1 到 5.2。也许不是即时的,但是快速的精选更新:社区反馈他们做不了的事,他们更新权重,放下一版。某种风味。更细的例子是 RLVR;你跑它,它更新。问题是你不能对每个人都那样做,因为给每个人更新权重太贵。即使 OpenAI 尺度、在建数据中心,也太贵。我觉得只有设备上、成本在消费者身上才可行。像 Apple 试图用 Apple Intelligence 模型做的:放在手机上,让它们从体验学习。

Continual learning — the updating of weights — we already have that in different flavors. The distinction is: do you do that on a personalized custom model for each person, or do you do it on a global model scale? We have that already with going from GPT-5 to 5.1 and 5.2. It’s maybe not immediate, but it is like a quick curated update where there was feedback by the community on things they couldn’t do. They updated the weights, released the next model. Kind of a flavor of that. Another even finer-grained example is RLVR; you run it, it updates. The problem is you can’t just do that for each person because it would be too expensive to update the weights for each person. Even at OpenAI scale, building the data centers, it would be too expensive. I think that is only feasible once you have something on the device where the cost is on the consumer. Like what Apple tried to do with the Apple Intelligence models, putting them on the phone so they learn from the experience.

现在记忆大多是上下文——把东西塞进上下文再召回。但即使缓存也贵,你在那上花 token。第二,你只能做那么多。更像偏好或风格。很多人解数学题时那样做。你可以加先前知识,也给某些偏好 prompt,「做我上次偏好的」。但不解锁新能力。为此人仍在用 LoRA adapter。

Right now memory is mostly like context — stuffing things into the context and then just recalling that. But again, it’s expensive because even if you cache it, you spend tokens on that. And second, you can only do so much. I think it’s more like a preference or style. A lot of people do that when they solve math problems. You can add previous knowledge, but you also give it certain preference prompts, like do what I preferred last time. But it doesn’t unlock new capabilities. For that, one thing people still use is LoRA adapters.

2:58:39长上下文Long context

Nathan Lambert

口语上公认的是:这是算力和数据问题。有时有小的架构东西,比如注意力变体。我们谈过混合注意力模型,本质上是 Transformer 里看起来像状态空间模型的东西。它们更合适,因为要建模最远的 token 花更少算力。但不是免费的,必须配大量算力或对的数据。世界上有多少 10 万 token 的序列,你从哪拿?缩放它们最后相当贵。

The colloquially accepted thing is that it’s a compute and data problem. Sometimes there are small architecture things, like attention variants. We talked about hybrid attention models, which is essentially if you have what looks like a state space model within your transformer. Those are better suited because you have to spend less compute to model the furthest along token. But those aren’t free because they have to be accompanied by a lot of compute or the right data. How many sequences of 100,000 tokens do you have in the world, and where do you get these? It just ends up being pretty expensive to scale them.

我们很快到了百万 token 的输入上下文长度。我预期会继续涨,今年到 200 万或 500 万,但我不预期到比如 1 亿。那会是真正的突破,我觉得那些突破可能。我把持续学习想成研究问题:可能有突破让 Transformer 在这上好得多而且便宜。科学注意力这么多,这些事可能发生。但摇把手,会是随时间一致的增长。

We’ve gotten pretty quickly to a million tokens of input context length. I would expect it to keep increasing and get to 2 million or 5 million this year, but I don’t expect it to go to like 100 million. That would be a true breakthrough, and I think those breakthroughs are possible. I think of the continual learning thing as a research problem where there could be a breakthrough that makes transformers work way better at this and it’s cheap. These things could happen with so much scientific attention. Turning the crank, it’ll be consistent increases over time.

Sebastian Raschka

看极端,没有免费午餐。让它便宜的一个极端是比如 RNN,单个状态保存之前所有东西。固定大小,内存永不真正增长。你把一切塞进一个状态,但上下文越长忘得越多,因为你不能把一切压进一个状态。另一边 Transformer 试图记住每个 token。你想查找特定信息时很好,但非常贵,因为 KV cache 和点积会涨。然后像你说的,Mamba 层有点同一问题。像 RNN,你试图把一切压进一个状态,只是更有选择。又是那种 Goldilocks 区:NVIDIA Nemotron 3 找到了你需要多少注意力层来做全局信息——一切可及——相对这些压缩状态的好比例。我觉得我们会更多地通过在那个 Goldilocks 区找更好比例来缩放:够便宜能跑,够强有用。

Looking at the extremes, there’s no free lunch. One extreme to make it cheap is to have, let’s say, an RNN that has a single state where you save everything from the previous stuff. A specific fixed-size thing, so you never really grow the memory. You are stuffing everything into one state, but then the longer the context gets, the more information you forget because you can’t compress everything into one state. On the other hand you have the transformers, which try to remember every token. That is great if you want to look up specific information, but very expensive because you have the KV cache and the dot product that grow. Then, like you said, the Mamba layers kind of have the same problem. Like an RNN, you try to compress everything into one state, and you’re a bit more selective there. This Goldilocks zone again with NVIDIA Nemotron 3; they found a good ratio of how many attention layers you need for the global information where everything is accessible compared to having these compressed states. I think we will scale more by finding better ratios in that Goldilocks zone between making it cheap enough to run and making it powerful enough to be useful.

再推一下:递归语言模型论文是试图处理长上下文的论文之一。他们发现,本质上,与其把一切塞进这个长上下文,如果你拆成多个更小任务,你省内存,准确率还能比 LLM 一次全做更好。新范式;我们会看有没有其他风味。我们仍会在长上下文上进步,但像 Nathan 说的,问题是预训练本身,我们没有那么多长上下文文档。所以更难看那个尺度上 LM 怎么表现。

One more plug: the recursive language model paper is one of the papers that tries to address the long context thing. What they found is, essentially, instead of stuffing everything into this long context, if you break it up into multiple smaller tasks, you save memory and can actually get better accuracy than having the LLM try everything all at once. It’s a new paradigm; we will see if there are other flavors of that. I think we will still make improvement on long context, but like Nathan said, the problem is for pre-training itself, we don’t have as many long-context documents as other documents. So it’s harder to study how LMs behave on that level.

Nathan Lambert

有些经验法则:你预训练一个语言模型——像 OLMo,我们在 8K 上下文长度预训练,再通过训练扩到 32K。经验法则是训练上下文长度翻倍大约要 2 倍算力,然后你通常可以再把上下文长度 2 到 4 倍。我觉得很多最后是预训练上算力受限。人人都在谈今年顶级实验室算力大增,那应该反映在一些更长的上下文窗口上。

There are some rules of thumb where, essentially, you pre-train a language model — like OLMo, we pre-trained at an 8K context length and then extended to 32K with training. There’s a rule of thumb where doubling the training context length takes about 2X compute, and then you can normally 2 to 4X the context length again. A lot of it ends up being compute-bound at pre-training. Everyone talks about this big increase in compute for the top labs this year, and that should reflect in some longer context windows.

后训练侧有些更有意思的。我们有智能体,智能体会自己管理这个上下文。现在重度用 Claude Code 的人怕 compaction:Claude 把整整 10 万 token 的工作压成项目符号列表。但下一模型会做的——我确信已有人在做——是模型能控制何时压缩、怎么压缩。所以你可以训练你的 RL 算法,压缩是一个动作,它缩短历史。然后问题表述会是:我希望评估分最高,同时模型把历史压到最短。因为那样你做这种复合自回归预测所需的 token 最少。其实有相当漂亮的问题设置,这些智能体模型学会用上下文的方式和只是往前犁不同。

On the post-training side, there’s some more interesting things. As we have agents, the agents are going to manage this context on their own. Now people who use Claude Code a lot dread the compaction, which is when Claude takes its entire 100,000 tokens of work and compacts it into a bulleted list. But what the next models will do — I’m sure people are already working on this — is the model can control when it compacts and how. You can essentially train your RL algorithm where compaction is an action, where it shortens the history. Then the problem formulation will be, I want to keep the maximum evaluation scores while the model compacts its history to the minimum length. Because then you have the minimum amount of tokens that you need to do this kind of compounding auto-regressive prediction. There are actually pretty nice problem setups in this where these agentic models learn to use their context in a different way than just plowing forward.

3:04:54机器人Robotics

Lex Fridman

该说 AI 空间里还有很多令人兴奋的东西。我最近脑子很在机器人上,今天几乎完全没谈机器人。图像生成、视频生成也有很多。公平地说,强度和热度上最令人兴奋的研究工作在 LLM 空间,所以我们聚焦正在谈的 LLM 是站得住的。但引入某些可能有用的东西会很好。比如世界模型——那方面兴奋在涨。你觉得接下来一年世界模型在 LLM 空间会有用吗?

There’s a lot of exciting stuff going on in the AI space. My mind has recently been really focused on robotics, so today we almost entirely didn’t talk about robotics. There’s a lot of stuff on image generation and video generation. I think it’s fair to say the most exciting research work in terms of intensity and fervor is in the LLM space, which is why I think it’s justified for us to focus on the LLMs we’re discussing. But it’d be nice to bring in certain things that might be useful. For example, world models — there’s growing excitement about that. Do you think there will be any use in this coming year for world models in the LLM space?

Sebastian Raschka

LLM 这边有意思的是:如果我们解锁更多 LLM 能力,也自动解锁所有其他领域,因为进展更快。很多研究者和工程师用 LLM 写代码。即使他们做机器人,你优化这些帮写代码的 LLM,也划得来。世界模型有意思。基本上是让模型跑世界的模拟——真东西的一个小玩具版——可以解锁 LLM 不知道的数据。它可以模拟东西。LLM 碰巧靠预训练和下一个 token 预测工作得好,但我们可以用更精巧的方式做。有一篇论文,我觉得是 Meta 的,叫 Coder World Models。他们基本上把世界模型概念用到 LLM:不只是下一个 token 预测和可验证奖励检查答案对不对,他们也确保中间变量对。模型基本上在学一个代码环境。我觉得很有道理,只是做起来贵。通过建模整个过程而不只是结果,让事情更精巧,能加更多价值。

With LLMs, if we unlock more LLM capabilities, it also automatically unlocks all the other fields because it makes progress faster. A lot of researchers and engineers use LLMs for coding. Even if they work on robotics, if you optimize these LLMs that help with coding, it pays off. World models are interesting. It’s basically where you have the model run a simulation of the world — like a little toy version of the real thing — which can unlock capabilities like data the LLM is not aware of. It can simulate things. LLMs happen to work well by pre-training and doing next-token prediction, but we could do this in a more sophisticated way. There was a paper, I think by Meta, called Coder World Models. They basically apply the concept of world models to LLMs where, instead of just having next-token prediction and verifiable rewards checking the answer correctness, they also make sure the intermediate variables are correct. The model is basically learning a code environment. I think this makes a lot of sense; it’s just expensive to do. It is making things more sophisticated by modeling the whole process, not just the result, and that can add more value.

我读研究生时有个比赛叫 CASP,做蛋白质结构预测。预测还没解出来的蛋白质结构。某种意义上这很好,我觉得 LLM 也需要那种:你做基准,但没人知道解,直到事后有人揭晓。AlphaFold 出来时碾碎了这个基准。多次迭代,但我记得第一个明确建模物理相互作用和分子物理。还有不可能的角度这类。下一版我觉得他们去掉这个,只用蛮力放大。LLM 我们目前处在这种蛮力缩放,因为它碰巧能用,但我觉得某时点把这种方法带回来也许有道理。世界模型那方面也许相当酷。当然对机器人,那和 LLM 完全相关。

When I was a grad student, there’s a competition called CASP where they do protein structure prediction. They predict the structure of a protein that is not solved yet. In a sense this is actually great, and I think we need something like that for LLMs also, where you do the benchmark but no one knows the solution until someone reveals it after the fact. When AlphaFold came out, it crushed this benchmark. Multiple iterations, but I remember the first one explicitly modeled the physical interactions and the physics of the molecule. Also things like impossible angles. Then in the next version I think they got rid of this and just used brute force, scaling it up. With LLMs, we are currently in this brute-force scaling because it just happens to work, but I do think at some point it might make sense to bring back this approach. With world models, that might be actually quite cool. And of course, for robotics, that is completely related to LLMs.

Lex Fridman

机器人非常明确。有运动或操作的问题。运动在学习域里解决得多得多。但有很多价值,就像最初的蛋白质折叠系统……把传统基于模型的方法带进来。不太可能端到端学会操作或全身局部操作问题。那是梦。但你看人手的魔力、真实世界的复杂度,你意识到很难一路学到底——AlphaFold 2 大概也没有。

Robotics is very explicit. There’s the problem of locomotion or manipulation. Locomotion is much more solved, especially in the learning domain. But there’s a lot of value, just like with the initial protein folding systems, bringing in the traditional model-based methods. It’s unlikely that you can just learn the manipulation or the whole-body local manipulation problem end-to-end. That’s the dream. But then you realize when you look at the magic of the human hand and the complexity of the real world, you realize it’s really hard to learn this all the way through the way I guess AlphaFold 2 didn’t.

Nathan Lambert

我对机器人学习空间很兴奋。我觉得它被语言模型整体的兴奋和投资集体增压。训 Transformer 的基础设施——通用建模的东西——正在变成世界级工业工具。机器人以前哪里受限,现在好得多。算力多得多。他们拿这些语言模型当中枢,你可以围绕已经能用的东西做有意思的探索工作。然后我看见它在浮现,有点像我们谈的 Hugging Face transformers。我在 Hugging Face 时试图让这发生,但太早。Hugging Face 上这些开源机器人模型让人能贡献数据、微调它们。我觉得我们现在近得多,机器人和自动驾驶的投资相关,也在启用这个。一旦到这种生态,有人可以下载机器人模型、微调到自己的机器人,或在世界范围共享数据集。这个领域有些工作,比如几年前的 RTX,人开始那样做。一旦他们有这个生态,看起来会很不同。然后这整场 ChatGPT 之后的繁荣把更多资源放进去,我觉得是非常好的研究领域。

I’m excited about the robotic learning space. I think it’s collectively getting supercharged by all the excitement and investment in language models generally. The infrastructure for training transformers, which is a general modeling thing, is becoming world-class industrial tooling. Wherever there was a limitation for robotics, it’s just way better now. There’s way more compute. They take these language models and use them as central units where you can do interesting explorative work around something that already works. I see it emerging, kind of like we talked about Hugging Face transformers. When I was at Hugging Face I was trying to get this to happen, but it was too early. These open robotic models on Hugging Face enable people to contribute data and fine-tune them. I think we’re much closer now that the investment in robotics and self-driving cars is related and enables this. Once you get to the point where you have this sort of ecosystem, someone can download a robotics model and fine-tune it to their robot or share datasets across the world. There’s some work in this area like RTX from a few years ago where people are starting to do that. Once they have this ecosystem, it’ll look very different. Then this whole post-ChatGPT boom is putting more resources into that, which I think is a very good area for doing research.

Lex Fridman

这也带来更好、更准、更真实的模拟器,在机器人空间里缩小 sim-to-real 差距。但你提到很多兴奋和投资。炒作周期的下行是——我个人相信,大多数机器人的人也相信——机器人不会在被隐式或显式承诺的时间尺度上被解决。所以当这些机器人公司冒出来、然后没有能用的产品,会有兴奋崩盘,让人紧张。希望有别的东西冲进来,让其中一些想法能继续发展。

This is also resulting in much better, more accurate, more realistic simulators being built, closing this sim-to-real gap in the robotic space. You mentioned a lot of excitement and investment. The downside of that, which happens in hype cycles — I personally believe, and most robotics people believe — is that robotics is not going to be solved on the timescale being implicitly or explicitly promised. So what happens when all these robotics companies spring up and then they don’t have a product that works? Then there’s going to be this crash of excitement, which is nerve-wracking. Hopefully something else will swoop in so that the continued development of some of these ideas keeps going.

Sebastian Raschka

这也和持续学习问题相关。真实世界那么复杂,而 LLM 你其实不必为用户学,因为有很多事人人都要做——人人也许想修邮件语法或代码。更受限,所以你可以为此准备模型。但为真实世界准备机器人更难。你有机器人基础模型,可以学抓取这类,但每栋房子都不同。不同到机器人基本上得在职学习。我觉得现在的瓶颈是即时定制。

I think it’s also related to the continual learning issue. The real world is so complex, whereas with LLMs you don’t really need to have something learn for the user because there are a lot of things everyone has to do — everyone maybe wants to fix their grammar in their email or code. It’s more constrained, so you can prepare the model for that. But preparing a robot for the real world is harder. You have robotic foundation models, and you can learn things like grasping, but every house is different. It’s so different that the robot would have to learn on the job, essentially. I think that is the bottleneck right now: customizing it on the fly.

3:14:04通往 AGI 的时间线Timeline to AGI

Lex Fridman

专门谈谈时间线:到 AGI 或 ASI。作为起点,说没人真正同意 AGI 和 ASI 的定义,公平吗?

Let’s talk about timelines specifically: timelines to AGI or ASI. Is it fair, as a starting point, to say that nobody really agrees on the definitions of AGI and ASI?

Nathan Lambert

分歧很多,但我在收到反驳:它是能复现大多数数字经济工作的东西。远程工作者是相当合理的例子。我觉得 OpenAI 的定义某种程度上相关——能做一定数量经济上有价值任务的 AI——我不太喜欢当定义,但可以当锚点。今天的语言模型虽然极强,不是这种远程工作者即插即用。有些 AI 能做的事比远程工作难得多,比如找到你甚至提不出来的意外科学发现,那会被人称为人工超级智能问题。或吞下所有医疗记录,发现人不知道的某些疾病关联,或发现某种常用药能治某种小众癌症。他们会说那是超级智能的事。所以这些是自然层级。我的问题是它深深缠进对 AI 意义的追寻、这些宗教面向。你可以走不同的路。

There’s a lot of disagreement, but I’ve been getting pushback where people say it is something that could reproduce most digital economic work. The remote worker is a fairly reasonable example. OpenAI’s definition is somewhat related to that — an AI that can do a certain number of economically valuable tasks — which I don’t really love as a definition, but it could be a grounding point. Language models today, while immensely powerful, are not this remote worker drop-in. There are things an AI could do that are way harder than remote work, like finding an unexpected scientific discovery that you couldn’t even posit, which would be an example of something people call an artificial superintelligence problem. Or taking in all medical records and finding linkages across certain illnesses that people didn’t know, or figuring out that some common drug can treat a niche cancer. They would say that is a superintelligence thing. So these are natural tiers. My problem is that it becomes deeply entwined with the quest for meaning in AI and these religious aspects. There are different paths you can take.

Lex Fridman

我甚至不知道远程工作是不是好定义。我喜欢原来叫 AI2027 的报告。他们更聚焦代码和研究品味,目标是超人程序员。他们有几个里程碑系统:超人程序员、超人 AI 研究者、然后超智能 AI 研究者、然后完整 ASI。你做出超人程序员之后,其他都很快跟上。任务是完全自主、自动化的编码,所以做研究所需的任何编码都完全自动化。从那里,人会和那个系统一起做 AI 研究,很快能发展出真正能替你做研究的系统。那是想法。最初他们预测 2027 或 28,现在推迟了三到四年,均值预测到 2031。我的预测大概还在 2031 之后,但至少你可以具体想完全自动化编程有多难。

I don’t even know if remote work is a good definition. I liked the originally titled AI2027 report. They focus more on code and research taste, so the target there is the superhuman coder. They have several milestone systems: superhuman coders, superhuman AI researcher, then superintelligent AI researcher, and then the full ASI. After you develop the superhuman coder, everything else follows quickly. The task is to have fully autonomous, automated coding, so any kind of coding you need to do in order to perform research is fully automated. From there, humans would be doing AI research together with that system, and they will quickly be able to develop a system that actually can do the research for you. That’s the idea. Initially their prediction was 2027 or ’28, and now they’ve pushed it back by three to four years to 2031, mean prediction. My prediction is probably even beyond 2031, but at least you can think concretely about how difficult it is to fully automate programming.

Nathan Lambert

我不同意他们一些关于会如何展开的前提和动态,但我觉得他们做了好工作,定义具体里程碑来讲一个有用的故事。所以这份 AI 2027 文档的触及远远超出硅谷——因为他们讲了好故事,做了很多严谨工作。我所属的阵营是:AI 所谓锯齿状,某些事极好、某些事真的很差。我觉得当他们接近这个自动化软件工程师时,它会擅长的是传统 ML 系统和前端——模型在那些上极好——但分布式 ML,模型其实相当差,因为大规模分布式学习这类的训练数据那么少。这是我们已经看见的,我觉得这会被放大。然后这些权衡更乱,然后你怎么想 AI 研究怎么工作等等。

I disagree with some of their presumptions and dynamics on how it would play out, but I think they did good work in defining concrete milestones to tell a useful story. That’s why the reach of this AI 2027 document well transcended Silicon Valley — because they told a good story and did a lot of rigorous work. The camp I fall into is that AI is so-called jagged, which will be excellent at some things and really bad at some things. When they’re close to this automated software engineer, what it will be good at is traditional ML systems and front end — the model is excellent at those — but the distributed ML, the models are actually really quite bad at because there’s so little training data on doing large-scale distributed learning and things. This is something we already see, and I think this will just get amplified. Then it’s kind of messier in these trade-offs, and then there’s how you think AI research works and so on.

Lex Fridman

所以你觉得超人程序员几乎不可达成,因为东西的锯齿性质,你总会有能力缺口?

So you think a superhuman coder is almost unachievable, because of the jagged nature of the thing, you’re just always going to have gaps in capabilities?

Nathan Lambert

我觉得是在给某样东西赋完整性,而模型在某些类型的代码上已经有点超人,我觉得会继续。人有创造力,他们会利用这些不可思议的能力去填模型的弱点,走得非常快。很长一段时间总会有这种舞:人启用模型做不到的事,最好的 AI 研究者是能启用这种超能力的人。跟我们已经看见的那些线比……我觉得像 Claude Code 建网站,你可以几小时立起漂亮网站或做数据分析。但整件事会在这些上继续变好,我们会沿路捡起一些新的代码技能。连到大科技在发生的事,这份 AI 2027 报告倾向奇点想法,而我觉得研究是乱的、社会的、很大程度上在数据里,AI 模型处理不了。但我们今天有的真的很强,这些科技公司集体用数百亿美元投资买进这个。所以我们会得到好得多的 ChatGPT 版本,比我们已有的好得多的 Claude Code。

I think it’s assigning completeness to something where the models are kind of superhuman at some types of code, and I think that will continue. People are creative, so they’ll utilize these incredible abilities to fill in the weaknesses of the models and move really fast. There will always be, for a long time, this dance between the humans enabling this thing that the model can’t do, and the best AI researchers are the ones that can enable this superpower. Those lines, compared to what we already see… Claude Code for building a website, you can stand up a beautiful website in a few hours or do data analysis. But the whole thing is going to keep getting better at these things, and we’ll pick up some new code skills along the way. Linking to what’s happening in big tech, this AI 2027 report leans into the singularity idea where I think research is messy and social and largely in the data in ways that AI models can’t process. But what we do have today is really powerful, and these tech companies are all collectively buying into this with tens of billions of dollars of investment. So we are going to get some much better version of ChatGPT, a much better version of Claude Code than we already have.

很难预测那会去哪,但那个未来的明亮清晰,是为什么世界上一些最有权的人把那么多钱放进去。我觉得只是小差异——我们其实不知道更好的 ChatGPT 是什么,但它能自动化 AI 研究吗?我会说大概不能,至少在这个时间框架里。大科技花 1000 亿美元,会比我们得到一个启用 AI 研究奇点的自动化 AI 研究者快得多。

It’s just hard to predict where that is going, but the bright clarity of that future is why some of the most powerful people in the world are putting so much money into this. It’s just kind of small differences — we don’t actually know what a better version of ChatGPT is, but also can it automate AI research? I would say probably not, at least in this timeframe. Big tech is going to spend $100 billion much faster than we get an automated AI researcher that enables an AI research singularity.

Lex Fridman

所以你的预测会是,如果这甚至是有用的里程碑,超过 10 年?

So you think your prediction would be, if this is even a useful milestone, more than 10 years out?

Nathan Lambert

软件侧我会说短于那个,研究这类事会长于那个。到今年年底,会被自动化的软件量会非常高。但会是那种:你试图用 RL 训模型,需要好几组 GPU 互相通信。那仍会难,但我觉得会容易得多。

I would say less than that on the software side, but I think longer than that on things like research. By the end of this year, the amount of software that’ll be automated will be so high. But it’ll be things like you’re trying to train a model with RL and you need to have multiple bunches of GPUs communicating with each other. That’ll still be hard, but I think it’ll be much easier.

3:21:20AI 会取代程序员吗?Will AI replace programmers?

Nathan Lambert

软件工程会更多地被推向系统设计和对结果的目标。过去几周这已经在发生:一个月前人还在说「哦对,智能体有点 slop」——Karpathy 的名言——到软件工业化,任何人都可以用他们的指纹创造软件。我觉得我们更接近那边,需要方向、理解系统怎么工作,才能从语言模型里挤出最好的。很难接受软件开发会变多少,以及有多少更多人可以做事而从不看代码。

I think software engineering will be driven more to system design and goals of outcomes. This has been happening over the last few weeks, where people have gone from a month ago saying, “oh yeah, agents are kind of slop,” which is a famous Karpathy quote, to the industrialization of software when anyone can just create software with their fingerprints. I do think we are closer to that side of things, and it takes direction and understanding how the systems work to extract the best from the language models. It’s hard to accept the gravity of how much is going to change with software development and how many more people can do things without ever looking at the code.

Sebastian Raschka

有意思的是想这些系统会不会独立。我毫不怀疑 LLM 某时点会像计算器解决计算那样解决编码。某时点人做了一个工具,你再也不需要人去算那个数;你打进去,它是算法。编码大概一样。但问题不是……会发生的是你会说「建那个网站」,它会做出很好的网站,然后你也许精炼。但会不会独立做事,你会不会仍有人在问 AI 做某事?会不会有人说「建那个网站」?还是会有 AI 就去建网站?

What’s interesting is to think about whether these systems will be independent, in the sense that while I have no doubt that LLMs will at some point solve coding in the way calculators solve calculating. At some point humans developed a tool that you never need a human to calculate that number; you just type it in, and it’s an algorithm. I think that’s the same probably for coding. But the question isn’t… what will happen is you will just say, build that website, and it will make a really good website, and then you maybe refine it. But will it do things independently where will you still have humans asking the AI to do something? Like will there be a person to say, build that website? Or will there be AI that just builds websites?

Lex Fridman

谈建网站太简单了。网站和网页、HTML 那些,对 slop 非常有弹性。它会给你看 slop,它擅长展示 slop。我宁愿想安全关键系统,比如让 AI 端到端生成管理物流的东西,或管理车、车队。它端到端给你生成那个。

Talking about building websites is too simple. The problem with websites and the web, HTML and all that, it’s very resilient to just slop. It will show you slop. It’s good at showing slop. I would rather think of safety-critical systems, like asking AI to end-to-end generate something that manages logistics — or manages cars — a fleet of cars. So it end-to-end generates that for you.

Nathan Lambert

更中间的例子是 Slack 或 Microsoft Word。如果组织允许,AI 可以很容易端到端实现功能,对你想试的东西做得相当好。你想在 Slack 加一个你想用的新标签,我觉得 AI 能做得相当好。

A more intermediate example is take something like Slack or Microsoft Word. If the organizations allow it, AI could very easily implement features end-to-end and do a fairly good job for things that you want to try. You want to add a new tab in Slack that you want to use, and I think AI will be able to do that pretty well.

Lex Fridman

那真是很好的例子。我们离那还有多远?

That’s a really great example. How far away are we from that?

Nathan Lambert

像今年。我也不知道生产代码库有多糟,但我觉得几年量级内,很多人会被推向更像设计师和产品经理,你有多个这些智能体可以替你试东西,它们也许要一两天实现一个功能或试图修一个 bug。你会有这些仪表盘——我觉得 Slack 其实是好仪表盘——智能体跟你说话,然后你给反馈。但像我做一个网站,「你想做一个过得去的 logo 吗?」这些有凝聚力的设计东西和风格对模型会很难,以及决定下一步加什么。

Like this year. I guess I don’t know how bad production codebases are, but I think that within, on the order of a few years, a lot of people are going to be pushed to be more like a designer and product manager, where you have multiple of these agents that can try things for you, and they might take one to two days to implement a feature or attempt to fix a bug. You have these dashboards — Slack is actually a good dashboard — where your agents will talk to you and you’ll then give feedback. But things like, I make a website and it’s like, do you want to make a logo that’s passable? These cohesive design things and the style is going to be very hard for models, and deciding on what to add next.

Lex Fridman

我跟很多程序员混,有些总体偏怀疑——就是那种氛围。我觉得给复杂系统加功能涉及很多复杂度。比如浏览器,Chrome。如果我想加一个功能,标签不要在上面,要在左边。我觉得这不是明年的事。

I hang out with a lot of programmers and some of them are a little bit on the skeptical side in general — that’s just the vibe. I just think there’s a lot of complexity involved in adding features to complex systems. Like if you look at the browser, Chrome. If I wanted to add a feature, if I wanted to have tabs as opposed to up top, I want them on the left side. I think we’re not… this is not a next year thing.

Nathan Lambert

今年某次 Claude 发布,他们的一个测试是:我们给它一份软件,让 Claude 跑去完全重建它,它已经几乎能从零重建 Slack,只给软件的参数,放在沙盒环境里做。

One of the Claude releases this year, one of their tests was we give it a piece of software and leave Claude to run to recreate it entirely, and it could already almost rebuild Slack from scratch, just given the parameters of the software and left in a sandbox environment to do that.

也许更小、更新的公司有优势,他们像「我们不必有那种膨胀和复杂度,因此这个功能存在。」如果你跟实验室的人谈,他们在训练和生产代码里用这些。Claude Code 是用 Claude Code 建的,他们都大量用这些东西。达里奥讲 Claude 的代码有多少……这些人在能力上略微领先。

It might be that the smaller and newer companies are advantaged and they’re like, we don’t have to have the bloat and complexity, and therefore this feature exists. If you talk to people at the labs, they use these in their training and production code. Claude Code is built with Claude Code, and they all use these things extensively. Dario talks about how much of Claude’s code… These people are slightly ahead in terms of the capabilities.

Sebastian Raschka

这接到你提的有些人怀疑。我觉得不是因为 LLM 不能做 X、Y、Z。是因为人不想它用这种方式做。

This gets to the point that you mentioned that some people you talk to are skeptical, and I think that’s not because the LLM can’t do X, Y, Z. It’s because people don’t want it to do it this way.

Lex Fridman

有些可能是人这边的技能问题。不幸我们必须对自己诚实。有些可能是规格不足。编程,有点像关系和友谊里的沟通问题。你假设 LLM 不知怎的该读你的心。这就是规格驱动设计真的重要的地方。你就是用自然语言,规格你想要什么。

Some of that could be a skill issue on the human side. Unfortunately we have to be honest with ourselves. And some of that could be an underspecification issue. Programming, this is like a communication type of issue in relationships and friendships. You’re assuming the LLM somehow is supposed to read your mind. I think this is where spec-driven design is really important. You just, using natural language, specify what you want.

3:39:51AGI 的梦在死吗?Is the dream of AGI dying?

Lex Fridman

但我觉得现在人人在追的是对所有人都有用的通用系统。好,如果那不是……那可以平台期,对吗?

I think what everybody’s chasing now is a general system that’s useful to everybody. So, okay, if that’s not… that can plateau, right?

Nathan Lambert

我觉得那个梦其实有点在死。就像你谈专门模型……多模态经常是……视频生成完全是另一回事。

I think that dream is actually kind of dying. As you talked about with the specialized models where it’s like… and multimodal is often… like, video generation is a totally different thing.

Lex Fridman

「那个梦有点在死」是很大的陈述,因为我不知道它是不是在死。你问真正前沿实验室的人,他们仍在追,对吗?

“That dream is kind of dying” is a big statement, because I don’t know if it’s dying. If you ask the actual Frontier Lab people, they’re still chasing it, right?

Sebastian Raschka

我觉得他们仍在赶着把下一模型放出来,会比上一版好得多。「多」是相对的,但会比上一版好。我看不见他们放慢。我只是觉得收益会更多通过不只缩放模型来感受,现在……我觉得有很多技术债。就像「把更好的模型塞进去,更好的模型,更好的模型。」现在人说「好,同时把周围一切也改进。」上下文工程、推理缩放。大实验室仍会继续做。现在小实验室也会赶上,因为他们在招更多人。会有更多人。LLM 有点像一个圈。它们也让他们更高效,就是放大器。我觉得我们可以预期的是放大,不是范式改变。我不觉得那是真的,但一切会一直被放大、放大、放大,我能看见那持续很久。

I do think they are still rushing to get the next model out, which will be much better than the previous one. “Much” is a relative term, but it will be better than the previous one. I can’t see them slowing down. I just think the gains will be made or felt more through not only scaling the model, but now… I feel like there’s a lot of tech debt. It’s like, well, let’s just put the better model in there, and better model, better model. And now people are like, okay, let’s also at the same time improve everything around it too. Like the engineering of the context and inference scaling. The big labs will still keep doing that. And now also the smaller labs will catch up to that because now they are hiring more. There will be more people. LLMs, it’s kind of like a circle. They also make them more productive and it’s just like an amplifier. I think what we can expect is amplification, but not a paradigm change. I don’t think that is true, but everything will be just amplified and amplified and amplified, and I can see that continuing for a long time.

Nathan Lambert

我说梦在死,取决于你精确觉得它会做什么。Claude Code 是能做很多事的通用模型,但很依赖集成和其他东西。我打赌 Claude Code 做你的邮件能做得相当好,最难的部分是搞清楚怎么给它信息、怎么让它能发你的邮件这类。但我觉得回到「一个模型统治一切」的精神,就是云里有个东西处理你整个数字生活,比所有人都聪明得多。从 Claude Code 变成那个——某种意义上有一些路径——是有意思的信仰之跃,但我觉得行业的修辞有点不同。

My statement with the dream is dying depends on exactly what you think it’s going to be doing. Claude Code is a general model that can do a lot of things, but it depends a lot on integrations and other things. I bet Claude Code could do a fairly good job of doing your email, and the hardest part is figuring out how to give it information and how to get it to be able to send your emails and stuff like this. But I think it goes back to what is the “one model to rule everything” ethos, which is just like a thing in the cloud that handles your entire digital life and is way smarter than everybody. So it’s an interesting leap of faith to go from Claude Code becomes that — which, in some ways, there are some avenues for that — but I do think that the rhetoric of the industry is a little bit different.

Sebastian Raschka

我觉得我们作为普通人用 LLM 接下来马上会感到的,大概和很琐碎的东西相关,比如做图。现在 LLM 做图糟透了。是因为我们被服务的是推理算力更少的便宜模型,幕后更多?也许拧几下我们已经能得到更好的图,但你今天让画 X、Y、Z 的流程图,大多数时候糟透了。对人类这几乎是很简单的任务。我觉得有时画东西几乎比写东西容易。

I think the immediate thing we will feel next as a normal person using LLMs will probably be related to something trivial, like making figures. Right now LLMs are terrible at making figures. Is it because we are getting served the cheap models with less inference compute than behind the scenes? Maybe with some cranks we can already get better figures, but if you ask today to draw a flowchart of X, Y, Z, it’s most of the time terrible. And it is kind of a very simple task for a human. I think it’s almost easier sometimes to draw something than to write something.

3:46:40AI 怎么赚钱?How AI will make money?

Lex Fridman

如果它存在,也充满——说到迪士尼乐园——广告 slop。世界上任何城市,你问「最该做的 10 件事?」问 LLM 就是比互联网上任何东西好得多。

And if it does exist, it’s full of — speaking of Disney World — ad slop. Like any city in the world, if you ask “what are the top 10 things to do?” An LLM is just way better to ask than anything on the internet.

Nathan Lambert

目前是因为它们被大规模补贴,最终它们会靠广告付钱。在来了。

Well, for now, that’s because they’re massively subsidized, and eventually they’re going to be paid for by ads. It’s coming.

Lex Fridman

天哪。我希望那个语境里什么是广告、什么不是,有非常清楚的标示,但——

Oh my goodness. I’m hoping there’s a very clear indication of what’s an ad and what’s not an ad in that context, but—

Sebastian Raschka

这是我几年前提过的。比如你在找新跑鞋,耐克碰巧排第一是巧合吗?也许是,也许不是。我觉得这方面有明确法律。你必须说清楚,但我觉得那是人人怕的:里面那种微妙信息之类。这也把我们带到广告这个题目。我觉得这是一件事,希望他们试图 2025 年推出,因为我觉得他们现在用别的方式仍不赚钱……比如里面真正的广告位。然后问题是他们做不到,因为有没有广告的替代,人会涌向其他产品。而且他们互相加码、花那么多钱只为抢用户,也crazy。

That’s something I mentioned a few years ago. If you are looking for a new running shoe, is it a coincidence that Nike maybe comes up first? Maybe, maybe not. I think there are clear laws around this. You have to be clear about that, but I think that’s what everyone fears. It’s like the subtle message in there. This brings us to the topic of ads where, I think this was a thing, hopefully they try to launch in 2025 because I think they’re still not making money in that other way right now, like having actual ad spots in there. Then the thing, though, is they couldn’t because there are alternatives without ads and people would just flock to the other products. And it also is just crazy how they’re one-upping each other, spending so much money just to get the users.

Nathan Lambert

有些 Instagram 广告——我不用 Instagram——但我理解付钱给平台去找会真正喜欢你产品的用户的吸引力。那是 Instagram 广告最好的情况。十年往后看,广告的命题是:你通过有那么多用户在广告上赚那么多钱,你可以用来资助更好的研发、做更好的模型,所以 YouTube 在主导市场。Netflix 怕 YouTube。我不知道——我每月付 28 美元高级会员。他们至少每月从我和很多人身上赚 28 美元,他们在视频上创造如此主导的位置。所以命题是广告可以给你在每用户花费上持续优势。但现在里面钱那么多,有人启动那个飞轮是可怕的,因为这是长期赌注。

Some Instagram ads — I don’t use Instagram — but I understand the appeal of paying a platform to find users who will genuinely like your product. That is the best case of things like Instagram ads. If we go 10 years out, the proposition for ads is that you will make so much money on ads by having so many users that you can use this to fund better R&D and make better models, which is why YouTube is dominating the market. Netflix is scared of YouTube. I pay $28 a month for premium. They make at least $28 a month off of me and many other people, and they’re just creating such a dominant position in video. So I think that’s the proposition, which is that ads can give you a sustained advantage in what you’re spending per user. But there’s so much money in it right now that it’s like somebody starting that flywheel is scary because it’s a long-term bet.

3:51:022026 年的大收购Big acquisitions in 2026

Lex Fridman

你觉得今年商业上会有一些疯狂的大动作吗?比如 Google 或 Apple 收购 Anthropic 这类?

Do you think there’ll be some crazy big moves this year business-wise? Like Google or Apple acquiring Anthropic or something like this?

Nathan Lambert

达里奥永远不会卖,但我们开始看到某种整合,Groq 估值 200 亿美元,Scale AI 几乎 300 亿。还有无数其他交易的结构方式实际上对硅谷生态有害——这些许可交易,不是所有人都被带上,而不是让普通员工股票兑现受益的完整收购。那是硅谷文化要处理的大问题,因为创业生态是命脉。你加入创业公司,即使不那么成功,你的创业公司很可能以便宜溢价被收购,你会拿到股权兑付。这些许可交易本质上很多时候在拿走顶尖人才。我觉得 Groq 到英伟达的交易传闻对员工更好,但仍是这种规避反垄断的东西。我觉得这种整合趋势会继续。我和很多我尊敬的聪明人一直预期整合更早发生,但似乎事情开始转了。但同时你有公司因为我不理解的原因融离谱的钱。我像「我不知道你为什么拿那笔钱。」所以今年也许是混的,但一些整合压力在开始。

Dario will never sell, but we are starting to see some types of consolidation, with Groq being valued at $20 billion and Scale AI for almost 30 billion. There are countless other deals structured in a way that is actually detrimental to the Silicon Valley ecosystem — these licensing deals where not everybody gets brought along, rather than a full acquisition that benefits the rank-and-file employees by getting their stock vested. That’s a big issue for Silicon Valley culture to address because the startup ecosystem is the lifeblood. If you join a startup, even if it’s not that successful, your startup very well might get acquired at a cheap premium and you’ll get paid out for your equity. These licensing deals are essentially taking the top talent a lot of the time. I think the deal for Groq to NVIDIA is rumored to be better for the employees, but it is still this antitrust-avoiding thing. I think this trend of consolidation will continue. Me and many smart people I respect have been expecting consolidation to have happened sooner, but it seems like things are starting to turn. But at the same time, you have companies raising ridiculous amounts of money for reasons that I don’t understand. I’m like, I don’t know why you’re taking that money. So it’s maybe mixed this year, but some consolidation pressure is starting.

Lex Fridman

你觉得我们会看到哪种令人惊讶的整合?你说 Anthropic 是「永不」。Groq 是大的——顺便说是带 Q 的 Groq。

What kind of surprising consolidation do you think we’ll see? You say Anthropic is a “never.” Groq is a big one — Groq with a Q, by the way.

Nathan Lambert

就是有很多创业公司,AI 创业公司溢价非常高。所以可能有很多 100 亿美元量级的收购,对也许一年前才创办的创业公司来说是非常大的收购。我觉得 Manus.ai——这家总部位于新加坡、八个月前创办然后 20 亿美元退出。我觉得会有其他几十亿美元的大收购,比如 Perplexity。

There’s just a lot of startups and there’s a very high premium on AI startups. So there could be a lot of $10 billion range acquisitions, which is a really big acquisition for a startup that was maybe founded a year ago. I think Manus.ai — this company based in Singapore that was founded eight months ago and then had a $2 billion exit. I think there will be some other big multi-billion dollar acquisitions, like Perplexity.

Lex Fridman

像 Perplexity,对吗?我们一直在谈代码。也许有人收购 Cursor。

Like Perplexity, right? We’ve been talking about code. Maybe somebody acquires Cursor.

Nathan Lambert

他们处境那么好,因为他们有那么多用户数据。我们谈了持续学习这类;他们有最有意思的博客之一。他们提到新的 Composer 模型是这些来自中国的大型 Mixture of Experts 模型之一的微调。你可以从八卦知道,或因为模型有时用中文回答,美国模型都不那样。他们有篇博客说:「我们每 90 分钟根据人使用的真实世界反馈更新模型权重。」这是发生在模型上的最接近真实世界 RL 的东西,就在他们一篇博客里。

They’re in such a good position because they have so much user data. We talked about continual learning and stuff; they had one of the most interesting blog posts. They mentioned that their new Composer model was a fine-tune of one of these large Mixture of Experts models from China. You can know that from gossip or because the model sometimes responds in Chinese, which none of the American models do. They had a blog post where they said, “we’re updating the model weights every 90 minutes based on real-world feedback from people using it.” Which is the closest thing to real-world RL happening on a model, and it was just right there in one of their blog posts.

Lex Fridman

那不可思议。顺便说我经常用 Composer,好处之一是快。还会有一些 IPO 可能。你觉得 Anthropic、OpenAI、xAI?

That’s incredible. By the way I use Composer a lot because one of the benefits it has is it’s fast. And there’ll be some IPOs potentially. You think Anthropic, OpenAI, xAI?

Nathan Lambert

他们都能那么容易融那么多钱,不觉得有必要……只要融资容易,他们就不会 IPO,因为公开市场施加压力。

They can all raise so much money so easily that they don’t feel a need to… So long as fundraising is easy, they’re not going to IPO because public markets apply pressure.

3:55:34OpenAI、Anthropic、Google DeepMind、xAI、Meta 的未来Future of OpenAI, Anthropic, Google DeepMind, xAI, Meta

Lex Fridman

你觉得十年后一些前沿模型公司还会在吗?Anthropic、OpenAI?

You think 10 years from now some of the frontier model companies are still around? Anthropic, OpenAI?

Nathan Lambert

我绝对不觉得会赢家通吃,除非真的有某个人找到算法秘密让这个飞轮转起来。因为他们的发展路径都那么像。Google 和 OpenAI 有完全一样的产品,Anthropic 更聚焦,但你跟人谈,听起来他们在解很多同样的问题。所以我觉得……供给会铺开。正在做的蛋糕非常大,人会从里面拿钱。

I definitely don’t see it to be a winner-takes-all unless there truly is some algorithmic secret that one of them finds that lets this flywheel. Because the development path is so similar for all of them. Google and OpenAI have all the same products, and Anthropic’s more focused, but when you talk to people, it sounds like they’re solving a lot of the same problems. So I think… there’s offerings that’ll spread out. It’s a very big cake that’s being made that people are going to take money out of.

Lex Fridman

我不想trivial化,但 OpenAI 和 Anthropic 主要是 LLM 服务提供商。其他一些公司像 Google 和连着 X 的 xAI,还做别的。所以很有可能如果 AI 更商品化,只提供 LLM 的公司会死。

I don’t want to trivialize it, but OpenAI and Anthropic are primarily LLM service providers. And some of the other companies like Google and xAI, linked to X, do other stuff too. And so it’s very possible if AI becomes more commodified that the companies that are just providing LLMs will die.

Sebastian Raschka

我觉得他们的优势是用户很多,他们会转型。像 Anthropic,我觉得转型了。我不觉得他们原先计划做代码,但碰巧发现「好,这是个不错的利基,现在我们在这个利基里舒服,我们推这个利基。」我能看见同样的事一旦……假设说,我不确定会不会真,但假设 Google 拿走通用聊天机器人的全部市场份额。也许 OpenAI 然后会聚焦某个其他子题目。他们用户太多,可预见的未来不会消失。

I think the advantage they have is they have a lot of users, and I think they will just pivot. Like Anthropic, I think, pivoted. I don’t think they originally planned to work on code, but it happened that they found, okay, this is a nice niche and now we are comfortable in this niche and we push on this niche. And I can see the same thing once… Let’s say hypothetically speaking, I’m not sure if it will be true, but let’s say Google takes all the market share of the general chatbot. Maybe OpenAI will then be focused on some other sub-topic. They have too many users to go away in the foreseeable future, I think.

Lex Fridman

我觉得 Google 随时准备说「hold my beer」,用 AI mode。

I think Google is always ready to say, “hold my beer,” with AI mode.

Nathan Lambert

问题是公司能不能撑住估值。我会看见 AI 公司某种程度上被看成 AWS、Azure 和 GCP,都在同一空间竞争,都是非常成功的业务。有一种可能是 API 市场那么不赚钱,他们沿栈上下走到产品和硬件。他们现金那么多,可以建电厂、建数据中心,现在是耐久优势。但也有合理结果:这些 API 对开发者那么有价值、那么灵活,变成类似 AWS 的东西。但 AWS 和 Azure 也会有这些 API,所以五六个人在 API 市场竞争很难。也许那是他们被挤出去的原因。

I think the question is if the companies can support the valuations. I’d see the AI companies being looked at in some ways like AWS, Azure, and GCP, which are all competing in the same space and all very successful businesses. There’s a chance that the API market is so unprofitable that they go up and down the stack to products and hardware. They have so much cash that they can build power plants and build data centers, which is a durable advantage now. But there’s also just a reasonable outcome that these APIs are so valuable and so flexible for developers that they become the likes of something like AWS. But AWS and Azure are also going to have these APIs, so five or six people competing in the API market is hard. So maybe that’s why they get squeezed out.

Lex Fridman

你提到「RIP LLaMA」。Meta 有没有赢的路径?

You mentioned “RIP LLaMA.” Is there a path to winning for Meta?

Nathan Lambert

我觉得没人知道。他们动得很多,所以他们在和 Black Forest Labs 签许可——图像生成——或 Midjourney。所以某种程度上,产品和面向消费者的 AI 前线,现在下结论太早。我觉得他们有一些出色、非常有动力、靠近扎克伯格的人。所以仍有故事要展开。Llama 有点不同,Llama 是组织最聚焦的表达。我不觉得 Llama 还会被支持到那个程度。我觉得它对它们是非常成功的品牌,所以他们仍可能做开源生态的某种参与,或把 Llama 品牌续进不同服务,因为人知道 Llama 是什么。

I think nobody knows. They’re moving a lot, so they’re signing licensing deals with Black Forest Labs, which is image generation, or Midjourney. So I think in some ways, on the product and consumer-facing AI front, it’s too early to tell. I think they have some people that are excellent and very motivated being close to Zuckerberg. So I think that there’s still a story to unfold there. Llama is a bit different, where Llama was the most focused expression of the organization. And I don’t see Llama being supported to that extent anymore. I think it was a very successful brand for them, so they still might do some part of participation in the open ecosystem or continue the Llama brand into a different service, because people know what Llama is.

Lex Fridman

你觉得会有 Llama 5 吗?

You think there’s a Llama 5?

Nathan Lambert

不是开权的。

Not an open weight one.

Sebastian Raschka

有意思。稍微回顾,Llama 是开权模型的先驱——Llama 1、2、3,很多爱。但我觉得然后发生的,只是假设或推测,是 Meta 的领导人、上层高管对 Llama 非常兴奋,因为他们看见社区里它多受欢迎。然后问题是试图用开源制造更大水花。感觉被迫,像开发这些非常大的 Llama 4 模型只为站上基准顶部。但我不觉得 Llama 模型的目标是站上基准顶部去打比如 ChatGPT 或其他模型。我觉得目标是有一个人们能用、能信任、能改、能理解的模型。那包括有更小的模型;它们不必是最好的模型。发生的是这些模型——当然基准暗示它们比实际更好,因为我觉得他们有针对偏好训的特定模型,好在基准上表现好。有点过拟合,强迫它当最好。但同时他们没做人们能用的小模型,也没人能跑这些大模型。然后有点怪。我觉得就是因为人对推前沿的头条太兴奋了。我觉得就是这样。

It’s interesting. Just to recap a bit, Llama was the pioneering open-weight model — Llama 1, 2, 3, a lot of love. But I think then what happened, just hypothesizing or speculating, is that the leaders at Meta, like the upper executives, got very excited about Llama because they saw how popular it was in the community. Then I think the problem was trying to use the open source to make a bigger splash. It felt forced, like developing these very big Llama 4 models just to be on the top of the benchmarks. But I don’t think the goal of Llama models is to be on top of the benchmarks beating, let’s say, ChatGPT or other models. I think the goal was to have a model that people can use, trust, modify, and understand. That includes having smaller models; they don’t have to be the best models. What happened was just these models were — of course, the benchmarks suggest that they were better than they were because I think they had specific models trained on preferences so that they performed well on the benchmarks. That’s kind of this overfitting thing to force it to be the best. But then at the same time, they didn’t do the small models that people could use, and no one could run these big models. And then there was kind of a weird thing. I think it’s just because people got too excited about headlines pushing the frontier. I think that’s it.

4:08:08AI 曼哈顿计划Manhattan Project for AI

Lex Fridman

你喜欢 2025 年美国的 AI Action Plan 吗?里面包括开源。白宫 AI Action Plan 有专门一节标题是「鼓励开源和开权 AI」,定义这类模型,论证它们对创新和创业公司有独特价值。

Do you like the 2025 America’s AI Action Plan? That includes open source stuff. The White House AI Action Plan includes a dedicated section titled “Encourage Open-Source and Open-Weight AI,” defining such models and arguing they have unique value for innovation and startups.

Nathan Lambert

喜欢。AI Action Plan 只是一份计划,但我觉得它也许是这届政府出来的最连贯的政策文件,我希望它大体成功。我认识参与过的人。挑战是把政策做成真的,作为 AI 研究者我完全不知道怎么做,但那里面很多东西非常真实。这个国家有巨大的 AI 建设,虽然人会听到从用水到随便什么的问题,我们应该能在这个国家建东西而不把地方毁了。值得花精力。我觉得那是联邦政府的角色。他们定议程。把议程定成开权应是第一考虑,是他们能做的很大一部分,让人开始想这件事。

Yeah. The AI Action Plan is just a plan, but I think it’s maybe the most coherent policy document that has come out of the administration, and I hope that it largely succeeds. I know people that have worked on it. The challenge is taking policy and making it real, and I have no idea how to do this as an AI researcher, but largely a lot of things in that were very real. There’s a huge build-out of AI in the country, and while there are issues people hear about, from water use to whatever, we should be able to build things in this country without ruining places in the process. It’s worthwhile to spend energy on. I think that’s a role for the federal government. They set the agenda. And setting the agenda so that open-weight should be a first consideration is a large part of what they can do to get people thinking about it.

Sebastian Raschka

对教育和人才也很重要。否则如果只有闭源模型,下一代贡献者怎么来?你只能加入公司之后才学,但到那一步,你怎么识别和雇有才华的人?我觉得开源对教育人口、训练下一代研究者是根本必要的。是唯一的路。

Also, for education and talent, it’s very important. Otherwise, if there are only closed models, how do you get the next generation of people contributing? You would only be able to learn after you joined a company, but at that point, how do you identify and hire talented people? I think open source is essential for educating the population and training the next generation of researchers. It’s the only way.

Nathan Lambert

我本可以把这个讲得更病毒的方式是讲一个中国 AI 与威权国家整合、变成 ASI、接管世界的故事,因此我们需要自己的美国模型。但我谈美国的创新和科学是非常有意的,因为我觉得那既是更现实的结果,也是我想显化的世界。

The way that I could’ve gotten this to go more viral was to tell a story of Chinese AI integrating with an authoritarian state, becoming ASI and taking over the world, and therefore we need our own American models. But it’s very intentional why I talk about innovation and science in the US, because I think it’s both more realistic as an outcome and it’s a world that I would like to manifest.

我的论点是我们应该处在领先位置。但值得说得这么简单,因为 AI 生态里仍有声音说我们应考虑因安全风险禁止放开源模型。值得补充的是,实际上那不可能,除非美国有自己的防火长城,而那出了名不太好用。训这些模型的成本,无论是一百万到一亿美元,对世界上大量想有影响力的人是够得着的,所以这些模型会在全世界被训。我们想让这些信息和工具在全世界自由流动、流进美国,好让人能用、能从中学。拦住那会是对我们互联网如此大的重组,看起来不可能。

My argument is that we should be in a leading position. But I think it’s worth saying it so simply because there are still voices in the AI ecosystem that say we should consider banning the release of open models due to the safety risks. And I think it’s worth adding that, effectively, that’s impossible without the US having its own Great Firewall, which is known to not work that well. The cost for training these models, whether it’s one to a hundred million dollars, is attainable to a huge amount of people in the world that want to have influence, so these models will be getting trained all over the world. We want this information and these tools to flow freely across the world and into the US so that people can use them and learn from them. Stopping that would be such a restructuring of our internet that it seems impossible.

Sebastian Raschka

你觉得也许来自中国的大开权模型其实对美国公司是好事?你之前提他们开源放出的通常落后一代。比如 gpt-oss-120b 也许不是刀刃模型,或 Gemini 3 也许不是,因为他们想确保它安全。但当这些公司看见 DeepSeek-V3.2 真的很棒、被用、没有反弹或安全风险,那可能鼓励他们放更好的模型。也许那是非常正面的事。

Do you think maybe the big open-weight models from China are actually a good thing for US companies? You mentioned earlier they are usually one generation behind in terms of what they release open source. For example, gpt-oss-120b might not be the cutting-edge model, or Gemini 3 might not be, because they want to ensure it is safe. But when these companies see that DeepSeek-V3.2 is really awesome and is being used with no backlash or security risk, that could encourage them to release better models. Maybe that is a very positive thing.

Nathan Lambert

百分之百。这些中国公司发动了一些我觉得如果他们不都在放模型、可能不会发生的事。我几乎确定那些讨论领导层有过。

A hundred percent. These Chinese companies have set things into motion that I think would potentially not have happened if they were not all releasing models. I’m almost sure that those discussions have been had by leadership.

4:14:42英伟达、GPU 和 AI 计算集群的未来Future of NVIDIA, GPUs, and AI compute clusters

Lex Fridman

硬件这边我们提了很多次英伟达。你觉得黄仁勋和英伟达会继续赢吗?

On the hardware side, we mentioned NVIDIA a bunch of times. Do you think Jensen and NVIDIA are going to keep winning?

Sebastian Raschka

我觉得他们的下行是必须大量迭代、大量制造。他们在创新,但总有可能有人做根本不同的事,非常走运然后做成。但问题是采用。英伟达的护城河大概不只是 GPU,更像 CUDA 生态,演化了二十年。我读研究生时就在实验室做生物物理模拟、分子动力学,那时我们就有 Tesla GPU 只为计算。现在十五年了。他们建了很久,那是护城河。不是芯片本身。虽然他们现在有钱迭代、建造、缩放,真正在兼容性上。如果你作为公司到那个规模,为什么要走风险的东西,他们一年只能做几块芯片?你走大的那个。但然后我觉得有了 LLM,设计类似 CUDA 的东西会更容易。花了 15 年因为难,但现在有 LLM,我们也许能复制 CUDA。

I think they have the downside that they have to iterate a lot and manufacture a lot. They do innovate, but I think there’s always the chance that there is someone who does something fundamentally different, who gets very lucky and then does something. But the problem is adoption. The moat of NVIDIA is probably not just the GPU; it’s more like the CUDA ecosystem, and that has evolved over two decades. Even back when I was a grad student, I was in a lab doing biophysical simulations, molecular dynamics, and we had a Tesla GPU back then just for the computations. It was fifteen years ago now. They built this up for a long time and that’s the moat. It’s not the chip itself. Although they have the money now to iterate, build, and scale, it’s really on the compatibility. If you’re at that scale as a company, why would you go with something risky where it’s only a few chips that they can make per year? You go with the big one. But then I do think with LLMs now, it will be easier to design something like CUDA. It took 15 years because it was hard, but now that we have LLMs, we can maybe replicate CUDA.

Lex Fridman

我想知道训练和推理算力会不会分开,随着我们更稳定、推理需要越来越多算力。

I wonder if there will be a separation of the training and the inference compute, as we stabilize a bit more and more compute is needed for inference.

Nathan Lambert

那本该是收购 Groq 的要点。也是 Vera Rubin 的一部分——他们有一块几乎没有高带宽内存、或很少的新芯片,而那是最贵的零件之一。它为 pre-fill 设计,推理里你本质上做大量矩阵乘法的那部分,然后你只在做自回归生成、有 KV cache 交换时才需要内存。所以他们有这块为那个特定用例设计的新 GPU,每 flop 的拥有成本其实低得多。但我觉得英伟达的命运仍在于 AI 的扩散。他们最大的客户仍是这些超大规模公司,无论 Google——显然能做 TPU——Amazon 做 Trainium,还是 Microsoft 试图做自己的东西。只要 AI 进展节奏高,英伟达的平台最灵活,人会要那个。但如果停滞,做定制芯片就有更多时间。

That’s supposed to be the point of the Groq acquisition. And that’s why part of what Vera Rubin is — where they have a new chip with no high-bandwidth memory, or very little, which is one of the most expensive pieces. It’s designed for pre-fill, which is the part of inference where you essentially do a lot of matrix multiplications, and then you only need the memory when you’re doing this autoregressive generation and you have the KV cache swaps. So they have this new GPU that’s designed for that specific use case, and then the cost of ownership per flop is actually way lower. But I think NVIDIA’s fate lies in the diffusion of AI still. Their biggest clients are still these hyperscale companies, whether it’s Google — which obviously can make TPUs — Amazon making Trainium, or Microsoft trying to do its own things. As long as the pace of AI progress is high, NVIDIA’s platform is the most flexible and people will want that. But if there’s stagnation, then with creating bespoke chips, there’s more time to do it.

Lex Fridman

有意思的是英伟达相当积极地试图开发各种各样的不同产品。

It’s interesting that NVIDIA is quite active in trying to develop all kinds of different products.

Nathan Lambert

他们试图创造会用掉大量 GPU 的商业价值区域。

They try to create areas of commercial value that will use a lot of GPUs.

4:22:48人类文明的未来Future of human civilization

Sebastian Raschka

我觉得仍会是计算,伞形术语「计算」。我不必然觉得即使 100 或 200 年后会是 AI。很可能仍是计算机。我们现在更好地利用计算机,但事实是计算。

I think it would still be computing, like the umbrella term “computing.” I don’t necessarily think that even 100 or 200 years from now it would be AI. It could still well be computers. We are now taking better advantage of computers, but it’s the fact of computing.

Lex Fridman

基本上是摩尔定律那种讨论。甚至 CUDA 和 GPU 的细节都不会被记住,也不会有所有这些软件动荡。会只是,显然,计算。

It’s basically a Moore’s Law kind of discussion. Even the details of CUDA and GPUs won’t even be remembered, and there won’t be all this software turmoil. It’ll be just, obviously, compute.

Nathan Lambert

我大体同意,但是互联网的连接性和计算能合并吗?还是两者都是?

I generally agree, but is it the connectivity of the internet and compute able to be merged? Or is it both of them?

Sebastian Raschka

我觉得互联网大概会和通信相关——可能是电话、互联网或卫星。计算更像缩放那一面。

I think the internet will probably be related to communication — it could be a phone, internet, or a satellite. And compute is more like the scaling aspect of it.

Nathan Lambert

我觉得人的连接对它非常根本。你想找到世界上某件事最好的那个人,他们在世界某处。能有那种信息流动——AI 也会依赖这个。我一直盯着我说关于一个中心模型的梦死了;正在演化的是人有很多智能体做不同任务。人已经开始对不同任务用不同的 Claude。被描述为数据中心里许多 AGI,每一个管理、它们互相说话。那如此依赖网络和计算之上的信息自由流动。但网络,尤其和 GPU,是计算缩放的一部分。GPU 和数据中心需要互相说话。

I think the connection of people is very fundamental to it. You want to find the best person in the world for something, they are somewhere in the world. Being able to have that flow of information — AIs will also rely on this. I’ve been fixating on when I said the dream was dead about the one central model; the thing that is evolving is that people have many agents for different tasks. People already started doing this with different Claudes for different tasks. It’s described as many AGIs in the data center where each one manages and they talk to each other. That is so reliant on networking and the free flow of information on top of compute. But networking, especially with GPUs, is such a part of the scaling of compute. The GPUs and the data centers need to talk to each other.

Lex Fridman

你觉得神经网络这件事有非常特定、单一的突破吗?像天才之举,你基本上以非常粗糙的方式复制人脑、人心的结构?

Do you think there’s something very specific and singular to the fact that it’s neural networks that’s seen as a breakthrough? Like a genius move where you’re basically replicating, in a very crude way, the structure of the human brain, the human mind?

Sebastian Raschka

我觉得没有人心,我们大概不会有神经网络,因为它是它们的灵感。但另一端,我觉得就是那么不同。数字对生物,所以大概会更多地被归为一种算法。碰巧是这个更高效、更好用。本可以是遗传计算,像遗传算法,只是并行化。

I think without the human mind, we probably wouldn’t have neural networks because it was an inspiration for them. But on the other end, I think it’s just so different. It’s digital versus biological, so I think it will probably be more grouped as an algorithm. It could have well been genetic computing, like genetic algorithms, just parallelized. It just happens that this is more efficient and works better.

Nathan Lambert

如果你想 100 年,社会可以被更多算力和智能改变得更多,因为自主。但看这个,工业革命我们记住什么?我们记住引擎——大概相当于这里的计算机。但还有很多其他物理转变人知道,像轧棉机和所有这些机器仍被知道——空调、冰箱。AI 带来的有些东西仍会被知道;「transformer」这个词很可能仍被知道。我猜深度学习一定仍被知道,但 transformer 也许 100 年后会随着到处都是 AI 研究者而演化离开。但我觉得深度学习很可能是被记住的术语。

If you think of it over 100 years, society can be changed more with more compute and intelligence because of autonomy. But looking at this, what are the things from the Industrial Revolution that we remember? We remember the engine — it is probably the equivalent of the computer in this. But there’s a lot of other physical transformations that people are aware of, like the cotton gin and all these machines that are still known — air conditioning, refrigerators. Some of these things from AI will still be known; the word “transformer” could still very well be known. I would guess that deep learning is definitely still known, but the transformer might be evolved away from in 100 years with AI researchers everywhere. But I think deep learning is likely to be a term that is remembered.

Lex Fridman

我想知道 AI 带来的未来的空调和制冷是什么。如果我们往前走 100 年,你觉得什么不同?世界看起来怎样?首先,你觉得还有人吗?你觉得到处都有机器人在走吗?

I wonder what the air conditioning and the refrigeration of the future is that AI brings. If we travel forward 100 years from now, what do you think is different? How does the world look? First of all, do you think there’s humans? Do you think there’s robots everywhere walking around?

Sebastian Raschka

我觉得会有为某些任务的专门机器人。也许半人形。我们会看。某些事会有人形机器人,因为就是对环境友好。但对某些任务也许没道理。更难想象的是我们怎么和设备交互、人用它们做什么。我相当确定不会是手机或笔记本。会是植入物吗?

I do think there will be specialized robots for certain tasks. Maybe half-humanoid. We’ll see. I think for certain things, yes, there will be humanoid robots because it’s just amenable to the environment. But for certain tasks, it might not make sense. What’s harder to imagine is how we interact with devices and what humans do with them. I’m pretty sure it will not be the cellphone or the laptop. Will it be implants?

Lex Fridman

必须是脑机接口,对吗?100 年后,必须——以我们现在看见的进展——必须有,除非我们对如何与现实交互有真正彻底的改变。

It has to be brain-computer interfaces, right? 100 years from now, it has to — given the progress we’re seeing now — there has to be, unless there’s legitimately a complete alteration of how we interact with reality.

Sebastian Raschka

另一边,你想车,车比 100 年老,对吗?仍是同一界面。我们没有用别的东西取代车;我们只是把它们做得更好。但仍是方向盘,仍是轮子。

On the other hand, if you think of cars, cars are older than 100 years, right? And it’s still the same interface. We haven’t replaced cars with something else; we just made them better. But it’s still a steering wheel, it’s still wheels.

Nathan Lambert

我觉得我们仍会随身带一块物理的计算砖——因为人想有某种私人界面的能力。你也许不会像手机那样频繁用它,但有个东西你可以有属于你的私人信息,作为你和互联网其余部分之间的界面,我觉得仍会存在。也许看起来不像 iPhone,也许用得少得多,但我仍预期人会带东西走。

I think we’ll still carry around a physical brick of compute — because people want some ability to have a private interface. You might not engage with it as much as a phone, but having something where you could have private information that is yours as an interface between you and the rest of the internet is something I think will still exist. It might not look like an iPhone, and it might be used a lot less, but I still expect people to carry things around.

Sebastian Raschka

关于那个还有一件:我也觉得让我们和 AI 非常不同、我为什么不担心 AI 接管的,是你说的意识。我们人,我们决定我们想做什么。AI 以目前的实现,我看不见它改变。你必须告诉它做什么。所以你仍有主体性。它不从你这里拿走主体性,因为它变成工具。你告诉它做什么。会比之前的工具更自动。当然比锤子更强,它能把事情搞清楚,但仍是你在管,对吗?所以 AI 不在管,你在管。你告诉 AI 做什么,它为你做。

One thing about that is also what I do think makes us very different from AI and why I don’t worry about AI taking over is, like you said, consciousness. We humans, we decide what we want to do. AI in its current implementation, I can’t see it changing. You have to tell it what to do. And so you still have the agency. It doesn’t take the agency from you because it becomes a tool. You tell it what to do. It will be more automatic than other previous tools. It’s certainly more powerful than a hammer, it can figure things out, but it’s still you in charge, right? So the AI is not in charge, you’re in charge. You tell the AI what to do and it’s doing it for you.

Lex Fridman

所以在奇点之后、人和机器的末日战争里,你是说人值得为之战斗?

So in the post-singularity, post-apocalyptic war between humans and machines, you’re saying humans are worth fighting for?

Sebastian Raschka

百分之百。Terminator 电影他们 80 年代就拍了。我能看见出错的唯一事情当然是,如果东西被明确编程去做有害的事。

100%. The movie Terminator, they made in the ’80s, essentially, and I do think the only thing I can see going wrong is, of course, if things are explicitly programmed to do things that are harmful.

Lex Fridman

其实在 Terminator 那种设置里,我觉得人赢。我觉得我们太聪明。很难解释我们怎么搞清楚,但我们会。我们大概会用本地 LLM、开源 LLM,来帮着打机器。抱歉这种荒唐。Nathan,我做你粉丝很久了。Sebastian,我也做你粉丝很久了,终于见到你们是荣幸。谢谢你们放到世界里的一切。谢谢你们在写的出色的书。谢谢你们教我们。谢谢今天的谈话。这很好玩。

I think actually in a Terminator type of setup, I think humans win. I think we’re too clever. It’s hard to explain how we figure it out, but we do. And we’ll probably be using local LLMs, open source LLMs, to help fight the machines. I apologize for the ridiculousness. Nathan, I’ve already been a big fan of yours for a long time. And I’ve been a big fan of yours, Sebastian, for a long time, so it’s an honor to finally meet you. Thank you for everything you put out into the world. Thank you for the excellent books you’re writing. Thank you for teaching us. And thank you for talking today. This was fun.

Sebastian Raschka

谢谢邀请我们来这里,有这种人类连接,这其实——

Thank you for inviting us here and having this human connection, which is actually—

Lex Fridman

极其宝贵的人类连接。感谢收听这场与 Sebastian Raschka 和 Nathan Lambert 的对话。用爱因斯坦的话结束:「并不是我有多聪明,而是我在问题上待得更久。」谢谢收听,希望下次见。

Extremely valuable human connection. Thanks for listening to this conversation with Sebastian Raschka and Nathan Lambert. And now let me leave you with some words from Albert Einstein: “It is not that I’m so smart, but I stay with the questions much longer.” Thank you for listening, and hope to see you next time.