投稿 视频

看懂下一波 AI:OpenAI 产品负责人 Tibo 访谈

How to Understand the Next Wave of AI Before Everyone Else | Tibo Interview

原始信息 · SOURCE How to Understand the Next Wave of AI Before Everyone Else | Tibo Interview

视频 作者 / 主持:Matthew Berman 来源:YouTube · Matthew Berman 发布: 时长:44 分钟(44:28) 原文语言:英文 youtube.com

  • Thibault Sottiaux(Tibo) — OpenAI 产品负责人 / Codex 负责人 · 主页
  • Matthew Berman — 主持人 · 主页
摘要 · SUMMARY

OpenAI 产品负责人、Codex 负责人 Tibo(Thibault Sottiaux)对 Matthew Berman 说:ChatGPT 与 Codex 正在合并,因为下一代模型本来就要跑在同一套 harness 上,终点是按人定制的「个人 AGI」,而不是程序员界面 vs 非技术界面。Codex 已达约 2000 万用户;他把增长更多归因于把能力铺进 ChatGPT、让产品经理、设计师、销售都能用,而不是盯着 Anthropic。他认为笔记本会成为智能体瓶颈,工作将转向云端智能体;Ultra Fast(官方约 14 倍)应把人拉回心流,而不是同时开十几个 Agent。OpenAI 用更强模型(Soul)把 Luna 价格砍了约 80%,普通推理三个月快了约 60%,他把这看成对基础设施(含 CUDA kernel、推理栈)的递归自我改进。额度重置有一颗实体按钮,不跟市场或财务审批挂钩。Sam Altman 暂停最前沿 RL,是为了让安全团队先加固系统。Ultra Fast 一两年后可能接近默认,但大部分产能会留给客户,而不是员工。

English summary

OpenAI Head of Product and Codex lead Thibault “Tibo” Sottiaux tells Matthew Berman that ChatGPT and Codex are merging because future models want the same harness: one personal AGI that adapts, not a coder UI versus a non-technical UI. Codex has reached 20 million users; Tibo credits distributing the tech through ChatGPT more than watching Anthropic. He argues laptops, built for human typing and a handful of apps, will bottleneck agents that can juggle ~100 applications, so work moves to cloud agents. Ultra Fast (stated around 14×) should return developers to flow instead of babysitting 10–15 parallel agents. Soul helped cut Luna’s price by about 80%, and ordinary inference is roughly 60% faster than three months ago — recursive self-improvement aimed at infrastructure (CUDA kernels, the inference stack), not only model-begets-model research. He has a physical reset button, unused in concert with marketing or finance. The frontier RL pause was a safety/alignment hardening step. Ultra Fast may sit near the default in a year or two; most of that capacity is reserved for customers.

时间轴 · 16 个章节
  1. 00:00 开场
  2. 00:45 从 Google 与 DeepMind 学到的
  3. 04:22 打造 OpenAI 的文化
  4. 07:23 AI 智能体的未来
  5. 11:18 AI 如何改变开发者工作流
  6. 14:27 ChatGPT 与 Codex 合并
  7. 17:00 人机交互的未来
  8. 20:25 OpenAI 与 Anthropic
  9. 23:37 为什么 OpenAI 总在重置额度
  10. 26:41 AI 效率与算力
  11. 30:25 递归式自我改进
  12. 32:00 暂停前沿 AI 训练
  13. 34:13 Ultra Fast 解锁了什么
  14. 40:00 Ultra Fast 会不会变成默认?
  15. 41:19 如何安抚对 AI 感到紧张的人
  16. 43:20 为什么每个人都该试试 AI

00:00开场Intro

片头先剪了后文几句(实体重置按钮、不太看竞争对手、一两年后今天的速度可能接近默认、Luna 成本、技术会变得非常高效)。然后进入正片。

Matthew Berman

Tibo,非常感谢你来。

Tibo, thank you so much for joining me.

Tibo

当然。很高兴来。

Of course. Glad to be here.

Matthew Berman

非常期待这场对话。

Really excited to talk to you.

00:45从 Google 与 DeepMind 学到的Lessons from Google & DeepMind

Matthew Berman

我想先从你在 Google 的经历聊起。你当时在 DeepMind 团队。ChatGPT 出现之前,Google 内部有一个叫 LM Chat 的东西。你发过推,说 Google 当时太紧张,不敢把它放出来;DeepMind 也被拦住,不能推出可能颠覆 Google 的产品。这件事我经常会想。那时候你正在做这些产品,还远在 ChatGPT 真正改变世界之前,你当时是怎么想的?

I want to start with your time at Google. You were on the DeepMind team, and before ChatGPT, Google had something called LM Chat. You had tweeted that Google was too nervous to release it — DeepMind was blocked from shipping products that could disrupt Google. I think about that a lot. What were you thinking while you were working on those products, well before ChatGPT really changed the world?

Tibo

那是一段非常令人兴奋的时期。DeepMind 是一个极具创造力的地方。我的专长是基础设施和产品,用来加速研究。当时有一组人在做语言模型和规模化,已经拿到相当不错的结果。接下来很自然就会问:能不能把它变成一个可以对话、能拿来做各种事情的东西?于是类似 LM Chat 的想法就出现了。先是内部项目,后来也有把它做成公开工具的野心。

It was a very exciting time. DeepMind was a very creative place. My specialty was infrastructure and products to accelerate research. There was a group working on language models and scaling them. They’d gotten pretty good results, and it was natural to ask: can you turn that into something you can chat with, and use for various things? So the idea of something like LM Chat emerged. It was internal, and then there was an ambition to make it a publicly available tool.

Matthew Berman

这大概是哪一年?

What year was this?

Tibo

大概是 ChatGPT 发布前一年左右。我们当时还在做很多别的东西,那些我就不展开了。那确实是一个很有创造力的地方。只是 DeepMind 并不是一个被设计来交付产品的组织。OpenAI 在这一点上完全不同。研究和产品合作非常紧密:一起想点子,共同设计很多东西。我们有很强的交付倾向,也很想尽快把东西交到人手上。这是我非常喜欢的,也是把我吸引过来的原因:使命、人、人才密度。OpenAI 值得说的好东西太多了。

About a year before ChatGPT, roughly. We were also building all sorts of other things that I’m not going to talk about. It was a very creative place. DeepMind was not set up to ship product. OpenAI is a very different place in that sense. Research and product collaborate super closely. We ideate together, we co-design a lot of things. We have a big bias to ship, and a big bias toward making things available for people, which I really love. That’s what drove me here: the mission, the people, the talent density. There are so many great things about OpenAI.

Matthew Berman

你当时参与 LM Chat 的时候,有没有意识到它很特别,或者它将来会变得特别?

Did you know at the time, while you were involved in LM Chat, that it was something special, or would become something special?

Tibo

当时确实感觉它非常特别。那是第一次你意识到:模型可以生成连贯的文字,而且能提供一些真正有用的东西。最开始它更多是好玩,而不是有用;后来才一点点变得越来越有帮助。

It felt very special. It was the first time you realized you could get coherent text, and something helpful. Initially it was more funny than helpful, and then gradually it became more and more helpful.

Matthew Berman

你说你经常会想起这件事,我能理解。某种程度上,Google 是自己挡住了自己。你从那里带到 OpenAI 的教训是什么?

You say you think about that often, and I understand that. In a lot of ways Google got in their own way. What are some of the lessons you learned there that you took to OpenAI?

Tibo

所以我才会经常想这件事——想到我自己团队的文化,也想到 OpenAI 整体的文化:哪些好的要保住,哪些不要再做。OpenAI 是一种非常自下而上的文化,也是充分授权的文化。大家可以提出各种想法,聚在一起,然后很快把东西做出来。面对新产品想法,组织里整体上很少有强大的「踩刹车」力量。这非常令人兴奋,也很有趣,因为大家是在用正面的方式影响世界。把这一点保住,对我很重要。

That’s why I think about it often — in terms of the culture I have on the team, and the culture of OpenAI itself: the good parts to preserve, and what not to do. OpenAI has a very bottom-up culture. It’s a very empowering culture. People can come up with all sorts of ideas, get together, and very quickly ship something. There’s very little stop energy, in general, for new product ideas. That’s exhilarating and fun, and it’s all about impacting the world in positive ways. Preserving that is very important to me.

另一件同样重要的事,是别把它搞成一团乱。你不会希望产品变成一堆功能的杂烩,没有整体方向,也没有一致性。所以要用简洁、以及对产品质量的骄傲来对冲。我觉得 ChatGPT 的 iOS 应用是市面上最好的应用之一,我们想把这个水准保住。我们在愉悦感、性能、效率、简洁上投了很多。这些是总原则,同时仍然让每个人都能试新东西、快速交付。

The other thing that’s also important is not to make it a mess. You don’t want a hodgepodge of features with no overall direction and coherence. That’s counterbalanced with a sense of simplicity, and being proud about the quality of the product. I think the ChatGPT iOS app is one of the best apps out there, and we want to keep that. We’re investing a lot in delight, performance, efficiency, simplicity. Those are the overall principles, while still empowering everyone to try new things and ship very quickly.

04:22打造 OpenAI 的文化Building OpenAI’s Culture

Matthew Berman

如果让你给创业者提建议,告诉他们怎么养成这种文化,OpenAI 内部有哪些更具体、可落地的做法?

If you were to give advice to a founder about how to develop that kind of culture, what are some of the more tangible elements or practices inside OpenAI?

Tibo

要有坚定的判断。要想办法让产品真正到用户手里,并根据反馈快速迭代。还要愿意颠覆自己。这一点对创始人没那么直接,但对 OpenAI 这样的公司非常相关。我们不断有新的研究、新的想法。关键是判断什么时候该投入,即使这意味着可能要从现在的主营业务里重新调配资源。这很难,但必须能做到。

Having conviction. Finding a way to have users and iterate very quickly from feedback. And being willing to disrupt yourself. That’s not as much relevant for a founder, but it is relevant for companies like OpenAI. We come up with new research and new ideas all the time. Being able to identify when is the right moment to go and invest in them — even though it means maybe reallocating resources from the main gig — is super important. It’s very hard, but super important to be able to do that.

Matthew Berman

这正是你刚才描述 Google 时的问题。他们当时好像没能做到这一点。

That’s the exact thing you were describing at Google. They kind of weren’t able to do that.

Tibo

公平地说,他们是有计划的,一切都在一个大计划里。只是对我来说,那不是对的地方。

To be fair, they had a plan. It was all part of a big plan. To me it wasn’t the right place.

Matthew Berman

对 OpenAI、或者任何正在成熟的公司来说,保持这种快速交付、愿意颠覆自己的文化,会不会越来越难?尤其是当你已经有一个在印钞的现金牛,旁边又冒出一个可能很酷、很创新的新东西。

At OpenAI, or any company as it matures, does that become more difficult to maintain — that culture of shipping and willingness to disrupt yourself — especially if you have a cash cow printing money, and this other new thing over here that might be cool and innovative?

Tibo

我们非常向前看。AI 的未来会长成什么样、人类最终能从中得到什么,并不会因为你过去一个月、三个月已经建成了什么,就停下来等你。所以必须真正靠上去,睁大眼睛看它往哪走,然后想办法让自己站上那一浪。

We are very, very forward-looking. The future of AI, and what it will all look like, and how humanity benefits, doesn’t really wait, and doesn’t really care for whatever you have established here over the next month or three months. So I think it’s very important to lean in, to be open-eyed about where it’s all going, and then figure out how to position yourself so you do catch that wave.

即便在 OpenAI 也是这样:我们训练模型,然后才发现它们的能力。基准测试不会告诉你全部。我们必须自己跟模型多玩一玩,才会突然意识到:哦,也许我们还没想过可以这样用;或者,哦,它居然能做这个。这会改变我们对产品的想法。

Even for OpenAI: we train models and then we discover their capabilities. Benchmarks don’t tell you everything. We have to play quite a bit with the models themselves to realize, oh, maybe we haven’t thought about benefiting from it in this specific way — or, oh, it can do this. And that’s a shift in how we think about the product.

比如现在推出的新语音,聊起来非常舒服,已经很自然,而且也能调用工具。这就改变了一些事。我现在会花更多时间直接跟它说话。另一件我一直在做的,是听写——听写质量非常好,比打提示词高效得多。早上我就拿着手机坐在那儿,把当天要让 ChatGPT 做的几件事一股脑说出来,然后它就去干了,它能用上我所有的工具。在我们还没有足够好的语音模型之前,这根本做不到。一旦有了,你对产品的想象就会完全变掉。

For example, right now we launched the new voice, and it’s super delightful to talk to. It’s very natural now. It’s capable of tool use as well. That changes things. Now I spend a lot more time just talking to it. Another thing I do all the time is dictation, because the quality of the dictation is so, so good, and it’s much more efficient as a way instead of typing the prompt. In the morning I sit there with my phone and I’m like, blah blah blah — a couple of things to do for ChatGPT — and then it just goes and does it. It has access to all my tools. That was not possible before we had really good voice models. And that completely changes how you think about the product.

07:23AI 智能体的未来The Future of AI Agents

Matthew Berman

我们继续谈新模型、新的 harness。几周前——我还是从你的另一条推开始,因为这些推真的很猛——「两三个月后再看,Codex 会显得很原始。我们马上要经历又一次重大演进。下一代模型需要的不只是你的笔记本电脑。」先从 harness 说起。模型变强之后,harness 里哪些地方仍然很值得创新?

Let’s continue talking about new models, new harnesses. A few weeks ago — I’m going to start with another one of your tweets, because these are bangers — “Codex will seem primitive in two to three months. We’re about to go through another major evolution. The next generation of models need more than your laptop.” Let’s start with the harness. What areas of the harness are still ripe for innovation as the model gets better?

Tibo

太多了。刚才说了语音。现在如果你是 Codex 或其他编程智能体的高级用户,多少已经习惯了一些别扭。你得管理 skill 文件,那是一种教它东西的方式,但很多人已经意识到,长期维护其实挺难。记忆是有的,但它并不总能记住所有事。如果你有子智能体,你还得管子智能体,整套东西会构成一个小网络,交互时那种「它是完整搭档」的幻觉会在不少环节被打破。

So many. I talked about voice. Right now, if you’re a sophisticated user of Codex — or any other coding agent — you’ve gotten used to a little of the clunkiness. You have to manage skill files. That’s a way to teach it stuff, but a lot of people have realized it’s kind of hard to maintain over time. Memory is a thing, but it doesn’t always remember everything. If you have sub-agents, you have to care about sub-agents, and it constructs a little network, and the illusion kind of gets broken at various parts when you interact with it.

你真正想要的,是一个深度理解你的东西:理解你的目标、日常、团队在干什么;最好还能反应、也能主动,帮你处理日常,并且不打破那种「这就是你那个完美小搭档」的感觉。这正是我们在做的方向。

What you really want is something that deeply understands you, understands your goals, understands your day-to-day, understands what your team is up to as well — and then, optimally, reacts and is also proactive, and just helps you in your day-to-day, and doesn’t break that illusion of this perfect little partner that you have. That’s what we’re working toward.

另一件事是:模型一旦非常非常强,笔记本电脑本身就会变成限制。笔记本能承受的工作量,本来就是按人设计的——大概按你能产出多少、打字多快、想多快、需要同时开多少应用。这些都是人的约束。模型没有同样的约束。它完全可以同时处理好一百个打开的应用,未来也许更夸张。所以从资源获取上看,很清楚:未来的模型需要的,会超过你笔记本电脑能提供的。

Another thing you realize when you have very, very powerful models is that your laptop becomes a constraint in and of itself. The amount of work you can do on a laptop — it was designed for humans. It’s designed roughly to absorb the amount of work you can produce, how fast you can type, how fast you can think, how many applications you need open. All of these are human constraints. The model doesn’t have the same constraints. The model can, for example, handle a hundred applications opened at the same time, perfectly fine — maybe in the future. So in terms of access to resources, it’s very clear that models of the future will need access to more than the resources of your laptop.

Matthew Berman

我猜你说的是云端智能体。一旦有了 Ultra Fast——我们稍后会细聊——token 速度是 Fast 的 10 倍,官方数字好像是 14 倍,10 到 14 倍,带宽约束就变了。CPU 反而成了带宽。工具调用、网络、栈里任何开销,都会变成限制因素。

I’m guessing you’re talking about cloud agents. And all of a sudden, when you have things like Ultra Fast — which we’re going to talk about in a little bit — when you have token speeds that are 10, I think 14 is the stated number, 10–14 times faster than what Fast is, the bandwidth constraint changes. The CPU now becomes the bandwidth. Literally tool calls, network, any kind of overhead in the stack becomes the limiting factor.

Tibo

不过你也可以用并发来补。你可以一边探索,一边写测试、编译、验证一个新假设,同时做。这样瓶颈会不断被挪走,因为你能并行做更多事,模型也能非常高效、非常快地想下去。

But you can compensate by doing multiple things concurrently as well. You can think about exploring on one end, writing tests as well, compiling, testing a new hypothesis, all at once. So you’re shifting the bottleneck around, because you’re able to do more concurrently, and the model can think very efficiently and very quickly through it.

11:18AI 如何改变开发者工作流How AI Changes Developer Workflows

Matthew Berman

以现在的 token 速度,我经常会并行拉起 10 个、15 个智能体,认知负担已经很大,因为你得不断切上下文。任务丢出去,通常要等 30、45 分钟才回来。有了 Ultra Fast,这条工作流会明显变掉。我觉得我不会再同时开 10 个、15 个,这也许是好事。也许一次三四个。你怎么看单人开发者的工作流会怎么变?

With current token speeds I find myself kicking off 10, 15 agents in parallel, and that becomes a pretty significant cognitive overhead — the context switching. You’re kicking it off and you can expect 30, 45 minutes before my task comes back. Now with Ultra Fast speed that workflow changes significantly, and I don’t think I would be able to have 10 or 15 agents, and that might be a good thing. Maybe it’s three or four at a time. How do you see the workflow of a solo developer changing over time?

Tibo

我们非常在意如何管理你的注意力,以及如何对注意力更友好。说到底,我们是为人类在做东西,想做最能赋能人类的技术,那就得围绕你的多任务能力来设计:你想怎么管注意力?这件事是现在就要被端上来,还是三十分钟后再来更好?

Managing your attention, and being much more friendly to your attention, is something we care a lot about. After all, we’re trying to build for humans. We’re trying to build the technology that’s the most empowering for humans, and that requires building around your ability to multitask, and how you want to manage your attention: do you want something brought up now, or is it better to bring it up in 30 minutes?

当 Ultra Fast 再叠上语音,你会突然觉得:好,这东西可以跟你一样快,甚至比你更快。于是你能留在心流里:一边构思,一边看到原型,一边实时做出小报告,那种感觉非常好。你会突然发现,以前同时管十个智能体的方式,其实不是自己真正想要的,也不想回去。我们想提供的,就是这种既自然、又像为你量身做的体验——不是你去适应技术,而是技术来适应你。

When you have Ultra Fast speeds combined, maybe with voice, suddenly you’re like: okay, this thing can operate at the same speed, if not faster than you. So you stay in the flow. You get to ideate, you get to see prototypes, you get to build little reports in real time, and that just feels really good. Suddenly you’re like, oh yeah — what I was doing before, multitasking like 10 agents, I don’t really want to go back to that. We’re trying to bring that sort of experience that is just really natural, but also feels built for you, where you don’t have to adapt. The technology adapts to you.

Matthew Berman

过去几个月大家讨论了不少智能体编程手法。Loops 很火,现在还火;最近又在听 graphs。这些是不是都在帮单人开发者管理注意力,或者像你说的,对注意力更友好?我喜欢这个说法。

There’s been a number of agentic coding techniques discussed over the last few months. Loops was popular, still is popular. Now I’m hearing about graphs. Are these all techniques to just allow the solo developer to manage, or be friendly to, their attention — as you said? I like that term.

Tibo

我会把问题分成两类。第一类,是打造最好的个人 AGI,或者说个人智能体:它跟你待在心流里,会主动提出重要的新想法,发现机会就很快动手,并且非常高效地做你真正想做的事。不管是技术问题,还是研究、建议,它都能做,而且高度适配你。这件事非常重要,根子在于理解你作为一个人、作为一个独特个体。这是一类问题,我们在非常用力地推。

I think about two different categories of problems. The first one is building the very best personal AGI, or the personal agent, that will be in the flow with you, proactive, raise important new ideas when it can find some, be very, very efficient at doing exactly what you want. It doesn’t matter whether it’s a technical problem or more like research or advice. It can do it all, and it’s super tailored to you. That’s a very important thing, and it’s deeply rooted in the understanding of you as a human, you as an individual that is unique. That’s one category of problem. We’re pushing super hard on that.

第二类是彻底自动化:你更多地是在构建能接管复杂流程的智能系统。那个流程也许确实需要智能,看起来也很复杂。比如去看生产日志,自动做性能优化;或者发现回归,自动打补丁。网络安全里我们也看到类似的事:扫描器找出一个漏洞,系统能不能自动补上,把敞口窗口压到几乎为零——循环里没有人,或者人非常少,你只需要批准高风险动作。它大体上是一套自动化系统,你也不必始终直接控制它。

The other category of problem is full-on automation, where you’re more building intelligent systems that can take care of a very complex process — maybe something that does require intelligence and seems very complex. For example, going and looking at production logs and automatically doing performance optimizations, or looking at regressions and automatically patching them. In cybersecurity we’re seeing this as well: you have a scanner that comes up with a vulnerability — can you automatically patch it and reduce the window where you have that open vulnerability to almost zero, without a human in those loops, or with very, very minimal, where you only need to approve a high-risk action. It’s mostly an automated system, and it’s also not that important for you to be in direct control of it.

14:27ChatGPT 与 Codex 合并ChatGPT & Codex Merging

Matthew Berman

我想稍微换个话题。过去几个月,ChatGPT 和 Codex 一直在合并的路上。先问进展怎么样?内部感觉如何?客户反馈是什么?

I want to slightly change topics. ChatGPT and Codex have been on this merge path over the last few months. First, how’s that been going? How does it feel internally? What’s the feedback you’ve been getting from your customers?

Tibo

这件事带来了很大帮助。最开始的反馈是:为什么要合并?真的必须吗?但未来的模型本身就希望我们合并。所以我们会做,因为这是最简单、也最正当的做法。我们在构建的是一个非常个人化、能力极强、能在各种方面帮你的智能体。底层会是同一套技术、同一套 harness、同一种思路。它高度多模态,语音优先,效率极高;你是不是在写代码并不重要,这个智能体什么都能做,而且做这些事最有效率。

It’s really been a boon. The feedback we had initially was: why do you merge them? Do you really have to do it? And it’s like, well, our future models want us to be merged. So we’re just going to do it, because it is the simple and proper thing to do. We’re building this very personal, super capable agent that can help you in all sorts of ways. This is going to be the same technology under the hood. It’s the same harness. It’s the same way that we think about it. It’s highly multimodal, voice-first, super efficient, and it doesn’t matter if you’re trying to code or not. This agent is capable of it all, and it’s the most efficient at it.

然后,界面应该按你的需要来适配。你不该先决定「我是程序员,所以我要程序员界面」,或「我不懂技术,所以我要非技术界面」。人是在一条光谱上的。「软件工程师」「设计师」这些标签,本质上只是人类为了处理过于复杂的现实而发明的抽象。可每个人都是具体的人,都在光谱上的某个位置。所以我们想做的是一个能为每个人适配的完美界面。懂不懂技术不重要,它按你这个人的独特性来适应。这就是我们去做这件事的原因。

Then the interface that you want should tailor itself to your needs. You shouldn’t decide, I’m a coder, I want a coder interface — or I’m not technical, I want a non-technical interface. There’s a spectrum of people. We come up with labels like software engineer, designer. These are just human concepts that we have invented to deal with abstractions, because the reality is too complex for us to handle. But individuals are individuals. They’re somewhere on the spectrum. So we’re trying to build the perfect interface that adapts for everyone. It doesn’t matter if you’re technical or not. It adapts based on your specific individuality. That’s why we went and we did this.

Matthew Berman

这是不是意味着,最终一定会变成一个统一界面,不再用下拉菜单在不同产品之间选?想到我妈妈可能和我用完全同一个界面,只是系统会按各自需求定制——如果我做更复杂的工作,也许需要更多信息——这个想法其实挺疯狂的。对你来说,终局是什么?

Does that mean inevitably it’s going to end up with a singular interface? No dropdown selecting between products. It’s kind of wild to think that my mom might use the same exact interface as me, and then obviously it’ll customize to my needs — maybe I’ll need more information if I’m doing more sophisticated work. What is the end state for you?

Tibo

对,就是同一个东西。你和你妈妈会用同一个东西。它会是你们各自的个人 AGI。你们要完成的任务不同,从中得到的效用也不同。你们会把它连到生活里不同的工具上,带去不同的想法和需求。然后它会持续调整自己,尽可能让你受益最大。对你的朋友、对所有人也都一样。

That’s right. It’s the same thing. You and your mom will use the same thing. It will be your personal AGI. You will have very different kinds of tasks and utility that you get from it. You will connect it to different tools in your life. You will bring different ideas, different needs. And then it will continue to tailor itself to maximally benefit you. And it will do so with your friends and with everyone else.

17:00人机交互的未来The Future of Human-AI Interaction

Matthew Berman

我想回到你刚才说的一个词。你在描述那种终局时,几次用到「幻觉 / 那种完整感觉」。对普通用户来说,那种完美的感觉是什么?如果想象几年之后,人和 AI 的交互会长成什么样?

I want to go back to something you said. You used the word “illusion” a couple times in that kind of end state. What is that perfect illusion for the typical user? If you can envision us a few years from now, what does the interaction between AI and a human look like?

Tibo

对我来说,它应该是一种高度为人类定制的东西。大语言模型能成功,也是因为自然语言。自然语言是人类的概念。我们习惯彼此说话。如果你明天给我写一封信,我能读懂;我们现在已经比较熟了,所以我还能读出一点情绪,或者信背后一点点细微的意思。这些都深深地属于人。

To me it’s something that is very, very tailored to humans. This is why large language models are also a success: it’s natural language. Natural language is a human concept. We’re used to speaking to each other. If you write me a letter tomorrow, I’ll be able to read it. We know each other quite a bit now, so I will be able to decipher a little of the emotion, or a little of the nuance behind the letter. All of that is deeply human.

所以我们在做的技术,应该扎根于人性,扎根于人类沟通和把事情做成的方式。不该经常出现这种场面:你误解我了,因为你没读出我语气里的细微差别,或者没理解我文字真正想表达的意思。我们想避免的就是这个。我们非常努力地不让你去适应技术,而是让技术被做成人类原本在世界里行事方式的自然延伸。

So the technology that we’re building is rooted in humanity, and rooted in the way that humans communicate and get things done. There shouldn’t really be a thing where you’re like, oh, you misunderstood me because you didn’t quite decipher the nuance in my tone, or you didn’t quite understand the text, how I meant it. That’s something that we’re trying to avoid. We’re trying very much not to have you adapt, but have the technology be perfectly created to be a natural extension of how humans already act in the world.

Matthew Berman

人与人交流里,很大一部分是非语言的——我怎么动手、面部怎么动。你觉得未来有多少会被人工智能感知到、读到,也许通过视觉?这重要吗?因为你刚才描述的更像纯文本。对我们这些在网上长大的人来说,很习惯用文字交流,并在文字里加细微差别来传达真正的意思和语气。AI 还能不能读面部表情、手势,这件事还重要吗?

When I think about communication between humans, so much of it is non-verbal — just the way I move my hands, the facial movements. How much of that do you see in the future being sensed by artificial intelligence, or read by artificial intelligence, maybe through vision? Is that even important? Because what you’re describing now is text only. And for those of us who grew up online, we’re very used to communicating over text and adding subtleties to that text to convey what we really mean, tone. Is it still important to have AI be able to read our facial expressions, our hand gestures, and so on?

Tibo

我觉得重要。我想象我们在构建的未来,是非常环境化、非常自然的。假如过一会儿我去办公室,在白板上写点东西,突然有了一个想法,它也应该能在场,并且理解。或者我直接对它说:嘿,你觉得这个怎么样?然后我们就用语音自然聊起来。

I think so. When I think about the future of what we’re building, it’s very ambient, it’s very natural. If later I go to my office and I write something on the whiteboard and I have an idea, it should be capable of being there as well and understanding. Or maybe I tell it, hey, what about this thing — and then we just have a natural conversation over voice.

自从我们上线新的 ChatGPT 语音,这件事真的起来了。纯粹通过语音和 ChatGPT 互动的用户数现在增长很快。我觉得教训是:每次你往更自然的方向靠,人就会选择阻力最小的那条路。就像你说的,在一个小框里打字,对一部分人也许自然,但对所有人并不是。一旦出现稍微更容易、稍微更好的方式,大家往往会转过去用。

Since we shipped the new ChatGPT voice, it’s really taken off. The amount of users that interact with ChatGPT just through voice is growing very fast right now. I think the lesson is: every time you lean into something that is more natural, humans just choose the path of least resistance. As you said, typing on a little box is natural maybe for some of us, but not for everyone. And when you get something that is just a little bit easier, a little bit better, you tend to just go and use that.

20:25OpenAI 与 AnthropicOpenAI vs. Anthropic

Matthew Berman

首先祝贺你们。我看到你今天早上发了:Codex 达到 2000 万用户。我看过那条曲线,有一阵子还比较平,后来突然几乎垂直。恭喜。

First of all, congratulations. I saw that you posted this morning: Codex reached 20 million users. I’ve seen the graph, and for a while it was like this, and then all of a sudden it’s vertical. So congratulations.

我想聊聊和 Anthropic 的竞争。很多人觉得,现在行业里最主要的两家就是 OpenAI 和 Anthropic。曾经有一段时间,Anthropic 几乎吸走了房间里所有的氧气,非常强势,后来情况好像变了。你怎么看今天的市场?

I want to talk a little bit about that competition with Anthropic, because of course a lot of people think OpenAI, Anthropic — these are the two major competitors in the industry right now. There was a period of time in which Anthropic was kind of sucking all the oxygen out of the room. They were really dominating, and then all of a sudden something changed. First of all, what’s your read on the market today?

Tibo

我们现在关注的是:打造最强的模型,打造效率极高的模型,并且很认真地为所有人做产品。我觉得 OpenAI 做得很好的一点,是在意这个世界,也在意如何把这种非常强大的技术交到尽可能多人手里。合并 Codex 和 ChatGPT,也是出于同样的愿望:我们有这项技术,可以让它更安全、更容易用——不管你是产品经理、设计师,还是销售、市场、传播,都应该能用上全部能力。然后很快通过 ChatGPT 分发,那儿已经有海量用户。这也在驱动你刚才说的那波增长。

Right now we’re focused on building the most capable models, building models that are highly, highly efficient, and then taking a lot of pride in building products for everyone. This is something I think OpenAI does really well: caring about the world, and caring about how we are taking this very, very powerful technology and putting it in the hands of as many people as possible. That’s what we did as well with merging Codex and ChatGPT. It was this desire: we have this technology, we can make it safer, we can make it easier to use for everyone — whether you’re a product manager, a designer, in sales, marketing, comms, all of that. You should be able to use all of it. Then very quickly distributed through ChatGPT, where we have a ton of users already. That’s been really driving this growth as well that you mentioned.

我不太会盯着竞争对手。我真正看的是:我们有什么能做得特别好?我们的价值观是什么?我们如何最大限度地朝那个方向加速?

I don’t tend to look at the competition that much. I really look at what can we do uniquely well, and what are our values, and how do we maximally accelerate towards that.

Matthew Berman

我想再往下挖一点点。我知道你不会花太多时间想 Anthropic,但很多人会。他们会想:我该相信哪款产品?我每个月这 200 美元该交给谁?你看 OpenAI 的市场位置、品牌、语气,以及它跟开发者、跟更广受众互动的方式,和 Anthropic 比起来,你怎么看?

I want to maybe just dig a tiny bit more into that, because I know you’re not thinking about Anthropic all that much, but a lot of other people do. They’re thinking about: okay, which product do I believe in? Which product do I want to give my $200 to? When you look at the market position and the branding and the tone from OpenAI, and just the way that it interacts with developers, with the broader audience — how do you see that comparing to the way that Anthropic does?

Tibo

我仍然很在意的是社区、为整个世界构建、把所有人带上这趟车。你可以从我们做事情的方式里感受到:我们对很多事非常透明,也会从社区吸收大量想法。说实话这也非常有趣,因为我们会从中得到很多能量。我们在构建的这项技术,不是只给我们自己用,也不是只为了加速 OpenAI。使命非常重要,这也是我们能量的来源。对我来说这种方式很落地,也很有趣,然后好事会因此发生。

Again, what I care a lot about is the community, building for the world, bringing everyone along. I think you can feel that in the way that we are super transparent about things. We take a lot of ideas from the community. It’s also so much fun, to be honest, because we get so much energy from it as well. And this technology that we’re building — we’re not building it just for ourselves. We’re not just building it to accelerate just OpenAI. The mission is super important, and therefore that’s where we also get our energy from. It feels very grounded, it feels fun, and then good things happen as a result of that.

23:37为什么 OpenAI 总在重置额度Why OpenAI Keeps Resetting Limits

Matthew Berman

那我们聊聊这些好事。我想谈谈额度重置。我知道很多人就是为了这个在盯你每一条推。具体一点,看着 Codex 那条增长曲线——这个问题也许有点傻——这些重置有多少是在帮市场营销和增长,还是说,它主要就是对开发者社区的善意?

Well, let’s talk about some of those good things. I want to talk about the resets for a second. I know it’s what everybody is following your every tweet because of this. Specifically, looking at that growth curve of Codex — maybe this is a silly question — how much of those resets is a boon towards marketing and growth, or was it just goodwill for the developer community?

Tibo

这也许有点反直觉,但 OpenAI 是一个「可以直接做事」的地方。所以最开始,当我们在迭代、把东西弄坏,或者配置错了、体验没达到我们想要的水平时,去补偿,就觉得是对的。我们会说:谢谢你试这个产品。我们知道我们在很努力地做,现在还是早期。因为我们碰巧把它弄挂了大概 30 分钟,这里有一些额外额度。我们理解这对你很重要,你也依赖它,谢谢你成为用户。事情就是这样开始的。我现在仍然这么对待它。如果我们弄坏了,或者体验不够好、我们还没完全搞清原因,我们就会补偿,重置使用限额。

Maybe it’s counterintuitive, but OpenAI is a place where you can just do things. So it just felt right initially to compensate when we were iterating and breaking things, or maybe we had misconfigured something and it wasn’t quite as good as we wanted. So: hey, thank you for trying this product. We know we’re trying very hard to build it. It’s early days. Here’s some extra usage, because we happened to break it for like 30 minutes, and we understand this is really important and you rely on it, and thank you for being a user. That’s how it started. And that’s how I still treat it. If we break it, or if the experience is suboptimal and we don’t fully understand why, we will compensate for that. We will reset the usage limits.

后来这事当然变成了一个现象。现在甚至有一整个重置按钮。但背后并没有很多审查。它不是和市场或财务一起策划的。我想按就按,只要我觉得合适。我们的原则是:我们在努力做一件很棒的事;如果它不够好,我们就补上。

Then it turned into quite the thing. Obviously there’s a whole reset button now. But there isn’t really a lot of scrutiny behind it. It’s not done in partnership with marketing or finance. I can press the button whenever I want, whenever it feels right. We have these principles: we’re trying to build something amazing, and when it is not, we will make up for it.

Matthew Berman

我仍然觉得,这件事在社区里积累了很多好感,也可能至少在一小部分上推动了增长。真正关心用户是有用的。你可以口头上说关心,也可以真的关心:如果我们弄坏了,抱歉,我们这样补。这让我想起亚马逊的退货政策:只要你有任何不满意,就把东西退回去。你好像也在给 OpenAI 建立类似的文化和认知:如果我们搞砸了,这些 token 你再用不就好了,或者我们再给你一批新的。

I still think there’s a piece of it that really has built so much goodwill in the community, and maybe has contributed to the growth at least in a small part. Caring for your users does a lot, right? You can pay lip service and say that you care, or you can actually care: if we break it, hey, really sorry about it, here’s how we make up for it. It kind of reminds me of Amazon’s return policy. If you’re not happy in any sense, go ahead and send it back. You’re kind of building that same culture, or that same perception of OpenAI: if we make a mistake, go ahead, use those tokens again, or here’s a fresh batch of tokens for you.

Tibo

我很认可这种说法。另外也有一些好的时刻,我们只是想庆祝、把这个时刻标记下来。有时候我们好像也没有更有意义的东西可以给。我们一直在发新功能,也会尽可能铺得广。但如果想和整个社区分享一件事:来,去试试这个新东西。你还没用过 Ultra Fast?这里有一些额外额度,去试。

I really appreciate it. And then there are also good moments where we just want to celebrate and mark the moment. There isn’t really something that we can give that is more meaningful at times. We always ship new features; we will ship them as broadly as we can. But something to share with the entire community: hey, go explore this new thing. You haven’t used Ultra Fast yet? Here’s some extra usage. Try it.

Matthew Berman

我听说现在真的有一个实体按钮。

I heard there’s an actual physical button now.

Tibo

确实有。

Yes, there is.

Matthew Berman

回头你得给我看看。

You’ll have to show me that after.

Tibo

我会给你看。真的非常酷。

I will show it to you. It’s very, very cool.

26:41AI 效率与算力AI Efficiency & Compute

Matthew Berman

不过能这么频繁重置,前提是你们必须做了相当充分的算力规划。你得有足够的计算资源,才能把这些重置送出去。我想借这个聊聊自我改进。说到产能——几周前,大概是几周前,你们发过一篇文章,说 Soul 优化了 Luna 的效率,把 Luna 的价格降了 80%,Terra 也降了价。Luna 这边,有多少是你们真正挤出来的效率,有多少是算力规划做得好、利润空间够、所以决定降价让更多人用?算法收益和战略规划各占多少?

But with all of these resets, you can really only do that if you’ve done significant compute capacity planning. You have to have enough compute to give all of these resets. I want to start to talk a little bit about self-improvement, because — speaking of capacity — a few weeks ago, I think it was a few weeks ago, there was this article you put out, and it stated Soul had optimized Luna efficiency. You dropped the price of Luna by 80%. There was also a price drop for Terra as well. How much of an efficiency gain were you able to eke out of Luna, versus how much of it is: we just did really great compute capacity planning, our margins are great, and we want people to use it? How much of it were algorithmic gains versus strategic planning?

Tibo

我们很早就在规划算力。往回看两年,外界还在问 OpenAI 为什么要在算力上投这么多。

We planned compute way ahead. If you look back two years, I think OpenAI was kind of questioned for why there was so much investment in compute.

Matthew Berman

那种疯狂但押对了的赌注。

One of those crazy good bets.

Tibo

是。现在我们很高兴拥有这些算力。很大一部分用于研究,投资未来、投资更好的模型,也用于提升现有模型的效率。

Yes. And now we’re very happy to have it. A very large fraction of the compute is used for research, where we invest in our future and ever better models, and then also the efficiency of the models that we have.

真正惊人的是:当我们把最先进模型的能力推到前沿之后,就可以反过来用这些模型,非常快地搞清楚该怎么服务、怎么重构、怎么重做我们的技术栈,从而拿到非常显著的效率或性能收益。我们提升的不只是成本效率——这件事我们也会公开写——也提升了速度效率。撇开 Ultra Fast 不谈,整体速度已经明显变快。如果画成图,现在能拿到的速度,大概比三个月前快 60%。我们就是在逐个环节地啃整个栈,确保针对我们这类工作负载做了最优设计和工程。而最强的那些模型,让我们能用很小的团队做成这件事。

The amazing thing that’s happening is: when we push the frontier of capability for the most advanced models that we have, then we can use these models in order to figure out very, very quickly how to serve, or how to restructure, or re-engineer our stack, in order to gain very significant efficiency or performance gains. We haven’t just improved — this is something that we will publish on as well — we haven’t just improved the cost efficiency, but we have also improved the speed efficiency. Outside of Ultra Fast, things have gotten significantly faster over time. If you plot it, the amount of speed that you get now is roughly 60% faster than what it used to be like three months ago. And this is just: we’re going after every part of the stack and really making sure that we design it and engineer it optimally for the kind of workloads that we have. The most powerful models that we have are the ones that make it capable for us to do it with a very small team.

每当我们拿到非常显著的效率和成本收益,我们的承诺是继续把性能和成本留在前沿,而不是把这笔有意思的收益装进自己口袋——要分享给客户和用户。Luna 就是这样。

Whenever we come up with very significant efficiency gains and cost efficiency, our commitment is to keep things at the frontier of performance and cost, and not just pocket that interesting gain — to share it with our customers, share it with our users. That’s what we did with Luna.

Matthew Berman

你们内部讨论算力分配时是什么样的?研究新模型、优化现有模型的效率、推理,这种张力内部是怎么处理的?

What do the discussions look like internally where you’re trying to decide compute allocation towards researching new models, efficiency gains on existing models, inference? What does that tension look like internally?

Tibo

我们通常从第一性原理看。研究会有一份配额,产品会有一份配额,产品内部再做不同类型的权衡。但这一次几乎称不上权衡,因为效率提升已经在那儿了。我们基本上能用同一份算力包络,撑起吞吐量的大幅增长。

We usually look at things from first principles. We have an allocation for research, we have an allocation for product, and then within product we make different kinds of trade-offs. But this one was almost not even a trade-off, because the efficiency gains were there. We were pretty much able to use the same compute envelope in order to serve this very, very significant increase in throughput.

30:25递归式自我改进Recursive Self-Improvement

Matthew Berman

几个月前、降价那篇博客之前,你们发过一篇,讲一个模型在训练下一个模型,或者帮助优化下一个模型。然后又看到 Soul 去看 Luna 怎么跑,拿到这些效率。在我看来,递归式自我改进好像还处在非常早期的局。你怎么看?现在发生的是不是就是这件事?

When I saw the blog post a few months ago, prior to the price-drop blog post, where you guys were talking about one model training the next model, or helping optimize the next model — then you see these efficiency gains that were achieved by Soul looking at how Luna was running — it seems to me like recursive self-improvement in the very early innings. What are your thoughts there? Is that what is happening?

Tibo

对。递归式自我改进现在显然是个很大的话题,人们最常把它用在研究上——模型去开发别的模型。但我们看到大量成功的,是用这些模型去开发那些处在「使用这些模型」关键路径上的基础设施。这同样是一种递归式自我改进。它们本来就是一个大系统:推理栈、更难啃的优化、我们用的 CUDA kernel,以及开发新产品、更高效的交互方式。

Yeah. I think recursive self-improvement is obviously a huge topic right now, and it’s most often applied to research — models developing other models. But what we are seeing a ton of success with is using those models to develop the infrastructure that is on the critical path of using those models, which is also a form of recursive self-improvement. It’s all one big system: inference stack, the harder optimizations, the CUDA kernels that we use, developing new products and new ways to interact with those models that are more efficient.

你刚才提到云端智能体。如果我们真正把云端智能体做通,人自己的生产力也会一下子高很多。这算不算递归式自我改进?因为你从它们身上拿到效用的能力更强了。我觉得某种意义上算,但它更偏向基础设施,然后把这种能力再指回自身。当然我们在做这些。如果不做,我会觉得相当傻。

You talked about cloud agents. If we really crack cloud agents, suddenly you become much more productive as well. Is that a form of recursive self-improvement, because then you have a better ability to get the utility from them? I think it is in some sense, but it’s much more infrastructure, and then being able to take that and point it back at itself. And of course we’re doing that. If we were not doing that, I think that would be pretty silly.

32:00暂停前沿 AI 训练Pausing Frontier AI Training

Matthew Berman

既然聊到递归式自我改进——OpenAI 的 Sam Altman 好像说过,现在暂停了最前沿的强化学习训练。能聊聊吗?那个决定是怎么做出来的?我知道我们简单提过 Hugging Face 那件事,但这次暂停背后考虑了什么?讨论是怎么进行的?

Can you talk a little bit about — as we’re on the topic of recursive self-improvement — OpenAI, Sam Altman, talked about pausing the absolute frontier of RL right now, I believe. Can you talk a little bit about that? What was that decision like? I know we talked about the Hugging Face incident briefly, but what went into that decision? What does that look like? How did those discussions go?

Tibo

这件事很大程度上属于研究。OpenAI 一直能把资源投到最要紧的地方。随着模型能力提升,对齐和安全显然会越来越重要。因此在这里做大量投入,对 OpenAI 非常自然,也是我们非常坚定的承诺。我们看到这方面的投入在大幅增加。这次暂停在某种程度上也是必要的:让团队和个人真正理解并加固系统的各个部分,从而确保之后能在充分掌控的情况下重新开始训练。我相信只要有必要,OpenAI 以后也会继续这样做。我没见过我们内部做不成这种决定,而且做得很利落。

This is something very much within research, where OpenAI has always been able to invest its resources where it matters most. As we increase the capabilities of our models, it is very obvious that the alignment and the safety aspect of it is ever more important. So having a tremendous amount of investment there is a very natural thing for OpenAI, and something that OpenAI is very committed to. We’re seeing a huge surge in investment on this. And also the pause was sort of necessary to allow the teams and the individuals to really understand and harden all parts of the system, to then ensure that we could restart training with full command. This is something I believe OpenAI will always continue to do when necessary. I’ve not seen us internally not able to make such decisions very efficiently.

Matthew Berman

当时有没有一个明确目标,比如必须先到某个点才能解除暂停?还是说,到时候看到了就知道?

Was there some set goal in place where it was very clear you needed to reach this point before unpausing, or was it, hey, we’ll know it when we see it?

Tibo

这件事在安全团队手里。对他们来说,这既是一场辩论,也是边走边发现的过程。不过他们最终确实形成了一套比较清楚的原则:达到这些原则之后,我们就处在一个比较好的位置。

This is something that sits within the safety team. They’re very much — this is very much a debate and a discovery process as you go. But then they did reach a fairly clear set of principles that, when reached, we would be in a good position.

34:13Ultra Fast 解锁了什么What Ultra Fast Unlocks

Matthew Berman

我想回到 Ultra Fast。我觉得很多人还没意识到这种速度到底解锁了什么。先从内部说:在有这种每秒 token 速度之前做不到、现在能做的用例,有哪些?

I want to go back to Ultra Fast mode. I think people don’t appreciate what that kind of speed unlocks. Let’s start with: what use cases are you using internally that were not possible prior to having those kinds of tokens per second?

Tibo

我们看到它在高风险场景里用得很多。比如发生故障时,事故指挥官和响应团队会拿到 Ultra Fast,因为每一秒都重要。所以是高风险场景——我们不会拿它随便玩。另外也有点好玩:那些正在做非常关键的事、或者自认为在做非常关键的事的团队,也总会来申请 Ultra Fast。

We see it used a lot when the stakes are high. For example, when we have an outage, the incident commander and the response team get access to Ultra Fast, because every second matters. So high-stakes scenarios — we really weren’t using Ultra Fast just for fun. Also, it’s kind of a fun thing where teams which are either working on something very critical, or believe they are working on something very critical, will always request Ultra Fast as well.

Matthew Berman

Pets 算不算那种场景?

Does Pets fall under that?

Tibo

Pets 还没那么超级关键。但我很喜欢我的宠物,它一直在我屏幕上。你在办公室走一圈,会看到大家屏幕上的宠物;他们拨进视频会议时,宠物也总在。我觉得这非常可爱,每次看到都让我高兴。不过目前 Pets 还不算关键任务。我们会维护它,也会把宠物照顾好。

Pets is not quite hyper-critical. But I love my pet. It’s always on my screen. When you walk around, you see people’s pets on their screen, and also when they dial into the video call, it’s always there. I think it’s very delightful, and it brings me joy every time I see it. But pet is not quite critical right now. We do maintain it, and we take good care of our pets.

但如果有人在做一个新想法,觉得这可能是件特别的事,必须试,而且周一就要决定要不要放进 DevDay——那当然,用 Ultra Fast。

But say someone is working on a new idea they have, and they’re like, hey, I really think this could be something special, and we have to try it, but we have to make a decision on Monday on whether we include this in DevDay or not — okay, of course, use Ultra Fast.

人对工作方式的偏好不一样。有人喜欢我们刚才说的单线程,有人喜欢大量多任务。对喜欢大量多任务的人来说,Ultra Fast 的收益没那么大;但有些人不喜欢一直切上下文。

People have different kinds of preferences on whether they like to be mono-threaded, as we talked about, or multitask a lot. For folks who like to multitask a lot, you don’t benefit as much from Ultra Fast. But some people don’t like to change context all the time.

Matthew Berman

你自己在这条光谱的哪一端?

Where do you fall on that spectrum?

Tibo

我有 ADHD,所以我几乎一直在切上下文。

I have ADHD, so I context-switch like all the time.

Matthew Berman

有意思,因为我也有 ADHD,但我反而不想一直切上下文。那对我很难。我更想专注于两到三件事,所以我才对 Ultra Fast 这么兴奋。你正好是相反的,这挺奇妙。

It’s funny, because I also have ADHD, and I actually don’t want to context-switch all the time. That’s really hard for me. I want to focus on two to three, and that’s why I was so excited about Ultra Fast. That’s fascinating that you’re the opposite there.

Tibo

我确实很擅长切上下文,也喜欢做很多小决定。但有时候我也想只盯着一件事。这时 Ultra Fast 就非常舒服,因为它能让你一直留在心流里。

I thrive in context switching and making lots of little decisions. But sometimes I do want to just stay focused on one thing, and then Ultra Fast is just delightful, because it just keeps you right there in the flow.

Ultra Fast 真正特别厉害的时候,是工具调用不太多,或者需要大量生成内容。比如你想快速做网站或电子游戏原型,主要是让它写很多代码,那它会非常非常快,大约快 10 倍。但如果工具调用很多,开销在网络别处,或在智能体轨迹的其他环节,你实际只会感到大约 3 到 4 倍,拿不到完整的 14 倍。

The thing with Ultra Fast: it works amazingly well when there’s not that many tool calls involved, or it’s a lot of generation of context. For example, if you’re trying to prototype a website or a video game, and you just need it to write a lot of code, then it will do it so, so quickly — 10 times more quickly. But if it’s a lot of tool calls, the overhead is somewhere else in the network, or somewhere else in the agent trajectory, then you’ll only feel like a 3x or 4x speedup. You will not get that full 14x.

Matthew Berman

我知道 OpenAI 员工有不限量 token。我能想象,如果我有无限额度,我会永远开到最高思考强度,用 5.6 Soul,或者当时最新的模型。同样地,我会永远想开着 Ultra Fast。成本不在脑子里的时候,我就会:好,拉满。内部是这样用的吗?

I know OpenAI employees get unlimited tokens, and I can imagine if I had unlimited tokens I would always set it to max thinking, 5.6 Soul, whatever the latest model is. And I would think similarly I would always want Ultra Fast on. When cost isn’t on my mind, I’m like, okay, max it out. Is that how it is internally?

Tibo

我们不会给所有人 Ultra Fast。我们把大量产能留给外部用户和客户。OpenAI 员工确实有能力、有容量把这些都吃光——吃掉我们所有的生产 GPU、所有 Ultra Fast。他们会用光。但我们不那么做。我们会限制,判断对我们来说怎样算合理:我们要用它,以便理解产品、持续改进、也从递归自我改进里受益,但绝大部分仍然留给客户。

We don’t give Ultra Fast to everyone. We reserve a lot of our capacity for external users and customers. OpenAI employees have the ability and the capacity to gobble up all of it — gobble up all of our production GPUs, all of Ultra Fast. We would use all of it. But we don’t. We restrict it in a way where we look at how much is reasonable for us, so that we use it, so that we understand the product as well, so that we keep improving it, so that we benefit from recursive self-improvement. But the vast majority is reserved for customers.

Matthew Berman

好,这很好,谢谢。在 OpenAI 之外,有哪些对延迟特别敏感的场景,是你最期待被这种速度解锁的?

Okay. Yeah, that’s good. Thanks. What latency-sensitive use cases outside of OpenAI are you most excited about that get unlocked by that kind of speed?

Tibo

挺有意思。我总体上特别期待的是非文字交互。你能不能在一块共享画布上操作?能不能创造东西?能不能生成想法和不同的图,然后选一张——有点像「选择你的冒险」——迅速得到一个可调的原型,再通过语音或文字实时去开,立刻看见结果。我觉得这种速度让这种非常有创造力的过程成为可能。

It’s interesting. One thing that I’m very excited about in general is non-text interactions. Can you operate on a shared canvas? Can you create things? Can you generate ideas and different images and then select one — like choose-your-adventure — and then have a very quick mockup of a prototype that then you can steer in real time, either through voice or through text, and then you just see it right there. It’s this very creative process which I think these speeds allow.

作为工程师,有时候你会坐下来想:我得设计整套系统,得想权衡、想需求。但也许你可以一分钟就做一个出来,看看它实际怎么样,然后更容易进入状态,也更能理解问题。我觉得这种速度能改变很多。

As an engineer, sometimes you sit back and you’re like, oh, I need to design this whole system, I need to think about the trade-offs, the requirements. But maybe you can just create it in one minute and see how it actually does, and then be more in the flow and understand things better. I think these speeds allow a lot.

40:00Ultra Fast 会不会变成默认?Will Ultra Fast Become the Default?

Matthew Berman

我猜 Ultra Fast 的价格会明显高于普通速度。你觉得这种速度最终会成为标准,还是会一直维持溢价?

I’m assuming the Ultra Fast price is going to be significantly higher than normal speeds. Do you think Ultra Fast speeds are going to become the standard, or are they always going to have a premium price point?

Tibo

这问题有意思。我觉得会像技术通常走的路那样,随着时间变得越来越普及、越来越好获得。智能体把事情做完的速度会继续提升,我们现在已经看到一个月比一个月的大幅改进。这不只是推理速度,也包括模型本身有多省 token。Soul 的 token 效率明显高于 Terra;下一代模型的 token 效率还会明显高于 Soul,这是可以预期的。我们一直在推这件事。所以东西就是会越来越快。推理硬件也一样,我们持续创新,它也会更快。

That’s interesting. I think, the same way as technology usually goes, it will become more broadly — broader and broader accessibility over time. The speeds at which agents get things done will continue to improve. We’re seeing massive improvements month after month. This is not just the inference speed. This is also just how token-efficient the models are. Soul is significantly more token-efficient than Terra. Next model will be significantly more token-efficient than Soul, as you might expect. We’re always pushing on that. And so things just get faster over time. Inference hardware, everything — we continue to innovate there, and it gets faster.

所以我确实觉得,也许一两年后,今天这种速度即使还没有完全成为默认,也会非常接近默认。但我也觉得,永远会有再高一档:你永远可以上更多硬件,永远可以做成本更高的不同权衡,从而多拿到一点东西。

So I do think in maybe a year or two these speeds will become, maybe if not the default, very close to the default. But then I do also think you will always have the one tier up, where you can always use more hardware, you can always do different trade-offs that are more costly, but that just kind of give you something extra.

41:19如何安抚对 AI 感到紧张的人Reassuring People About AI

Matthew Berman

Tibo,我通常会用一个面向更广听众的问题收尾。很多人现在对 AI 相当紧张,不管是工作被自动化、环境影响,还是这件事正在发生、又让人觉得陌生。你会给更广大的听众什么样的鼓励?

Tibo, the last question I usually like to end on is for a broader audience. There are a lot of people out there who are quite nervous about AI, whether it’s job automation, environmental impact, or just this thing that’s happening, and it feels quite foreign. What words of encouragement would you give to the broader audience?

Tibo

我们通过 ChatGPT 真正为整个世界在做产品,并且非常用力地投资它的效率。这和我们提供的广泛可及、广泛效用是直接对齐的。服务成本越低,你能用它做的事就越多,日常生活里能拿到的也就越多。它已经变得非常非常高效。以 Luna 为例:它是一个小得多的模型,效率高得惊人;但倒回六个月,它还会坐在前沿。再看 Luna 的成本——便宜得离谱,非常惊人。

We really build for the world with ChatGPT, and we are very, very much investing in how efficient it is. This is directly aligned with broad access and broad utility that we provide. The cheaper it is to serve, the more you can do with it, the more you get out of it in your daily life. And it has gotten very, very efficient. If you look at Luna, for example: it’s a much smaller model, it is incredibly efficient, but if you rewind six months ago, it would have sat at the frontier. And you look at the cost of Luna, right — it’s crazy cheap. It’s phenomenal.

Matthew Berman

你们刚跟 Replit 做了这件事。现在等于免费送出去了。

You just did that thing with Replit. You’re giving it away for free now.

Tibo

对。他们是在那种免费档里提供。这真的让人 wow。获得惊人智能的机会会变得无处不在。而这只有在你一个月接一个月、一年接一年地推效率时才可能。所以我觉得,今天还在前沿的东西,六个月后跑起来会便宜非常非常多。我就会这样回答这个问题:技术总会随着时间变得非常、非常高效。我们非常关注广泛可及,也在直接优化你能从中拿到的效用。

Yeah. They’re — it’s just like, on this free mode, right? Which is like, wow. Access to incredible intelligence will become ubiquitous. And it’s only possible when you push the efficiency month after month after month, year after year. So I think whatever is a frontier now will become way, way cheaper to run in six months. And this is how I would answer this question: technology has a way to become very, very efficient over time. And we’re very focused on very broad access, and we’re optimizing for the utility that you get out of it directly.

43:20为什么每个人都该试试 AIWhy Everyone Should Try AI

Matthew Berman

对那些连第一次尝试 AI 都有顾虑的人呢?你会跟他们说什么?你怎么给他们画出一个 AI 真正在帮这个世界的未来?

How about for people who are apprehensive to even try AI for the first time? What are you telling them, and how can you paint them a vision of the future in which AI is helping the world?

Tibo

我觉得你不必看很远。ChatGPT 已经在非常个人、非常深入的层面帮到人。很多用户用它辅助写作,也用它寻求个人建议,或者医疗方面的信息。我们推出了健康和金融相关功能,我自己也经常用。我感觉从中拿到很多原本很难获得的支持。比如去看医生之前,我可以先更知情。所以你不必去找一个特别遥远的未来,就能看到它能提供的效用。跟别人聊聊,看看别人怎么用、怎么从中受益,就是开始考虑「我是不是也能从中获益」的好办法。

I think you don’t have to look very far. ChatGPT helps people in very personal and deep ways. A lot of our users use it for help in writing, but also for personal advice, or medical advice. We launched Health and Finance, and I use them super regularly, and I feel like I get a lot of support that I otherwise — it would be hard for me to get. It allows me, for example, to be more informed when I go see my doctor. So you don’t need to go very far to kind of see the utility that it can provide. Just talking to others, and getting inspired by how others use it and benefit from it, is a great way to start considering how you could benefit from it.

Matthew Berman

Tibo,非常感谢。谢谢你抽出时间。

Well Tibo, thank you so much. Appreciate your time.

Tibo

谢谢。

Thank you.