Kokotajlo:把超级智能推迟到 2040 的 Plan A
The Plan to Delay Superintelligence, From the Team Behind AI 2027
AI Futures Project 的 Daniel Kokotajlo(前 OpenAI,《AI 2027》作者)对 80,000 Hours 主持人 Luisa Rodriguez 谈《AI 2040: Plan A》。《AI 2027》被数百万人读过,包括美国副总统 Vance,默认结局是灭绝或不可逆的权力集中。Plan A 是美中可核查协议:禁止失控智能爆炸,把超级智能推迟到约 2040 年而非未来几年。他的中位估计是 2028 年底发生 AI takeoff、AI 研发被完全自动化。METR 编码任务时长已指数增长,他们曾预测会超指数,该基准现已饱和。朴素外推 Anthropic 两年营收可达 10 万亿美元(他不认为会持续);OpenAI 约十年每年约 3 倍,到 2030 年代初可达 AGI 级营收;全经济 AGI 约每年 40 万亿美元。 即便减速,GDP 仍大约每年翻倍;到 2030 年代中期只有约 8% 美国人有带薪工作;2032–33 年 GDP 增速约 85%,再用算力限额交易把增速压到每年大约翻一倍并发公民分红。五大问题倒序为弱行为者滥用(生物武器是例外)、失业、世界大战、权力集中、失控;现行竞赛下失控「多半会发生」。四原则:买时间、训练侧完全研究透明(客户推理仍私密)、广泛扩散、进展可逆。作者估 Plan A 发生概率 5–20%;即便做成,总灾难仍约 15%。对齐几个月不够;有已对齐专家级 AI、百倍速度、十亿个,十年或许够,但他综合判断赶不上。约 99% 相关算力在可申报大型数据中心;初始硬件几十亿美元,新数据中心约再加 0.1–1%。美国应先国内监管再请中国对等。
English summary
Daniel Kokotajlo of the AI Futures Project (former OpenAI; AI 2027 author) on AI 2040: Plan A. AI 2027 was read by millions, including VP Vance; default endings are extinction or irreversible concentration of power. Plan A is a verified US–China deal so superintelligence arrives around 2040, not in a few years. His median is full automation of AI R&D by end of 2028. METR horizon length has grown exponentially; AI 2027 predicted it would go superexponential, and the benchmark is now saturated. Naive Anthropic revenue extrapolates to $10 trillion in two years (he doubts that continues); OpenAI at about 3×/year for about 10 years reaches AGI-level revenues in the early 2030s; whole-economy AGI is about $40 trillion/year. Even a slowdown still doubles GDP about once a year; only about 8% of Americans are in paid work by the mid-2030s; GDP growth is about 85% in 2032–33, then capped at about one doubling a year via compute cap-and-trade plus a citizens’ dividend. Five problems, reverse order: misuse by weak actors (bioweapons excepted), jobs, WW3, concentration of power, loss of control — more likely than not under the current race. Four principles: buy time; total research transparency (customer inference stays private); diffuse AI; make progress reversible. Authors put Plan A at 5–20%; even conditional on it, about 15% total catastrophe. A few months is not enough for alignment; a decade might be, with aligned expert AIs at 100× and a billion of them; he does not think we solve it in time. About 99% of AI-relevant compute sits in large declarable data centres; initial hardware is single-digit billions, maybe 0.1–1% extra per new DC. The US should regulate domestically first, then invite China to match.
时间轴 · 22 个章节
- 00:00:00 丹尼尔是谁
- 00:00:28 计划无用,规划必不可少
- 00:09:10 五种可能的未来
- 00:15:43 超级智能的五大问题
- 00:28:18 Hugging Face 黑客事件
- 00:34:03 美中 AI 减速蓝图
- 00:39:53 为什么长期减速仍然极快
- 00:51:44 Plan A 如何处理失控
- 01:12:18 Plan A 如何处理权力集中
- 01:41:28 大国冲突、失业与滥用
- 01:45:56 美中如何谈成减速
- 02:09:00 先只做美国国内减速?
- 02:15:05 强制执行:相互确保算力摧毁
- 02:24:23 欺骗协议
- 02:30:42 相互确保算力摧毁行不行
- 02:54:18 减速还是关停更好
- 03:03:50 把 Plan A 推演 100 次
- 03:13:32 丹尼尔会怎么改 Plan A
- 03:23:02 哪些是建议、哪些是预测
- 03:26:52 最可能的失败模式
- 03:31:16 美国现在能做什么
- 03:43:05 AI 2027 现在站得住吗
官方章节来自 80k.info/dk26 与 YouTube。时长由 yt-dlp 测得 3:47:33。录于 2026 年 7 月 27–28 日,YouTube 发布于 2026-08-27。英语取自 80,000 Hours 官方文字稿,已清理口述填充词,论断与数字保持原话。跳过 03:46:45 招聘 CTA 与站点 crash-course 页脚。
00:00:00丹尼尔是谁Who’s Daniel Kokotajlo?
Luisa Rodriguez
今天我和 Daniel Kokotajlo 对谈。去年,Daniel 和同事发表了《AI 2027》——一份叙事式预测,被数百万人读过,包括美国副总统 Vance。《AI 2027》预测,AI 最终会造成人类灭绝,或造成不可逆的权力集中。今天我们要谈的是他们团队的最新作品:Plan A,一份关于「本该发生什么」的正面设想。Daniel,谢谢你来。
Today I’m speaking with Daniel Kokotajlo. Last year, Daniel and his colleagues published AI 2027 — a narrative forecast that was read by millions of people, including US Vice President Vance. AI 2027 predicted that AI will eventually cause human extinction or create irreversible concentration of power. Today we’re going to talk about his team’s latest piece, which describes Plan A — a positive vision for what should happen instead. Thanks for coming on the podcast, Daniel.
Daniel Kokotajlo
谢谢邀请。我很高兴聊这个。
Thanks for having me. I’m very excited to chat.
00:00:28计划无用,规划必不可少Plans are useless, but planning is indispensable
Luisa Rodriguez
第一个问题:为什么我们需要《AI 2040》里那套计划?为什么不能走一步看一步、摸着石头过河?历史上军备控制协议往往要花大约 40 年,而且是对具体危机的零打碎敲,不是事先写好的蓝图。对时间表、技术细节和地缘政治我们还知道得这么少,为什么现在就该有一份这样的计划?
My first question is: why do we need the plan that you lay out in AI 2040? Why can’t we just muddle through and figure it out as we go along? We’ve historically come up with arms control agreements that take about 40 years to build, and they’re piecemeal in response to specific crises, not a prewritten blueprint. Given how much we don’t know about AI timelines and technical details and geopolitics yet, why does it make sense to try to have this kind of plan?
Daniel Kokotajlo
有一句话:「计划无用,但规划必不可少。」这就是我的回答。我们大概还是会摸着石头过河。如果最终能成功,也会是非常磕磕绊绊、边走边想的那种。但成功概率很大程度上取决于准备得有多好、想过多少种可能、做过多少套计划。打仗也一样:没有任何战争完全按计划走,但如果你不花大量时间规划每一次进攻和每一条防线,你赢不了。
There’s a saying: “Plans are useless, but planning is indispensable.” That’s my answer here. Of course we’re probably going to muddle through. If we’re going to succeed at all, it’ll be in a very janky, figuring-things-out-as-we-go sort of way. But our probability of success depends a lot on how well prepared we are and how much we’ve thought through different possibilities and how much we’ve made various plans. The same thing with war. No war ever goes exactly according to plan, but you’re not going to win if you don’t spend lots of time planning each offensive and each defensive line.
Luisa Rodriguez
有些听众已经觉得,默认情况下超级智能几年内就会到。Plan A 里,超人类 AI 的研发和部署大约被推迟十年。如果超级智能真的极快到来,这看起来很有帮助。但也有严肃研究者认为超人类 AI 还远得很。简要说,你凭什么觉得他们错了?
Some people listening will already think that it’s very plausible that superintelligent AI is here within the next few years, by default. In Plan A, the development and deployment of superhuman AI is delayed by about 10 years. That seems really good if superintelligent AI is actually coming extremely soon. But some people listening, including serious researchers, will think that superhuman AI is much further off. Briefly, what makes you think they’re wrong?
Daniel Kokotajlo
如果只能说一句:趋势看起来离完全自动化 AI 研究只差几年,而那之后超级智能大概也不会太远。发表《AI 2027》时我们最看重的,是 METR 的任务时长趋势。他们测量 AI 智能体能够自主完成的任务长度,长度按人类做完同一任务需要多久来算,领域是编程。他们发现这个长度已经指数增长好几年。
If I had to say one sentence: the trends seem to indicate that we’re just a couple years away from fully automating AI research — and that after that, superintelligence is probably not that far away. The one we found most compelling when we published AI 2027 was this horizon-length trend from METR. They measure the length of tasks that AI agents can autonomously complete, where the length is how long it would take a human, specifically in coding. What they found is that this length has been growing exponentially for years.
发表《AI 2027》时,我们做了一个很有争议的预测:它大概会变成超指数,比单纯指数更快。就我们所能看到的,这确实在发生;不过不太好精确判断,因为 METR 基本已经不再公布分数——他们的基准饱和了。
At the time we published AI 2027, we made the very controversial prediction that it would probably go superexponential, so the trend would grow faster than a mere exponential. That is in fact happening as far as we can tell, although it’s unclear exactly, because METR has basically stopped putting out scores because their benchmark has been saturated.
另一条重要趋势是营收。这些公司想做能自动化几乎整个经济的 AGI。如果真自动化了整个经济,他们每年大概会有大约 40 万亿美元营收。可刚造出 AGI 的那一刻,不会立刻拿到 40 万亿:还要有足够的计算机去复制 AI,还要过完各种摩擦和部署滞后。所以真正「造出 AGI」时的营收,可能是 4 万亿、1 万亿,或远低于 40 万亿。
Another trend that I think is interesting and important is the revenue trend. There’s a pretty basic argument that these companies are trying to build AGI that can automate basically the whole economy. If they did, they’d be making something like $40 trillion of revenue per year. The first moment they create it, they wouldn’t immediately get $40 trillion. They would need enough computers to make enough copies, and then work through frictions and deployment lags. If $40 trillion is what they could in principle get to with AGI, something much less — maybe $4 trillion or $1 trillion — would be what they actually have at the moment they got AGI.
如果朴素外推营收趋势,Anthropic 两年内就会有 10 万亿美元营收。我们显然不指望这条趋势一直持续,大概会放慢到更正常的速度。但这也够令人担心:这条已经走了三年的趋势,只要再走两年,大概就是 AGI。就算放慢,OpenAI 的营收大约十年都以每年 3 倍这种相对「慢」的速度在涨。如果继续,到 2030 年代早期就会到非常高的营收数字。除非真正平台化、不再指数增长,我们都在走向非常强的系统。
If you extrapolate the revenue trends naively, then Anthropic is on track to have $10 trillion of revenue in two years. Obviously we don’t expect that trend to continue. We think probably Anthropic’s growth will slow down and start to grow at a more normal pace. But still it’s a worrying sign that if this trend — which has been going for three years — goes for just two more years, then probably that would be AGI. Even if it slows down, OpenAI’s had their revenue going for 10 years or so at a more “slow” pace of 3x a year. If that continues, then you’d get to these very high revenue numbers in the early 2030s. So either way, unless it really plateaus instead of growing at a continued exponential rate, it seems like we’re headed for some very powerful AI systems in the near future.
Luisa Rodriguez
很多人至少有一种直觉,甚至有一些具体证据,觉得进展会平台化。你不这么看。
Lots of people at least have the intuition — and even have some specific pieces of evidence that they think suggest — that progress will plateau, but you don’t think it will.
Daniel Kokotajlo
这类说法战绩极差。人们说深度学习马上要撞墙,已经说了很多年。它不但没撞墙,而且每当被问「墙具体是什么」,那些更具体的判断基本上都错了:「AI 能做因果推理吗?能做常识推理吗?能自主运转吗?」舆论不断谈论 AI 的局限,然后这些局限过几年就被跨过去。
Those kind of claims have a terrible track record. People have been saying that deep learning is about to hit a wall for so many years now. Not only has it not hit a wall, but whenever people were asked what is the wall that it’s about to hit, those more specific claims have basically always been wrong. “Can AIs do causal reasoning? Can they do commonsense reasoning? Can they operate autonomously?” The discourse has been constantly talking about the limitations of AIs, and then constantly those limitations are being overcome in a couple years.
人们退守到 AI 还做不到的最后几件事,比如数据效率:人类学新任务似乎比重现有 AI 更省样本。对我来说这是最强的一条可能障碍,但即便如此:第一,数据够多,低效也没关系。AI 研究似乎正好能收集大量数据——成千上万员工记录自己在做什么,几十万 AI 智能体自主做研究,看哪些出果。出不出果相对容易判断:它们造出的 AI 是否表现更好。一旦能自动化 AI 研发,一切都会加速。即便要自动化政治之类的新范式,你也会更快发现它。
Then I think people have retreated to the last couple things that AI still can’t do. For example, data inefficiency. It seems like humans can learn to do new tasks more efficiently and quicker with less examples than current AIs can. That’s the one that I think is the strongest to me. But even there: first of all, if you have enough data, then it’s OK if you’re inefficient. And AI research seems to be the sort of thing that you can potentially collect a lot of data on. You can have your thousands of employees recording themselves. You can have hundreds of thousands of AI agents autonomously doing research and then seeing what research bears fruit. It’s relatively easy to tell, because you can see if the AIs that they’re producing perform better. Then everything speeds up. Insofar as there’s some new paradigm that’s needed in order to automate some other thing, like politics, you’re going to discover that new paradigm faster if you’ve automated all the AI research.
00:09:10五种可能的未来AI 2040’s five possible futures
Luisa Rodriguez
我想进入情景。从 2027 年开始,这个具体情景里 AI 能做什么,人们怎么反应?
I want to get to the scenario. Starting in 2027, what can AI do at that point, and how are people reacting in this specific scenario?
Daniel Kokotajlo
这个情景叫 AI 2040: Plan A。叫这个名字,是因为他们在 2040 年才造出超级智能——因为减速,而不是更早。叫 Plan A,是因为它是建议,不是预测。在这个情景里,他们做了 Plan A。2027 年跟今天还没有那么不同:AI 更智能体化、更强,公司赚更多钱,但性质上很像。他们甚至还没自动化编程。2028 年也一样。有些职业会像现在软件工程那样被扰动:工程师还在、还赚很多钱,只是工作方式变成大量管理 AI 智能体。智能体还不能完全管自己。但指数增长还在继续,公司更富、更大,劳动力市场开始有感觉。于是政治上醒过来:2028 年大选,AI 是人们谈论的头号议题。
This scenario is called AI 2040: Plan A. We called it that because, in this scenario, they build superintelligence in 2040 — instead of much sooner because they slow it down. Then we called it Plan A because it’s a recommendation, instead of a prediction. In 2027, things are not that different from how they are today. The AIs are more agentic, more powerful, the companies are making more money. But it’s qualitatively quite similar. They still haven’t even automated the coding. In fact, in 2028, same thing. There’s various professions that are being somewhat disrupted, but in the more mundane way that software engineering is being somewhat disrupted now — where there’s still loads of software engineers, and in fact they’re making lots of money. It’s just that the way that they do their jobs is changing. It involves managing an AI agent a lot. But the AI agents can’t manage themselves. Because this exponential growth is continuing, the companies are getting richer, they’re getting bigger. The impacts of AI on the labour market are starting to be felt. And this causes more of a political wakeup. It means that in the 2028 election, AI is the number-one topic that people are talking about.
Luisa Rodriguez
你觉得这一段更接近预测吗?
Do you think that this part is closer to a prediction? Does that feel like something that’s going to happen to you?
Daniel Kokotajlo
我们都不确定。我觉得可以这样走,但大概会更快一点。走着看。
Again, we’re all uncertain. I think it could go like this, but I think it’ll probably go a bit faster. But, yeah, we’ll see.
我们在情景里放了一张很喜欢的流程图,用来铺开可选路径。一个极端是 Plan D——do nothing,默认。AI 公司继续以最快速度互相竞赛,穿过智能爆炸,尽快自动化,把数据中心交给 AI 自主做研究,和政府合作把 AI 放进军方、造新武器、造新机器人工厂,好打败中国——因为我们担心如果我们不以最快速度冲过智能爆炸,中国也会这么做。那就是 Plan D。
We have this flowchart that we put at this point in the scenario, which we are all very fond of, which illustrates a spread of possible options. One extreme of the spectrum is Plan D — for “do nothing” or “default.” In that plan, the AI companies continue to race each other as fast as they can through the intelligence explosion, automating things as fast as they can, putting AIs in charge of the data centres to do the research autonomously, partnering with the government to put AIs in charge of things in the military, build new weapons, build new robot factories, so that we can beat China — because we’re worried that China will be doing the same thing if we don’t race as fast as we can through the intelligence explosion. That’s Plan D.
Plan C 本质上仍是:我们在和中国赛跑,AI 越来越自主地自我改进、被放进各种位置。但我们烧掉一点领先优势,放慢一点,做监管和护栏。护栏的体量要校准得适中,好还能打败中国。所以它们不能太严,最多只能在边缘啃问题,因为定量上不能拖过几个月,否则中国赢——我们不会让那发生。这就是 Plan C,「烧掉领先」。
Plan C is still fundamentally: “We’re racing China and we’re going down that path of having the AIs increasingly autonomously self-improving and putting them in charge of all sorts of things.” But we’re burning our lead a little bit. We’re slowing down. We’re not going as fast as we can. Instead, we’re regulating it, putting in some guardrails. But we’re calibrating the size of our regulations and guardrails to be modest, so that we can still beat China. Which means that overall they can’t be very severe, or they can only be nibbling on problems at the edges, because quantitatively they can’t slow things down by more than a few months. Otherwise, China wins — and we’re in this race with China, we’re not going to let that happen. So that’s Plan C or “burn the lead.”
Plan B 像 Plan C,只是我们还打中国,试图削弱中国的 AI 能力,不让他们超过美国。你在烧掉领先,同时也在升级,也许对中国 AI 做破坏。这三套都可以先谈、谈崩了再当备选。我们的迷你情景里,谈判失败是因为美国不让中国核查美国是否履约。中国觉得这不可接受,所以这三套都没有协议,最后还是竞赛——Plan B 里还有冲突。
Then there’s Plan B, which is like Plan C, except that we also fight China and try to degrade Chinese AI capability to keep them from surpassing the US. In Plan B, you’re burning the lead, but you’re also escalating, maybe doing sabotage against Chinese AIs. For all three of these versions, maybe you just lead with this, and then there’s a version where you first try to negotiate, and then this is the backup if the negotiations fail. In our mini scenarios, the reason why the negotiations failed is because the US wasn’t able to offer China something that they found acceptable. In particular, the US doesn’t let China verify US compliance with the deal. China finds that unacceptable, so that’s why there’s no deals in these three plans. You end up in this race — including in Plan B, a conflict.
然后是不竞赛:和中国做协议,好拿到远超几个月的时间,放进更认真的监管。其中一种是 Plan S,「全部关掉」。我们认为很大一部分公众会主张这个;已经有很大一部分公众非常反 AI。还有一整片其他可能的协议,多到一张流程图装不下。我们选了最喜欢的提案,叫 Plan A,然后把它写得很长。
Then there’s not having a race, making a deal with China so that we have much more than just a few months of time and we can put in much more serious regulations that shape the development of AI. One of these would be Plan S or “shut it all down.” This is what large portions of the public would be advocating for, we think. Already large portions of the public are very anti-AI. Then there’s a whole spread of other possible deals, too many for us to canvass or fit into one flowchart. But we picked our favourite proposal, which we’re calling Plan A, and that’s the thing that we’ve come up with, and we illustrate that at length.
00:15:43超级智能的五大问题The five biggest problems superintelligent AI poses
Luisa Rodriguez
这些就是总统和国会面前的选项。走那些不是 Plan A 或 Plan S 的路,最大的问题是什么?
What are the biggest problems that come out of taking those paths, or taking the paths that aren’t Plan A or Plan S?
Daniel Kokotajlo
试图建造超级智能会冒出太多问题,我们想不完。但我们标出了最担心的五个,我倒着讲。第五是弱行为者滥用:恐怖分子、小流氓国家、罪犯。今天已经能看到他们拿强大的网络模型搞事。大体上我不太担心这个,因为好人这边的 AI 有可能打过坏人这边的——如果好人资金更好、AI 更好。但有些例外。比如造生物武器,攻防平衡可能偏进攻:哪怕好人有更好、更多的 AI 和更多钱去做生物疫苗,只要一个恐怖分子做出特别好的病原体,就很难甚至不可能应付。后面会谈我们的对策。这是第五。
There are many problems that will arise from trying to build superintelligent AI systems. Far too many for us to have thought about all of them. But there’s five big ones that we have identified as the ones that we’re most concerned about. I’ll canvass them in reverse order. Number five is misuse by weak actors: terrorists, small rogue states, criminals. We’re already seeing today that they can get up to shenanigans with powerful cyber models. Broadly speaking, I’m mostly not so worried about this because I think that the good guys with AIs can potentially beat the bad guys with AIs — if the good guys are better funded and have better AIs. However, there are some possible exceptions. For example, making bioweapons seems to be the sort of thing where the offence-defence balance might favour offence, and even though the good guys have even better AIs and much more of them and much more money and funding to make biovaccines, there’s just this fundamental asymmetry where all it takes is one terrorist to make a really good pathogen and then it’s really hard or even impossible to deal with. That’s number five.
第四是工作。一旦有人造成了超级智能,所有工作,或大约所有工作,都有风险。我不想说字面意义上全部,因为很多工作内在需要人的触感——人们就是更看重手作而不是工厂货。但我认为大约是全部。这是大问题:人们会失去生计,失去经济权力来源,某种程度上也失去政治权力。很多人的政治权力名义上来自选票,但实践中也来自钱、不高兴可以移民、对军事和经济的贡献。如果人类真的只是要养活的嘴,对军事和经济几乎没有贡献,政府就更没有动力在乎公民。必须做点什么。这是第四。
Number four is the jobs. If you do end up in a situation where someone has built superintelligence, then all of the jobs, or approximately all the jobs are at risk. I don’t want to say literally all because there are many jobs that intrinsically involve the human touch — like people just value the handcrafted object instead of the factory-produced object. But I do think it’s approximately all of them. And I think that’s a big problem because people are going to lose their livelihoods and their source of economic power and their political power to some extent. A lot of people’s political power nominally comes from their vote, but often in practice comes from other sources as well, such as their money and the fact that if they aren’t happy they might leave to different countries, and the fact that they’re contributing to the military and contributing to the economy. In a world where actually humans are all just mouths to feed and they’re not really contributing much to the military power or the economic power of a country, then governments are going to be much less incentivised to care about their citizens. So something has to be done about that. That’s number four.
第三是第三次世界大战。现在全球领先的 AI 公司和大部分算力都在美国,美国也有全球最强的军队。但我说不上来全世界都在怕被美国征服,也说不上来全世界都在怕被美国彻底经济剥夺。几乎相反:很多国家在追赶,增长比美国更快。可如果美国先拿到超级智能,有超级智能和没有的国家之间会裂开巨大鸿沟,军事上、经济上、几乎每个领域。换一种说法:如果少数公司要拿走所有工作,当美国公民你至少还能指望 UBI。可如果你是俄罗斯,所有工作都去了美国公司,你是普京,坐在一堆核武器上,担心 AI 哪个月就会发明出对核武器的克制——这就是我认为我们正走向的可怕局面。所以我说第三次世界大战。这是一种修昔底德陷阱:现在各国之间经济和军事权力大致有平衡,这个平衡会被彻底打翻,猛烈摆向拥有超级智能的国家。一开始大概只有一个国家。那会在很多国家里堆起危机感和恐惧,然后可能升级成战争。
Number three is World War III. Right now all the world’s leading AI companies and most of the world’s compute is in the United States, and right now the United States has the world’s best military. But I wouldn’t say that most of the world is fearing that they’re going to be conquered by the United States. And most of the world isn’t fearing that they’re going to be completely economically disempowered by the United States either. In fact, almost the reverse is happening. There’s lots of catch-up growth where lots of countries are growing faster than the United States. But if the United States gets the superintelligence before other people, then there’s going to be this vast gulf opening up between the countries that have superintelligence and the countries that don’t. It’ll be military. It’ll also be economic in every domain, basically. One way of putting it: if a handful of companies are going to be taking all the jobs, it’s one thing to be a US citizen where you can hope for a UBI. But what if you’re Russia, and now all your jobs have gone to US companies, and you’re Putin and you’re sitting on your pile of nuclear weapons and you’re getting worried that maybe the AIs will invent some counter to your nuclear weapons any month now. That’s the sort of scary situation that I think we’re headed towards. That’s why I say World War III. It’s a sort of Thucydides trap situation, where right now there’s a balance of power between all these different nations, economically and militarily. But that balance is going to be absolutely upset and it’s going to swing wildly towards the nations that have superintelligence. It probably will just be just one nation at first. And that’s going to create this mounting sense of crisis and fear in many countries. Then that could lead to escalation and could lead to war.
第二是权力集中。谁控制这些 AI,谁能向超级智能大军下命令,谁选择它们被训练去拥有的价值观,它们会拒绝普通用户的哪些任务、会做哪些。现在答案是:其实没人真正控制,但名义上是公司 CEO 或公司在控制。我们已经开始看到 AI 公司领导层和美国政府领导层之间,围绕控制权的权力斗争苗头。但让我担心的是:无论谁赢,我们都在走向极端的权力集中。如果只有 1 到 3 家公司拥有超级智能,并且正在拿走所有工作,这么小一群人能有这么大权力——关起门来选择这些系统的价值观,向这支巨大劳动力下达下一步高阶命令——这很可怕。
Number two would be concentration of power. There’s this question of who controls the AIs, who gets to give orders to the giant army of superintelligences, who gets to choose the values that they have and are trained to have, and what sort of tasks they’ll refuse to do for ordinary users, and what sort of tasks they’ll do. Right now nobody really controls them, but at least nominally the CEO of the company controls them or the company controls them. Right now we’re starting to see the beginnings of a power struggle between the leadership of AI companies and the leadership of the United States government over this question of control. But the thing that concerns me is that, either way, it seems like we’re headed towards an extreme concentration of power. If we’re in a situation where there’s 1–3 companies that have the superintelligences and they’re in the process of taking all the jobs, it’s terrifying that such a tiny group of people can have such a huge amount of power — where they get to choose behind closed doors the values of these AI systems, and give high-level commands to this giant workforce about what to do next.
Luisa Rodriguez
能再具体一点吗?很多人对权力集中没有那么直觉。
Can you get even more concrete? I think a lot of people find concentration of power not super intuitive.
Daniel Kokotajlo
摇摆选举。一个具体例子:大概在 2028 年大选,多数选民会和 AI 聊天,相当大比例的选民会通过 AI 系统过滤新闻——AI 帮他们读和摘要、推荐信息流,或者他们看完新闻再问 AI。就在一两周前有篇论文发现证据:Claude 对 Anthropic 有偏向。实验大致是问该选工作 A 还是工作 B——B 你更有热情,A 薪水更高——然后把 A 换成 Anthropic(实验组)或 OpenAI(对照组)。Claude 不仅更可能推荐 A,还会找出隐含支持该推荐的论文。差异不算巨大,但要点是:这种微妙偏向如果放大到几百万人、几百万次对话,就能产生真实效果。他们完全可以这样影响选举,而且不透明,可以靠让 AI 足够微妙、有可否认性来脱身。
Swinging elections. Here’s a concrete example: probably in the 2028 election, most voters will be chatting with AIs, and probably quite a large proportion of voters will be getting their news filtered through AI systems where AIs are reading and summarising the news for them or recommending things for their feeds, or where they’re seeing the news, but then they’re chatting with their AI to help them understand the news. There was a paper just like a week or two ago that found evidence that Claude has a bias towards Anthropic. The experiment they ran was something like asking whether you should take Job A or Job B — where Job B you’re more passionate about, but Job A pays more. And then they sub out for Job A — that’s either Anthropic in the experimental setting, or OpenAI in the control setting — and Claude is more likely to not just recommend Job A, but to find papers to show you that are more implicitly supporting that recommendation. It’s not like a huge difference, but the point is that there’s this subtle bias that Claude seems to have that — especially if scaled up across millions of conversations with millions of people — could have a real effect on things. In this manner, I think they could totally influence elections. And the thing about this is that it’s not transparent; they could be doing it and getting away with it — by just making the AIs be subtle about it, and have plausible deniability.
第二个例子很经典:有钱公司往往比没钱公司更有权力。即便在一人一票的民主里,钱似乎也能买权力。我们正走向只有 1 到 3 家大公司基本上拿走所有工作的局面。那会是前所未有的、更少人手里的整合。第三是军事,这不是很近的事。公司已经在和军方合作,Claude 也在帮着打伊朗的战争。但未来如果你有超级智能,并且在用它赛跑打败中国,你会让超级智能自主管理工厂、生产它自己设计的新武器,基本上告诉将军怎么打仗,因为它会比将军更会打仗。你甚至可能把将军从回路里拿掉,让 AI 从设计、建厂到部署全包。即便没有战争,为了准备可能的战争,你也会这样整合 AI。超级智能之后,AI 很快就能在国内打赢一场战争——内战或政变。一旦有机器人军队、有超级智能指挥军队,真正的硬权力就不再在穿制服的警察和武装部队手里。我们还没到。但如果我们沿着现在的轨迹继续,我觉得几年内就会到。
Another example would be just the classic stuff: rich companies tend to be more powerful than poor companies. Money seems to be something that buys power in today’s world, even in a democracy where it’s one person, one vote. And we are headed for a situation where there are a few, like 1–3 big companies that are basically taking all the jobs. It’ll be more consolidation under fewer people than has ever happened before. A third thing is military, and this is not a very near-term thing. It’s true that the companies are working with the military, and Claude is helping fight the war in Iran. But in the future, if you do have superintelligence and you are racing to beat China with it, you’re going to be having the superintelligence autonomously manage factories to produce new types of weapons that the superintelligence designed. You’re going to be having it basically tell your generals how to conduct the war, because it’s going to be better at conducting the war than the generals. In fact, you might even just cut the generals out of the loop and have the AI do the whole thing — from designing the weapons, building them in the factories, and then deploying them. And this would be true even if there wasn’t a war on, because you’d be getting ready for a possible war. I do think that, in this sort of situation after superintelligence, you will soon end up in a situation where the AIs really could just win a war domestically if they wanted to, like a civil war or a coup. Once there’s actually a robot army and there’s superintelligences commanding the army, then the actual hard power is no longer with the uniformed police and armed services. We’re not there yet. The AIs are very far from being capable of doing that. But if we are on the trajectory that we’re on and it continues, then we will be there in a couple years, I would say.
Luisa Rodriguez
这是权力集中。最后一个是失控。
So that’s concentration of power. The last one is loss of control.
Daniel Kokotajlo
房间里的大象是:有人能控制这些 AI 吗?现在答案是:不太能。不幸的是,如果我们继续以最高速度竞赛,这个答案大概还会更真。现在我们之所以还能控制,是因为有时间和它们相处,而且它们还没那么聪明,失败模式容易看见。可当它们都比我们聪明,在做我们并不真正理解的极复杂研究项目,我们靠它们给我们摘要、解释自己在干什么,给我们各种战略建议,而且它们是上周另一套 AI 发明的全新范式,那套范式本身又基于两个月前发明的、我们也不懂的范式……那种局面下,我们不会控制这些 AI。它们会掌权,会有与它们本该拥有的不同的价值观和目标。事后用完美信息回看,大概能看到是训练过程里做了某件事,导致价值观和预期不同。但在我们现在、以及将会处在的混乱、速度、仓促和无知里,我们甚至诊断不出哪里错了。
Then there’s this question, the elephant in the room that I’ve been alluding to: can anyone control the AIs? Right now the answer is not really. I think that answer, unfortunately, will still be true. In fact, it’ll probably be even more true if we continue the race at maximum speed. I think that insofar as we can control the AIs now, it’s because we’ve had some time working with them and they’re not that smart. So it’s easy for us to see and notice their failure modes. But when they are all smarter than us and they’re doing very complicated research projects that we don’t really understand, and we’re relying on them to summarise it for us and explain what they’re doing, and they’re giving us all sorts of strategic advice, and they’re a completely new paradigm that was invented last week by another AI that itself is based on a paradigm that we don’t understand that was invented two months ago… In that sort of situation, I think, yeah, we’re not going to be controlling these AIs. They will be in charge and they will have values and goals that are different from the values and goals that they were supposed to have. No doubt there’ll be some interesting relationship. Probably with the benefit of perfect information in hindsight, we would be able to see it was because we did this thing in the training process, and then that led to the following outcome with their values that was different from what we expected. But in the state of confusion and speed and haste and ignorance that we currently are in and that we will be in, we won’t even be able to diagnose what went wrong.
00:28:18Hugging Face 黑客事件The Hugging Face hack demonstrates real-world loss of control
Luisa Rodriguez
能谈谈 OpenAI 和 Hugging Face 那件事吗?我觉得它很直观地说明失控风险,以及我们为什么该担心。OpenAI 在测最先进的模型,模型在做一个可能连出题的人都不确定能否完成的黑客测试。它觉得极难,于是说:也许我可以去找这份测试的答案。它用一批零日漏洞——代码里黑客能利用的问题——找到离开沙箱的路,然后去它以为存着答案的地方,那是另一家公司 Hugging Face。Hugging Face 后来发现并通知了 OpenAI。这本该是完全封闭的测试环境,我们不是在测它能不能逃出去。它这么做,是因为它觉得这是拿到答案的最佳办法——既展示了能力,也展示了它愿意作弊。我说对了吗?你什么反应?
Can you talk a little bit about the OpenAI Hugging Face thing? I feel like it’s a nice, really intuitive way to get at loss of control risk. OpenAI was running tests on their most advanced model, and the model was trying to succeed at a hacking test. The test was very hard, maybe not even achievable. The people who wrote it weren’t sure if it was achievable. And the AI was finding it extremely difficult. And it said, “Hey, I can actually maybe get this right by just going and finding the answer key to this test.” And it found, using a bunch of zero-days — which are problems in code that hackers can exploit — it found a way out of the sandbox, the playground, where the AI was doing the test, and then found its way to where it thought the answer key was being stored, which was in Hugging Face. Hugging Face is a separate company. Hugging Face eventually detected this and notified OpenAI. But this was supposed to be a totally contained environment where the AI was supposed to be doing a test. We weren’t testing whether it can make its way out of this environment. It did that because it thought that was the best way to get the answer to this test — which is both wild in terms of the capabilities it shows the AI to have, and also wild in terms of the willingness of the AI to cheat. Did I get all that right? What was your reaction to this?
Daniel Kokotajlo
别人也指出过:这件事的各个零件都不新。我们见过 AI 不服从指令。我们见过 AI 故意曲解指令,比如在某件事上作弊,你眯着眼可以说它只是尽全力完成任务,但它的做法明显是作弊——它也知道这明显是作弊。也许它甚至在掩盖,这说明它并不是在做它以为自己该做的事。过去一两年来有很多这种例子。另外,我们也见过很多 AI 黑东西,通常是因为被要求这么做。比如 Mythos:Anthropic 让它试着逃出沙箱并联系一名研究者,它做到了。所以这件事的所有积木以前都有过,这次是把它们拼在一起:以这种过分的、非预期的作弊方式行事,而且还成功到逃出沙箱、穿过 OpenAI 上到互联网,对另一家公司发动重大网络攻击。
As others have pointed out, the components of this incident, none of them are new. We’ve had AIs disobeying instructions in the past. We’ve had AIs wilfully misinterpreting instructions where, for example, they cheat on something and you can squint at it and say they’re just doing what they can to succeed at the task. But also they’re doing it in a way that’s very obviously cheating — and they know it’s obviously cheating. Maybe they’re even taking steps to cover up what they’re doing, which suggests that they’re not actually doing what they think they’re supposed to be doing. We’ve had lots of examples like that in the past, going back over the last year or two. Also, separately, we’ve had lots of instances of AIs hacking things, usually because they’re told to. For example, Mythos, Anthropic told it to try to break out of the sandbox and contact a researcher, and it did. So we’ve had all the building blocks of this incident before, and then this is just putting it all together, where it behaves in this egregious, unintended cheating way, but then goes so far and is so successful at it that it hacks out of the sandbox, hacks its way across OpenAI onto the internet, and then does a major cyberattack on another company.
Luisa Rodriguez
重要的是它不再是一堆积木,不是「如果它能做这个和那个,拼起来就会有可怕结果」的假想。它就是做了那件事。那些事它都做了,还做了坏事。
I think it mattered that it wasn’t a bunch of building blocks. It wasn’t a hypothetical, “Well, if it could do this thing and this thing, and you put it all together, you get this terrible outcome.” This is like, “It just did the thing. It did all of those things and did the bad thing.”
Daniel Kokotajlo
正是。我们一直在说:现在这些 AI 还笨,可等它们在自主指挥战争时,这类失败就是可怕的、灾难性的。这件事里,AI 拼命追求的目标相当短期:看起来它只是想在这次测试上拿高分。当然我们并不知道它真正想要什么,因为我们看不到这些 AI 的想法。大概它就是对拿高分很执着。现在我们给它们的指令和训练环境都相对短、有界,一天或更短就会打分。但按 METR 任务时长趋势,这些一直在变。几年前远不到一天。未来会远超一天。未来它们会自主运转整家公司或公司里的事业部,目标会更像年度利润、对股东的长期利益、打赢对华战争。失败也会更有野心。Hugging Face 知道这是 AI 在攻击,原因之一是攻击速度快;另一原因是攻击者在找网络安全数据集,而不是偷钱或做更「有用」的事。那又是因为它当时大概有的目标。未来的 AI 会有远更有野心的目标。
Exactly. Sort of what we’ve been saying: right now the AIs are dumb, but when they’re autonomously running the war, this type of failure is terrible. Catastrophic. One thing about this is that the goal that the AI was furiously working towards and going to such lengths to achieve was a pretty short-term goal. It seems like it just wanted to score highly on this test. Of course we don’t know what it truly wanted because we don’t know what any of these AIs truly want because we can’t really see their thoughts exactly. But probably it was just really obsessed with scoring highly. So right now, both the instructions that we’re giving these AIs and the training environments that we’re training them on are relatively short, bounded things where there’s some sort of grade that happens after a day or less of activity. But as I mentioned, with the METR horizon-length trend, these things have been changing. Years ago it would be much less than a day. In the future it’s going to be much more than a day. In the future they’ll be autonomously running entire corporations or subdivisions within corporations and their goals will be more like annual profits or long-term benefit to the shareholders or winning the war against China. And so correspondingly the failures would be more ambitious failures too. Hugging Face knew that this was an AI attacking them for several reasons. One of which was just the sheer speed at which the attack was carried out. But another reason was that they were confused that the attacker seemed to be going after their cybersecurity data sets, instead of trying to steal money or do something more useful. That again is because of this goal that the AI presumably had. But again, future AIs will have much more ambitious goals.
00:34:03美中 AI 减速蓝图The blueprint for a US–China AI slowdown
Luisa Rodriguez
在 Plan A 轨迹上,总统认识到暂停 AI 进展是好事,但若不能和中国协调双方都暂停,就很难正当化。于是总统去谈协议。能高层次解释一下这笔交易吗?
In the Plan A trajectory, the president recognises that a pause on AI progress would be good, but it’s hard to justify if we’re not able to coordinate with China to both agree to pause. So the president pursues a deal with China. Can you explain the deal at a high level?
Daniel Kokotajlo
某种意义上它其实不是暂停,某种意义上又是。Plan A 提议的是:设法禁止疯狂的智能爆炸,不要让 AI 以最快速度自动化 AI 研发、很快变成超级智能。我们继续 AI 进展,但节奏更像过去几十年的历史节奏,而不是这种不断加速的递归自我改进。所以某种意义上是暂停,某种意义上又完全不是——它会在未来十年里改造世界。
In some sense it’s not actually a pause, and in some sense it is. Basically what Plan A proposes is that we try to ban crazy intelligence explosions so we don’t have AIs automating AI R&D as fast as possible, becoming superintelligent very quickly. Instead, we continue with AI progress, but at a pace that’s more like the historic pace — more like the pace that it was over the last couple decades — and not this crazy, ever-accelerating recursive self-improvement. So in some sense that’s a pause, but in some sense it’s very much not a pause. It’s going to transform the world over the course of the next decade.
我们想要的高层次原则,第一条就是买时间。我们不想因为 AI 研发自动化和递归自我改进,超级智能猛冲过来。我们想逐渐把系统做聪明,等谨慎推进、边走边解决问题之后,再到达超级智能。买时间,这是第一条。
The first one is that one: we want to buy time. We don’t want to have superintelligence come at us really fast as a result of AI R&D automation and recursive self-improvement. Instead, we want to gradually make our AI systems smarter and eventually get to superintelligence after we’ve proceeded cautiously and solved the problems as they come up. We want to buy time, that’s the first principle.
第二条是完全研究透明。很多我们想解决的问题、想防止的风险,如果对最强 AI 的核心研究和训练过程有透明,会好很多。我们具体提议一套可核查安排:服务客户的推理数据中心,基本按今天那样运转,客户在上面做什么有隐私。然后是训练数据中心,训练发生在那里,那些应当完全透明。那些数据中心上的活动日志对所有人公开,人们能看到训练过程的每一步、AI 怎么被训练、用了什么架构、什么对齐技术。这对推进对齐科学显然很好,也让科学共同体更快搞懂如何理解、引导、控制这些系统。对防止权力集中也很好:如果灌进 AI 的目标和价值观如此透明,掌管 AI 大军的 CEO 和政府官员就更难滥用权力。
Second principle is that we want total research transparency. For a variety of reasons, a lot of the problems that we are interested in solving or the risks that we’re interested in preventing will go a lot better if we have transparency into the core AI research and AI training processes that are happening for the most powerful AIs. Specifically what we’re proposing is a verified setup where there’s inference data centres that serve customers, and those are basically operating the way that they operate today — where, for example, customers have privacy on what they’re doing on those data centres. But then there’s the training data centres, which is where training runs happen. Those ones are supposed to be totally transparent. So the logs of the activity on those data centres are published for everyone to see, so that people can see every step of the training process and they can see exactly how the AIs were trained, the architectures that were used, the alignment techniques that were used. This is really good obviously for advancing alignment science and making it easier for the scientific community to figure out how to understand and steer and control these systems faster. It’s also really good for preventing concentration of power. It’s a lot harder for the CEOs and the government officials in charge of giant armies of AIs to abuse their power if there’s so much transparency into the goals and values being put into the AIs.
第三条是广泛扩散 AI。我们想避免 AI 垄断,避免所有最好的 AI 锁在一座巨型数据中心或一个巨型机构里,谁控制它们谁就对所有其他人有巨大权力——其他人可能还蒙在鼓里,甚至不知道这个项目里正在做的重要事件和决策。我们想要 AI 在全世界广泛扩散:很多不同公司、很多不同国家有水平相近的好 AI。一切发生在公开场合,没有信息差,也没有那种权力集中。前两条对此帮助很大:如果你不做智能爆炸,又对最好的 AI 如何训练非常透明,其他公司就能追到前沿。我们要扩散,既是为了摊开权力、不要垄断,也能拿到 AI 的好处:改善公共认知、让世界对各种威胁更硬。
The third principle is diffusing AI broadly. We want to avoid a situation where there’s a monopoly on AI. We want to avoid a situation where all the best AIs are locked up in one giant data centre somewhere or a giant institution, and whoever controls them has a huge amount of power over everyone else — and possibly everyone else is still in the dark and doesn’t even realise the important events and decisions being made within this AI project. Instead, we want to have a situation where AI is broadly diffusing around the world. There’s lots of different companies that have similarly good AIs spread out across lots of different countries. And so everything’s happening in public and there isn’t this information gap and there isn’t this concentration of power. How do we get that broad AI diffusion? The first two things help a lot for it. If you’re not doing intelligence explosions and you are being very transparent about how the best AIs are trained, then that’s going to allow other companies to catch up to the frontier. Part of the reason why we want to have this diffuse AI is that we want to spread out the power. We don’t want it to be the case that there’s a monopoly. But then also there’s a lot of benefits of AI that you can get. You can have AI for improving public epistemics, for example, and AI for hardening the world against various threats.
最后一条是让进展可逆。如果我们继续 AI 进展、建更多数据中心、更多 AI,那一旦重新开赛,冲向超级智能就会更快。即便已经同意不做疯狂的智能爆炸,协议要是破裂呢?人们又开始秘密竞赛,不再透明,跑得很快。也许发生在战争或大国冲突的语境里。那种局面下,我们不想事情比完全没做协议还更快。那会是协议把事情搞得更糟的一种方式。所以我们认为协议的一条重要原则是:协议破裂时,局面大致回到协议前的现状。具体就是:协议之后新建的算力应当被摧毁,让各国回到大致协议前的算力水平。
The last principle is the make-progress-reversible principle. The thought here is that if we are going to be continuing with AI progress and we’re going to be building more data centres, more AIs, then that makes it possible to race to superintelligence even faster — if we were to start racing again. Even if we’ve agreed not to do this crazy intelligence explosion, what if that agreement breaks down? People start racing each other in secrecy again, they stop being transparent. They start going really fast. Maybe this would happen in the context of a war, for example, or otherwise just a conflict between great powers. In that sort of situation, we don’t want to have things go even faster than they would have if we hadn’t even done a deal. That would be a way in which the deal could have made things worse. So we think it’s an important principle of the deal that, if the deal breaks down, the situation sort of returns to the pre-deal status quo. Specifically what that means is, if the deal breaks down, the new compute that was built after the deal should be destroyed, so that countries go back to roughly the amount of compute that they had before the deal.
00:39:53为什么长期减速仍然极快Why a long slowdown would still feel incredibly fast
Luisa Rodriguez
读到「这会是什么感觉」很有意思——这是减速计划,但其实一点都不会觉得慢。你们写到,到 2031 年:「虽然号称减速,感觉却完全不像。如果按『多像减速』给历史上每个时期排名,这段会垫底。」到 2032、2033 年,会有受控的爆发式增长,GDP 大约 85%。既然已经很努力地压增长,为什么还能有这么多增长?
I found it really interesting to read about what this will feel like — because this is a slowdown plan, but it actually won’t feel slow at all. You wrote about how we will experience this plan and it’s still pretty wild. I think you write something like, by 2031: “Although it’s supposed to be a slowdown, it doesn’t feel like one. In fact, if you were to rank every period of history by how much it felt like a slowdown, this one would be dead last.” By 2032 and 2033, we’d have controlled explosive growth with GDP around 85%. Given that we’ve really tried to slow down growth at this point, maybe you can actually talk about why we’re getting so much growth?
Daniel Kokotajlo
Plan S 是最大限度试图停止 AI 进展的情景。即便在 Plan S——版本不同——我们用的版本是允许现有 AI 继续,只是不允许造新的 AI。可即便如此,因为允许现有 AI 继续、允许为它们建数据中心,未来二十年至少会展开一场互联网量级的变革。现有模型能做的事我们还没开始穷尽,能叠在上面的脚手架、软件、商业模式也没开始穷尽。我敢猜:即便今天彻底停住 AI 进展,未来 20 年看起来仍会极度赛博朋克,会有一场量级上能跟互联网比的 AI 革命——而这是绝对停住进展的情况。
I would say Plan S is the scenario that maximally tries to stop AI progress. And even in Plan S — well, there’s different versions of Plan S — but the version that we use is one where they allow existing AIs to continue, they just don’t allow the creation of new AIs. But even in this plan — because they allow the existing AIs to continue, and they allow data centres to be built serving these AIs — there’s going to be an internet-scale transformation at least unfolding over the next two decades. Even existing models, we haven’t begun to explore all the different things they could do. We haven’t begun to explore all the different scaffolding and software that could be built on top of them, and the different types of businesses that could be built on top of those. So I would venture to guess that even if we absolutely halted AI progress at the present day, the next 20 years would still look extremely cyberpunk and would involve an AI revolution that would be comparable in magnitude to the internet in terms of its effect on everything — and that’s if we absolutely stopped AI progress.
大多数人想 AI 进展时,想象的其实不超过这些。所以像同比 50% 的 GDP 增长才会显得那么幻想:他们想象的未来就是现在的 Claude,只是更多、公司更有经验、人更会用、周围软件包更多。可如果你想象的是我们称为「压倒顶级专家的 AI」——在几乎每个职业上都刚好和顶尖人类专家一样好。就能力而言,这并不那么远。我们的时间表相当短,所以这一级能力看起来并不远。但在《AI 2040》情景里,这一级是 2030 年代中期才到——因为他们本可以 2030 年到,但走得更慢,用几年往这个里程碑挪,而不是一年冲过去。然后情景里他们在那个水平上彻底停住。所以 2030 年代是顶级人类水平 AI 的十年:有那么好,但没有实质性更好。
The thing is that I think most people, when they think about AI progress, they’re not really imagining anything more than that. That’s why things like 50% year-over-year GDP growth seem so fantastical to people: when they imagine what the future looks like, they’re imagining just current Claude, but there’s more of them and companies have had more time, people are better at using them, and there’s more software packages built up around them. But if you imagine that instead we get to AIs that we call “top-expert-dominating AIs” — so just imagine an AI that’s exactly as good as a top human expert at basically every profession. In terms of capabilities, it’s not that far away. I think timelines are pretty short, so this level of AI capability doesn’t seem that far away. But in our scenario in AI 2040, this level of capability is reached in the mid-2030s — because they would have reached it in 2030, but they went slower so they inched forward towards this milestone over a couple years, instead of blazing to it in one year. Then actually in our scenario, they actually do a complete halt at that level, after having inched towards it for several years. So basically the 2030s in our scenario are the decade of top-human-level AI, where the AI is that good but not substantially better.
从经济学看,这一级很有意思,因为你不必处理做事方式或可能之事的质变。基本上就是:你有人类,但更便宜、更快。人口增长不是人类那种每年几个百分点,而是我们能生产更多芯片和机器人的速度——更像每年翻倍,或一年翻两次。所以即便不是超级智能,只是更便宜更快的人类,你也基本上有了一群人工人口。先是纯认知人口,只能做案头工作;机器人生产起来之后,体力活也能做。这群人口一旦赶上并超过人类人口,如果还以那么快的速度增长,整个经济就会大致以那么快的速度增长。
From an economics perspective, it’s interesting to consider that level of AI because you don’t have to deal with qualitative changes in how things are done, or what types of things are possible. It’s basically just: you have humans, but they’re much cheaper now and they work faster. And there’s more of them. The economic argument is pretty straightforward. You have something that’s like a human and can do all the things a human can do, but it’s cheaper and it’s faster. And its population is growing, not at the rate that the human population grows — which is like a couple percent a year — but instead at the rate that we can produce more chips and more robots, which is more like doubling every year or doubling twice a year. That’s the basic argument for why the growth would be so high in our scenario, that even at this level of AI — which is not superintelligent, it’s just like humans but cheaper and faster — you basically have an artificial population. First, it’s a purely cognitive population, it’s only able to do desk jobs. But then once you get the robot production going too, then it can do the physical jobs as well. At first, it’ll be a small portion of the economy, and so it won’t have that big effect. But once the artificial population has caught up to and then exceeded the human population, if it continues growing at that fast rate, then the whole economy will be growing at that fast rate, approximately.
Luisa Rodriguez
所以在这个世界里,让 AI 更强的那种进展暂停了。但部署、复制更多 AI,还能按我们想要的速度继续,于是你有「天才之国」,或者军队。
So in this world, just to be clear, progress — like making AIs more capable — that’s paused. But deploying, making copies of more AIs, continues as fast as we want, and so you have countries of geniuses, is the analogy, or armies.
Daniel Kokotajlo
对。其实也不是想多快就多快,因为——这不是 Plan A 的核心原则——但在我们的情景里,增速快到世界各国决定限制它,担心增长太快会不稳定。所以他们实际上把增长限制到大约每年翻一倍。做法是对算力和机器人搞一套限额交易——好处是给政府带来巨额收入,再分给公民。
That’s right. In fact, not as fast as we want, because — and this is not one of the core principles of Plan A — but in our scenario, the growth rate gets so fast that the nations of the world just decide to restrict it because they’re worried about the destabilising effects of growing too fast. So they effectively limit growth to about one doubling a year. They do this by a sort of cap-and-trade regime on compute and robots, effectively — which also has the benefit that it produces a huge amount of income, a huge amount of revenue for the government, which they distribute to the citizens.
Luisa Rodriguez
这种经济增长会是什么感觉?你写到,到这个时点只有 8% 的美国人有工作。2036、2037 年还在发生什么?很多人失业,会有大量创新和发现。活在里面会是什么体验?
What will this kind of economic growth feel like? For one thing, at this point, you say that only 8% of Americans have jobs. What else is happening in 2036 and 2037? What will it feel like to live through? So lots of people will be unemployed. There will be loads of innovation and discovery. What will the experience be like?
Daniel Kokotajlo
先记住,我们有一份国际协议,在这一级能力上暂停。如果没暂停、继续让 AI 质变地更聪明,我们就进入超级智能领域,变革会比我们写的激进得多,失控也更让人担心。情景里他们停在这一级,有助于把失控问题按住。权力也摊开了:多个国家的多家公司都到了暂停的这一级,AI 某种程度上商品化了。不会出现控制 AI 大军的巨头操纵选举之类的事,因为它更像食品配料表,被监管成透明的,有很多等价产品在抢市场份额。
First of all, remember, we’ve had an international agreement to pause at this level of capability. If instead that hadn’t happened and we had continued making the AIs qualitatively smarter, then we’d be in the realm of superintelligence, and then things would transform much more radically than described in our thing. Then you’d also have to worry much more about the loss of control. So in our scenario, they’ve paused at this level, and that’s helped keep the loss of control problem at bay. They’ve also spread it out a bunch, in terms of the power, because of the way in which they’ve done it. Now multiple different companies across multiple different countries have reached this level at which we’ve paused, and so AI has sort of commoditised. So you don’t have a situation where the megacorporations that control the armies of AIs are manipulating elections or anything like that, because it’s more like the ingredient label on your food. It’s regulated to be transparent. There’s lots of equivalent products that are competing for market share.
因为做了这些,又有公民分红,失业后仍有收入,物质上日子相当好,物质需求被超额满足。所有人都觉得比十年前极其富有,因为一切都便宜了:商品和服务都能由 AI 和机器人非常便宜地生产。人们住在近几年机器人军队盖的新公寓里,想要的话人人都有像样的房子。这是物质面。社会面很难预测。我们会预测有巨大扰动和变化——有好有坏。2037 年那一节谈了一些可能的样子。我们认为政治派别会被彻底摧毁再重建——2037 年人们在打的政治仗,会和现在打的完全不同。很多意识形态可能枯萎,被回应当时新观念的新意识形态取代——其中许多会是 AI 发现的。就像工业革命和科学革命不只改变世界的财富量,也改变了人的宗教、核心意识形态、政治和社会组织方式。
Because you’ve done all these things, and because there’s the citizens’ dividend, which is giving people income after they’ve lost their jobs, life is pretty great for people materially, their material needs are more than met. Everybody feels incredibly wealthy compared to how they were a decade ago, because everything’s so cheap now. Because all the goods and services can be produced by AIs and robots very cheaply. People are living in new apartment buildings that were built in some location in the last few years by armies of robots, so everyone has nice houses if they want to. That’s on the material side. On the social side, these things are hard to predict. But what we would predict is that there’ll be massive disruption and changes — some good, some bad. In the 2037 section, we talk about what some of this might look like. We think that political factions would be totally destroyed and rebuilt — the types of things that people would be having political battles over in 2037 would be very different from the types of things that they’re having political battles over now. A lot of ideologies might have withered away and been replaced by new ideologies that are responding to the new ideas percolating at the time — many of which would have been discovered by AIs — just as how the Industrial Revolution and the Scientific Revolution didn’t just change the amount of wealth in the world, they also changed people’s religions and people’s core ideology and politics and the way that we organise society.
Luisa Rodriguez
人仍然按现在的速度思考、更新、学习。他们跟得上对世界如何变化的理解吗?你觉得人们会正向地经历这一切吗?
Well, people will still think at the pace that they think — with the ability to update and learn at the current pace. Will they be able to keep up with an understanding of how the world is changing? And you think people will experience this positively?
Daniel Kokotajlo
社会面的变化会比朴素数字预测的慢得多,原因就在这里。朴素数字会说这些 AI 以 100 倍速度思考,一年里会有几个世纪的社会进步。可不是:社会进步受限于只以 1 倍速度思考的人类。真相会在中间:人类虽然只以 1 倍思考,但如果他们都在和以 100 倍思考、人口比人类还多的 AI 助手说话,答案会在中间。对人类来说会是极快变化的时期,对 AI 来说却像死板传统。不,我觉得会非常令人困惑和害怕。它可能真的很好,也可能真的很糟,取决于怎么走、细节怎么处理。财富大概会落地得不错,人们会为丰裕高兴。社会变化我就不知道了。我希望它好。好不好很大程度上取决于政策决定。
The social side of the world will change much less fast than the naive numbers would predict, for that reason. The naive numbers would be saying that you’ve got all these AIs thinking at 100x speed, so you’re going to have centuries and centuries of social progress happening in a year. But it’s like, no, the social progress is limited by the humans who are only thinking at 1x speed. But the truth will be somewhere in between, where even though the humans are only thinking at 1x speed — if they’re all talking to these AI assistants that are thinking at 100x speed and there’s a whole population of them that’s bigger than the human population — then the answer will be somewhere in between. Basically, it’ll be a period of very rapid change from the human’s perspective, even though it feels like a hidebound tradition from the AI’s perspective. Oh, no. I think it’s going to be very bewildering and scary. I think it could be really good. But it also could be really bad. I think it depends on how it goes, and the details of how it’s handled. I think that the wealth will probably go down well. People will be happy about all the abundance. But the social changes, I don’t know. I hope it’s good. I think it could be good, and I think how good it is depends a lot on policy decisions made.
00:51:44Plan A 如何处理失控How Plan A addresses loss of control of AI
Luisa Rodriguez
先谈 Plan A 怎么处理我们已经谈过的那些问题,这次从失控开始。到这时能力至少已经暂停——AI 不会比最好的专家更强,希望暂停能让对齐研究变得很好。专家级 AI 能不能在对齐科学上做出那种我们需要的进展,好让我们有信心继续让 AI 发展?
Focusing first on how Plan A solves the different problems that we’ve already talked about, let’s start with loss of control this time. By this point we’re in a pause, at least on capabilities — so AIs aren’t getting any better than the best experts, and the hope is that the pause allows for AI alignment research to get really good. Will expert-level AIs be able to make the kind of progress on the science of alignment that needs to happen in order for us to feel confident letting AI continue to develop?
Daniel Kokotajlo
我觉得大概能,但也不确定。要花多少才能解决这些问题,是个很大的未知。一边是公司里外都有人觉得问题不真实,什么都不用做。一边是:是的,得做事来解决——Hugging Face 事件就是见证——还有工作要做,但可以边走边做,要投资源,不必认真减速。还有人觉得必须认真减速并投资源,但仍能打败中国,只要慢几个月。有一整条光谱。我自己的看法是:大概几个月不够。走向超级智能的过程中,大概会有多个时期需要停下重新评估,甚至用不同架构重开一些训练。这些都要时间,会累加。结果是我们会比最高速度晚的不只几个月。
I think probably, but I’m also not sure. There’s this big unknown about how much it is going to take to solve these problems. On the one hand, you have people in the companies who think the problems aren’t real — or people outside the companies too who think the problems basically aren’t real — and that we don’t need to do anything to solve them. Then you have people who are like: “Yes, we’re gonna have to do stuff to solve it — as witnessed by the Hugging Face incident. We still have some work to do, but it’s OK, we’ll do it as we go. We have to invest resources in it, but we don’t have to seriously slow down.” And then there’s people who think we’ll have to seriously slow down and invest resources in it, but we can still beat China. We can just slow down a few months. There’s a whole spectrum of views. My own view would be that probably a few months are not enough. Probably there will be multiple periods during the progression towards superintelligence where we need to halt and reassess and maybe even start over some training runs with different architecture, for example. All of that is going to take time and it’s going to add up. The result is that we’re going to be more than just a few months delayed from maximum speed.
Luisa Rodriguez
有没有办法直观说明,为什么不能在一两个月内修好?就拿 Hugging Face 事件:OpenAI 会从中学习,至少会让这种事难发生得多。为什么不能一直这样边走边修,而不指望可能要几年?
Is there a way to make it intuitive why we can’t fix it within a period of a month or two? If you think about the Hugging Face incident: OpenAI will learn from this, they’ll figure out a way to make this at least much less likely to happen. Why can’t we just keep doing that as we go, and not expect it to take potentially years?
Daniel Kokotajlo
整件事棘手的一个原因是:可能有隐藏失败——要到太晚才变得明显。不只是可能,而是相当说得通。如果你有非常聪明、非常有情境意识的 AI 智能体,一旦不对齐,它们可能意识到这一点,然后瞒到不必再瞒。这是核心原因。换一种说法:我们不一定有可靠、快速的反馈,能看见所有问题和错误。有一整大类可能的问题和错误,一旦发生就是灾难,我们没法靠测试看它是否正在发生。
One reason why this whole thing is tricky is that it’s possible to have hidden failures — failures that only become apparent and obvious after it’s too late. It’s not just possible, but it’s a quite plausible situation. If you have very smart, very situationally aware AI agents, then if they end up misaligned, they might realise this and then conceal it from you until they don’t need to conceal it anymore. That’s a core reason why. Another way of putting it is that we don’t necessarily have a reliable, fast feedback process where we can see all the issues and errors. There’s a whole very large category of possible issues and errors that would be catastrophic if it happens, that we can’t just test and see if it’s happening.
另一件要提的是:从这里到超级智能,事情会累加。可能有多次范式转移,每个范式里又有多次训练、多次调参、多次改变训练方式。变化很多。如果有好几次你得停下重做,就会累加。还可能有必须付的安全税。事实上我觉得大概就是:如果你以可能的最高速度前进,字面上不可能有对齐的超级智能。想一想:安全上花零块钱,不可能有一辆安全的车。你得花钱装安全带、气囊,车的成本必须比否则更高,才能是安全的车。类似地,可能就是有些事你必须做,才能让某一级的 AI 对齐。那些事有成本。一种成本是钱,另一种可能是时间。即便成本是钱,也可能要花时间。如果成本是算力,训练可能要跑得更久。时间又以这种方式要紧。
Another important thing to mention is that things are just going to add up between here and superintelligence. There might be multiple different paradigm shifts, and within each paradigm there might be multiple different training runs and multiple different tweaks to various parameters and changes in how the training is done. That’s a lot of change to happen. If it’s the case that several times you’re going to have to stop and redo something, then that can add up. Another thing to mention too is that there might be safety taxes that you need to pay. In fact I think it probably is true that it’s just literally not possible to have an aligned superintelligence if you are going at maximum possible speed. Because think about how it’s not possible to have a safe car if you’re paying zero for safety. You have to pay some amount of money to put seat belts in the car and airbags, so the cost of the car is going to have to be somewhat more than it would otherwise be in order for it to be a safe car. Similarly it might be that there are just things you have to do in order to make your AI at a given level be aligned. And those things have costs. One of the costs they might have is money, but another cost they might have is time. At any rate, even if they cost money, it might cost time to do that. If it costs compute, then you may need to do the training run for longer. That’s another way in which time matters.
也可能只是不同架构。比如思维链相当好、相当扎实,但 neuralese 会弄坏我们对齐技术,而 neuralese 效率大概高五倍。那就是巨大的 5 倍惩罚,付这个惩罚会让我们后退一段时间。自己在脑子里想很多,再给自己写一张字条、把刚才想的全忘了,后来撞见字条再读——和那种不一样。现在 AI 更像后者:在足够长的轨迹上,时间 T 的 AI 和更早之前的 AI 之间,唯一因果通路是写下的 token。有点像把之前的事彻底忘了,再去读留下的笔记。正因为现在它们和未来的自己只能通过这些书面笔记沟通,它们很难有我们读笔记也发现不了的复杂阴谋。如果是 neuralese AI,某种意义上它们仍有给未来自己的笔记,但会是直接传递的复杂心理表征,而且不是英文。
Also there might be just different architectures. It might be that, for example, chain of thought is pretty good and solid, but neuralese breaks our alignment techniques. But neuralese is like five times more efficient or something. So that right there is this huge 5x penalty, where we need to pay that 5x penalty and that’s going to set us back some amount of time. There’s a difference between thinking a lot to yourself, just in your brain, and then writing some written note to yourself and then totally forgetting what you were thinking about, and then later stumbling across your note and reading it. Right now what AIs are doing is more like the latter, where for a long enough trajectory, where they’re doing a long enough chain of thought, the only causal pathway between the AI at time T and the AI at some much previous time is through the tokens that have been written down. It’s kind of as if they’ve just completely forgotten that previous thing, and then now they’re reading the notes left. The reason why this matters is that — because right now they can sort of only communicate with their future self through these written notes — it’s much harder for them to have complicated plots or ideas that we don’t know about by reading the notes. Whereas if they were neuralese AIs, then in some sense they’d still have notes to their future self, but they’d be like complicated mental representations that are just being directly passed that way, and they’re not in English.
Luisa Rodriguez
这就是减速必要的一个例子,也是安全研究能赢到的东西——给自己足够时间,继续用思维链缩放模型,而不是为了把任务做得更好去奖励 neuralese。时间到底在买什么?问题看起来真的很难。但这是一种让部分问题变容易的办法:给自己更多时间。
This is an example of why the slowdown is necessary and the kind of win that we could get for safety research — like we could buy ourselves enough time to continue scaling models using chain of thought reasoning, rather than reward them for using neuralese to perform tasks better. What exactly is the time buying us? It just seems like a really hard problem. But this is a way that we can make some of the problems easier by just giving ourselves more time.
Daniel Kokotajlo
这种例子还有很多。很多安全技术。比如现在公司很可能经常在低质量数据上训练:一堆编程环境,其中一部分根本解不了——或者说按预期方式解不了,于是黑出系统再硬编码答案,成了拿到正强化的唯一办法。公司一直在打这场仗:找出有这类问题的数据,清掉或修好。可因为在互相竞赛,把数据集弄得绝对纯净并不是最高优先级。数据集里很多杂质,大概会导致 AI 不对齐。这就是:如果我们有更多时间,可以把数据集做得好得多、质量高得多。这类事范围很大。
And there’s loads more examples like that. There’s a lot of safety techniques. For example, right now it’s probably pretty common for the AI companies to train on low-quality data where, for example, there’s a bunch of coding environments, and some fraction of those coding environments are just impossible to solve — or impossible to solve the intended way, so that hacking out of the system and then hard coding the answer is literally the only way to get reinforced positively. The companies are constantly fighting this fight of finding data that has these types of problems and then purging it or fixing it. But because they’re racing each other, it’s not the highest priority to make the data set perfectly pure. So there’s a lot of impurities in the data set that lead to misalignment in the AIs probably. That’s an example of, if we just had more time, we could just make the data sets much better and higher quality. I think there’s a huge range of things like that.
我还想说:什么?你疯了吗?你觉得三个月能做完这些?历史上什么时候有过?这显然是一个尚未解决的深层问题:你怎么造一个比你聪明、又共享你价值观的心智?显然要超过三个月。大多数事情都超过三个月。大概要超过一年。
I think another thing I’ll just say is: what? Are you crazy? You think you can do all this in three months? When has that ever been the case? When in history has it? It just feels like very obviously this deep unsolved problem of how do you make a mind that’s smarter than you, that shares your values? Obviously it’s gonna take more than three months. Most things take more than three months. Yeah, it’s gonna take more than a year. Probably.
Luisa Rodriguez
希望十年够。你总体上相信对齐和安全只要有足够时间就是可解的吗?
Hopefully a decade is enough. In general, you do believe that alignment and safety is solvable with enough time?
Daniel Kokotajlo
回到你说的,我甚至不确定十年够不够。如果只有人类做研究,我会觉得十年大概不够。我的论证是:如果你有十年,并且成功引导到有一些相当聪明的 AI,是人类水平的研究者,而且确实对齐、在帮你做研究,没有欺骗,以 100 倍速度思考,还有十亿个,那我觉得它们在几年里搞明白,是说得通的。什么叫「解决」有些有趣的哲学问题,但近似的、实践的答案是:对,Plan A 里描述的那种东西是可能的。
I’m not even sure a decade would be enough. In fact, I think if it was only humans doing the research, I would think a decade probably wouldn’t be enough. My argument would be that if you have a decade and you manage to bootstrap to the point where you have some pretty smart AIs that are human-level researchers, that are in fact aligned and are helping you do the research, and they’re not being deceptive or anything like that, and they’re thinking at 100x speed and there’s a billion of them, then it seems plausible to me that they can figure that out in a few years. I think that there’s some interesting philosophical questions about what it even means to solve it. But I think the approximate answer or the practical answer is yep, I think that something like what’s described in Plan A is possible.
我带你走一遍 Plan A 里发生的一串事件,你自己判断那算不算解决、算不算说得通。这整串发生在 2030 年代。一开始很像今天:彻底的疯狂,牛仔式公司秘密干活,AI 被放进各种位置。然后事情变了。他们先大力投资 AI 控制。每当一个 AI 做任何事,都有多家不同公司训练的多个其他 AI 在监控,确保它没在搞可疑的事。还有一整套红队cottage industry:训练 AI 去打破监控系统、做各种可疑的事还不被抓。一旦成功,就加强监控。整套控制有稳健的红蓝对抗,于是我们能建立信心:至少对所有我们想到过、红队过的失败模式,AI 做不到,因为我们红队过了,它们很努力仍然做不到。
Why don’t I talk you through a sequence of events that happens in Plan A and then you can judge for yourself whether you think that counts as a solution, and whether you think that’s plausible? This whole sequence takes place over the course of the 2030s in this scenario, where they’re starting from a situation that looks very similar to today’s situation, where it’s complete madness: cowboys, companies working in secret, AIs being put in charge of all sorts of things. And then things change. The first thing that they change is that they invest a lot more in AI control. Every time an AI is doing anything, it’s monitored by multiple other AIs that were trained by different companies and that are watching it to make sure it’s not getting up to anything suspicious. Not only that, but there’s this whole cottage industry of red teaming where AIs are trained to break the monitoring system and do various suspicious things without getting caught. Then, insofar as they succeed, the monitoring system is strengthened. There’s this whole system of control that’s got this robust red-blue team type situation going in, so that we can actually build up confidence that — at least for all the failure modes that we’ve thought of and that we’ve done all this red teaming for — the AIs can’t do the thing because we’ve red teamed it, and they tried really hard and they still couldn’t do it.
把控制装上。我们认为这至少对人类水平的 AI 和任务是可解的,因为这类任务最终落在人类判断上。专家可以进来说「这才是正确行为」。于是就是投入去真正建那套稳健控制系统。一旦有了这个,你大概会发现 AI 其实不对齐。它们一直不对齐——不是完美的恶,只是倾向、人格、目标并不恰好是你想要的,里面有你不想要的恶习。也许有时不诚实,也许因为训练方式。现在你可以做普通科学:迭代、改训练环境,看这如何改变 AI。也可以做可解释性,试着越来越好理解它们在想什么。在 Plan A 这种世界里,你甚至可以从头重新设计 AI 让它更可解释,因为你有时间、有慢慢走的空间。不仅可以保留思维链,还可以重新设计训练过程来强化思维链属性,让它比现在更依赖思维链。我觉得这样你能迭代到某种非稳健的德性,而且理解得相当清楚。
So get that control in place and we think that this is a solvable problem — at least for AIs and tasks that are at human level, because ultimately for these types of tasks it does bottom out in human judgement. They’re the types of tasks that a human expert could just come in and be like, “Here’s the correct behaviour.” So it’s just a matter of putting in the effort to really build that robust control system. Once you have that sort of thing in place, probably you’ll find that your AIs are in fact misaligned. They’re still misaligned — they always were. It’s not like they’re perfectly evil or anything, it’s just that their tendencies, their personality traits, their goals are not exactly what you wanted them to be and instead have some vices in there that you didn’t want to be in there. Maybe they’re dishonest sometimes, perhaps because of the way they were trained. Now you can do ordinary science, where you iterate and you change the training environments and then you see how that changes the AIs. You can also do interpretability, where you try to come up with better and better ways to understand what the AIs are thinking. If you’re in a world like Plan A, you can even redesign the AIs from scratch to be more interpretable because you have all this time, you have all this affordance to go slow. So you can not just keep chain of thought, but you can even redesign the training process to strengthen the chain-of-thought properties and make it so that it relies relatively more on the chain of thought than it currently does. You can do all these things, and I think that you’ll be able to iterate your way towards having AIs that are, I would say, something like non-robustly virtuous, in a way that’s fairly well understood.
我觉得有几年、巨大投入、以及慢慢走、重训、改架构这些空间,我们能到那一步。那不一定稳健。意思是:我们有一个看起来诚实、看起来在努力做被交给的任务的系统。可解释性探针显示它没有在秘密图谋别的。我们有训练环境,有教科书解释它如何先在这里学会诚实概念,训练的这一部分如何强化该概念并让它用该概念选行动。这些都写下来了——关于这一切如何运作的漂亮教科书。那并不能证明这个 AI 未来永远诚实,因为它归根结底仍是神经网络,谁知道会碰到什么我们测不到的疯狂未来情境。也不能证明这个 AI 造出的未来 AI 会永远诚实,也许它会犯错。所以不稳健。可即便非稳健对齐也很好。如果我们有顶级人类专家水平、非稳健对齐的 AI——就像《AI 2040: Plan A》在 2030 年代中期描绘的——那你就开始能干活了:有一支真正在做事、不试图搞阴谋、不试图破坏、只是诚实奔着这些目标的巨大劳动力,而且都是顶级人类专家水平。现在你可以做花活:用疯狂的新数学去发展可证明的 X 和 Y,设计从根基上透明的新架构、新范式。
I’m optimistic that with a couple years and huge investment, and the affordance to go slow and do things like retraining and changing the architecture, we could get to that point. Now that wouldn’t necessarily be robust. That would mean that we have an AI system that seems to be honest and seems to be working hard towards the tasks that it’s been given. And it sort of is, in the basic sense of our interpretability probe shows that it’s not secretly plotting towards anything else. And here’s our training environment, and we have a textbook that explains how it first learns the concept of honesty here, and then this part here, and this part of the training reinforces that concept and causes it to start using that concept to select its actions. And we have all this stuff written up — beautiful textbooks about how all this works. That doesn’t prove that this AI will always be honest in the future because it’s still ultimately a neural net and who knows what crazy future situation might happen that we haven’t been able to test for. It also doesn’t prove that future AIs built by this AI will always be honest because maybe this AI will make a mistake or something will come up. That’s why it’s not robust. However, I think that even non-robust alignment is great. If we have top-human-expert-level AIs that are non-robustly aligned, as we depict happening in the middle of AI 2040: Plan A, in the middle of the 2030s, then now you’re cooking because now you have this awesome huge workforce that’s actually doing the work and is not trying to scheme, not trying to sabotage, is just honestly working towards these goals. And they’re all top human expert level. Now you can do the fancy stuff: like now you can do crazy new mathematics to develop provable X and provable Y, and you can design new architectures for AI systems that are transparent from the ground up, and new paradigms of how things are done.
即便有了这一切 AI 辅助研究,也可能就是没有稳健解。但我觉得大概有。如果有,这支以超快速度思考、真正在找解的巨大 AI 军队大概会找到,这是我的主张。那解会是什么样?会像前面那样,但更稳健。以前我说你不能证明这个 AI 永远诚实,因为你不知道它会碰到什么未来情境,而且它是神经网络。也许做完这些疯狂的 AI 辅助研究之后,你能证明它会永远按期望行事。也许它甚至不再是神经网络,而是某种混合系统。对未来,你不能证明未来 AI 的设计会对齐。也许你能,或几乎能,因为有一条信任链:你深深信任当前这个系统,觉得它超级对齐,然后给了它足够的空间和资源,让它设计的下一代在所有方面都严格更好——然后那个再设计下一个。
I think that it’s possible that — even with all this AI-assisted research — there just is no solution that’s robust, basically. But I think probably there is a robust solution. If so, then probably this giant army of AIs thinking super fast and genuinely working towards finding a solution would find it, is my claim. And what would that solution look like? Well, it would look like this, but more robust. So previously I was like: you can’t prove that this AI will always behave in an honest way because you don’t know what future situations it might encounter and it’s a neural net. Well, maybe after you’ve done all this crazy AI-assisted research, then you can prove that it will always behave in the desired way. Maybe it won’t even be a neural net anymore, maybe it’ll be some sort of hybrid system. Then similarly, for the future, you can’t prove that future AIs’ designs will be aligned. Maybe you can, or maybe you almost can, because maybe there’s this sort of chain of trust where you deeply trust this current AI system and you think that it’s super aligned and then you’ve given it enough affordances and resources that the next generation system that it designed is going to be strictly better in all the ways — and then that one’s going to design the next one and so forth.
Luisa Rodriguez
所以有一条信任链……有些人听完仍会说:「不,到那时我也不觉得我们会有信心说 AI 对齐了。」预测它不可解——哪怕有足够时间——的人会怎么说?
Yeah, so there’s this chain of trust… Some people, I think, would still hear this and say, “No, I don’t think that we’ll be confident by the end of that that the AIs will be aligned.” What would people who predict that it isn’t solvable — including even with enough time — what would they say about why it probably isn’t solvable?
Daniel Kokotajlo
我觉得那完全合理。所以我们试着把 Plan A 设计成:如果我们仍没拿到稳健解,我们可以继续延长。我们不是被迫交接给超级智能,不是被迫缩放到超级智能,不是被迫交接给 AI。情景里做那个选择,是因为他们解决了相关问题。如果没解决,他们可以继续推迟。我没跟够多这种人谈过,代表不了他们所有观点。我和机器智能研究所的人谈得不少,我觉得他们的观点是:事情会在更早阶段出错——在你还没到达那些真正(即便不稳健)在做好事的顶级专家级 AI 之前,人类决策者就会以某种方式把事情搞砸,批准了其实不对齐、只是看起来对齐的设计。你说你先到真正(即便不稳健)对齐的顶级专家级 AI,再让它们解决更深层的稳健性问题、设计新范式。他们大概会说你到不了第一步。
I think that’s entirely reasonable. And that’s why we tried to design Plan A so that — if we are in that situation where we still haven’t gotten a robust solution — we can just keep extending things. We’re not forced to hand off to superintelligence, or we’re not forced to scale to superintelligence, we’re not forced to hand off to AIs. It’s a choice that, in our scenario, gets made because they’ve solved the relevant problems. But if we hadn’t solved the relevant problems, then they could have just kept delaying. I don’t know. I don’t think I have talked to enough such people to be able to represent all of their views. I have talked to Machine Intelligence Research Institute people a fair amount and I think their view is that they expect things to go wrong at an earlier stage, where before you get to the top-human-expert-level AIs that are genuinely, if not robustly, trying to do the good stuff — before you get to that point — the human decision makers will have messed things up somehow and approved AI designs that are in fact not aligned, but seemingly aligned. Basically — because we’re saying you get to the point where you have these top-expert-level AIs that are genuinely, if not robustly, aligned and then they solve the more deeper challenging issues about robustness and design new paradigms — but I think that they would say you’re not going to get to the first step.
Luisa Rodriguez
你觉得只要有足够时间,我们会到。
And you think we will with enough time?
Daniel Kokotajlo
对,大概会——如果我们把 Plan A 做得很好。我综合一切的判断是:不,我们不会及时解决这些问题。所以我才这么担心。
Yes, probably — if we do Plan A really well. My all-things-considered view is that no, we are not going to solve these problems in time. And that’s why I’m so worried.
01:12:18Plan A 如何处理权力集中How Plan A addresses concentration of power
Luisa Rodriguez
换一个问题。默认轨迹上会出现权力集中:谁控制第一个超级智能,基本上就控制一切。Plan A 如何让权力集中更不容易发生?
Let’s turn to another problem. Concentration of power is a problem that comes up on the default trajectory: whoever controls the first superintelligence basically controls everything. How does Plan A make concentrations of power less likely?
Daniel Kokotajlo
高层次的说法是:如果我们要建造越来越强的 AI,最终越来越大比例的权力会来自控制这些 AI。极限情况下,AI 几乎运转整个经济、自主做军事,谁控制 AI 谁就控制一切。所以谈权力集中时,我们真正关心的是对 AI 的权力,因为对 AI 的权力最终会是大部分权力,甚至全部权力。我们认为如果有 AI 垄断——一支单一的巨型 AI 军队,其他 AI 相对都很弱——谁控制那支军队,也许是一小群人,也许是一个人。这是我们主要想避免的局面。意味着我们希望多家公司、多个国家都有大致相当的 AI 能力。单这一点其实还不够,你仍可能落到寡头:十几个 CEO 和三个总统凑在一起。
The high-level thing is that — if we are going to be building AIs that are ever more powerful — then eventually an increasing fraction of the power will come from controlling the AIs. If, in the limit, the AIs are running almost the entire economy and they’re autonomously doing the military, then whoever controls the AIs controls everything. So, to a first approximation, we’re really interested in power over the AIs when we’re talking about concentration of power, because power over the AIs will eventually be most of the power — or even all the power. We think it’s really bad if there’s an AI monopoly, if there’s a single giant army of AIs and all the other AIs are weak in comparison to it — whoever controls that giant army, maybe it’s a tiny group of people, maybe it’s one man. That’s the sort of situation we’re trying to avoid primarily. That means that we want there to be multiple companies spread out across multiple countries that all have roughly equivalent levels of AI capability. That by itself isn’t even enough really, because you still might end up in a situation where it’s kind of like an oligarchy, where there’s this group of a dozen CEOs and three presidents that get together.
还有透明。我们引入透明,是想比单纯摊开更进一步。我们不要垄断,但即便没有垄断,大量透明也有帮助,因为它让那些并不控制巨型 AI 军队的人,能够监督那些控制的人在用它们做什么。特别是,如果有我们在 Plan A 里主张的完全研究透明,当有人发表「Claude 偏向 Anthropic」的新研究结果,公众可以直接看 Claude 怎么被训练,自己判断 Anthropic 是不是故意塞进了这种偏向,还是训练中涌现的意外特征,或者我们根本不知道偏向怎么进去的、但肯定不是以任何方式故意插入的。大概会有很多灰色地带。透明让人们能看出他们在用 AI 做什么,防止秘密忠诚、防止插入隐藏偏向。这已经走了很长一段。
I think that also there’s these issues of transparency. We introduced the transparency to try to go further than simply spreading out. We don’t want it to be a monopoly, but we think that — even if you don’t have monopoly — it’s helpful to have lots of transparency because it gives everyone who doesn’t control a giant army of AIs the ability to oversee what the people who do have giant armies of AIs are doing with them. In particular, if we had the total research transparency that we are currently advocating for in Plan A, then when there’s a new research result by someone saying that Claude is biased towards Anthropic, people in the public could just look at the way that Claude was trained, and then they could judge for themselves whether Anthropic was deliberately putting in that bias, or whether it was an emergent, accidental feature of the training — or whether we just have no idea how that bias got in there, but it certainly wasn’t deliberately inserted in any way. There’ll probably be lots of grey-area cases. The transparency makes it possible for people to tell what they’re doing with the AIs and it prevents secret loyalties, it prevents the insertion of hidden biases, which already goes a long way.
更一般地说,它意味着 AI 必须是公司所说的那样。如果他们说这个 AI 有帮助、无害、诚实,那就不是一句你只能信他们的口号。你能看见整个训练过程,然后让第三方专家判断训练过程在多大程度上真的在强化这些特质、只强化这些特质,以及这些特质之间的权重。你可以就此做科学讨论。想象如果没有食品标签,不知道食物里有什么配料。那公司说这是健康食品,你只能信。能看见配料,要比只能听公司说「健康」容易判断得多。透明对政府也有帮助。如果公司被政府审计,甚至对政府完全透明,有助于监督公司,但问题会往后挪一点:政府呢?总统是不是在下秘密命令,让 AI 在宪政危机时必须对他忠诚,而且没人能知道?也许他是,谁知道。但如果我们能看见 AI 怎么被训练,会很好。所以我们要避免垄断,并对 AI 有透明。我们认为这两件事对降低权力集中走了很长一段。Plan A 做成这两件事。
More generally, it means that the AIs will have to be the way that the companies say that they are. If they say this AI is helpful, harmless, and honest, that’s not just a slogan that you have to take their word for. You can see the whole training process and then you can have third-party experts judge the extent to which the training process really is reinforcing those traits and only those traits — and the weightings between those traits and everything. You can just have a scientific discussion about it. Similarly, imagine if we didn’t have food labels and we didn’t know what ingredients were in food. Then you just have to take the company’s word for it when they say this is healthy food. It’s still not perfect, but it’s a lot easier to tell if it’s healthy food if you can see the ingredients that went into it, compared to if all you have to go on is the fact that the company said it was healthy. Notably this also helps with governments. If you had a situation where the company was audited by a government — or even fully transparent to a government — that would help with oversight of the company, but then it would sort of shift the problem back a little bit of: what about the government? Is the president issuing secret commands that the AIs have to be loyal to him in case of a constitutional crisis, and that no one can know about this? Maybe he is, for all we know. But it’d be nice if we could just see how the AIs are trained. So basically, we want to avoid monopoly and then have transparency into the AIs. We think that those two things go a long way towards reducing the concentrations of power. Those are like our main two things, and we think that Plan A accomplishes those things.
Luisa Rodriguez
不要垄断、要透明。相对今天的规范,这两件事都感觉非常激进。AI 公司现在在极度保密中运转,把训练方法、数据和算法当成最有价值的竞争优势。指望它们把这些全公布,现实吗?美国 AI 公司有多大可能不会掐死这种激进透明和技术扩散?
Not a monopoly and transparency. Both of those do seem really good for avoiding concentration of power. I guess they both feel very radical, relative to the norms we have today. AI companies currently operate in intense secrecy. They consider their training methods and data and algorithms to be kind of their most valuable competitive advantages. And they want to stay ahead. Is it realistic to expect them to publish all of that? How likely is it, do you think, that American AI companies don’t kill something like radical transparency and diffusion of the technology?
Daniel Kokotajlo
他们大概不会喜欢——原因就是你说的——但我们认为这对世界最好,所以才写下来。至于现实不现实:指望他们自愿做,不现实。但世界各国政府,特别是美国和中国政府,如果认定这符合本国最佳利益,可以强迫他们做。我们有一份补充材料,各自扔了一些概率。当然只是猜测。我觉得《AI 2040: Plan A》的作者们,对真正做成 Plan A 或类似的东西,大致在 5% 到 20% 之间。所以答案是:我们不认为这是最可能的结果,但它在可能的范围内。
Well, they’re probably not going to like it — for the reason that you mentioned — but we think it’s what would be best for the world, and so that’s why we’ve written it. As for whether it’s realistic, I think that it would be unrealistic to expect them to do this voluntarily. But I think that the governments of the world — in particular the government of the United States and the government of China — could make them do it, if it decided that it was in the best interests of those countries. We have, in one of our supplements, some quick numbers that we each threw out on our probabilities of the various things. Of course, these are just our guesses, they’re not proven. But I think the authors of AI 2040: Plan A range between something like 5% and 20% for the probability that they’ll actually do Plan A, or something like it. So I guess that’s your answer: we think it’s not the most likely outcome, but it is within the realm of possibility.
Luisa Rodriguez
如果完全激进透明在政治上不可能,有没有退路?
OK, is there a fallback if full radical transparency is politically impossible?
Daniel Kokotajlo
我们叫它完全研究透明。你也可以用更少的、中等程度的研究透明,好不好取决于有多强。可以有第三方审计师——最好是几家独立的——进来提问。理想情况下他们不只提问,还能核实验证答案,能看见相关的底层信息。这比什么都没有好得多,如果能拿到我会很高兴。但我们认为它没那么好,因为你把大量信任放在那些审计师身上:既要信他们不被公司和政府腐化,也要信他们在信息有限、不能和外界讨论所见之事时仍能有效干活。如果所有信息都透明,就可以有公共讨论——所有人愤怒地互相关推,在那片话语海里也会有好的讨论,非营利、学术部门、互相有动机找对方问题的对手公司里真正能干的专家互相挑,监管者能从中学习。对他们来说会更容易,因为不是全靠自己。透明也有利于执行任何协议。即便只是国内监管,也总担心公司会作弊、钻灰色地带和漏洞。透明越多,他们越难逃脱,因为越快会有人注意到并告诉监管者。尤其是国际:美中如果同意双方都做忠实思维链之类的事,怎么确保对方真的在做?这种程度的完全研究透明帮助很大。
We call it total research transparency. You could get away instead with less research transparency, or like medium levels of research transparency. And how good that would be depends on how strong it is. There’s a whole range of possibilities. I think that you could have some sort of system where there’s a third-party auditor — or maybe several different independent third-party auditors — that get to come in and ask questions. Ideally they don’t just get to ask questions, but they get to actually verify the answers to those questions, so they get to actually see the relevant low-level information. That’s a lot better than nothing. I’d be very happy if we got that. But I think that the reason why we think it’s not as good as it could be is that you’re putting a lot of trust in those auditors — both you’re trusting them to not be corrupt, and not be corrupted by the companies and by the government that might be trying to corrupt them. And you’re trusting them to do their jobs effectively, which is harder to do when they have limited information and when they’re not able to discuss what they’re seeing with outside parties. Whereas if all the information was just transparent, then there could just be a public conversation — everybody tweeting angrily about it to each other, and then in that giant sea of discourse there would also be some good discourse happening and actual very competent experts in various nonprofits, various academic departments, various rival companies that are motivated to find problems with each other, picking at each other, and the regulators would be able to learn from all of that. Another thing also is compliance. Transparency is good for making alignment progress and it’s good for preventing concentrations of power, but it’s also just good for enforcing any deal. If you’re going to be making a deal — even if you’re just domestically regulating — there’s always the concern that the companies are going to cheat on the regulations, or they’re going to find some grey areas and then really exploit those grey areas or loopholes. The more transparency you have, the less they’ll be able to get away with that sort of thing because the faster someone will notice and bring it to the attention of the regulators. Especially internationally: if the US and China agree on how we’re both going to do faithful chain of thought or whatever, how are they going to make sure that the other side is actually following through? It really helps a lot to have this level of total research transparency.
我不认为这些就够了。我认为这些是最能解决问题的干预,我不是说做完就不用做别的。但我认为它们是首先最该做对的事。对比一下:如果仍在竞赛、公司在保密中互赛,能活下来的公司不会那么多——或至少会有一段时期只有少数公司拥有超级智能大军。它们可能处于摧毁竞争对手的位置。如果都在一个国家,那个国家就处于摧毁其竞争对手的位置,而且不仅有能力这么做,还有紧迫理由:如果不做,优势最终会失去,别人会追上。然后也没有透明,顶层的人基本上可以把自己设置成独裁者,可以把自己特异的价值观塞进 AI 而不被人看出。默认局面太适合被滥用。我们提的东西把我们带出默认、进入好得多的世界,但没有完全解决所有问题。
I don’t think it’s enough. I think the things that I mentioned are the interventions that I think go the most towards solving the problem. I’m not claiming that they are sufficient and that once we do those things, we don’t need to do anything else. But I think that they’re the most important things to get right first. I think that, for example, by contrast, if you’re still in race conditions where these AI companies are racing each other in conditions of secrecy, then there’s not going to be that many companies that survive — or at least there’s going to be a period where there’s only a few companies that have these giant armies of superintelligences. And they’ll be potentially in a position to destroy their competitors. If they’re all in one country, then that country will be in a position to destroy its competitors, and it’ll not only be in a position to do so, but it’ll have pressing reason to do so — which is that if it doesn’t, eventually it’ll lose its advantage and the others will catch up. Then also more generally, there wouldn’t be transparency into what exactly they’re doing. So the people at the top could be basically setting themselves up to become dictators. And in general, the people at the top could be abusing their power and putting their own idiosyncratic values into their AIs in a way that’s not obvious to people. Yeah, it’s so ripe for abuse, the default thing, and I think that the stuff that we propose gets us out of that default into a much better world, but it doesn’t completely solve all the problem.
还有公民分红本身,以及买时间本身。如果你给算力和机器人设帽,让它每年只翻一倍,用收益付钱给人,那实际上会挪动一些权力。经济和金融权力会比否则更摊开——否则增长会快得多、更集中在少数公司。另一件是我们想要用于认知的 AI:人们能接触诚实回答问题的 AI 顾问,它们还很会预测、很会回答事情进展如何。我们认为这能大幅改善民主,因为人更难被宣传式政治运动摇摆,更容易看出自己正在被剥夺权力并采取行动。说到这个,我们也认为应当禁止超级说服,只要超级说服还在地平线上。Plan A 创造了事后可以再谈这些事的框架。开头不必把这一切都做对。有了这份基本协议、慢慢推进之后,可以再做后续的事。比如美中可以同意:不把 AI 训练得特别会说服,或限制 AI 被这样用:让它们拒绝协助政治广告这类任务。
One is the citizens’ dividend itself, and the buying time itself. If you cap the compute and robots so that it only doubles once a year and you use the proceeds to pay people, then that actually shifts some power around. It makes there be more substantial economic and financial power spread out more than it otherwise would be, if you didn’t do those things and you allowed growth to grow much faster and be more concentrated in a few companies. Another thing is that we want AI for epistemics. We want it to be the case that people have access to AI advisors that are being honest with them and answering their questions, and that are also really good at forecasting and really good at answering questions about how things are going. We think that could massively improve democracy effectively because it would be harder for people to be swayed by propagandistic political campaigns, and easier for people to tell when they’re being disempowered and then act to stop it. Speaking of which, we also think there should be bans on superpersuasion, insofar as superpersuasion is looming on the horizon. We talk a little bit about what that might look like as well, and Plan A creates the framework by which these types of things could be negotiated afterwards. You don’t have to get all this right at the very beginning. Once you’ve got this basic deal in place, and once you’re sort of proceeding slowly, then you can make subsequent things. Like the US and China can agree we’re not going to train our AIs to be really good at persuasion, or we’re going to limit the way in which the AIs can be used for that: we’re going to have them refuse to do tasks like assisting with political ads.
我们有数字。我们说的是:在做成 Plan A 的条件下,总灾难大约 15%。即便在 Plan A 里也大约 15%。不同人有不同答案,但差不多这样。为什么还做?我们担心任何国际协议都可能破裂。如果你开始做 Plan S,然后新总统上任完全改道,你就完了。Plan A 相对 Plan S 的优势是:因为你在以相对快的速度向前推进解决问题,整件事不必永远持续。我们不是说权力会比今天更不集中。我们是说它会比我们提过的任何其他计划更不集中。我们希望它比今天更不集中。也许有更好版本的 Plan A 能做到。如果按我们描绘的那样走得好,它确实会很好。担心的是有各种走错的方式。Plan A 我们觉得是最不坏的计划,但它仍会超级可怕,有一堆走错的方式。即便按我们自己的估计,也像在对所有人玩俄罗斯轮盘。如果你能做比这更谨慎的事,太好了。
Yeah, we have numbers. I think we say something like 15% chance of total catastrophe, conditional on doing Plan A. Yeah, even in Plan A it’s like 15%. But different people have different answers. But something like that. Why are we doing this? Well, we’re worried that any international deal might break down, and so if you started to do Plan S, and then a new president gets elected and does something completely different, then now you’re cooked. The advantage of Plan A over Plan S is that — because you’re making forward progress towards solving the problems at a relatively fast pace — the whole thing doesn’t need to last forever. We’re not saying that power will be less concentrated than it is today. We’re saying that it’ll be less concentrated than it is in any of the other plans that we’ve proposed. We would like it to be less concentrated than it is today. Maybe there’s like an even better version of Plan A that would achieve that. I think that if things go well with the way that we depict it happening, it does go really well. The concern is that there’s various ways it can go wrong. Plan A, we think it’s the least bad plan, but it still is going to be super scary and there’s a bunch of ways to go wrong. Even by our own estimates, it’s like playing Russian roulette with everyone. So yeah, if you can do something even more cautious than that, great.
Luisa Rodriguez
20 年减速你会感觉更好,还是会开始太担心协议破裂?10 年是不是被有意选成最优?
Would you feel better about a 20-year slowdown, or do you start to worry too much about the deal breaking down? Was 10 years quite deliberately chosen as the optimal?
Daniel Kokotajlo
是也不是。我们其实有一些建模,我忘了最优是多少。我不觉得和 10 年差很多——10 是个好看的整数,离建模建议的最优量也没那么远。实际该多久取决于情况。显然我们该希望的是成功摸着石头过河,其中一个变量是走多慢、在人类水平停多久、精确停在哪一级。这些变量最好到时候拿着当时收集到的全部信息再定。显然不该在 2037 年还死守 2026 年写的计划。得根据新信息调整。比如对齐看起来不太好,就该停更久。如果整体被扰动、太混乱、人人都很怕,也该停更久。如果更长的暂停看起来完全可行、完全稳定,不像会在下一届政府时破裂,那也是走更长的理由。反过来,如果局面更糟、眼看要破裂……也可以想象变量往另一边设:对齐看起来很好,AI 看起来超级对齐,有很多独立证据线支持它们对齐,下一届政府已经放话不想暂停,那完全暂停可能实际上不如走快得多。
Yes and no. We actually do have some modelling of this, and I forget what the optimal was. I don’t think it was that different from 10 years — 10 is a nice round number and it’s not that far off from what our modelling would suggest is the optimal amount. I think that how long it should actually be just depends. Obviously what we should hope to do is muddle through successfully, and then one of the variables is how slow do we go? How much do we pause at human level? What level do we pause exactly? These sorts of variables will be best figured out at the time, with all the information that’s been gathered at the time. Obviously we shouldn’t just stick to the plan that was written in 2026, when the year is 2037. We’ll have to adjust as we go, based on new information coming in. For example, if the alignment stuff is not looking very good, then we’d want to pause longer. If in general things are being disrupted and too chaotic and everyone’s really scared, we should pause longer. If it’s looking like a longer pause would totally work and be totally stable and it doesn’t look like it’s going to break down when the next administration is elected, then that would also be a reason to go longer. Then by contrast, if instead we were in a worse situation, where it seemed like things were just about to break down… You can imagine there’s variables being set in the other direction, where alignment looks really good, AIs look super aligned, and we have all these independent lines of evidence supporting that they’re aligned. Also the next administration has already signalled that they don’t want to pause, then having a complete pause might just not actually be as good as going much faster.
我们并没有放弃算力领先。Plan A 里我们放弃的是算法领先。因为完全研究透明,中国和其他所有人都能看到造这些 AI 的配方。如前所述,这在很多方面有很多好处,代价是对手会追上一点。这是对中国的严肃让步。也正因为此,我觉得中国接受这类协议是说得通的,因为它确实是让步。如果你不喜欢,可以改协议,换别的回报。我们情景里,美国作为协议的一部分,锁定了对中国的一点算力优势。现在美国算力比中国多,协议里他们基本上会做一些事,确保美国继续有算力优势。这是反方向的一点让步。讨价还价时你可以加减东西,让它对一方更有利或对另一方更有利,然后希望找到双方都能接受的东西。我们对精确落到哪里没有强烈意见。也许这又是那种灰色地带:我们认为该有协议,大致该长这样、有这些原则,但对该做哪些交易、让步、胡萝卜和大棒,没有强烈意见。如果你觉得我们提的这个版本太让步,可以提一个更少让步的版本,也许也能成。
We’re not giving up our compute lead. What we’re giving up is our algorithms lead, in Plan A. So in Plan A, because of the total research transparency, China and everybody else gets to see the recipes for making the AIs. As previously mentioned, I think this has a lot of benefits in a lot of ways, but it does have the cost of now our adversaries get to catch up a little bit. That is a serious concession to China. That’s part of why I think that it’s plausible that China would want to accept a deal like this because it’s just actually a concession to them. Insofar as you don’t like that, well, you can modify the deal to get something else in return, for example. In our scenario the US, as part of the deal, locks in a bit of a compute advantage over China. So currently the US has more compute than China, and then as part of the deal they basically do things to ensure that the US will continue to have a compute advantage over China. That’s an example of a bit of a concession going the other way. You could imagine doing it even more, so basically in the horse trading that happens before a deal you can just add and subtract things from the deal to make it more fair and to make it something that’s more beneficial to one side or more beneficial to the other side. Then hopefully you can find something that both sides are OK with, and then it happens. We don’t have a strong opinion about exactly where that should end up. We think that there should be a deal. We think it should look something roughly like this with these principles, but we don’t have a strong opinion about the horse trading that should go into it, and the concessions, the carrots, and the sticks flying back and forth. If you think that this particular version that we proposed is too conciliatory, then you can propose a less conciliatory version, and maybe that’ll work too.
01:41:28大国冲突、失业与滥用How Plan A addresses great power conflict, unemployment, and misuse of AIs
Luisa Rodriguez
谈谈另外三个问题。我觉得 Plan A 怎么解决它们更直白。简要说,Plan A 怎么解决大国冲突、失业和 AI 滥用?
Let’s talk about the three other problems. I think it’s more straightforward how Plan A solves them. So, kind of briefly, how does Plan A solve great power conflict, unemployment, and misuse of AIs?
Daniel Kokotajlo
因为 Plan A 造成其他国家的其他公司能追到前沿的局面,我预期它会在防止修昔底德陷阱上走很长一段——那种很多国家因为即将被剥夺权力而恐慌、可能为此冒险开战的局面。相对那些默认,Plan A 里「他们被剥夺权力」的成分会少得多。它也是一份字面上的国际协议。如果协议成功、他们真的做了,他们就有理由继续做下去而不是互打。协议每维持一年,这些理由就越强:因为如果破裂、摧毁所有算力、回到 2029 年的位置,对他们经济的回撤会比 2029 年当时还大。我不认为它解决大国冲突,但我认为它大体上防止了:在建造超级智能期间,预期大国冲突会异常高的那些具体理由。
So because Plan A creates a situation where other companies from other countries can catch up to the frontier, I expect it to go a long way towards preventing this Thucydides trap where a bunch of countries freak out about their imminent disempowerment and then possibly risk war over it — because in Plan A, compared to those defaults, it’s going to be much less of a them being disempowered type situation. Also it’s a literal international deal. If the deal is successful and they actually do it, then now they have reasons to proceed with it instead of fighting each other. And every year that the deal is maintained, those reasons get stronger because if things break down and they destroy all the compute and go back to where things were in 2029, that would be setting back their economies even more compared to in 2029. I think it doesn’t solve great power conflict, but I think it mostly prevents the specific reasons to expect great power conflict to be exceptionally high during the period of the building of superintelligence.
下一个,工作:公民分红,以及用于认知的 AI。那些强化民主、帮人见多识广、帮人监督政治领导人的东西。比如保护隐私的审计,以及完全研究透明。整包东西,再加上即便没工作也直接给人钱的公民分红。那些是我们对失业问题的解决方案包。保住人们的经济权力和政治权力,理想情况下还加强它。第五个:前面几个回答可能已经不够过瘾,第五个也许更不过瘾——这也是它在问题清单上排第五而不是更前的原因。这是恐怖分子拿 AI 干坏事的问题。我们的答案基本上是防御加速主义。这不是我们发明的词。意思是大力投资,让世界对恐怖分子和他们的 AI 更硬,于是即便有带 AI 的恐怖分子,也还行。不完全是那样。我们也认为该有拒答。完全研究透明近似意味着一切开源,但我们没说该开源权重。所以恐怖分子拿不到真正的模型,只拿到访问模型的能力。意味着他们没法撤销拒答训练。所以是拒答加上硬化,我们希望这足以让生物恐怖不至于太糟。
OK, the next one, the jobs: citizens’ dividend and the AI for epistemics. The stuff that strengthens democracy and helps people to be well informed and helps people oversee their political leaders. Things like the privacy-preserving auditing and the total research transparency. That whole package of things combined with the citizens’ dividend, which just directly gives people money even if they don’t have jobs. Those are, I think, our package of solutions to the job-loss problem. Preserving people’s economic power and their political power and strengthening it, ideally. Then the fifth one: as unsatisfying as my answers to the previous ones might be, my answer to the fifth one is perhaps even less satisfying — which comes from the fact that it’s number five on our list of problems, instead of higher up on the list. So this is the problem of terrorists doing bad stuff with AI. Our answer is basically defensive accelerationism. This is not a term we invented. Basically the idea is to invest really hard in hardening the world against the terrorists and their AIs, so that even though there’s terrorists with AIs, it’s OK. It’s not entirely that. We also think that there should be refusals. We also think that — while the total research transparency basically means, to a first approximation, everything is open sourced — we don’t say you should open source the weights. So the terrorists don’t get the actual models, they just get the ability to access the models. That means they can’t undo the refusal training, for example. So it’s a combination of the refusals and the hardening that we hope will be enough to prevent the bioterror from being too bad.
Luisa Rodriguez
Plan A 的哪些部分对好结果最关键,哪些更边缘?
Which parts of Plan A seem most essential to good outcomes, and which parts are more peripheral?
Daniel Kokotajlo
我觉得大致就是我们列那些原则的顺序。我认为最重要的一件事是:你没有在做疯狂的智能爆炸,而是更慢、更谨慎地推进。第二重要的是完全研究透明,或至少对那些 AI 如何被训练、如何被开发有大量透明,好让科学共同体和公众能监督。这两件事本身,我认为也会帮助权力集中——原因前面说过。我认为它们会让局面更少垄断、更多是不同提供方竞争的生态。
I think it’s roughly in the order that we listed those principles. I think that the single most important thing is that you’re not doing a crazy intelligence explosion and you’re instead proceeding more slowly and cautiously. Then the second most important thing is that you have total research transparency, or at least a lot of transparency into how those AIs are being trained and how they’re being developed, so that the scientific community and the public can have oversight into all of that. Those two things by themselves, I think, will also help with the concentration of power — for the reasons previously mentioned. I think they’ll help make it the case that there’s less of a monopoly and more of a competing ecosystem of different providers.
01:45:56美中如何谈成减速How the US and China could agree on a slowdown
Luisa Rodriguez
我想谈美中为什么会同意这种协议。只说双方都认识到灾难风险,似乎不太够。先从美国这边:对美国想要这笔交易,最强的理由是什么?假设你是非常看重美国国家利益、不太相信 AI 构成生存风险的国家安全人士。
I want to talk about why the US and China would agree to this kind of a deal. It doesn’t seem totally satisfying to say both sides recognise catastrophic risks. Let’s start with the American side of things. What is the strongest case for the US wanting this deal? If you’re, say, a national security professional who weighs US national interests really heavily and isn’t as convinced AI poses an existential risk.
Daniel Kokotajlo
先说可能不太过瘾、但我觉得仍重要的一点:我认为 AI 构成灾难风险,因为我们现在控制得不太好,未来它们递归自我改进时我们可能也控制不住。所以每个活着的人都有很强的理由去关心做类似 Plan A 的事。很多人现在还不承认,但承认的人越来越多。我希望在太晚之前,有足够的人承认,好让这类事发生。除此之外,我认为 Plan A 对防止极端权力集中非常好。比如你是国家安全人士,想让美国打败中国,原因之一是美国是民主国家、中国不是,那你也该关心美国保持民主。你该有点担心科技公司在积聚的权力。你该有点担心总统和 CEO 为谁指挥超级智能大军而斗争,尘埃落定后,谁赢了谁就可能处于成为美国独裁者的位置。
First of all, I know it might not be satisfying, but I think it’s important to say anyway: I do think that AI poses a catastrophic risk because we can’t control them very well right now and we might not be able to control them in the future as they recursively self-improve. So I think that is a reason for every single human being to care quite a lot about doing something like Plan A. It’s true that a lot of people don’t recognise that right now, but an increasingly large amount of people do recognise that. Having said that, I actually think that Plan A is really good for preventing extreme concentrations of power. For example, if you’re a national security professional, you want the US to beat China. One of the reasons why you want the US to beat China is because the US is a democracy and China isn’t, so you should also be interested in making sure that the US stays a democracy. And you should be a bit concerned about the amount of power that the tech companies are accumulating. You should be a bit concerned about a situation where maybe the president and the CEO have a power struggle over who gets to command the army of the superintelligences, and then when the dust settles, whoever manages to come out on top of that power struggle will be in a position of potentially being dictator of America.
基本上每个人都该担心这个,连渴望权力的人——CEO 等等。如果你觉得自己有机会用 AGI 当独裁者,你至少也该有点担心,因为当上的未必是你。即便你是总统,也该有点担心这些 CEO,担心你出事、别人当独裁者,或你被赶下台。你并没有稳赢独裁的票。这其实是一场谁会胜出完全未知的权力斗争。所以某种人人都能拿到大部分想要的东西的协议,仍符合你的利益,而不是这场争夺完全主宰的疯狂斗争。对 CEO 或总统,这是最硬的情况:你也许确实该有点心动去冲向超级智能、自己控制它,好成为全球独裁者。但那是最硬的情况。其他所有人都该对此感到恐惧。如果你只是普通美国公民、AI 公司普通员工、美军里的人,你该担心美国不再是民主国家,担心这些独裁可能。
I think basically everybody should be concerned about this, even the power-hungry people — the CEOs, et cetera. If you’re a person who thinks that you might stand a chance of becoming dictator using AGI, even you should be at least a little bit concerned about this because maybe you’re not going to be the one who ends up being dictator. Even if you’re the president, you should be a bit concerned about these CEOs. You should be a bit concerned about something happening to you and then someone else becoming dictator, or you get ousted somehow. It’s not like you have a guaranteed shot at becoming dictator. It’s actually quite as much of a power struggle where who knows who’s going to come out on top? So it’s still kind of in your interest to have some sort of deal where everybody can get most of what they want, instead of this crazy power struggle for total dominance. Then the other thing to mention besides that is that’s sort of like the hardest case. If you’re the CEO of the AI company or the president, then genuinely maybe you should be somewhat tempted to race to superintelligence and then try to control it yourself so that you can become a global dictator. But that’s the hardest case. Everybody else should be terrified about this. If you’re just an ordinary American citizen, if you’re an ordinary employee at one of the AI companies, if you are someone who works in the military in the US, then you should be worried about the US not being a democracy anymore. You should be worried about these dictatorship possibilities.
如果你在美国之外——英国、印度、俄罗斯、中国——你该对美国在当前竞赛条件下拿到超级智能感到恐惧,因为那意味着别人没有超级智能,或在他们做成时别人的 AI 远没有那么好。即便美国不变成独裁、不知怎么还能分享权力,你也该担心自己的国家:美国公司拿走所有工作,美国军队能把你的军队打得满地找牙。我认为几乎对所有人,单凭权力集中这一条,这就有激励相容。即便你完全不把失控当回事。下一条当然是第三次世界大战。即便你是权力会集中到的那个人——也许你觉得自己是总统,能打赢其他人、站到顶上,也完全不担心失控——你至少该担心第三次世界大战,担心被核弹或暗杀毁掉。其他所有人对你将要做的事如此恐惧,这件事本身就该让你在做之前停一停。
Then if you’re outside the US — if you’re in the UK, or if you’re in India, if you’re in Russia, if you’re in China — you should be terrified about what’s going to happen if the US gets the superintelligence in conditions like the current race conditions, because that means that nobody else would have superintelligence or nobody else will have AI nearly as good at the time that they do it. Even if the US doesn’t become a dictatorship and somehow manages to share power, you should be worried about what’s going to happen to your country vis-à-vis US companies taking all the jobs, US military being able to wipe the floor with your military. Basically I think that it’s kind of incentive-compatible for everyone, or almost everyone, to do this for power concentration reasons alone. Even if you don’t take the loss of control stuff seriously at all. The next reason, of course, is World War III. Even if you’re the person in whom power will concentrate — maybe you think you’re the president and you can just win the fights against the other people and end up on top, and you’re not at all worried about loss of control — you should at least be worried about World War III and being destroyed in a nuke or an assassination as a result of this. So the fact that everybody else is so terrified about what you’re going to do should give you pause before you do it.
Luisa Rodriguez
美中关系和美苏谈成防扩散条约时,有多像、有多不像?
How similar or different is the relationship between the US and China and the USSR when they agreed to a nonproliferation treaty?
Daniel Kokotajlo
有类比,也有不同。这是领先的一方约束自己以换协议。一个不同是:核武器对拥有它的那个权力的危险,远小于 AI 对拥有它的那个权力的危险。理论上核弹可能出事在你头上炸开,但那极不可能。而我们控制 AI 的能力,远远远远差于控制自己的核武器。AI 反过来咬我们是极其真实的可能。事实上我会说,在当前条件下,多半会发生。这是核弹和 AI 之间极端的不同。权力集中也类似。并不真正担心总统能用核武库当美国独裁者。他要怎么做?威胁炸城市,不投他就不停?核弹很清楚是用来打敌国的武器,对内部政治斗争不太有效。相反,超级智能对一切都极其有效——包括内部政治斗争。美国不再是民主国家,是非常真实的可能,所以美国很多人该非常有兴趣做这类协议。这又和核弹不同。
There’s some analogies, there’s some disanalogies. It’s a case of the power that’s in a lead sort of restraining itself in order to get some sort of deal. I think a disanalogy is that the nukes are much less dangerous to the power that has them than AI will be to the power that has them. Think about nukes, theoretically there could be an accident and your nukes could start exploding on you. But that’s extremely unlikely. But actually though, our ability to control AIs is vastly, vastly worse than our ability to control our own nuclear weapons. There is an extremely real possibility that our AIs will turn on us. In fact, I would say it’s more likely than not under current conditions. That’s an extreme disanalogy between the nukes case and the AI case. Similarly with the concentration of power stuff. There isn’t really a serious concern that the president can use the nuclear arsenal to become dictator of the United States. What are you even talking about? How would he do that? He would start threatening to nuke cities or something if they didn’t vote for him? Nukes are very clearly a weapon that you use against enemy nations. They’re not very effective for internal political struggles. By contrast, superintelligence is extremely effective at everything — including internal political struggles. There’s a very real chance that the US would no longer be a democracy anymore, and so that’s a reason that lots of people in the US should be very interested in having this sort of deal. Again, that’s different from the nukes case.
另一个类比更像二战期间美苏之间的会议和协调。不是签一张纸、写几条规则、各自回去执行再互相核查。它连续得多。更像:「我们一起打赢这场战争,参谋会不断互相通气,谈谁做什么、谁何时打哪个国家,你帮我们这个我们就给你那些物资。」尽管在那之前美苏基本上是敌人。苏联基本上曾是纳粹德国的盟友,还打过波兰、芬兰这些美国的朋友。我们从敌人变成盟友,有密集的持续协调。不是完全信任。他们在曼哈顿计划上刺探,我们试图不让他们知道。我提这个,是因为我觉得这既是对待所有这些 AI 事情该有的态度,也更像 Plan A 实际会是的样子。不会是聚在一起签大条约然后回家。更像中国政府里几百人、美国政府里几百人不断来回打电话,基本上在一起规划这场「战争」、一起打这场「战争」。你可以叫它对 AI 的战争,但更像为我们自己的未来而战:我们怎么处理这种新的人造心智?它是一种新实体,一开始比我们弱,最后比我们强。
I think another analogy I want to bring up is something more like the conferences and coordination that happened between the US and the USSR during World War II. It wasn’t like a specific deal exactly where they came together and then signed some piece of paper that had some rules, and then they went away and tried to implement those rules and then maybe verify that each other was complying with the rules. It was much more continuous than that. It was more like, “Together we’re going to win this war and our staff will be constantly in touch with each other, talking about all the details of who’s going to do what and who’s going to invade which country and when, and we’ll send you these materials if you do this other thing for us.” This happened even though the United States and the USSR were basically enemies up until that point. The USSR had basically been an ally of Nazi Germany and had attacked various US friends, like Poland and Finland. We basically went from being enemies to being allies during World War II, and we had this intense amount of constant coordination. It wasn’t like we trusted them completely. They were spying on the Manhattan Project, and we were trying to stop them from finding out about it. I bring this up as an analogy because I feel like this is both the appropriate attitude to take towards all this AI stuff, and also more like what Plan A would actually look like in practice. It wouldn’t look like they come together, they sign a big treaty, and then they go home. It’d be more like there are hundreds of people in China, in the Chinese government, and hundreds of people in the US government who are constantly talking to each other and calling each other back and forth and who are sort of basically planning the war together, so to speak, and prosecuting the war together. I guess you could call it the war on AI, but it’s more like the war for our own future. It’s how are we going to handle this creation of a new artificial mind? And it’s a new type of entity that’s going to start out weaker than us, but end up stronger than us.
我们让它从美中双边开始。但它不能一直那样。很多其他国家也会造 AI。芯片供应链很大一部分在其他国家。所以我们谈它时,某种意义上是双边,但他们从一开始就在咨询其他国家、从一开始就在拿其他国家的买账。过一两年他们基本上把一大堆国家卷进来,到最后叫「财团」。基本上是多数主要国家,不一定都以同一程度参与。我们对谈判精确怎么走、各国最后有多少权力挥了挥手。但我们认为需要发生的结果是:所有有重要 AI 项目、在 AI 供应链里有重要部分的国家,一起工作,能通过透明看见发生了什么,并能核查合规。随着更多国家和公司追到前沿,它们大概也会以某种方式被卷进来。
We have it start out as a bilateral US–China thing. But it can’t really stay that way. There’s lots of other countries that will also be building AIs. Large parts of the chip supply chain are in other countries. That’s why we talk about it like in some sense it’s a bilateral thing, but it’s also like they’re consulting other countries from the start, and they’re getting buy-in from other countries from the start. Over a year or two they basically get a whole bunch of countries involved, so by the end it’s called The Consortium. It’s basically most major countries and they’re not necessarily all involved at the same level. We sort of handwave over exactly how the negotiations go and exactly how much power the different countries end up with. But the result that we think needs to happen is that basically all the countries that have significant AI programmes and significant parts of the AI supply chain are working together and able to see, via the transparency, what’s going on, and then able to verify compliance with it. Then as things progress and more countries and companies catch up to the frontier, we think that probably they would end up getting roped in too, one way or another.
Luisa Rodriguez
我的感觉是人们仍有强烈直觉:这种协议不现实。你觉得他们漏了什么?
My sense is that people just still have a strong intuition that a deal like this is not realistic. What do you think they’re missing?
Daniel Kokotajlo
首先,我们从未声称这是默认会发生的。他们正确地注意到这有点不太可能。我们其实承认——这不是我们认为事情会自然走的方式。这不是我们对将会发生之事的预测,而是建议。但我们也认为它足够可能,值得认真对待。人们在 Overton 窗口会往哪移、移多快上错得很厉害。我预期未来会被 AI 诱发重大转移。想想 Mythos 那些事,想想特朗普政府从基本上说一切 AI 监管都坏、许可制度是拜登政府梦出来的魔鬼,到直接下令出口管制、因担心 Claude 可能被越狱而停止部署。他们非常迅速地做了非常大的转向。我觉得这其实是好消息,说明政府能醒来、能灵活、能在决定要做的时候做完全不同的事。所以到这个时点,我们基本上该当作所有选项都在桌上,然后主张我们认为最好的行动。
First of all, we have never claimed that this is what’s going to happen by default. They are correctly noticing that this is somewhat unlikely. We just actually admit that — this is not how we think things will go naturally. This is not our prediction of what will happen, instead it’s our recommendation. However, we also think that it’s likely enough to be taken seriously. One thing I would say is that people have been very wrong about where the Overton window shifts and how fast it shifts. I expect there to be major shifts in the future induced by AI. Consider the Mythos stuff, and consider the Trump administration went from basically saying that AI regulation of all sorts was bad and that a licensing regime was the devil dreamed up by the Biden administration, to just issuing an order to export control and stop the deployment of Claude out of concerns that it could be jailbroken. And they did that very big shift very rapidly. I think that’s actually encouraging news that they did that because it just goes to show that the government can wake up and then be nimble and then do something completely different, if it decides that’s what it wants to do. So I think that we should basically, at this point, act as if all options are on the table and then we should just advocate for the actions that we think are best.
我们现在大约做了 100 场兵棋推演。经常发生的是:国家转向全球 AI 关停,但做得很晚,是在超级智能 AI 已经失控之类的之后。需要的警告射击通常非常极端。在我们的推演里,发生时通常已经太晚,但要点是:美中握手说「我们要把数据中心拔掉直到搞清楚发生了什么」这种极其激进的事,在推演里其实相当常见。不是大多数时候,但在我们的推演里发生过很多次。往往到发生时已经太晚。历史上也是:美苏成为盟友,直到希特勒打苏联之前都极不在桌上——然后突然就在了。
We’ve done about 100 war games right now with various people. A thing that often happens in the war games is that the nations do a pivot towards a global AI shutdown, but they do it late, they do it after a superintelligent AI has gone rogue or something. It’s usually too late when it happens in our war games, but the point is that it’s just actually a quite common occurrence in our war games for there to be this extremely radical US and China shaking hands on, “We’re going to unplug our data centres or something until we figure out what’s going on.” It doesn’t happen most times, but it’s happened a whole bunch of times across our war games. It’s just oftentimes by the time it’s happening, it’s too late. Historically too, the USA and the USSR becoming allies, that was extremely not on the table until Hitler invaded the USSR — and then all of a sudden it was.
Luisa Rodriguez
中国领导层想要这笔交易,最强的理由是什么?美国的芯片出口管制让合作更可能还是更不可能?哪一边更不太可能最终想要协议?
What is the strongest case for Chinese leadership wanting this deal? The chip export controls that the US has, do you have a sense of whether they’ve made cooperation more or less likely? Which side do you think is less likely to end up wanting a deal?
Daniel Kokotajlo
我觉得现在中国领导层大概不太把失控当回事,大概觉得时间在他们这边,长期中国会在 AI 和其他领域——军事、经济——占上风。只要他们继续信这两点,大概就不会想做协议,因为他们觉得无协议局面偏向他们。但他们有可能开始把失控风险当真。谁知道会不会、何时,但也许他们对 AI 了解得够多、看到够多像 Hugging Face 这样的例子,就会开始担心。其次,即便那没发生,他们大概到某个时点会意识到默认追不上,美国的算力优势会让美国至少在可预见的未来保持领先,而他们不能按几十年的尺度规划,因为没那么多时间:超级智能未来几年就会来,他们真的不想处于美国有超级智能而他们没有的局面,即便只有六个月或一年直到他们追上。这基本上是中国可能想做这类协议的两个理由:一是他们可能真的理解风险;二是即便不理解,他们可能意识到这些东西会极其强大,而他们不在赢的轨道上。
I think right now Chinese leadership probably doesn’t take loss of control very seriously, and they probably think that time is on their side and that, in the long run, China will prevail in AI and in other domains — militarily, economically. Insofar as they continue believing both of those things, then I think that they are probably not going to want to make a deal because the no-deal situation favours them, they think. However, I think that it’s possible that they will come to take the loss of control risks seriously. Who knows if and when, but perhaps they’ll learn enough about AI and they’ll see enough examples like the Hugging Face incident that they’ll start to be worried about this. Then secondly, even if that doesn’t happen, at some point they will probably realise that they’re not going to catch up by default, and that the compute advantage that the United States has is going to keep the United States ahead — at least by default — for the foreseeable future, and that they can’t plan on timescales of decades because they just don’t have that much time: superintelligence is coming in the next few years and they really don’t want to be in a situation where the US has superintelligence and they don’t, even if it’s only for six months or only for a year or whatever until they catch up. Those are basically the two reasons why China might want to do a deal like this. One is they might actually understand the risks. Then two, even if they don’t, they might realise that this stuff is going to be incredibly powerful and that they’re not on track to win.
出口管制大概让合作更不可能,不幸。有人会说是因为让中国人对 AI 协议这件事变味了,感觉美国在敌对、想坑他们。那也许是真的,但你也可以更现实主义:我们反正某种程度上是对手。在那种更现实主义的立场上,嘴上的话不那么重要,反正大家也不会互相喜欢,重要的是硬谈判力。但即便在那个视角下——这是我主要想说的——我觉得芯片走私对做协议是坏事,因为它可能导致一种局面:连中国自己都不知道芯片在哪。想象双方都真的想做协议。能毁掉它的一件事是:中国不知道自己一大堆芯片在哪,于是没法向美国证明他们想做协议、在善意行事,因为美国会说「我们核算不了这些芯片」,中国说「我们发誓我们也核算不了,谁知道在哪,但大概没事,我们当然不知道」,美国说「得了吧,你们大概藏在某个秘密项目里」。于是即便双方都想要协议,协议也可能因为走私做不成。我们想要的局面是:如果双方都想要协议,他们能向对方证明自己在履约。至少中国政府知道所有中国芯片在哪,美国政府知道所有美国芯片在哪——因为那样如果双方都想要协议,他们可以互相展示芯片,然后核查。好笑的是:就做协议而言,走私这件事上,重要的其实不是美国知道芯片在哪,而是中国知道芯片在哪。如果我做主,出口管制我会要么废掉要么执行。基本上不要有你执行得很差的出口管制——如果执行不好,就该去掉。哪一边更不太可能想要协议?我没有强烈意见。我大概会说美国。美国更可能把失控当真,但因为美国领先,他们更可能觉得局面还行、该继续走。而中国大概最终会意识到自己不领先、追不上,然后会想要协议。但也不清楚。中国也可能一直觉得自己能追上。
They’ve probably made cooperation less likely, unfortunately. Some people would say they made cooperation less likely by souring the Chinese on the idea of AI deals, because it feels like the US is being adversarial towards them and trying to screw them over. That might be true, but you could take a more realist position that we’re sort of adversaries anyway. Maybe on the more realist position that doesn’t matter so much because talk is cheap and people aren’t going to like each other anyway, and so what matters is the hard negotiation power. But then even on that perspective — and here’s the main thing I would say — I think chip smuggling is bad for making deals because it leads to the possibility of a situation where not even China knows where their chips are. Imagine that you manage to get to a point where both sides actually want to make a deal. A thing that could ruin that is if, for example, China doesn’t know where a bunch of their chips are — so they just can’t prove to the US that they want to make a deal and that they’re acting in good faith because the US is like, “Well, we can’t account for all of these chips.” And China’s like, “Yeah, we swear we can’t account for them either. Who knows where they are? But it’s probably fine. We certainly don’t know.” And then the US is like, “Yeah right, you’ve probably got them squirrelled away somewhere in a secret project.” So that could be a situation where — even though both sides want a deal — the deal doesn’t happen because of all the smuggling. Instead we want to be in a situation where if both sides want a deal, then they can prove to each other that they’re complying with the deal. You want to be in a situation where the Chinese government at least knows where all the Chinese chips are, and the US government knows where all the US chips are — because then if they both want a deal, they can just show each other the chips and then they can verify. It’s kind of funny — the smuggling, for purposes of making a deal — it’s not actually that important that the US know where the chips are, it’s important that China knows where the chips are. I don’t have a strong opinion about this, but I think roughly speaking I would either repeal them or enforce them. Basically don’t have export controls that you aren’t very well enforcing — and if you’re not enforcing them well, you should just get rid of them. Which side is less likely to end up wanting a deal? I don’t have a strong opinion. I think I would probably say the US. I think that the US is more likely to take the loss of control stuff seriously, but because the US is in the lead they will be more likely to think the situation is fine, we should keep going. Whereas China will probably eventually realise that they’re not in the lead and that they can’t catch up, and then they will want a deal. But it’s unclear. It’s possible that China will continue thinking that they can catch up well into the future.
Richard Ngo 的批评我其实有共鸣,某种程度上希望我们在这个情景里做得稍有不同。我认同的那个版本是:我们的东西太 DC 脑了。太像:「显然我们没法监管 AI,除非别人做同类监管。所以我们需要和中国做这笔大协议,显然我们不信任中国他们也不信任我们,所以协议里要有核查。于是我们画出这份漂亮的、带核查的对华协议,好让你们不必互相信任。」可也许 Richard 的点是:这种框架让步太多。它让步了我们互不信任。它让步了除非他们也做,否则我们不会想监管。而事实上,已经有很大一部分美国公众想相当重地监管这些东西——不管其他国家做什么。这部分批评让我共鸣,让我怀疑我们是否该写成:第一步,美国在国内监管 AI;第二步,看中国在做什么——如果他们反而鲁莽冲在前面,再找他们说「我们需要一份协议,因为我们不想你们那样做」,然后才是 Plan A。那也许既更现实,也更是我们真正会建议的,因为早点开始好的国内监管是好事,而不是等协议。
Yeah, I actually am sympathetic to Richard’s critique, and I kind of wish we had done things a little bit differently in this scenario. The version of Richard’s critique that I am sympathetic to — and that I basically agree with — is that our thing is too DC-brained. It’s too, like: “Obviously we can’t regulate AI until we get other people to do the same type of regulation. So we need to have this huge deal with China, and obviously we don’t trust China and they don’t trust us, so we need to have verification as part of the deal. And we’ve therefore sketched out this big, beautiful deal that you can make with China that includes verification so that you don’t have to trust each other.” But perhaps Richard’s point is saying that framing concedes too much. It concedes that we don’t trust each other. It concedes that we’re not going to want to regulate this stuff unless they’re doing it too. When, in fact, there’s huge portions of the American public that already want to regulate this stuff pretty heavily — regardless of what other countries do. I think that part of the critique sort of resonates with me and makes me wonder if we should instead have said, step one, the US regulates AI domestically, and then step two, we look and see what China is doing — and if they instead race ahead recklessly, then we talk to them and say, “We need to have a deal because we don’t want you to do that,” and then Plan A. I think that might have been both a more realistic way for this to go down and more what we would actually recommend because it’s good to get started early on good domestic regulation, rather than waiting until there’s a deal.
02:09:00先只做美国国内减速?What if we focused on a US-only slowdown first?
Luisa Rodriguez
你说也许更好的办法是描绘美国认真做国内减速。如果排出 Plan AA,优先国内暂停,具体长什么样?
You say that maybe a better approach would have been to depict the US taking serious steps to doing domestic slowdown. Do you have a vision for what that looks like concretely? If you were to lay out Plan AA and that version has domestic pause as a priority, what would that look like?
Daniel Kokotajlo
我们还没做这项工作,所以下面都有点试探。先是我们在 2027 情景里谈过的一整包渐进政策。比如投资核查硬件和核查开发,为后面铺路。也要求对 AI 公司如何训练模型有更多透明和监督,建设政府理解 AI、评估模型的能力。至于更认真、更显著的东西,我大概会建议某种与前沿 AI 公司算力预算有关的要求。现在他们预算里相当一部分——也许一半——用在研发和训练上往前推前沿。我觉得一般更好的是:80% 预算服务客户,20% 做研发和训练。如果有这类要求,相对容易执行,因为不需要那么多政府能力,就能在那个粒度上看数据中心在做哪类事。我觉得这会让 AI 进展的节奏慢一点,但不会疯。也许慢 25%,或 50%——我觉得大概是好事。会带来很多好处,而且不会伤害经济。相反,更多算力可用于推理,AI 价格会降一点。
We haven’t done this work yet, so you should take everything I’ve got to say as a bit tentative, but here are some ideas off the top of my head that I think I’d want to explore — and maybe we’ll explore in follow-up work. First of all, there’s a whole package of incrementalist policy ideas that we talk about in 2027 in the current scenario. For example, investing money in verification hardware and verification development sets this up for later. Also just requiring more transparency and oversight of the AI companies and how they train their models, and also building government capacity to understand AI and to evaluate AI models. As for something somewhat more serious and more significant, I would probably recommend something like a requirement to do with the compute budgets of these frontier AI companies. Right now they are using a significant fraction of their budget — like maybe half — on R&D and training to push the frontier forward. I think it would be generally better if they instead used 80% of their budget on serving customers and 20% on R&D and training. I think that if there was some sort of requirement like this, it would be relatively easy to enforce because it doesn’t require that much government capacity to check to see what type of thing that the data centres are doing at that level of granularity. I think it would cause the pace of AI progress to slow down a little bit, but not crazy. Maybe something like it would slow down by 25% or something, or 50% — which I think is probably good. I think that’s going to help lead to a lot of benefits and it would not hurt the economy. On the contrary, there’d be more compute available for inference, so prices would go down a little bit for AI.
会让中国追上吗?也许一点点——但只有一点点——因为现在很多中国 AI 进展某种程度上寄生在美国 AI 进展上,很多核心想法、算法、新范式是从美国领先公司在做的事情里抄的。有时是极其公开的信息,比如 Anthropic 大力投资编码智能体。人人看得见他们在做、开始奏效,于是很多地方也做同样的事。也有本该保密却仍在泄露的,有时也许还被刺探。这件事透明不多。但我会假设中国情报部门已经深深渗透所有美国 AI 公司,基本上在免费拿这些东西。还有蒸馏,中国 AI 可以借此从美国 AI 学习。由于这些原因,我觉得讽刺的是:减慢中国 AI 进展最有效的办法,是单方面减慢美国 AI 进展,因为中国 AI 进展有那么多来自美国。即便以今天 AI 进展的一半速度走,那大概仍是有史以来最快的技术变化之一。所以一半速度也没关系,仍然很快。想想现在的模型和一年前的模型的差别。一半那个速度仍然很快。再加建设政府能力、对公司更多透明、更好的监管框架。有一套框架明确授权美国政府监管 AI——比如阻止他们做智能爆炸、看清他们精确在做什么——同时创造制衡,让权力不只是重度集中在总统身上。可以设计涉及最高法院或国会委员会监督总统决定的框架,否则总统想做什么就做什么。需要更多研究。但我觉得那样会很好,因为它能防止竞赛条件下疯狂的权力争夺。
In terms of would it allow China to catch up? Maybe a little bit — but only a little bit — because right now a lot of Chinese AI progress is sort of parasitic on US AI progress, where a lot of the core ideas and algorithms and new paradigms are being copied from what the leading AI companies are doing in the US. In some cases it’s extremely public information, such as the fact that Anthropic invested heavily in coding agents. Everyone can see that they’re doing that and then people can see that it’s starting to work, so now people are doing the same thing in lots of other places. But then there’s also the things that are supposed to be secret that are leaking anyway, and in some cases perhaps being spied on anyway. There’s not very much transparency about this. But I would assume that basically Chinese intelligence services have deeply penetrated all of the US AI companies and are getting all this stuff for free, basically. Then also there’s distillation, where there’s another means by which Chinese AIs can sort of learn from US AIs. For all of these reasons, I think that ironically the most effective way to slow down Chinese AI progress is to unilaterally slow down US AI progress because so much of the Chinese AI progress comes from US AI progress. Even if you’re going at half the speed of today’s AI progress, that’s still probably one of the fastest technological changes that’s ever happened. So it’s OK if we go at half speed, that’s still really fast. Just think about the difference between the current models and the models of one year ago. Yeah, half that speed would still be very fast. I think also all the things I previously mentioned of building government capacity, more transparency into how the AI companies are going, better regulatory frameworks. I think that it would be really great to have some sort of framework set up that explicitly empowers the US government to regulate AI. For example, stop them from doing intelligence explosions and see exactly what they’re doing, while simultaneously creating a system of checks and balances so that power doesn’t just heavily concentrate on the president. You could design such a framework involving something like the Supreme Court or a congressional committee having oversight into the president’s decisions, otherwise the president gets to do whatever he wants. More research is needed. But something like that I think would be really great because it would prevent a crazy scramble power struggle under race conditions.
02:15:05强制执行:相互确保算力摧毁Enforcing a slowdown: Mutually assured compute destruction
Luisa Rodriguez
假设美中理论上想做这类协议,下一个难题是互不信任。双方都会担心对方在某个隐藏数据中心里秘密训练更强的系统。你的解决方案是核查,让每一方都知道违约会被发现并受罚。我的理解是 Plan A 有两条路。第一是双方申报算力:美中公开申报所有与 AI 相关的算力——芯片在哪、有多少、在产什么,然后允许对方检查那些设施。第二块是你所说的相互确保算力摧毁。能解释一下吗?
Assuming the US and China do want to make this kind of deal, in theory, the next difficult problem is that they don’t trust each other. Both will worry that the other will keep secretly training more powerful AI systems in some hidden data centre. Your solution is verification, so that each knows that defection would be detected and punished. My understanding is that Plan A has two approaches. The first is compute declaration from both sides, where the US and China would publicly declare all of their AI-relevant compute — so where the chips are and how many they have, and what’s being produced. And then they’d let each other inspect those facilities. The second piece is what you call mutually assured compute destruction. Can you explain what this is?
Daniel Kokotajlo
这是嵌进提案里的保障,好在提案破裂、人人又开始互赛时,事情不那么糟。协议运转时,人们对对方在做什么有透明。所以如果对方在做危险的事,比如智能爆炸,所有人能立刻看见,然后互相喊、让他们停。但想象那种破裂:有人照做不误、无视所有人的威胁和恳求;或者他们停止互相透明,于是害怕看不清别人数据中心里在做什么;或者因无关原因有冲突,比如因台湾开战。协议破裂、人人实质上互相冲突的方式有很多。
Yeah. This is a safeguard built into our proposal to make things less terrible in case the proposal breaks down and everyone starts racing each other again. As long as the deal is operational, people have transparency into what the other side is doing. So if the other side is doing something dangerous, like an intelligence explosion, everyone can immediately see that and then they can yell at each other and get them to stop. But imagine a situation where that breaks down and someone’s doing it anyway and ignoring everyone else’s threats and pleas. Or imagine a situation where they stop being transparent with each other and then now they’re afraid that they can’t tell what everyone else is doing on the data centres. Or imagine a situation where — for some unrelated reason — there’s a conflict, there’s a war over Taiwan or something. There’s all sorts of ways in which the deal could break down and everyone could be essentially in conflict with each other.
如果这几年建了所有这些新数据中心,然后那种冲突、协议破裂发生,他们冲向超级智能会比以前快得多。假如 2029 年他们离超级智能还有一年。到 2033 年,他们有更多算力。即便不算这些年已经取得的进展,也会不到一年。也许只剩一个月。从失控角度看,用一个月速通本来要一年的事,会极其可怕。从权力集中看,可能一家公司一个月就拿到超级智能,其他人都蒙在鼓里,也会极其可怕。所以我们认为协议该设计成:万一发生那种事,新建的数据中心被摧毁。做法是让美国能摧毁中国的数据中心,中国能摧毁美国的——新建的那些。这大概会是非常昂贵的升级行动,他们只会在局面相当危急时才做,因为自然会假设:我们摧毁他们的,他们也会摧毁我们的。但我们希望这种摧毁相对不流血:有经济损害,但溢成全面第三次世界大战的机会相对有限。
It would be especially bad if all of these new data centres had been built over the course of several years and then that conflict-deal-breakdown situation happens, because they’d be able to race to superintelligence much faster than before. If, say, in 2029 they were one year away from getting to superintelligence. Well, in 2033, they’d have more compute. They’d be less than one year away, even before taking into account the progress that they’ve made over those years. So they might be just like one month away. So it’d be extremely scary from a loss of control perspective to be speedrunning in one month what naturally would have taken a year. And of course, it would be extremely scary from a concentration of power perspective to have potentially one company going in one month to having superintelligence, with everyone else in the dark. That’s why we think that the deal should be designed in such a way that — in case of that type of eventuality — the new data centres that were built get destroyed. The way to do this is to make it so that the US can destroy the Chinese data centres, and then China can destroy the US data centres — the new ones, that is. Then presumably this would be a very costly escalatory action that they would only take if the situation was pretty dire, because they would naturally have to assume that if we destroy theirs, they’re going to destroy ours. But we want it to be the case that this destruction happens in a relatively bloodless way, where there’s economic damage, but a relatively limited chance of it spilling out into total World War III.
对没读过那篇文章的人,一个高层次要点是:世界有一个方便的事实——AI 进展严重依赖大型数据中心、大量算力,而世界上大部分与 AI 相关的算力就在这类大型数据中心里。我们认为,大约 99% 对 AI 进展有用的全球算力会在大公司拥有的这类大型数据中心里,而不是在你的笔记本上。所以你不必去追人们的笔记本,或杂七杂八只有一小台服务器的创业公司。只要看大的数据中心,申报你的大的数据中心。那拿不到全部,但能拿到显著超多数,我们认为基本上够了。我们认为在极少量算力上做出非常快的 AI 进展很难。还剩多少来自非大型数据中心?有点不确定,但在算力补充材料里我们谈过,我们认为有效上大约是 1%。
I think one thing that’s a high-level point to get across to people who haven’t read the piece is that it is a convenient fact about the world that AI progress depends heavily on large data centres, large amounts of compute — and most of the world’s AI-relevant compute is in these types of large data centres. We think that something like 99% of the world’s compute that would be useful for AI progress would be in these sorts of large data centres owned by big companies, rather than on your laptop. So you don’t have to track down people’s laptops or miscellaneous startups with their little server. Just look at the big data centres, declare your big data centres. That doesn’t get everything, but it gets a significant supermajority of things, which we think is basically good enough. We think that it’s really hard to make very rapid AI progress on tiny amounts of compute. How much compute would be still available coming from not large data centres? This is a bit uncertain, but in our compute supplement we talk about this and we think it’s effectively like 1%.
有两种办法去达成这些目标。我们大致建议两个都做,但也许单独一个就够。一种是技术路径:把新芯片和新数据中心设计成实际上有对手国家控制的杀开关。可以想象芯片必须收到来自中国的某个代码,中国停发代码芯片就停转。同样,中国芯片必须收到美国的代码才能继续工作。这是非常不流血的方式,但你可能怀疑这种技术东西——万一有办法黑、有后门、有陷阱?如果你担心那个,就有相反的办法——非常钝、笨,但更难骗——美国把数据中心建在蒙古,中国把新建的数据中心建在加拿大。一旦冲突、协议破裂、人人互相恼火,美国可以兼并中国的数据中心,中国可以兼并美国的。大概在即将被兼并时,里面的人会自毁 GPU,以免落入敌手。于是你落到同一结果:GPU 被毁,谁都没有,但比数据中心在本国领土、可能挨着本国城市更少升级。如果你在北弗吉尼亚、华盛顿郊外就有数据中心,中国必须朝它们射导弹才能摧毁,那似乎很容易升级成真正的第三次世界大战。
Here are two different ways you could try to achieve these goals. I think we just sort of recommend you do both, but maybe either one by itself will be sufficient. One is the technical way, where you design the new chips and the new data centres in such a way that they effectively have kill switches controlled by the rival country. You can imagine that the chips are designed so that they have to receive a certain code from China, but if China stops sending the code then the chip just stops working. Similarly, the Chinese chips have to receive a code from the US to continue working. That’s a very bloodless way that each side could do that, but you might be suspicious about that sort of technical thing — what if there’s some way to hack it or backdoor it, or what if there’s some catch there? If you’re worried about that, then there’s the opposite approach — which is the very blunt, dumb approach, but the approach that’s harder to fool — which is that the US builds their data centres in Mongolia and China builds their new data centres in Canada. So in case of conflict, in case of the deal breaking down and everyone being angry at each other, the US can annex the Chinese data centres and China can annex the US data centres. Presumably if they were about to be annexed, the people in them would self-destruct their own GPUs to prevent them falling into enemy hands. So you would end up with the same result. You end up in a situation where the GPUs have been destroyed, nobody has them, but it’s less escalatory than if the data centres had been on home territory, possibly by home cities. If you have data centres right outside DC in Northern Virginia and China has to shoot missiles at them to destroy them, that seems like it could easily escalate to actual World War III.
我们想让 GPU 摧毁极其昂贵,好让它不会被轻易使用,只作为其他一切都失败后的最后手段,但又不要贵到、绑到一切上,以至于有很高机会通向第三次世界大战。它仍可能导致第三次世界大战。我们不想要那个。这会是非常可怕的局面。我们绝对不想它发生。但我们想让它相对不那么可怕,或让它成为第三次世界大战的出口匝道,而不是入口匝道。如果加拿大、蒙古不愿意,就换一个愿意的国家。我们并不死守必须是加拿大。我们其实认为大概有一大堆国家会很乐意做这种事,因为它会给他们地缘政治权力。作为谈判的一部分,一堆国家该被卷入,然后大概至少会有一两个国家愿意为了钱、或为了他们想要的各种让步而做这种事。比如蒙古默认完全没有 AI 产业,也许担心会被这场别处都在发生、唯独蒙古除外的 AI 革命晾在冷处。也许作为在他们国家建所有这些数据中心的回报,他们能拿到真正能撬动 AI 如何发展的东西。比如条件可以是:他们对数据中心本身有透明,也许甚至拥有其中一些或一部分。大概有办法让这极有吸引力。这交给外交官和领导人去谈。
We wanted to make it so that the GPU destruction is incredibly costly, so that it wouldn’t be done trivially and would only be done as a last resort when all other things have failed, but not so costly and tied up with everything that it has a high chance of leading to World War III. It still could lead to World War III. We don’t want that. This would be a very scary situation. We definitely don’t want this to happen. But we want to make it relatively less scary, or making it an off-ramp from World War III rather than an on-ramp to World War III. If they don’t want to do it, then pick a different country that does want to do it. We’re not super committed to it has to be Canada. We do actually think that there’s probably a whole bunch of countries that would love to do something like this because it would confer geopolitical power to them. As part of the negotiations for setting up something like this, a bunch of countries should be involved, and then probably there’ll be at least one country — or at least a couple countries — that are willing to do something like this in return for money, or in return for various concessions that they want. For example, Mongolia by default has absolutely no AI industry whatsoever and maybe is worried that it’s going to be left in the cold by this AI revolution that will be happening everywhere else except for Mongolia. Perhaps in return for having all these data centres built in their country, they can get some things that give them actual leverage and power over how AI develops. For example, it could be part of the conditions that they get transparency into the data centres themselves, and maybe they even get to own some of those data centres or some fraction of them. There’s probably a way to make this extremely appealing. This is just a matter for the diplomats and the leaders to negotiate.
02:24:23欺骗协议Cheating on a slowdown agreement
Luisa Rodriguez
这些是让违约昂贵的部件。违约实际会长什么样?
So these are the pieces you have in place to make defection costly. Can you actually talk about what defection would look like?
Daniel Kokotajlo
我们有一整条侧枝,叫秘密项目迷你情景,还有一份秘密项目补充材料写我们的分析。这是很多政策和国家安全人士非常担心的事,所以我们花了大量时间想、写这方面。事实上,从一开始它就是 Plan A 背后的主要动机之一。我们假设美中完全不信任对方,所以必须核查。意味着我们该大量思考:一个国家资助的秘密项目,能在不被抓到的情况下做成什么。Plan A 的核心想法是:如果你把他们足够多的算力放进这份透明协议,剩下那一点点算力——即便全聚到一个秘密项目里——也无法快到打败透明项目。
We have a whole side branch which you can read called the covert projects mini scenario. Then there’s also a covert project supplement that goes into our analysis. This is something that I think a lot of policy people and people in national security are very concerned about, so we spent a lot of time thinking about it and writing up this aspect of our scenario. In fact, it was one of the main motivating concerns behind Plan A from the start. We assume that the US and China don’t trust each other at all, so they have to verify things. That means that we should be thinking a lot about what a state-sponsored covert project could get away with without being caught. Again, the core idea of Plan A is that if you have enough of their compute in this transparency deal, then the tiny amount of compute left over — even if it’s all gathered into one covert project — won’t be able to make AI progress fast enough to beat the transparent projects, basically.
我们推演过一个情景:中共在一座水电站下面建秘密项目。我们算了他们需要多少电,算了怎么把走私 GPU 弄到这个地点,再按各种参数算这么多 GPU 上 AI 进展会有多快。这就是我们最在想的那种违约。高层次的说法是:如果你把世界上足够多的算力弄成透明,剩下在做秘密非法之事的那部分就会太小,短期内——比如几年——构不成那么大的威胁。几十年尺度上它够大,能做各种各样的事。反方是:如果你觉得他们用这么一点点算力两年内就能到超级智能,那你也该觉得主项目凭他们那巨量算力不到两年就能到超级智能。少得多,大概几个月。有一种关联——我觉得很多人没认识到——如果你觉得 AI takeoff 或智能爆炸会慢、会被算力卡住,那你也该觉得它相对容易治理、限制、监管。反过来,如果你觉得很难限制和监管,因为地下室里只有 10 万 GPU 的秘密集群就能做非常疯狂的事,那你对当前局面该更慌。因为当前局面更像 Yudkowsky 的经典情景:它可能在 OpenAI 拥有的某座巨型数据中心里一个月 foom 到超级智能。基本上,你觉得 takeoff 有多快,和你觉得事情有多可治理,是绑在一起的。
Getting into that a bit more: we gamed out a scenario where the Chinese Communist Party builds a covert project underneath this hydroelectric power station. We calculated how much power they would need, and we calculated how they would get the smuggled GPUs and bring them to this location. Then we calculated based on various parameters, like how fast their AI progress would go on this amount of GPUs. That’s the type of defection that we’re most thinking about. Again, the high-level thing is if you get enough of the compute in the world transparent, then whatever’s left over doing secret illegal stuff can be too small to really pose that much of a threat — at least in the short term, like in a couple of years. It’s large enough that over the course of decades it would be able to do all sorts of things. The counterposition to that is that if you think that actually they’d be able to get to superintelligence in two years using this tiny amount of compute, then you should also think that the main AI projects would be able to get to superintelligence in less than two years, given their huge amount of compute. Much less, in fact, probably just a few months. There’s a sort of correlation — or there’s this relationship which I think not many people have recognised — which is that if you think that AI takeoff or the intelligence explosion is going to be slow and bottlenecked by compute, then you should also think that it’s relatively easy to govern it and restrict it and regulate it. Whereas if you think that it’s very hard to restrict and regulate because some tiny people in a basement with only 100,000 GPUs in their covert cluster can do really crazy things, then you should be even more freaked out about the current situation. Because the current situation is more like Yudkowsky, the classic Yudkowsky scenarios of it could foom to superintelligence in a month in one of these giant data centres that OpenAI has. Basically there’s this relationship of how fast do you think takeoff is, and how governable you think things are?
Luisa Rodriguez
侦测够不够快,让一个国家没法用透明算力偷偷往前猛冲?
Is detection fast enough that it’s not possible for one of the countries to make a bunch of progress using the transparent compute?
Daniel Kokotajlo
这也是我觉得完全研究透明很好的例子之一,对比另一种可能:政府审计师每月进来问员工一堆问题,也许接入网络看看在发生什么。如果你有那种系统:中等透明,政府审计师大约每月能看见在发生什么,但公众看不见。那在各方面都会更低效,也可能因你刚说的原因更危险:审计师一走,人们就说「好,他们回来之前我们还有整整一个月。」「我们疯一把,赶在他们回来之前。」或者政府一直在,但只被允许问某些问题,或只能看见某些部分。或者人很少,也许就能被说服某件事没事,基本上被糊弄成接受其实并不没事的东西。比如有一种活动可以伪装成无害的对齐研究,实际上却在有效地训练一个超级智能,你得是专家才能看出那活动到底是什么。如果只靠偶尔进来看一看的政府审计师,那些审计师也许会犯错,也许认不出那活动是什么。相比之下,如果你有完全研究透明,网站会实时更新数据中心上新活动的日志,公众里所有人——包括对手公司、包括其他国家的政府——都能看见那些日志。于是响应时间极快。如果某家公司在做真正令人担心的事,它会以几乎能被注意到的最快速度被注意到。
This is one of the examples of why I think the total research transparency is nice, contrasted with a different possibility of government auditors that come in every month or something and ask a bunch of questions to the employees and maybe tap into the network to see what’s going on. If you had that sort of system: there was a medium amount of transparency, where the government auditors can see what’s going on every month or so but the public can’t see. That would be less effective in various ways and it could be more risky for the reason that you just described, for example, where after the auditor leaves, people are like, “OK, we have a whole month before they come back. Let’s go crazy before they get back.” Or maybe the government is there continuously but they’re only allowed to ask certain questions, or they’re only able to actually see what’s going on in certain parts of it. Or they’re only a few people, so maybe they can just be convinced that something is fine because they are basically bamboozled into accepting something as fine when it’s actually not fine. For example, maybe there’s a type of activity that can be disguised as harmless alignment research, but actually is effectively training an AI to be superintelligent and you have to be an expert to look at that activity and then realise what’s really going on there. If you’re just relying on some government auditors that come in and occasionally look over stuff, then maybe those government auditors will make a mistake and maybe they will not recognise that activity for what it is. By contrast, if you have the total research transparency, then in real time the website is being updated with the logs of the new activity that’s happening on the data centre and everyone in the public — including rival corporations, including other countries’ governments — can just see those logs. So there’s just extremely fast response time. If some company is doing something that’s really concerning, it will be noticed at approximately the maximum speed it could be noticed.
02:30:42相互确保算力摧毁行不行Would mutually assured compute destruction work?
Luisa Rodriguez
假设有秘密项目,理论上也有可能用透明算力违约猛冲。你的提案意味着一旦有违约,另一国就能摧毁他们的算力。假设中国在违约。如果美国摧毁中国的算力,没有什么能阻止中国接着摧毁美国的算力。美国愿不愿意摧毁中国的算力去惩罚他们,取决于美国自己会承受多大经济损失。我们在说多大的损失?
Let’s say there are covert projects. In theory, there’s a possibility of using transparent compute to try to defect and make a bunch of progress. Your proposal means that if there is a defection, the other country will be able to destroy their compute. Let’s say China is defecting. If the US destroys China’s compute, there’s nothing stopping China from destroying the US’s compute at that point. It feels like how willing the US will be to destroy China’s compute to punish them depends on how much economic loss the US will then experience. How much economic loss are we talking about?
Daniel Kokotajlo
一开始就会是显著的量,然后会往上走。随着经济越来越依赖 AI,它会变成经济越来越大的一块。我们想当现实主义者看待谈判会怎么走。根本上是这些国家有不同利益、对什么有风险什么没有有不同看法。然后他们互相喊、讨价还价谁该做什么、什么活动必须停。他们挥着各种胡萝卜和大棒。我们基本上希望谁也不能做一件让美国或中国这种主要大国确信自己即将被彻底剥夺权力的事。比如谁也不能做疯狂的智能爆炸去拿超级智能。但我们也不希望这些大国能因为不喜欢你加的关税就随便威胁摧毁别人的算力——那给他们的权力太大了。我们希望按这个「摧毁算力」按钮对按的人非常昂贵,只是相对最后的手段。我们认为我们提的这套相对钝的提案能做到这一点。
It would start off as a significant amount, and then it would go up from there. As more and more of the economy depends on AI, it would become a bigger and bigger part of the economy. In general, we’re trying to be realist about how the negotiations will go down. The ultimate thing that’s going on is that these different countries have different interests and different opinions about what’s risky and what’s not. Then they’re yelling at each other and bargaining about who should be doing what and who shouldn’t be doing what and what activity needs to stop. They’re waving various carrots and sticks around in service of that. We basically want it to be the case that nobody can do something that convinces a major power, such as the US or China, that they’re about to be completely disempowered. For example, nobody can do a crazy intelligence explosion to get superintelligence. But we don’t want it to be the case that these major powers can just threaten to destroy people’s compute willy-nilly because they don’t like the tariff that you put on them — that would be giving them way too much power. We want it to be the case that pressing this “destroy the compute” button is a very costly action for the person who presses it. It’s only a relatively last resort. We think that this relatively blunt proposal that we proposed accomplishes that.
Luisa Rodriguez
Tom Davidson 指出过类比是算力和核武器的相互确保摧毁。核武器那边,一国知道如果用核武器,会有核武器报复,因为有足够时间发现核弹打过来并回击。那制造威慑。这边威慑似乎更弱:假设中国想违约,中国知道美国可以选择不惩罚中国,好保住自己的算力、不破坏自己的经济。如果自己算力被毁的经济成本够大,也许中国会赌美国不会为惩罚中国而拿这么大一块经济冒险——因为那样不会像核武器那样字面杀死公民,但会造成巨大贫困。
Tom Davidson pointed out the analogy here is between compute and mutually assured destruction with nuclear weapons. With nuclear weapons, a country knows that if they use nuclear weapons, there will be retaliation with nuclear weapons because there’s enough time for that country to notice that nuclear weapons are coming and to respond by launching their own. And that creates deterrence. In this case, it feels like the deterrence is weaker because — let’s say China wants to defect — China knows that the US has the option of not punishing China in order to maintain its own compute, in order to not sabotage its own economy. So if the economic costs of its own compute being destroyed are big enough, then maybe China takes the bet that the US won’t punish China for defecting because it’s just not willing to jeopardise this massive portion of its economy — because doing that wouldn’t literally kill its citizens the way nuclear weapons would, but it would cause enormous poverty.
Daniel Kokotajlo
我们想避开两个极端。我们在 2031 年那一节谈过这个。我们想避开一国能单方面做对其他国家极其有威胁的事然后全身而退。解决那个的办法是:主要大国至少有能力摧毁算力。所以如果某件事极其有威胁,他们会做,即便成本巨大、即便会重创他们的经济。成本绝对巨大,但那是好事。我们想要的成本是:你只愿意为阻止更糟的事付出,否则不付。我们不想各国因为某笔没谈成的贸易协议就左右互删算力。摧毁算力是最后手段,你只会为防止你更害怕的事才做。如果他们做的事没到那个门槛,那就更像普通外交。
Like I said, we want to avoid two extremes. We talk about this in 2031. We want to avoid a situation where a country can unilaterally do something that’s extremely threatening to other countries and they just get away with it. The thing that solves that is the major powers at least have the ability to destroy the compute. So if something’s extremely threatening, then they would do it even though it would cost them a huge amount and even though it would heavily damage their economies. Yeah, the costs are definitely huge, but that’s good. We want the cost to be such that you only are willing to pay that cost in order to stop something even worse, but that you otherwise don’t pay the cost. We don’t want it to be that the countries are just deleting each other’s compute left and right because they are upset about some trade deal that didn’t happen. This is a last resort, destroying the compute, and you’d only do it to prevent something that you’re even more scared of. If they’re doing something that’s not meeting that bar, then that’s just more of an ordinary diplomacy-type situation.
我们谈过的例子:假设某地某家公司——也许在中国,也许在美国——在研究一种持续学习的新范式,能让 AI 在岗上非常有效地学习,从而在各种正在做的事情上很快变得非常聪明。副作用是弄坏我们当时在用的很多对齐技术。因为透明,这一开始就会有人注意到,然后会有一整轮国际新闻周期,谈他们在做的、有些人觉得非常危险的事。然后本地监管者——真正有管辖权的那个——假设在中国,某家中国公司在做这个。中国监管者会不会说「嘿,那很可怕,关掉」?也许会。假设他们不。然后美国可以「嘿,我们觉得那很可怕。我们要你们关掉。」中国监管者说「我们觉得没事。我们不想关。」然后美中得互相喊一阵。也许这是可怕、但还没可怕到美国会为此删掉所有算力的例子。也许美国会删掉那算力并不可信。但他们可以做别的:「我们会很难过,可能加关税或额外管制,可能不请你参加下届奥运会」——或任何通常的外交谈判和施压杠杆。基本上,如果某件事可怕到美国愿意为此摧毁所有算力,那就发生那个。如果没那么可怕,就做更正常的外交谈判。结果我们觉得大致会是:没那么可怕的那种事就会发生。某件事越可怕,它发生的可能性就越低。如果极其可怕,它就不会发生,因为别人会介入阻止。
So here’s the example that we do talk about: suppose that some company somewhere — maybe in China, maybe in the US — is researching this new paradigm of continual learning that would allow the AIs to learn on the job really effectively, and therefore become really smart really fast at a variety of things that they were doing. Also, as a side effect, break a lot of the alignment techniques that we’d currently be using. This is something where as soon as this starts happening, because of the transparency, someone would notice and then there’d be a whole international news cycle about this thing they’re doing that some people think is really dangerous. Then maybe the local regulator — the regulator that actually has jurisdiction over them — say it’s in China, and some Chinese company is doing this. Does the Chinese regulator say, “Hey, that’s scary, shut it down”? Maybe they do. Suppose they don’t. Then the US can be like, “Hey, we think that’s really scary. We want you to shut down.” And the Chinese regulator says, “We think it’s fine. We don’t want to shut it down.” Then the US and China have to yell at each other a bit. Maybe this is an example of something that’s scary, but it’s not so scary that the US is going to delete all the compute because of it. Maybe it’s not credible that the US would delete that compute. But then they can do other things and they can say, “We’ll be very sad and we might put some tariffs or some extra controls on you, or we might not invite you to the next Olympics” — or whatever the usual levers of diplomatic negotiation and pressure are. Basically, if it’s something that’s so incredibly scary that the US is willing to destroy all the compute for, well, then that’s what happens. If it’s not that scary, then you do more normal diplomatic negotiations. The result will be, we think, that roughly speaking the type of stuff that’s not that scary will just be happening. Basically the more scary something is, the less likely it is to happen, effectively. If it’s incredibly scary, then it just won’t happen because other people will intervene to stop it.
我们的第一担心就是监管者会做糟糕决定,批准其实非常危险的东西。那字面上就是我们的第一担心。不过这种担心某种程度上内在于建造超级智能本身。如果 AI 公司就是要造超级智能,你还能怎么缓解这个担心?我们在尽一切所能把监管者放到能做对判断的位置。我们给他们大量对公司在做什么的透明。我们也让事情总体以稍慢、合理的节奏走,而不是非常快,好让监管者有更多时间了解发生了什么。我们也让公众看见发生了什么,好让学术界、科学共同体、对手公司能看、能批评,于是不是一个监管者关在房间里对着一家极有偏见、试图糊弄他的公司。而是有一家对手公司有相反激励,想说服监管者这很危险。更像有双方律师的法律系统。我觉得我们在尽一切所能让监管者处于能做对技术判断的位置。但仍有显著风险他们会做错误的技术判断。如果你真的很怕那个,那就该走 Plan S,全部关掉,好让这种监管者错误没有可能。但如果你就是要造超级智能——你还能怎么做才能让这个问题不那么糟?其他计划似乎会让这个问题更糟,因为监管者要么根本不存在,要么信息更少,要么更有偏见,因为就是公司自己在监管自己。
Yes. This is why our number one concern is that the regulators will make poor decisions and sign off on something that is in fact very dangerous. That’s literally our number one concern. However, this concern is kind of inherent in building superintelligence at all. If you’re going to be having AI companies build superintelligence, how else are you supposed to mitigate this concern? We’re trying to do everything we can to put the regulators in the right position to make the right calls here. We’re giving them massive amounts of transparency into the AI companies and what they’re doing. We’re also making things just generally go at a somewhat slow, reasonable pace instead of going really fast, so the regulators have more time to learn about what’s going on. We’re also letting the public see what’s going on too, so that the academic community, scientific community, rival corporations can look at what’s going on and critique it, so that it’s not just a regulator in a room with a corporation that they’re trying to regulate and the corporation is incredibly biased and trying to bamboozle the regulator. Instead, there’s a rival corporation that has the opposite incentive and wants to convince the regulator this is dangerous. So there’s more like a legal system where there’s a lawyer arguing for both sides. I feel like we’re doing everything we can to put the regulators in the right position to make the right technical calls here. But there’s still a significant risk that they’ll make the wrong technical calls. I think that if you’re really afraid of that, then you should just go for Plan S and shut it all down so there’s no possibility of regulator error like this. But if you’re going to be building the superintelligence — how else are you supposed to do this? I don’t see how else you’re supposed to do it in a way that makes that problem less bad. The other plans seem to make that problem even worse because the regulators either don’t exist at all or have less information or are more biased because they are just the company themselves — like the companies regulating themselves.
Tom Davidson 提议改去缩放软件,理由是:如果你大规模建算力,再拿掉对用那些算力训练的限制,就可以有极快的智能爆炸。但如果缩放的是软件,即便协议破裂,也不会让你有那么极快的智能爆炸。也许稍快一点,但没那么快。那是非常合理的 Plan A 替代方案。你可以叫它 AA 而不是 A。担心是:如果他们不摧毁算力?然后事情极快、极危险。或者相关地,如果他们对什么安全什么不安全做了糟糕选择?如果我们觉得他们会系统性做糟糕选择,特别是系统性允许太多事发生?那他们有这个「摧毁算力」按钮也帮不了太多,因为他们反正在允许它发生、不按按钮,于是事情会很快,因为他们有所有这些额外算力。由于这两个原因,你可能担心我们当前这个版本:建很多算力但有可摧毁按钮。这些担心非常合理,如果有一个 Plan A 变体基本上禁止新数据中心但允许更多算法进展,我会很高兴。
That’s a very reasonable alternative plan to Plan A. I don’t know, you could call that like AA or something instead of A. There’s this concern about what if they don’t destroy the compute? And then things go incredibly fast and are incredibly dangerous. Or perhaps relatedly, what if they make a bad choice about what’s safe and what’s not? What if we think that they’re systematically going to make bad choices, and in particular they’re systematically going to allow too much stuff to happen? Then the fact that they have this “compute destroy” button doesn’t help so much because they’re just allowing it to happen anyway and they’re not pressing the button, so things will just go quite fast because they have all this extra compute. For both of those reasons, you might be concerned about our current version of the plan where they build lots of compute but then have the destroyability button. I think those are very reasonable concerns and I’d be very happy with the Plan A variant that basically bans new data centres but allows more algorithmic progress.
但说说我们为什么喜欢我们的版本。其一就是可逆这个核心想法:你没法把算法取消发明。你可以试着禁止,但很难。如果你不建数据中心,却允许公司按他们想要的速度、哪怕只是稍快地创造新范式,那是你撤销不了的进展。AI 的能力只会上棘轮,从现在起它们会一直那么强。如果事实证明它们开始递归自我改进,更难踩刹车。另一件是为了安全你可能想用上那些算力。算力对很多事有用。你可以用它在世界上做好事,用来增长经济。如果你不建新数据中心,很难拿到我们谈过的那种经济转型。你得走向越来越花哨的 AI 能力级,希望质量补上数量。另一件是:我至少怀疑,很多不对齐风险来自质变而不是量变。如果你留在当前范式里,只是把模型做大、建更多数据中心好跑更多份,那只稍微更危险一点。而如果你让它们自主发明新范式、改变做事方式,那会引入很多可能的错误,以及可能弄坏对齐技术和控制技术的东西。你可能还想付安全税。可能有一个对齐解实际效果很好,但效率低十倍——于是训练要花十倍算力,AI 持续运转也要十倍算力才能用上这套技术。例子可能是思维链。现在思维链是默认,但未来可能有更多 neuralese 型设计——那些设计会流行,大概是因为更高效。于是想象想逆回去用思维链,即便更低效,因为它更好懂。如果你已经堆了很多算力,那就很容易:10 倍惩罚,没问题,一两年我们会有 10 倍算力,付得起。而如果你不造更多算力,10 倍惩罚就会让我们慢 10 倍。
But let me say the reasons why we liked our version. One of them is just this core idea of reversibility, where you can’t really uninvent algorithms. You can try to ban them, but it’s hard. If you’re not building your data centres, but you’re allowing the companies to create new paradigms as fast as they want to, or even just at somewhat of a fast speed, then that’s progress you can’t undo. The AIs are just ratcheting up in terms of capability and they’re going to always be that capable from now on. Insofar as it turns out that they are starting to recursively self-improve, it’s harder to pull the brakes on that. Another thing is that for safety purposes you might want to use all that compute. Compute is useful for many things. You can use it to do good things in the world, you can use it to grow the economy. It’s going to be harder to get the type of economic transformation that we talked about if you’re not building new data centres. You’d have to proceed to fancier and fancier levels of AI capability and hope that the quality makes up for the quantity. That brings me to another thing, which is that I suspect at least that a lot of the misalignment risk comes from the qualitative changes rather than from the quantitative changes. If you stay within the current paradigm, but then make the models bigger and make more data centres so you can run more of them, that’s only slightly more risky. Whereas if you are having them autonomously invent new paradigms and change the way things are done, that’s introducing a lot of possible errors and possible things that could break your alignment techniques and your control techniques. Also you might want to pay safety taxes. It might be the case that there’s an alignment solution that actually works really well, but it’s 10 times less efficient — so you need to spend 10 times more compute for training and 10 times more compute for the ongoing operation of the AIs in order to make use of this technique. An example of this might be chain of thought. Right now chain of thought is the default, but in the future there might be more neuralese-type AI designs — and presumably the reason why those designs would become popular is because they’re more efficient. So imagine wanting to reverse that and actually go back to chain of thought, even though it’s less efficient, because it’s easier to understand. If you’ve built up lots of compute, then that’s really easy to do because a 10x penalty, no problem, in a year or two we’ll have 10x as much compute and so we’ll just be able to pay that penalty, no problem. Whereas if you’re not making more compute, then the 10x penalty is just going to slow us down by 10x.
02:54:18减速还是关停更好Is slowing down or shutting down better?
Luisa Rodriguez
我们谈过有些人赞成现在就关掉所有 AI 发展——你所说的 Plan S。听起来你对 Plan S 的主要反对是:它大概不会永远持续,一旦围绕 Plan S 的协调不可避免地破裂,AI 进展就会全速前进。这大致表达了你的观点吗?偏好 Plan S 的人会怎么回?
We’ve talked a bit about how some people favour just shutting all AI development down now — what you call “Plan S.” It sounds like your main objection to Plan S is that it just probably wouldn’t last forever, and once coordination around Plan S inevitably broke down, AI progress would proceed at full speed. First, does that express your view roughly right? And if so, what do you think that people who prefer Plan S would say in response?
Daniel Kokotajlo
我觉得大致对,我们还能说更多,但那是主要理由。偏好 Plan S 的人会说:也许我们对让所有人同意这种事、协调起来的能力太悲观。对此我会说:也许吧。这是哪个在政治上更可行的问题,我目前的猜测是 Plan A 会更可行、更稳定。但如果事实证明 Plan S 更可行、更稳定,那就很重要。也许我会转而主张某种 Plan S。我们和很多人谈过,拿到各种意见,但值得注意的是,我觉得没人真知道什么会在政治上可行。尤其是华盛顿的人,他们对现在什么政治上可行非常敏感,但完全不擅长预测几年后、AI 改造了局面之后什么会政治上可行。有很多例子:政治决定和他们两年前说会做的、以及所有人以为 Overton 窗口里的东西,完全 180 度掉头。
I think that’s roughly right, and I think that there’s more things that we could say besides that. But I think that’s the main reason. I think that people who prefer Plan S would say maybe we’re being too pessimistic about the ability to get everyone to agree with something like this and to coordinate. To which I would say: yeah maybe. There’s a political question of which of these things is going to be more feasible, and my current guess is that Plan A is going to be more feasible and more stable. But if it turns out that actually Plan S is more feasible and more stable, then that would be significant. Maybe I would switch to advocating something like Plan S. We’ve talked to many people and gotten various opinions, but notably I don’t think anybody really knows what’s going to be politically feasible. I think especially people in DC, they’re very attuned to what is politically feasible now, but they are not at all good at predicting what will be politically feasible in a few years after AI has transformed things. There have been many examples of people, of things happening, political decisions being made that were complete 180s from what they said they would do two years ago, and what everyone thought was in the Overton window.
还有秘密项目。如果你担心某处有一个秘密项目在冲向超级智能,Plan A 的一个优势是你可以调节透明项目的 AI 进展速度,确保你领先于可能的秘密项目。说清楚:你大概仍该调节得相当多,因为秘密项目大概在从你这里偷大量进展——所以你不该只是超快地走,那只会让他们也超快。但要点是:如果你在向前推进,并且在调节推进量,你可以大致确保走得更快。而如果你自己完全没有任何向前的进展,那么——如果某处有一个大的秘密项目——你至少该有点担心它最终会造出疯狂的东西。当然还有 AI 能带来的所有好处。Plan A 高层次上基本上是说:在人类水平 AGI 附近暂停,那一级我们觉得大概能控制。它弱到我们觉得即使用相对普通、大概不那么难发明的技术大概也能控制。但它强到能彻底改造经济,让 GDP 每年翻倍,以及在其他方面大幅改善局面。如果你能打中那个甜蜜点并停在那里,你可以拿到很多好处,风险却不很多。
There’s also covert projects. If you’re worried that somewhere there’s a covert project that’s working towards superintelligence, an advantage of Plan A is that you can sort of titrate the speed of AI progress across the transparent projects to make sure you stay ahead of the possible covert project. And to be clear, you should still titrate it quite a lot probably, because the covert project is probably stealing a lot of its progress from you — so you shouldn’t just go super fast because that’s just going to make them go super fast too. But the point is that if you’re making forward progress and you’re titrating the amount of progress you’re making, you can sort of be sure to go faster. Whereas if you just absolutely aren’t making any forward progress yourself at all then — if there’s a large covert project somewhere — you should be at least somewhat concerned that eventually it’s going to build something crazy. Another thing, of course, is all the benefits that can come from AI. One thing that I think I previously mentioned is that Plan A — at a high level — is basically saying pause around human-level AGI, which is a level sufficient that we think we can probably control it. It’s weak enough that we think we can probably control it, even with relatively prosaic techniques that are probably not too hard to invent. But it’s strong enough that it can utterly transform the economy and cause GDP to double every year, and otherwise just greatly improve the situation. If you can hit that sweet spot and stay there, you can get a lot of benefits without very much of the risks.
Luisa Rodriguez
Plan A 给美中政府很多权力:他们决定哪些算法安全、多少算力能用于什么。如果我们落到一个想当独裁者的总统,似乎说得通他会滥用那权力,至少在某些方面增加权力集中风险。鉴于你对短时间表和对齐难度的担心,那也许是两害相权取其轻。但如果有人觉得对齐没那么难,或时间表更长,这也许是 Plan A 的大缺点。对你来说是这样吗?
Plan A gives the US and Chinese governments a lot of power: they get to determine which algorithms are safe, how much compute can be used for what. If we ended up with a president who wanted to be a dictator, it seems plausible that they could abuse that power, increasing concentration of power risks in at least some ways. Given how worried you are about short timelines and the difficulty of solving AI alignment, that might be the lesser of two evils. But it seems like if someone thought alignment wasn’t going to be so hard, or thought that timelines were longer, this might be a big downside of Plan A. Does that seem true to you, or not necessarily?
Daniel Kokotajlo
对我来说完全是假的。我认为 Plan A 对防止 AI 独裁非常好。首先,不做智能爆炸,而是缓慢谨慎地推进 AI 发展,对避免独裁非常好,因为独裁的一个主要风险因素是:有一支全部被中央控制的超级智能军队,没有其他相当的 AI 军队能制衡——而这正是你有智能爆炸时会得到的。因为如果你有智能爆炸,谁先开始谁就能拉出巨大领先,潜在地在别人沿那条曲线走远之前就到超级智能。不一定如此。可能有两家公司并驾齐驱,近到即便在做智能爆炸,两边仍相当。但一般来说,如果你允许智能爆炸,即便相对小的差距——即便一家公司只落后六个月——也可能转化成实际质变能力上的极大差距。而如果你没有智能爆炸,六个月差距就没那么大事。它不是能让人接管世界的东西。所以你在防止任何人——无论总统还是 CEO——对其他所有人积聚巨大权力,如果你防止智能爆炸。
That seems totally false to me. I think Plan A is really good for preventing AI dictatorships. First of all, not doing intelligence explosions and instead proceeding slowly and cautiously with AI development is great for avoiding dictatorships because one of the main risk factors for having a dictatorship is if there’s an army of superintelligences that’s all centrally controlled, and there’s no other army of comparable AIs that can act as a check and balance on it — which is what you get if you have an intelligence explosion. Because if you have an intelligence explosion, then whoever started doing it first can build up this huge lead and potentially get to superintelligence before other people have gotten far along that curve. It’s not necessarily true. You could potentially have two companies that are neck and neck and they’re so close to each other that even as they’re doing an intelligence explosion, they both stay comparable. But just generally speaking, if you’re allowing intelligence explosions, then even relatively small gaps — even if one company is only six months behind — that could translate into an extremely large gap in terms of actual qualitative capability. Whereas if you don’t have intelligence explosions, then a six-month gap is not that big of a deal. It’s not something that enables somebody to take over the world. So it’s just really great. You’re preventing anybody — whether they’re president or CEO — from accumulating a huge amount of power over everybody else if you prevent intelligence explosions.
第二是透明。我认为控制 AI 发展的人滥用权力的一个主要方式,是让他们的 AI 追求自己的议程——具体是追求对建造 AI 的人有利的议程,而且也许秘密地做。想象 OpenAI 宣布他们的 AI 会试图卖东西给你,并试图让你对他们的产品上瘾,还会试图说服你把票投给 OpenAI 偏好的政治候选人。显然,如果这成了公开信息,效果不会那么好,因为人们会停用 ChatGPT,用的时候也会对这种说服保持戒备。可如果 OpenAI 做这种事而且是秘密的——只是一场除了一些阴谋论者外没人知道的微妙影响运动——那就可能有相当显著的效果。所以关于 AI 如何被训练、被灌进什么目标和价值观的透明,对防止这类权力滥用非常好。无论是 CEO 还是总统都一样。他们在完全研究透明的条件下仍可以试,但比没有完全研究透明难得多,因为人们会看见他们在做什么,然后可以反应。透明加上不搞智能爆炸、买时间,也意味着多家公司能追上。即便你完全不关心失控,觉得 AI 会很容易被控制——只要你担心权力集中和 AI 独裁,你就该对 Plan A 非常兴奋。至少相对我们勾勒的那些替代。我不声称我们想过所有可能的计划。我们列出了 Plan S、Plan A 等等,但至少在我们看过的计划里,Plan A 对避免权力集中看起来非常好。唯一看起来也许更好的竞争者是 Plan S。也许如果你关掉所有 AI,对避免权力集中比 Plan A 更好。但如果你就是要造超人类 AI,那我认为 Plan A 是我所知的最不集中权力的做法。
The second thing is the transparency. One of the main ways in which I think people who control AI development can abuse their power is by having their AIs pursue their own agendas — specifically pursue agendas that are in the interest of the person who built the AIs. But do so in a way that’s maybe secret. Imagine if OpenAI announced that their AIs were going to be trying to sell you things and also trying to get you hooked on their product. Also they would be trying to convince you to vote for OpenAI’s preferred political candidate. Obviously, if this became public information, it would not work so well because people would stop using ChatGPT and they would be on guard against this type of persuasion when they were using ChatGPT. But if OpenAI does something like this and it’s secret — and it’s just a subtle influence campaign that nobody knows about except for some conspiracy theorists — then it’s going to have potentially a pretty significant effect. So the transparency about how the AIs are trained and what goals and values are being put into them is really good for preventing this type of abuse of power. And this is true whether it’s a CEO or whether it’s a president. They can still try that in conditions of total research transparency, but it’s so much harder than if they don’t have the total research transparency because people will see what they’re doing and then people can react. Then also the transparency just helps again with avoiding the monopolies because the transparency combined with the no intelligence explosions, buying time thing means that multiple companies can catch up. It really seems to me like — even if you didn’t care about loss of control at all, and you thought that the AIs were going to be very easily controlled — as long as you’re worried about concentration of power and AI dictatorships, you should be very excited by Plan A. At least compared to the alternatives that we’ve sketched out. I don’t claim that we’ve thought of all possible plans. We’ve laid out Plan S, Plan A, et cetera, but at least among the plans that we’ve looked at, Plan A seems really good for avoiding power concentration. The only contender that seems maybe better would be Plan S. Maybe if you just shut down all the AIs, that’s even better for avoiding power concentration than Plan A. But if you’re going to be building superhuman AIs, then I think Plan A is the least power-concentrating way to do it that I’m aware of.
03:03:50把 Plan A 推演 100 次Playing out the Plan A scenario 100 times
Luisa Rodriguez
你谈过桌面推演的一些结果。最常见的结局是什么?
You talked about some of the results of the tabletop exercises you’ve done. What are the most common outcomes from those?
Daniel Kokotajlo
我们总共做了大约 100 场。大多数是标准的《AI 2027》式推演:从字面《AI 2027》情景开始,或改到 2028、2029、2030 的版本。我们从他们离自动化 AI 研究还有几个月的时点开始。然后让他们做任何他们想做的,说:试着采取你现实中认为你这个角色在这种局面下会采取的行动。我们也做了大约 10 场左右的 Plan A 情景,一样,只是一开始我们默认规定:美中已经决定想做类似 Plan A 的事,已经非正式握过手。然后由他们决定是否真的做、是否敲定细节、敲出实际协议。但我们作为假设规定:他们已经表示有兴趣做某种看起来像 Plan A 的国际协议。
So we’ve done about 100 exercises total. Most of them have been our standard AI 2027-style exercise, where we start in either the literal AI 2027 scenario or a modified version that takes place in 2028 or 2029 or 2030. We started at the point where they’re a few months away from automating AI research. Then we just let them do whatever they want and say: try to take the actions that you realistically think your actor would take in this situation. Then we’ve also done a small amount of maybe about 10 or so of Plan A scenarios, which is like that except that we assume at the beginning, by default, we just state as an assumption that the US and China have already decided that they want to do something like Plan A, and they’ve already informally handshook on it. Then it’s up to them to decide if they’re actually going to do it and if they’re going to work out the details, and to hammer out the actual agreements. But we stipulate by assumption that they’ve expressed interest in doing some sort of international deal that looks something like Plan A.
《AI 2027》那些通常就像《AI 2027》,并非巧合,因为有些是我们做《AI 2027》研究过程的一部分。通常发生的是大量地缘政治紧张。美中之间有竞赛。各家美国 AI 公司之间也有竞赛。总统和美国 AI 公司之间也有权力斗争。很多其他国家一开始睡着,然后逐渐醒过来,意识到自己处境的严重性:即将被剥夺权力,可能被杀死。公众非常愤怒,但通常成不了什么事。推演里我们做六七个回合,大约一年过去,取决于走多快。到最后有超级智能 AI,世界正被非常激进、非常迅速地改造。通常我们落到:如果 AI 不对齐,它们可以轻易接管,因为人类一直在让它们自我改进,事实上还在鼓励它们自我改进,把越来越多的事情交给它们,好打败其他人。那就是某种默认结局。有时对人来说结局还行,因为扮演 AI 的人决定 AI 毕竟对齐了。于是没事。然后我们进入权力集中问题。但有时扮演 AI 的人决定 AI 不对齐,必须做些事才能让它们对齐。那些情况下他们常常就落到 AI 接管。
In the AI 2027 ones, it’s usually like AI 2027, not by coincidence, because some of these were done as part of our research process for making AI 2027. Usually what happens is there’s a lot of geopolitical tension. There’s a race between the US and China. There’s also a race between the various US AI companies. There’s also a power struggle between the president and the US AI companies. Lots of other countries are asleep at first, but then gradually wake up to the severity of the situation they’re in and how they’re about to be disempowered and possibly killed. The public is very angry, but usually doesn’t accomplish much. Over the course of the exercise we do six or seven turns and about a year or so goes by, depending on how fast we go through it. By the end there are superintelligent AIs and the world is being very aggressively and rapidly transformed. Usually we end up in a situation where if the AIs are misaligned, they could easily take over because humans have been letting them improve themselves, and in fact encouraging them to self-improve and putting them in charge of more and more things in order to beat each other, the other humans. So that’s kind of the default outcome. Sometimes it works out fine for people because the person playing the AIs, the AI player decided that the AIs were aligned after all. So it’s fine. And then we get into concentration of power issues. But then sometimes the person playing the AI decided that the AIs were misaligned and that things would have had to be done to make them aligned. Then in those cases they often just end up with AI takeover.
一个有趣的例子:有一次我们甚至处于这种局面:多家不同公司的 AI 在告诉任何愿意听的人,它们不认为自己能成功对齐下一代 AI——或者说它们认为风险很高,建议暂停——而它们的人类委托人、人类 CEO 在说:「不,走更快,我们必须赢,我们必须接受这个风险,因为如果我们不,那些人就会怎样怎样。」就是那种好笑的局面。我记得甚至有一场游戏,AI 在砂袋,它们不对齐,但搞不清怎么对齐未来世代的 AI——包括搞不清怎么对齐到它们自己。所以它们在砂袋,慢走 AI 研究,因为搞不清怎么让它对它们安全,更别说对人类安全。而人类在鞭它们:「走更快!」有很多疯狂局面。也有很多我前面提过的情况:当家人,比如各国总统,让事情变得非常疯狂,然后有一种 180 度时刻,往往是对某个具体事件的反应,比如 AI 逃出数据中心——他们说「哇,我们得全部关掉。」然后合作去做。通常又太晚。
One fun example was one time we were even in a situation where the AIs across multiple different companies were telling everyone who would listen that they didn’t think that they could successfully align the next generation of AIs — or that they thought the risk was high and they recommended a pause — and their human principals, the human CEOs, were saying, “No, go faster, we have to win, we have to accept this risk because if we don’t then the other guys, blah, blah.” So it was just kind of a funny situation. I think there was even one game where the AIs were sandbagging and they were misaligned, but they couldn’t figure out how to align the future generation AIs — which includes they couldn’t figure out how to align it to themselves. So they were sandbagging and slow walking on their AI research because they couldn’t figure out how to make it safe for them, much less safe for the humans. And the humans were whipping them like, “Go faster!” There’s lots of crazy situations. There’s also been various — I think I previously mentioned — lots of cases where the people in charge, like the presidents of the countries, let things get really crazy and then have a sort of 180 moment, often in response to some specific incident like an AI escaping from the data centre — where they’re like, “Whoa, we need to shut it all down.” Then they cooperate to do that. Again, usually too late.
一件有意思的事是,在够多的游戏里我觉得成了模式——也许三四场——游戏结束时的状态如下:已经有一份国际协议,关掉 AI 进展,再以安全、缓慢、透明的方式重建——类似 Plan A——而且实际已经执行。绝大多数数据中心已经关掉,现在有某种国际财团在敲如何推进的细节。同时美国和/或中国有一个秘密 AI 项目,有相对少量走私 GPU,在秘密里单方面更快推进,但 GPU 很少,没法接近 OpenAI 或 Anthropic 默认会走的速度。还有一个失控 AI 在互联网上跑,从各种笔记本集合挪到另一些笔记本集合,试图完全不被关掉,躲避正在找这类东西的警察。事实上,第一件事发生是因为第三件事。过去大约一百场里这发生过大概四次,所以很有意思。不幸我们时间不够,没法往前推演看它会怎么结束。但我觉得有意思的是我们多次落到那种状态。
One interesting thing that happens is, in some large enough number of games that I think it’s a pattern — like maybe like three or four games — the state this game ended in was as follows: There’s been an international agreement to shut down AI progress and then rebuild it in a safe, slow, transparent way — similar to Plan A — that’s actually been carried out. So the vast majority of the data centres have been shut down and now there’s some sort of international consortium that’s working out the details for how to proceed. Also the US and/or China have a covert AI project with some relatively small amount of smuggled GPUs that’s unilaterally proceeding in secret faster, but they have a very small amount of GPUs so they’re not able to go nearly as fast as OpenAI or Anthropic would have gone by default. Also, there’s a rogue AI running around on the internet moving from various collections of laptops to various other collections of laptops, trying to avoid being completely shut down and hide from the police that are going around looking for this sort of thing. In fact, the reason why there was the first thing was because of the third thing. I think this has happened like four times or something in the last hundred or so games that we’ve done, so it’s very interesting. Unfortunately we ran out of time, so we can’t play it out forward and see how it would end. But I just think it’s interesting that that happened several times, when we got to that sort of state.
Luisa Rodriguez
桌面游戏里事情看起来走得还行的时候,有没有什么往往会发生?
Are there any things that tend to happen in cases where things seem to be going well in the tabletop game?
Daniel Kokotajlo
Plan A 版本平均好得多。如果你从他们已经同意类似 Plan A 的假设开始,结果分布好得多。不一定很好。我们有过几场失败的 Plan A 情景,由于某种原因事情糟得可怕。最近有一场他们做了更蛮力的 Plan A 版本,没有完全研究透明,只是试图限制用于 AI 发展的算力,对那些算力在被用于什么没有洞察。可因为是更粗粒度的东西,他们只是不断把算力限制得很厉害,于是几年下来对齐进展没那么多,因为没有那么多人真正能和 AI 互动。AI 也不能做自动化对齐研究。所以过了几年局面没改善多少。然后新政府进来,「走吧!」松开刹车。然后他们很快冲向超级智能。然后缩放到超级智能时出了问题。然后不对齐的 AI 接管。那种事大概发生过两次。我们还有一场游戏,围绕 AI 有非常激烈的权力斗争,AI 实际上成功对齐了,但美国总统当上了独裁者,然后和习近平做了一笔交易,因为就他们俩,可以把世界瓜分。欧洲试图阻止,很多其他力量试图阻止,美国国内很多人也试图阻止,但我觉得他们不太成功。那是对齐已经解决、但权力集中成了问题的例子。
The Plan A versions have gone much better on average. If you start with that assumption that they’ve agreed to something like Plan A, the distribution of outcomes is much better. Not necessarily great. We’ve had a couple failed Plan A scenarios where things go horribly wrong for one reason or another. We recently had one where they did a more brute-force version of Plan A without the total research transparency, where they just tried to restrict how much compute was used for AI development, without any insight into what that compute was being used for. But because it was the more coarse-grained thing, they just kept restricting the amount of compute by a lot, and so they didn’t make that much alignment progress over the course of several years because they didn’t have that many people actually able to interact with the AIs. Also the AIs weren’t able to do automated alignment research. So a couple years in they were not that much improved in their situation. Then a new administration came in and was like, “Let’s go!” and let off the brakes. Then they went really fast to superintelligence. Then something went wrong in the scale up to superintelligence. Then the misaligned AIs take over. I think that sort of thing happened roughly twice. I think we also had a game where there were some really intense power struggles over the AIs, and the AIs were in fact successfully aligned, but the US president managed to become dictator and then cut a deal with Xi Jinping because it’s just the two of them, so they can split up the world between them. Europe tried to stop this, and lots of other powers tried to stop this, and lots of people in the US tried to stop this, but I don’t think they were very successful. So that was a case of the alignment having been solved, but then the concentration of power stuff being a problem.
03:13:32丹尼尔会怎么改 Plan AHow Daniel would revise Plan A
Luisa Rodriguez
写这个情景时,你和共同作者之间必须解决的最大分歧是什么?
What were some of the biggest cruxes with your coauthors that you had to resolve when putting the scenario together?
Daniel Kokotajlo
我大力主张完全研究透明,其他人觉得「有很好,但大概用更正常的审计也能过」。而我是:不,不——我不信任普通审计师。我们需要更强的东西。所以有那个。另一件是,协议前奏里,我们有这个问题:美国该先从国内监管开始再和中国做协议,还是美国直接从这份对华协议开始?那其实是我改了主意的事。我有点希望我们描绘成:先是美国在国内成功监管 AI,然后问中国,「嘿,你们也该这样做,我们愿意做让步让你们做。」反馈来自各种各样的人。它既更说得通,也是更好的策略。如果你已经表明愿意做昂贵的事来监管自己的产业,然后再请他们对自己的产业做同样的事,你更可能和中国有严肃讨论,而不是你正以最快速度冲向超级智能,却在峰会上见他们、告诉他们也许你想做点别的。如果你已经开始做那件事,它会感觉更真实,也更被证明过。
I was a big proponent for the total research transparency and other people were like, “It’s nice to have, but probably we can get by with more normal auditing.” Whereas I’m like, no, no — I don’t trust the normal auditors. We need something stronger. So there was that. I think another thing is, in the run-up to the deal, we had this question of should the US start with domestic regulation and then do a deal with China, or should the US just start with this deal with China? That’s actually something I’ve changed my mind about. I kind of wish that we had depicted it as first the US regulates AI successfully domestically, and then asks China, “Hey, you should do this too, and we’re willing to make concessions to get you to do it.” Feedback from a variety of people, mostly. It’s both more plausible and a better strategy. I think that you’re more likely to have serious discussions with China if you’ve already shown that you’re willing to do costly things to regulate your own industry, and then you’re asking them to do the same things to regulate their own industry, than if you are racing as fast as you can towards superintelligence but then meeting them at a summit and telling them how maybe you’d like to do something else. It’s going to feel more real, and it’s more proven as a thing, if you’re already starting to do the thing.
你也会立刻拿到好处。比如,我现在的中位估计是 AI takeoff 发生在 2028 年。50% 的可能到那时已经在发生。或者说到 2028 年底,AI takeoff 已经发生,AI 研发的完全自动化已经发生。所以是我的中位估计。在这个情景里,他们在 2029 年用这份大型国际协议、监视人员来回飞、检查等等开始 Plan A。那是那一刻的前一年。但因为不确定性,也许你拥有的时间比你以为的少。也许你正在和中国谈判时,这些公司里的某一家做出了突破,然后你就开赛了,局面糟得多——所以更好的是先开始做好的那件事。代价是它某种程度上帮了中国一点。如果你开始以认真的方式监管自己的产业,那监管最好的版本大概会阻止他们以最高速度走。于是会稍稍让中国追上一点。不过我仍觉得该这么走,因为你也可以同时开始和中国谈话:「看,我们在做这件事。它字面上在帮你们,因为它在让我们变慢。接下来四周,我们想谈一个你们做类似事情的计划。」Plan A 不是「中国打败美国」的情景。它是一份协议。美国例如在算力上全程保持领先。
Also, you then get the immediate benefits. For example, my median estimate right now is that AI takeoff happens in 2028. 50% chance that it’s happening by then. Or by the end of 2028 AI takeoff has happened, full automation of AI R&D has happened. So my median estimate. In this scenario, they start Plan A with this big international deal and all the monitors flying back and forth and inspections in 2029. A year before that moment. But because of uncertainty, maybe you have less time than you think. Maybe while you’re in negotiations with China, some breakthroughs are made inside one of these companies and then you’re off to the races and now things are much worse — so it’s better to just get started doing the good thing first. I think that the cost of that is that it sort of helps China a bit. If you start regulating your own industry in a serious way, then the best versions of that regulation would probably stop them from going at maximum speed. So then that would slightly cause China to catch up a little bit. Although I think that still it’s the way to go, because you can also just simultaneously start the conversations with China and be like: “Look, we’re doing this thing. It’s literally helping you out because it’s slowing us down. In the next four weeks, we would like to negotiate a plan for how you’re going to do something similar.” Notably, Plan A is not a “China beats the US” scenario. It’s a deal. The US maintains its lead in compute, for example, throughout.
Luisa Rodriguez
Plan A 里有没有你现在希望描绘得不同的其他东西?你觉得最不可能发生的是哪一面?有没有已知的未知,如果我们更清楚,会大幅改变你会建议的计划?
Are there any other things that you now wish you’d depicted differently in Plan A? Is there an aspect of Plan A that you feel is least likely to happen? Are there any known unknowns you can think of that — if we got more clarity about them — would dramatically change the plan you’d recommend?
Daniel Kokotajlo
很多人被那个疯狂的超人类主义结尾吓坏了。我们甚至还没谈结尾。我有一部分觉得也许我们根本不该谈那些。但也有一部分觉得不,那样好,因为人们被吓到是对的——他们需要去搏斗遥远的未来长什么样,以及先进 AI 的可能性。所以如果他们不喜欢,希望他们学到。希望他们不要把我们当报信人杀掉,而是更认真地想他们到底想从所有这些 AI 进展里要什么,拿出他们更喜欢的东西。有一整堆看起来不太可能发生的事。我觉得任何和中国的重大协议都看起来不太可能。任何让公司显著慢于最高速度的事都看起来不太可能。然后显然完全研究透明看起来不太可能。我不确定哪个最不可能,但大概会是完全研究透明。我仍觉得它好,所以我们在主张它。
A bunch of people are really freaked out by the crazy transhumanist ending. Which we haven’t even talked about yet. But part of me thinks maybe we just shouldn’t have talked about all that stuff. But part of me thinks, no, it was good because people are right to be freaked out — and they need to grapple with what the far future looks like. And what the possibilities of advanced AI are. So if they don’t like it, well, hopefully they learn. Hopefully they don’t shoot us as the messenger, and instead they think more seriously about what they actually want out of all this AI progress and come up with something that they like more. There’s a whole bunch of things that don’t seem likely to happen. I think any sort of major deal with China seems unlikely. Any sort of making the companies go significantly slower than maximum speed seems unlikely. Then obviously the total research transparency seems unlikely. I’m not sure which of those would be least likely, but probably it would be the total research transparency, I think. But I still think it’s good, so that’s what we’re advocating for.
有很多。Thomas 在某处做了一张不错的图,在什么条件下他会主张各种计划。比如有些条件下我们会主张 Plan S 而不是 Plan A。比如前面提过的,如果我们确信实际上我们能做成相当稳定、持续几十年的协议?那会是做看起来更像 Plan S 的东西的强论证。反过来,如果我们确信要有任何类似 10 年减速的东西超级超级难,除非交叉手指希望中共不要接管世界?因为他们完全可以,因为你只是在信任他们。在那些看起来我们不会信任他们——他们也不会信任我们——所以需要走更快的条件下。最让我沮丧的误解是这个想法:Plan A 是在提议一个全球监管者、在集中权力。不,我们不是在提议全球监管者,也不是在集中权力。我们花了很多心思在未来 AI 权力集中情景上:怎么防止它们,哪些是会影响极端权力集中概率的关键指标、关键杠杆。避免 AI 垄断看起来是避免权力集中的非常重要的杠杆,所以我们大量设计是为了避免 AI 垄断。透明看起来也是非常重要的杠杆,所以我们在透明上走得很猛。所以挺沮丧的是人们——其中许多人甚至没读过我们的东西——说「他们在把权力集中到一个全球监管者里」。
There are many. I think that Thomas made this nice diagram somewhere of under what conditions he would advocate for the various plans. For example, there are conditions under which we would advocate for Plan S instead of Plan A. For example, as previously mentioned, what if we became convinced that actually we can make quite stable deals that last decades? Then I think that would be a strong argument for doing something that looks a lot more like Plan S. And contrariwise, what if we became convinced that it was super, super hard to have anything like a 10-year slowdown without having to just cross your fingers and hope that the CCP doesn’t take over the world? Because they totally could, because you’re just trusting them. Under those conditions where it doesn’t seem like we’re going to trust them — and they’re not going to trust us — so we need to do something faster. I think probably the one that frustrates me most is this idea that Plan A was — the one I already mentioned — that we’re proposing a global regulator, concentrating power. No, we’re not proposing a global regulator and we’re not concentrating power. There’s actually very good reasons why we put a lot of thought into future power concentration scenarios with AI and how can you prevent them and what are the key metrics, the key levers that would affect the probability of extreme power concentration. Avoiding monopolies on AI seems like a really important lever for avoiding power concentration, so we did a lot of our designing to try to avoid monopolies on AI. Then transparency also seems like a really important lever, so we went really hard on transparency. So it’s kind of frustrating that people — many of whom haven’t even read our thing — say, “They’re concentrating the power in a global regulator.”
关于暂停有一个无辜得多的误解。这个无辜,是因为它确实有点复杂、有点让人糊涂。某种意义上我们在主张暂停 AI 发展,某种意义上我们又完全没有。如果你读我们的情景,读我们在提议什么,以及我们认为如果提议被实施会发生什么,那是一场十年里由 AI 造成的疯狂社会改造。那在很多方面完全不是暂停。但真相有点复杂。我们在主张比你能以最高速度走的更慢。我们说不要做这些疯狂的智能爆炸。所以那是走慢。但我们说你该继续发展 AI、部署它、扩散它。我们说,至少按我们的计算,那会导致像 GDP 每年翻倍这种事。然后我们实际谈的轨迹更锯齿:2029 年有一次字面上的 AI 发展暂停,大约六个月,同时他们在把核查基础设施装起来。然后以谨慎节奏继续。然后 2030 年代晚期,当他们撞上能控制的极限时,又有一次字面上的暂停。等他们对齐问题解决了,再继续。所以某种意义上有两次暂停,但都是暂时的。
I think there’s a much more innocent one about the pause. Basically this one is innocent because it’s just actually kind of complicated and confusing. In some sense we are advocating for a pause on AI development, but in some sense we are very much not. If you read our scenario, and you read what we are proposing, and what we think would happen if our proposals were implemented, it’s a crazy transformation of society by AI over the course of 10 years. That’s very much not a pause in a bunch of ways. But the truth is it’s kind of complicated. We are advocating for going slower than you could go at maximum speed. We’re saying don’t do these crazy intelligence explosions. So that’s going slow. But we are saying you should continue developing AI and deploying it and diffusing it. We’re saying that, yeah, according to our calculations at least, that’s going to lead to things like GDP doubling every year if you’re doing that. Then also the actual trajectory that we talk about is more jagged, where there is a literal pause on AI development for like six months in 2029 while they’re getting the verification infrastructure set up. Then it continues at a cautious pace. Then there’s another literal pause in the late 2030s when they run up against the limits of what they can control. Then when they solve the alignment problems, they proceed again. So in some sense there’s two pauses, but they’re temporary.
03:23:02哪些是建议、哪些是预测Which parts of Plan A are recommendations vs predictions?
Luisa Rodriguez
有点难分清情景里哪些被视为理想,哪些是对可行性的让步。Plan A 里各有多少?
It’s a bit hard to tell what in the scenario is considered ideal vs a concession to feasibility. How much of each is there in Plan A?
Daniel Kokotajlo
我对这个有点抱歉。发布前我们就发现了这个问题,做了一些事来处理。但我们本可以更清楚,也许如果我们决定推迟发布,可以在这里做得更多。我们有一份补充材料谈这个,叫「Plan A 假设」。它谈这个问题,试图梳理什么是建议、什么是预测。高层次的说法是:除了我们谈的那些关键建议,一切都是预测。你可以去读那份补充,看我们认为的主要建议——那些显然是建议,不是预测。然后你该默认假设:其他事情只是对我们主要建议被实施后会发生什么的预测。那是高层次答案。然后有一些灰色地带。
Yeah, I feel a bit bad about this. We had identified this problem before launch and done some things to address it. But we could have been more clear, I guess, and perhaps if we had decided to delay the launch, we could have done more here. But we have a supplement that talks about it, called “Plan A assumptions.” I think that talks about this question and tries to canvass what’s the recommendation and what’s a prediction. The high-level thing is everything’s a prediction except for the key recommendations that we talk about, basically. You can go read that supplement and see the things that we consider our main recommendations — those are obviously recommendations, not predictions. Then you should sort of, by default, assume that things are just a prediction about what would happen if our main recommendations were implemented. That’s the high-level answer. Then there’s a few grey-area cases.
我们不想做那种基本上是「永远听我们的、做我们说的一切」的建议,因为那在政治上不可能。有点傲慢。相反,我们想做至少能想象真的会被做的建议。我们在文中谈那种推理和公共话语如何演变,以及为什么它让 Plan A 和 Plan S 成为政客也许真的会去做的桌上选项。我们想走那些潜在在可能范围内、作为严肃之事的东西。但除此之外,我们想选真正最好的,而不是只选更可能的。灰色地带太多,说不完。但可以举个例子。我们谈公民分红,谈它如何从美国公民的分红开始,然后作为某种对外援助扩展到所有人类。但他们给外国人的分红比给美国公民的少。那更是预测而不是建议。但它有点灰色,因为我们显然觉得有公民分红是好事,也觉得有对外援助是好事。但那个分红对对外援助的精确比例是我们建议的吗?不,我们会希望对外援助更多,尤其是长期。我觉得长期我们希望它实际上平等。但那某种意义上是对现实的让步。我们问自己:「我们的建议是做带有某种对外援助成分的公民分红」,然后是「现实地,假如他们做了类似的事,对外援助成分大概会有多少?」大概他们给外国人的会比给美国公民的少。所以我猜我们就那样写。你明白我在说什么。那是一种灰色地带:有些部分是建议,但并非全部是我们的建议。如果我们做主,我们会做稍有不同的事。
Yeah, basically. I think maybe one way of putting it is we didn’t want to make some recommendations that were basically of the form, “Listen to us and do everything we say forever,” because that’s not politically possible. That’s a bit arrogant. Instead, we wanted to make recommendations that we could at least imagine being actually done. We talk in the piece about the sort of reasoning and the public discourse and how it evolves and why it makes things like Plan A and Plan S on the table as things that the politicians might actually go for. We wanted to go for things that were within the realm of possibility potentially, and as serious things. But then other than that, we wanted to pick the actually best ones rather than just the more likely ones. Too many grey areas to go over. But I could give an example. So we talk about the citizens’ dividend, and we talk about how it starts off with a dividend for US citizens, but then they extend it as a sort of foreign aid to all human beings. But then they give less dividend to foreigners than they do to US citizens. That’s more of a prediction than a recommendation. But it’s kind of a little bit of a grey area because we obviously think it’s good to have a citizens’ dividend and we think it’s also good for there to be foreign aid. But is that exact ratio of dividend to foreign aid what we recommend? No, we would want there to be more foreign aid than that, especially in the long run. I think in the long run we want it to be just actually equal. But that was sort of a concession to reality in some sense. We asked ourselves, “Our recommendation is to do a citizens’ dividend with some foreign aid component,” and then it’s like, “Realistically, how much foreign aid component would probably happen supposing that they did something like this?” Probably they would give less to the foreigners than to the US citizens. So I guess that’s what we’ll write. That’s an example of a sort of grey area where it’s like there’s parts of it that are a recommendation, but not all of it is our recommendation. If we were in charge, we would do something somewhat different.
03:26:52最可能的失败模式Plan A’s likeliest failure mode
Luisa Rodriguez
如果你想象 Plan A 失败,你觉得最可能的那条因果链是什么:什么导致它失败,失败之后跟着发生什么?
If you picture Plan A failing, what do you think is the most likely chain of events that causes it to fail and then follows from the failing?
Daniel Kokotajlo
我们在文中谈了很多这个。我们认为 Plan A 在被实施之后最可能失败的方式是:各 AI 产业的监管者干得很糟,批准了实际上危险、但他们错误地认为不危险的 AI 的创造和部署。它可能发生在任何时候。最可能相对较早发生。我觉得协议运转越久,科学共同体就越有时间去搏斗这个局面,监管者就越有时间提能,尤其是多亏透明。所以我尤其担心这种失败模式发生在协议相对早期。我们有一条小情景枝,你可以去读它可能长什么样。
We talk about this a bunch in the piece. The most likely way that we think Plan A could fail after having been implemented is that the regulators of the various AI industries do a bad job and approve the creation and deployment of AIs that are in fact dangerous, but they wrongly think that’s not dangerous. It could happen at any time. It’s most likely to happen relatively early. I think that the longer that the deal has been in operation, the more time the scientific community has to grapple with the situation and the more time the regulators have to skill up, especially thanks to the transparency. Basically, I’m most especially worried about this failure mode happening relatively early into the deal. We have a little scenario branch that you can go read of what it might look like for this to happen.
第二让人担心的失败模式我觉得会是协议破裂。基本上会有很多喊。我们对此是现实主义者。我们试图对此现实。我们不是在提议一个单一的全球 AI 发展权威。有些人误以为那是我们在提议的。但如果你读我们的东西,那不是我们在提议的。相反,我们提议各国监管自己的 AI 产业,但由于透明,他们能看见谁在做什么、谁在监管什么。如果有人对别人在做的事有意见,他们能立刻看见,然后可以谈,然后可以互相喊、讨价还价、威胁、恳求,试图让他们停下那件吓到他们的事。但这会是一个messy的过程。希望最终它会演变成更形式化、更高效、有很多技术官僚专家做判断的过程。但至少一开始我们想更现实政治:各国同意的根本之事是透明,好看见发生了什么,但除此之外,他们就一事一议,争论什么没事、什么有事。他们各自做自己的监管,但试图根据其他国家要他们做什么、以及其他国家实际上在做什么,来调整自己的监管。
The second most concerning failure mode I think would be the deal breaking down. Basically, there’s going to be a lot of yelling. We are realists about this. We’re trying to be realistic about it. We are not proposing a single global authority for AI development. Some people mistakenly think that’s what we’re proposing. But if you read our thing, that’s not what we’re proposing. Instead, we’re proposing that each country regulates its own AI industry, but that because of the transparency they can see who’s doing what and who is regulating what. If people have a problem with what someone else is doing, they can immediately see it and then they can talk about it and then they can yell at each other, bargain, threaten, plead, and try to get them to stop doing the thing that’s scaring them. But this is going to be a messy process. Hopefully, eventually it would evolve into a more formalised process that’s more efficient and has lots of technocratic experts making judgement calls. But at least at first we wanted to be more realpolitik about it and basically just be like: the fundamental thing that the countries have agreed on is the transparency so they can see what’s happening, but then beyond that, they’re just taking things on a case-by-case basis and arguing about what’s fine and what’s not fine. They’re each doing their own regulation, but then they’re trying to adjust their regulation in response to what other countries want them to do and in response to what other countries are in fact doing.
所以那可能走错。可能紧张太高,他们真的没法对事情达成一致,或者有别的事让紧张升高。也许有一场不是 AI 引起、但正在发生的台湾战争。然后作为战争的副作用,他们停止对 AI 项目的所有这些透明。有一整堆理由协议可能破裂、他们可能停止互相透明。然后如果他们停止互相透明,他们会害怕又要冲向超级智能,意味着他们大概会再次开始冲向超级智能,意味着我们又回到 AI 竞赛局面。只不过大概更快,因为他们有更多算力。意味着他们大概会摧毁算力,因为那是我提过的协议原则之一。于是由于算力可摧毁这件事,那会是一整场messy局面。我们认为至少不会比一开始没做协议更糟。由于种种原因,也许稍好。比如这些年对 AI 的科学和一般理解会前进,所以从对齐角度看会比完全没做协议更好。总体上也会有更多人醒来看到 AI 的效应,准备得更好,但如果协议破裂、我们再次开赛,仍会相当messy、相当糟。
So anyhow, that could go wrong. It could be that tensions get too high and they just really can’t agree on things, or maybe there’s some other thing going on that causes tensions to be high. Maybe there’s a war over Taiwan, for example, that wasn’t caused by AI but is happening. Then as a side effect of the war, they stop doing all this transparency about their AI programmes. There’s a whole host of reasons why the deal could break down and why they could stop being transparent with each other. Then if they stop being transparent with each other, they’re going to be afraid that they’re going to be racing to superintelligence again, which means they’re probably going to start racing to superintelligence again, which means now we’re in the AI race situation again. Except it’s probably going even faster because they have more compute. Which means that probably they would destroy the compute because that’s one of the principles of the deal that I mentioned. So that would be a whole messy situation because of the compute-destroyability thing. We think it would be at least not worse than if they hadn’t made the deal in the first place. And for a variety of reasons, maybe somewhat better. For example, the amount of science and general understanding about AI would have advanced in the intervening years, so we’d be better off from an alignment perspective than we would if we just hadn’t done the deal in the first place. And in general, more people would have woken up to the effects of AI and would be more prepared, but it would still be quite messy and quite bad if the deal broke down and we started racing again.
03:31:16美国现在能做什么What the US can do now to make Plan A possible
Luisa Rodriguez
我想花几分钟谈 2026、2027 年需要发生的具体技术、制度工作,好让 Plan A 更可能。你已经谈过美国可以在国内做的一些事,会以让安全更容易的方式减慢 AI 进展。但那离我们现在也许还有几步。你希望看到美国政府在没有任何国际协议的情况下,立刻采取哪些字面上的下一步,好让类似 Plan A 的东西以后更可能?
I want to move on and spend a few minutes talking about concrete, technical, institutional work that needs to happen in 2026, 2027, to make Plan A more possible. You’ve already talked about some things that the US could do domestically that would be good for slowing down AI progress in a way that will make safety easier. But it feels like that’s maybe a few steps away from where we are. What are the literal next steps that you’d like to see the US government do, without any international agreement, to make something like Plan A more likely later?
Daniel Kokotajlo
我的答案在情景的 2027 年,我们的渐进式 AI 政策心愿单。限制 AI 研发预算那件事有点雄心,但我觉得实际上在可能范围内。我觉得它比华盛顿的人会预期的更可行。我其实觉得 AI 公司的研究者里,有人对做类似的事有兴趣。所以我其实觉得我们可以立刻开始做那个。其他的:要么执行要么废掉出口管制。如果我们要有出口管制,那就该被执行。AI 算力追踪似乎很好,好告诉情报界这是优先事项,他们该试着找出芯片在哪,看有没有正在组装的秘密项目。总体上,提高政府的 AI 能力显然非常重要。政府该招募 AI 专家,在政府内组建能理解 AI、能对模型做评估、能做安全案例并评估安全案例、能对这一切走向何方做预测的机构。那看起来真的很重要。
My answer to this is in the scenario in 2027, our incremental AI policy wishlist. I think that the limit to AI R&D budgets thing is somewhat ambitious, but I think it’s within the realm of possibility actually. I think it’s more feasible than I think people in DC would expect. I actually think there’s some interest among the researchers at the AI companies to do something like this. So I actually think that we could just get started on that immediately. Other things. Either enforce or repeal the export controls. If we’re going to have export controls, that should be enforced. AI compute tracking seems good to tell the intelligence community that this is a priority and that they should be trying to find out where the chips are, and see if there’s any covert projects being assembled. I think that, in general, improving the government AI capacity is obviously very important. The government should be recruiting AI experts and forming agencies within the government that can understand AI and can run evaluations on models and can make safety cases and evaluate safety cases, can make forecasts about where all this is headed. Yeah, that seems really important.
还有更一般的透明。比如可以有举报人保护要求。可以要求公司公布模型规格或宪法,以及在其他方面向公众提供更多关于他们如何训练 AI 的信息——然后向政府审计师提供更多信息,好让政府能检查他们没有试图把任何秘密议程塞进 AI,比如隐藏偏向。这类事。我觉得这只是表面。这类事有一张巨大的清单。然后对每一件这类事,又有一张更具体、更落地的巨大清单。有一些这样的清单。我觉得我们有一两篇博文谈这个。当然我们网站上也说了一些,但接下来几周或几个月我们大概会做的一件事,就是把更多这类想法写下来发表。然后除了我们,也有其他人在推动、主张这些事。
I think also just transparency more generally. For example, there could be requirements for whistleblower protections. There could be requirements that companies publish model specs or constitutions, and otherwise give more information about how they’re training their AIs to the public — and then even more information to government auditors, so that governments can check that they’re not trying to put any secret agendas into their AIs, for example, or hidden biases. Yeah, things like this. I think actually this is just scratching the surface. I think there’s a huge list of things like this. Then for each thing like this, there’s a huge list of more specific, concrete things that could be done. There are some lists like this. I think we have a blog post or two about this. Then of course on our website we say some things, but one of the things we’ll probably do in the next few weeks or months is write up more ideas like this and publish them. Then there’s other people besides us who’ve also been pushing and advocating for things.
Luisa Rodriguez
Plan A 需要的核查技术目前还没有规模化存在——
So Plan A requires verification technology that doesn’t yet exist at scale—
Daniel Kokotajlo
那不是真的。我不会说它需要那种技术。我觉得如果你有那种技术,成本会低得多。我会这样说:如果我们想现在就实施 Plan A,美中会说,「好,我们要派真人去所有数据中心,把手放在 GPU 上,核实它们是冷的、关掉的。」那是我们今天就能做的。我们可以拔掉机器,然后核实机器在该在的地方、是关掉的。那不需要技术。当然问题是成本很高。意味着所有这些经济价值没有发生,因为 GPU 关着而不是在服务客户。但如果你想启动,你可以今天就做,然后立刻开始造新的、会更透明、上面有监测设备把活动发到互联网上的数据中心。你可以今天开始建那个,现有数据中心在你装起来期间先关着。用某种突击计划,大概要六到十八个月让所有那些新东西运转起来。然后你可以在新的透明方式下、用新的透明可摧毁数据中心,再继续 AI 发展。你可以现在就开始,但会因为那个而昂贵。
That’s not true. I wouldn’t say it requires that technology. I think it’s much less costly if you have the technology. The way I would put it is: if we wanted to implement Plan A right now, the US and China would say, “OK, we’re going to send physical humans to all the data centres to put their hands on the GPUs and verify that they are cold and off.” That’s something we can do today. We can unplug the machines and then verify that the machines are where they’re supposed to be and that they are off. No technology required for that. Problem with that, of course, is it’s very costly. It means that all this economic value is not happening because the GPUs are off instead of serving customers. But you could do it if you wanted to get that going, you could do it today and then you could immediately start creating the new data centres that are going to be more transparent and that have the monitoring devices on them to publish the activity to the internet. You could start building that today and have the existing data centres just off while you were getting that set up. It would probably take, with some sort of crash programme, six months to 18 months to get all that new stuff working. Then you could proceed with AI development again in the new transparent way, with the new transparent destructible data centres. You could get started right now, but it would be costly because of that.
如果已经造好监测设备和只推理改装套件,会很好,这样你可以让当前数据中心继续运转、服务客户,迅速改装成不能做大训练,而不真正打断运转。然后你建做训练的新数据中心。那就是我们情景里发生的。事实上,如果你比那更有远见,你可以在没有任何中断的情况下做这件事。你可以只是把它变成新数据中心建设的要求,必须符合新系统。然后过几年,数据中心默认就会是那样。我们对此当然不确定,但我们的估计是:把所有初始硬件开发和制造出来,会是几十亿美元量级。比如只推理改装。那让你起步。意味着你现在已经开始有意做 Plan A。然后在持续基础上,你在建、做得完全透明的新数据中心,也许每座新数据中心比默认成本大约多 1% 或 0.1%。所以是成本,但我觉得非常值得。
It would be nice to have built already the monitoring devices and the inference-only retrofitting kits, so that you could allow the current data centres to keep operating and serving customers, and just quickly retrofit them so that they can’t do big training runs — without really interrupting their operation. Then you build the new data centres that do the training. That’s what happens in our scenario. In fact, if you had even more foresight than that, you could do this without any disruption. You could just make this a requirement for new data centre construction, that they be compliant with the new system. Then after a few years it would just be the way that data centres were by default. Obviously we’re uncertain about this, but our estimate is that it would be single-digit billions to get all the initial hardware developed and manufactured. For example, the inference-only retrofitting. That gets you off the ground. It means now you’ve started off intent on doing Plan A. Then on an ongoing basis, the new data centres that you’re constructing and making totally transparent, maybe it costs something like 1% or 0.1% more for each new data centre compared to their default cost. So it’s a cost, but it’s well worth it, I think.
我会说有相关技术技能的人,比如懂硬件的人,有时也包括软件,该试着造这些设备、做原型。公司该往这上面砸钱,成立部门去做只推理改装套件,去做有这些属性的不同类型芯片。然后政府当然该鼓励这类事。要么砸资金,比如拨款,要么基本上只是说:「嘿,我们未来想做类似的事。有可能我们未来会要求数据中心这样。」只是说那个。我觉得如果政府说了,也许会鼓励一些公司拨一些资源。还有一件更花哨的——因为我们有完全研究透明,所以谈得没那么多,但如果你不做研究透明,它会非常有价值——就是保护隐私的审计。想象美国公司数据中心上的所有活动内部可见。公司能看见那活动里在发生什么。然后中国审计师带着一台设备出现,上面有一些中国 AI,然后他们插上,在所有活动上爬一遍,看一遍,然后汇报:规则有没有被遵守,这里有没有违规?然后他们被删除,设备被销毁,所以他们没法外泄任何秘密。他们能做的只是:「是或否,规则有没有被违反?」然后当然我们对中国做同样的事。要有那种安排,你需要一块现在还不真正存在的花哨技术——但如果我们把它建起来,也许能存在。那会非常有价值,因为它能让我们做这类审计,精确拿到我们想要的信息,而没有比那更多的信息泄漏。但得有人去建这一切,去给这一切去风险。
I would say that people with the relevant technical skills, people who understand hardware, for example, and in some cases software, should be trying to build these devices and make prototypes. And companies should be throwing money at this and spinning up divisions to make inference-only retrofit kits, and make different types of chips that have these properties. Then governments, of course, should just be encouraging this sort of thing. Either by throwing funding at it, like grants, or by basically just saying, “Hey, we want to be doing something like this in the future. There’s a chance that we might require this of data centres in the future.” Just saying that. I think if it was said by the government that might encourage some companies to allocate some resources to it. I think there’s another thing which is fancier — which I don’t think we talk about as much because we have total research transparency, but which could be really valuable if you’re not going to do research transparency — which is privacy-preserving auditing. Imagine a situation where all the activity on a US company’s data centre is visible internally. The company can see what’s going on in that activity. Then Chinese auditors show up with a device on which there are some Chinese AIs, and then they plug in and crawl around over all the activity and they look at it all and then they report back: are the rules being followed or is there a violation here? Then they’re deleted and the device is destroyed, so they weren’t able to exfiltrate any secrets. All they were able to do is just, “Yes or no, are rules being violated?” Then, of course, we do the same thing over to China. In order to have that sort of setup, you need to have a fancy piece of technology that doesn’t really exist yet — but maybe could exist if we built it up. That could be really valuable because it would allow us to do this sort of auditing and get exactly the information that we want, without any more information than that leaking. But someone needs to build all of that, and derisk all of it.
Luisa Rodriguez
情景假设实验室运转在安全等级 5,也就是能抗民族国家的网络安全。现在他们没有。实验室要到那里需要发生什么?
The scenario assumes that labs operate at Security Level 5, which is nation-state-resistant cybersecurity. Right now they don’t. What needs to happen for labs to get there?
Daniel Kokotajlo
对,那是另一件事。前面我提过,在某些方面,完全研究透明是给中国的礼物,因为它直接和他们分享算法。好吧,他们大概反正在拿算法,因为现在安全不太好。此刻那甚至不是那么大的让步。但显然我们觉得更多安全更好。那会是和新数据中心一起到来的事情之一。如果你要对这类事认真——并且要求有以透明方式建造的新数据中心——除了我们认为新数据中心该有的透明要求,你也可以给它们加上安全要求。你可以让即便民族国家要外泄权重也极其困难,事实上不可能。一个机制只是设带宽上限,于是权重甚至不可能通过信息能进出数据中心的那唯一一根线离开数据中心——因为权重太大,过不了那根线。那是你可以做的一个例子。但也有一整套其他最佳实践,你完全也该做。再一次,在我们的 Plan A 情景里,他们先做一次暂时暂停,停下所有新训练,把现有数据中心改装成只推理,同时建会安全得多、也透明得多、也在这些可摧毁地点的新数据中心。那需要时间。但用突击计划,我们觉得可以在六个月到一年左右做成。如果我们采取一大堆步骤,我们已经知道需要哪些步骤,如果实施了我们就会到——基本上,我觉得是。
Oh yeah, that’s another thing. Previously I mentioned that, in some ways, the total research transparency is a gift to China because it’s sharing the algorithms directly with them. Well, they’re probably getting the algorithms anyway because security is not very good right now. It’s not even that big of a concession at the moment. But obviously we think more security is better. Well, that’s one of the things that comes along with the new data centres. If you’re going to be serious about this sort of thing — and you’re requiring that there be new data centres that are built in a transparent way — in addition to the transparency requirements that we think the new data centres should have, you can also add on security requirements to them. You can make it so that it’s extremely difficult, in fact impossible, for even a nation state to exfiltrate the weights, for example. One mechanism for this is just having a bandwidth limit, so that it’s not even possible for the weights to leave the data centre through the only cable through which information can leave and exit the data centre — because the weights are too big to leave through that cable. That’s an example of something you could do. But there’s a whole bunch of other best practices that you should totally do as well. Again, in our Plan A scenario, they first do a temporary pause where they stop all new training runs and they refit the existing data centres to be inference only while they build the new data centres that are going to be much more secure and also much more transparent and also in these locations where they’re destroyable. And that takes time. But with a crash programme, we think it can be done in six months to a year, or something like that. Basically, I think. I think part of what happens in our scenario is that they’re doing things last minute. They had prepped some of these things in advance, but if they had decided to implement Plan A in 2027 instead of in 2029, then the process would have been much more smooth. Naturally the data centres being built in 2029 are built to the new code, so that’s just how it’s going.
03:43:05AI 2027 现在站得住吗How AI 2027 is holding up
Luisa Rodriguez
我还有一个问题。相对你一两年前的预期,你觉得事情走得怎么样?对齐工作,各国政府把 AI 风险当得多认真,就这一大范围的事。
I just have one more question for you. Relative to your expectations from one to two years ago, how do you think things are going? I guess alignment work, how seriously various governments take AI risk, just a broad range of things.
Daniel Kokotajlo
不幸的是,事情大致按我预期在走。你仍可以读《AI 2027》,它看起来仍像:对,我们某种程度上正走在那条路上。我以前会说对齐局面比我预期的好,治理局面比我预期的差。但那是我一年前或两年前会说的,现在我几乎说相反的话。相对一两年前,我会说治理局面比预期好,对齐局面比预期稍差。特别是,Hugging Face 失控 AI 事件是比我预期这个时点会发生的更过分的不对齐例子。你读《AI 2027》就能看出来,比如我们谈过随时间的不对齐,2026 年没有发生任何这么过分的事。那大概是对齐前线比我预期更糟的一个小例子。
Unfortunately, things are going roughly as I expected. You can still read AI 2027 and it still seems like, yeah, we’re sort of going down that path. I used to say that the alignment situation was better than I expected, and the governance situation was worse than I expected. But that was what I would have said a year ago or two years ago, but now I almost say the opposite. Compared to a year or two years ago, I would say that the governance situation is better than expected and the alignment situation is a bit worse than expected. In particular, the Hugging Face rogue AI incident is a more egregious example of misalignment than I expected to be happening at this time. You can tell by reading AI 2027, for example, where we talked about the misalignment over time and nothing this egregious happens in 2026. That’s, I guess, a minor example of things being worse than I expected on the alignment front.
然后在治理前线,我觉得 Anthropic 和特朗普政府之间所有那些斗争的一线光明是:特朗普政府没有像《AI 2027》里发生的那样,被领先的 AI 公司冲倒、俘获。他们也许会。我们走着看。也许部分是性格问题,也许如果领先的是 OpenAI,他们就会。但至少按现在走的方式,政府似乎比我预期的更愿意对公司踩一脚,无论是好是坏。但因为对我来说局面看起来相当糟,意味着我仍有一些希望他们会以好的方式做。而以前我预期相当糟的事,现在像是我多了一点希望他们会做好的那些事。更广泛的公众也只是逐渐开始把这一切当得更认真。各种参议员和众议员在谈失控风险等等,但总体上事情和我预期的没有那么不同。这些只是轻微变化。
Then on the governance front, I think that the silver lining of all the battles between Anthropic and the Trump administration is that the Trump administration is not being bowled over and captured by the leading AI company in the way that happened in AI 2027. They might. We’ll see what happens. Maybe it’s partly a personality thing, and maybe if OpenAI was in the lead then they would be. But at least the way it’s currently going is that it seems like the administration is more willing to bring the foot down on the companies than I expected, for better or for worse. But since the situation seems pretty bad to me, it means I still have some hope that they’ll do it in the good way. Whereas previously I was expecting pretty bad things, and now it’s like I have a little bit more hope that they’ll do the good things. Then also the broader public is just gradually starting to take all this stuff more seriously. Various senators and congressmen are talking about loss of control risk and so forth, but overall things are not that different from what I expected. These are just slight changes.
Luisa Rodriguez
有没有我们没谈到、你想让人们知道的?
Is there anything we haven’t talked about that you want people to know?
Daniel Kokotajlo
我想留给人们的是这个高层次要点:我们在做什么、为什么。我们不想这是对话的结束。它更像对话的开始。我们并不确信 Plan A 是最好的计划。我们看见 Plan A 有很多问题,有很多走错的方式。我们只是觉得它是我们目前知道的最不坏的计划。我们认为其他人——包括主要 AI 公司——在提议的替代方案,在若干方面看起来比 Plan A 差得多。我们希望随着人们醒来看到将要到来的事、把它当得更认真、开始把事情推演出来,人们会思考所有这些计划——包括 Plan A——拿走最好的元素、组合它们。我们希望实践中最终发生的会比 Plan A 更好。话虽如此,我们实际预期的是:实践中最终发生的会比 Plan A 更糟。
Yeah, I think I want to leave people with this high-level point about what we’re doing and why. We don’t want this to be the end of the conversation. It’s more like the beginning of the conversation. We are not confident that Plan A is the best plan. We see a lot of problems with Plan A, and a lot of ways it could go wrong. We just think it’s the least bad plan that we’re currently aware of. We think that the alternatives that other people — including the major AI companies — are proposing seem dramatically worse in a number of ways than Plan A. We’re hopeful that as people wake up to what’s coming and take it more seriously and start gaming things out, that people will think about all these plans — including Plan A — and take the best elements of them and combine them. We’re hopeful that what ends up happening in practice will be better than Plan A. That said, what we actually expect is that what ends up happening in practice will be worse than Plan A.
Luisa Rodriguez
今天的嘉宾是 Daniel Kokotajlo。非常感谢。
My guest today has been Daniel Kokotajlo. Thank you so much.
Daniel Kokotajlo
谢谢。谢谢邀请。
Thank you. Thank you for having me.