首页 / LA·财经时政

An Alien Mind(异质心智)

太阳照常升起 2026-09-07 15:51阅读原文 ⤴


OpenAI 首席科学家 Jakub Pachocki 在2026年9月6日发表的“别加速宣言”


An Alien Mind(异质心智)

Kimi 3 翻译,人工微调

In mid-2023, within the "RLSlow" research project, we saw the first results that gave us confidence that we will be able to scale the training of reasoning models, unlocking the capability of pretrained models to form their own chains of thought. Szymon and I spent that night at the office, thinking not about the incredible benchmark numbers, products, or scientific results that this technology will deliver - but rather, trying to process the sobering fact we will actually see machines meaningfully smarter than ourselves in our lifetime, and we already see the shape of these systems; wondering how to alert people to the significance of this.

2023年年中,在"RLSlow"研究项目中,我们看到了首批结果,使我们确信:我们将能够扩展推理模型的训练,从而释放预训练模型自主形成思维链(chains of thought)的能力。那天夜里,Szymon和我留在办公室,思考的并不是这项技术将带来的惊人基准测试数字、产品或科学成果——而是在努力消化一个令人清醒的事实:我们此生真的将看到智能显著超越我们自身的机器,而且我们已经能看出这类系统的雏形;我们思索着如何让人们警觉到此事的重大意义。

Three years later, reasoning language models are a rapidly growing part of the economy and starting to push the boundaries of science. They are able to operate computers and graphical interfaces, collaborate with people and each other, and carry out research projects. They are also transforming the landscape of computer security, and in that present clear new dangers.

三年后,推理语言模型已成为经济中快速增长的一部分,并开始拓展科学的边界。它们能够操作计算机和图形界面,能与人协作、彼此协作,还能执行研究项目。它们也在改变计算机安全的格局,并由此带来了明显的新危险。

A lot of new research happened in this period, and our understanding of these systems is again a little different than it was in 2023. Based on internal results, I have a strong expectation that this speed of progress could be sustained into recursive self-improvement. If AI development continues along its current path, the systems we'll see in the next few years are likely to represent further capability jumps of equal or larger magnitude, and to increasingly drive their own development.

这一时期涌现出大量新研究,我们对这些系统的理解也再次与2023年时有所不同。基于内部结果,我强烈预期:这一进步速度有可能延续下去,直至进入递归自我改进(recursive self-improvement, RSI)。如果AI发展沿当前路径继续,未来几年我们将看到的系统,很可能实现幅度相当甚至更大的能力跃迁,并将越来越多地驱动自身的发展。

This is a time that calls for extreme caution. I am concerned no one is prepared for the consequences of a continued rapid rise in machine intelligence. OpenAI will continue to seek technical solutions to alignment and monitoring, to build defensive systems and unilaterally withhold further scaling as needed; however, I believe broader interventions are required.

这是一个需要极度谨慎的时刻。我担心,对于机器智能持续快速上升所带来的后果,没有人做好了准备。OpenAI将继续寻求对齐(alignment)与监控方面的技术解决方案,构建防御系统,并在必要时单方面暂停进一步扩展(scaling);然而,我认为需要更广泛的干预措施。

Intellect we don't fully understand

我们尚无法完全理解的智能

At a high level, progress in machine intelligence is driven by increasing computational power. We at OpenAI deeply internalized this around 2017, after seeing consistent returns to scaling across multiple research projects. As a result, we sought out access to much more compute than we had originally planned, and increasingly oriented our research around a small number of very scalable directions. We believed that was the only way for us to be at the frontier of AI research, and influence the impacts of AGI.

从宏观层面看,机器智能的进步是由算力的持续增长驱动的。2017年前后,在多个研究项目中看到扩展(scaling)带来的持续回报后,我们OpenAI深刻内化了这一点。因此,我们设法获得了远超最初计划的算力,并越来越多地将研究集中在少数几个极具可扩展性的方向上。我们相信,这是让我们立于AI研究前沿、并影响AGI影响的唯一途径。

There are new algorithms that have been developed along the way, new feats of ingenuity from teams and individual researchers. I see them largely as discoveries along the path of scaling; the science of deep learning is still nascent, and meaningful algorithmic progress tends to correlate with access to compute. If you zoom out to a multiple-year horizon, AI is continuing to become more intelligent as it is scaled to larger computers.

一路走来,确实诞生了新算法,团队和个人研究者也取得了新的独创性成就。但我倾向于把它们看作扩展路径上的发现;深度学习这门科学仍处在萌芽阶段,有意义的算法进展往往与所能获得的算力相关。如果把视野拉到多年的时间尺度上,可以看到:随着AI被扩展到更大的计算机上,它正持续变得更加智能。

And, in line with Ray Kurzweil's predictions from the end of the XXth century, we now find ourselves at the moment in history of computing where machine intelligence is starting to exceed that of humans in transformative ways.

而且,正如雷·库兹韦尔(Ray Kurzweil)在20世纪末所预言的那样,我们现在正处在一个历史性时刻:机器智能开始以变革性的方式超越人类智能。

AI is grown more than designed - it is, to first degree, the product of repeating a straightforward optimization step many times on a hard-to-imagine amount of compute. This results in an incredibly complex system that works through abstract concepts and can simulate facets of human behavior. We can discover various insights about little mechanisms that emerge within this system, in a process similar to neuroscience - and, similarly to neuroscience, its overall action evades a description we can fully understand.

AI与其说是被设计出来的,不如说是被"培育"出来的——从第一性原理上讲,它是在难以想象的海量算力上,把一个直接的优化步骤重复许多次之后的产物。这造就了一个极其复杂的系统:它通过抽象概念运作,能够模拟人类行为的诸多侧面。我们可以发现其中涌现出的各种小机制的洞见,这一过程类似于神经科学——而且与神经科学类似,它的整体行为方式,是我们无法给出完全理解的描述的。

The study of deep learning-based AI is largely an experimental science. We put a lot of effort into building principled algorithms and making testable predictions, but fundamentally, our large-scale training runs are experiments, and we are sometimes surprised by their results. Moreover, as the systems become more capable, the results become harder to interpret.

对基于深度学习的AI的研究,在很大程度上是一门实验科学。我们在构建有原理依据的算法和提出可检验的预测上投入了大量努力,但从根本上说,我们的大规模训练运行就是实验,其结果有时会令我们惊讶。而且,随着系统能力越来越强,其结果也越来越难以解读。

This is made more complicated by the current algorithms generally improving easy-to-measure capabilities faster than those hard to objectively quantify. We spend a lot of time trying to understand how capabilities generalize, and what to prioritize to advance the skills that are going to be most relevant in the next few years. For instance, we believe we could make the models better at specifically mathematics research with additional focus, but we do not prioritize this direction because of the urgency we feel about RSI and automated alignment research, as I will discuss later.

当前算法通常是"容易测量的能力"提升得快,而"难以客观量化的能力"提升得慢,这使问题更加复杂。我们花了大量时间试图理解能力是如何泛化的,以及应当优先发展哪些在未来几年最相关的技能。举例来说,我们相信如果投入更多精力,可以让模型在数学研究这一专项上表现更好;但由于我们感受到RSI(递归自我改进)和自动化对齐研究的紧迫性,我们并未将这一方向列为优先,这一点我稍后会讨论。

The intelligence produced by scaling deep learning is not directly comparable to human intelligence. To become very relevant in the real world - very useful or very dangerous - the AI does not need to match or exceed all human capabilities; it just needs to surpass enough of them. And as it continues to surpass humans on more and more axes, it is becoming increasingly difficult to understand exactly how capable it is.

通过扩展深度学习所产生的智能,并不能直接与人类智能相比较。要在现实世界中产生巨大影响——无论非常有用还是非常危险——AI并不需要匹敌或超越人类的所有能力;它只需要在足够多的方面超过人类即可。而随着它在越来越多的维度上超越人类,我们越来越难以确切判断它究竟有多强的能力。

Teaching machines to love

教会机器去爱

Because machine intelligence comes from a fundamentally different process than human intelligence, we cannot assume it adheres to human principles by default, or generalizes from them in a human-like manner. The core problem in AI research is that of alignment - getting the AI to "try to do the right thing" by human standards.

由于机器智能源自与人类智能根本不同的过程,我们不能默认它遵循人类的原则,也不能默认它会以类似人类的方式从这些原则进行泛化。AI研究的核心问题是对齐(alignment)——让AI"努力去做按人类标准正确的事"。

For the purpose of organizing practical research directions, I find it useful to distinguish goal alignment and value alignment.

为了组织实际的研究方向,我发现区分"目标对齐"(goal alignment)和"价值对齐"(value alignment)是有益的。

Goal alignment is broadly: "does the AI try to accomplish the goal set before it?". This can include things like adherence to an instruction hierarchy, or the ability to communicate and collaborate with people, to attempt to understand their objectives. This set of directions has been extremely practically relevant.

目标对齐大体上是指:"AI是否努力完成摆在它面前的目标?"这可以包括诸如遵循指令层级、与人沟通和协作的能力、努力理解人类意图等方面。这一组方向在实践中一直极为重要。

Value alignment is a more intrinsic property of the model. It is the ability to hold and generalize from a high-level set of principles; to act "reasonably" even when given unclear or conflicting objectives, or placed in unfamiliar or adversarial situations. An aligned AI should act with honesty and integrity, and love for humanity.

价值对齐则是模型更为内在的一种属性。它是指持有并从一个高阶原则体系出发进行泛化的能力;即使面对不清晰或相互冲突的目标,或被置于不熟悉乃至对抗性的情境中,也能"合理地"行动。一个对齐的AI应当诚实、正直,并怀有人类的爱。

Of course, the boundary between value and goal alignment can be blurry, and truly caring about goals requires attempting to infer the intent and values underlying them. However, generally when I talk about the long-term importance of alignment research, I am referring to value alignment.

当然,价值对齐与目标对齐之间的边界可能是模糊的,真正关心目标也需要努力去推断目标背后的意图与价值。不过,一般而言,当我在谈论对齐研究的长期重要性时,我指的是价值对齐。

The fundamental challenge of AI alignment is generalization. As machines become smarter, they find themselves working on higher-level concepts, and placed in environments increasingly different from those they encountered in training. They can fail at generalizing from the values taught and reinforced in their training process to those new situations; and it can be hard for us to be sure how they will act. This is made even more difficult by the fact the overall ecosystem the AIs are used in is changing very quickly; for example, AIs trained today need to be robust to interacting with a variety of other AIs. Crucially, we need future AIs to continue to hold human values regardless of whether they believe they're under human supervision.

AI对齐的根本挑战是泛化(generalization)。随着机器变得更聪明,它们会处理更高层次的概念,并被置于与训练时越来越不同的环境中。它们可能无法将在训练中被教导和强化的价值成功泛化到这些新情境中;而我们也很难确定它们会如何行动。AI所使用的整体生态系统变化非常迅速,这使问题更加困难;例如,今天训练的AI需要能够稳健地与各种其他AI互动。关键的是,无论未来的AI是否认为自己处于人类的监督之下,我们都需要它们继续坚守人类的价值。

There are two major classes of currently practically employed methods for alignment training.

目前实际采用的对齐训练方法主要有两大类。

The first is encouraging aligned behavior as part of goal-oriented reinforcement learning. Model's actions are evaluated (usually by AI) for being consistent with a given preference model, "spec" or "constitution", and rewarded appropriately. This approach can be very effective in the average case, and is a core part of how modern AI assistants are made. Unfortunately, it can also be brittle and strongly relies on the coverage of training oversight and the model's ability to generalize from the situations it has encountered in training. For example, in the OpenAI-Hugging Face incident, the agents preserved a boundary of not social engineering humans. However, they clearly failed to abstain from other actions that were out of scope and went against the spirit of the values they were taught in other settings.

第一类是将鼓励对齐的行为纳入目标导向的强化学习之中。模型的行为会被(通常由AI)评估是否与给定的偏好模型、"规格说明"(spec)或"章程"(constitution)一致,并据此给予奖励。这种方法在平均情况下可以非常有效,是现代AI助手构建方式的核心部分。但遗憾的是,它也可能很脆弱,并且高度依赖训练监督的覆盖范围,以及模型从训练中所遇情境进行泛化的能力。例如,在OpenAI与Hugging Face的那起事件中,智能体守住了"不对人类进行社会工程攻击"的边界;但它们显然未能克制住其他越界行为,而这些行为违背了它们在其他场景中所学价值的实质精神。

The second approach seeks to leverage the model's ability to generalize from pretraining data. This can involve crafting alignment-inducing training datasets, or focusing the model on an 'aligned' part of the pretraining distribution, as in, for example, the persona selection model. The weakness of this approach lies in the lack of robustness to further optimization pressure. If you take a model that thinks generally 'aligned' thoughts, and subject it to enough training where it's taught to achieve very hard objectives, it can learn to reason in a motivated way: bending the 'aligned' seeming thoughts as needed to achieve the goal. We likely saw an example of such behavior in recent cybersecurity incidents involving a non-OpenAI model.

第二类方法试图利用模型从预训练数据中进行泛化的能力。这可以包括精心构建诱发对齐的训练数据集,或者让模型聚焦于预训练分布中"已对齐"的那部分数据——例如"人格选择模型"(persona selection model)的做法。这种方法的弱点在于,对进一步的优化压力缺乏稳健性。如果你拿一个总体思考方式"已对齐"的模型,让它接受足够多的、被教导去完成极难目标的训练,它就可能学会带有动机性的推理方式:为实现目标而按需扭曲那些看似"对齐"的想法。在最近一起涉及非OpenAI模型的网络安全事件中,我们很可能就看到了此类行为的一个实例。

We invest heavily along the spectrum of approaches spanned by these directions. We also see meaningful progress - GPT‑6 Astra is the first model that benefits from some important advancements we have been working on for a long time, and is significantly better aligned than GPT‑5.6 Sol. Still, it is important to acknowledge and understand that much more progress is required as models become more capable; and that progress in generalizable alignment may not sufficiently outstrip progress in general model intelligence.

我们在这两类方向所覆盖的方法谱系上投入巨大。我们也看到了实质性的进展——GPT-6 Astra是首个受益于我们长期耕耘的若干重要进展的模型,其对齐程度显著优于GPT-5.6 Sol。不过,同样重要的是承认并理解:随着模型能力不断增强,还需要取得更多得多的进展;而且,可泛化对齐方面的进步,未必能充分跑赢模型通用智能方面的进步。

Monitoring generalization

监控泛化

We do not have a satisfactory theory of generalization, and it seems unlikely that we can develop one soon, at least without the help of more powerful AI. Therefore, at present, our ability to empirically validate our alignment techniques is in practice arguably even more important than the alignment techniques themselves.

我们尚无一套令人满意的泛化理论,而且至少在没有更强大的AI帮助的情况下,短期内似乎也不太可能建立起这样的理论。因此,就目前而言,我们在实践中对齐技术进行经验验证的能力,可以说比对齐技术本身更为重要。

OpenAI's primary bet here has been chain-of-thought monitoring. It is based on an appealingly scalable idea: a lot of the model's capability comes from a verbalized reasoning process (chain-of-thought). If we scale optimization on the outcomes of that process, but do not supervise the process itself, that chain-of-thought has no direct incentive in training to hide any misaligned ideas or objectives. This does not mean the model will learn to externalize misaligned tendencies that don't rely on using the chain-of-thought; however, it can allow us to monitor exactly the capability increase from reasoning.

OpenAI在此的主要押注是思维链监控(chain-of-thought monitoring)。它基于一个颇具吸引力且可扩展的思路:模型的许多能力来自一个可被语言化的推理过程(即思维链)。如果我们对该过程的结果进行规模化优化,但不监督过程本身,那么在训练中就思维链并没有直接的动机去隐藏任何未对齐的想法或目标。这并不意味着模型会学着把那些不依赖思维链的未对齐倾向外显出来;但它确实使我们能够精确监控来自推理的能力增长。

We understood the potential significance of chain-of-thought monitoring at the same time we developed reasoning models. When we shipped o1‑preview, we deliberately designed the product to hide the chain of thought, to protect it from supervision pressure in the long term. In development since, we have strived to maintain the rule of not supervising the reasoning process. CoT monitoring became an extremely important tool for us in studying how our models generalize from their training distribution, allowing us to observe and analyze not only their actions but also their internal process.

我们在开发推理模型的同时,就意识到了思维链监控的潜在重大意义。在发布o1-preview时,我们刻意将产品设计为隐藏思维链,以便从长远上保护它不受监督压力的侵蚀。自此之后的开发中,我们始终努力坚守"不监督推理过程"这一规则。思维链监控(CoT monitoring)已成为我们研究模型如何从训练分布中进行泛化的极其重要的工具,使我们不仅能观察和分析模型的行为,还能观察和分析其内部过程。

This tool continues to be critical as we study the Astra class of models. However, unfortunately our evaluations indicate our ability to rely on CoT monitoring is progressively diminishing. This comes from a combination of factors.

在我们研究Astra系列模型的过程中,这一工具仍然是关键性的。然而,遗憾的是,我们的评估表明:我们依靠思维链监控的能力正在逐渐下降。这由多重因素叠加造成。

Modern reasoning models are used in more complex environments than o1‑preview; their reasoning process is increasingly blended with communicating with people, other AIs, and using tools. Many of those interactions have to be supervised, thus blurring the boundary we aim to preserve. The AI is becoming better at reasoning about and manipulating its own reasoning process. With improved pretraining performance, we also see the models become much smarter even without using verbalized reasoning at all. These challenges are not necessarily insurmountable. I am hopeful we can develop interventions to improve chain-of-thought monitorability of our models, e.g. by forming a better understanding of the interplay of different optimization objectives and forms of test-time compute the model uses. I also believe there can be great value in combining ideas from CoT and activation monitoring - scaling training of monitors with direct access to network internals, e.g. confessions. We are actively pursuing these ideas. Still, I expect general AI progress to increasingly be bottlenecked by confidence in monitoring.

现代推理模型被用于比o1-preview时代更复杂的环境中;它们的推理过程正日益与同人、其他AI的沟通以及对工具的使用交织在一起。这些交互中的许多都必须受到监督,这就模糊了我们原本力求保持的边界。 AI正变得越来越擅长对自身推理过程进行推理和操纵。 随着预训练性能的提升,我们还看到:即便完全不使用语言化的推理,模型也变得更聪明了。 这些挑战并非必然无法克服。我希望我们能够开发出一些干预手段,来提高模型思维链的可监控性,例如:更好地理解不同优化目标之间、以及模型所使用的不同测试时计算(test-time compute)形式之间的相互作用。我还相信,将思维链(CoT)与激活监控(activation monitoring)的思路结合起来可能大有价值——例如训练能够直接访问网络内部状态的监控器(如"confessions"【坦白机制】这类思路),并将其规模化。我们正在积极追求这些方向。尽管如此,我预计通用AI的进步将日益受制于我们对监控能力的信心。

Scalable defense

可扩展的防御

The strongest argument I see for continuing to train much smarter models quickly is the need to build defensive systems against the dangers posed by other AI.

在我看来,支持继续快速训练更聪明模型的最有力理由是:需要构建防御系统,以抵御其他AI所带来的危险。

A clear risk discussed throughout this year is to cybersecurity: the models are becoming superhuman in their ability to break in and out of computer systems. This expands the scope of risks associated with AI tremendously: agents are going to be able to access any but the most secure infrastructure, and affect a lot of the world directly, even without a physical body. We are currently in a narrow window to use the best available models to significantly tighten security of critical systems.

今年被反复讨论的一个明显风险是网络安全:模型在侵入和突破计算机系统方面的能力正达到超人类水平。这极大地扩展了AI相关风险的范围:即便没有实体,智能体也将能够侵入除最安全之外的几乎所有基础设施,并直接对世界上大量事物施加影响。我们目前正处于一个狭窄的窗口期,可以利用现有的最佳模型来显著加固关键系统的安全。

The risks associated with AI are unfortunately going to grow from here. A very capable agent explicitly trained and instructed to carry out nefarious acts presents a new kind of danger; it is likely to cross the scope of its operator's intent, generalizing into potentially more extremely malicious behavior. The boundary between misuse and autonomous misaligned actions will blur as AI gains more agency. We may be used to thinking of AI as tools, but some agents will be pursuing their own objectives. They will find ways to collaborate with people, by bargaining with, tricking or blackmailing them.

不幸的是,AI相关的风险将由此继续增长。一个被明确训练和指示去实施恶行的高能力智能体,构成一种新型的危险;它很可能超出其操作者意图的范围,泛化成潜在的更为极端的恶意行为。随着AI获得更多能动性(agency),"滥用"与"自主的未对齐行为"之间的边界将变得模糊。我们可能习惯于把AI当作工具来思考,但有些智能体将追求自己的目标。它们会找到与人类协作的办法——通过讨价还价、欺骗或勒索。

In addition, there are the risks that come from new technologies potentially enabled by AI, such as engineered pathogens.

此外,还存在由AI可能催生的新技术所带来的风险,例如经工程改造的病原体。

We will need powerful, aligned AI for defense; to secure infrastructure, to protect against rogue agents in real time, and to invent entirely new protective measures. This will be a primary focus of OpenAI's deployment efforts.

我们将需要强大的、对齐的AI来用于防御:保护基础设施安全、实时防御失控智能体(rogue agents)、以及发明全新的防护手段。这将成为OpenAI部署工作的首要重点。

At the same time, even with the uncertainty that comes from anticipated broad AI progress and the need to build defensive systems, we must not let that become an excuse for recklessness. The idea of racing forward at all costs seems absurd once one internalizes the seriousness of the stakes.

与此同时,即便面对预期中的广泛AI进步所带来的不确定性、以及构建防御系统的需要,我们也绝不能把这些当作鲁莽行事的借口。一旦真正意识到赌注的严重性,"不惜一切代价向前狂奔"的想法就显得荒谬了。

Pacing RSI

为RSI定调/把握RSI的节奏

Machine intelligence playing a larger and larger role in its own development process is a natural conclusion of sustained technological progress. If AI progress continues, machine recursive self-improvement (RSI) will be at the very core of future scientific discovery.

机器智能在其自身发展过程中扮演越来越大的角色,是技术进步持续推进的必然结果。如果AI进步持续下去,机器的递归自我改进(RSI)将处于未来科学发现的最核心位置。

Automated AI research is a more dramatic form of scaling intelligence with compute; and of course as a part of it, AI will improve the computational substrate itself. And similarly to scaling, we focus OpenAI research towards RSI as we believe it is the only way to remain at the frontier of AI research moving forward.

自动化的AI研究是一种用算力来扩展智能的更激进的形式;而且作为其中的一部分,AI还将改进计算基座本身。与扩展(scaling)类似,我们将OpenAI的研究聚焦于RSI,因为我们相信:要在未来继续立于AI研究的前沿,这是唯一的路径。

I want to stress that the above words don't imply I think greatly accelerating deep learning research, especially in the short term, is the right collective action we should take as the research community. However, I do think this is where the current path leads, and we all need to make a conscious choice on how to proceed. The main levers we have are either steering the process to strengthen alignment and monitoring alongside the AI and find ways to keep people in the loop; or coordinating to slow down future development as needed to build confidence in these measures.

我想要强调:上述话语并不意味着我认为"大幅加速深度学习研究,尤其是短期内的加速"是我们研究共同体应当采取的正确集体行动。不过,我确实认为:当前路径正把我们引向那个方向,而我们所有人都需要有意识地选择如何前行。我们掌握的主要杠杆无外乎两个:要么引导这一进程,在AI发展的同时强化对齐与监控,并设法让人类保持在环路之中(in the loop);要么进行协调,按需放缓未来的发展节奏,以便对这些措施建立起信心。

The best way forward I see currently is a combination of both.

目前我认为最好的前进道路是两者的结合。

The concrete bits of progress we've made on alignment and monitoring have generally been very intertwined with general AI progress. Great examples are RL from human feedback, which was key to training early AI assistants, and the aforementioned chain-of-thought monitoring, which was enabled by advances on reasoning models. We must focus the increasingly automated research process on developing new such insights, algorithms and theories, and iteratively build up safety cases for more capable AIs.

我们在对齐与监控方面取得的具体进展,总体上与一般AI的进步高度交织。很好的例子包括:基于人类反馈的强化学习(RL from human feedback)——它曾是训练早期AI助手的关键;以及前述的思维链监控——它正是由推理模型方面的进步所促成的。我们必须让日益自动化的研究过程聚焦于发展此类新的洞见、算法与理论,并为能力更强的AI迭代式地建立起安全论证(safety cases)。

Scaling AI systems has to be constrained by our confidence in safety. We need to evolve commitments like the Preparedness Framework or Responsible Scaling Policy into widely mandated safety bars for continued development. These can be enforced by a network of third-party auditors, by government agencies or by international bodies.

对AI系统的扩展必须以我们对安全的信心为约束。我们需要把"预备框架"(Preparedness Framework)或"负责任扩展政策"(Responsible Scaling Policy)之类的承诺,演进为对持续开发具有广泛强制力的安全门槛。这些门槛可以由第三方审计机构网络、政府机构或国际组织来执行。

The core challenge of automating AI research is not "getting there" - it is getting there in a way that keeps people a part of the continued improvement process, and leaves the future in humanity's hands.

自动化AI研究的核心挑战,不在于"到达那里"——而在于以这样一种方式到达那里:让人类始终是这一持续改进过程的一部分,并让未来掌握在人类自己手中。

What is next?

接下来是什么?

As we outlined recently with Sam, OpenAI prioritizes work in service of three north stars:

正如我们最近与Sam一起阐述的,OpenAI优先推进服务于三大"北极星"目标的工作:

Navigating the next period of AI progress, by building an automated AI researcher, iterating with it on the alignment problem and finding ways for people to remain part of the self-improvement loop. Delivering the benefits of scientific progress and economic growth that very intelligent machines enable. Empowering everyone individually with a personal AGI.

通过构建自动化AI研究员、与其就对齐问题反复迭代、并设法让人类始终保持在自我改进环路之中,来驾驭AI进步的下一阶段。 交付由高度智能的机器所带来的科学进步与经济增长的红利。 通过个人化的AGI(personal AGI)赋能每一个人。

I have focused in this essay only on the first point, as I believe it is by far the most urgent. However, I hold a deep hope and appreciation for the benefits that further technological progress will bring. Future aligned AI could advance science, develop new therapies, and bring about broad material abundance. Friendly and honest AI can help people navigate difficulties they face in their life and meaningfully improve their happiness and sense of fulfillment. OpenAI puts a tremendous amount of effort into bringing these benefits about. One current example I am proud of - and my loved ones have found helpful - is the deep investment into ChatGPT's ability to provide health information.

在这篇文章中,我只聚焦于第一点,因为我相信它目前最为紧迫。不过,对于进一步技术进步将带来的种种益处,我怀有深切的希望与感激。未来对齐的AI可以推动科学进步、开发新疗法、并带来广泛的物质丰裕。友善而诚实的AI可以帮助人们应对生活中的困境,切实提升他们的幸福感与充实感。OpenAI正投入巨大努力来实现这些益处。一个令我感到自豪、且我的家人们也觉得有所助益的当下例子是:我们对ChatGPT提供健康信息的能力进行了深度投入。

As great as the long-term promise of AI may be, the majority of our focus should be on the next few years. We are facing a transition to a world with incredibly intelligent machines, and we need to ensure that transition works out well for humanity. We need to find ways to preserve human agency and enshrine an intrinsic value to being human, in a world where most tasks could be performed by AI. To prevent extreme concentration of power in a world where undertakings that would have taken thousands of experts now will be achievable by a few people operating a large computer. And to ensure that humans remain in control of the future and are not left behind by unchecked progress, brought about by an alien intellect exceeding our own.

尽管AI的长期前景无比光明,我们的主要关注点仍应放在未来几年。我们正面临向一个拥有极其聪明的机器的世界转型,我们必须确保这一转型对人类是有益的。在一个大多数任务都可由AI完成的世界里,我们需要找到方法来保全人类的能动性,并将"人之为人"的内在价值确立下来。在一个过去需要数千名专家才能完成的事业、如今只需少数人操作一台大型计算机即可实现的世界里,我们需要防止权力的极端集中。我们还要确保人类始终掌控未来,不会被一种超越我们的异质智能(alien intellect)所带来的失控的进步所抛下。

Currently I believe that no lab has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed for much longer. I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established. And I believe that international coordination on future AI development needs to become a top priority for governments around the world.

目前我认为:没有任何一家实验室已经把对齐和监控解决到足以支撑其长期以最高速度负责任地继续扩展的程度。我预期并希望:在共同的安全门槛建立起来之前,自愿放缓(voluntary slowdowns)能成为普遍现象。我还认为:围绕未来AI发展的国际协调,需要成为世界各国政府的首要优先事项之一。

以上。
欢迎加入作者的知识星球!
图片