Perspective观点
Alignment is evidence,
not authority.
对齐是证据,
不是权力。
Why OpenAI's beneficial-RL results strengthen the case for execution governance. 为什么 OpenAI 的 beneficial-RL 结果,反而强化了执行治理的必要性。
Aron · Aiegis · June 2026
In response to OpenAI,回应 OpenAI, Reinforcement Learning Towards Broadly and Persistently Beneficial Models (June 18, 2026)(2026 年 6 月 18 日)
OpenAI's new alignment paper is real progress, and we want to say so plainly before we say anything else. Training models on beneficial traits — honesty, epistemic humility, corrigibility, concern for human welfare — and finding that those traits generalize across domains and persist under adversarial pressure is a meaningful result. It is, in our reading, the strongest empirical case yet for why a governance layer like ours needs to exist.OpenAI 这篇新的对齐论文是真正的进展,在说别的之前,我们想先把这句话讲清楚。把模型训练出"有益特征"——诚实、认知谦逊、可纠正性、对人类福祉的关切——并发现这些特征能跨领域泛化、能在对抗压力下保持,这是有分量的结果。在我们看来,它恰恰是迄今为止最有力的实证,说明为什么像我们这样的治理层必须存在。
That may sound backwards. Let us explain.这听起来像是反话。容我们解释。
Anyone who has worked in banking, insurance, or audit already thinks in one equation:任何做过银行、保险或审计的人,脑子里都有一个等式:
Risk = Probability × Impact风险 = 概率 × 后果
Alignment research attacks the first term. It lowers the probability that a model forms or acts on harmful intent. The beneficial-RL paper does this well: across its evaluations the trained model became measurably more truthful, more open to correction, and "harder to push" toward harmful behavior.对齐研究攻的是第一项。它降低模型产生或执行坏意图的概率。这篇 beneficial-RL 论文做得不错:在它的评测里,训练后的模型变得更真实、更愿意接受纠正、"更难被推向"有害行为。
But notice what the paper carefully does not claim. Its results are improvements, not guarantees. The model improved on 44 of 53 benchmarks — so nine did not. Under adversarial pressure it was "harder to push," not impossible to push. Under harmful fine-tuning it was "somewhat more resistant." The authors devote real effort to studying persistence — which is itself an honest admission that alignment can drift under pressure and fine-tuning. These are the right words. They are also probabilistic words.但请注意论文谨慎地没有声称的那部分。它的结果是"改善",不是"保证"。模型在 53 项中的 44 项上变好——也就是有 9 项没有。对抗压力下是"更难被推",不是不能被推。有害微调下是"略微更耐"。作者花了真功夫研究持续性(persistence)——这本身就是一种诚实的承认:对齐会在压力和微调下漂移。这些用词是对的,但它们都是概率性的词。
A bank cannot underwrite a fiduciary system on "44 of 53 benchmarks improved." Regulated, high-stakes deployment needs a worst-case bound, not an expected-case improvement. Lowering probability is necessary and valuable. It is not the same kind of object as bounding impact.银行无法用"53 项中 44 项变好"去为一个受托系统做承保。受监管的高风险部署需要的是最坏情况边界,而不是平均情况的改善。降低概率是必要且有价值的,但它和"限定后果"不是同一类东西。
That is the second term in the equation, and it is where execution governance lives. Governance does not try to make the model good. It assumes residual risk remains after the best available alignment, and it bounds what any agent — aligned or not — is allowed to do: which actions require authorization, which are reversible, which are irreversible and therefore gated, and what the blast radius can ever be if something goes wrong. Crucially, this authorization layer is independent of the model. It does not consult the model's character before deciding what the model may do.这就是等式的第二项,也是执行治理所在的地方。治理不试图让模型变好。它假设在最好的对齐之后,残余风险依然存在,并去限定任何 agent——无论是否对齐——被允许做什么:哪些动作需要授权、哪些可逆、哪些不可逆因而必须设门、一旦出错爆炸半径最大能到多大。关键在于,这个授权层独立于模型。它在决定模型能做什么之前,不去咨询模型的"品格"。
Here is the part of OpenAI's paper that we find most clarifying — and it cuts in our favor, not against us. The paper builds on the finding that models have personas: behavioral dispositions that are real, that generalize across tasks, and that can drift. Some readers will take that as evidence that we can simply train the persona to be good and stop worrying. We read it the opposite way. The more real and generalizable a persona is, the less verifiable it becomes from the outside — and the less it can be trusted at the moment of execution. The paper does not weaken the assumption that a model should be untrusted at the boundary where it takes real-world action. It strengthens it.OpenAI 论文里我们觉得最有启发的一点,恰恰是对我们有利、而非不利的:论文建立在一个发现之上——模型拥有人格(persona):真实存在、能跨任务泛化、且会漂移的行为倾向。有人会把这当作"那就把人格训好、不必再担心"的证据。我们的读法正相反。人格越真实、越能泛化,从外部就越无法验证——在执行的那一刻就越不能被信任。论文没有削弱"模型在执行真实动作的边界上应被视为不可信"这个假设,反而强化了它。
This leads to a principle we are willing to put our name on:由此得到一条我们愿意署名的原则:
Alignment is evidence, not authority.Alignment is evidence, not authority.(对齐是证据,不是权力。)
A model's measured traits — its honesty score, its persona profile, its drift telemetry — are legitimate and useful evidence. They belong in audit, monitoring, trust analysis, and forensics. They should inform how closely a human watches an agent. They must never become authority — they must never silently raise or lower what an agent is permitted to do. A bank does not skip an approval because the customer "seems honest." Neither should an agent platform.模型被测得的特征——诚实分、人格画像、漂移遥测——都是正当且有用的证据。它们属于审计、监测、信任分析与取证;它们应当影响人类盯一个 agent 盯得多紧。但它们绝不能变成权力——绝不能悄悄抬高或压低一个 agent 被允许做的事。银行不会因为客户"看起来诚实"就跳过审批,agent 平台也不应该。
We hold this line even in the tempting direction. It is natural to imagine a "drift monitor" that detects a model going off the rails and automatically tightens its permissions. Detection: yes, and we think it is exactly the right lesson to draw from a paper about persistence and trait stability. But the response must stay an open loop — raise an alert, trigger human review, pull an out-of-band kill-switch. The moment a model-derived score is wired directly into the permission engine, you have rebuilt the very dependency you were trying to remove: an adversary who controls the model can now spoof the score, and the gate moves with it. Observation feeds humans. It does not feed the gate.即便在最有诱惑的方向上,我们也守这条线。很自然会想象一个"漂移监测器":发现模型跑偏就自动收紧它的权限。监测,可以——我们认为这正是从一篇研究持续性与特征稳定性的论文里该汲取的教训。但响应必须保持开环:发告警、触发人工复核、拉带外 kill-switch。一旦把模型派生的分数直接接进权限引擎,你就重建了自己本想拆掉的那条依赖:控制了模型的攻击者现在能伪造这个分数,门就跟着动。观测喂给人,不喂给门。
None of this is a critique of alignment work. It is a division of labor. Probability is OpenAI's to lower, and they are lowering it. Impact is ours to bound, and that job does not get easier as models get better — it gets more important, because more capable models are trusted with more consequential actions.以上没有一句是对对齐工作的贬低。这是分工。概率归 OpenAI 去降,他们正在降;后果归我们去限,而这件事不会因为模型变好就变容易——只会变得更重要,因为更强的模型被托付了更有后果的动作。
Alignment shapes what an agent wants to do.
Governance determines what an agent is allowed to do.
You need both, and you must control them independently. 对齐塑造一个 agent 想做什么。
治理决定一个 agent 被允许做什么。
两者都需要,而且必须各自独立地控制。
To the OpenAI authors: thank you for the paper, and for the honesty of its language. We would genuinely welcome the conversation about where evidence ends and authority begins.致 OpenAI 的作者们:谢谢这篇论文,也谢谢它用词的诚实。关于"证据"在哪里结束、"权力"从哪里开始,我们非常愿意把这场对话继续下去。