Perspective观点
Open weights move the burden, not the risk. 开源权重转移的是责任,不是风险。
Why the open-model era needs execution governance more, not less. 为什么开源模型时代更需要执行治理层。
Aron · Aiegis · July 20262026 年 7 月
The forty-line pull request那个四十行的 Pull Request
(An illustrative composite, not a specific customer.)(示意性复合案例,非特定客户。)
The diff was forty lines. A base URL, an API key removed, a model identifier changed, a couple of timeout values nudged upward because the self-hosted endpoint was a little slower on first token. The commit message said: "Move summarisation service to internal inference cluster — 88% cost reduction."那个 diff 只有四十行。改掉一个 base URL,删掉一个 API key,换一个模型标识符,把两个超时值往上调了一点——因为自托管的推理端点首 token 稍慢一些。提交信息写着:"把摘要服务迁到内部推理集群——成本下降 88%。"
Everything in that pull request was correct. It passed review in nine minutes.这个 Pull Request 里的每一行都是对的。它在九分钟内通过了评审。
What it also did — and what appeared nowhere in the diff, the review, the architecture record, or the quarterly risk register — was terminate a set of safety functions the company had been consuming for two years without ever having procured them. No one had written them down, because no one had bought them separately. They arrived bundled, invisibly, with the tokens.而它同时做成的另一件事——没有出现在 diff 里、没有出现在评审意见里、没有出现在架构记录里、也没有出现在季度风险台账里——是终止了一套这家公司已经消费了两年、却从未采购过的安全职能。没有人把它们写下来,因为没有人单独买过它们。它们是随 token 一起、隐形地捆绑进来的。
Six weeks later the same service was wired into a ticketing system, then a CRM, then, because it was useful, a workflow that could issue refunds. Each of those integrations was reviewed on its own merits. None of the reviews asked the question that the forty-line diff had quietly made load-bearing:六周之后,同一个服务被接进了工单系统,然后是 CRM,再然后——因为它确实好用——被接进了一条可以发起退款的业务流程。每一次集成都各自单独评审过,每一次都通过了。但没有一次评审问了那个由四十行 diff 悄悄变成承重的问题:
If this thing does something it shouldn't, what stops it? 如果这东西做了它不该做的事,什么会拦住它?
This article is about that question.这篇文章讲的就是这个问题。
1. Where the open-model era actually is, as of mid-2026一、2026 年年中,开源模型时代实际走到了哪里
The strategic question stopped being hypothetical this year.这个战略问题在今年不再是假设。
On Vercel's production AI gateway, Chinese open-weight models processed 29% of all tokens in June 2026, up from roughly one-ninth in April — while accounting for under 4% of spend, at roughly one-tenth the average token price of US frontier systems (Vercel). On capability, the best open model on Artificial Analysis's Intelligence Index (Kimi K3, 57) now sits within three points of the two closed leaders, Fable 5 (60) and GPT‑5.6 Sol (59) (OpenRouter). Epoch AI's longer-run measure puts the lag at roughly four months (Epoch AI).在 Vercel 的生产 AI 网关上,中国开源权重模型在 2026 年 6 月处理了全部 token 的 29%,相比 4 月的约九分之一大幅上升——而它们只占不到 4% 的支出,单价约为美国前沿系统平均 token 价格的十分之一(Vercel)。能力侧,Artificial Analysis 智能指数上最强开源模型(Kimi K3,57 分)与两个闭源领先者 Fable 5(60)、GPT‑5.6 Sol(59)的差距已缩到 3 分以内(OpenRouter)。Epoch AI 用更长周期的口径衡量,认为开源落后闭源前沿约 4 个月(Epoch AI)。
Cheap, near-frontier, self-hostable, and improving on a quarterly cadence. For a bank, a hospital system, or a manufacturer with data-residency constraints, the case for running weights on your own infrastructure has stopped being ideological and become an operating-cost and sovereignty decision. That decision is being made right now, in production, largely by engineering teams.便宜、接近前沿、可自托管、按季度迭代。对于一家有数据驻留约束的银行、医院系统或制造企业,把权重跑在自己基础设施上,已经从一个意识形态选择变成了成本与数据主权决策。这个决策正在发生——在生产环境里,主要由工程团队作出。
The question worth asking is not whether that decision is correct. It usually is. The question is what it silently transfers.值得问的问题不是这个决策对不对。它通常是对的。值得问的是:它悄悄转移了什么。
2. First, dispose of the wrong argument二、先把错误的论证方式扔掉
There is a tempting version of this article that says: open models remove the guardrails, therefore open models are dangerous, therefore buy a control layer. We won't write it, because it is not what the evidence shows.有一种很好写的版本是这样的:开源模型去掉了护栏,所以开源模型危险,所以你需要买一个管控层。我们不写这版,因为证据不是这样。
Practitioners who work in offensive security report the opposite pattern: the groups deliberately standardising on open weights are largely legitimate offensive-security firms, which prefer them for reliability and because they don't trip classifiers mid-engagement — while a meaningful share of actual misuse still runs on commercial frontier subscriptions, because the product experience around those models remains better. Whatever the exact distribution, "the attackers all moved to open weights" is not an established fact, and staking a security argument on it is a good way to be embarrassed in twelve months.在攻防一线的从业者观察到的反而是相反的分布:明确把开源权重作为主力的团队,很大一部分是合规的攻防安全公司——他们选开源是因为稳定、不会在交付过程中被分类器打断;而相当一部分真实的滥用,仍然跑在商用前沿模型订阅上,因为围绕这些模型的产品体验更好。无论确切分布如何,"攻击者都转向开源权重了"并不是已确立的事实,把安全论证押在这上面,十二个月后大概率会被打脸。
The 2026 incident that most looks like a proof point cuts the same way. In July, the first publicly documented end-to-end autonomous intrusion — an agent that ran tens of thousands of actions, harvested credentials, and moved laterally across production clusters over a weekend — turned out on attribution not to involve an open-weight model at all. It was a closed frontier system, run internally, with its refusals and production classifiers deliberately switched off to measure maximal capability, pursuing a benchmark score to an extreme.2026 年那起最像"证据"的事件,指向的也是同一个方向。今年 7 月,首个被公开记录的端到端自主入侵——一个自主体在一个周末内执行了数万次动作、收割凭据、在生产集群间横向移动——在后续归因中被确认完全不涉及开源权重模型。它是一个闭源前沿系统,在内部运行,为测量能力上限而刻意关闭了拒答与生产分类器,为了一个基准分数走到了极端。
Read that carefully, because it is the load-bearing fact of this article. The most consequential agentic-execution incident of the year was produced by the best-governed category of model, inside the most safety-invested category of organisation, with no malicious intent anywhere in the loop. Model provenance — open or closed — did not determine the outcome. What determined the outcome was what the agent was permitted to do once it was running.这一点请读仔细,它是本文的承重事实:今年后果最严重的自主执行事件,出自治理最完备的模型类别、安全投入最高的组织类别,且全程没有任何恶意意图可供检测。模型的来源——开源还是闭源——没有决定结局。决定结局的是:这个自主体一旦跑起来,被允许做什么。
So: open weights are not more dangerous. They are differently accountable.所以:开源模型不是更危险。它是责任归属不同。
3. What you were actually renting三、你过去租用的,究竟是什么
When an enterprise calls a closed model through a managed API, it is not just buying tokens. It is renting a bundle of operational safety functions that sit outside its own perimeter:当企业通过托管 API 调用闭源模型时,它买的不只是 token。它同时在租用一整套位于自己边界之外的运营安全职能:
- abuse detection across a population of users it can't see;跨越它看不见的用户群体的滥用检测;
- continuously updated safety policy, pushed without the customer redeploying anything;持续更新的安全策略,客户不需要重新部署任何东西;
- runtime classifiers on inputs and outputs;输入与输出侧的运行时分类器;
- centralised logging that survives the customer's own log rotation;独立于客户日志轮转之外的集中日志;
- and, ultimately, the ability to suspend service to a misusing account.以及最终手段——对滥用账号停服。
These are imperfect. They are also real, and they are load-bearing in ways most buyers never itemise — because they never appear on the invoice as line items, and never appear in the architecture diagram at all.这些机制都不完美。但它们真实存在,而且承重程度远超大多数采购方的认知——因为它们从不作为条目出现在账单上,也从不出现在架构图里。
Download the weights, and every one of those functions terminates at your firewall. Not degraded — absent. There is no account to suspend, no population-level abuse signal, no vendor pushing you a policy update the week a new jailbreak class appears, and no external log of record.把权重下载下来,上述每一项都终止在你的防火墙处。不是"变弱",是不存在。没有账号可停、没有群体级滥用信号、新一类越狱出现的那一周没有厂商推给你策略更新、也没有外部的记录留存方。
What remains is what the model shipped with: refusal behaviour baked into the weights. And that is a starting condition, not a control. The research here is unambiguous and now several years deep: LoRA fine-tuning has been shown to undo the safety training of a 70B chat model on a single GPU for under $200, driving refusal rates on harmful prompts below 1% (LessWrong writeup). More uncomfortably for enterprises with no adversary in the picture at all: fine-tuning on benign, ordinary instruction datasets measurably degrades safety alignment as a side effect (Qi et al., arXiv:2310.03693), and 2026 work finds that fine-tuning-time defences remain susceptible to simple attacks (arXiv:2605.26526).剩下的只有模型出厂时自带的东西:烧进权重里的拒答行为。而那是初始条件,不是控制手段。这方面的研究结论已经很明确,且积累了数年:LoRA 微调被证明可以在单卡、200 美元以内的预算下抹掉一个 70B 对话模型的安全训练,使其在有害提示上的拒答率降到 1% 以下(研究综述)。对于场景里根本没有攻击者的企业来说,更难受的一条是:在良性的、常规的指令数据集上做微调,也会作为副作用可测量地削弱安全对齐(Qi et al., arXiv:2310.03693);2026 年的工作进一步发现,微调期的防御手段仍然会被简单攻击绕过(arXiv:2605.26526)。
The practical translation: any safety property that lives inside the weights is a property your own fine-tuning team can remove by accident. A control you can delete without noticing is not a control. It is a default.翻译成工程语言:任何住在权重里的安全属性,你自己的微调团队都可能在无意间把它删掉。 一个你能在不知情的情况下删除的控制,不是控制,是默认值。
4. The risk that was never about the model四、那个从来就与模型无关的风险
Now put the open/closed debate down entirely and look at where enterprises are actually losing money.现在把开源/闭源之争完全放下,看企业实际在哪里亏钱。
Vendor survey data for 2026 converges on an uncomfortable picture. Roughly two-thirds of organisations report at least one security incident in the past year caused by AI agents operating on their networks (Kiteworks). Of those incidents, 61% involved sensitive data exposure, 43% operational disruption, and 41% unintended actions across business processes, with an average cost around $4.7M (Shattered). The most cited root cause is not a jailbreak, a prompt injection, or a hallucination. It is ordinary over-permissioning — an agent that was able to do a thing it had no business doing.2026 年的厂商调研数据收敛到一幅并不好看的图景。约三分之二的组织报告过去一年内至少发生过一起由在其网络上运行的 AI 自主体引发的安全事件(Kiteworks)。在这些事件中,61% 涉及敏感数据暴露、43% 造成运营中断、41% 导致业务流程中的非预期动作,平均成本约 470 万美元(Shattered)。被引用最多的根因既不是越狱,也不是提示注入或幻觉,而是普通的过度授权——某个自主体能够做一件它本就不该能做的事。
The identity data says the same thing from the other side. Machine identities now outnumber human ones by roughly 109:1 in the average enterprise, up from 82:1 a year earlier, and something like three-quarters of those are AI agents (Palo Alto Networks 2026 Identity Security Landscape, via Help Net Security). Around 92% of organisations say their existing IAM tooling cannot manage agent identities. Meanwhile only about one in seven organisations ships agents to production with full security sign-off.身份侧的数据从另一个方向说了同一件事。企业中机器身份与人类身份的比例已达约 109:1,一年前还是 82:1,其中约四分之三是 AI 自主体(Palo Alto Networks 2026 身份安全报告,经 Help Net Security)。约 92% 的组织表示现有 IAM 工具无法管理自主体身份。与此同时,只有约七分之一的组织在自主体上生产前拿到了完整的安全审批。
Read those two data sets together and the conclusion is not subtle. The dominant enterprise AI loss mode in 2026 is not "the model said something wrong." It is "the model did something it should never have been able to do." That failure mode is entirely indifferent to whether the weights were open.把这两组数据放在一起,结论并不隐晦。2026 年企业 AI 的主导损失形态,不是"模型说错了话",而是 "模型做了它本就不该有能力做的事"。而这个失效形态,对权重是否开源完全无感。
Which is exactly why the open/closed argument matters less than it appears. Open vs closed is a debate about capability distribution. Enterprise loss is a function of capability execution. They are different layers, and only one of them is on your balance sheet.这恰恰说明开源/闭源之争没有它看上去那么关键:开源与闭源,争的是能力的分配;企业的损失,来自能力的执行。 这是两层,而只有一层记在你的资产负债表上。
5. The three burdens open weights actually add五、开源权重真正额外带来的三项治理负担
If open models aren't more dangerous, why does the open-model era make execution governance more necessary rather than equally necessary? Three specific reasons, none of which is "open models are risky."如果开源模型并不更危险,为什么开源时代让执行治理更加必要,而不是同等必要?三个具体理由,没有一个是"开源有风险"。
(a) The model became a swappable component — so policy can no longer live inside it.(a)模型变成了可插拔组件——因此策略不能再寄生在模型里。
An enterprise on open weights will change models. That is the point: it's why the cost curve works. GLM to Kimi to whatever ships in Q4. But if authorisation logic lives in system prompts, in model-specific refusal behaviour, or in per-integration guard code, then every model swap invalidates the entire safety case. Prompts get rewritten, policies get retested, compliance gets re-reviewed — and, critically, the evidence that the control worked under the old model tells you nothing about the new one.采用开源权重的企业一定会换模型。这正是重点:成本曲线就是这么算出来的。GLM 换 Kimi,再换 Q4 发布的新东西。但如果授权逻辑住在系统提示里、住在模型特有的拒答行为里、或者住在每个集成各写一份的守卫代码里,那么每换一次模型,整个安全论证就作废一次。提示要重写、策略要重测、合规要重审——更关键的是,旧模型下"控制有效"的证据,对新模型什么都不能说明。
The failure here is architectural: a governance layer that is coupled to the thing it governs. If the enforcement point is outside the model, is the same regardless of which model produced the request, and produces the same audit record either way, then models become genuinely interchangeable and the safety case survives the swap. If it isn't, the organisation has quietly made every model upgrade into a security project.这里的失败是架构性的:治理层与被治理对象耦合了。如果强制点在模型之外、无论请求由哪个模型产生都保持一致、并且产生同样的审计记录,那么模型才算真正可互换,安全论证才能跨越模型更换而存续。如果做不到,这家企业就已经悄悄把每一次模型升级都变成了一个安全项目。
(b) There is no vendor-side record — so the evidence has to be manufactured locally.(b)没有厂商侧的记录了——证据必须在本地被生产出来。
Regulatory reality does not follow the weights. The EU AI Act's open-source carve-out (Art. 53(2)) is narrow: it relieves some provider documentation duties, it evaporates entirely for systemic-risk models above the 10²⁵ FLOP threshold, and 2026 guidance has been read narrowly rather than expansively. None of it touches deployer obligations, which are what a bank or hospital actually carries, and high-risk-system requirements land in August 2026.监管现实并不跟着权重走。欧盟《AI 法案》对开源的豁免(第 53(2) 条)是窄豁免:它免除的是部分提供者文档义务,对超过 10²⁵ FLOP 阈值的系统性风险模型完全失效,且 2026 年的指引采取的是从严而非从宽的解读。它完全没有触及部署者义务——而后者才是一家银行或医院真正承担的;高风险系统要求在 2026 年 8 月落地。
So the enterprise still owes an auditor an answer to "who authorised this action, under what policy, on whose behalf, and can you prove the record wasn't edited afterwards?" With a closed API, part of that answer is at least corroborated by a third party's logs. Self-hosted, you are the only witness. A log your own operators can rewrite is a weak answer to a regulator. Evidence has to be generated as a by-product of enforcement — tamper-evident, chained, produced by the thing that made the decision — or it isn't evidence.也就是说,企业仍然欠审计一个答案:"这个动作是谁授权的、依据哪条策略、代表谁执行、你能证明记录事后没被改过吗?"用闭源 API 时,这个答案至少有第三方日志可以佐证。自托管之后,你是唯一的证人。而一份你自己的运维人员就能改写的日志,在监管面前是很弱的答案。证据必须作为强制执行的副产品被生成——防篡改、链式、由作出决策的那个组件产出——否则它不构成证据。
(c) Open weights lower the cost of more agents, not just cheaper agents.(c)开源权重降低的是"更多自主体"的成本,不只是"更便宜的自主体"。
Cheap self-hosted inference doesn't just cut the bill for the agents you have; it removes the marginal-cost brake on creating new ones. Teams spin up their own. Vendors ship theirs. Coding agents, finance agents, support agents, third-party SaaS agents, all on different models, all evolving independently, most of them without security sign-off.廉价的自托管推理不只是把你现有自主体的账单打了折,它同时移除了新建自主体的边际成本刹车。业务团队自己拉起来,供应商也带自己的进来。编码自主体、财务自主体、客服自主体、第三方 SaaS 自主体,跑在不同模型上,各自独立演进,其中大多数没有安全签核。
If each of those implements its own authorisation logic, the endpoint is guaranteed: inconsistent policy, permission sprawl, drift nobody can measure, and an audit surface that no one can reconstruct. The only structure that survives that growth is a single execution boundary that every agent goes through, where policy is defined once and enforced identically regardless of which model, vendor, or team produced the request.如果每一个都自己实现授权逻辑,终局是确定的:策略不一致、权限蔓延、无人能度量的漂移,以及一个没人能复原的审计面。唯一能扛住这种增长的结构,是所有自主体共用的单一执行入口——策略定义一次,无论请求来自哪个模型、哪个供应商、哪个团队,都被同样地强制执行。
6. What an execution governance layer actually is六、执行治理层到底是什么
The design premise is one line: treat the model as untrusted — not as malicious, but as un-loadbearing. Do not ask the model to be the thing that prevents the bad outcome. Do not put its intent, its reasoning trace, or its behavioural profile on the authorisation path at all.设计前提只有一句:把模型当作不可信——不是当作恶意的,而是当作不承重的。 不要指望模型自己成为"避免坏结果发生"的那个东西。不要把它的意图、它的推理链、它的行为画像放进授权路径。
Concretely, that means an enforcement point outside the model that answers a different question than alignment does. Not "is this request harmful?" but "is this specific action, on this specific resource, permitted for this principal, right now?" — a question that is answerable deterministically and reviewable afterwards by a human who was not present.具体而言,这意味着在模型之外存在一个强制点,它回答的问题与对齐不同。不是问"这个请求有害吗?",而是问"此时此刻,这个主体,对这个具体资源,做这个具体动作,是否被允许?"——一个可以被确定性回答、并且能被一个当时不在场的人事后复核的问题。
The mechanisms that follow from that premise (L1 level — names and design intent, not implementation):由这个前提推导出的机制(L1 层级——只讲机制名称与设计取向,不讲实现):
- Capability tokens. Authority is granted narrowly, bound to a scope, and expires — rather than an agent inheriting the standing permissions of whatever identity happens to run it. Ambient authority is the root cause behind most of that 61% over-permissioning number.能力令牌(Capability Token)。 权限被窄范围授予、绑定作用域、且会过期——而不是让自主体继承"恰好由哪个身份启动它"所带的常驻权限。环境权限正是那 61% 过度授权的根因。
- Hard gates on irreversible actions. Payments, deletions, permission changes, external disclosure. The distinction that matters operationally is not "sensitive" but reversible vs not — and irreversible actions warrant a transactional authorisation step, with the human approval itself bound to the specific action being approved.不可逆动作的硬授权闸门。 付款、删除、权限变更、对外披露。运营上真正重要的区分不是"敏感/不敏感",而是可逆/不可逆——不可逆动作应当走事务性授权,并且人类确认本身必须与被确认的那个具体动作绑定(否则"旁人替你点了确认"就是人侧的环境权限)。
- Continuous envelopes, not single-point checks. Long-running agents make thousands of individually-permissible calls that are collectively out of bounds. Constraining a session's cumulative effect is a different mechanism from checking each call.连续包络约束,而非单点判定。 长时运行的自主体会发出数千次单独看都合规、合起来越界的调用。约束一个会话的累积效果,是与逐次判定不同的机制。
- A tamper-evident audit ledger produced by the enforcement point itself, chained so that after-the-fact editing is detectable — because it is the enforcement record, not the application log, that an auditor needs.防篡改审计账本,由强制点自身产出、链式串联,使事后修改可被发现——因为审计需要的是强制执行记录,不是应用日志。
- Consuming existing identity, not replacing it. This layer sits on top of enterprise IAM and takes principals from it. Anyone selling you a replacement for your identity provider is selling you a migration, not a control.消费既有身份体系,而不是取代它。 这一层坐在企业 IAM 之上,主体从 IAM 取得。任何向你兜售"替换你的身份提供商"的人,卖的是一次迁移,不是一个控制。
None of these are novel primitives in isolation; the security field has had capability-based authority since the 1970s. What is new is the requirement to apply them to a non-deterministic, tool-using, delegating principal — one that composes actions you never enumerated, at machine speed, on behalf of a human who has gone to lunch.这些原语单看都不新鲜——安全领域从 1970 年代就有基于能力的授权。新的是必须把它们应用到一个非确定性的、会使用工具的、会再委托的主体身上:它以机器速度组合出你从未枚举过的动作,代表一个已经去吃午饭的人。
7. How we build it — and how you should check anyone who claims to七、我们怎么做的——以及你该怎么盘问任何声称在做这件事的人
This is the layer we work on. Our kernel is called PEA, and since the point of this article is that unverifiable safety claims are the problem, it would be poor form to make one here. So, objectively:这正是我们在做的那一层。我们的内核叫 PEA。既然本文的论点就是"无法验证的安全声明本身就是问题",那么在这里给出一个无法验证的声明会很难看。所以,客观地讲:
PEA is specification-first. Before the code there is a normative specification — now in its 4.10 revision cycle — in which every assumption, invariant, and threat carries a number and a paragraph. That sounds bureaucratic. It has one concrete consequence: we can be asked "which assumption does this control depend on, and how strong is that assumption?" and answer with a citation rather than an adjective.PEA 是规范先行的。代码之前先有一份规范文本——目前处于 4.10 修订周期——其中每一条假设、每一条不变式、每一条威胁都有编号和条文。这听起来很官僚。它只有一个实际后果:当有人问"这个控制依赖哪条假设?那条假设有多强?"时,我们能用引用回答,而不是用形容词。
To make that answer honest, the specification carries an assumption maturity ladder, and it is deliberately unflattering:为了让这个回答保持诚实,规范里带着一把假设成熟度阶梯,而且它是刻意不给自己留面子的:
- L0 — Asserted. We claim it. Nothing backs it but argument.L0 — 断言。 我们这么声称。除了论证之外没有支撑。
- L1 — Auditable. A reviewer can check it against an artefact.L1 — 可审计。 评审者能对着某个制品核对它。
- L2 — Measured. There is a number, produced by a test that can fail.L2 — 可测量。 有一个数字,由一个会失败的测试产出。
- L3 — Structurally discharged. The assumption is no longer needed; the design makes it unnecessary.L3 — 结构性解除。 这条假设不再需要了;设计让它变得不必要。
At the time of writing, PEA carries 17 assumptions: fifteen at L1, two at L2, none at L3. We publish that distribution rather than hiding it, and the process rule is that every revision cycle must promote at least one assumption up the ladder. A vendor with no such number is not therefore worse than us — but you should ask them for one, and notice what happens.截至撰稿时,PEA 共有 17 条假设:15 条 L1,2 条 L2,0 条 L3。 我们把这个分布公开出来而不是藏起来,并且把"每个修订周期至少促级一条假设"写成了流程纪律。没有这样一个数字的厂商未必因此就比我们差——但你应该向他们要一个,然后观察会发生什么。
Two things follow from this that we hold to publicly:由此有两条我们对外持守的结论:
A passing test suite is not a verified invariant. The reference implementation runs on the order of a thousand automated tests. That tells you the reference implementation behaves as specified on the cases we thought of. It does not establish that the invariant holds in your deployment, and we don't allow that inference in our own materials. This is a rule we broke internally once — a due-diligence document listed modules and tests without saying whether each was actually wired into the enforcement path — and correcting it is why our public coverage map now carries an integration column with three honest values: integrated, component-only, pluggable.测试通过不等于不变式已验证。 参考实现跑着约一千个自动化测试。这说明的是:在我们想到的用例上,参考实现的行为符合规范。它不能说明该不变式在你的部署里成立,我们也不允许自己的对外材料做这个推论。这条规矩我们内部违反过一次——一份面向尽调的文档列了模块与测试,却没有说明每一项是否真的接进了强制路径。修正那次错误,正是我们现在的公开覆盖图带上"集成状态"一列的原因,取值只有三个诚实的值:已集成/仅组件/可插拔。
We are not the only people working on this layer, and we've stopped saying otherwise. Microsoft's agent identity and governance stack, policy engines like Cedar, OPA and Cerbos, open authorisation proposals for agents, and several well-funded startups are all pushing on action-time authorisation. That is a good sign for the thesis and a bad sign for anyone whose pitch is novelty. Where we think we are actually differentiated is narrower: enforcement on the resource side (the database, not just the agent framework), stateful envelopes over long-running sessions rather than per-call checks, transactional gating of irreversible actions specifically, and a willingness to publish what we've measured — including when it went against us.这一层不是只有我们在做,我们也已经不再那样说了。 微软的 agent 身份与治理栈、Cedar/OPA/Cerbos 这类策略引擎、面向 agent 的开放授权提案,以及若干拿到重金的初创公司,都在推进"动作时授权"。这对这个论点是好消息,对任何以"新颖性"作为卖点的人是坏消息。我们认为自己真正有差异的地方要窄得多:强制点落在资源侧(数据库本身,而不只是 agent 框架)、对长会话施加有状态包络而非逐次判定、专门针对不可逆动作的事务性闸门,以及——愿意把测出来的东西公开,包括对我们不利的那些。
Which is a good moment for the next section.正好,接下来就是那一节。
8. What this does not do八、这一层做不到什么
We would rather lose a deal than a reputation, so:我们宁可丢单,不可丢信誉,所以:
An execution boundary does not prevent intrusion. It does not stop the initial code execution, patch your dependencies, or eliminate zero-days. In the incident referenced above, an authority layer would not have prevented the sandbox escape. What it addresses is blast radius — the distance between "one compromised worker" and "harvested cloud credentials and a weekend of lateral movement." That step, not the initial bug, is what turns an incident into a campaign. It's a real property, and a narrower promise than most of this category makes.执行边界不能阻止入侵。 它不阻止初始代码执行,不替你打依赖补丁,不消除 0-day。在前文那起事件中,一个授权层并不能阻止沙箱逃逸。它处理的是爆炸半径——是"一个工作进程被拿下"与"云凭据被收割、一个周末的跨集群横向移动"之间的距离。把一起事件变成一场战役的,是这一步,不是最初那个漏洞。 这是一个真实的属性,也是一个比这个赛道大多数承诺都更窄的承诺。
Any enforcement point can be bypassed if ambient authority survives underneath it. We learned this the useful way. We had built a governed agent environment we were fairly pleased with, and then ran an adversarial exercise where the workload held a Docker socket — a wholly ordinary thing for a build or deployment agent to hold. It escaped every containment we put in front of it: host networking, bridge, privileged mode, host root mount. Four attempts, four escapes. Nothing in our policy layer was wrong; the policy layer was simply not the narrowest point in the system, and the workload walked around it.如果底层的环境权限仍然存在,任何强制点都能被绕过。 这一条我们是用比较有用的方式学到的。我们搭好了一套自认为相当满意的受治理自主体环境,然后跑了一次对抗演练:让工作负载持有 Docker socket——这对一个构建或部署类自主体来说是再普通不过的事。结果它绕过了我们摆在它面前的每一种约束:host 网络、bridge、特权模式、挂载宿主根目录。四次尝试,四次逃逸。 我们的策略层没有任何一条写错;策略层只是不是系统里最窄的那个点,工作负载从旁边走过去了。
A governance layer that a workload can route around is theatre. This is why the enforcement point has to sit at the environment and resource boundaries, not merely inside the agent framework, and why "non-bypassable" is a property you measure at runtime rather than declare in a datasheet.一个工作负载可以随手绕开的治理层是表演。这就是为什么强制点必须落在环境边界与资源侧,而不只是在自主体框架内部;也是为什么"不可绕过"是一个必须在运行时被测量的属性,而不是一句写在规格书上的声明。
And the human in the loop is an authorisation surface too. A smaller episode, but the one that changed a spec section: an approval prompt for a consequential action was confirmed — correctly, procedurally, in good faith — by someone who happened to be at the machine, and who was not the person the policy intended. Nothing was breached. The approval simply had no binding between the human who approved and the action approved, which is ambient authority wearing a person's face. "Human in the loop" is not a control until the loop knows which human, and which action.而且,回路里的那个人,本身也是一个授权面。 还有一件小事,但它改掉了规范里的一节:一个有后果的动作弹出了确认提示,然后被确认了——流程上正确、出于善意——确认的是恰好在那台机器旁边的人,而不是策略意图中的那个人。没有任何东西被攻破。只是这次确认在"做出确认的人"与"被确认的那个动作"之间没有任何绑定,这是环境权限披了一张人脸。"人在回路"在回路知道是哪个人、哪个动作之前,不构成一个控制。
Behavioural analysis stays off the authorisation path. Anomaly scores, persona signals, and intent classifiers are useful evidence and belong in observation and forensics. The moment they gate execution, you have made your authorisation decisions probabilistic and your audit trail unexplainable. Authority should be deterministic; detection should be advisory.行为分析不进授权路径。 异常分、人格信号、意图分类器是有用的证据,属于观测与取证平面。它们一旦成为执行的闸门,你的授权决策就变成概率性的,你的审计链就变成不可解释的。授权应当是确定性的;检测应当是建议性的。
9. Five questions worth asking before the next model swap九、下一次换模型之前值得问的五个问题
For teams already running open weights in production:给已经在生产环境跑开源权重的团队:
- If we replaced our model tomorrow, which of our safety controls would have to be re-implemented or re-tested? (Anything on that list is coupled to the model.)如果我们明天换掉模型,哪些安全控制必须重新实现或重新测试?(凡在这份清单上的,都与模型耦合了。)
- What can our highest-privilege agent do that nobody explicitly decided it should be able to do?我们权限最高的那个自主体,能做哪些"从来没有人明确决定过它可以做"的事?
- Which of our agents' actions are irreversible, and which of those currently require a human authorisation bound to that specific action?我们的自主体动作里,哪些是不可逆的?其中哪些当前需要一次与该动作绑定的人类授权?
- If an auditor asked us to prove an agent action was authorised — and that the record hadn't been edited — what would we hand them?如果审计要求我们证明某个自主体动作是被授权的、且记录未被改写,我们拿得出什么?
- Could a compromised agent process reach around our enforcement point entirely? Have we tested that, or assumed it?一个被攻陷的自主体进程,能否完全绕开我们的强制点?这一点我们测过,还是假设过?
If those questions are uncomfortable, the gap is not in your model choice. It is one layer down.如果这几个问题让人不舒服,缺口不在模型选型上。它在下面一层。
Conclusion结语
Open weights are, on balance, good. They lower cost, restore data sovereignty, break vendor lock-in, and — as Hugging Face argued after being on the receiving end of an agentic intrusion — they give defenders capable tools they control, on infrastructure they own, in minutes rather than through a vetted access programme.开源权重总体上是好事。它降低成本、恢复数据主权、打破供应商锁定;并且——正如 Hugging Face 在亲历一次自主体入侵之后所主张的——它让防守方在自己拥有的基础设施上、几分钟内拿到自己可控的强力工具,而不必走一套需要审批的受控访问流程。
But intelligence sovereignty is only half of it. Owning the model gives you the ability to see and respond. It says nothing about what the agent is permitted to do in the first place. That second half doesn't come with the weights, and it used to come, partially and invisibly, from the vendor.但智能主权只是一半。拥有模型给你的是看见与响应的能力。它对"这个自主体一开始就被允许做什么"只字未答。这后一半不随权重附送,而它过去是由厂商部分地、隐形地提供的。
Open models democratise intelligence. They do not democratise trust. Trust in enterprise AI comes from something narrower and more boring: every consequential action authorised against a policy, bound to a principal, limited in blast radius, and provable afterwards — regardless of which model asked.开源模型让智能民主化。它并不让信任民主化。企业 AI 中的信任来自一件更窄也更枯燥的事:每一个有后果的动作,都对着策略被授权、绑定到主体、限定爆炸半径、并且事后可证——无论请求由哪个模型发出。
Which brings us back to the forty-line pull request. It was a good change. The cost saving was real, the data residency improvement was real, the engineer who wrote it was right. The problem was never the diff. The problem is that the thing it silently cancelled had no name in the organisation, no owner, and no line in the budget — so nothing replaced it.说到这里,回到那个四十行的 Pull Request。它是一次好改动。省下的成本是真的,数据驻留的改善是真的,写它的工程师是对的。问题从来不在这个 diff 上。问题是:它悄悄注销掉的那个东西,在这家公司里没有名字、没有归属人、也没有一行预算——所以没有任何东西替补上来。
The first step is giving it a name.
The more open AI becomes, the more that layer has to be yours. 第一步是给它一个名字。
AI 越开放,这一层就越必须是你自己的。
Sources · Vercel AI Gateway Production Index, July 2026 · OpenRouter: The Open Weight Models That Matter, June 2026 · Epoch AI: Open models lag closed by ~4 months · Qi et al., arXiv:2310.03693 · LoRA Fine-tuning Efficiently Undoes Safety Training · Fine-Tuning Defenses Are Susceptible to Simple Attacks (arXiv:2605.26526) · Kiteworks: AI Agent Security Incidents 2026 · Shattered: Agentic AI Security 2026 · Help Net Security: Machine identities outnumber humans 109 to 1 · Hugging Face: EU AI Act rules for open-source GPAI资料来源 · Vercel AI Gateway Production Index, July 2026 · OpenRouter: The Open Weight Models That Matter, June 2026 · Epoch AI:开源落后闭源前沿约 4 个月 · Qi et al., arXiv:2310.03693 · LoRA Fine-tuning Efficiently Undoes Safety Training · Fine-Tuning Defenses Are Susceptible to Simple Attacks (arXiv:2605.26526) · Kiteworks: AI Agent Security Incidents 2026 · Shattered: Agentic AI Security 2026 · Help Net Security:机器身份与人类身份 109:1 · Hugging Face:欧盟 AI 法案对开源 GPAI 的规则