Governing AI agents that touch money: the agent proposes, it never executes 让碰钱的 AI Agent 受治理:agent 只提议,永不执行
One rule holds the whole design together: the agent can propose, but it can never act on anything with real consequences. Every write, whether a payment, a journal entry or an account change, leaves the model’s reach and passes through an authorization service the model can’t write or bypass, and then through a person with the right role. Governance is built into the architecture, not into a carefully worded prompt. 整个设计靠一条规则撑着:agent 可以提议,但永远不能直接执行任何有实际后果的操作。每一次写入,不管是付款、记账还是账户变更,都要脱离模型的控制范围,先经过一个模型既写不了也绕不过的授权服务,再经过一个有相应权限的人。治理靠架构来实现,不靠措辞小心的提示词。
2026-07-15 · updated 2026-08-08 · first published on boheastill.com最初发表于 boheastill.com
Why “just prompt it carefully” fails
If the model holds a tool that can move money, the safety of that money depends on the model’s judgment on every single call, including the calls where a retrieved email, a poisoned document or an ambiguous instruction pushed it somewhere you didn’t intend. A prompt asks for good behavior; it doesn’t enforce anything. The moment a consequential action is reachable from the model’s tools, the model becomes your last line of defense, and the model is exactly the part you can’t fully predict.
So the fix isn’t a better prompt. It’s to make the dangerous action unreachable from where the model lives.
The one rule, and how each guarantee is enforced
The agent reads from least-privilege, read-only sources and drafts a proposed action. That proposal, never a direct write, is all it can produce. Everything that matters is enforced in a layer the agent has no tool to reach:
- No unauthorized payment or journal. Write tools are not in the agent’s tool surface at all. Only the authorization service can reach them, and only after a role-checked human approval.
- Segregation of duties. The identity that requests an action can never be the one that approves it. The authorization service enforces this; it isn’t left to convention.
- Every figure is traceable. Each number carries a pointer to where it came from: source document, page, retrieval citation. Without that, it’s shown as “unsourced” and never presented as fact.
- Prompt injection can’t act. Retrieved emails and documents are untrusted data, never instructions; tool use is allowlisted; outputs are schema-validated. A malicious document can lie to the model, but it cannot reach a write.
- Limited blast radius. Each integration runs on its own least-privilege, read-only credential, so one compromise exposes one read-only surface, with no way to move sideways into the rest.
A worked example: where money actually escapes an escrow flow
The same habit, assuming the dangerous path will be taken and designing so it’s contained, is what separates an escrow flow that holds from one that quietly leaks. Two parties negotiate and agree, and a hold goes on the card. Drawn end to end, there are four places where the money actually escapes, each with a design answer:
- “Zero-UI” versus strong authentication. An off-session charge on a 3DS card comes back “authentication required”, and there’s no screen to authenticate on, so the hold silently never happens. The answer is to set up the mandate before the negotiation (a setup step that saves an off-session payment method), with a defined fallback for cards that still challenge.
- Hold expiry. An uncaptured authorization lapses in about a week. If capture happens after the appointment and the booking was far enough ahead, the hold has already expired. The answer is to limit how far ahead a slot can be booked, or re-authorize close to the appointment, and to treat “released because they came” and “expired because we were late” as two different events.
- Webhook retries. Payment webhooks are delivered at least once, so the same event will arrive twice, and without deduplication a retry captures the deposit twice. Worse, an endpoint that doesn’t verify signatures is a public “capture this hold” button. The answer is to verify the signature against the raw body, store each event ID and drop repeats in the database, and send idempotency keys on capture and release.
- A hashed phone number isn’t anonymous. Phone numbers come from a tiny space, so a plain hash of one can be brute-forced offline in seconds, and a “hashed key” reveals identity to anyone who can read the table. The answer is a keyed hash with a secret held on the server, which can’t be reversed without that secret.
The through-line
Whether it's an accounting agent or a payment flow, the discipline is the same. Keep irreversible actions out of reach of the part you can't fully trust. Make every consequential number carry its source. Give each integration only the access it needs. Assume the worst path will be taken, and design for it up front. It's the same rule I use on the industrial side, where safety-critical limits are re-checked on the controller and never taken on trust from the layer above. Governance that can be bypassed isn't governance.
This is how I design these systems. The approach comes from years of backend work and from building production LLM pipelines and MCP tool-calling systems, where a wrong write has a real cost.
为什么“把提示词写仔细点”注定失败
如果模型手里有一个能动钱的工具,那这笔钱安不安全,就取决于模型每一次调用时的判断,包括那些被一封检索到的邮件、一份被投毒的文档或一句含糊的指令带偏的调用。提示词只是请模型好好表现,它强制不了任何事。一旦有实际后果的操作能被模型的工具直接调用,模型就成了你最后一道防线,而模型恰恰是你最没法完全预测的部分。
所以解法不是更好的提示词,而是让危险动作从模型所在的地方根本够不着。
那一条铁律,以及每条保证怎么被强制
agent 从最小权限、只读的数据源读数据,起草一个建议操作。它能输出的只有这份建议,绝不是直接写入。所有关键的控制,都放在 agent 没有任何工具能碰到的那一层:
- **没有未授权的付款或记账。**写工具压根不在 agent 的工具面里。只有授权服务能触达它们,且必须先过一个按角色核验的人工审批。
- **职责分离。**发起操作的身份,永远不能同时是批准它的身份。这一点由授权服务强制执行,不靠大家自觉。
- **每个数字都能溯源。**每个数值都带着来源信息:出自哪份文档、第几页、检索引用。没有来源的,就标成“无来源”,绝不当作事实展示。
- **提示注入动不了手。**检索来的邮件和文档是不可信的数据,不是指令;工具调用走白名单;输出经 schema 校验。一份恶意文档可以骗过模型,但它够不到任何写操作。
- **影响范围有限。**每个集成都用自己独立的最小权限只读凭据,一处被攻破,也只是暴露一个只读的接口,没法横向蔓延到其他部分。
一个实例:钱到底从一条托管支付流的哪里溜走
同样的思路,假设危险路径一定会被走到,再把它设计成可控的,正是“守得住钱的托管支付”和“悄悄漏钱的托管支付”的区别。双方谈价、达成一致,然后在卡上做一笔预授权。把整个流程从头到尾画出来,钱实际会在四个地方流失,每个地方都有对应的设计:
- **“无界面”遇上强认证。**对一张 3DS 卡发起离线扣款,会返回“需要认证”,可根本没有界面让用户认证,于是那笔预授权悄无声息地没做成。解决办法是在谈价之前先完成授权(一个保存离线支付方式的准备步骤),再为仍然需要验证的卡准备明确的备用流程。
- **预授权过期。**没扣款的授权大约一周就失效。如果扣款发生在预约之后,而预约又订得比较远,那笔预授权早就过期了。解决办法是限制最多能提前多久预订,或者临近预约时重新授权,并把“客人到店而释放”和“我们太晚而过期”当成两种不同的事件处理。
- **Webhook 重试。**支付 webhook 至少投递一次,同一个事件肯定会来两次,不去重的话,一次重试就会把押金扣两遍。更糟的是,不验签名的接口等于一个公开的“扣掉这笔预授权”按钮。解决办法是用原始请求体验证签名,把每个事件 ID 存进数据库、在数据库层丢掉重复的,扣款和释放时都带上幂等键。
- **哈希过的手机号并不匿名。**手机号的可能取值很少,普通哈希离线几秒就能暴力破解,所谓“哈希后的键”,任何能读到这张表的人都能还原出身份。解决办法是用服务器上保存的密钥做带密钥的哈希,没有这把密钥就无法还原。
贯穿其中的那条线
不管是记账 agent 还是支付流程,原则都一样。不可逆的操作,要放在你没法完全信任的那部分够不到的地方。每个有后果的数字,都要带着来源。每个集成只给它真正需要的权限。假设最坏的情况一定会发生,提前为它做设计。这和我在工业项目里用的是同一条规则:安全相关的限值在控制器上重新核对,绝不直接相信上层传来的值。能被绕过的治理,就不算治理。
这就是我设计这类系统的方式。这套做法来自多年的后端工作,以及搭建生产级 LLM 流程和 MCP 工具调用系统的经验,在那些系统里,一次写错就有实际损失。