AI extrapolates fluency. I extrapolate emotion. AI 外推流畅,我外推情绪
I use AI to draft client correspondence, and I don’t forward what it writes. The popular division of labour, where AI handles the technical part and the human handles the human part, is backwards, and I think it will catch a lot of people out. AI and people fail in opposite directions. That isn’t a weakness of either; it’s why the pair works. 我用 AI 起草给客户的邮件,但从不直接转发它写的内容。现在流行的分工是技术交给 AI、与人沟通交给人,我认为这是反的,而且以后会让很多人栽跟头。AI 和人出错的方向正好相反。这不是谁的缺点,恰恰是两者搭配能成立的原因。
2026-08-08 · first published on boheastill.com最初发表于 boheastill.com
Two failure modes, opposite by nature
Watch either side long enough and the pattern is consistent. Both kinds of failure are extrapolation, filling a gap with something plausible, but they fill different gaps.
| The AI | Me | |
|---|---|---|
| Fails by | Fluency extrapolation. Writing in the completed tense about something that never happened. Inventing a timeframe on someone else's behalf. Producing a comparison that quietly undercuts a named person's work, just because the sentence flowed that way. | Emotion extrapolation. Reading enthusiasm as impatience, a delay as going cold, an escalation as a power grab. Filling silence with the worst story that fits it. |
| How the other side corrects it | Every "has / will / should" gets checked back against the code and the record before it goes out. | Evidence before action: go re-read the actual thread, line by line, before responding to a feeling. |
In practice both columns come up regularly, and both have caught real damage. I’ve caught drafts claiming work was done when it wasn’t, tripled time estimates offered on someone else’s behalf, and sentences that read as a dig at a colleague who would have been on the email. Going the other way, I’ve been talked out of treating a client’s silence as displeasure, and shown, by going back through the thread, that everything I’d read into it was contradicted by what was actually written.
One detail matters more than the rest: the AI had the full context and still missed things. The sentence that undercut a colleague was written with every relevant file in its context. So the review step isn’t there because one side lacks information. It’s there because every layer has its own blind spot, and the only reliable fix for a blind spot is a second layer that’s blind somewhere else.
The reading time is the product, not a tax
Spending two hours understanding a draft I could have forwarded in five minutes looks like waste until you ask what the client is paying for. They’re paying for a person who will put their name on the result. You can’t stand behind something you don’t understand. When they call tomorrow and ask what a particular margin figure means, answering clearly is part of the deliverable, and those two hours are why you can.
And the investment compounds. Once you understand a mechanism, it stays with you, and the next letter that touches it costs almost nothing. The cost for a given client falls on its own as templates and shared shorthand build up, so don’t judge the steady state by what a phase-boundary letter costs.
Here’s what I’d bet on: in two years the market will be full of people forwarding AI drafts. They’ll fail in exactly the ways in the AI column above: the invented completion, the borrowed estimate, the sentence that undercuts someone. Not because their model is worse, but because someone who just forwards has no second layer. The process is the advantage, and an unusually cheap one.
Tiered review: spend where the stakes are
Reviewing everything at full depth doesn’t scale and isn’t necessary. Three tiers, chosen by what the letter can cost you:
- S: acceptance, quotes, scope, conflict, first impressions. Understand every detail, however long it takes. With a new client the first several letters are all tier S, because the first impression sets the price of everything that follows.
- A: substantive technical exchanges. Thirty to sixty minutes on three things only: (1) every commitment, checking tense, numbers and dates; (2) a scan for anything that undercuts a person, speaks in absolutes, or estimates on someone else’s behalf; (3) can I restate the core mechanism in my own words? If I can’t, I didn’t understand it, and it doesn’t go out.
- B: confirmations, status updates, short replies. A two-minute scan, then send.
The restatement test in tier A does most of the work. It isn’t a stand-in for understanding; it is understanding, and it’s the only one of the three checks you can’t fake by skimming.
What the AI owes the review
The pair only works if the reviewer knows where to look. So every draft comes with a list of review points: the two or three places the AI is least sure about, the full list of commitments the draft makes, and every sentence that names a person or carries a tone. That one habit cut my review time by about a third without reducing what I caught, because the time now goes where the problems are most likely to be.
The rule I’d keep if I could keep only one
Every “X should happen” must be preceded by X actually happening once. Checklists, procedures, promised behaviors: all of them get written after the thing has been done at least once, never before. Written the other way round, they read exactly the same, and they’re fiction. It’s the same discipline as the rest of my work: don’t automate a flow you haven’t run by hand many times, and don’t claim a result you haven’t measured.
Written from how I actually handle client correspondence. This note was itself produced by the pairing it describes, and the first draft contained three commitments I had to go back and verify. No specific client, project or person is described.
Related: the same “make it structural, don’t rely on anyone being careful” approach applied to a machine, in AI-written code always has bugs nobody has found yet, and to money, in Governing AI agents that touch money. On keeping communication overhead low in the first place: One email per milestone (on boheastill.com).
两种失效模式,天然对偶
观察久了,两边的规律都很稳定。两种错误都是“顺着往下推”,用听起来合理的内容去填空白,只是填的空白不一样。
| AI | 我 | |
|---|---|---|
| 失效方式 | 顺着文字往下推。把没发生的事写成已经做完。替别人随口报一个工期。写出一句暗中贬低某个具体人工作的对比,只因为句子顺下来就是这样。 | 情绪外推。把热情读成不耐烦,把延后读成变凉,把升级读成夺权。用“最坏的那个说得通的故事”去填沉默。 |
| 对方怎么纠 | 每一个“已 / 将 / 应该”,发出前都回到代码和记录里核对一遍。 | 先查证据再动作:在对一种感觉做出反应之前,逐行回去读真实的对话记录。 |
实际用下来,两边的问题都经常出现,也都真的拦住过麻烦。我抓到过把没做完的事写成已完成的草稿、替别人报出翻了三倍的工期,还有读起来像在挖苦某位同事的句子,而那位同事就在收件人里。反过来,我也被劝住过,别把客户的沉默当成不满来回应;回头把往来邮件逐条看一遍,发现我脑补的每一点都和对方实际写的相反。
有一个细节最重要:**AI 拿到了全部上下文,依然会漏。**那句贬低同事的话,是在所有相关文件都给了它的情况下写出来的。所以审查这一步存在,不是因为哪一方信息不够,而是每一层都有自己的盲区,而弥补盲区唯一可靠的办法,是加一层盲区在别处的检查。
理解所花的时间是产品,不是税
花两个小时去弄懂一份五分钟就能转发的草稿,看起来是浪费,直到你想想客户到底在买什么。他们买的是一个愿意为结果署名的人。**你没法为自己没弄懂的东西担保。**明天客户打电话问某个余量数字是什么意思,你能讲清楚,本身就是交付的一部分,而你讲得清楚,靠的就是那两个小时。
而且这份投入会累积。一个机制弄懂一次就记住了,下一封涉及它的邮件几乎不花时间。随着模板和彼此的默契越来越多,服务同一个客户的成本会自然下降,所以别拿阶段交接时那封邮件的成本去估算平时的情况。
我敢打赌的是:两年后,市场上到处都是直接转发 AI 草稿的人。他们出错的方式,正好就是上面 AI 那一栏:凭空写成已完成、替别人报工期、贬低别人的句子。不是因为他们用的模型差,而是因为**只会转发的人没有第二层检查。**这套流程就是优势,而且成本低得出奇。
分层审查:把力气花在赌注所在处
对每封信都做完整深度审查,既不可扩展也没必要。按“这封信可能让你付出什么代价”分三档:
- **S 级:验收、报价、范围、冲突、第一印象。**每个细节都要弄懂,花多久都值得。新客户的前几封邮件全是 S 级,因为第一印象决定了之后的一切。
- **A 级:实质性的技术往来。**花三十到六十分钟,只查三件事:(1) 每一个承诺,核对时态、数字和日期;(2) 有没有贬低别人、把话说死、或替别人估时间的句子;(3) *我能不能用自己的话把核心机制讲一遍?*讲不出来,就是没懂,那就不发。
- **B 级:确认、进度更新、简短回复。**看两分钟,然后发出。
A 级里“用自己的话复述”这一条最管用。它不是衡量理解的间接指标,它本身就是理解,也是三项检查里唯一没法靠扫一眼糊弄过去的。
AI 欠这场审查什么
这种搭配要起作用,审查的人得知道该看哪里。所以每份草稿都附一份审查要点:AI 自己最没把握的两三处、草稿里做出的全部承诺,以及每一句提到具体人名或带情绪的话。就这一个习惯,我的审查时间少了约三分之一,抓到的问题却没少,因为时间都花在了最容易出问题的地方。
只能留一条规矩的话,我留这条
**每一条“应该做 X”,都得先真的做过一次 X。**清单、流程、承诺的做法,都要在事情至少实际做过一次之后才写,绝不提前写。提前写出来的东西读起来一模一样,但那是编的。这和我其他工作的原则一致:没有亲手跑过很多遍的流程不去自动化,没有测过的结果不去声称。
根据我实际处理客户往来邮件的方式整理。这篇笔记本身就是用文中这种搭配方式写出来的,初稿里有三处承诺,我不得不回头逐一核实。文中没有描述任何具体的客户、项目或个人。
相关笔记:同样是“靠结构保证,不靠谁小心”的思路,用在机器上,见AI 写的代码总有没找到的 bug,我照样大量用它;用在钱上,见让碰钱的 AI Agent 受治理。至于怎么从一开始就把沟通成本压低,见每个里程碑,一封邮件(在 boheastill.com)。