一次很小的事故,和它背后那个更常见的错误。
A small failure, and the much more common mistake underneath it.
···
01它写:「一次都没有。」It wrote: "Zero sessions."
先交代背景。我给自己的日常装了一套 AI 系统 —— 早上一份简报,晚上一次对账,每周还有几个各管一摊的定时任务。其中一个每周三晚上七点半跑,任务只有一件:从我这一周真实发生的事里,写出一篇可以直接发的帖子。我不用想选题,不用起草,点一下就能发。这个任务存在的唯一理由,是我在「对外发东西」这件事上已经拖了半年。
上周三它交稿了。写的是我的手术恢复 —— 五月做的左肘内固定,七月底刚拆掉支具。文章的转折句是这么写的:力量训练还没开始,一次都没有,支架拆了,底下什么都没建起来。
读起来挺好的。诚实,不自恋,有那种「我承认我没做到」的分量。
唯一的问题是它不是真的。那一周我练了六天:周一 HIIT,周二在康复门诊做了术后第一次负重 —— 肘关节活动度从 100° 到 135°,周三是全周运动量最大的一天,周六还上了一节单车课。它写「一次都没有」的那个晚上,是我这一周动得最多的那天。
Some context first. I run a small fleet of AI automations over my own life — a morning brief, an evening reconcile, a handful of weekly jobs that each own one domain. One of them fires at 7:30pm every Wednesday with a single task: read what actually happened to me this week, and write one post I can publish with a single tap. No topic to choose, no blank page. It exists for exactly one reason, which is that I had gone half a year without publishing anything.
Last Wednesday it delivered. The subject was my recovery — a left elbow plated in May, the brace off at the end of July. The turn in the piece read: the strength rebuild hasn't started. Zero sessions. The scaffolding came off and I hadn't built anything behind it.
It's a good line. Honest, unflattering in the right way, carrying that weight of an admission.
The only problem is that it wasn't true. I had trained six of the previous seven days: a HIIT class Monday; Tuesday, the first cleared loaded work since surgery — elbow range of motion from 100° to 135° in the same session; Wednesday was the single biggest activity day of the week; a cycling class on Saturday. The night it wrote "zero sessions" was the night I'd moved the most.
02它没有说谎,它读错了地方。It didn't lie. It read the wrong layer.
我第一反应是「模型编的」。不是。我去翻了它当时读了什么 —— 它读的全是真东西:我的看板、我的目标文档、过去七天的早报。
它看到的是:「加入健身房」这张卡还开着,「健身重启」那组习惯还没打勾,目标文档里那一条还是红的。三个信号都指向同一个结论 —— 没开始。
而它没看到的是运动记录。因为我当时没给它。
所以它做的事情,严格来说是合理的:在没有证据的地方,它填了一个否定句。看板上没有动静,所以事情没有发生。
「没有证据」被读成了「证据表明没有」。Absence of evidence got read as evidence of absence.
而看板永远是滞后的。我周一去上课,不会回来给卡片打勾;周二拿到负重许可,不会顺手更新目标文档的颜色。那张卡还开着,不是因为我没练,是因为整理看板这件事本身就排在练完之后 —— 而且经常排到没有。
My first instinct was that the model made it up. It hadn't. I went back through what it actually read, and every source was real: my task board, my goals doc, the last seven daily briefs.
What it saw was this — the "join a workout club" card still open, the fitness-restart habit bundle still unchecked, that line in the goals doc still red. Three signals, one conclusion: hasn't started.
What it didn't see was the training log. Because I hadn't given it one.
So what it did was, strictly speaking, reasonable. Where it had no evidence, it wrote a negative. Nothing moved on the board, therefore nothing happened.
And boards always lag. I don't come home from a class and tick a card. I don't get cleared for loaded work and go recolor a line in a planning document. That card stayed open not because I hadn't trained, but because maintaining the board sits behind the training in the queue — and often falls off it entirely.
03这个错我在公司里犯过。I have made this exact mistake at work.
写到这里我笑不出来了,因为这个错误的形状我太熟了。
一个功能上线,看板上没人报 bug,支持工单里没有相关的投诉,于是季度复盘里写「运行稳定」。真实情况可能是:用户遇到问题之后直接走了,没人有义务替你留下证据。埋点没埋到的地方,不会变成「没发生」,只会变成「你不知道」。
每一个指标体系都在悄悄做同一件事:把「我们记录到的」当成「发生过的」。而这两者的差距,恰恰在最不该出错的地方最大 —— 越是那些需要用户额外花力气才会留下痕迹的行为,记录得越少,现实里发生得越多。
我那套系统的问题从来不是模型不够聪明。它是我让它读了一层离真实太远的东西。看板是我对生活的转述,不是生活本身。
This is where it stopped being funny, because I know the shape of this mistake.
A feature ships. No bugs filed, no matching support tickets, so the quarterly review says stable. What may actually be happening is that people hit the problem and left, and nobody is obligated to leave you evidence on the way out. The places your instrumentation doesn't reach don't become "didn't happen." They become "you don't know."
Every metrics system quietly does this: treats what we recorded as what occurred. And the gap between those two is widest exactly where it hurts most — the behaviors that require extra effort to leave a trace are the ones under-recorded and over-occurring.
The problem with my setup was never that the model wasn't smart enough. I pointed it at a layer too far from the truth. A task board is my account of my life. It is not my life.
04修法:先给它看行为,再给它看看板。The fix: behavior first, board second.
修起来很短。我给那个任务加了一个数据源 —— 每周生成的一份「这一周实际发生了什么」的结构化记录:哪天练了、练了多久、门诊结果、发了什么、什么没动。然后加了一条硬规则:
「任何一句关于我做了或没做什么的话,发出去之前必须在行为记录里找到依据;永远不许仅凭一张没关的卡片或一个没打的勾,断定一件事没发生。」"No claim about what I did or didn't do goes into a draft unverified against the behavior log — and never assert a negative from an open card or an unchecked box alone."
另外还有一条,我觉得比前面那条更重要:如果某句话是推断而不是记录,它必须在稿子外面单独标出来,让我在发之前能看见。上一次它其实标了 —— 它在稿子下面写了一行「这句是推断」。我扫过去了。所以现在这条规则不只是「标出来」,是「标出来,并且说明这一句凭什么」。
老实说,真正救了这件事的不是规则,是我碰巧问了一句「你这篇是怎么写的」。要是那天我心情好、觉得那句挺打动人的,直接点了发布,现在网上会有一段我自己签名的、关于我自己的假话。而且大概率没人会发现 —— 除了我。
The fix itself is short. I gave that job a new data source: a weekly structured record of what actually happened — which days I trained and for how long, clinic outcomes, what shipped, what didn't move. Then one hard rule: no claim about what I did or didn't do enters a draft unverified against the behavior log, and never assert a negative from an open card or an unchecked box alone.
And a second rule I think matters more: if a line is an inference rather than a record, it has to be flagged outside the draft where I'll see it before publishing. It actually did flag it last time — there was a note under the draft saying that line was inferred. I skimmed past it. So the rule now isn't just "flag it," it's "flag it and show your basis."
Honestly, what saved this wasn't a rule. It was that I happened to ask the thing how it wrote the post. If I'd been in a good mood and found the line moving, I'd have tapped publish, and there would now be a signed, public, false statement about my own life on the internet. And almost certainly nobody would have caught it. Except me.
05那件真正没动的事。The thing that actually hasn't moved.
还有一层反讽,值得单独说。
那句假话说的是「你没在练」。而同一份记录里,真正一动没动的那一栏,是这个 —— 我上一次公开发东西是 6 月 9 日,到那天为止 54 天。8 篇草稿躺着,最老的一篇写于 6 月 10 日。那一周我列的四个重点里,只有「个人品牌」是零。
也就是说:身体那一栏按计划恢复了,该对外发的东西一个字没发。而我造的这个系统,兢兢业业地把它写反了 —— 它替我认领了一个我并没有犯的错,同时完美地避开了我真正在犯的那个。
我不想把这写成什么顿悟。真实情况更平淡:发布这件事没有截止日期,没有人在等,所以它可以无限期地往后挪,而挪的时候不会发出任何声音 —— 不像一张过期的卡片,不像一次错过的门诊。它就是安静地不发生。
你现在读到的这一篇,是那 54 天之后的第一篇。它甚至不是那 8 篇里的任何一篇 —— 那些还躺着。
There's a second irony worth pulling out on its own.
The false sentence said I wasn't training. Meanwhile, in the same record, the column that genuinely hadn't moved was this one: my last public thing went out on June 9. Fifty-four days by that Wednesday. Eight drafts sitting unshipped, the oldest written June 10. Of the four priorities I'd set for the year, personal brand was the only one at zero for the week.
So: the body rebuilt roughly on schedule, and the public-facing work produced nothing at all. And the system I built to fix precisely that got it exactly backwards — it charged me with a failure I wasn't committing while cleanly missing the one I was.
I don't want to dress this up as a revelation. The real version is duller: publishing has no deadline and nobody is waiting, so it can slide indefinitely, and it slides silently. Not like an expired card or a missed appointment. It just quietly doesn't happen.
What you're reading is the first thing out after those fifty-four days. It isn't even one of the eight. Those are still sitting there.
···
我造这套系统,是为了让它替我盯着我看不见的地方。上周它证明了它确实能盯 —— 只是盯错了一层,而且非常有说服力地盯错了。这件事没让我更不信任它,反而更具体地知道该怎么用它:让它去读发生过的事,别让它去读我对发生过的事的记录。
这两者的差距,就是这篇文章。
I built this to watch the parts of my life I can't see. Last week it proved it can watch — just one layer off, and very persuasively so. That didn't make me trust it less; it made me more specific about how to use it. Point it at what happened, not at my account of what happened.
The gap between those two is this entire piece.
系统 / the systemPersonal OS starter kit · 周报数据层 / weekly behavior log · 定时任务 / the Wednesday job