The False Claim I Introduced While Fixing a False Claim
Published on: August 29, 2026
Ready for your "Oh" moment?
Ready to accelerate your breakthrough? Send yourself an Un-Robocall™ • Get transcript when logged in
Send Strategic Nudge (30 seconds)Published on: August 29, 2026
Ready to accelerate your breakthrough? Send yourself an Un-Robocall™ • Get transcript when logged in
Send Strategic Nudge (30 seconds)Green in-lane · amber a little out · red drift. Every panel is a real commit, byte-identical on recompute. Tap any panel to open its shareable receipt.
On 28 August I published a false claim about fire insurance, a reader caught it, and the correction I shipped the same hour was false too — a second error introduced while fixing the first, inside the same sentence. The repair cited Factory Mutual, the carrier whose entire model is mandating sprinklers as a condition of coverage, to argue that insurance never prevents anything. A chief actuary with twenty-five years and signed reserve opinions read the corrected version and named that one line as the thing that most convinced him we do not understand his industry. All 1,130 tests in this repository passed on both versions, and that is the part that should cost you something today: a guard fails on a forbidden token, and a false sentence that sounds true contains no token to forbid. The bill is not embarrassment. It is every paragraph standing near the bad one.
A reader who catches you inventing history in their own field stops crediting the claims they cannot check — and you never find out which deal that ended or what it was worth.
Fixing an error is where the next error gets introduced, because a repair is written in the one state most likely to produce a confident wrong sentence: concentrated, mildly defensive, moving fast to close a fault you have just been shown. What found it was not a smarter test. It was six readers, dispatched in parallel, each told to be a specific person and asked to think out loud — and the cheapest useful discovery of the week is that a reader is a measuring instrument, with a cost of about four minutes and a failure mode that does not overlap with your test suite at all.
Each course's predicted sentence was committed to the repo before the prose existed — a claim to grade, not a wish. The win condition: this post succeeds only if you run something against your own work, and fails if you finish it agreeing.
The maître d', presenting: Repaired Porcelain, Warm — sent back once for a chip and returned with a hairline crack the length of the rim, still smelling faintly of the glue. Sweet on the first touch of the lip, then grit. The second fault was introduced by the repair, in the same hand's width of clay, and nobody in the kitchen noticed until a guest turned it over.
Here is the whole sequence as a shape, dated, so the rest of this meal is expansion rather than reveal:
28 Aug 14:02 v1 "insurance never prevented a fire; it made fire a number"
↑ senior engineer, in a monologue: Factory Mutual wrote the
automatic-sprinkler standard and ran loss-prevention
engineering for a century. Contestable history. TRUE CATCH.
28 Aug 14:31 v2 "fire became countable FIRST, and the counting funded
the inspectors" ← my repair
↑ chief actuary, 25 yrs: (1) fire insurance predates rigorous
loss counting by 100+ years — the London fire offices wrote
on judgment; counting MATURED the market, it did not create
it. (2) you just cited the carrier most famous for MANDATING
behaviour to argue insurance does not change behaviour.
guard suite 1,130 tests PASS on v1 PASS on v2
reader 6 monologues caught v1 caught v2
The verdict on the repaired sentence, verbatim: a real underwriter reads that paragraph and stops trusting the rest of the history on the page. Two errors, thirty minutes apart, one of them mine and made while concentrating.
The defect was not in the draft. It was in the correction — introduced by the act of being careful, in the same sentence, while I was paying attention.
The maître d', presenting: Smoke Detector, Battered and Fried — crisp, salted, faintly metallic on the back of the tongue. Certified, fully functional, and mounted directly above a flood. The house recommends it without reservation for fires.
This repository carries 1,130 guard tests, and I wrote one for that page specifically — tests/hiring/hiring-ad-holds.test.mjs, sixteen assertions, green. It works, and later in this meal you will see the two things it caught. But look at what a test is: a predicate over text, written by someone who already knows what to forbid. Ours bans a promise about the machine's behaviour, an unsourced statistic, a link to a post that was never deployed. Every one of those is a token or a shape somebody enumerated in advance, which is the definition: a guard fails on a forbidden token. Now hold the four errors the readers found against that — a theorem cited for a claim it does not make, a strawman about an industry, a benchmark reported without its spread, and a false piece of insurance history. Not one contains a forbidden word. Each is a well-formed English sentence asserting something untrue about the world, and no pattern distinguishes that from a well-formed English sentence asserting something true. The gap is not that our tests are weak; it is that this class of defect has no syntactic marker at all — the same reason, one level up the stack, that no procedure decides a program's semantic properties. What closes it is a person who knows the domain reading the sentence and saying no, that is not what happened.
The maître d', presenting: Mirror, Polished on One Side — chilled, served facing the guest, condensation beading on the silver. The reverse is rough and unfinished, as it has never been the side anybody inspects.
Think about the last thing you corrected under time pressure — a reserving note, a rate filing, a flagged paragraph, a patch after an incident. It got attention on the way in. Now the narrower question: did anything check whether the correction itself introduced a defect? In most workflows the fix inherits the credibility of the person fixing it, precisely because they have just demonstrated care. Review looks hardest at new work and softest at repairs. In this repository, across 20,367 commits, the number of automated checks that distinguish a repair from a feature is zero — I looked on 2026-08-29, with git log --oneline | wc -l and a grep over the guard suite, because this post required me to. Repairs are the least-instrumented artifact we produce and I would bet the same holds for you, which is a bet you can settle in about ten minutes by grepping your own tracker for how many post-fix reviews exist as a distinct step.
The maître d', presenting: The Complaint, Served Warm — sharp, vinegared, unpleasant at the front of the palate and clean afterwards. The house pays for these. It has found them cheaper at the table than in the dining room, and very much cheaper in the street.
The most valuable thing anyone gave us this week cost them one paragraph. Three generalist readers praised the insurance framing; the count that matters is that it took the one reader who had signed reserve opinions to see the history was invented. That ratio — three to one, on the same file, the same day — is the finding, because term-of-art errors are invisible to a smart outsider by construction: the outsider's only available check is whether it sounds like the field, and a fabricated sentence passes that check by design. So here is what you can hand over, and it costs you a page. If you price risk, read what we claim about your trade at thetadriven.com/hiring and tell us where it is wrong; the errors above are the ones we already know about. An objection is worth most at the moment before it becomes a paragraph somebody else has read — which is also the moment it is cheapest for you to make, because nothing has been defended yet and nobody has to lose.
The maître d', presenting: A Flight of Six — poured at once rather than in sequence, six glasses sweating on the pass, on the grounds that a tasting served one at a time is not a tasting, it is an afternoon.
The mechanics are unremarkable, which is the point. Six readers, each given one person to be — a self-funded builder, a founder with no need of a salary, a senior engineer at a frontier lab, a working pricing actuary, a chief actuary of twenty-five years — dispatched in a single message so they run concurrently instead of one after another, each asked for a running monologue and three numbers: how well the page let them predict the work, how much it changed their next two days, how confident they were the people behind it are real. The absolute values are soft; the deltas against a known change are not. The hardest engineering reader moved 40 / 20 / 40 to 62 / 30 / 65 across one revision. The funded reader moved 58 / 40 / 62 to 70 / 55 / 79. The pricing actuary returned the highest confidence anyone gave the page, 82, and the lowest actionability, because she hit a wall in the application she could not climb. Total elapsed time for all six: under four minutes, dispatched as one message in 2026 rather than a for loop with an await inside it — the sequential version of this same read took roughly five minutes per document and was abandoned for exactly that reason. Read those as a series and the one number that refused to move is louder than the ones that climbed.
The maître d', presenting: A Faultless Dish, Sent to Table Nine — seasoned to the gram, timed to the second, still steaming, and ordered by table four. The kitchen has no procedure that detects this and the guest at nine eats it anyway.
Here is the failure the readers could not catch, and it is the most useful thing in this post. For four increasingly careful passes — commits e22a49635c through f4d34d88d9, all on 29 August — that page was aimed at software engineers. Every fix in those four was real. The target was wrong the whole time: the person who fits this work is an actuary who has started doing their own analysis with a model and cannot say exactly what it did. Not one of six readers could have told me, because each was told who to be before they read. A persona is an input to the instrument, so the instrument cannot interrogate it — ask a simulated senior engineer whether a senior engineer is the right audience and the answer arrives from inside the assumption. That correction came from the only participant with no assigned role. The operational residue, if you take one thing: the parameter you never vary is the one to suspect, and it is usually the audience, because it is set once and then silently inherited by every downstream decision.
An instrument cannot audit its own configuration. Six careful reads of a page aimed at the wrong person produce six careful reports about the wrong person.
The maître d', presenting: Two Small Saves, Unsalted — plain, dense, faintly sour, served without ceremony on a cold plate. The house notes that a tool doing one narrow thing reliably outlasts a tool described as comprehensive.
None of this makes tests useless, and the honest accounting matters more than the confession. The guard caught both places the page quoted the one sentence we are least allowed to say — a promise about what the machine will do — first inside a code sample, then again while narrating the bug in prose. Both were mentions rather than assertions, so the fix was to let the page declare a quoted specimen explicitly, putting the author's intent in the diff instead of leaving a pattern to guess at it. And its own firing test found a defect in the guard on the first run, on 29 August: the original version waved a match through whenever any negation sat within forty characters, so We guarantee the AI will not deviate passed — the "not" negating the machine's behaviour read as a negation of our claim. The one sentence we are least permitted to publish would have shipped under a green check. Tests and readers fail on disjoint sets. Buy both and know which is which: a test tells you a forbidden shape is absent, a reader tells you a present sentence is false, and neither will ever answer the other's question.
The maître d', presenting: The Second Bill — thin paper, still warm from the printer, itemising not the meal but the corrections to it. Most kitchens do not print this one. It is the shorter document and the more interesting.
There is a version of seniority that consists of making fewer mistakes, and a better one that consists of knowing what your fixes cost. Almost nobody occupies the second, because the data has never existed: corrections are produced under pressure, reviewed by the person who caused the fault, and absorbed into the record with no marker saying this part was a repair. A person who can state what their last ten corrections cost is holding something their entire field lacks. That is not a character trait — it is an instrument and a habit, and both were available this afternoon for four minutes and one command. The record of having done it also outlives the specific job, because a documented history of measured repair is not revocable by an employer, a funding round, or a market. The argument for why a coordinate you actually stood in cannot be taken back is You Cannot Clone a Coordinate, and it is a strange fact that almost nobody has started one.
The maître d', presenting: The Ledger, Open — brought face up, smelling of ink and cheap paper. The house finds that guests handed the arithmetic rarely ask to see it.
Only now the authority, earned by the confession above it. What made any of this auditable is the thing we sell: a record produced by something other than the process being examined. An account of a process, produced by that process, cannot contain what the process displaced — which is exactly why re-reading my own correction could never have found the fault, and why the fix is structurally a reader who did not write the sentence. That argument, with its ten sources checked, is The Record You Evicted Is Unpurchasable. Rice's theorem, 1953, closes the other door: no procedure decides a non-trivial semantic property of the function a program computes, so is this good has no answer while where did this land does — the load-bearing version is Undecidability Is the Asset, and the book works the distinction from scratch in the sixty-second test. That we grade our own prose against a reader at all comes from the walk that convicted our own marketing voice while we slept.
The ingredients, presented as material rather than conclusion: six monologues across four personas on 29 August; four errors found, one of which I introduced while repairing another; two guard trips, both correct, on a guard whose own firing test found a bug in it; scores of 40 / 20 / 40 to 62 / 30 / 65 and 58 / 40 / 62 to 70 / 55 / 79; and the page they were all reading, still up and still wrong in ways nobody has caught yet, at thetadriven.com/hiring.
The maître d', presenting: The Bill, Itemised — three lines on warm paper, none of them a subscription.
Three actionable things, in ascending order of cost. One command: npx -y thetacog-mcp@latest attest-demo — a one-minute local run on your own machine that uploads nothing and reads none of your files — returns the deterministic half of everything above: where a unit of work landed against the lane it declared, recomputable by you to the same answer, no model anywhere in the verdict. One habit: take the last correction you shipped, hand it to a reader who knows the domain and did not write it, and ask what is false rather than what is unclear — then run it against your own draft before the next one ships. One grade owed back: the win condition of this post was declared before it was written and it was never your agreement — the meal wins only if you recompute something against your own work, and fails if you leave nodding. Count how many of the nine predicted sentences actually fired in your head; they were committed to the repository before a word of prose existed, in the claims and predictions sidecar. The ones that missed are worth more to me than the ones that landed, and the argument for that is the same theorem the book opens on. Tell me which missed and you will have done to this post exactly what the chief actuary did to the sentence that started it.