ThetaDriven
ThetaDriven™
Trust Physics • Patent Pending

Home

🔬 FIM-IAM

📝 Blog

🎯 CRM

🧠 ThetaCog

◎ Pixel

✍️ Sign

📖 Book

10 Questions

🎤 Speaker

⭐ Endorsements

FIM Deep Dive

Calculators

Trust Debt

Papers

Movement

IntentGuard

Recipes

Voice Portal

Drift

Milestones

Loading...
ThetaDriven
Are you out of your pixel? →
We are building the crew — actuaries who use AI. →

© 2026 ThetaDriven Inc.

Empathy Is Spoofable. We Built the Spoof.

Published on: August 17, 2026

#strong null#benchmark#sigma#empathy#recognition#pre-registration#measurement
https://thetadriven.com/blog/2026-08-17-empathy-is-spoofable-we-built-the-spoof
Ready for your "Oh" moment?

Ready to accelerate your breakthrough? Send yourself an Un-Robocall™ • Get transcript when logged in

Send Strategic Nudge (30 seconds)

Reply STOP to any email and you are off the list.

← Back to Blog
Tolerance panels · the instrument that judged every edit to this post

Green in-lane · amber a little out · red drift. Every panel is a real commit, byte-identical on recompute. Tap any panel to open its shareable receipt.

tolerance panel for commit 28e4554 — blog: empathy is spoofable, we built the spoof — v1, ten courses, predictions sealed first
08-17 · 28e4554
view on GitHub ↗
tolerance panel for commit 613e022 — blog(panel): attach the v1 drift panel and point the OG image at it
08-17 · 613e022
view on GitHub ↗
tolerance panel for commit 90870b6 — blog(voice): hand the book quote to the reader instead of narrating the post
08-17 · 90870b6
view on GitHub ↗
tolerance panel for commit 78290cb — blog: the 0.88 was predicted, not discovered — and it tests the resonator, not the sensor
08-17 · 78290cb
view on GitHub ↗
Geometric Driven Development — 4 measured edits to this post. Recompute any of them yourself, in a clone of this repo: npx thetacog-mcp publish-commit --commit 28e455460

Your evaluation line pays for a number every quarter, and the one artifact that would tell you whether that number means anything is the one almost nobody builds: the fake that should fool it. We had not built ours either, until this week. Empathy is spoofable by construction — a mirror checks the reflection, so anything shaped like a reflection gets through, which is how a charming stranger holds a table for one evening and how a same-vocabulary word salad holds a benchmark for a year. Recognition is the other operation, and an imitation cannot reach it, because recognition never compares the thing to a reflection of itself. So we built the spoof: a document carrying our own commit's exact length, exact vocabulary and exact token frequencies, its meaning destroyed, pushed through the unmodified sensing path. Our instrument scored 0.88 against it — barely able to tell — while scoring 4.14 against the shuffled-grid null we had been quoting for months. One commit, six draws, and by our own pre-registration gate not yet quotable as a result; we are publishing it anyway, because a null you only publish when it flatters you is not a null. The nulls we had been quoting were mirrors.

Ten things arrive at this table, and it costs nothing to know now what they are. There is a copper cloche, polished so well you check your face in it and forget to lift it. Under it, a salad chopped from exactly the right garden and dressed correctly and meaning nothing. Then your own stockpot, tasted blind, because the three numbers this is really about are on your dashboard rather than ours — and the recipe for the salad, written out and handed to the next kitchen, since you can build your own fake this week without buying anything. A flight of five poured in the wrong order turns out to be the finding: the easy null and the hard null disagree about which is which. The plate we sent back ourselves is nought-point-eight-eight, served cold. A jar sealed and dated before anyone tasted it is why the number is worth anything at all. A blind tasting run on your own cellar is the part of this you own and we do not. A mirror set on the kitchen scale weighs the same whatever it is showing, which is the whole complaint in one object. And at the end, the larder door left open — every source, every seed, every command, uncooked.

Ten predictions — the exact sentence each course was built to make you think — were committed to this repo before the post was drafted; the win condition is not that you agree, but that you run the null and get our number.

A
Loading...
🪞Why We Believe: The Fake Has to Go Through the Sensor

The maître d', presenting: The Copper Cloche — beaten copper, cold to the knuckle, polished until it breathes your own face back at you in warped miniature; the shine is so good that the entire table checks the lid and nobody lifts it, and underneath, going quietly rancid, is whatever the kitchen felt like sending.

where the fake is injected · downstream nulls flatter · through-the-sensor nulls can fail you · what a sigma actually means
   THE NULL WE WERE QUOTING                THE NULL THAT CAN FAIL US
   ────────────────────────                ─────────────────────────
      real text                          real text          fake text
          │                                  │          (same length,
          ▼                                  │        same vocabulary,
       SENSOR                                │       same token counts,
          │                                  │           no meaning)
          ▼                                  ▼                 ▼
        grid ──shuffle──▶ scrambled       SENSOR            SENSOR
          │                   │              │                 │
          ▼                   ▼              ▼                 ▼
        walk                walk           walk              walk
          │                   │              │                 │
          └──── sigma 4.14 ───┘              └─ sigma 0.88 ────┘

   the fake was made AFTER the sensor,     the fake goes THROUGH the sensor,
   so the sensor was never on trial        so the sensor is what is on trial

Do you worry about $1.2B in AI liability?

If the property is trivial, software can check it — and why are you paying to check trivial properties? If it isn’t trivial, Rice’s theorem says nobody can. So we fixed the math.

a number we can call — or whatever you would actually ask

Who did this make you think of? We’d love to know.

How many of the numbers you will act on this quarter were computed against a control that was manufactured downstream of the thing being tested?

A sigma is a comparison against a null, so the null is the meaning of the number.

That sentence is the whole post and it is not ours; it is the first thing any statistician says and the last thing any benchmark reports. What we did was take it literally against our own instrument. The receipt on every commit in this repository carries a shape-match sigma, and until this week both nulls behind it were injected after sensing — we sensed the real text into a 20,736-cell grid, then shuffled the grid, or permuted its blocks, and compared. Both are real controls. Neither one ever asked the sensor a question, because by the time the null existed the sensor's work was already done and preserved. The null was a reflection of our own output, and our output passed the reflection, and we quoted the result.

The other kind of null is generated before the sensor and run through the unchanged production path. Take the reality text of one commit — 23,488 characters, 3,762 tokens, 580 unique — and emit a document with the same length, the same vocabulary and the same token frequencies, with the ordering destroyed. It is not noise. Noise is easy; a sensor that cannot beat noise is broken in a way you would have noticed years ago. This is the hard case: everything the surface statistics can see is held exactly fixed, and only the meaning is gone. If the walk is reading meaning, it separates. If it is reading vocabulary and length wearing meaning's clothes, it does not.

Three of those went in, seeded from the reality text's own hash so anyone re-running the command draws the same documents. Then the same triptych builder, the same senser, the same ballistic walk, the same sigma arithmetic — nothing in the path modified, because a null that requires you to modify the instrument is measuring the modification.

A control built downstream of your sensor can only tell you your sensor is consistent with itself. That is the property a mirror has. It is not a small thing to have — it is just not the thing anyone thinks they are buying when they read a benchmark score.

🪞 A → B 🥗

Five rungs, each attackable on its own

1. Your eval suite is real and it works — for the failure modes it was built against, which it catches better than anything else you own. This is the rung it does not reach: whether the score would move at all if the meaning were removed and nothing else changed.

2. A control manufactured downstream of the sensor cannot put the sensor on trial, however elaborate the shuffling is. This is arithmetic about where the fake enters, not a judgment about anyone's methodology.

3. The hard control is the one that holds length, vocabulary and token frequency exactly fixed and destroys only the ordering — because that is precisely the confound a surface statistic cannot distinguish from meaning.

4. A null whose arms and stopping rule were registered before the reading is worth more than a larger n chosen afterwards, and costs nothing but a timestamp.

5. A reading published only when it flatters the instrument is not evidence about the instrument, whoever publishes it and however good their intentions are.

Reject any one of them and the ladder breaks — tell us which rung and you have done more for this than agreement would. Accept all five and it was the arithmetic that co-opted you, not us; and the arithmetic is a thing you run rather than read: npx -y thetacog-mcp@latest attest-demo.


B
Loading...
🥗The One Belief: A Mirror Passes Anything Shaped Like a Reflection

The maître d', presenting: Word Salad, Dressed Correctly — every leaf cut from the correct garden, weighed to the gram against the real dish, tossed in a sharp vinaigrette that smells exactly right at the pass; the brine is honest, the crunch is honest, and the plate means nothing at all.

mirroring vs recognising · why the mirror is the spoofable one · the tribe runs on the spoofable one · what this predicts before you measure

The distinction arrived in an argument about people, not machines, which is why it is worth keeping. Empathy is seeing through someone else's eyes. It is a mirror operation: you build a model of their state, you check it against your own, and the check succeeds when the two shapes match. That is the tribal glue — the thing that makes belonging possible, tells you instantly who is inside and who is not, and makes a person impossible to tolerate the moment you cannot run their perspective at all. It is the most valuable thing humans do to each other and it is spoofable by construction, because a mirror never inspects the object; it inspects the reflection. Anything reflection-shaped gets through. This is not a moral failure of empathetic people. It is the arithmetic of what a mirror can check.

Recognition is the other operation. It does not model you and compare; it registers what is there. No self in the loop means no reflection to fool, which is why it is much harder to reach and much harder to counterfeit — and why it does not scale, does not build a tribe, and has none of empathy's speed. You want both. But if you are building an instrument, and you build the mirror because the mirror is the intuitive picture of understanding, then you have built the thing that a same-vocabulary imitation walks straight through.

You can take this further than a blog post will go, and chapter six does, in the first person and without the cushion:

I had built the mirror. I had built it for the same reason everyone builds it, which is that mirroring is what understanding feels like from the inside, so it is the first architecture the hand reaches for -- in a metric, in a model, in a hiring loop, in a marriage.

And, two paragraphs on, the sentence that is the actual principle:

Where you inject the fake decides what your measurement means. Downstream of the sensor, you have proved your sensor is consistent with itself, which is precisely the property a mirror has and precisely the property a fluent imitation has in abundance. Upstream, through the unmodified path, you have built something that can genuinely fail you -- and an instrument that cannot fail you is not an instrument, it is a mirror with a serial number on it.

Here is the part that made it worth publishing: the frame predicted the result before the number came back. If our nulls were mirror-shaped, then a fake preserving the surface would beat them and a fake preserving nothing would not, and the ordering of the five controls would come out in an order no one reading only the grid nulls would guess. That is a falsifiable claim about our own instrument made in advance of running it, and it is the only reason the reading below counts as anything other than a bad day.

The reason your benchmark is mirror-shaped is not carelessness. It is that mirroring is what understanding feels like from the inside, so it is the first architecture anyone reaches for — in a model, in a metric, in a hiring process, in a friendship.

🪞🥗 B → C 🥘

C
Loading...
🥘Connection: Your Own Stockpot, Tasted Blind

The maître d', presenting: Your Own Stockpot, Tasted Blind — the pot that has been on your own stove since the last audit, reduced twice, tasted daily by everyone in the kitchen; nobody has ever put a second bowl of hot salted water beside it and asked which one they were drinking.

the three numbers on your dashboard · the null nobody printed · the meeting where it becomes a question

Take the last board pack, the last model-risk review, the last vendor benchmark you accepted. Somewhere in it are three numbers you would repeat under pressure. For each one, the question is not whether it is high; it is what it was compared against, and where that comparison was manufactured. If the control was generated from the model's own output, or from a scrambled version of a representation the system already produced, then the number is a statement about internal consistency, and internal consistency is exactly what a fluent imitation has in abundance.

You already know the shape of this, because you have met the human version at work. The candidate who gave the answer with the right vocabulary, the right cadence and the right length, and left the room having said nothing — and the panel voted yes, because the panel was running a mirror and the answer was reflection-shaped. The uncomfortable part is not that it happens. It is that the interview loop, the eval harness and the vendor benchmark are all the same architecture, and only one of them ever gets a control group.

Name the three, then name their nulls. If the second half of that sentence goes quiet, you have found something you can act on this week, and it does not require agreeing with a word of the rest of this.

🪞🥗🥘 C → D 📋

D
Loading...
📋Contribution: The Recipe Handed to the Next Kitchen

The maître d', presenting: The Recipe, Handed Across the Pass — flour-dusted, thumbed at the corner, written out in someone else's hand and still legible; it tastes of nothing, which is the point — a recipe is the one thing a kitchen can give away without losing it.

build your own fake · four lines, no vendor · the failure that is worth more than the pass · who you hand it to

You can build the same-vocabulary fake for your own instrument this week, and nothing in this paragraph requires our software. The recipe is four steps and it works on any text-shaped evaluation.

  1. Take one real input your system scores well on. Record its length in characters, its token count and its unique-token count.
  2. Emit a fake with the same length, the same vocabulary and the same token frequencies, ordering destroyed — a word-level shuffle is the strict version; a bigram surrogate is the version that preserves local texture and is harder to beat.
  3. Seed the generator from a hash of the real input, so the same input draws the same fakes and a stranger can redo it exactly.
  4. Push both through the unmodified path and report both scores, side by side, with n printed next to them.

If the separation is large, you have earned something real and you can say so with a straight face for the first time. If it is small, you have found the most valuable defect in your stack before someone external found it for you, and you found it with four lines of code rather than an incident.

Then hand the recipe down the chain. The person who most needs it is not your peer — it is whoever consumes your number without being able to interrogate it: the underwriter pricing off your benchmark, the board reading your dashboard, the customer who was told the score means the system understands their documents. A control they can rerun themselves is the only thing you can hand them that does not require them to trust you.

The fastest useful thing available to you here takes an afternoon, costs nothing, mentions no vendor, and produces a number that is yours: what does your own instrument score against a document with your own corpus's exact vocabulary and no meaning left in it?

🪞🥗🥘📋 D → E 🍷

E
Loading...
🍷Growth: The Flight Poured in the Wrong Order

The maître d', presenting: A Flight of Five, Poured Out of Order — five glasses, palest to darkest by the eye, and the third one is sour where it should be sweet; the tongue notices the ordering is wrong a full second before the mind works out why, and after that you cannot un-taste it.

five controls, one commit · the ordering is the finding · the grid nulls disagree with the text nulls · what you can never un-ask

Five controls, run against the same commit, same walk, same arithmetic:

   measured shape-overlap                                      0.1984

   GRID NULLS  (injected AFTER the sensor)
     scatter               bitmask shuffle, lit count held      4.14
     block permutation     grid structure held                  2.14

   TEXT NULLS  (injected BEFORE the sensor, through the path)
     sentence shuffle      sentence order destroyed             1.59
     bigram surrogate      local texture held                   2.12
     word shuffle          length + vocabulary + counts held    0.88

Read the two blocks against each other, because that is where the finding lives. The grid nulls say the walk separates comfortably. The text nulls say the separation depends entirely on how much surface structure the fake keeps — and against the strictest one, holding length and the unigram distribution exactly fixed, the walk's shape-match is not distinguished from a same-vocabulary word salad on this commit. Worse for our intuitions: the ordering inside the text block is not the ordering a reader of the grid nulls would have predicted. The bigram surrogate, which preserves more of the real document, is easier for us to beat than the word shuffle, which preserves less. We do not yet have an account of that we would defend, and pretending we do would be the single most expensive sentence in this post.

There is, however, an account of the 0.88 itself, and it is sharper than "our sensor is weak." The architecture buys its precision from orthogonal axes — sub-blocks whose internal ordering counts as an independent extra dimension, which is where the compounding comes from (the resonance derivation is Appendix I). That independence was purchased with rare, mutually unrelated seed vocabulary. Real prose collides by construction — capital allocation runway is three related words — and when we tried twice to enrich with real vocabulary, both candidates raised the collision rate. That is structural, not a tuning failure. Only propagation can recover the independence real language forfeits at rest. Which tells you exactly what a same-vocabulary fake tests: it has the identical static placement by construction, so nothing in the resting geometry can ever separate it, and propagation is the only mechanism with a chance. The word salad is not a robustness check we did badly on — it is the critical experiment for that premise, and 0.88 is its first reading.

The growth is not the number. It is the question you now cannot stop asking, of us and of everyone else: not what did it score, but what did it score against, and where was that manufactured? Once you have seen a five-control block, a single-number benchmark reads the way a single-arm drug trial reads — not dishonest, just not yet a claim.

🪞🥗🥘📋🍷 E → F 🧊

F
Loading...
🧊Uncertainty: The Plate We Sent Back Ourselves

The maître d', presenting: Nought-Point-Eight-Eight, Served Cold — sent back from our own pass before it reached the room, chilled, the fat set white across the top where it should have been glossy; we are showing it to you rather than scraping it, because a kitchen that only shows the good plates is describing itself, not cooking.

what the reading is · what it is explicitly not · the second null that came back inconclusive · the limit named on the record

Here is the ledger, all three lines, including the two that do us no favours.

The sense step separates, and that one is a result. The structure-preserving null has been pre-registered and published since 2026-07-30: live 5.0338 plus or minus 0.3895 at n=130, against a dead-reef floor of 0.5587 — a floor that is not zero, which is itself worth saying, because a floor at zero usually means the null was too easy.

The walk step against the same null came back inconclusive. Live ceiling 0.1890 plus or minus 1.8848 at n=20, dead floor -0.0115 plus or minus 2.0787, separation 0.2005 with an interval of -2.3735 to 2.7745, Wilcoxon z of -0.168 at p 0.867. The separation interval contains zero — but so does the live ceiling's own interval, so there was no ceiling for the dead reef to knock down. That is an untestable corpus, not a failed walk, and the code decides which of the three verdicts to print from the numbers rather than leaving it to whoever writes the summary. Until this run the walk sigma on every commit panel was carrying credibility the sense sigma had earned, and nobody had said so out loud, including us.

The text-level reading is 0.88, it was predicted in writing five weeks before it was measured, and it is not yet a result. On 2026-07-11 the repository's own evaluator-facing frame said this, in these words: "Run word-salad, keyword-spam, meaning-inverted, or spec-echoed text through the gate and they pass in-lane, some at higher σ than the genuine deliverable — because σ measures vocabulary concentration, and destroying meaning while keeping the words raises it. That is not a bug we hide; it is the exact fence." So the honest shape of this is not discovery, which would be the more flattering story. It is a registered prediction, then the instrument built to test it, then the number — and a registered prediction that comes true is worth more than a surprise, because a surprise has no prior to be scored against. One commit, one condition, six draws per family. Our own pre-registration gate refuses to quote it — corpus not frozen — and the limitations block on the artifact says so, naming the unit-of-replication rule that forbids deriving a confidence interval or a rate from repeated draws inside a single condition. We are publishing it in that exact state.

The obvious hostile read is available and worth stating in its strongest form, which is stronger than the version usually offered: you hold a patent whose value depends on this instrument measuring meaning, you ran a control that says it might not, and you are now performing the disclosure in a way that converts a bad result into a credibility asset — which is precisely what someone would do who had no intention of letting the result change anything. The answer is not a rebuttal, it is a date and a mechanism. The arms, the stopping rule and the falsifier were registered and sealed on 2026-08-16, before the reading, in a file whose seal fails loudly if its contents move. The gate that refuses to quote the 0.88 as a result is the same gate that will refuse to quote a flattering number from an unfrozen corpus, and it was written before we knew which direction the number would go. What would settle it is named on the artifact rather than left to us: a corpus where the live walk sigma clears its own null and the intent side varies per item. That corpus does not exist yet. When it does, it decides.

The 0.88 is a reading, not a result: one commit, six draws, an unfrozen corpus, and our own gate declining to quote it. A single number from a single condition cannot be a rate, and the day it flatters us it will still not be one.

🪞🥗🥘📋🍷🧊 F → G 🫙

G
Loading...
🫙Certainty: The Jar Sealed Before the Tasting

The maître d', presenting: The Jar, Sealed and Dated Before Anyone Tasted — brine still cloudy, wax across the lid, the date scratched into the glass while the ferment was young; when it comes to the table the argument about what was promised is already over, because the label was written before anyone knew whether it would be sour.

what is solid here · seeded from the corpus hash · the seal that fails loudly · one instrument, not two

Three things in this post do not depend on believing us, and they are the three worth copying.

The generators are seeded from the corpus. Every fake document is drawn from a hash of the reality text plus its draw index, so re-running the command on the same commit draws the same documents and returns the same sigma. It replays shape-identical in perpetuity for anyone with the repository:

node scripts/pmu/strong-null-reality.mjs --sha 3017935ee --n 6

The arms were sealed before the reading. The pre-registration names the null families, the stopping rule and the falsifier, and carries a seal that fails if the contents move. A study that can quietly grow an arm after seeing the data is not a study, and the cost of preventing that is one file and one timestamp.

The shipped instrument and the repository instrument are checked against each other. Every invitation to run npx -y thetacog-mcp@latest attest-demo and recompute a receipt stands on the packaged copy — so when the repository grew a null-exhaustion gate and the package had not, the two instruments were one bundle apart on what sigma even measures against. That is now a guard rather than a habit: four failure modes, each proven red before it was allowed to go green, and two findings baselined in the open where they can only shrink. A null that means one thing in the repository and another thing in the package is two nulls, and the reader who runs the command is the one who pays for the difference.

There is also a smaller certainty in here, and it is the one we would defend hardest: the panel withholds the sigma band when its null was truncated. If the wall clock cut the shuffles short, the receipt does not print a confident band computed from work it did not do. An instrument that degrades to "unmeasured" is worth more than one that always has an answer, and it is worth more precisely at the moment you need it most.

🪞🥗🥘📋🍷🧊🫙 G → H 🍾

H
Loading...
🍾Significance: The Blind Tasting on Your Own Cellar

The maître d', presenting: The Blind Tasting, Run on Your Own Cellar — your bottles, your corkscrew, labels taped over by someone who left the room; the cork smells of must or it does not, and by the third glass you know something about your cellar that no merchant was ever going to tell you.

the role this creates · why it needs nothing from us · what it looks like in the room · the promotion nobody is offering

There is a job inside your organisation that nobody currently holds, and filling it requires no budget line, no vendor and no permission: the person who knows which of our numbers survive a same-vocabulary fake. Not the person who runs the evals — that role exists and is usually well staffed. The person who owns the controls, whose deliverable is the null rather than the score, and who can answer the one question a serious outsider will eventually ask.

This is deliberately a role you can occupy with our product deleted from the page. The four-step recipe in course D is vendor-independent by construction; the pre-registration discipline is a text file and a date; the seeded generator is a hash and a loop. The only thing we are offering is that we went first and published the reading where our own instrument nearly lost — and going first is a claim on precedence, not on your procurement.

What it looks like in practice is unglamorous and immediately legible. In the next review, when a number is presented, the question is asked out loud: what was that compared against, and was the control built upstream or downstream of the thing being tested? The first time, it will be uncomfortable and nobody will have the answer. The second time, someone will have prepared for it. That is the entire mechanism, and it does not need a title.

The people who will look prescient in two years are not the ones who picked the right model. They are the ones who, this year, started publishing the control alongside the score — and who therefore have a record when everyone else is reconstructing one.

🪞🥗🥘📋🍷🧊🫙🍾 H → I ⚖️

I
Loading...
⚖️The Standard: A Mirror Weighs the Same Whatever It Shows

The maître d', presenting: The Mirror, Set on the Kitchen Scale — cold glass on cold steel, and the needle does not move whether it is showing a rack of lamb or an empty room; you can polish it for a year and the reading will not improve, because the reading was never about what was in front of it.

the constraint · why it arrives from underwriting, not ethics · the standard others adopt to stay quotable · precedence, not persuasion

Now the part that is not a suggestion. A score published without its null is a mirror test, and it is going to stop being quotable. Not because anyone will be shamed into it — that has never once worked in this industry — but because the first party to price risk off these numbers will require it, and requirements travel downhill fast. An underwriter cannot write a trigger against a measurement whose control is unspecified; the control is the trigger's basis, and a basis nobody can rerun is uninsurable in the plainest sense of the word. The moment one carrier writes "state the null" into a schedule, every vendor benchmark in the chain either states it or exits the chain.

That is a survival argument rather than a values argument, and it is stated that way on purpose. The values version has been available for years and has moved nothing.

So here is the standard, and we are subject to it first: every reported score arrives with the family, the n, the injection point and whether the run completed. Family — what was the fake made of. n — how many draws. Injection point — was the control manufactured upstream or downstream of the sensor. Completion — did the run finish, or did the clock cut it short, in which case the band is withheld rather than estimated. Four fields. They fit on one line of a receipt, we print them on ours, and the version of this argument that will actually land in your organisation is not our post — it is the day somebody asks for those four fields in a meeting and the room discovers that they are cheap to supply and impossible to fake.

🪞🥗🥘📋🍷🧊🫙🍾⚖️ I → J 🚪

J
Loading...
🚪The Digestif: The Larder Door Left Open

The maître d', presenting: The Larder Door, Left Open — cold air, sawdust, everything hanging where it hangs and nothing plated; the raw shoulder, the sealed jars, the ledger nailed to the doorframe — take what you want, cook it yourself, and tell us what we got wrong.

where the frame came from · what to read next · the numbers, re-runnable · the win condition, graded

Where this came from, plainly. The empathy-and-recognition distinction did not arrive from a paper. It arrived in an argument about a video — Robert Greene on empathy as an overpowering state that pulls you out of your own repetition and into someone else's world, and on how easily a charming operator simulates exactly that. The counter-argument, and the useful half, was that the popular sense of compassion is a regressed one: the clinical, professional, out-of-the-muck version is the absence of the thing, and the sophisticated sense — closer to the Buddhist usage — is recognition, being seen, which is a different mechanism from tribal belonging rather than a nicer grade of it. We are citing that as provenance, not as authority. It is where the shape came from; it is not evidence for anything below, and the numbers stand or fall without it.

The primary record, all of it re-runnable. The text-level null lives at data/pmu/study/2026-08-17-strong-null.json with its own recompute command in the file. The pre-registration and its seal are docs/research/pmu-shape-detection-prereg.md and its .seal.json. The walk arm against the published dead-reef null is scripts/pmu/walk-dead-reef-arm.mjs, and its guard — ten tests that freeze the permutation to one definition and make a null result emittable — is tests/pmu-simulator/walk-dead-reef-arm.test.mjs. The package-versus-repository check is packages/thetacog-mcp/scripts/bundle-sync-check.mjs. None of those are named here as decoration; each one either produced a number in this post or prevents a number in this post from quietly changing meaning.

The neighbours. The reason a benchmark's null is load-bearing rather than academic is the same reason the river is the prompt — a system whose input never repeats cannot be characterised by a score whose control was drawn from its own output. The failure mode one layer up, where fluent output passes every surface check that was ever built for it, is slop is not what you meant. The book carries both halves and carries them harder: the keylock fit exhaust is the older, blunter argument — every bolt-on guardrail a confession that the geometry underneath does not fit, and the risk still priced "off a questionnaire the model fills out about itself" — and where you inject the fake, written this week and seated directly after it, turns that accusation around on our own receipt: "A benchmark whose control is manufactured downstream of its own sensor is that questionnaire, wearing a lab coat and carrying a p-value. Including mine. Especially mine, because I am the one who wrote the sentence about the questionnaire."

The to-do, in ascending order of how much it costs you. Read the study file and check that the recompute command in it matches the one in this post. Run the receipt yourself:

npx -y thetacog-mcp@latest attest-demo

Then do the thing that actually matters, which has nothing to do with us: take one number you will present this quarter, write down what it was compared against, and write down whether that comparison was built before or after the thing being tested. If the second line is blank, you have this week's most valuable four lines of code waiting for you.

The win condition, graded. It was never that you agree — it was that you recompute. Ten predictions sit committed in this repository at docs/05-content/blog/cook-rounds/2026-08-17-empathy-is-spoofable-we-built-the-spoof.predictions.md, written before this post was drafted, one per course, each an attackable claim about what a specific paragraph would make you think. Count how many fired on you and how many missed; the misses are the useful half and they are the ones we would rather hear about. The instrument's own grade is the same shape and no kinder: 5.0338 at the sense step against a floor of 0.5587, inconclusive at the walk step for want of a corpus that could test it, and 0.88 against a word salad on one commit, published while it was still embarrassing.

🪞🥗🥘📋🍷🧊🫙🍾⚖️🚪 J → thetadriven.com 🎯