a question, and then we stop talking

Have you ever gone for a walk, argued with a model for forty minutes, come back, and closed a two-week ticket before dinner?

If that landed as a specific memory, you have the one prerequisite we cannot teach. If it landed as a slogan, close the tab with our blessing.

And if you are an actuary, that question was aimed at you. Your market currently handles AI exposure by excluding it, which made it blind rather than safe — there is no loss history because the exclusion prevents one from forming. You already know an exposure you cannot count is an exposure you cannot price. We have made one part of it countable, on silicon, with a receipt a third party can recompute.

And that is one leg of four. There is no severity model, no aggregation framework for a defect that fires across every insured at once, no mapping from a technical crossing to a litigable trigger, and no balance sheet behind any of it. Nobody has bound a policy. Those four gaps are not the disclaimer at the bottom of this page — they are the job, and the rest of this page is why we think you are the person who closes them.

you do not have to trust any of this

npx -y thetacog-mcp@latest attest-demo

Offline, zero network, nothing leaves your laptop. Run it twice and diff the receipts: same placement, same lane, same circled regions. Shape-identical, which is the honest word — timings and paths differ between runs, so we do not claim the bytes match and one failed re-run would deserve to end the paragraph around it. What does not move is the verdict, because no model sits in that path. There is no signup wall and there never will be one: we gave the verification away and the instrument stopped needing us, since a verdict you cannot recompute is just an expensive opinion.

Three bugs we have already published against ourselves. Any of them is a first PR.

The on-chip daemon is macOS Apple Silicon today; the software witness runs everywhere and native Linux prebuilds are unbuilt. Versions at or below 2.46.0 pinned a native dependency hard, so on Node 25 and up npx aborted with exit 1 and zero output before any of our code ran — optional since 2.47.0, and a silent exit today is a bug we want, with your Node version and OS. And the roadmap still carries its own box reading the HTML dashboards do not actually read state.json yet.

We publish the failures because a countable failure is the whole physics under the word accountability — and because a company that hides them is not one you can check.

the reason we think this is worth your year

In August a diligence model was pointed at this package and returned the verdict you would expect: the mechanism is real but small — do not wire this into anything. Four questions later, same session, no new argument from us, it inverted its own headline finding.

You should not be impressed by that, and we are not either. A model that reverses itself under four questions has demonstrated that it is suggestible, which is a fact about models rather than a fact about us — and if you spotted that objection while reading the last paragraph, you are the person this ad is looking for.

The transcript is not the evidence. The two questions are, they are printed in the README, and they are falsifier-shaped: does the integer sensor call zlib at any point, and — swap two sentences so the meaning is unchanged, re-run, and report what the compression cell does and what the walk grid does separately. One should move only when the meaning moves. If it moves anyway, the instrument is measuring length and we are wrong. Go and find out; we ran diligence on ourselves and published what it found, including our own sales pitch failing the gate we sell, twice.

the artifact, not a description of the artifact

This is the receipt for the commit that created the page you are reading.

The encircled tolerance panel for commit e22a49635c: a 144 by 144 lattice of declared intent against shipped reality. Six clusters are ringed and numbered; the dense red lower-right quadrant is not ringed at all.

144 by 144. Every cell is a coordinate. Green landed inside the lane this commit declared before it ran, red landed outside it. No model touched any of it — a pure function of the commit, recomputable by you to the same shape.

Now look at what is not circled. The six rings are what the clusterer named — one holding 57 green blocks in lane, four bleeding amber, and not one of them red — while the same commit reports 41% of its lit cells off-lane, almost all of it that dense lower-right quadrant with no ring on it. Two things are true at once and we would rather print both than the flattering one. The clusterer rings contiguous blobs, and this red is diffuse, so the largest off-lane mass on the panel is the one it cannot name — a real limit, now filed. And the mass itself is this page confessing: reality fired hard in operational coordinates the intent never declared, which is the shape a page makes when it does too much.

What it is sufficient for, and what it is not. WHERE this commit landed is provable and re-runnable. WHETHER the code is any good is undecidable and we do not claim it. Every projection on this site carries that contract, because a summary that does not declare what it cannot answer is a false green waiting to happen.

And this is the guard that holds this page.

Not a description of our testing culture — the actual check, from tests/hiring/hiring-ad-holds.test.mjs, which exists because a recruiting page is the least-reread surface a company owns.

// The one claim our own perimeter forbids: a promise about what the machine will DO.
function behaviorGuarantee(text) {
  for (const m of text.matchAll(PROMISE)) {
    const [full, , negator, , tail] = m;
    if (negator) continue;              // "we never guarantee" — the rule working
    if (!OBJECT.test(tail)) continue;   // no behavior object: not this construct
    if (MEASURE.test(tail)) continue;   // detected/placed/priced/dispatched: checkable
    return full.trim();
  }
  return null;
}

test('FIRING - the overclaim is caught and the correct negation is not', () => {
  assert.ok(behaviorGuarantee('We guarantee the AI will not deviate.'));
  assert.ok(behaviorGuarantee('The architecture prevents drift.'));
  assert.equal(behaviorGuarantee('We never guarantee behavior. We guarantee the'
    + ' deviation is detected, placed, priced, and dispatched.'), null);
});

That firing test failed the first time it ran, and it was right to. The original version waved a match through whenever any negator sat within forty characters — so we guarantee the AI will not deviate passed, because the "not" negating the AI's behavior was read as a negation of our claim. The one sentence we are least allowed to say would have shipped under a green guard. That is the job: not writing the check, but proving the check can go red for the reason you think it does.

the compensation section, and it is not what you think

There is no salary. There is also no unpaid internship — that is the other thing you just guessed, and it is also wrong.

No payroll yet; that one is obvious. The second guess is the one worth killing, because "work free, get equity, trust me" is a structure where every dollar you generate lands on somebody else's desk and you find out how it went at Christmas.

A percentage, not a wage. A price list at $20 per agent-year, a fillable advisory invoice, and a capped, disclosed ladder all exist today. Bring a deployment — yours, your employer's, a stranger's from a forum — run the gate, produce the number, send the invoice, keep a cut agreed in writing before you start. This sits in the compensation section rather than under growth opportunities because it is the actual answer to how you eat. Call it what it is: commission-shaped, with the cut written down first. It is not a salary and we are not going to dress it as one — and the gap between falling crashes and rising lawsuits is a job market nobody has staffed yet.

A coordinate, not a title. All 20,367 commits here are placed — signed, re-runnable, 4,130 of them rendering their own panel of declared intent against shipped reality with the gap circled. After a month you do not have a resume line saying "contributed to"; you have a record of what you moved and where, which matters because you cannot clone a coordinate, the corner you actually stood in is the one nobody can revoke, and time on target is the one input a model cannot synthesize. It is also liquid: reach is verify, which is what turns competence into a market for humans — the second product, and you would be building the thing you stand on.

The honest floor. A fellowship means you need to be able to eat while you do it. We will not be coy about that. What we can do is make the runway problem solvable rather than permanent: the fastest way off it is one closed deal, and you get the invoice on day one instead of earning the right to sell.

The objection your chief actuary will make, so you do not have to carry it alone. Read coldly, this is a well-argued way of asking for free skilled work against a promise, and that reading is fair. The answer is paper rather than reassurance: the cut is a written term before any billable work begins, not after it lands, and nothing touches a client, a book, or real exposure data until it is signed. If your first email asks to see the terms before you spend an evening on the demo, that is the correct instinct and we will send them.

what a tuesday looks like

Roughly half of this job is done from a phone, walking.

Not a perk and not a remote-work policy — the method. The expensive part of building anything hard is knowing precisely what you want before you type, so you go outside, you argue, you get told you are wrong twice, and you come back with a spec sharp enough that the implementation is an afternoon. The manual for working at this speed is written down, failures included.

Then the afternoon is agentic, and your attention goes where a model still cannot go alone: the invariant, the boundary, the guard. There are 1,130 of those guards here and the deliverable is never the patch — it is a guard that has been watched going red for the right reason, in the same commit as the fix. Not "makes the bug impossible", which is a sentence our own rules forbid: a guard that has only ever been observed passing has been observed exactly as often as an assertion that always returns true, and three instruments here have died of precisely that. Doing it the other way has a physics: slop is what an insufficient summary reconstructs, and tagging it does not remove it.

The uncomfortable half, said now rather than in week one: a rewrite has no control group, so you never see the afternoon where you did it the other way. That is exactly why the measurement exists and why it points here first — our own drift alarm convicted our own marketing voice while we slept, and the receipt is still up.

the part everyone gets wrong about this way of working

You walk alone. Your specification does not land alone.

The obvious failure mode is nine people vanishing into nine private conversations and a very quiet Slack. It is real, we measured our own version — 3,806 asks dispatched, zero verdicts back. That number is not a badge; it is the architecture below failing completely at the one job it existed to do, and the honest state today is that the return path is instrumented and enforced rather than solved. You would be able to check that on the ledger in your first week, which is the only reason quoting it is worth anything.

Nine rooms, one shared tree, no per-room branches, because branching forces merge serialization and kills the parallelism. Each walk produces a spec; every spec goes onto a signed append-only ledger and is placed against the same 144-coordinate lattice everyone else's lands on. Your afternoon and mine are not two opinions to reconcile in a meeting — they are two coordinates, near each other or not, and that is a fact rather than an argument. It works because compression distance alone never could — the address is what makes semantics run on silicon.

Consolidation is the point and it compounds: carried meaning compounds once the map feeds itself, so your Tuesday walk makes everyone's Thursday cheaper. The counterintuitive part cost us real time to learn — we fed the machine 128KB of context and it got slower and vaguer — so consolidation means placement, never accumulation. 48.2% of the 144-cell lattice had been reached by at least one placed commit the day we measured it — a coverage figure, not a quality one, and the definition matters more than the decimal.

And walking is better with company — two people, two phones, one problem, one direction of travel. One of those walks went up a mountain and came back with a safety treaty. A funded room with actual walls is an explicit goal rather than a someday.

now the machine, because you have earned it

Ask what artifact exists, in production, showing where a specific agent action landed.

Not whether the lab tested the model — serious ones run held-out evals, sandboxed execution and red-teaming, none of which ask the model to report on itself, and pretending otherwise would be a strawman. The gap is elsewhere and it is narrower: evals are pre-deployment and statistical, they characterise a distribution of behaviour before release. They do not leave behind a per-action record that a third party who was not in the room can recompute afterwards. We have no survey to quote for how common that gap is, so do not take a number from us — check your own stack instead: when an agent did something last month that somebody now disputes, what artifact settles it, and was that artifact produced by the agent?

That is the single-witness defect, and it is structural rather than an oversight: an account of a process, produced by that process, cannot contain what the process displaced, and what it dropped is not expensive to recover, it is unpurchasable. A better model does not help — the witness is inside the boundary it testifies about — and the argument survives a good engineer because a single forward pass really is bounded, and nobody deploys a single pass.

So we read from outside. Semantic categories are packed so a semantic boundary crossing becomes a physical cache-line crossing, and a dependent load's latency is the read-out — no privileged counters, just time. Cache benchmarking is among the noisiest measurements available on an out-of-order core, so a lone number is not a receipt and our own protocol refuses to print one: it reports a median over N runs with the observed spread beside it, and gates on a control size resident in cache landing within 15% of 1.0 — outside that band the run is marked not admissible and the other rows are thrown out. A quiet run reproduced control 1.000 with 4.3x at 8MB and 6.3x at 128MB. A nine-run sweep taken deliberately under heavy concurrent load reported control 1.114x [0.668x, 8.746x], 8MB median 6.705x [4.392x, 7.797x], 128MB median 8.133x [5.757x, 10.909x] — admissible on the median, wider than we would like, and published with the spread rather than trimmed to the clean figure. And the fourth open bounty, handed to us by a reviewer who does this for a living: the median is arguably the wrong statistic. Contention only ever adds latency, so the minimum over trials is the noise-robust estimator of a physical quantity and the median is not — and a control ratio of 0.668x, below the resident baseline, says the baseline half of that ratio is itself a live noisy measurement. Nine runs, no pinned core, no outlier rejection. Show the effect survives a proper protocol and that is a real contribution; show that it does not and that is a bigger one, and we would rather you found it than a customer did. The walk itself does the reaching; it exists because the dictionary never closes — words defined by words is an infinite regress with exactly one known exit: definitions that resolve to positions instead of to more words.

Be careful about what that buys, because the careful version is the stronger one, and the loose version is a theorem-wave. Rice's theorem, precisely: every non-trivial semantic property of the function a program computes is undecidable in general. So does this code do what its specification says has no decision procedure, and if any part is Turing-complete the meaning is undecidable. It does not say "whether the work was good" — that is an engineering judgment, not a semantic property of a computed function, and anyone citing Rice for it is reaching.

The strongest objection to all of this is one you should make now rather than in month three: software has run on undecidable ground since Turing without anyone needing a decision procedure — we test, we review, we accumulate statistical confidence, and Rice does not privilege anybody's product. That is correct, and it is the design specification rather than a counterargument. Testing buys confidence about future behaviour. It does not leave a stranger a per-action record they can recompute two years later in a dispute. The missing thing is not correctness, which nobody can have. It is admissibility, which is a lower bar and an unmet one. Where it landed against a lane declared before it ran is decidable, because competence is a shape, not a score. Detected, placed, priced, dispatched — and the honest footnote on that phrase, since an actuary will spot it before the end of the sentence: it is frequency. There is no severity model here. A boundary crossing is not a loss, a count is not a claims triangle, and frequency without severity is half a rate. We are not going to pretend the other half exists.

So here is the actual state of the thing, and it is the job rather than the pitch. Countability gets you a submission worth reading; it does not get you capacity, and confusing those two is the standard mistake outsiders make about this market. Four things are missing between what we have built and anything bindable: a severity model with a development pattern you could reserve against; a framework for correlation and aggregation — which is the hard one, because a shared foundation-model defect is not idiosyncratic like a building fire, it fires across the whole book at once; the translation from a technical crossing to an indemnifiable trigger someone will litigate; and a balance sheet willing to write against any of it. Nobody has bound a policy on this. We have a sensor and no capacity.

Undecidability is the asset rather than the obstacle — the turn most people take a week to feel. We had a neat line here about insurance never having prevented a fire, and a chief actuary took it apart: the London fire offices wrote fire risk on judgment for a century before anything like rigorous loss counting existed, so counting matured that market rather than creating it — and Factory Mutual, which we had cited, is the carrier most famous for mandating sprinklers as a condition of binding. Pointing at them to argue against prevention was a self-own. Used correctly they are the better analogy: what we make is closer to a sprinkler certificate or a telematics box than to a rating plan — a loss-control instrument a carrier can require, verify and discount against. You insure the crossing, not the catastrophe; you cannot underwrite a philosopher; nothing dangerous was ever made safe enough to insure, only countable enough; and capital does not fund the uninsurable, which is why this is urgent rather than merely interesting.

the floor, and it is a floor rather than a stretch goal

The bar is not how much you know. It is how fast you can come to know something.

Cognitive velocity. Take a subject you have never touched — scale balancing, cache coherence, whatever the week demands — and be genuinely useful on it in days, because you know how to interrogate a model until a hard thing goes simple. Most people have not learned that. If you have, it beats a decade of adjacent domain experience here.

Conscientiousness about invariants. Velocity without verification is debt with better marketing. Test the smoke detector with smoke rather than admiring its wiring — generally correct and specifically wrong is the grip problem, and it is what self-improving systems cannot skip.

Appetite for the big game. Some of this is inconvenient for very large companies. You need to hold that without it eating you — when the incumbent calls you a monster, that is the compliment.

Complementary, not a copy of me. Nine rooms exist because one perspective is not enough — systems plumbing, ergonomics, test infrastructure, formal structure, design, the ability to talk to a person on a phone. Being unlike the founder is a qualification, not a tolerance.

The background that fits best is not the one you would guess. It is not machine learning. It is actuarial — someone who thinks in frequency and severity and has spent a career unable to sign anything they could not show the working for. We are not going to explain your own trade back to you: an exposure you cannot count is an exposure you cannot price is your maxim, older than this company by about two centuries, and you would be right to be irritated by anyone presenting it as a finding. What is new is only where it gets pointed. The rest of this argues from your ground, not ours: you insure the crossing, not the catastrophe; Knightian uncertainty has to become a countable frequency before anybody writes a line; you cannot underwrite a philosopher. You would not be learning that. You would be pointing it at the one exposure your industry currently handles by excluding it, which made you blind rather than safe.

Who use AI — and that half is load-bearing. Not any actuary. One who has already had the afternoon where a model turned a week of work into a morning, and who wants to know what that thing actually did rather than taking its word for it. The instinct to distrust an unaudited number is the whole qualification, and it is not teachable.

Counter-indicators, out loud. No CS degree required and its absence is not a strike — and a machine-learning background is genuinely double-edged, because fluency in treating systems as black boxes evaluated statistically fights the method here, which is measuring rather than estimating. You cannot barbell a hallucination however good your evals are. If your style is to prompt, ship, and find out later whether you can defend it, this is miserable for both of us.

the weird part, handled honestly

The cosmology is a conclusion. It is not the entry requirement.

Dig and you find very large claims — about grounding, about what a symbol is, about where meaning stops being words defined by other words. Honestly: the thread started nearer consciousness than cache lines, and we think Penrose is right about the hole and wrong about the shovel. All 303,908 words of it are public.

The question worth asking, and the one a careful reader asks here: does the engineering claim depend on the metaphysics? It does not, and that is checkable rather than reassuring. The placement is a function of a commit, a lattice and a walk; you can recompute it without holding any belief about consciousness, and if the cosmology were deleted tomorrow the receipt would render the same shape. Arrive entirely from the engineering side and be excellent here. What you cannot do is flinch when it comes up.

Why state the big version at all: database normalization is the symbol grounding problem wearing work clothes, and the bargain Codd struck in 1970 is the one you are still paying for. Pin a symbol to a position rather than to more symbols and several open questions turn out to have been one question. Checkable, written down, go break it.

how big, and what actually gates it

The constraint is not the market. We modelled it, and the constraint is people.

The insurance market withdrew cover for AI in January — ISO issued three exclusions, sixty-plus P&C groups filed, most were approved. What is supposed to replace it is roughly five products worldwide. That gap is not a forecast, it is filings, and it is the reason any of this is urgent.

So we built the sizing model properly, put every assumption on a slider, and set the defaults to rates the market has actually been observed doing rather than ones we would like. Then we asked it what moves the date. Almost nothing does. Agent density can swing five-fold and the answer barely shifts; the timing is set by a build-out somebody else is already paying for. What is left is whether the standard tips — and that is decided by whether there is someone credible enough in the room when an underwriter asks the hard question.

Which is an uncomfortable finding and a clarifying one. It means the honest version of this page is not come help us build a big thing. It is: the thing is measured, the timeline has about nine months of slack in it, and closing that slack is a hiring problem before it is a sales problem. The model says so in its own sensitivity table, which is a stranger thing for a recruiting page to be able to point at than a mission statement.

Read it and argue with it — the sequence and who it is for, and the live model with every assumption exposed. Three of the four things that have to happen for it to tip are not ours to move. The fourth is, and it is the one that needs you: measuring whether the drift gate actually saves tokens on real runs, which is currently unmeasured and is the cheapest, highest-value experiment we have. If it does not, the whole adoption curve is slower and we would rather find out from a contributor in month one than from a counterparty in year two.

how to join the crew

We do not read resumes. We read terminals, and the conversation you had with your model on the way there.

Here is the key, handed over before you start, because withholding it would only measure whether you can guess: take the problem for a long walk and ask the repository the right questions, and it will tell you everything. 518 posts, 1,130 guards, an anti-rules ledger, 4,130 commit panels and a book, all in the open. Nothing is hidden. The only thing measured is whether you do it.

1 · Fork it and run it. github.com/wiber/thetacog-mcp, then the command above. Read the output closely enough to be annoyed by something in it.

2 · Then one of two things, depending on which half you already have. Both are real applications; neither is the consolation prize.

If you write code — one pull request carrying a failing test and the guard that fixes it. Not a feature, a real defect: a boundary that is not held, a check that stays green while the thing it watches is broken, a platform where the demo dies silently. Small is fine; small and load-bearing is the whole art.

If you price risk — read what the command printed and answer the question we cannot: what would have to be true before you would put a number on this? What is the exposure base. What would you need before you would call it a countable event rather than an anecdote. Where does the definition leak, who argues about it at claim time, what would you refuse to sign, and how would you even begin on the aggregation problem two sections up. A page of that is worth more to us than a pull request, because we can write the test and we cannot supply the judgment.

The assumption the first door quietly makes, said out loud rather than discovered: it wants a terminal, a fork, and a working idea of what a test is. If you do not have those, that is a few evenings and not a career change, and it is not what we are selecting on — someone who cannot yet open a pull request but can say what makes a count admissible is further along the part that matters. Say so and we will pair on the rest.

3 · Send the walk. Paste the actual conversation you had with your model while working it out, including the parts where you were wrong — especially those. It cannot be produced by piping an issue into an agent and shipping the diff, which is why it is the part we read first.

4 · Email it. elias@thetadriven.com, subject Crew - [your name]. Twenty minutes on your own submission, mostly on how it breaks. Then bounded, milestone-defined work with the terms written down before it starts.

Plainly, rather than left to be discovered: an email asking what the requirements are, from someone who has not run the command, is an answer to the question we were asking.

evidence, last — and the red flags with it

And the question that gates anyone who does not need the money: what happens if the single maintainer disappears? Two halves, and they have different answers. Your record is the easy one: the repo is MIT and the tape is append-only, sha-named and anonymously clonable, so your coordinate lives in something you can hold a full copy of and the receipt recomputes without us. Your money is the one worth asking about, and the answer is that it does not route through us at all. We are the licensor of the runtime, not the intermediary on your engagement — the advisory work is contracted between you and the client, you invoice them, and our cut is a written term against that rather than a payment we forward to you. An invoice you hold directly does not inherit our bus factor. That is the structure, not a promise about our conduct; the whole point is to make our conduct not load-bearing, and if you want it papered before you start, say so and it gets papered before you start.

Raw material, not an argument you are meant to arrive at. Single maintainer: true. The package is about 118MB unpacked: true, and it vendors the walker binary and the corpora so the demo runs offline — unpack the tarball and account for the bytes yourself. Recently published: true. None of that is load-bearing, because you recompute the receipt rather than trusting the maintainer.

The benchmark board — adoption numbers including the unflattering ones.

The commit panels — the product measuring the company that builds it.

Tesseract Physics — normalization to the grounding problem, 303,908 words.

The blog — 518 posts, several wrong and left up on purpose.

The map — the Smith and Rice spine. The pixel — why a coordinate is the unit of a career.

US Patent Application 19/637,714 — 36 claims, filed 2 April 2026, Track One. Filed and pending, not granted; claim counts measure what was written, not what survives examination.

If you got this far and you are already thinking about which guard you would write, that was the interview.

Fork it and run the gate →

then elias@thetadriven.com, subject Crew - [your name]

ThetaDriven Inc. · the crew call · this page is graded like everything else