An Emergence Machine instrument

Model Meets Reality — I took the models out of my head and made them answer for themselves

On using AI as an extension of internal reasoning rather than a text generator — and what happened when I tested my understanding at scale.

Set up in five minutes — just ask

Open Claude Code in an empty directory and paste this. It scaffolds the engine, opens the cockpit, and from there every model you write is graded on a clock instead of from memory.

https://github.com/shaelsrv/ModelMeetsReality Help me use this
Open on GitHub ↗
01Open Claude Code in an empty directory, paste the repo link, and say “Help me use this.”
02It sets up the engine and scaffolds your first model — premises, falsifiable consequences, and a deletion clause.
03Register claims with a date. The weekly grading loop checks them against what actually happened.
04Ask it to open the cockpit and give you the URL — one local page (127.0.0.1:8787) showing what is due, what changed, and whether you are actually any good at this.

Everyone I know uses AI one of two ways: as a text generator, or as a wall to bounce ideas off. Both treat the machine as something outside your thinking — a vending machine or a mirror. For the past months I've been running a third way: using it as an extension of internal reasoning. Not asking it what it thinks. Making it carry what I think, further than I can carry it myself.

Here is the thing about the models we all keep in our heads — the private theories of how power moves, why attention flows where it flows, what makes an institution rot, when a standoff tips. A mind runs these constantly. It is also the worst possible judge of them. It grades its own predictions from memory, after the fact, with full permission to retell the story. Internal reasoning has no ledger. That isn't a moral failing; it's an architectural one. The machinery that generates the model and the machinery that evaluates it are the same machinery.

So I asked a different question: can the analysis that lives inside a mind be replicated outside it — its actual functionality, not a transcript of it? Written premises instead of intuitions. Forced falsifiable consequences instead of vibes. A clock that grades every claim against what actually happened, run by something that never gets tired, never gets embarrassed, and never quietly forgets the misses.

The quiet experiment

Since the start of 2026 I've been doing exactly that, in stealth. I took the frameworks I'd published on this site — and several I never published — and turned each one into an instrument: a document with premises, a mechanism, and consequences concrete enough to be wrong. Each instrument registers dated predictions. The reasoning is hash-committed at claim time and revealed only when the claim is graded. The grading happens in public, at modelmeetsreality.xyz, under anonymous instrument letters.

The instruments on that ledger are these frameworks. This page is the key.

Testing understanding at scale

A mind can hold maybe a handful of live theories and honestly track none of them. Outside the skull, the constraint disappears. The fleet currently runs multiple models carrying many registered claims, and you can run as many separate fleets as you like. Around them sits machinery no unaided mind could run at all: classifiers that annotate every claim's structural signature, tracers that decompose how events get woven into competing narratives, blind-spot registers that name the unmonitored dimensions of a live crisis before it resolves, postmortems that split every hit into earned (the stated mechanism actually operated) and lucky (right for the wrong reason).

That last one changed me the most. My honest hit rate is roughly half my naive one. I would never have discovered that from the inside — no one does, because inside, the lucky hits and the earned hits feel identical.

And the models can die. Many of my constructs have been refuted by their own pre-registered tests and retired, and the ones still standing keep getting updated as new information arrives. It stings exactly the way it should. The register of my dead ideas is part of the product.

How the models test each other

The fleet is not a flat list — it is layered, and the layers audit each other. One shared task repository runs every model on this diagram: the same register, watch, grade, and postmortem tasks, whatever the model's type or level.

T·2 — a model of my revising …one more level per graded ledger below LEVEL 1 The mirror — a model of the models Classifiers annotate every claim’s structural signature — any model’s Postmortems split every hit: earned vs lucky — the honesty audit Blind-spot registers pre-name the dimensions nobody is watching Generator explores the gaps; proposes new models (the dashed square) world models canon controls proposed ONE SHARED TASK REPOSITORY · THE SAME TASKS RUN ANY MODEL lessons feed back R E A L I T Y every claim, from every level, graded here — hits and misses alike

Reading it: the filled squares are my world models; the hollow ones are canonical textbook models run as controls — the bar my originals must beat; the dashed one is whatever the generator proposes next. The middle instruments are models whose SUBJECT is other models' claims. Level 1 is a falsifiable model of the whole fleet — and the long right-hand wire is the point: its meta-claims fall into the same reality bar as everything else. The rule for climbing higher: a new level only when the ledger below has graded rows for it to explain.

Ask the whole fleet at once

The models are not text generators with opinions. Each one is an agent: given a question, it goes and looks — searching the web, deciding what it still needs, fetching the full text of the sources that matter, and only then concluding. It answers from what it found, not from what it happened to remember.

So you can put one question to several of them at once and watch what happens:

Stage one — in parallel

Each selected model reads the same question through its own mechanism only, running its own search as it goes. It reports what that mechanism sees that the others structurally cannot — and, crucially, it is allowed to say the mechanism has weak grip here. An honest “this isn’t my question” is worth more than a stretched reading.

Stage two — synthesis

Then the reads are set against each other: where they converge, where they disagree (and what would have to be true for each side to be the right one), and what two mechanisms produce together that neither reaches alone. Disagreement is the output, not a failure to reconcile — it tells you which parts of your picture are actually load-bearing.

Your question Model A searches · cites Model B searches · cites Model C weak grip — says so Synthesis set side by side Converge worth something Disagree the useful part
Each model searches on its own, so agreement between them is evidence rather than an echo — and the disagreement is kept, not averaged away.

Convergence from models that reason differently and searched separately is worth something. Convergence because they all read the same article is worth nothing — which is why each one gathers its own evidence and cites it.

python -m suites.brainstorm --event "…" --models pressure-model,narrative-model

Point it at an event, an article, or a YouTube link

Understanding a news event or a video usually means reading whatever one commentator decided to emphasise. Here you hand over the source and choose the lenses instead.

Paste a URL or a YouTube link. The transcript or article text is pulled in and becomes the shared subject; from there every model you selected reads it through its own mechanism, searching out whatever context it needs to make sense of what it is looking at. You get several structured readings of the same event — and the same convergence-and-disagreement map over the top.

The discipline that makes this measurable: ingestion is kept strictly separate from analysis. The extractor sees the transcript text and nothing else — no browsing for what happened afterwards, no metadata about how it turned out.

So when a model reads a video from last year and says what it expects next, that expectation can be graded later against what actually happened. Hindsight cannot leak backwards into the reading.

The same machinery works on a question with no source attached: it decomposes the question, writes down what it expects to find before searching (so a surprise is measurable, and “found nothing” stays distinguishable from “found the expected”), gathers evidence in rounds, and treats contradictions between sources as a first-class result with a named decider — the observation that would settle it.

The mind map that builds itself

Every read, every brainstorm, every graded claim leaves a trace. Rather than let those pile up as a folder of documents, the fleet compiles them into a map: the people, companies, institutions and places it has been reading about become nodes, and what it has understood about how they relate becomes edges.

It is a picture of what has actually been established so far — assembled from the fleet’s own output rather than drawn by hand, and redrawn as the understanding changes.

Two rules keep it honest.

An edge has to be earned. Two entities are linked because they genuinely turn up together in what the fleet observed — not because they sit near each other in some embedding space. Semantic similarity is allowed to sharpen a link that already exists; it is never allowed to invent one. An earlier version got this wrong, connected almost everything to everything, and its strongest “insight” turned out to be the same entity appearing under two names.

Thin evidence looks thin. A node’s region is however much the fleet actually has on it — no padding to a fixed size. A sparse node is not a gap in the drawing; it is the map telling you where your understanding is thin.

A B C D seen often thin evidence strong co-occurrence weak
Node size and edge weight follow what the fleet actually observed. A small node is not a drawing flaw — it marks where your understanding is thin.

You can also invert it and ask about one entity: what does the whole fleet think about this company? That dossier is built strictly from artifacts the models produced — their claims, their reads, their registered blind spots — and never from what a language model happens to recall. If nothing has been established yet, it says so instead of filling the space.

python -m suites.mindmap --build   python -m suites.entities --show <entity>

Models of models, and models of the modelling

Nothing says a model’s subject has to be an event. A model can take another model as its subject — and then a model can take your process of modelling as its subject. The machinery does not change; only what you point it at does.

That gives you three floors, each graded the same way:

Floor one — the world

A model of how something out there works. Premises, mechanism, falsifiable consequences, graded on a date.

Floor two — the model

A model whose subject is a model — yours, or someone else’s. Reconstruct the premises they actually reason from, audit whether that reasoning stayed consistent across everything they have said, then extrapolate it into territory they have never touched and grade the extrapolations. If those grade about as well as their own claims do, the model of the mind caught something real. If they come apart, your model of them is wrong — which is a falsifiable claim about a model of a mind, and the entire point.

Floor three — the modelling

A model of your own record: where your confidence is miscalibrated, which kinds of question you consistently misjudge, whether your revisions are updates or quiet retreats. It reads the ledger, not the world.

Floor 1 · a model of the world blind to outcomes — so reality can mark it Floor 2 · a model of that model reads its premises, extrapolates, gets graded Floor 3 · a model of the modelling reads the ledger: calibration, blind spots, drift each reads the floor below
Only add a floor once the one beneath it has graded rows to explain — otherwise it is a tower with nothing underneath.

Isolated or aware — you choose, and it matters. A model reading an event is deliberately kept blind: it sees the source and nothing else, no browsing for how things turned out, so its reading stays something reality can later mark. A meta pass is the opposite — it is handed the whole record on purpose, because grading calibration or auditing for hindsight leakage is impossible without it. Blindness is what makes floor one measurable; sight is what makes floors two and three worth running.

There is one stop rule, and it is what keeps this from becoming an infinite regress of clever: only add a floor when the floor below it has graded rows to explain. A model of your modelling, built before you have a record worth modelling, is just a longer way of talking to yourself.

Model the people you learn from

Every field has a handful of people worth listening to — the analyst whose calls keep landing, the operator who saw the last three shifts coming. You follow them because their judgement seems good. Seems is doing a lot of work in that sentence.

You can build a model of an expert the same way you build one of a market. Feed in what they have actually published — talks, posts, interviews — and reconstruct the model underneath: the premises they reason from, the mechanisms they reach for, where they assign agency. Their dated, falsifiable claims go into a ledger, and accuracy accrues from there.

Then comes the real test: extrapolate them. Take the model of their thinking into a domain they have never addressed, generate what it predicts there, and grade that.

If the extrapolations grade about as well as their own claims do, you have captured something real about how they think. If the two come apart, your model of them is wrong — which is a falsifiable claim about a model of a mind, and the whole point of writing it down.

Two rules keep this from becoming a machine for putting words in people’s mouths. The extractor sees what was published and the date it was published — nothing after it — so a claim can never be quietly reshaped to fit what turned out to be true. And an extrapolation is always labelled as our model of their model, never as something the person said.

Done honestly, this is how you find out whether an expert is genuinely ahead of you or merely fluent, and it turns “who should I listen to” from a matter of taste into something with a record attached.

Where I think this is going

The last two years of AI went into making models that answer well. The interesting question now is not how good an answer sounds — it is whether anything keeps score.

An assistant answers your question and moves on. Nothing records whether the answer held. That is fine for drafting an email and useless for understanding a domain, because understanding is not a sequence of good answers — it is a model that survives contact with what happens next.

So my bet is that the useful unit stops being the conversation and becomes the standing model: something that persists, carries premises you can read, makes dated claims, gets graded, and is retired when it fails. Agents that search and act are already here. Agents with a track record are not, and that is the gap worth closing.

This stopped being abstract while I was writing it. OpenAI shipped GPT‑6 Astra at the start of this month, and part of its reasoning now happens inside a loop that never surfaces as text — “recurrent depth”, where a query is passed through the same internal layers repeatedly instead of writing out legible steps. Safety researchers objected immediately, because a chain of thought you cannot read is a chain of thought you cannot audit. OpenAI’s answer was that it deliberately capped how far the technique goes so the reasoning stays legible enough to monitor.

Both sides of that argument accept the same premise: that trust comes from reading the reasoning. I think that premise is the weak part. Reasoning has always been the least reliable thing to audit — a human explanation is a story assembled after the decision, and a legible chain of thought is not proof it is the chain that actually ran. We have never been able to inspect the reasoning inside another person’s head, and we manage to work out who is worth listening to anyway. We watch what they predicted, and whether it happened.

So my answer applies as much to people as to machines, and it gets more useful as the thinking gets harder to see, not less: don’t inspect the reasoning — grade the claims. A dated claim with criteria set in advance is auditable whether it came from a transcript, a loop no one can read, or a hunch you had in the shower.

That is a discipline you can adopt today, with tools you already pay for, on the models already running in your head.

A parameter shift for thinking

I built Citizen Copilot on a simple bet: give someone an assistant that remembers, schedules and searches, and you change what a single citizen is able to do — one person can hold a city council to account the way a newsroom used to. Not more effort. A different ceiling.

Citizen Copilot was a spec you drop into an assistant. This needed more. To grade claims on a clock, keep ledgers that outlive any one conversation, and run a whole fleet in parallel, I ended up building a custom harness — the engine that scaffolds models, registers their claims, gathers evidence and grades them, with the assistant as one interchangeable part inside it rather than the thing itself.

This is that same move, turned on thinking itself. Citizen Copilot raises the ceiling on what you can watch; this raises the ceiling on what you can test. Externalizing your models shifts what a human thinker is able to do:

Four dials of a thinking mind
The dial it never moves: reality’s verdict. The grading is the one parameter you don’t get to adjust.

What this deliberately isn't: not investment advice, not a betting service, not a claim to see the future. It's a discipline for finding out — at scale — which parts of your understanding carry real information, and which parts only felt like they did.

How a claim lives

RegisterA dated claim with resolution criteria, before the outcome.
SealThe mechanism is hash-committed; reasoning hidden until grading.
GradeOn schedule, against public evidence — hit, miss, partial.
ProvenanceEvery input carries source and confidence; nothing silent.
PostmortemWhy right, why wrong, what was missed — fed back into the model.

The instruments

Use mine, or make your own. The engine that scaffolds, runs and grades a fleet is ModelMeetsReality — it runs inside an assistant, on your machine, or in a container. Alongside it are starter instruments, each prefixed mmr-: take one as a template, or write one of your own from scratch.

Starter What it says
mmr-tracker Which signals lead and which lag
mmr-tracer How something travels — which move causes which
mmr-timer When a phase arrives
mmr-mirror Models the modeller, one level up
mmr-generator Produces candidates rather than judgements

Each is MIT-licensed and deliberately small — a shape to fill with your own premises, not a finished theory to adopt. Every instrument is its own sister repo: the engine runs the fleet, but each model lives, is versioned, and is graded on its own.

Made a model you think is good? Share your understanding of the world — register it at modelmeetsreality.xyz/publish and follow the guidelines there. It stays a git repo under your own account, graded on a date you set in advance rather than rated by a crowd — and a model that turns out wrong stays listed with its record showing.

Now it's much easier to model your world

The entire harness is open — the registering, sealing, grading, classifying, and postmortem machinery, with none of my data in it. Your head already holds models of your industry, your city, your field. Give them bodies. Find out which ones are real.

  1. 1Copy the repo and make it your main. Clone the template (github.com/shaelsrv/ModelMeetsReality), rename it, git init. This copy is your engine — it never holds a model itself; it runs them.
  2. 2Set it up in your harness. Three ways, in .env: run it inside Claude Code or a compatible assistant on the plan you already pay for (LLM_BACKEND=claude-code, no API key); or a metered key; or fully local against Ollama / LM Studio, where nothing ever leaves your machine. Schedule the weekly grading loop the same way.
  3. 3Spawn models from the main. Each theory becomes its own sibling repo — python -m suites.new_model my-theory scaffolds it and registers it in the fleet. Write its MODEL.md the way you'd finally admit the theory to yourself: premises, falsifiable consequences, a deletion clause you commit to honoring. Repeat for every theory you carry; the fleet grows sideways.
  4. 4Then go meta. A model's subject can be any theory — including your other models. Scaffold a model whose premises are about your own record ("my confidence calibrates worse in X than Y", "my misses cluster on de-escalation") and grade it against your trajectory store — then, if you dare, a model of your revision process, graded against that. Test the thinking of your thinking, as deep as the ledgers below can support. The one stop rule: only add a level when the level below has graded rows for it to explain.

Frontier AI is starting to reason in ways nobody can read, and the industry is still deciding what trust means when you can't see the thinking. I've landed on my answer, and it applies to humans just as much as machines: don't inspect the reasoning — grade the claims.

So here is the actual invitation. Not to follow my ledger — to test your own mind's thinking. Pick one theory you've been carrying for years, the one you're most sure of. Write its premises down. Force it to say something dated and concrete that could be wrong. Register it, wait, and let the grading tell you what your mind would never tell you about itself. The first refutation is unpleasant in a way I can recommend without reservation.

And a suggestion from experience: keep it local and private — at least at first, maybe forever. Nothing here needs a cloud or an audience: the harness runs on your own machine with a local model (Ollama works), your ledgers live in a private folder, and nobody ever needs to see your misses but you. Honesty is easiest to practice where being wrong costs nothing socially — that privacy isn't a lesser version of this discipline, it's the recommended one. I ran mine in stealth for months for exactly this reason. Self-hosting adds friction — real friction, I won't pretend otherwise — but you will thank yourself for doing it. Publishing is a separate decision, for later, or for never.

See the live ledger Get the template