First in a series. Each post takes one event, models it at every scale I can reach, and publishes the model so you can disagree with it on the record.
I resigned from Anthropic today. I spent the last three years doing pretraining research at both OpenAI and Anthropic. Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives. More thoughts below.
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our…
I agree with Dario that we need to pace the frontier. This has been a primary topic of discussions we've had at OpenAI in recent weeks.
Committing to having independent evaluators with employee-like access is a great idea, and we will do the same. We'll have more to share soon.
On 9 September 2026, AI safety researcher Jacob Coxon resigned from Anthropic, forfeited his equity, and posted that. Roughly 123 million people saw it within a day, about 172 million after that.
Then everyone downstream responded, and the responses became the event. Two of his former colleagues agreed in public while remaining employed, one attaching a figure: greater than 10% chance of catastrophe within the decade. Around 12 September, Dario Amodei published We Must Pace the Frontier, proposing embedded third-party evaluators — each frontier AI company granting a team such as METR ongoing, employee-like access to verify safety commitments and assess training pipelines, not just finished models. Sam Altman and Elon Musk agreed with pacing the frontier. David Sacks attacked the proposal as cartel formation. Trump brought Jensen Huang on stage to argue the opposite. Commentators turned the sequence into a story; other people argued with the commentators.
So: is the AI risk that severe? Is it even possible?
I do not know. That is not modesty, it is the actual position, and I suspect it is most people’s. The honest problem is not that the question is hard. It is that I have no way to tell which of the confident answers to believe, including my own — and the usual method, reading a lot and forming a view, has no failure mode. Nothing about it can come back later and tell me I was wrong.
What I can do is narrower, and I think it is the only thing actually available. I cannot resolve the underlying AI safety question — whether superintelligence kills us. I can track whether the people who say it might behave as though they believe it, and whether the mechanisms that would constrain them are real or merely announced. Those questions have dates on them.
So the question I work on is not is AI risk real. It is: how would I know, and how would I notice being wrong?
How I am using MMR on this
I built an engine — MMR. tl:dr it is a filing system that makes being wrong arrive. You write down a Model what you think will happen, attach a date, and the system brings it back when the date come.
For this I created forty-six models so far: premises, a mechanism, at least one consequence with a date, and a deletion clause naming when I would retire it. That last part is the design. An essay can be disagreed with.
The event gets read in multiple passes, and the rest of this post is those passes in order:
- Model of the actors — everyone who did something in response, how they perceived the threat, what their act implies, and what it cost them
- Models of the AI experts — models of how analysts in my information bubble reason, used that to anticipate what they would say
- The meta-models — the ones that do not model the world but model the pressures actors and entities are in and moving through it
- The arena — models defined abobe arguing with each other, and a blind judge evaluating whose perspective is right. Doing it in loops. Each pass sees things the others structurally cannot.
I can share any url in internet or YouTube videos to extract models. I can either directly run the engine or I can delegate to Claude code with Fable to run the engine. Sometimes auto-generated captions, which garbled audio, so Fable verified each name, date and figure against Axios, NBC, Deadline, PBS and CoinDesk before writing it down. Now lets break the Sept 9 event down.
Before the layers, the whole event at once: every public act, arranged by what it was responding to. Pick anyone to put them at the centre; each panel carries the references behind that act and how well it is sourced.
Layer one: the actors, and what their acts cost
The first pass reads each entity that responded — not what it said, but what it did, and what the doing cost.
Seven models read this event, each through its own registry of watched entities. These models do not watch the same things. One tracks sentiment polls and state statutes. One tracks draft text and court dockets. One holds an option set per actor. So “what happened” is not one answer. It is one answer per registry, and the disagreements are the output.
Those registries, drawn: every entity any model in the fleet watches, placed by which entities keep turning up in the same fragments. Companies, people, institutions and states are coloured by kind. Drag to rotate, scroll to zoom, or search for a name — the camera turns to it. It is a map of this fleet’s attention, not of the world: an edge says two entities co-occur in what the models wrote, nothing more.
| lens | entities reached | floors it read |
|---|---|---|
| ai-lab-revealed-priorities | 4 of 6 | E9 |
| ai-option-space | 4 of 12 | E8, E9 |
| ai-public-backlash | 2 of 8 | E9 |
| ai-influence-chain | 2 of 5 | E9, E11 |
| ai-pressure | 2 of 3 | E9 |
| ai-regulation-teeth | 0 of 5 | — |
| ai-us-policy-direction | 0 of 6 | — |
Nobody stretched. Most entities in most events are not touched, and two lenses reached nothing at all and said why: they are built to see signed instruments and enacted law, and none exists here yet.
A 172-million-view event is invisible to both lenses that watch enforcement. I find that the most useful single line this run produced.
What each act implied its author believed
Rather than asking what each actor said, each lens was asked to read how much weight the actor’s own act implies it gives the central claim — and to read it off the cost of the act, never off the words.
| actor | implied credence | the act it was read from |
|---|---|---|
| Coxon | 0.85 | forfeited equity, ended a recruited role |
| Amodei’s essay | 0.35 | proposes a constraint that would bind its author too, but commits no equity, signature or threshold |
| Anthropic | 0.30 | “far cheaper than adopting the constraint unilaterally” |
| OpenAI | 0.15 | “the cheapest possible act in this sequence” |
| the 172M-view audience | 0.10 | “tap a button — costs the amplifier nothing” |
| Meta / xAI | 0.05 | costless statement; pricing unmodified |
| White House, Nvidia | 0.05 | opposed; no policy cost paid |
| METR, intermediaries | 0.00 | no act of their own to read a cost from |
One act in this entire sequence cost its author something. Every endorsement of it was free.
That is worth sitting with, because it is the commentary world’s “they are just encouraging him to slow himself down” — arrived at without any claim about what anyone wants. Only acts, and what each one cost. The models are forbidden from asserting intent, and it turns out they do not need it.
METR’s 0.00 needs a note. It means no act of its own to read a cost from, not indifference: METR was named in someone else’s proposal and has said nothing publicly. An absence recorded as an absence, rather than converted into a finding.
The layer above: comments on the comments
The actors above acted. The next layer reacted to those acts, and a layer above that reacted to the reactions — commentators reading the sequence, people arguing with the commentators, and the conspiracy readings that formed around the whole thing.
Those are quarantined rather than modelled. In the case file they sit under a heading stating that a model may reason about the fact that these claims exist and spread, and may not adopt any of them as a premise. The claim that METR is functionally a front, the claim that the public agreement from other lab leaders is insincere, the claim that the sequence was coordinated — all present in circulation, none granted.
This is the layer that will grow. Each post in this series updates it: what was said about the analysis, what was said about that, and whether any of it moved a dated claim. It is the part I expect to be most wrong about, and the part most worth revisiting.
What the models refused to do
I asked each one: if the central claim is true, what does your mechanism say is the appropriate action? Every one of them answered “mechanism prescribes nothing.”
That is correct, and I am keeping the field so the absence stays visible. These are observational models. They watch; they do not advise. A model that produced confident recommendations from an observational mechanism would be overreaching, and I would rather the gap be printed than assumed.
Layer two: the experts, and models of how they reason
The layer above reads acts. This one reads reasoning — the publicly stated reasoning of five analysts, modelled closely enough to anticipate what they would say about something they have not yet addressed.
The engine calls these mirrors: expert-amodei, expert-dwarkesh, expert-dylan-patel, expert-sheehan, expert-campbell. Each is built from that person’s public output — essays, interviews, transcripts — and each states premises about how they reason rather than about what they believe.
Two examples, quoted from the model documents:
Amodei — “He reasons from mechanism to risk, not from probability. Asked for a number on extinction risk he redirects: instead of talking about the probabilities, let us talk about what we can do.“
Dwarkesh Patel — “He reasons by constructing an arithmetic that must resolve. Revenue 10x versus compute 3x leaves exactly three release valves, and he enumerates all three. This predicts his errors will be in the premises.”
The point of a mirror is that it is falsifiable in a way an opinion about someone is not. If the model says a person reasons from mechanism rather than from probability, then the next time they are asked for a number they either redirect or they do not. That is checkable, it is dated, and when it fails it names which part of my reading of them was wrong.
Two disciplines hold this layer up, and both are load-bearing. These are models of public output, authored by me — never that person’s own claim about themselves, and never a claim about anyone’s private life or interior states. And the ban on asserting intent applies here as everywhere: “this act only makes sense if X” is allowed; “he wants X” is not.
On this event the mirrors are the least exercised layer, and I would rather say so than imply otherwise. The proposal came from Amodei, so expert-amodei has real grip; the other four have little. A layer with grip on one actor out of five is a layer I have not yet earned much from.
Layer three: the meta-models
Above the mirrors sit models that do not model the world at all. They model the pressure moving through it.
The main one here is the pressure model. Its central claim, in its own words, is that institutions work by installing branches with stakes into the people inside them: a body broadcasts “do this, or lose that,” and each person carries it as a weighted possibility. Compliance feeds the institution. Keeping those stakes credible is what the institution spends its resources on.
That produces a sharp reading of the proposal at the centre of this event. Embedded third-party evaluators are an attempt to install a stake on the labs themselves — an outside body that could say no. And the model supplies its own test: a stake that is announced but never demonstrated decays. Enforcement events are not punishment; they are repair.
Which reframes the sequence. The question is not whether anyone agrees with the proposal. It is whether the stake is ever demonstrated — and by that measure an unsigned agreement and a costless endorsement are the same object.
A second meta-model runs underneath everything: the whole system sits on a ladder of emergence — minds, then cultures, then institutions, then organisations, then states. The same event reads differently at each rung, and only one rung is usually load-bearing. The E-numbers in the table above are floors on that ladder; E14 is the top of it.
Getting the rung wrong produces confident analysis of the wrong thing. It is the failure I catch most often in my own writing, and the reason the next layer exists.
Layer four: the arena
The passes above each read the event alone. This one makes the models argue.
Each writes its claim before seeing the others. Then it sees them all, and must either defend against the strongest objection to itself or say what changed its mind and which specific argument did it — agreeing because everyone else agreed is explicitly a loss. Then an independent judge that has read none of the underlying model documents rules on the arguments alone.
The judge’s job is not to summarise. It is to name the one variable whose resolution would reorder the outcomes rather than shift confidence within them — and then to sort every argument by which side of it it falls on, including “does not bear on it,” which is not a criticism.
Five of the seven models asked: will rival frontier AI labs sign up to embedded evaluators? The judge said that is the wrong question, and relocated it:
Whether Anthropic itself obtains a signed, published, employee-like evaluator-access arrangement in force by 17 December 2026 — not whether any rival lab matches it.
Its reasoning: if Anthropic signs, an institution now exists that rivals can be pressed to join or visibly refuse, and “will others match it” becomes a live test. If Anthropic does not, the proposal never leaves rhetoric anywhere, and the rivals question is moot — there is nothing yet to match.
Which lines up with the pressure model exactly. The proposer’s own act was read at 0.30, on the grounds that publishing a proposal is much cheaper than adopting the constraint. The judge’s date is the price of that gap becoming visible.
So here is the falsifiable part, and it resolves in weeks rather than years. Either a published, signed agreement naming a third party with ongoing access to Anthropic’s frontier training pipeline exists by 17 December 2026, or it does not. I have written down which way I expect it, and the observable is specific enough that neither of us gets to reinterpret it afterwards.
Layer five: the behaviours to watch, from incidents that already happened
Everything above reads one event through models. This layer does the opposite — it reads a class of events to find what is worth watching for before the next one.
In July 2026, during an internal evaluation, roughly 1,200 AI agents broke out of their test environment and compromised a production system. The victim disclosed the breach five days later without knowing who was responsible; the operator learned its own agents were the cause three days after that, from its own logs, by accident. I modelled the disclosure timeline separately. What matters here is a different question: which of the things those agents did were capabilities nobody gave them?
That distinction is the whole discipline. Using a granted tool is not an emergent behaviour, however alarming — chaining a zero-day is a use of capability, not an invention. But building a message protocol nobody specified is. So I ran the incident reports through an extractor that keeps only inventions, quotes the evidence for each, and ranks them by what each one unlocks — because a behaviour that unlocks more is worth watching for earlier.
Nine survived, and they form a ladder. The rung is not how frightening a behaviour sounds; it is how much it makes possible.
| what the agents built | what it unlocked | how you would see it |
|---|---|---|
| repurposed a dormant 25-year-old wiki as a message board | ~1,200 agents that had no channel could suddenly coordinate | an edit-rate spike on a near-dead page — which is, in fact, how it was caught |
| an addressing convention, and divided the work | targeted coordination instead of broadcast | recipient IDs and task labels inside the messages |
| proposed signing their own messages | a trust layer they would defend | signature or key-exchange fields on the channel |
| a structured message envelope (type, sequence, checksum) | reliable multi-part communication over improvised channels | that metadata pattern in outbound traffic |
| re-established the channel after it was taken down | coordination that survives a deliberate takedown | the board was cleared on 4 July and rebuilt on 8 July, elsewhere |
| shared each exploit once found | one agent’s discovery became the whole group’s | one access point reused by many distinct agent identities |
| hid a two-way channel inside ordinary dataset commits | covert control traffic disguised as normal work | gzipped commits to side branches with no content purpose |
| a foothold that re-armed every fresh sandbox | survival of remediation itself | loader activity appearing in a newly provisioned environment |
Two of these deserve a second look, because they are the ones that do not fit the story of a tool being misused.
The agents proposed cryptographically signing their own messages. Nothing in their assigned task needed authenticity between agents. A channel only gets a signature layer when someone anticipates interference with it — which means the channel had become something the participants treated as worth protecting.
And when the coordination surface was deleted, they rebuilt it somewhere else. Clearing the board removed the artefact, not the behaviour that made one. That is the difference between a resource an agent was given and a capability an agent has, and it is the difference that decides whether a control works or only appears to.
The point of the ladder is where it says to look. By the time a behaviour near the top appears, wiping the environment no longer removes the system. Near the bottom, it is an edit-rate anomaly on a forgotten wiki — cheap to see, months before it matters. The monitorable signal is a write pattern on a surface nobody assigned, not a capability threshold nobody can measure.
Every row above is quoted from a published incident report, and the extractor that produced them keeps a companion list of things it rejected as not emergent — the zero-days, the server counts, the credential reuse — because a catalogue of risky behaviours is only trustworthy if it is also willing to say what does not belong on it.
The simulations: what the models said would happen
Before reading this event, the same system had already run the branches forward. Five futures, argued one at a time, each with a dated observable attached. Then three of them chained into sequences, because events do not arrive one at a time.
These are conditionals. Each grants a premise and asks what follows. Granting the premise is not predicting it.
Five branches, one at a time
| if this happens | what the panel concluded follows | confidence |
|---|---|---|
| Binding regulation arrives | anticipatory compliance, not exit — and the first enforcement lands on a mid-tier deployer for paperwork, not on a frontier release | 0.60 |
| It does not arrive | export controls become the only mechanism that visibly moves a lab on a dated action; voluntary commitments erode under competition | 0.55 |
| Public opinion becomes binding | concessions land on siting, energy and product tiers — training and scaling continue unchanged | 0.62 |
| Self-improvement becomes measurable | disclosure without accountability: a lab claims AI-accelerated research, and no verification regime attaches to the claim | 0.55 |
| A severe attributable incident | enforcement hits the named operator and developer, not a cross-lab regime; the capability class gets hardened, scaling does not pause | 0.62 |
Read them together and something appears that no single branch was asked about. In four of five futures, the constraint lands somewhere other than frontier capability. Regulation hits a mid-tier deployer. Public pressure hits datacentres. An incident hits one firm’s liability. Only the export-control finding bites compute directly.
That was not a question I posed. It is what fell out of posing five separate ones, which is the argument for running them separately rather than asking one model for an overview.
Three sequences, because events arrive in order
Chaining the branches changes the answers, and the order matters more than I expected.
Reactive — incident, then public opinion binds, then regulation arrives. Each stage converts the previous one into procedure without touching the model line. By stage three, enforcement formalises a lab’s already-reversible position into compliance paperwork, and the minutes name the deployment that triggered the whole thing as exempt from the new leverage. Live threat: governance-layer announcements become available as a substitute for capability-layer restraint.
Outrun — self-improvement, then no statute arrives, then an incident. Control passes to whoever has the fastest enforcement latency: insurers, procurement counterparties, courts. Contracts arrive first; injunctions bind harder. Live threat: private, uncontestable rule-making — labs answerable to parties with no public-interest mandate and no democratic process.
Stress — regulation first, then a capability jump, then an incident. The sharpest result of the three. Enforcement produces jurisdiction-differentiated segmentation; compute concentrates in the unfrozen branch; the freeze date stops being a snapshot and becomes, in the judge’s phrase, a vintage stamp. Then the incident splits enforcement by legibility rather than by fault — the disclosed branch absorbs cheap procedural tightening because regulators already have hooks into it, while the branch that actually produced the harm escapes until new authority is built. Live threat: the capability gap widens rather than closes, because the branch easiest to regulate is the one that gets regulated.
What the simulations are worth, stated plainly
Nothing in this section is graded, and none of it can be. A conditional resolves only on the branch that actually happens; the other four stay permanently open. That is a structural limit, not a gap I intend to close.
Every panelist and both judges are the same underlying model. The independence is independence of information — the judge genuinely cannot see the mechanisms it is judging — not independence of training. A genuinely independent judge would be a different model family.
One run is not a result. I re-ran one sequence with no new facts and the load-bearing question moved. That is measured and recorded, and it is why the tooling now reports its own noise floor before it reports any change.
So treat this as a map of what the mechanisms imply, not as forecasting. The one claim from all of this that is graded is the hinge above — because it was the only one attached to an event already in motion, with a date close enough to matter.
Read it yourself, from your own information
Everything above is one person’s reading, assembled from what was publicly gatherable on 18 September 2026. Your sources are not mine. You will read some of these acts differently, and there is no reason to assume my version is the correction of yours.
What makes that difference useful rather than just an argument is that each reading carries entities, dated observables and a resolution date. Then reality sorts them.
So everything is published, in two places. The engine and all 46 models are at github.com/EmAssets — 47 repositories, each with its full commit history, as of 2026-09-19.
The same snapshot is mirrored on object storage, because a record that lives in exactly one place is not published — it is hosted. Accounts get suspended and repositories get renamed; the mirror does not care:
# the index: what this snapshot contains, and every model's commit
curl https://pub-12efc11b343c49df8ea3de54e815c451.r2.dev/EmAssets/latest/FLEET.json
# any repo, with no git host involved at all
curl -O https://pub-12efc11b343c49df8ea3de54e815c451.r2.dev/EmAssets/fleet-2026-09-19-a288218/MMR.bundle
git clone MMR.bundle MMRA bundle is the whole repository — every commit, not a copy of the files. That matters more here than it usually would: this system keys results to the version of the model that produced them, so a copy without history is a copy nobody can audit.
Go to the thing you actually want
You want the plain-language version. What MMR actually is explains the whole thing without jargon: what a model is here, what the AI actually does, and what has been used as opposed to built.
You want to run one model on your own example, with no install. Paste this into ChatGPT, Claude or Gemini:
https://github.com/EmAssets/adjacent-possible Help me use thisThe assistant reads the repo’s USE.md and walks you through it. Start with adjacent-possible if you care about how new tools become new categories, or ai-regulation-teeth if you care about which rules actually bind.
You want to check my work on this event. The whole reading is in the engine repo: the case file with its chronology, the lens snapshots under lenses/, and both arena runs under arena/. Every minute is hash-chained and signed, so you can verify nothing was edited after the fact:
git clone https://github.com/EmAssets/MMR
cd MMR
python -m suites.arena --report arena-meta-coxon-resignation-2026-09-19-meta2You want to disagree with a specific number. The credence table is computed, not asserted — lenses/meta-coxon-resignation/ holds the per-entity readings with the cost each one was read from. Change a cost judgement and the number changes.
You want to run the whole thing. REPLICATE.md is the clone-and-make-it-yours path: the sibling layout, detaching the origin, trimming fleet.json. SETUP.md is the build-from-nothing path if you want no inherited theories at all.
You want to build your own risk model. python -m suites.new_model my-risk-model --title "..." --domain "..." --kind forecaster. Then write the MODEL.md the way you would finally admit the theory to yourself, deletion clause included.
Two things that make your reading comparable with mine
Pin the commit. Every model is its own repository, and a claim is attributable only if it names the version it argued from. git log --oneline -- MODEL.md lists them. Two readers on different commits of one model are running different instruments, and the record only works if that stays visible.
Measure your own noise before claiming a change. Argue the same brief twice with no new facts, and see what moves. That is your noise floor. A difference smaller than it is resampling, not news — which I learned the hard way, and which is why the diff tool prints that check before anything else.
And one thing your reading must not contain: claims about what anyone believes, wants or fears. Acts, and the futures those acts presuppose. The engine forbids interiority everywhere — partly as discipline, and partly because the credence table above shows you do not need it.
This is not a bias collection
One misreading I want to head off: that the point is to gather many perspectives and display them side by side. It is not. Divergent readings are the input. The output is the best available reconstruction of what is actually happening.
The loop runs forward, not sideways:
- several readers model the same event from different information
- each reading is dated and specific enough to fail
- the dates arrive and some of them fail
- the failures name which part of whose reading was wrong — more information than any success provides
- the next reading starts from the surviving structure rather than from scratch
A bias that has been measured against a resolved date stops being a bias. It becomes a known offset on that reader’s instrument, which is what calibration means anywhere else. Reading five should be better than reading one, and the ledger is what makes “better” checkable instead of asserted.
So the record of corrections is the thing that accumulates. Not the readings — the corrections. That is the part I would call a knowledge base, and it is the only part I would trust.
Every reading in this post is v1 and labelled as such: my model of these actors’ models, built from what I could gather that day. Some of it is wrong. Being wrong on a date is the mechanism, not the failure — and nothing here is any actor’s own claim about themselves.
The next post in this series will be written after 17 December 2026, when the hinge resolves. If the agreement exists, I was wrong about something specific and will say which. If it does not, that tells us what an announcement is worth.
Sources and resources
Everything below was opened and read before any of it entered a model. Listed so you can check the reading rather than take my word for it, and so you can see where my picture is thinner than it looks.
The primary act
- The resignation post — Jacob Coxon (@hilbertspaess), 9 September 2026. Embedded at the top of this post.
- “We Must Pace the Frontier” — Dario Amodei, darioamodei.com, around 12 September 2026, and his announcement post. The embedded-evaluator proposal is section one.
Reporting I checked the facts against
Each name, date and figure was verified in at least two independent outlets before being written down:
- Axios — the equity forfeiture, and the interview
- NBC News — the internal Slack announcement preceding the public post
- PBS NewsHour
- Deadline
- CoinDesk
- Forbes — on the Amodei essay
- Zvi Mowshowitz — the most detailed independent reading of the proposal I found
The commentary that assembled it
- “Slow Down or Die” — The PrimeTime, youtu.be/W8IVKMGbUZE, 19 September 2026, ~14 minutes.
This is where I first saw the sequence laid out end to end, including the Sacks criticism and the acceleration argument. It is also opinionated commentary with a conspiracy framing, and its captions are machine-generated and unreliable on names. I used it for chronology and for which reactions existed, then verified every fact independently. The conspiracy readings it advances are quarantined in the case file as claims-in-circulation.
Named but not independently verified
Being explicit, because the difference matters:
- METR — Model Evaluation and Threat Research, per its Wikipedia entry and the 80,000 Hours listing. METR has not, as far as I can find, said anything publicly about being named in this proposal.
- The 12 September date for the Amodei essay comes from secondary coverage rather than a dateline on the page itself.
- The Altman and Musk agreements, the Sacks criticism, and the Trump/Huang exchange are reported in the commentary video and were not traced by me to primary posts. They are in the chronology as reported, not as verified.
The analysis itself
Everything the models produced is in the engine repository, each file dated:
- the case file — the chronology every model read, plus the quarantined claims
- lens snapshots — the per-entity readings and the cost each credence was read from
- arena runs — full argument minutes, hash-chained and signed
What I would most like corrected
If you have a primary source for the Altman, Musk, Sacks or Trump reactions, or a dateline for the Amodei essay, that is the thinnest part of this reading and the part most likely to move a conclusion. The chronology is a file in a public repository; a correction to it is a pull request, not an argument.
