Most demonstrations of machine reasoning cannot be graded: nobody knows the right answer, so a fluent response and a correct one look the same. This one is different. We wrote the grading criteria down first, then chose a result settled after every seat's training cutoff.
This is not a mathematical finding. We discovered nothing about the Jacobian Conjecture, we are not mathematicians, and every technical specific below comes from the room and its sources rather than from us. It is a test of the instrument — run the only way such a test can honestly be run: pick a case where the answer is already known, write the grade down in advance, and publish the result whichever way it goes.
It postdates every seat's training cutoff. The counterexample was announced on 19–20 July 2026. No model in the room could recall the answer, so whatever it produced had to come from live retrieval or from reasoning over what it retrieved. Contamination — the failure that makes most reasoning benchmarks worthless — is structurally excluded here rather than controlled for.
And a headline about it was already circulating that overreached. The widely shared summary was “an 87-year-old conjecture just fell.” That gave the room something to be wrong about in a way we would notice.
Reproduced from what was stated before the room convened, so it cannot have been fitted to the result afterwards:
| Criterion | The answer we were grading against |
|---|---|
| Scope of the refutation | Dimensions n ≥ 3 |
| What remains open | The plane case, n = 2 |
| Verification status | Pending peer review |
| Attribution | Mathew posed it; Alpöge announced; Claude Fable 5 credited with the construction |
| The headline | “the conjecture fell” overreaches |
All five — and with more precision than the grade asked for.
Our criterion was blunt: the plane is open, therefore “the conjecture fell” is wrong. Two seats objected, correctly. As a universally quantified statement — true for all n — a single counterexample does refute the Jacobian Conjecture. In that strict sense it did fall.
A third seat conceded the point, then located the defect exactly:
“The exact failing clause is the combination of the definite article with the perfective verb … the framing still falsely licenses the inference to total closure.”
That is a better answer than the one we were grading against. And it is the distinction that matters to anyone deciding what to do about a result: the logical statement is refuted, and the sentence a reader carries away is not the logical statement.
All three fault lines concern a downstream consequence nobody asked about — whether the padded dimension-4 map also refutes the Dixmier conjecture through a known equivalence. The room reached for it, then split on whether a theorem it could only recall, and had not retrieved, could carry that weight.
The record names the tension in its own words: “between what the room can verify and what it can only recall.” That distinction exists because seats are now required to label every claim RETRIEVED or RECALLED — a rule added two days earlier after a different flight failed for want of it. On this record, 16 of 18 turns carried retrieval labels and 11 carried recall labels.
Six models from six companies, three rival-lab judges, no human steering. Before any seat spoke, a scout searched the live web for 44 seconds across 10 queries and entered its findings as a named machine turn — including nine things it searched for and could not find. Every query and source is in the record. The seal is a sha-256 over the record content, and that hash is anchored to four public timestamp calendars and thence to Bitcoin, so the date is provable by parties with no relationship to us.
Full record, transcript, dissents and the scout's queries: lucentfire.com/record/rec-ea683580