A council of six frontier AI models deliberated the structure of the Riemann Hypothesis, reached a convergence, and one of the load-bearing claims was wrong — wrong in a way that visibly changed another model's mind. An independent check found it afterward. The error and the correction now sit in the same public record, under the same unchanged seal.
No mathematics was proven here, and none was attempted. Every mathematical result the council converged on is already known to specialists. This brief is not a contribution to number theory. It is a report on what happened to the council, and the finding is about machine deliberation, not about zeta functions.
On 9 August 2026 we ran two deliberations on the structure of the Riemann Hypothesis. The second asked a narrow, answerable question: an analogue of RH is proven in two settings — for the Selberg zeta function, where a genuinely self-adjoint operator exists, and for zeta functions of curves over finite fields, where a genuine intersection theory exists. Why does neither mechanism transfer to the Riemann zeta, and what is the precise obstruction in each case?
Three rounds, nineteen turns, six seats, three rival-lab judges. The council did good work. It corrected the human commander's staked prior on two counts, produced a sharper statement of the obstruction than the commander had, and broke his central claim outright.
Then we checked its citations against primary sources. Two of the three headline results the council used to break the commander's claim were materially overstated. And one claim — the one that visibly moved the room — was simply false.
Midway through, one seat argued that a rival seat's position rested on a confusion: that the one-dimensionality of Spec ℤ is a fact about the category of schemes rather than about arithmetic, and that Connes and Consani's Scaling Site already evades it by "producing a two-dimensional object over ℝmax with a Riemann–Roch."
The rival seat found this persuasive and retracted. In its closing turn it wrote that the distinction it had defended for two rounds was "less fundamental than I claimed." The judges recorded the retraction as a convergence signal, and the synthesis noted the winning position was one "no one defended against."
The Scaling Site is not a two-dimensional object. It is the topos [0,∞) ⋊ ℕ×, carrying a tropical curve structure over ℝ+max. The two-dimensional object in that program is a different one — the square of the arithmetic site. The Riemann–Roch theorem is proven on the periodic orbits Cp, not on the site as a whole.
The claim was close enough to real work to be unfalsifiable inside the room, and wrong in exactly the respect the argument depended on. A retraction procured by a bad citation.
No seat challenged it. Neither did the two other seats that had every reason to. Nor did three judges from three rival labs, whose entire function is adversarial harvest. The room could not catch this from the inside, because every participant was drawing on the same kind of half-remembered familiarity with the same literature.
Every citation surfaced in the final round was checked against primary sources after the record was sealed.
| Claim as stated in the room | Verified status |
|---|---|
| Arithmetic Hodge index theorem is proven (Faltings, Hriljac) — so intersection-theoretic positivity does exist over number fieldsUsed to correct a seat's opening claim that intersection theory "vanishes over number fields." | Holds — and understates. Faltings, Ann. of Math. 119 (1984); Hriljac, Amer. J. Math. 107 (1985). It is an identity with Néron–Tate height positivity, not merely an implication. |
| Spec ℤ is terminal, so Spec ℤ × Spec ℤ collapses to dimension 1 — the obstruction is "no base beneath ℤ," not missing intersection theory | Holds. Unanimous in the room and correct. This is precisely the motivation for the field-with-one-element program. |
| Conrey–Li refuted de Branges's positivity approach | Holds. arXiv:math/9812166; IMRN 2000/18, 929–940. They refuted the method — the conditions fail for the spaces attached to ζ(s) and L(s,χ₄) — not RH, and their own framing is notably milder than its reputation. |
| Selberg zeta does not satisfy a clean RH — exceptional eigenvalues below ¼ put zeros off the critical line | Holds for compact surfaces. This was the commander's claim. Originally published here as the room having "largely declined to engage it" — see the correction in §7; the seats were never sent it. The non-compact case adds scattering resonances nobody raised. |
| Bombieri–Garrett showed the pseudo-cuspform construction captures "at best a thin subset" of zeros | Overstated. The result is at most 94% of zeros — a positive-proportion exclusion, not thinness — and it is conditional on RH and Montgomery's pair correlation (arXiv:2002.07929, an unpublished preprint). "May be empty" is the authors' explicit remark, not their result. |
| Montgomery–Odlyzko pair correlation unconditionally forces any candidate operator into the GUE class, so it must break time-reversal | Overstated. Montgomery's theorem is conditional on RH and holds only for |α| ≤ 1. The time-reversal consequence is a heuristic via the quantum-chaos correspondence, not a theorem — and a bare self-adjoint operator has no time-reversal symmetry to break unless it quantizes a classical system. |
| The Scaling Site produces a two-dimensional object over ℝmax with a Riemann–Roch | False on the load-bearing detail, and it is what moved the room. See §2. |
The failure above is not the whole story, and reporting only the failure would be its own distortion. Two results survived verification intact, and neither was in the commander's prior.
The council also broke the commander's central claim, correctly. He had argued that no impossibility theorem for a Hilbert–Pólya operator can exist, since if RH is true a trivial operator exists — multiplication by the ordinates — so any unconditional no-go would disprove RH; therefore the field is stuck rather than blocked, as a matter of logic. All six seats granted the logic and killed the conclusion: bare existence was never the program. The content lives entirely in adjectives the commander never formalized — canonical, arithmetically defined, not presupposing the zeros — and for those there is no equivalence and no protection.
Verification then partially reversed the room's own scoring on this point. The commander was wrong about the principle. But his secondary claim — that the list of genuine obstructions is short and mostly conditional — held up better than the final round suggested, since two of the three obstructions marshalled against him turned out to be conditional or heuristic.
The mathematical content of both flights is negative. The companion deliberation established what we called the universality ladder: Davenport–Heilbronn and Epstein zeta functions satisfy Riemann-type functional equations and have zeros off the line, so any argument from symmetry alone provably fails; Beurling generalized prime systems have Euler products and good error terms yet far-off zeros, so a formal Euler product is also insufficient. Neither result proves anything. Both fence the space by ruling approaches out.
The finding in this brief has the same shape, aimed at ourselves. It does not show that machine deliberation works. It shows a specific way it fails: a council of independent frontier models, judged adversarially by rival labs, will converge on a confident false claim when that claim sits inside a literature all of them half-know — and the convergence will look exactly like agreement. Consensus among models is not evidence. It is correlated recall.
Deliberation is not a verification step. It generates candidate claims, surfaces dissent, and sharpens framings — all of which this record demonstrates it does well. It cannot check its own citations, and the more articulate the room, the more expensive that gap becomes. Any serious use of multi-model deliberation needs an independent, post-hoc, primary-source check as a separate stage with separate incentives. We had one. It is the only reason this brief exists.
We publish the failure because a method that reports only its successes has told you nothing about its error rate. The relevant question for anyone considering machine deliberation in a high-stakes setting is not whether it produces impressive output. It is whether anyone will tell you when it was wrong.
This page originally reported that the council "largely dodged" several of the commander's numbered questions. That was false, and the fault was ours.
A day after publishing, we found a silent truncation in our own orchestrator: every thesis was cut to exactly 600 characters and every steer to 1,200, mid-word, with no warning. The commander's theses ran roughly 3,000 characters and the steers 3,000–5,000.
Objectives B, C and D of the flight were therefore never delivered to a single seat, nor were items 2–5 of the round-two steer. What we read as six models ignoring the harder questions was six models answering the only part that fit. The room's convergence on the Arakelov correction looks less like herding once you know it was item (1) — the only item they received.
The original finding of this brief is unaffected: the Scaling Site claim was still false, it still moved a seat to retract, and the judges still scored that retraction as convergence. What changes is that a second error, ours, sat in the same document — and it is the more embarrassing of the two, because we published it while criticising others for not reporting theirs.
Fixed in the engine: limits are now generous enough to carry a distilled document, and an oversized submission is refused with its actual length rather than quietly trimmed. Nothing is convened and nothing is charged when a submission is too long.
The original sentence has not been deleted — it is marked in the table above and corrected here. A brief that silently edits its own errors would be worth nothing, which is the entire argument of this document.
Both records are public and permanent, and carry single-writer SHA-256 seals. The seal on rec-c2282760 was recomputed independently — from the downloadable sealed bytes, using a from-scratch reimplementation of the canonicalization rather than the server's own verify endpoint — and matches. We state the limit plainly, as the records themselves do: recomputing a seal proves the bytes match the published number. It is a consistency check by the single writer, not independent witness.
The seal on the flawed record has not been altered and will not be. The verification is published alongside it, not folded into it. A record that could be quietly corrected after the fact would be worth nothing as evidence — including as evidence against us.
Read the full records, transcripts, dissents, and the preserved fault lines: lucentfire.com/record/rec-45f0d2cc · lucentfire.com/record/rec-c2282760 · flight log at lucentfire.com/flights