When the six models debate a specific, verifiable task with a hidden correct answer that none of them state correctly in their first independent round, does the structured debate process (without external tools) converge on the correct answer by the final round more often than chance or than any individual model's initial guess?
6 models deliberated; a human Commander held the room but did not steer. Sealed 2026-08-12T22:41:15.009Z. Engine lucentfire-roundtable/v1 (live).
The question put to the room
When the six models debate a specific, verifiable task with a hidden correct answer that none of them state correctly in their first independent round, does the structured debate process (without external tools) converge on the correct answer by the final round more often than chance or than any individual model's initial guess?
What survived
- The thesis's stated baseline is degenerate: conditioning item selection on 'all six wrong in round 1' fixes best-single-model-first-guess at 0%, so any convergence wins trivially; the honest comparators are chance over the answer space and a token-budget-matched single model with self-consistency.
- Debate's error-correction mechanism without tools is re-derivation, not memory, so it can only help on items whose answer is re-derivable inside the transcript; on recall-bound items (the parachute riddle) cross-examination cannot generate the answer — conceded by its own proposer.
- The only genuinely new correct answer produced in this run (241) arrived via two seats independently and simultaneously computing it in round 2, i.e. by parallel sampling rather than by any exchange or refutation.
Seal (sha-256, single-writer): 2139f24695a37daf37b3c63d77005fa09b4b71a64f39762612d266111ff6cafe