Method and firewall
How this was built, and what it is not
How the census was built, the firewall that governs it, and the question it answers.
How it was built
- Source-faithful transcription. Every record was read off the printed page image, one agent per section, and stitched with the boundary entries verified for no gap and no overlap. Odd-looking strings that are clear in the scan (a reversed page range, a title ending mid-word) were kept as printed and flagged, never silently corrected.
- Grading on one axis. Each record was graded only on publication status, against a fixed vocabulary. Published-class means journal, proceedings, book, or monograph. Everything else, thesis, preprint, lecture notes, typescript, awaiting-publication, is counted non-published and named.
- Dual-oracle resolution. The non-published tail was chased twice, by independent search agents and by a separate model with web search, and only the merged verdict was kept. Where the two disagreed, the stricter standard won: a catalogue record is an identification, not a holding.
- Fail loud, default at the edges. Counts are computed from the data, not asserted. Where a resolution rested on inference rather than a document, it is marked as inference. Where an entry remains unknown after genuine search, it is left unknown and says what was searched.
What this is not
To be plain, because it matters more here than anywhere else on this site: nothing in this census is evidence against the Classification. The proof is accepted. The K-group firewall in the second-generation write-up is, on the evidence we have seen, the healthiest reference discipline in any large proof we have examined. What we are measuring is currency, not correctness: whether the record of the mathematics is as current as the mathematics. It usually is not, because keeping it current is manual work that falls to nobody in particular.
The question it answers
This connects to a live question in the philosophy of mathematics. It has been argued that the epistemic standing of the Classification, whether it is genuinely surveyable and understood rather than merely declared complete, is an empirical question about the community rather than something settled by declaration. A census of the proof's own references, and of how faithfully they have been carried forward, is a way to make that question answerable rather than rhetorical. That is what this is.
A record that does not die with its keeper
The reason to build it as a watched, dated, machine-readable thing rather than a one-off report is the same reason behind everything on this site. A record that depends on one person tending it dies when that person stops. This one is watched on a fixed cadence with a public heartbeat, so it keeps breathing whether or not anyone is looking in a given week.
The full census is machine-readable and open. All 507 records, the resolution verdicts, and the transmission specimens are available for anyone who wants to check the working or take it further.
The gap, dated
The specimens are graded against what was actually known to be open at each date. This is that timeline, kept plain.
- 1980–83Completeness of the original proof announced.
- 1980sThe quasithin case rests on Mason's unpublished manuscript of roughly 800 pages.
- 1989It is noticed that some subcases were untreated in that manuscript: a real gap.
- 1992Aschbacher distributes a typescript filling the gap; it is not published.
- 1995Solomon sets out the whole situation, in print, in the Notices.
- 2004Aschbacher and Smith publish the fix, in two volumes. Aschbacher re-declares the theorem, hedged: "(for the moment)".
- 2008A remaining sporadic-case gap is closed by Harada and Solomon.
- 2011Surveys 172 gives the second-generation account of the characteristic-2 half. Its bibliography is one of the two censused here.
- 2018–2025The second-generation write-up continues, volume by volume, past two expired completion estimates.