Method
Every claim on this site carries a class, eight component scores and a band. The band is not a judgement typed in by hand — it is derived from the components, so a reader who disagrees can disagree with one line rather than with a verdict.
The numbers on this page are computed from the corpus at build time, not written into the HTML. If a lane is re-scored, this page changes on the next build.
Lane sizes: CASES_EXP 145, ANCIENT 100, SWEEP_B 95, PEOPLE 85, SWEEP_A 82, CRAFT 80, MARS 72, SPECIES 72, SCIENCE 50, SITES 20.
Before anything is scored it is classed, because the three kinds of statement fail in different ways and a single confidence number cannot express that difference.
| Class | Rule | Count | Share |
|---|---|---|---|
| FACT | Checkable against a record that exists independently of the speaker: a document, a published paper, a court filing, a confirmed job title, a dated event. | 578 | 72.2% |
| INFERENCE | A conclusion someone draws from facts — including a sound one. Reasoning quality does not promote it. | 142 | 17.7% |
| SPECULATION | Asserted without a checkable basis, or unfalsifiable as stated, regardless of how sincerely it is held. | 81 | 10.1% |
Each component is filled independently. A total without its components is rejected on ingest. The right-hand column is what the corpus actually scored, which is a different question from what the scale allows.
| Component | Range | Scale | Corpus mean | How much was used |
|---|---|---|---|---|
proximity | +0 … +30 | primary document in hand 30 · firsthand witness 22 · named secondhand 12 · hearsay 5 · anonymous 3 | +18.9 | 63% of the credit available |
corroboration | +0 … +20 | independent sources: 1→0, 2→8, 3→13, 4→17, 5+→20. The same person repeating it in another episode is not corroboration. | +4.7 | 23% of the credit available |
verifiability | +0 … +15 | identifier exists and was retrieved 15 · named but not retrieved 8 · describable but unnamed 3 · unverifiable 0 | +8.8 | 59% of the credit available |
falsifiability | +0 … +10 | specific and testable 10 · testable in principle 6 · vague 2 · unfalsifiable 0 | +7.5 | 75% of the credit available |
track_record | -10 … +10 | position independently confirmed 10 · plausible and undisputed 5 · claimed only 2 · previously found unreliable −10 | +6.1 | 61% of the credit available |
physical_evidence | +0 … +15 | instrumented or measured data 15 · photo/video with provenance 9 · photo/video without 4 · none 0 | +1.3 | 9% of the credit available |
corpus_consistency | -15 … +10 | independently repeated by another guest +10 · consistent +4 · unmentioned 0 · contradicted elsewhere −10 · −15 when an independent source contradicts it | +2.3 | 23% of the credit available |
penalties | -30 … +0 | known hoax link −10 · direct financial incentive −5 · retracted or walked back −10 · chain of custody broken −10 · single low-fidelity source −5, stackable | -1.3 | 4% of the deduction available |
physical_evidence averages 1.3 out of 15 and
corroboration averages 4.7 out of 20. Most of what this corpus contains is a
single person saying something once, with nothing instrumented behind it.
A high score does not mean a claim is true. It means it is defensible in writing to an editor.
| Band | Score | What it asserts | Count | Share of corpus |
|---|---|---|---|---|
| documented | 75–100 | a primary record was retrieved and quoted, not merely named | 93 | 11.6% |
| well-supported | 60–74 | several independent sources, or one document plus firsthand testimony | 164 | 20.5% |
| contested | 40–59 | real support and real contradiction, both on the record | 255 | 31.8% |
| thin | 20–39 | asserted by someone in a position to know, with nothing retrievable behind it | 219 | 27.3% |
| unsupported | 0–19 | no checkable basis, or unfalsifiable as stated | 70 | 8.7% |
The median claim in this corpus is contested. For a body of work about disputed phenomena that is the expected shape; a corpus of the same material reporting a median of documented would be evidence of a broken rubric, not of a stronger case.
Scored against these, not against a feeling. Anchors exist so that two people scoring the same claim a month apart land in the same band.
| Example claim | Expected | Why |
|---|---|---|
| “The 1995 Harvard press release on John Mack exists and says X” | documented | primary document retrieved and quoted verbatim |
| “Nimitz radar operators recorded an object with no visible propulsion” | well-supported | multiple firsthand accounts plus instrumented data |
| “A craft is buried under Skinwalker Ranch” | thin | firsthand assertion, no retrievable instrument record, financial interest present |
| “Nuclear war destroyed a Martian civilisation” | unsupported | published by a credentialed physicist, rejected by the field, and largely falsified |
These three checks run over all 801 claims on every build. They are published whether they pass or not.
| Consistency check | Result |
|---|---|
| Every band follows from its own number | 0 failing |
| Every total equals the sum of its components, held inside 0–100 | 0 failing |
| Every total sits inside the 0–100 scale | 0 failing |
The interface had been computing the band client-side from the number all along, so the site was displaying these claims correctly while the stored field disagreed. Nobody would have caught it by looking. It was only visible by reading the data against the rule it claims to follow — which is the argument for running the check on every build rather than once.
The same instrument is pointed at three different kinds of material on this site. The Mars claim ledger applies it to every claim the show aired about Mars. The news dossiers apply it to live stories, where the band feeds a calibrated probability that is then reconciled against a prediction market. The control room renders the whole graph, computing each band from its number at display time.
Two scoring boards on this site were built to test the same idea against data nobody disputes: Mars Arena ranks cave candidates on complete records, and Moon Arena scores every lunar pit twice — once forgiving the gaps, once charging for them — to show how much of a ranking rests on measurements nobody has taken.
This rubric decides what a claim is worth once you have looked at it. The prior question — which of roughly a hundred and fifty weekly candidates is worth looking at at all — is scored by a separate instrument built on the same principles: the daily triage, whose hard rule is that nothing reaches “open a dossier” on machine scoring alone.
Rubric text: 05_graph/CONTRACT.md, version 1.2. The v1.1 amendment widened the three negative components after the principal's earlier scoring system was recovered and found to price independent falsification well above independent repetition; the change was backwards compatible, so no claim needed rescoring.