Every mistake we know about, including ours
One row per mistake, not one per paper - a paper can be wrong in eight distinct ways. Some rows are mistakes in published science; 118 are mistakes the Kaimen Rigor engine made. Both are here, counted the same way, because a verification service that publishes only the first is publishing an advertisement.
System of record
Committed, diffable, reviewed in a pull request, frozen at the commit the engine is built from. Readable with no credentials and no network. This is what a benchmark scores against - a denominator that can change under a running benchmark is not a denominator.
Quote the small number beside the total, never instead of it. A corpus reporting 463 without the 69 is reporting the size of a register export.
Capture tier
Every candidate mistake from every run, plus the judgements recorded at review time. Grows without bound, and is never a scoring denominator on its own. Promotion runs capture → record, as a reviewed commit; nothing is ever written back.
The capture tier could not be reached for this request, so its counts are absent rather than zero. The system of record beside it needs no credentials and is complete.
How often the engine is right, where anyone has checked
Measured over findings a reviewer judged one at a time - one verdict per thing the engine said, anchored to the pass, check and locator it came from. A report-level decision is one verdict for a whole report however many findings it holds, and it has nowhere to point; only the finding level can measure precision.
What this corpus cannot yet measure
Not a health score. Each line is a specific claim the corpus is currently unable to support, derived from the corpus itself rather than written by hand - so it shortens when the evidence arrives and not before.
- 1reader. Every independent judgement was made by the same person, so inter-rater agreement is unmeasurable and any bias in it is uncorrected.
- 0disagreements. An empty disagreement count is not agreement - it means almost nothing has been judged twice.
- 5scored problem classes hold zero independent human judgement. These scored problem classes hold zero independent human judgement, so a recall number reported for one of them is a number about silver labels. citation_integrity, unsupported_conclusions, cites_retracted, resource_integrity, data_availability.
- 44groups sit at the same place on the same paper and are not linked. These groups sit at the same place on the same paper and are not linked: some are one mistake judged twice, some are genuinely distinct, and nothing infers which.
Who judged it, and what that is worth
The tier is the whole argument. Anyone can count rows; what decides whether a row may score the engine is who formed the judgement and whether they had seen the engine’s answer first.
human-independent69A named person judged the artifact on its own terms. The only tier that can overturn the engine, and the only one counted as expert judgement.test-verified46Nobody’s opinion: a deterministic test fails on the old code and passes on the new. Stronger than any human tier for an engine defect, and deliberately not counted as expert judgement.human-endorsed84A person accepted or annotated something the engine produced. Evidence it is presentable; not independent confirmation that it is true.agent-adjudicated7An AI agent working on the engine reached this verdict. Scoring the engine against it is scoring the engine against itself.notice-grounded257A model read the retraction notice and quoted it verbatim. The problem is real; whether it is visible in the paper is a different question.external-fact0A public register decides it. Free, plentiful, and says nothing about what is wrong.The engine’s own record
118 of the rows above are mistakes the engine made, not mistakes it found. A fix is only recorded as fixed when it carries a receipt naming a test that still runs - a claim that something was fixed, with nothing re-checking it, is a memory.