Skip to content
Method

How we assess benchmark governance

What we measure, what we refuse to measure, and why no per-model reliability finding is published today.

Version v12026-07-29Loupe

Current status: no per-model reliability finding is published. The publishing path is closed in source control, not merely unused.

The short version

We assess how complete our own governance record is for each benchmark. We do not score benchmarks, and we do not publish per-model reliability findings. The machinery for the second thing is built and tested; the switch that would turn it on is a constant in source control, and it is off.

Governance coverage is a statement about our record

For each benchmark we record five governance facts: what contamination controls the maintainers describe, how often the set is updated, its licence, how it is accessed, and whether it is saturated. Coverage says how many of those five we hold.

This is a fact about us, not a grade on the benchmark. Anyone can check it against the register. A complete record does not mean the benchmark is reliable, and we never say that it does; an incomplete record does not mean the benchmark is weak. The reader draws whatever inference they think the record supports. We publish the record.

  • Governance record not assessed. We hold none of the five facts. This is a gap in our work, not a mark against the benchmark.
  • Governance record incomplete. We hold some of them. We list which ones are absent, because the absences are the honest half of the record.
  • Governance record complete. We hold all five. This is a statement about our coverage and nothing more.

There are three bands because a record has exactly three states: empty, partial, complete. There is no number attached to any of them. A coverage percentage would be a certainty dial with the label removed, and research on decision support found that pairing hedged wording with a numeric confidence measurably increases reliance on wrong recommendations.

Incidents are reported beside the record, never folded into it

Where a benchmark has documented incidents, we count them by kind and show them next to the coverage band. We do not subtract them from it. A benchmark with a detailed incident log is usually one whose maintainers disclose problems, and an index that quietly penalised disclosure would reward silence instead.

What a per-model finding would have to clear

None of these is published today. They are recorded here so that the bar is visible before anything meets it, and so that a reader can hold us to it later.

  • A correction for multiple comparisons. Fifty models across twenty benchmarks is a thousand simultaneous tests. At a 5 percent threshold with no correction, roughly fifty models would be flagged by chance alone. We use the Benjamini and Hochberg false discovery rate procedure across the whole grid, and the adjusted value travels with the finding.
  • A reproduction record. The harness, the figure we reproduced, and its confidence interval. A figure we cannot reproduce is not evidence of anything.
  • The effort each figure was measured at. Vendors quote results at maximum effort; users get the API default. A gap between a claim and a reproduction is uninterpretable unless both efforts are stated, so we record them separately and allow them to differ.
  • Fourteen days of notice and a right of reply. The vendor sees the finding before anyone else does, and any response we receive is published in full alongside it.
  • Language bound to the evidence. A measured gap is a measurement and we will state it as one. An assertion about intent is not a measurement, and no amount of statistical signal converts one into the other.

What we deliberately do not use

Membership inference and perplexity based detectors are excluded, and not as a matter of caution. A 2026 replication across several large models found these methods performing at roughly the level of chance, and separate work found that baselines ignoring the model entirely matched or beat the published state of the art. A signal at chance level cannot support a public statement about a named product, so it is absent from our tier vocabulary rather than merely discouraged.

We also do not combine several weak detectors into a single index. Averaging uncertain signals produces a confident looking number without producing any more evidence, and it is the same certainty dial in a different costume.

We use licensed access or permissively licensed public data. We do not rotate credentials or disguise requests to work around a rate limit or a ban, under any circumstances.

Corrections, and how versions work

Each version of this document is fixed once published. We do not edit a published version in place, because an assessment stored against a version is only meaningful if that version still means what it meant. A change to the method or to this text is a new version with a changelog entry, and earlier assessments keep their original version label.

If we get something wrong we correct it in the open: the correction is added, the original stays readable, and nothing is quietly removed.

Changelog

v1 · 2026-07-29

Initial version. Governance coverage bands, the evidence bar for a per-model finding, and the exclusions. No finding is published.