Evidence
Everything we can show, each item labelled and sourced. Open any of it; you do not need to ask.
What you can open now
- Public data
Analyser change explorer
34 method comparisons at 14 stand-in sites, from 12 open-access studies, 10,579 paired results. Pick any one to see the comparison line, the bias at its decision limits and how many results crossed a limit against repeat testing alone. Of the 34 method comparisons, 21 get a full re-test band and 5 an indicative one, because they have fewer pairs than the minimum we set; the other 8 are shown as a comparison only or are too small for a band. In 6 of them the change moved fewer results across a limit than repeat testing alone would, and the explorer says so.
- Public data
Public data we re-analyse
The 12 open-access studies behind the explorer, each with its authors, a link to the source, its licence and what we changed.
- Sample
Sample analyser-change report, CD4
A sample report built from one public platform replacement of 1,885 paired CD4 counts. Reported without adjustment, 535 results change CD4 category against about 238 expected from repeat testing alone; the report shows how the comparison line and a re-test band manage the switch. Nobody commissioned it.
- Live demo
MedEval-1
A working version of a versioned evaluation standard for medical AI, built by NakedSignal. Its held-out track mints, seals, scores, rotates and audits a test set on a public stand-in corpus. 48 of 48 published scores were re-derived by our own audit from de-identified bundles and manifests. The same instrument scores 9 domains, inside and outside medicine, without renaming a dimension.
- Designed
Analyser-Change Registry
A pooled record of what analyser changes do to patient results, contributed by laboratories by choice. Designed, not started: no laboratory has joined yet. Its public tier is the explorer above.
Checks fixed before the run
All five checks we fixed before the run passed. The checks are ours, not an outside review. The run covered the 34 public method comparisons in the explorer.
| Check fixed before the run | Result |
|---|---|
| Results unchanged from the previous version (41 of 41)The program reproduces the earlier version's comparison lines and crossing counts. | Passed |
| Re-test budget held on the fitted dataOn each comparison's own pairs, the share of results in the re-test band stays within the budget fixed in advance. | Passed |
| Re-test band holds on unseen dataFor every comparison large enough for a full band, the re-test share on held-out data stays within the limit fixed in advance. | Passed |
| The small study that failed before now holdsThe small comparison on which the earlier version re-tested far too many results now stays within the limit on held-out data. | Passed |
| Validation cases (15 of 15)Awkward export files with missing columns, decimal commas, text results, duplicate IDs and similar problems each give the expected outcome. | Passed |
Source 12 open-access method-comparison studies, each cited with its authors, source and licence on the public data page. Run file sha256 5586cd892ef167b9b196e27b6b2490b90fc0f6ca3a66ee161cd8fd0b10534223, run date 25 September 2026.
Available on request
- Offer
Technical validation kit
A container you run offline. It rebuilds our headline result from public data, compares it with the sealed result and prints pass or fail. On request.
- Offer
Evidence library and console
A library of methods, runs and limits, and the console we run analyses from. Available on request.