Guide
Central lab, local lab, new device: measurement change in trials and in medical AI
In a trial, a laboratory value can decide who is eligible and whether an endpoint is met. In a model, it can decide a prediction. When the laboratory, the device or the site behind that value changes, some people move across the threshold for reasons that have nothing to do with them. This note sets out how to size that movement, and what we have and have not done.
Where a lab value decides
Many protocols turn on a laboratory number: an HbA1c range for eligibility, a haemoglobin floor for safety, a creatinine limit for dosing, a change in a marker as the endpoint. Each is a threshold, and a threshold converts a small measurement difference into a different decision.
Whether to measure locally or centrally is an old question with a measured answer. In a multicentre trial of children with newly diagnosed type 1 diabetes, HbA1c measured at the local centres was compared with the same children’s samples measured centrally: almost six hundred pairs from fifteen centres. There was no overall bias, yet only four in five local results were within ten percent of the central one, discrepancies were present at every centre, and the authors recommended central measurement for multicentre trials [Arch et al., 2016].
Eligibility is where it bites first. In one study of point-of-care haemoglobin meters, the authors note that anaemia is a frequent reason for exclusion from clinical studies, and that their evaluation began after they saw discrepant results between a meter and the laboratory during screening for a clinical study [Jaggernath et al., 2016]. The current good clinical practice guideline expects investigators to review data from external sources that can affect eligibility, treatment or safety, and names central laboratory data as an example [ICH E6(R3), 2025].
What changes between sites and periods
- Between sites: different analysers or methods for the same analyte, point-of-care devices against a central laboratory, capillary against venous blood, different operators.
- Within a site, over time: an analyser replaced mid-study, a new reagent or calibrator lot, a device firmware update, a move from local to central testing or back.
- Between a model’s development and its use: any of the above, behind the values the model reads.
None of these needs to show up as an average bias. The HbA1c study above found none overall, and still found individual differences wide enough to matter at a threshold.
Sizing crossings at the protocol threshold
The question a trial team needs answered is not whether two measuring systems agree on average, but how many participants would be on the other side of the protocol’s threshold if their sample had been measured on the other system. With paired results, that is three numbers per threshold: the bias at the threshold, read off the comparison line; the number of participants whose category differs between the systems; and the number that repeat testing on a single system would move anyway, from its imprecision. The excess over that baseline is what the change did. The guide to bias at clinical decision limits sets out the method.
Fix the check before unblinding
Offer
A comparability check that is designed after the outcome is known invites the suspicion that it was designed to find, or not find, something. The good clinical practice guideline expects a statistical analysis plan consistent with the protocol, criteria for including participants in an analysis set defined in advance, and any change made after unblinding documented, justified and reported; it also asks sponsors to set pre-specified acceptable ranges for risks to what is critical to quality [ICH E6(R3), 2025].
What we offer follows the same discipline. The thresholds, the analysis and the pass criteria are written down and their cryptographic hash recorded before the data are opened, and any later change is a dated amendment. The check runs inside the sponsor’s environment and only aggregate results leave it. This is an offer: it has not yet been run on a sponsor’s trial data. Our own public run works this way: all five checks fixed before the run passed. The checks are ours, not an outside review, and the configuration hash is printed under every figure below.
The same logic for models whose inputs are lab values
Offer
A model inherits every instrument that feeds it. Clinicians and developers have been warned that dataset shift can come from changes in technology as well as in populations and behaviour; one of the examples given is that adopting high-sensitivity troponin assays changes the clinical interpretation of detectable troponin levels [Finlayson et al., 2021]. The TRIPOD+AI reporting guideline asks model reports to define every predictor, including how and when it was measured, and to identify any differences between development and evaluation data in healthcare setting, eligibility, outcome and predictors [Collins et al., 2024].
When the analyser, device or site behind a model’s inputs changes, the method above sizes how far those inputs moved at the thresholds the model relies on, against measurement noise, before anyone looks at the model’s output. This is an offer: it has not been run on any model’s inputs or outputs.
Worked examples on public data
Public data
Neither example comes from a trial. Both are public method comparisons of the kind a trial would face when one site uses a point-of-care device and another a central laboratory. The decision lines are set in our program, not taken from any protocol, and repeat-test imprecision is assumed rather than measured for both.
The first compares a point-of-care haematology analyser with a central laboratory analyser on 201 samples from hospital and outpatient settings [Shean and Bennett, 2026], with haemoglobin lines at 8, 12 and 13 g/dL.
Comparison line with slope 1.035 (95% interval 1.015 to 1.056) and intercept -0.411, drawn against the line of no change. Decision limits drawn at 8, 12, 13 g/dL. Where the comparison line sits above the line of no change, the new system reads higher; below, lower.
Fig. 1. Central laboratory against point of care: the average difference is minus 0.1 percent, yet 24 results changed haemoglobin category against about 18 expected from repeat testing alone.
Source Shean RC, Bennett ST (2026). Clinical Evaluation of a Novel Point-of-Care Hematology Analyzer for Complete Blood Count With Differential. Int J Lab Hematol. doi:10.1111/ijlh.70032 (opens in a new tab); PMC12956494 (opens in a new tab). Licence: CC BY 4.0 (opens in a new tab). Changes: Paired results re-analysed by NakedSignal's program; instruments described generically; only derived results are shown. Repeat-test imprecision assumed, not measured.
Show the numbers
| Pairs | 201 |
|---|---|
| Slope with its interval | 1.035 (1.015 to 1.056) |
| Intercept | -0.41 |
| Changed category | 24 |
| Expected from repeat testing | 17.7 |
| Decision lines | 8, 12 and 13 g/dL |
An average difference close to zero does not mean results are interchangeable at a threshold. For a trial that used this haemoglobin value as a safety floor, the count against repeat testing is the number that says whether the sites can be pooled. Open this series in the explorer.
The second example sets two point-of-care HbA1c devices against a laboratory analyser, on fingerstick blood in a prospective accuracy study [Giachino et al., 2024], at lines of 5.7, 6.5 and 7 percent.
Point of care device A
Crossed after the change: 236
Expected from repeat testing alone: about 59
Point of care device B
Crossed after the change: 89
Expected from repeat testing alone: about 60
Fig. 2. Two point-of-care HbA1c devices against one laboratory analyser at 5.7, 6.5 and 7 percent: point of care device A moved 236 results across a line against about 59 from repeat testing and point of care device B moved 89 results across a line against about 60 from repeat testing.
Source Giachino M, Vetter B, Perone SA, Correia JC, Erkosar B, Heller O, Khanal VK, Lab B, Pataky Z, Poudel S, Rai M, Sharma SK (2024). Performance and usability of cardiometabolic point of care devices in Nepal: A prospective, quantitative, accuracy study. PLOS Glob Public Health. doi:10.1371/journal.pgph.0003760 (opens in a new tab); PMC11449279 (opens in a new tab). Licence: CC BY 4.0 (opens in a new tab). Changes: Paired results re-analysed by NakedSignal's program; instruments described generically; only derived results are shown. Repeat-test imprecision assumed, not measured.
Show the numbers
| Device | Pairs | Slope | Mean difference, as % of old | Changed category | Expected from repeat testing |
|---|---|---|---|---|---|
| Point of care device A | 352 | 1.103 | +18.5% | 236 | 59.2 |
| Point of care device B | 352 | 0.833 | −2.6% | 89 | 59.6 |
Two devices measuring the same people give different answers at the same lines. In a trial with sites using each device, the eligible population could differ by site before any participant was randomised. That is the kind of effect a pre-specified check would size before the data are pooled.
Held-out evaluation: MedEval-1
Live demo
Sizing input changes is one half of keeping a model honest; evaluating it on data it has not seen is the other. MedEval-1 is a working version of a versioned evaluation standard for medical AI, built by NakedSignal. Its held-out track seals, scores and rotates its test set, and 48 of 48 published scores were re-derived by our own audit from the published manifests.
What this does not cover
- Trial design and statistics. Sample size, randomisation and the primary analysis are the sponsor’s statisticians’ work. The check described here sits beside them.
- Regulatory submissions. Nothing here is regulatory advice or a claim about what any regulator will accept.
- Fixing a model. Sizing how far a model’s inputs moved does not retrain, recalibrate or validate the model.
- Individual participants. Counts describe groups of results. They are not advice about any one participant’s eligibility or care.
Sources
- Arch BN, Blair J, McKay A, Gregory JW, Newland P, Gamble C (2016). Measurement of HbA1c in multicentre diabetes trials - should blood samples be tested locally or sent to a central laboratory: an agreement analysis. Trials. pmc.ncbi.nlm.nih.gov/articles/PMC5078896/
- Jaggernath M, Naicker R, Madurai S, Brockman MA, Ndung'u T, Gelderblom HC (2016). PLoS ONE. pmc.ncbi.nlm.nih.gov/articles/PMC4821624/ The article’s title names the devices studied and is left out here.
- International Council for Harmonisation (2025). Guideline for Good Clinical Practice E6(R3). ICH harmonised guideline, final version adopted 6 January 2025. database.ich.org/sites/default/files/ICH_E6(R3)_Step4_FinalGuideline_2025_0106.pdf
- Finlayson SG, Subbaswamy A, Singh K, Bowers J, Kupke A, Zittrain J, Kohane IS, Saria S (2021). The Clinician and Dataset Shift in Artificial Intelligence. New England Journal of Medicine. pmc.ncbi.nlm.nih.gov/articles/PMC8665481/
- Collins GS, Moons KGM, Dhiman P, Riley RD, Beam AL, Van Calster B, et al. (2024). TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. pmc.ncbi.nlm.nih.gov/articles/PMC11019967/
- Shean RC, Bennett ST (2026). Clinical Evaluation of a Novel Point-of-Care Hematology Analyzer for Complete Blood Count With Differential. International Journal of Laboratory Hematology. pmc.ncbi.nlm.nih.gov/articles/PMC12956494/
- Giachino M, Vetter B, Perone SA, Correia JC, Erkosar B, Heller O, Khanal VK, Lab B, Pataky Z, Poudel S, Rai M, Sharma SK (2024). Performance and usability of cardiometabolic point of care devices in Nepal: A prospective, quantitative, accuracy study. PLOS Global Public Health. pmc.ncbi.nlm.nih.gov/articles/PMC11449279/
Each link was opened and checked on 26 September 2026.