NakedSignal

For trials, health data and AI

Site, lab and device differences, sized at your thresholds.

When eligibility or an endpoint rests on a lab value or a device reading, a change of platform between sites or periods can move patients across the protocol's threshold. We size that movement against measurement noise, with the check fixed before the data are opened.

01Evidence

What you can open now

Public data

HbA1c at 5.7, 6.5, 7 %

Crossed after the change: 236

Expected from repeat testing alone: about 59

Fig. 1. The same people measured on a central lab chemistry analyser and on a point of care device A. At the HbA1c lines of 5.7, 6.5, 7%, 236 results fall in a different category, against about 59 expected from repeat testing alone.

Source Giachino M, Vetter B, Perone SA, Correia JC, Erkosar B, Heller O, Khanal VK, Lab B, Pataky Z, Poudel S, Rai M, Sharma SK (2024). Performance and usability of cardiometabolic point of care devices in Nepal: A prospective, quantitative, accuracy study. PLOS Glob Public Health. doi:10.1371/journal.pgph.0003760 (opens in a new tab); PMC11449279 (opens in a new tab). Licence: CC BY 4.0 (opens in a new tab). Changes: Paired results re-analysed by NakedSignal's program; instruments described generically; only derived results are shown. Repeat-test imprecision is assumed, not measured. Open this comparison in the explorer.

run 25 September 2026 · config 6eb23df45893 · run file sha256 5586cd892ef1

Show the numbers
HbA1c: Central lab chemistry analyser to Point of care device A
MeasureValue
Paired results352
Decision limits (%)5.7, 6.5, 7
Results that changed category after the change236
Expected to change from repeat testing alone59.17
Mean difference, new minus old, as % of old18.49
Public data

Central laboratory versus point of care, on public data

Haemoglobin measured on a central lab haematology analyser and on a point of care haematology analyser: 201 pairs from an open-access deposit, with the comparison line, bias at the decision limits and results crossing each limit against repeat testing alone. The explorer also holds 10 HbA1c and creatinine comparisons, analytes that often set trial eligibility.

run 25 September 2026 · config 6eb23df45893 · run file sha256 5586cd892ef1

Live demo

MedEval-1

A working version of a versioned evaluation standard for medical AI, built by NakedSignal. Its held-out track seals, scores and rotates its test set. 48 of 48 published scores were re-derived by our own audit from the published manifests.

medeval-1.netlify.app

02Clinical trials

Multi-site trials

OfferSomething we do for you. Not yet run on a sponsor’s trial data.

For a lab value that sets eligibility or an endpoint, we run a pre-specified check inside your environment and report bias and crossings at those thresholds, against what repeat testing alone would move.

The check, its thresholds and its pass criteria are fixed and hashed before the data are opened. Changes after that are dated amendments. It is designed around the data-reliability expectations of ICH E6(R3) Good Clinical Practice.

Guide: Measurement change in trials and medical AI

Whether to test locally or send samples to a central laboratory is a known question in multicentre trials; one agreement analysis of HbA1c in diabetes trials compared the two directly.

Arch BN, Blair J, McKay A, Gregory JW, Newland P, Gamble C (2016). Measurement of HbA1c in multicentre diabetes trials - should blood samples be tested locally or sent to a central laboratory: an agreement analysis. Trials. doi:10.1186/s13063-016-1640-6 (opens in a new tab)

Guide: Point-of-care HbA1c against the laboratory

03Study devices

Device changes in trials

OfferSomething we do for you. Not yet run on a study’s device data.

Trials that take a measure from a device often change device model or firmware during follow-up, or run different device generations at different sites. An endpoint measured that way can then move for reasons that have nothing to do with the participant.

When a study replaces a device, updates its firmware or changes the software that derives an endpoint, we write the bridging protocol before the change and report afterwards whether the endpoint moved, set against measurement noise. Typical cases: a new device generation in a sponsor study, or a device or operating-system update during a multi-year study.

For each measure the protocol relies on, we hand over:

  • the size of the shift between device generations, with its uncertainty;
  • how many participants’ readings move across the protocol’s threshold, set against what measurement noise alone would move;
  • where participants were measured on both devices for a period, a bridging correction with its uncertainty, fixed before the endpoint data are unblinded.

What it needs: a change record (old and new device model, firmware and settings, and the month of the change) and either a period in which both devices were used or enough participants measured on each device. The analysis can run in the sponsor’s environment, and only aggregates leave it.

Our device-change work has so far been tested on public data only. We describe what it produces, not how it works; method details are shared under a confidentiality agreement.

Holders of study device data can also contribute device-change aggregates to the Analyser-Change Registry.

04Health data and medical AI

Model inputs and held-out evaluation

OfferNot yet run on a model’s inputs.

A model inherits every instrument that feeds it. When a site, device or analyser behind its inputs changes, the same method sizes how far those inputs moved at the thresholds the model relies on. We scope it with you as a pre-specified check.

For evaluation, MedEval-1 is the working example: a held-out test set that is sealed, scored and rotated, with every published score traceable to its manifest.

A threshold in your protocol or your model?

Tell us which value it rests on and where it is measured. We will say what a pre-specified check would show and what it needs.