NakedSignal

Guide

Are your analysers interchangeable? Comparability of results across instruments and sites

A laboratory that runs the same test on more than one analyser, or at more than one site, is expected to show that a patient’s result does not depend on which one was used. This note sets out what the standards ask, and how to report comparability where it matters: at the decision limits.

By NakedSignal · Updated 26 September 2026

What ISO 15189:2022 asks

ISO 15189:2022 sets the requirements for quality and competence in medical laboratories. Its clause on the comparability of examination results covers examinations done on different methods or equipment, or at different sites. As a published summary puts it, the laboratory needs a process to assess that comparability, and where differences are found, their impact on biological reference intervals and clinical decision limits should be evaluated and acted on, with users told of any clinically significant difference [Marrington et al., 2026].

Two things in that requirement are easy to miss. It is not limited to two analysers of the same model: a point-of-care device and the main laboratory analyser measuring the same analyte are in scope. And the test of success is clinical, not statistical: whether a difference changes what a result means at a reference interval or decision limit.

The EP31 approach

CLSI EP31 gives recommendations for verifying the comparability of quantitative results for individual patients between two or more instruments within one health care system, whether the instruments differ in method, in model, or are separate instruments of one type [CLSI EP31, 2025]. Its protocol compares patient samples, chosen at the concentrations of interest, across all the instruments, using a range test for up to ten instruments and modified strategies above that.

In routine use it works, with a practical caveat. A laboratory that applied it to two dozen chemistry tests across several analysers found the guideline useful for periodic verification, but noted that it may need acceptance criteria adjusted to the analytical performance of the technology available, and a number of measurements matched to the laboratory’s resources [Guiñón et al., 2024].

Patient samples, not only control material

Control and EQA material is designed to be stable, not to behave exactly like a patient sample. Two methods can agree on control material and still disagree on patient blood, and the difference can depend on the sample type as much as the instrument. That is why EP31 is framed around patient results, and why a periodic comparison with real samples, near the limits the laboratory reports against, is worth more than a control chart for this particular question.

Setting matters as well as sample. The same device used by a different operator, on capillary rather than venous blood, away from the laboratory, is in practice a different measuring system. The worked example below shows how much that can change.

Comparability at the decision limits

A comparability check usually ends in one number per analyte: a mean difference, or a maximum difference against an allowable limit. For the clinical question the standard asks, three further numbers are more useful, each one per decision limit.

  • Bias at the limit. Read from the comparison line at the limit itself, not averaged over the range. A proportional difference can be small at one limit and large at another.
  • Results that change category. Of the paired results, how many fall on one side of the limit on one instrument and on the other side on the second.
  • The repeat-testing baseline. How many of those changes one instrument measuring the same samples twice would produce anyway, from its imprecision. Only the excess is due to the instruments differing.

The third number needs the imprecision of each instrument. Where replicate measurements exist, use them; where a default is assumed, say so. The guide to bias at clinical decision limits works through the method in detail.

One analyser, three meters, two settings

Public data

A public study compared three point-of-care haemoglobin meters with one laboratory haematology analyser as the reference, first in a central laboratory and then in a community clinic, where the meters were used on finger-prick blood and the reference on a venous sample drawn at the same visit [Jaggernath et al., 2016]. Its data are deposited under CC BY 4.0 [Gelderblom, 2016]. The study also repeat-tested a subset of samples, so imprecision can be measured rather than assumed, which makes these the only comparisons in our public set where it is. We ran all six comparisons through our program with limits at 8, 12 and 13 g/dL; the study itself used 12 and 13 g/dL as its anaemia thresholds for women and men.

Public data

Meter A, central laboratory

Crossed after the change: 14

Expected from repeat testing alone: about 2

Meter B, central laboratory

Crossed after the change: 16

Expected from repeat testing alone: about 32

Meter C, central laboratory

Crossed after the change: 36

Expected from repeat testing alone: about 12

Meter A, community clinic

Crossed after the change: 36

Expected from repeat testing alone: about 7

Meter B, community clinic

Crossed after the change: 37

Expected from repeat testing alone: about 62

Meter C, community clinic

Crossed after the change: 65

Expected from repeat testing alone: about 25

Fig. 1. Against one laboratory analyser, three meters moved different numbers of results across the haemoglobin limits in the central laboratory and community clinic; for meter b, repeat testing alone would move more results than the comparison did.

Source Data: Huub Gelderblom (2016). Performance characteristics of three point of care hemoglobin meters in Durban-2.xlsx. figshare. doi:10.6084/m9.figshare.3119449 (opens in a new tab). Licence: CC BY 4.0 (opens in a new tab). Supplement to Jaggernath M, Naicker R, Madurai S, Brockman MA, Ndung'u T, Gelderblom HC (2016). PLoS One. doi:10.1371/journal.pone.0152184 (opens in a new tab); PMC4821624 (opens in a new tab) (title on the source page; it names the instruments). Changes: Paired results re-analysed by NakedSignal's program; instruments described generically; only derived results are shown. Measured: precision file (patient-sample replicates (5 per sample)), the larger of the two systems.

run 25 September 2026 · config 6eb23df45893 · run file sha256 5586cd892ef1

Show the numbers
ComparisonPairsSlopeMean difference, as % of oldChanged categoryExpected from repeat testingRe-test band
Meter A, central laboratory600.990 (0.946 to 1.000)+4.7%141.6Indicative
Meter B, central laboratory561.131 (1.000 to 1.369)−12.5%1632.4Indicative
Meter C, central laboratory590.800 (0.722 to 0.909)−24.1%3612.3Indicative
Meter A, community clinic1001.047 (0.946 to 1.177)−1.3%366.5Full
Meter B, community clinic1001.031 (0.938 to 1.143)−4.4%3762.3Full
Meter C, community clinic1000.727 (0.621 to 0.857)−23.6%6525.1Full

What the figure shows

Comparability belongs to a system in a setting, not to a device. A meter need not behave the same way in the laboratory and in the clinic: compare each meter’s slope and mean difference across the two settings under “Show the numbers”. The study’s authors reached the same conclusion from their own analysis: good performance of a point-of-care test in a central laboratory does not guarantee good performance in a community based clinic setting [Jaggernath et al., 2016]. A laboratory network verifying comparability should test each system where and how it is used.

Some disagreement is noise, not bias. For meter b, the number of results expected to change category from repeat testing alone is larger than the number that did. The expected count uses the larger of the two systems’ measured imprecision, so near the limits the count cannot separate a difference between the systems from that meter’s own scatter. Its imprecision has to be dealt with first, and a correction factor alone will not do that.

Small studies are graded, not hidden. The comparisons from the central laboratory phase have fewer pairs than our program requires for a full re-test band, so their bands are shown as indicative. A laboratory’s own periodic comparison will often be this size; the count against repeat testing is still informative, the band is less certain.

A second public comparison, from a clinical reference laboratory in Cameroon, sets a point-of-care analyser against a laboratory haematology analyser on 128 pairs [Sagnia et al., 2024]; it is in the change explorer with the six above.

What this does not cover

  • Traceability and calibration. Why two methods differ, and how metrological traceability and harmonisation programmes reduce it, is a separate subject.
  • EQA scheme design. How schemes choose materials and assess participants is not covered.
  • The text of the standards. ISO 15189:2022 is paraphrased here from a published summary and EP31 from its publisher’s description. The documents themselves are the authority.
  • Qualitative tests and individual patients. The counts apply to quantitative results and describe groups of results. They are not advice about any one patient or device purchase.

Sources

  1. Marrington R, Anderson M, Roch M, Davies S, Forster R, Williams G, d'Oyen-Fitchett J, MacKenzie F (2026). Use of EQA for assessing variation in post-analytical interpretation of results using a clinical case of ?adrenal insufficiency taking into consideration variation in analytical performance of cortisol. British Journal of Biomedical Science. pmc.ncbi.nlm.nih.gov/articles/PMC13461562/
  2. Clinical and Laboratory Standards Institute (2025). Verification of Comparability of Patient Results Within One Health Care System. CLSI guideline EP31, second edition. clsi.org/shop/standards/ep31-plus/
  3. Guiñón L, Illana FJ, Cuevas B, Canyelles M, Martínez-Bru C, García-Osuna Á (2024). Periodic verification of results' comparability between several analyzers: experience in the application of the EP31-A-IR guideline. Clinical Chemistry and Laboratory Medicine. doi.org/10.1515/cclm-2023-0994
  4. Jaggernath M, Naicker R, Madurai S, Brockman MA, Ndung'u T, Gelderblom HC (2016). PLoS ONE. pmc.ncbi.nlm.nih.gov/articles/PMC4821624/ The article’s title names the devices studied and is left out here.
  5. Gelderblom H (2016). Performance characteristics of three point of care hemoglobin meters in Durban-2.xlsx. figshare data set, CC BY 4.0. doi.org/10.6084/m9.figshare.3119449
  6. Sagnia B, Mbakop Ghomsi F, Moudourou S, Gutierez A, Tchadji J, Sosso SM, Ndjolo A, Colizzi V (2024). Accurate and reproducible enumeration of CD4 T cell counts and Hemoglobin levels using a point of care system: Comparison with conventional laboratory based testing systems in a clinical reference laboratory in Cameroon. PLoS ONE. pmc.ncbi.nlm.nih.gov/articles/PMC10954178/

Each link was opened and checked on 26 September 2026. ISO 15189:2022 is named but not linked because the publisher’s page could not be opened for checking.