Guide
Bias at clinical decision limits: what a method comparison misses
When a laboratory changes analyser, the method comparison asks whether the two agree. The clinician’s question is different: whose result now lands on the other side of a decision limit. This note sets out how to answer both.
Two questions, not one
A method comparison answers a numerical question: how far apart are results from the old and the new measurement procedure, and is that within what the laboratory will accept [CLSI EP09]. Near a decision limit there is a second, interpretive question: does the same patient land in the same category on both? The two can disagree. Numerically similar results can fall on opposite sides of a limit, and a difference outside the allowable error can leave a patient’s category unchanged [Cai et al., 2026].
The second question matters because clinical decisions are made at limits, not at averages. A laboratory that reports a difference in percent has answered the first question. Clinicians and quality managers usually need the second.
The comparison line: Passing-Bablok and Bland-Altman in brief
Passing-Bablok regression fits a straight line of new results on old without assuming either method is free of error, and is robust to outliers [Passing and Bablok, 1983]. Its slope says whether the difference grows with concentration (proportional) and its intercept whether there is a fixed offset (constant). Both come with intervals, usually from the method’s own procedure or from bootstrap resampling.
Bland-Altman analysis plots each pair’s difference against its mean and reports the average difference with limits of agreement, the range in which most differences fall [Bland and Altman, 1986]. It is the familiar summary of agreement across the whole range.
Both are standard tools for verifying comparability in a medical laboratory, alongside Deming regression and paired difference tests [Abdel Ghafar and El-Masry, 2021]. Neither, on its own, says what happens at a particular limit.
Bias at the limit, with an interval
With a comparison line in hand, the systematic error at any concentration can be read off it: that is how method-validation practice estimates error at medical decision concentrations rather than as one average [Westgard]. At a decision limit L, with slope b and intercept a:
bias at L = (b × L + a) − L
Expressed as a percentage of L, this is the figure to set against the laboratory’s acceptance limit at that concentration. Repeating the fit on bootstrap resamples of the pairs gives an interval for it. The point is not the formula, which is simple, but the habit: report bias at each limit the laboratory publishes, not once for the whole range. A proportional difference that looks small on average can be large at a high limit, and a constant offset matters most at a low one.
Counting patients across a line
Bias at a limit is still a property of the method. The number the clinician asks for is a count: how many patients were in one category on the old analyser and in another on the new. From the paired results this is direct. Place each pair’s old result and new result into the categories the limits define, and count the pairs whose category differs, separately for each limit and each direction.
Direction matters. A new analyser that reads lower moves patients down across a limit; one that reads higher moves them up; a change in slope can do both at different limits. A table of old category by new category shows all of it at once.
Why repeat testing sets the baseline
Some of those crossings would have happened with no change at all. Measure a sample twice on one analyser and a result close to a limit can land on either side of it, purely from imprecision. So a raw count of crossings overstates what the change did.
The baseline comes from the analyser’s repeat-test imprecision (its coefficient of variation). For each result, take the probability that a repeat measurement on the same analyser would put it in a different category; summed over all results, that is the number of crossings to expect from repeat testing alone. The change’s own effect is the excess of observed crossings over that expectation. If the laboratory has QC or duplicate data, the imprecision should be measured; if a default is used, the report should say so.
A re-test band follows naturally. Results that land within a narrow range of a limit on the new analyser are the ones most likely to have crossed because of the change or by chance. Re-testing them during the transition catches much of the excess for a known share of extra tests, and the share can be fixed in advance as a budget.
A worked example
Sample
The example uses a public CD4 method comparison: 1,885 paired counts from an open-access deposit in which one flow-cytometry platform replaced another [Coetzee and Glencross, 2017]. We ran it through our program as if a laboratory had exported it, with decision limits at 100, 200, 350, 500 cells/µL. It is a sample built from public data: no laboratory commissioned or reviewed it, and repeat-test imprecision is assumed.
Comparison line with slope 0.867 (95% interval 0.856 to 0.877) and intercept -1.2, drawn against the line of no change. Decision limits drawn at 100, 200, 350, 500 cells/µL. Where the comparison line sits above the line of no change, the new system reads higher; below, lower.
Fig. 1. Comparison line for the CD4 sample: slope 0.87, so the new platform reads about 14% lower on average, with the difference growing with the count.
Source Coetzee LM, Glencross DK (2017). PLoS One. doi:10.1371/journal.pone.0187456 (opens in a new tab); PMC5669480 (opens in a new tab) (title on the source page; it names the instruments). Licence: CC BY 4.0 (opens in a new tab). Changes: Paired results re-analysed by NakedSignal's program; instruments described generically; only derived results are shown. Public data, not a laboratory’s.
Show the numbers
| Pairs | 1,885 |
|---|---|
| Slope with its interval | 0.867 (0.856 to 0.877) |
| Intercept | -1.2 |
| Decision limits | 100, 200, 350, 500 cells/µL |
All decision limits combined
Crossed after the change: 535
Expected from repeat testing alone: about 238
Fig. 2. Across the decision limits, 535 results changed category against about 238 expected from repeat testing alone.
Source Same deposit, CC BY 4.0. Repeat-test imprecision assumed, not measured.
Reported without adjustment, the change moves 535 results across a limit where repeat testing alone would move about 238. Correcting with the comparison line and re-testing results inside a band around each limit removed 66% of that excess on pairs the band was not fitted to, against 46% for the correction line alone, while re-testing 15% of results. The sample report shows the bias at each limit read off the comparison line, and the re-test band.
What this does not cover
- Commutability. Whether control or EQA material behaves like patient samples on both methods. The approach here uses patient pairs and says nothing about control material.
- Reagent and calibrator lot changes. The same counting applies, but lot-to-lot studies are usually smaller and need their own acceptance rules.
- ISO 15189 verification. Counting patients across limits supports a laboratory’s verification and its communication with clinicians. It does not replace the verification the standard requires.
- Individual patients. Counts describe a population of results. They are not advice about any one patient.
Sources
- Cai X, Yang W, Lu Q, Lin Y (2026). Numerical and interpretive comparability of optical and mechanical coagulation analyzers: a Zlog-supported method comparison study. Diagnostics (Basel). pmc.ncbi.nlm.nih.gov/articles/PMC13511516/
- Clinical and Laboratory Standards Institute. EP09: Measurement procedure comparison and bias estimation using patient samples. CLSI guideline, third edition. clsi.org/standards/products/method-evaluation/documents/ep09/
- Passing H, Bablok W (1983). A new biometrical procedure for testing the equality of measurements from two different analytical methods. Part I. Journal of Clinical Chemistry and Clinical Biochemistry. doi.org/10.1515/cclm.1983.21.11.709
- Bland JM, Altman DG (1986). Statistical methods for assessing agreement between two methods of clinical measurement. The Lancet. doi.org/10.1016/s0140-6736(86)90837-8
- Abdel Ghafar MT, El-Masry MI (2021). Verification of quantitative analytical methods in medical laboratories. Journal of Medical Biochemistry. pmc.ncbi.nlm.nih.gov/articles/PMC8199534/
- Westgard JO. The comparison of methods experiment. Westgard QC, Basic Method Validation. www.westgard.com/lessons/basic-method-validation/lesson23.html
- Coetzee LM, Glencross DK (2017). PLoS One. doi.org/10.1371/journal.pone.0187456
Each link was opened and checked on 26 September 2026. ISO 15189:2022 is named but not linked because the publisher’s page could not be opened for checking.