NakedSignal

Guide

How many results cross a decision limit by chance?

Measure the same sample twice and a result near a decision limit can land on either side of it. This note shows how to turn imprecision and distance from the limit into an expected count of crossings, and why that count is the baseline any change of method has to beat.

By NakedSignal · Updated 26 September 2026

The question

When a laboratory changes analyser, reagent or method, the paired results tell it how many patients now fall in a different category. That count alone does not say what the change did, because some of those patients would have changed category on a simple repeat, with nothing changed at all. The useful number is the difference between what was observed and what repeat testing alone would produce.

The same arithmetic answers a question laboratories ask without any change in view: if imprecision rose, or a small bias crept in, how many patients would be classified differently? Simulation studies of fixed interpretive limits show the share of patients reclassified rising with both imprecision and bias [Loh et al., 2024].

One result near a limit

Take a result x, a decision limit L and the analyser’s repeat-test imprecision as a coefficient of variation (CV). A repeat measurement of the same sample varies around the first. Because the first result and the repeat both carry imprecision, the spread of the difference between them is larger than either alone by a factor of √2. Treating that difference as normally distributed:

spread of a repeat = √2 × CV × x

chance of crossing L = Φ(−|L − x| ÷ spread)

Φ is the standard normal distribution function. The chance depends on one ratio: how far the result sits from the limit, measured in units of that spread. A result exactly at the limit has an even chance of landing on either side. One spread away, the chance is about 16 in 100; two spreads away, about 2 in 100. So only results near a limit contribute much, and how near counts as near is set by the CV.

From one result to a count

Summing that chance over every result in the comparison gives the number of results expected to change category on repeat testing alone. Where a laboratory reports several limits, the chance for each result is the probability of leaving its own category, above or below. The sum is an expected value, not a prediction for any one patient, and it can be a fraction.

Two things drive it: how many results sit close to the limits, and how large the CV is there. A comparison in which most results cluster around a limit will show many crossings from imprecision alone, even with a perfectly matched new method. That is why a raw count of crossings is not a verdict on the change.

Bias against imprecision

Imprecision scatters results in both directions, so its crossings roughly cancel in the totals above and below a limit. Bias moves every result the same way. A new method that reads a few percent lower pushes all results just above a limit to just below it, and those moves add up in one direction.

In simulation, the effect of imprecision on reclassification lessens as bias grows, and at large biases bias becomes the dominant cause [Loh et al., 2024]. In practice this means a small change with little bias usually shows crossings close to the repeat-testing baseline, while a larger bias shows an excess that is mostly in one direction. Counting crossings separately for each limit and each direction makes that visible.

Measured or assumed imprecision

The expected count is only as good as the CV behind it. A larger CV widens the spread, so more results count as near a limit and the baseline rises; a smaller CV lowers it. An assumed CV that is too large makes a real change look like noise, and one that is too small makes noise look like a change.

Where the laboratory has replicate measurements of patient samples, or duplicate or quality-control data at concentrations near its limits, the CV should come from those. Where it does not, a default can be used, and the result should say so. CV also varies with concentration, so the value at the limit matters more than a single figure across the whole range.

Reference change values. The reference change value asks whether two results from one patient differ by more than analytical and within-subject biological variation would explain. Its formula carries the same √2 factor, because it too compares two results that each carry variation [Fraser, 2011]. Estimates of within-subject variation for many measurands are collected in the [EFLM Biological Variation Database]. The baseline here uses analytical imprecision only, because the question is what a repeat of the same sample would do, not what the patient would do over time.

Analytical performance specifications. The 2014 consensus of the European Federation of Clinical Chemistry and Laboratory Medicine sets out three models for deciding how good a method must be: the effect of analytical performance on clinical outcomes, biological variation, and the state of the art [Sandberg et al., 2015]. The first model includes indirect outcome studies of the effect of performance on clinical classifications, for example by simulation. Counting results that change category against the repeat-testing baseline is a direct, local version of that question, run on the laboratory’s own paired results rather than on a simulated population.

Worked examples on public data

Public data

We ran open-access method comparisons through our program as if a laboratory had exported them. Fig. 1 shows four of the 34 comparisons, chosen to show both outcomes: changes that moved more results than repeat testing alone would, and changes that moved fewer. The first is the public CD4 sample, where 535 results changed category against about 238 expected from repeat testing [Coetzee and Glencross, 2017]. The imprecision there is assumed, not measured.

In the second, a change of creatinine chemistry on one analyser, 23 results changed category against about 41 expected from repeat testing [Schmidt et al., 2015]. On these pairs, the change moved fewer results than a simple repeat would have, and the honest report is that there is nothing extra to manage at these limits.

The last two are haemoglobin meters compared with a laboratory analyser in a community clinic, the only public series we hold with imprecision measured from replicate patient samples [Jaggernath et al., 2016], data [Gelderblom, 2016]. With similar numbers of crossings (36 and 37), one meter sits far above its repeat-testing baseline and the other below it. The baselines differ because the imprecision measured for each meter differs. The crossing count alone would have ranked them the same.

Public data

CD4 absolute count: Flow cytometry platform A (predicate) to flow cytometry platform B (new), data set 2. Imprecision assumed.

Crossed after the change: 535

Expected from repeat testing alone: about 238

Creatinine: Central lab analyser, enzymatic creatinine to same analyser, kinetic Jaffe creatinine. Imprecision assumed.

Crossed after the change: 23

Expected from repeat testing alone: about 41

Haemoglobin: Laboratory haematology analyser (reference) to point of care haemoglobin meter A, community clinic setting. Imprecision measured.

Crossed after the change: 36

Expected from repeat testing alone: about 7

Haemoglobin: Laboratory haematology analyser (reference) to point of care haemoglobin meter B, community clinic setting. Imprecision measured.

Crossed after the change: 37

Expected from repeat testing alone: about 62

Fig. 1. Results that changed category after the change, against the number expected from repeat testing alone, in four public method comparisons. Two exceed the baseline; two fall below it.

Source Coetzee LM, Glencross DK (2017). PLoS One. doi:10.1371/journal.pone.0187456 (opens in a new tab); PMC5669480 (opens in a new tab) (title on the source page; it names the instruments). Licence: CC BY 4.0 (opens in a new tab). Schmidt RL, Straseski JA, Raphael KL, Adams AH, Lehman CM (2015). A Risk Assessment of the Jaffe vs Enzymatic Method for Creatinine Measurement in an Outpatient Population. PLoS One. doi:10.1371/journal.pone.0143205 (opens in a new tab); PMC4657986 (opens in a new tab). Licence: CC BY 4.0 (opens in a new tab). Data: Huub Gelderblom (2016). Performance characteristics of three point of care hemoglobin meters in Durban-2.xlsx. figshare. doi:10.6084/m9.figshare.3119449 (opens in a new tab). Licence: CC BY 4.0 (opens in a new tab). Supplement to Jaggernath M, Naicker R, Madurai S, Brockman MA, Ndung'u T, Gelderblom HC (2016). PLoS One. doi:10.1371/journal.pone.0152184 (opens in a new tab); PMC4821624 (opens in a new tab) (title on the source page; it names the instruments). Changes: Paired results re-analysed by NakedSignal's program; instruments described generically; only derived results are shown. Public data, not a laboratory’s.

run 25 September 2026 · config 6eb23df45893 · run file sha256 5586cd892ef1

Show the numbers

Crossings here are counted across all of each comparison’s limits together. The decision limits are configured in our program for each analyte; they are common clinical lines, not a recommendation for any laboratory. Every series in the figure has its own page with its comparison line and limits (links in the table), and all of them sit side by side in the change explorer.

What this does not cover

  • Pre-analytical variation. Sampling, transport and storage add variation that a repeat on the same sample does not capture.
  • Biological variation over time. Whether a patient’s change between visits is real is the reference change value’s question, not this one.
  • Non-normal imprecision. The formula assumes a normal spread proportional to the result. Near the detection limit, or for counts and ratios with skewed error, that can be wrong.
  • Individual patients. Expected counts describe a population of results. They are not advice about any one patient.

Sources

  1. Loh TP, Markus C, Lim CY (2024). Impact of analytical imprecision and bias on patient classification. American Journal of Clinical Pathology. doi.org/10.1093/ajcp/aqad115
  2. Sandberg S, Fraser CG, Horvath AR, Jansen R, Jones G, Oosterhuis W, Petersen PH, Schimmel H, Sikaris K, Panteghini M (2015). Defining analytical performance specifications. Clinical Chemistry and Laboratory Medicine. doi.org/10.1515/cclm-2015-0067
  3. Fraser CG (2011). Reference change values. Clinical Chemistry and Laboratory Medicine. doi.org/10.1515/cclm.2011.733
  4. European Federation of Clinical Chemistry and Laboratory Medicine. EFLM Biological Variation Database. EFLM. biologicalvariation.eu/
  5. Coetzee LM, Glencross DK (2017). PLoS ONE. doi.org/10.1371/journal.pone.0187456
  6. Schmidt RL, Straseski JA, Raphael KL, Adams AH, Lehman CM (2015). A Risk Assessment of the Jaffe vs Enzymatic Method for Creatinine Measurement in an Outpatient Population. PLoS ONE. doi.org/10.1371/journal.pone.0143205
  7. Jaggernath M, Naicker R, Madurai S, Brockman MA, Ndung'u T, Gelderblom HC (2016). PLoS ONE. doi.org/10.1371/journal.pone.0152184
  8. Gelderblom H (2016). Performance characteristics of three point of care hemoglobin meters in Durban-2.xlsx. figshare data set, CC BY 4.0. doi.org/10.6084/m9.figshare.3119449

Each link was opened and checked on 26 September 2026. Where a paper’s title names an instrument, the reference gives authors, year, journal and DOI only.