NakedSignal

Guide

Reagent and calibrator lot changes: verifying with patient samples at decision limits

A new lot of reagent or calibrator can shift patient results by an amount quality control does not reveal. This note covers why, what the CLSI protocol for lot changes asks, where it leaves gaps, and how to read a small patient-sample comparison at the limits clinicians use.

By NakedSignal · Updated 26 September 2026

Why a lot change can move results

Every new lot of reagent or calibrator is made to the same specification, but not to the same result. Lot-to-lot variation limits a laboratory’s ability to produce consistent results over time, and it has well-documented clinical consequences [Thompson and Chesher, 2018]. Most lot changes are small. The ones that matter are those that move results near a decision limit, where a small shift changes what the clinician does.

Verification of each new lot is part of monitoring the long-term stability of a measurement procedure. It is limited by the resources it needs and by uncertainty over the design and statistics that suit an individual laboratory [Loh et al., 2023].

Why quality control may not show it

The obvious check is to run quality control or external quality assessment material on both lots. The problem is commutability: that material often does not behave like patient samples, so a difference seen on it may not be present in patient serum, and a real difference in patients may not show on it. Native patient samples are therefore preferred [Thompson and Chesher, 2018]. The same review cites one large study across several platforms and analytes in which control material and patient serum gave significantly different results in 40.9% of lot-change events.

What EP26 asks

The Clinical and Laboratory Standards Institute guideline EP26, now in its second edition (2022), sets out how a laboratory can use patient samples to detect clinically important changes from the current lot [CLSI EP26]. It works in two stages. Once per analyte, the laboratory fixes its acceptance thresholds and its tolerance for risk: a critical difference tied to allowable total error, and the statistical power it wants. Then, for each new lot, it runs a small number of patient samples on both lots. In the first edition, as reviewed by Thompson and Chesher, the mean difference at each target concentration is compared with a rejection limit set as a share of the critical difference [Thompson and Chesher, 2018].

The strength of this design is that it is decided in advance. The laboratory states what size of change matters before it looks at the data, and the number of samples follows from that and from the method’s imprecision.

Where it leaves gaps

  • Sample numbers. The required number of samples depends on imprecision and on the critical difference, and can be large. One laboratory that compared the first edition of the protocol with its own 20-sample procedure across six thyroid assays found the two agreed on 9 of 12 lot evaluations; the protocol needed more samples for 4 of the 6 analytes, and its rejection limits were hard to set [Katzman et al., 2017].
  • Power. Published lot-evaluation protocols differ in statistical rigour, and some may be underpowered to detect a clinically meaningful change [Thompson and Chesher, 2018].
  • Cumulative shifts. Each lot is compared with the one before. Current protocols, the CLSI one included, will not detect small shifts that accumulate in one direction over several lots [Thompson and Chesher, 2018].
  • Patients rather than means. A pass or fail on the mean difference says whether the lot is acceptable. It does not say how many patients near a limit would now be classified differently.

Patient samples at decision limits

A lot comparison is a method comparison on a small scale, and the same habits apply. Choose patient samples spread around each limit the laboratory reports, not only across the analytical range. Estimate the difference at each limit, not once for the whole range: a proportional shift that is small on average can be large at a high limit, and a constant shift matters most at a low one. Then place each pair in the categories the limits define and count how many changed category, at each limit and in each direction.

Loh and colleagues’ simulations make the case for looking at classification directly: the share of patients reclassified rises with both imprecision and bias, and at larger biases bias dominates [Loh et al., 2024]. A lot shift is a bias, so it is exactly the case where counting patients adds information the mean difference does not.

Small studies and the repeat-testing baseline

A lot comparison is usually small, and samples chosen near limits are, by design, the ones most likely to change category on a simple repeat. A handful of crossings in twenty samples can therefore be entirely ordinary. The way to tell is to compare the count with the number expected from repeat testing alone, computed from the method’s imprecision and each result’s distance from the limit. Our guide to crossings by chance sets out that calculation.

The count should also be graded by size, not hidden. In our own run on public method comparisons, 5 of the 34 comparisons were too small for a full re-test band and are shown as indicative. A lot study will usually fall in that range, and should be read the same way: a direction and a rough size, not a precise estimate.

Drift across several lots

Because each verification compares a new lot with the one before, a series of small shifts in the same direction can pass every check and still move results a long way. Moving averages of patient results over several lot changes can reveal that drift, though they need software, careful tuning, and a stable patient population [Thompson and Chesher, 2018]. Collaborative verification between laboratories and patient-based monitoring are likely to improve detection further [Loh et al., 2023]. Keeping every lot comparison in one record, with the difference at each limit, makes the running total visible.

The nearest public example

Public data

We have no lot-change results of our own. No public data set we hold is a reagent or calibrator lot change. The nearest case is a change of creatinine chemistry on one analyser, from an enzymatic to a kinetic Jaffe method, in an open-access deposit [Schmidt et al., 2015]. It is a method change, not a lot change, but it shows the question a lot verification should answer: did the change move more patients across the limits than repeat testing alone would?

Public data

creatinine: Central lab analyser, enzymatic creatinine to Same analyser, kinetic Jaffe creatinine

Crossed after the change: 23

Expected from repeat testing alone: about 41

Fig. 1. A change of creatinine method on one analyser: 23 results changed category against about 41 expected from repeat testing alone. The change moved fewer results than a simple repeat would.

Source Schmidt RL, Straseski JA, Raphael KL, Adams AH, Lehman CM (2015). A Risk Assessment of the Jaffe vs Enzymatic Method for Creatinine Measurement in an Outpatient Population. PLoS One. doi:10.1371/journal.pone.0143205 (opens in a new tab); PMC4657986 (opens in a new tab). Licence: CC BY 4.0 (opens in a new tab). Changes: Paired results re-analysed by NakedSignal's program; instruments described generically; only derived results are shown. Repeat-test imprecision assumed, not measured. Public data, not a laboratory’s.

run 25 September 2026 · config 6eb23df45893 · run file sha256 5586cd892ef1

Show the numbers
Pairs529
Decision limits1.2 and 1.5 mg/dL
Changed category23
Expected from repeat testing alone40.8

The limits at 1.2 and 1.5 mg/dL are configured in our program for creatinine; they are common clinical lines, not a guideline and not a recommendation for any laboratory. The series opens in the change explorer with its comparison line.

What we offer

Offer

We run the same count on a laboratory’s or a maker’s lot-change pairs: the difference at each decision limit with its interval, patients crossing each limit in each direction against repeat testing, and a running record across lots. This has not yet been run on any laboratory’s or maker’s lot-change data. See diagnostics makers and EQA schemes and laboratories and networks.

What this does not cover

  • Setting acceptance criteria. What size of change is acceptable for an analyte is the laboratory’s decision, from clinical need or biological variation. This note does not set it.
  • Sample-size tables. The number of samples EP26 requires for a given imprecision and critical difference is in the guideline itself, not reproduced here.
  • Regulatory release of lots. Manufacturers’ release criteria and regulatory requirements are outside this note.
  • Individual patients. Counts describe a population of results. They are not advice about any one patient.

Sources

  1. Clinical and Laboratory Standards Institute (2022). EP26: User Evaluation of Acceptability of a Reagent Lot Change. CLSI guideline, second edition. clsi.org/standards/products/method-evaluation/documents/ep26/
  2. Thompson S, Chesher D (2018). Lot-to-Lot Variation. The Clinical Biochemist Reviews. pmc.ncbi.nlm.nih.gov/articles/PMC6223607/
  3. Katzman BM, Ness KM, Algeciras-Schimnich A (2017). Evaluation of the CLSI EP26-A protocol for detection of reagent lot-to-lot differences. Clinical Biochemistry. doi.org/10.1016/j.clinbiochem.2017.03.012
  4. Loh TP, Markus C, Tan CH, Tran MTC, Sethi SK, Lim CY (2023). Lot-to-lot variation and verification. Clinical Chemistry and Laboratory Medicine. doi.org/10.1515/cclm-2022-1126
  5. Loh TP, Markus C, Lim CY (2024). Impact of analytical imprecision and bias on patient classification. American Journal of Clinical Pathology. doi.org/10.1093/ajcp/aqad115
  6. Schmidt RL, Straseski JA, Raphael KL, Adams AH, Lehman CM (2015). A Risk Assessment of the Jaffe vs Enzymatic Method for Creatinine Measurement in an Outpatient Population. PLoS ONE. doi.org/10.1371/journal.pone.0143205

Each link was opened and checked on 26 September 2026.