NakedSignal

Guide

Least squares, Deming or Passing-Bablok: choosing and checking a method comparison regression

Every analyser change ends in a scatter plot and a line. Which line you fit, and whether the data can support any line at all, decides the bias you report at each decision limit. This guide sets out what each regression assumes, the checks that matter more than the choice, and how to read slope and intercept where clinicians decide.

What the line is for

A method comparison runs the same patient samples on the current and the new measurement procedure and estimates the bias between them [CLSI EP09]. A straight line of new on old splits that bias into two parts: a constant difference (the intercept) and a difference that grows with concentration (the slope). From the line you can read the expected difference at any concentration, which is what a laboratory needs at its decision limits.

Which line, though? The three in common use make different assumptions about measurement error. When the assumptions fail, the slope and intercept are wrong, and so is every bias read from them. The guide to bias at clinical decision limits takes a fitted line and counts the patients who cross each limit; this guide is about getting the line right first.

Ordinary least squares and the error in x

Ordinary least squares assumes the x values, here the current method’s results, are measured without error. In a method comparison they never are. The error in x pulls the slope towards zero and pushes the intercept the other way, so least squares reports a proportional difference that is not there. One early study found the error becomes significant when the standard deviation of a single x measurement is more than 0.2 of the standard deviation of the x values across the data set [Cornbleet and Gochman, 1979]. In other words, least squares is safe only when the range of results is wide compared with the method’s imprecision.

A simulation study later found that least-squares point estimates of bias were reliable when the correlation coefficient exceeded 0.975, but that its confidence intervals for bias were unreliable in every case studied [Martin, 2000]. So a high correlation can rescue the estimate, but not the interval.

Deming regression: say how noisy each method is

Deming regression allows error in both methods. It needs one extra input: the ratio of the two methods’ error variances. Many analyses set it to one by default. With a short range of values relative to the measurement error, a wrong ratio can produce a bias of up to two thirds of the least-squares bias, and misleading standard errors; the ratio should come from duplicate measurements or from quality-control data. Even with a misspecified ratio, Deming still did better than least squares [Linnet, 1998].

Simple Deming assumes each method’s standard deviation is constant across the range. Most methods have an error closer to a constant coefficient of variation, so the standard deviation grows with concentration. There, simple Deming’s intervals for bias become unreliable, and only an iteratively reweighted form produced unbiased estimates with reliable intervals in every case simulated [Martin, 2000]. Weighting also saves samples: at a range ratio of 10, a weighted approach cut the number needed by more than 50% [Linnet, 1999].

Passing-Bablok regression: robust, with conditions

Passing-Bablok regression makes no assumption about the distribution of the samples or of the errors, and gives the same answer whichever method is put on the x axis [Passing and Bablok, 1983]. It is robust to outliers, which is why laboratories reach for it. It does have conditions: continuously distributed data and a linear relationship between the two methods. A cumulative sum (cusum) test checks linearity, a residual plot shows outliers and curvature, and data that are not linear are not suitable for concluding that methods agree [Bilić-Zulle, 2011]. The original procedure tests linearity before it gives intervals for slope and intercept [Passing and Bablok, 1983].

Checks that matter more than the model

Comparing all of these models on real data, one group concluded that the quality of the analytical input data matters more than the choice of model [Stöckl et al., 1998]. When a fit looks poor, investigate the data before reaching for a different regression. The checks below start from their list:

  • Range. Spread the samples across the whole range the laboratory reports, with extra samples near each decision limit. The range ratio (highest over lowest result) drives power more than any other factor. To detect one standardised slope deviation at a range ratio of 2 took 544 samples; at a range ratio of 10, 64. The conventional 40 to 100 samples often need to be reconsidered, and very narrow ranges such as electrolytes need very large studies [Linnet, 1999].
  • Scatter against imprecision. Compare the standard deviation of the residuals with the two methods’ combined analytical imprecision. If the scatter is much larger, something sample-related, such as an interference or a matrix effect, is at work, and no regression will remove it [Stöckl et al., 1998].
  • Linearity. Look at the residual plot. A curve means one straight line cannot describe the relationship, and the bias at each limit should be estimated locally instead.
  • Outliers. Investigate them, do not just drop them. One rule removes points further than 4 times the standard error of the estimate from the line [Cornbleet and Gochman, 1979]; whatever rule you use, fix it before the data are seen and report what was excluded.
  • Correlation is not agreement. A high correlation coefficient says the range is wide relative to the scatter, not that the methods agree [Bland and Altman, 1986]. It is a screening check on whether a least-squares fit is usable [Stöckl et al., 1998], not a verdict on the change.

Reading the line at a decision limit

“Slope interval includes one, intercept interval includes zero” is a statistical statement about the whole line. It is not the same as “the bias at our decision limits is acceptable”. A small intercept matters a lot at a low limit; a slope a few per cent from one matters most at a high one. Read the line at each limit.

A published example shows why. In one hospital laboratory’s analyser change, the Deming line for serum creatinine was new = 0.944 × old + 0.165 mg/dL, from 44 patient samples averaging about 5 mg/dL [Bush et al., 2020]. Reading that line at three concentrations:

Expected new result and difference from the published line at three creatinine concentrations
Current method, mg/dLExpected on new method, mg/dLDifference, mg/dLDifference, %
0.800.920.12+15.0%
1.201.300.10+8.1%
5.004.88−0.12−2.3%

The same line gives a positive bias at low creatinine and a negative one at high creatinine; it crosses the line of identity at about 2.9 mg/dL. The average bias over the study says little about the concentrations where most eGFR decisions are made. And because the samples averaged about 5 mg/dL and included dialysis patients, the first question to ask before using the line at 0.8 mg/dL is whether the study had enough samples there to support it. A line is only as good as the samples near the point where you read it.

The bias at each limit should come with an interval, from the regression’s own procedure or from bootstrap resampling of the pairs, and the guide to eGFR after a creatinine bias shows what a bias of this size does to eGFR categories.

Choosing, in short

  • Wide range, high correlation, point estimate only: least squares gives a usable slope, but not a usable interval.
  • Error in both methods, known error ratio: Deming, weighted when imprecision is closer to a constant CV than a constant SD.
  • Outliers or unknown error structure, linear relationship: Passing-Bablok, with the linearity test and residual plot reported.
  • In every case: check range, scatter against imprecision, linearity and outliers first, then read the bias at each decision limit, not once for the whole range.

Run it on your own pairs

The change check fits a Passing-Bablok line of new on old to your paired results in the browser, gives Bland-Altman agreement, reads the bias at each decision line you set, and counts the results that crossed each line against those repeat testing alone would move. Your file is not uploaded.

To see what a published comparison line implies for one result, the free analyser comparability lookup converts a value from one analyser to another using open-access method comparisons, each with its source, and flags a value outside the range a study measured where the study reports that range.

Sources

  1. Clinical and Laboratory Standards Institute (2018). EP09: Measurement Procedure Comparison and Bias Estimation Using Patient Samples. CLSI guideline, third edition. clsi.org/standards/products/method-evaluation/documents/ep09/
  2. Cornbleet PJ, Gochman N (1979). Incorrect least-squares regression coefficients in method-comparison analysis. Clinical Chemistry. doi.org/10.1093/clinchem/25.3.432
  3. Martin RF (2000). General Deming regression for estimating systematic bias and its confidence interval in method-comparison studies. Clinical Chemistry. doi.org/10.1093/clinchem/46.1.100
  4. Linnet K (1998). Performance of Deming regression analysis in case of misspecified analytical error ratio in method comparison studies. Clinical Chemistry. doi.org/10.1093/clinchem/44.5.1024
  5. Linnet K (1999). Necessary sample size for method comparison studies based on regression analysis. Clinical Chemistry. doi.org/10.1093/clinchem/45.6.882
  6. Passing H, Bablok W (1983). A new biometrical procedure for testing the equality of measurements from two different analytical methods. Part I. Journal of Clinical Chemistry and Clinical Biochemistry. doi.org/10.1515/cclm.1983.21.11.709
  7. Bilić-Zulle L (2011). Comparison of methods: Passing and Bablok regression. Biochemia Medica. doi.org/10.11613/BM.2011.010
  8. Stöckl D, Dewitte K, Thienpont LM (1998). Validity of linear regression in method comparison studies: is it limited by the statistical model or the quality of the analytical input data?. Clinical Chemistry. doi.org/10.1093/clinchem/44.11.2340
  9. Bland JM, Altman DG (1986). Statistical methods for assessing agreement between two methods of clinical measurement. The Lancet. doi.org/10.1016/s0140-6736%2886%2990837-8
  10. Bush V, Smola C, Schmitt P (2020). Evaluation of the [maker and model] chemistry analyzer. Practical Laboratory Medicine. pmc.ncbi.nlm.nih.gov/articles/PMC6909053/