Papers
Topics
Authors
Recent
Search
2000 character limit reached

LOAM: Limits of Agreement with the Mean

Updated 9 July 2026
  • LOAM is a method for evaluating agreement by comparing individual measurements to the subject-specific mean rather than using pairwise contrasts.
  • It uses a two-way random-effects model with subject–observer interaction to separate variability into reproducibility and repeatability components.
  • The approach provides distinct limits for reproducibility and repeatability, aiding study planning and robust method comparison.

Searching arXiv for primary LOAM sources and related agreement-method papers. Limits of Agreement with the Mean (LOAM) are agreement limits for continuous measurements in which an individual observation is compared with a subject-specific mean across observers, rather than with a second method in a purely pairwise contrast. In the framework extended in “A statistical note on extending Christensen's limits of agreement with the mean” (Christensen et al., 22 Aug 2025), LOAM are defined for multi-observer studies with repeated measurements and are embedded in a two-way random-effects model that includes subject–observer interaction. This places LOAM within the broader limits-of-agreement tradition initiated by Bland–Altman methodology, while distinguishing it from standard pairwise limits of agreement for two methods and from other uses of the phrase “limits of agreement” in contingency-table agreement theory (Vallejos, 24 Apr 2026, Safak, 2020).

1. Conceptual basis and relation to classical limits of agreement

Classical Bland–Altman limits of agreement are formulated for two measurements XX and YY through the pairwise difference XYX-Y, summarized as

mean difference±1.96×SD of differences.\text{mean difference} \pm 1.96 \times \text{SD of differences}.

The corresponding graphical diagnostic plots the difference against the arithmetic average (X+Y)/2(X+Y)/2. In the review “Agreement coefficients for continuous variables: A review” (Vallejos, 24 Apr 2026), this construction is presented as a foundational device for assessing numerical similarity rather than mere linear association.

LOAM changes the target of comparison. Instead of asking how far one observer is from another observer, it asks how far one measurement is from the subject-specific mean over an observer panel. In the repeated-measures setting described in (Christensen et al., 22 Aug 2025), the core deviations are

YijkYˉiandYijkYˉij,Y_{ijk}-\bar Y_{i\cdot\cdot} \quad\text{and}\quad Y_{ijk}-\bar Y_{ij\cdot},

where YijkY_{ijk} is the kkth measurement on subject ii by observer jj, YY0 is the mean for subject YY1 across observers and repetitions, and YY2 is the mean for a fixed subject–observer pair across repetitions. This yields a single summary of agreement relative to a subject-level consensus, rather than a family of pairwise observer-specific limits (Christensen et al., 22 Aug 2025).

A central motivation is the multi-observer setting. With more than two observers, exhaustive pairwise Bland–Altman comparisons produce many limits but no single measure of how a typical measurement behaves relative to the subject’s overall average. LOAM is designed for that situation. A plausible implication is that LOAM is most natural when the operational reference is not an external gold standard but a panel-based subject mean.

2. Random-effects formulation and the role of interaction

The 2025 extension formulates LOAM under a balanced two-way random-effects model with repeated measurements and subject–observer interaction: YY3 with

YY4

Here YY5 is the overall mean, YY6 is the subject effect, YY7 is the observer effect, YY8 is the subject–observer interaction, and YY9 is the residual measurement error (Christensen et al., 22 Aug 2025).

The interaction term is the principal extension beyond the earlier Christensen framework described in (Christensen et al., 22 Aug 2025). Its function is to separate two distinct sources of disagreement. The component XYX-Y0 captures systematic variation in how particular observers measure particular subjects, whereas XYX-Y1 captures pure replicate-to-replicate error around the subject–observer mean. Without repeated measurements, these two sources cannot be disentangled. Repetition is therefore not an incidental design feature but a structural requirement of the extended LOAM decomposition.

The ANOVA decomposition used for estimation is based on the expected mean squares

XYX-Y2

XYX-Y3

XYX-Y4

XYX-Y5

These identities underwrite both the variance-component estimators and the subsequent LOAM estimators (Christensen et al., 22 Aug 2025).

3. Reproducibility LOAM and repeatability LOAM

The interaction-extended framework defines two distinct LOAM quantities, one for reproducibility and one for repeatability (Christensen et al., 22 Aug 2025).

Quantity Difference target LOAM
Reproducibility LOAM XYX-Y6 XYX-Y7
Repeatability LOAM XYX-Y8 XYX-Y9

The reproducibility LOAM is defined as

mean difference±1.96×SD of differences.\text{mean difference} \pm 1.96 \times \text{SD of differences}.0

Its explicit variance decomposition is

mean difference±1.96×SD of differences.\text{mean difference} \pm 1.96 \times \text{SD of differences}.1

This quantity measures how far a single measurement on subject mean difference±1.96×SD of differences.\text{mean difference} \pm 1.96 \times \text{SD of differences}.2 may deviate from the subject’s average over all observers and repetitions. It includes observer-to-observer variation, subject–observer interaction, and residual error, but it excludes the subject variance mean difference±1.96×SD of differences.\text{mean difference} \pm 1.96 \times \text{SD of differences}.3, because the comparison is centered on the subject-specific mean and the subject effect cancels (Christensen et al., 22 Aug 2025).

The repeatability LOAM is defined as

mean difference±1.96×SD of differences.\text{mean difference} \pm 1.96 \times \text{SD of differences}.4

This isolates within-observer, within-subject replicate variability. Once conditioning is on the subject–observer mean mean difference±1.96×SD of differences.\text{mean difference} \pm 1.96 \times \text{SD of differences}.5, the subject effect, observer effect, and subject–observer interaction all cancel, leaving only residual measurement error (Christensen et al., 22 Aug 2025).

The distinction between the two limits is substantive rather than merely notational. Small repeatability LOAM with large reproducibility LOAM indicates that observers are individually stable but differ from one another, or differ systematically on particular subjects. Large repeatability LOAM indicates noisy measurement even when subject and observer are fixed. This suggests that LOAM is not a single scalar notion of agreement, but a family of mean-referenced limits whose interpretation depends on which averaging structure defines the comparator.

4. Estimation, confidence intervals, and study planning

The extended LOAM paper uses ANOVA or method-of-moments estimators obtained from the expected mean squares: mean difference±1.96×SD of differences.\text{mean difference} \pm 1.96 \times \text{SD of differences}.6

mean difference±1.96×SD of differences.\text{mean difference} \pm 1.96 \times \text{SD of differences}.7

Substituting these estimates into the variance formulas yields the estimated reproducibility and repeatability LOAMs (Christensen et al., 22 Aug 2025).

Inference differs sharply between the two quantities. Because repeatability LOAM depends only on mean difference±1.96×SD of differences.\text{mean difference} \pm 1.96 \times \text{SD of differences}.8, it admits an exact chi-square-based interval. If mean difference±1.96×SD of differences.\text{mean difference} \pm 1.96 \times \text{SD of differences}.9, then a 95% confidence interval for the upper repeatability limit is

(X+Y)/2(X+Y)/20

with the lower limit obtained symmetrically by negation (Christensen et al., 22 Aug 2025).

Reproducibility LOAM does not have an exact closed-form distribution in the ANOVA formulation. The proposed interval is therefore approximate and is based on the Graybill–Wang method for confidence intervals on nonnegative linear combinations of variances. For the upper reproducibility limit, the approximate 95% interval is

(X+Y)/2(X+Y)/21

with (X+Y)/2(X+Y)/22 and (X+Y)/2(X+Y)/23 defined from (X+Y)/2(X+Y)/24-quantiles and the component sums of squares (Christensen et al., 22 Aug 2025).

Study planning is precision-based rather than power-based. The paper defines a confidence-interval width (X+Y)/2(X+Y)/25 for reproducibility LOAM and recommends using pilot estimates (X+Y)/2(X+Y)/26, (X+Y)/2(X+Y)/27, and (X+Y)/2(X+Y)/28 to solve numerically for the required number of observers (X+Y)/2(X+Y)/29, given fixed YijkYˉiandYijkYˉij,Y_{ijk}-\bar Y_{i\cdot\cdot} \quad\text{and}\quad Y_{ijk}-\bar Y_{ij\cdot},0 and YijkYˉiandYijkYˉij,Y_{ijk}-\bar Y_{i\cdot\cdot} \quad\text{and}\quad Y_{ijk}-\bar Y_{ij\cdot},1. For comparing LOAMs between two methods, such as CT and MRI, the proposal is a subject-level bootstrap in which resampling is performed at the subject level and each resampled subject carries all associated observers, repetitions, and methods. Non-overlapping confidence intervals are said to indicate a statistically significant difference, whereas overlapping intervals do not rule one out (Christensen et al., 22 Aug 2025).

5. Relation to Bland–Altman graphics and the problem of unequal precision

Although LOAM is a mean-referenced agreement framework for multiple observers, its interpretation remains closely tied to the logic of Bland–Altman analysis. In standard Bland–Altman work, agreement is inspected by plotting

YijkYˉiandYijkYˉij,Y_{ijk}-\bar Y_{i\cdot\cdot} \quad\text{and}\quad Y_{ijk}-\bar Y_{ij\cdot},2

Borga shows that this construction assumes that the two methods have approximately equal within-subject variance; otherwise YijkYˉiandYijkYˉij,Y_{ijk}-\bar Y_{i\cdot\cdot} \quad\text{and}\quad Y_{ijk}-\bar Y_{ij\cdot},3 and YijkYˉiandYijkYˉij,Y_{ijk}-\bar Y_{i\cdot\cdot} \quad\text{and}\quad Y_{ijk}-\bar Y_{ij\cdot},4 become mechanically correlated even when there is no true dependence of difference on measurement magnitude (Borga, 2022).

The proposed correction is to keep the difference on the vertical axis but replace the arithmetic mean by an inverse-variance weighted average: YijkYˉiandYijkYˉij,Y_{ijk}-\bar Y_{i\cdot\cdot} \quad\text{and}\quad Y_{ijk}-\bar Y_{ij\cdot},5 where YijkYˉiandYijkYˉij,Y_{ijk}-\bar Y_{i\cdot\cdot} \quad\text{and}\quad Y_{ijk}-\bar Y_{ij\cdot},6 and YijkYˉiandYijkYˉij,Y_{ijk}-\bar Y_{i\cdot\cdot} \quad\text{and}\quad Y_{ijk}-\bar Y_{ij\cdot},7 are within-subject sample variances. Under the model

YijkYˉiandYijkYˉij,Y_{ijk}-\bar Y_{i\cdot\cdot} \quad\text{and}\quad Y_{ijk}-\bar Y_{ij\cdot},8

with independent errors and no true magnitude-dependent bias, the covariance between the difference and a general weighted average reduces to

YijkYˉiandYijkYˉij,Y_{ijk}-\bar Y_{i\cdot\cdot} \quad\text{and}\quad Y_{ijk}-\bar Y_{ij\cdot},9

For the ordinary Bland–Altman mean, this becomes

YijkY_{ijk}0

whereas the inverse-variance choice YijkY_{ijk}1, YijkY_{ijk}2 gives zero covariance (Borga, 2022).

For LOAM, the immediate significance is methodological rather than definitional. If magnitude-dependent disagreement is assessed by importing Bland–Altman-type diagnostics into a LOAM workflow, unequal precision can create false trend or conceal real trend. Borga’s blood-pressure reanalysis and simulations show both failures: false positive trend when precisions differ but no true trend exists, and failure to detect a real trend when unequal precision masks it (Borga, 2022). This suggests that any LOAM-adjacent graphical analysis of bias versus measurement size should treat precision asymmetry explicitly.

6. Scope, limitations, and neighboring frameworks

The current LOAM extension is derived for a fully crossed, balanced design in which every subject is measured by every observer and each subject–observer pair has exactly YijkY_{ijk}3 repeated measurements. The model assumes Gaussian random effects and error, and the repeated-measures structure is essential because without YijkY_{ijk}4 the interaction variance YijkY_{ijk}5 cannot be separated from residual error YijkY_{ijk}6. Most confidence intervals are approximate rather than exact, and the ANOVA estimators

YijkY_{ijk}7

can be negative in finite samples. The bootstrap test for comparing LOAMs is proposed operationally rather than as a fully formalized theorem-based procedure (Christensen et al., 22 Aug 2025).

LOAM is also not interchangeable with every other “limits of agreement” construction in the literature. The review “Agreement coefficients for continuous variables: A review” surveys classical Bland–Altman limits, concordance correlation coefficients, probability of agreement, and repeated-measures or multivariate extensions, but it does not explicitly discuss LOAM (Vallejos, 24 Apr 2026). The MRMC paper “Three-way Mixed Effect ANOVA to Estimate MRMC Limits of Agreement” develops variance-component machinery for pairwise limits of agreement in multi-reader multi-case studies under a three-way mixed-effect ANOVA, but it does not derive limits against a subject-specific mean; its value for LOAM is therefore preparatory rather than direct (Wen et al., 2021). By contrast, “Min-Mid-Max Scaling, Limits of Agreement, and Agreement Score” uses “limits of agreement” to mean feasible lower and upper bounds on agreement in contingency tables under fixed marginals, with YijkY_{ijk}8, YijkY_{ijk}9, and the corresponding feasible range of Cohen’s kappa; that usage is conceptually adjacent but not a LOAM method for continuous paired measurements (Safak, 2020).

Within contemporary agreement analysis, LOAM therefore occupies a specific niche: agreement of continuous measurements made by multiple observers, referenced to the subject-specific mean, with explicit separation of reproducibility and repeatability under a random-effects model with interaction. Its principal strength is that it replaces a proliferation of pairwise observer comparisons with a mean-centered formulation whose variance decomposition remains interpretable. Its principal constraints are the need for balanced repeated-measures data, the reliance on variance-component assumptions, and the fact that graphical and inferential extensions beyond the core formulas remain an active area of methodological development.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Limits of Agreement with the Mean (LOAM).