BAR in Repeated Binary Measurements
- Binary Agreement Rate (BAR) is a measure of concordance between two binary outcomes, reflecting the probability that both methods classify subjects identically.
- A probit generalized linear mixed model (GLMM) adjusts for subject heterogeneity, temporal dependence, and rater variability to provide a robust assessment of method agreement.
- Model-based summaries like Bland–Altman plots and adjusted kappa statistics offer more reliable insights than raw percent agreement in complex longitudinal settings.
Searching arXiv for the cited paper and possible acronym ambiguity. Searching arXiv for (Wang et al., 2018). Binary Agreement Rate (BAR), in the sense of agreement between two binary measurement methods, is not explicitly defined or named in “Assessing method agreement for paired repeated binary measurements administered by multiple raters” (Wang et al., 2018). In that setting, the closest quantity to a BAR is the observed agreement proportion obtained after reducing repeated binary measurements to one model-based predicted binary outcome per subject per method. The paper’s central position is that a naive raw concordance proportion can be misleading when binary outcomes are repeatedly observed over time and depend on the rater as well as the method. It therefore replaces direct repeated-observation agreement counting with a probit generalized linear mixed model (GLMM), assesses interchangeability through equality of method fixed effects, and supplements that test with latent-scale and probability-scale agreement summaries.
1. Concept and scope
The paper studies agreement between two binary measurement methods under a design with repeated observations and multiple raters. It distinguishes three related objectives: agreement between two methods, agreement among raters, and the decision whether the two methods can be used interchangeably. Interchangeable use is tied to whether, after accounting for subject heterogeneity, longitudinal change over time, rater variability, and repeated-measure dependence, “it does not matter which method is being used to take the measurement as both give practically the same value.”
If BAR is understood as a simple binary concordance proportion, the closest quantities in the paper are the observed agreement in the table used for Cohen’s kappa, the model-based probability-scale summaries , and the model-based predicted binary classifications . The raw agreement proportion is
where are the cell counts in the subject-level contingency table formed from .
A simple BAR is usually the probability that two binary measurements agree,
The paper does not derive this probability analytically under its GLMM. Instead, it argues that direct percent agreement or a naive repeated-data kappa may be distorted because repeated measurements induce within-subject dependence, multiple raters introduce an additional source of variability, prevalence imbalance affects kappa, and time trends may shift response probabilities. In that sense, the paper treats raw BAR-like quantities as secondary summaries rather than primary inferential targets.
2. Statistical formulation for paired repeated binary measurements
The design indexes subjects by , raters by , time points for subject 0 by 1, and methods by 2. At each time point, each subject is assessed by both methods, and each method is administered by a different rater. The paired observations are written as 3, with each 4. The paper emphasizes that the data are often unbalanced because not all subjects are observed at all time points.
The proposed observation model is a probit GLMM:
5
with linear predictor
6
Here 7 is the fixed effect of Method 8, 9 is a regression function of a time-dependent covariate 0, 1 is a subject random effect, and 2 is a rater random effect within method 3. The random effects satisfy
4
and the rater variances are allowed to differ by method.
The latent-variable representation is
5
with thresholding rule
6
This latent Gaussian construction is the basis for adapting Bland–Altman analysis and intraclass correlation methods to binary repeated data.
Temporal dependence is introduced through correlated latent errors within method:
7
An explicit example is AR(1),
8
The resulting model assumes independence across subjects, independence across methods at the latent error level, within-method temporal correlation over repeated measurements, normality of latent random components, and a probit link. No method-by-subject interaction or method-by-rater interaction beyond method-specific rater variance is included.
3. Primary agreement criterion: equality of method fixed effects
The paper’s primary criterion for method agreement is not a raw agreement rate but equality of the fixed method effects. On the latent scale, the method difference for subject 9 at time 0 is
1
with marginal mean
2
This yields the central hypothesis test
3
The interpretation is explicit. If 4 is rejected, there is a significant systematic method difference and the methods do not agree. If 5 is not rejected, there is no evidence of a systematic difference, and the methods may be considered to agree; the extent of agreement is then examined further using limits of agreement, Bland–Altman-style displays, and kappa.
Inference is performed within the fitted GLMM, effectively as a Wald-type fixed-effect comparison. The reported output includes the estimate of 6, its standard error, a 7-value, and a 8 confidence interval. This criterion operationalizes interchangeable use at the population level. A plausible implication is that, in this framework, a BAR-like concordance proportion is descriptive, whereas the equality test on 9 and 0 is decisive for interchangeability.
4. Model-based agreement summaries closest to BAR
After the primary method-effect test, the paper develops several agreement summaries that are more closely aligned with what BAR often informally denotes.
For each subject 1 and method 2, the target subject-level latent summary after removing rater effect is
3
Its posterior distribution is given as
4
with posterior mean
5
In practice, the paper replaces 6 by the EBLUP 7. These subject-level latent summaries support a generalized Bland–Altman construction with one point per subject:
8
Theorem 1 states that under 9 and 0, 1 is uncorrelated with 2, and 3. This reproduces the classical Bland–Altman intuition that differences should scatter around zero without systematic dependence on the average.
The same summaries can be displayed on the probability scale as 4, or, if skewed, on the log-transformed probability scale:
5
Predicted binary subject-level outcomes are then defined by thresholding the latent summaries,
6
These 7 produce the 8 table from which observed agreement 9 and Cohen’s kappa are computed:
0
with
1
and
2
Within this construction, 3 is the clearest analog of a BAR. However, it is not formed from all repeated observed pairs directly; it is formed after dependency-aware model reduction to one predicted binary outcome per subject per method.
The paper also quantifies inter-rater reliability through a latent-scale ICC for each method,
4
This is not a direct method-agreement metric, but it measures how strongly rater variability can affect agreement assessment. As 5 increases, 6 decreases.
5. Estimation, simulations, and empirical application
The GLMM is fitted using PROC GLIMMIX in SAS 9.4, with a binary response, probit link, fixed effects for method and time, a random subject effect, a random rater effect grouped by method, and AR(1) within-subject correlation (Wang et al., 2018). The theoretical arguments rely on 7, but the paper notes that the method may still perform well with few raters if inter-rater agreement is reasonably high, that is, if 8 is small.
The main simulation uses 9 subjects, 0 raters, and 1 time points for all subjects. Two models are studied. In Model 1 the methods agree, with 2. In Model 2 the methods disagree, with 3 and 4. Other parameters are
5
and
6
The estimated fixed-effect difference 7 was 8 with 9 in the agreement model and 0 with 1 in the disagreement model, showing that the test distinguished agreement from disagreement. Estimated variance components were close to the true values. On latent and probability scales, the Bland–Altman diagrams showed random scatter around zero, about 95 of 100 points within the 2 limits of agreement, and mean differences close to the truth. The true ICCs were approximately 3 for Method 1 and 4 for Method 2, and the estimated ICCs were close.
The kappa results provide the clearest demonstration of why a naive BAR-like summary may fail in repeated binary data. When 5, the model-based kappa was 6 with 7 CI 8. When 9, the model-based kappa was 0 with 1 CI 2. By contrast, the naive kappa applied directly to the repeated observed data was 3 in the agreement case and 4 in the disagreement case. The paper attributes this near-invariance to correlated repeated measurements keeping the disagreement counts 5 similar.
The real-data application compares CAM and 3D-CAM for postoperative delirium. The dataset contains 42 paired readings from 20 patients, assessed by 6 raters, with up to 6 time points and unbalanced longitudinal structure. The reported fixed-effect difference was
6
with 7 and 8 CI 9, so there was no evidence of a method difference and the methods may be used interchangeably. The ICCs were 00 for CAM and 01 for 3D-CAM. On the latent scale the mean difference was 02. On the log-transformed probability scale the mean difference was 03, interpreted as indicating that the probability of a positive score by 3D-CAM was on average 04 times that of CAM, based on
05
The model-based kappa was reported as 06, indicating perfect model-based agreement over the 20 patients.
6. Interpretation, limitations, and acronym ambiguity
For a simple BAR or percent agreement to be adequate, the paper identifies a narrow setting: one binary measurement per subject per method, no repeated measures, no raters or raters acting effectively as recorders, negligible rater variability, and no severe prevalence distortion. The model-based framework is preferable when subjects are measured repeatedly over time, multiple raters administer the methods, raters influence outcomes, data are unbalanced or partially missing, time trend matters, or a raw agreement rate would confound dependence with concordance.
In that formulation, agreement should not mean merely the fraction of observed pairs that match. It should reflect whether, after adjusting for subject heterogeneity, time effects, repeated-measure correlation, and rater variability, the two methods have the same underlying tendency to classify subjects. This is why the primary criterion is
07
The strengths of the model-based approach relative to raw BAR are that it adjusts for repeated-measure dependence, accounts for rater heterogeneity, allows subject-level and population-level summaries, supports graphical display via a latent Bland–Altman plot, avoids misleading prevalence effects from naive repeated kappa, and can handle unbalanced longitudinal data. Its limitations are that it is more complex to fit and explain, relies on model assumptions including a probit link, Gaussian random effects, and a specified correlation structure, leans theoretically on a large number of raters 08, and may be less immediately interpretable than a raw percentage.
The acronym “BAR” is also overloaded. In distributed-systems work such as “Asynchrony and Collusion in the N-party BAR Transfer Problem,” BAR refers not to binary agreement rate but to the taxonomy Byzantine, Altruistic, and Rational, and the paper studies the N-party BAR Transfer problem in an asynchronous authenticated network (Vilaça et al., 2012). A plausible implication is that acronym-based literature searches for “BAR” can retrieve technically unrelated material unless the measurement-agreement context is specified.