- The paper presents B-CALM, a hierarchical Bayesian framework that uses latent encoders to align partially overlapping covariate spaces from RCTs and observational studies.
- It derives an effective sample size limit by controlling the external data’s influence via a comparative-bias prior, preventing over-precision.
- Empirical results on synthetic, semi-synthetic, and real-world datasets validate that B-CALM achieves near-nominal uncertainty quantification and robust bias sensitivity.
B-CALM: Bias-Limited Bayesian Borrowing for RCT-Anchored Treatment Effects under Covariate Mismatch
The "B-CALM: Bias-Limited Bayesian Borrowing for RCT-Anchored Treatment Effects under Covariate Mismatch" (2607.04036) addresses the statistical challenge of synthesizing information from randomized controlled trials (RCTs) and observational studies (OS) for treatment effect estimation, focusing on conditional average treatment effects (CATEs) when covariate sets are only partially overlapping. RCTs retain high internal validity but are often underpowered for accurate estimation of heterogeneity, whereas OSs offer larger, more diverse samples albeit with potential for substantial confounding and systematic biases due to lack of randomization and population differences.
Existing borrowing approaches—both frequentist calibrated models (e.g., R-OSCAR, MR-OSCAR, CALM) and Bayesian meta-analytic/commensurate priors—have thus far either focused on point estimation with ad hoc uncertainty quantification or have not provided explicit functional sensitivity to bias. Most do not address the high-dimensional, covariate-dependent treatment effect under mismatch in covariate measurement across sources.
B-CALM Model and Architecture
B-CALM (Bayesian Calibrated ALignment under covariate Mismatch) proposes a hierarchical Bayesian framework explicitly modeling the source-specific outcome surfaces anchored on a shared latent space, constructed via source-specific encoders operating on partially overlapping covariate blocks (Xr=(W,U) for RCT and Xo=(W,V) for OS, where W is shared).
Each encoder maps observed covariates into a common latent state Z; outcome models are then defined as:
- RCT: ηr(a,z)=μ(z)+aτ(z)
- OS: ηo(a,z)=μ(z)+aτ(z)+b0​(z)+abΔ​(z)
Here, b0​(z) accounts for baseline bias (control arm discrepancies) and bΔ​(z) for comparative bias (treatment effect discrepancies). These bias functions capture the deviation of OS outcome structure from the RCT-anchored estimand after adjusting for measurement and latent alignment.

Figure 1: B-CALM architecture and causal bookkeeping. Source-specific encoders qϕr​ and qϕo​ map partially overlapping covariate blocks into a shared latent state Xo=(W,V)0, supporting joint modeling of trial and OS outcome surfaces with bias decomposition.
Key distinguishing features of the model:
- The comparative-bias process Xo=(W,V)1 is assigned a prior whose variance, Xo=(W,V)2, serves as an explicit and interpretable "sensitivity knob"—modulating allowable OS-to-RCT information transfer. A tight prior enforces strong borrowing, while a diffuse prior limits transfer, recovering RCT-only inference.
- An effective-sample-size (ESS) formula, derived for both scalar and function-valued outcomes, demonstrates that the information contributed by OS to the RCT-defined contrast is upper bounded by the comparative bias prior’s precision, irrespective of OS sample size. Thus, information "saturates," precluding sample-size domination by large OS cohorts.
Statistical Theory and Bias-Limited Borrowing
The analytical core of B-CALM is the function-valued bias-limited information bound: in a linearized function model, posterior precision for the RCT contrast derived from the OS is limited by Xo=(W,V)3, the comparative-bias prior variance.
- Scalar case (ATE): The OS's ESS for the RCT ATE is
Xo=(W,V)4
increasing Xo=(W,V)5 cannot move beyond this ceiling.
- Function-valued case: The OS's Fisher information in any direction of the treatment effect surface is bounded by the prior precision on Xo=(W,V)6. Increasing Xo=(W,V)7 improves estimation up to the prior-imposed ceiling but does not permit arbitrary reduction in uncertainty.
This result codifies the notion that borrowing must trace an explicit, pre-specified region of plausible bias, with prior sensitivity analysis yielding meaningful assessments of decision robustness.
Experimental Results
Extensive empirical analyses validate the theoretical guarantees of B-CALM on synthetic, semi-synthetic (STAR-style), and real-world pediatric obesity datasets.
- Calibration and Negative Transfer: Across a grid of simulations varying covariate overlap and bias, B-CALM achieves near-nominal coverage (mean 0.939 for 90% intervals) with a low negative-transfer rate (0.023), outperforming alternatives whose intervals become over-confident and anti-conservative under bias.
- Sample Size Saturation: As OS sample size increases, interval widths for B-CALM contract only to the Xo=(W,V)8-imposed floor; pooled or OS-only models' intervals contract indefinitely, ignoring bias risk. Empirical effective sample sizes concur with theoretical predictions.

Figure 2: Comparative-bias sweep on the balanced DGP, showing CATE RMSE and empirical coverage for varying bias strengths. B-CALM uniquely maintains coverage near nominal regardless of increasing bias, contrasting with dramatic undercoverage for pooled and causal-forest approaches.

Figure 3: Interval width saturation as Xo=(W,V)9 increases (log scale). B-CALM intervals plateau at a W0-dependent floor, while pooled and OS-only methods' intervals shrink unchecked.
- Posterior Calibration: Posterior coverage across nominal levels is accurate for B-CALM, while non-bias-regulating approaches exhibit systematic undercoverage as bias increases.

Figure 4: Posterior calibration curves display superior empirical coverage for B-CALM versus alternatives under comparative bias.
- Real-World Application: In external control augmentation for a pediatric-obesity RCT, B-CALM delivered a moderate 9.4% reduction in the ATE interval width (maintaining conservative bias risk) compared to >28% for pooled/naive borrowing, but notably, the latter methods would have absorbed bias larger than their nominal width under plausible bias priors.


Figure 5: ATE interval width across baseline-bias prior W1 in the real-data application, illustrating borrowing sensitivity.

Figure 6: Baseline-bias recovery and prior calibration in the pediatric-obesity application. The credible interval is sensitive and appropriately expands as the prior permits more potential bias.
Practical and Theoretical Implications
B-CALM's explicit bias modeling framework yields several important implications:
- Principled Uncertainty Quantification: Uncertainty on CATE (or ATE) is guaranteed not to understate risk due to bias, even with large external control datasets. This is crucial for regulatory and clinical settings where false precision is particularly hazardous.
- Sensitivity Analysis as a Reporting Standard: B-CALM provides a rigorous mechanism for sensitivity analysis, with W2 serving as a transparent lever for users to explore conclusions' robustness to OS bias assumptions.
- Posterior Inference, Not Point Estimation: Unlike prior calibrated-borrowing approaches (R-OSCAR, MR-OSCAR, CALM), B-CALM provides full posterior inference, naturally propagating encoder and outcome model uncertainty.
- Function-valued Borrowing Adaptive to Covariate Mismatch: By utilizing latent alignment and explicitly modeling bias as a process on this space, B-CALM accommodates high-dimensional, partially-matched covariate sets, aligning with recent advances in representation learning for causal inference.
Theoretically, B-CALM tightens the connection between functional Bayesian borrowing and information-theoretic risk bounds, integrating PAC-Bayesian generalization theory, IPM-based alignment, and explicit calibration decompositions—a novel fusion in causal transfer learning.
Future Directions
Natural extensions include: time-varying treatments with temporally resolved bias models, decision-theoretic or asymmetric bias priors, and multi-source hierarchical borrowing frameworks. Technical challenges include scalable posterior inference (e.g., sparse GPs, amortized encoders) as applications move to EHR-scale data.
Conclusion
B-CALM (2607.04036) delivers a robust, theoretically principled framework for RCT-anchored treatment effect inference when augmenting with potentially heavily biased external data. By capping the OS's information contribution via explicit prior control on bias processes, B-CALM achieves empirically validated uncertainty quantification and robust negative-transfer protection, providing a model-driven sensitivity analysis protocol and a practical safeguard against sample-size-driven over-precision. The approach advances Bayesian causal transfer by making the trade-off between precision and bias sensitivity explicit and operational for regulators and clinical trialists.