Integrative learning of individualized treatment rules from multiple studies with partially overlapping treatments
Published 12 Apr 2026 in stat.ME and stat.AP | (2604.10712v1)
Abstract: An individualized treatment rule (ITR) tailors treatments to a patient's specific characteristics. However, randomized controlled trials (RCTs) are often underpowered to detect the treatment effect heterogeneity needed for reliable ITR estimation. To address this limitation, there is growing interest in leveraging information from multiple studies to improve statistical power and support individualized decision-making. A key challenge in this context is that available RCTs may not evaluate the same set of treatments. In this paper, we propose an integrative learning framework that synthesizes evidence across multiple RCTs that share a common comparator but differ in their alternative treatment arms. Our method integrates information through a regularized weighted misclassification risk function and adaptively determines the contribution of each study to the ITRs of the others. We rigorously study the excess risk of the resulting estimator. Simulation studies demonstrate that the proposed approaches improve the estimation of both value functions and benefit functions. We illustrate the utility of our methodology using data from two landmark studies of major depressive disorder: the Establishing Moderators and Biosignatures of Antidepressant Response in Clinical Care (EMBARC) study and the International Study to Predict Optimized Treatment in Depression (iSPOT-D) study, both of which include a selective serotonin reuptake inhibitor as a common treatment arm. We find that the separate learning method outperforms one-size-fits-all methods, and our integrative methods further improve performance.
The paper presents two integrative algorithms, IntLS and IntLF, that jointly estimate individualized treatment rules from multiple RCTs with partially overlapping treatments.
Simulation studies demonstrate that both methods achieve lower RMSE for value and benefit function estimates, particularly when studies are closely aligned.
Application to major depressive disorder trials highlights the framework’s ability to enhance patient-specific treatment selection and clinical relevance.
Integrative Learning of Individualized Treatment Rules from Multiple Studies with Partially Overlapping Treatments
Introduction and Motivation
This paper addresses the statistical and algorithmic challenges of learning optimal individualized treatment rules (ITRs) in the context of multi-study randomized controlled trial data with partially overlapping treatments, focusing specifically on patient populations with major depressive disorder. The standard strategy of developing ITRs within a single RCT is underpowered for detecting treatment effect heterogeneity at the individual level, particularly when trial arms vary across studies and direct comparisons among all interventions of interest are absent. Classical meta-analytic and integrative data analysis strategies often assume full overlap among treatment arms or yield limited interpretability, thus motivating the development of new integrative frameworks.
Methodological Innovation
The central contribution is an integrative algorithmic framework that enables joint estimation of study-specific ITRs across RCTs with a shared common comparator but differing alternative treatments. Let Study 1 compare treatments A versus B, and Study 2 compare treatments C versus B, with B as the common arm. The authors introduce two approaches:
IntLS: Integrative Learning using Separate rules. Each study-specific ITR is regularized toward the ITR learned independently from the other study, via a weighted misclassification risk objective incorporating a Laplacian-inspired penalty that enforces alignment of treatment recommendations for similar patient profiles.
IntLF: Integrative Learning using Full data. This method incorporates all individual-level data simultaneously, introducing cross-study regularization directly into the optimization, enforcing consistency of recommendations across studies and samples as a function of shared covariate space coverage.
Both approaches are implemented using a doubly robust outcome-weighted learning framework, leveraging both propensity-score-based inverse probability weighting and mean-outcome regression adjustment, thereby enhancing robustness to model mis-specification.
The unifying penalized risk function for IntLF, optimized using surrogate loss minimization (specifically, the Huberized hinge loss), yields study-specific decision functions fj​(⋅) that adaptively borrow information across studies according to the empirical utility of such transfer, tuned by cross-validation. When the study context is highly heterogeneous, cross-study penalties κj​ are adaptively shrunk, reducing potential bias from inappropriate pooling.
Theoretical Guarantees
The authors provide a detailed excess risk analysis for the proposed estimators. The primary result demonstrates that the asymptotic excess risk of the integrative estimator f​jIntLF​ is governed by the minimum of (i) the standard RKHS estimation-approximation-sample tradeoff, modified by the magnitude of the regularization and cross-study fusion parameters, and (ii) the external informativeness of the other study's ITR. Specifically, if the external rule is highly informative (fj′​ approximates fj∗​), the method can achieve excess risk of order B0, improving upon the rates available to single-study learning (2604.10712).
Simulation Studies
The proposed integrative strategies are validated through extensive simulations under scenarios with both linear and nonlinear treatment-covariate interactions. Key findings include:
Consistently lower RMSE: Both IntLS and IntLF yield lower root mean squared errors for estimation of value and benefit functions compared to Separate Learning (SepL), with IntLF generally achieving the best performance across both metrics.
Diminishing returns with increased heterogeneity: The benefit of integration declines as the data-generating processes for the two studies diverge (as parameterized by B1 and B2), but cross-validation-driven regularization prevents substantial negative transfer except in the most extreme cases.
Figure 1: RMSE outcomes for linear interaction scenario, illustrating decreasing marginal gains for integrative approaches as study alignment (controlled by B3) declines.
Figure 2: Performance of value and benefit function estimators under nonlinear interaction models across increasing divergence (B4 values).
Real-World Application: MDD Trials
The methods are empirically validated on two major depression studies: EMBARC (sertraline vs. placebo) and iSPOT-D (sertraline vs. venlafaxine XR). Both include rich clinical, demographic, and EEG phenotypic features, with sertraline (an SSRI) as the shared comparator. The cross-study integrative methods yield:
Superior average values and benefit scores: IntLF achieves the highest values and benefit metrics on post-treatment depression improvement in both datasets, outperforming one-size-fits-all and separate-learning comparators.
Enhanced clinical relevance: IntLF more effectively identifies subpopulations with differential benefit; patients with higher HRSD scores and lower EEG-derived alpha power are more likely to be recommended SSRIs, matching known biological correlates of antidepressant response.
Figure 3: Heatmaps of key tailoring variables, stratified by IntLF-derived optimal treatment assignments in both studies, supporting cross-study consistency.
Figure 4: Distribution of average scaled tailoring variables for subjects predicted to derive optimal benefit from SSRI therapy, reflecting underlying biomarker-driven heterogeneity.
Practical and Theoretical Implications
For empirical researchers, the framework extends the capacity to draw robust, patient-specific treatment rules from networks of RCTs, even in the absence of direct treatment contrasts. The adaptive penalization approach mitigates risks from population/policy shift—the method borrows strength from related studies only to the extent that such transfer is empirically supported. This is of particular relevance to mental health and other applications with fragmented comparative trials.
Theoretically, the excess risk results clarify the regime under which integrative rule learning offers nontrivial statistical improvement, and when its impact is neutral (reverting to separate learning in the worst-case). The extension to multi-study (beyond two arms) networks via pairwise fusion penalties is straightforward, and the modular optimization scheme facilitates scalability.
Future Directions
Promising avenues for extension include subject-specific and data-driven penalties for more granular adaptive fusion, high-dimensional regularization (e.g., B5/elastic net in place of B6 penalties for variable selection), generalized outcome models (e.g., binary or time-to-event outcomes), and explicit modeling of distributional or covariate shift for fairness or robust transportability. Integration of multi-domain outcome modalities and broadening the framework for multi-study multi-arm networks are natural next steps.
Conclusion
This work presents a rigorous and flexible framework for integrative individualized treatment rule learning across multi-study RCTs with only partially overlapping interventions, yielding principled statistical and practical improvements in estimating optimal treatment policies. The method is shown to be robust in practice and supported by non-asymptotic theoretical guarantees, with direct application to high-stakes decision-making in mental health and other domains (2604.10712).
“Emergent Mind helps me see which AI papers have caught fire online.”
Philip
Creator, AI Explained on YouTube
Sign up for free to explore the frontiers of research
Discover trending papers, chat with arXiv, and track the latest research shaping the future of science and technology.Discover trending papers, chat with arXiv, and more.