Synergy Area with FDR-controlled Evaluation (SAFE) to robustly assess safety profile in clinical trials
Published 4 May 2026 in stat.AP | (2605.03041v1)
Abstract: Safety assessment plays a fundamental role in developing a new drug via clinical trials for ethical considerations. Due to complexity, manual review is typically conducted on the totality of data to draw safety conclusions. There are some existing quantitative methods to facilitate or tailor further medical review, with a controlled error rate and integration of clinical knowledge. In addition to those two key aspects, we emphasize the importance of relying on substantial evidence to draw robust conclusions on safety. Motivated by these three important properties, we propose a two-layer Synergy Area with FDR-controlled Evaluation (SAFE) structural framework to robustly assess the safety profile in clinical trials. In the first layer of SAFE, we investigate each clinically meaningful Synergy Area (SA) based on compelling evidence. In the next layer, the false discovery rate (FDR) is controlled for potential findings across all SAs. Simulation studies show that SAFE properly controls error rates within and across SAs at the nominal level. We further apply the proposed approach to two case studies based on real data from the Historical Trial Data (HTD) Sharing Initiative of the DataCelerate platform. As compared to some direct methods, SAFE demonstrates an appealing feature of screening out extreme data and reaching solid safety conclusions. It can act as either a building block in another framework, or a platform to incorporate additional components.
The paper introduces the SAFE framework, a two-layer method for robust safety assessment in clinical trials by enforcing convergent evidence from within clinically meaningful groups of adverse events (AEs) while preserving family-wise error rate (FWER) control, via FDR control across groups.
Simulations confirm SAFE's ability to control false discovery rate (FDR) at 5%, across 23 SAs containing 259 AE variables. It avoids false signals from single AE violations by applying their stringent decision function within groups.
Two case studies demonstrate SAFE’s practical utility for clinical trial safety assessments by avoiding false positives in gastrointestinal disorders and psychiatrist targets, demonstrating its reliability.
Motivation and positioning
Quantitative safety assessment in clinical trials must balance three requirements: controlled error rates, incorporation of clinical and domain knowledge, and conclusions supported by substantial evidence. Existing multiplicity-adjustment frameworks—FWER- or FDR-based procedures applied across adverse event (AE) terms, and Bayesian flagging approaches that exploit biological relationships among AEs—address the first two requirements but generally do not enforce the third. SAFE is designed to satisfy all three simultaneously by requiring convergent evidence within clinically meaningful groupings of AEs before declaring a safety signal (2605.03041).
The two-layer structure
SAFE organizes the safety data into m Synergy Areas (SAs), each containing ni AE variables. The SA definition is flexible: MedDRA System Organ Classes, high-level groups, severity scales, causality, or action taken with study drug are all admissible choices. Within each SA, the framework replaces the conventional union alternative K1(i)=⋃jH1(i,j) with a more stringent intersection-based alternative
G1(i)=j<k⋃[H1(i,j)∩H1(i,k)],
so that an SA is flagged only when at least two AE variables show signals. This mirrors clinical diagnostic practice, such as myocardial infarction diagnosis requiring troponin elevation plus corroborating symptoms or ECG changes, and Hy's law criteria combining transaminase and bilirubin conditions. The decision function Di(α) uses the second smallest Holm-adjusted p-value pi(2) within the SA; because this condition is stricter than the standard Holm rejection rule based on pi(1), no additional adjustment is needed for FWER control at level α under arbitrary dependence. The construction is equivalent to the Bonferroni-type partial conjunction test of Benjamini and Heller under general dependency.
In the second layer, the within-SA statistics p˘i=pi(2) are combined via the Benjamini–Hochberg (BH) procedure to control FDR at level α across SAs; the Benjamini–Yekutieli (BY) procedure substitutes when arbitrary dependency is a concern. The full algorithm requires only raw p-values per AE variable, Holm adjustment within SAs, extraction of second-smallest adjusted p-values, BH q-value computation, and thresholding at ni0. A generalized variant SAFE(ni1) requires ni2 elementary findings per SA.
Simulation evidence
Simulations with ni3 SAs, ni4 AE variables per SA, compound-symmetric correlation structures spanning positive dependency (ni5), independence, and the negative-definiteness boundary (ni6), and 100,000 iterations at nominal ni7 support two claims. First, within-SA false rejection probabilities are at or below 5% in all settings, approaching the nominal level when some elementary nulls are easily rejected (e.g., an AE variable with mean shift of 6), consistent with Berger's (1982) results on multiparameter tests. Second, cross-SA FDR is controlled at 5% throughout—for example, 4.8–5.0% under all-null configurations with independent or negatively correlated variables, dropping to 3.0% and below as true alternatives accumulate. Under strong positive within-SA correlation (ni8), BH-based FDR can be slightly inflated, but BY restores exact control; this is the one scenario where the default configuration requires care. Supplementary analyses confirm robustness across further correlation structures and show that SAFE(ni9) maintains error control for all K1(i)=⋃jH1(i,j)0, with both false positives and power decreasing monotonically in K1(i)=⋃jH1(i,j)1.
A direct method using the smallest adjusted p-value instead of the second-smallest exhibits inflated FDR in some scenarios, since its implicit alternative K1(i)=⋃jH1(i,j)2 is less restrictive than K1(i)=⋃jH1(i,j)3. The power cost of SAFE relative to direct methods is acknowledged explicitly: direct methods yield higher true-positive probabilities, so SAFE trades power for evidential robustness.
Case studies from TransCelerate HTD data
Case study 1 (HS vs. PsO placebo arms). With K1(i)=⋃jH1(i,j)4 SOCs and 259 AE variables, SAFE flags exactly one SA—"gastrointestinal disorders"—with a q-value of 3.6%. Notably, "nervous system disorders" and "skin and subcutaneous tissue disorders" have extremely small smallest adjusted p-values (K1(i)=⋃jH1(i,j)5 values of −5.23 and −5.53) but near-zero second-smallest values, so SAFE does not flag them for lack of corroborating evidence. Direct Holm and BH procedures flag AEs in four SAs, including those two driven by single extreme findings. The flagged gastrointestinal finding is directionally consistent with prior observational literature associating HS with gastrointestinal comorbidity, though the authors correctly characterize it as exploratory rather than confirmatory.
Case study 2 (two AD trials). Here one SA has a strikingly small smallest adjusted p-value (K1(i)=⋃jH1(i,j)6), yet its second-smallest value is large, and all q-values exceed 5%; SAFE declares no safety difference. Both direct methods flag that single AE. In a setting where similar safety profiles are expected between the two trials, SAFE's behavior is the more defensible one—it screens out an isolated extreme observation that direct methods convert into a "finding."
Together, the case studies demonstrate SAFE's central practical property: it suppresses single-AE-driven discoveries that lack internal corroboration, at the cost of missing isolated true signals.
Limitations and open questions
The authors state plainly that SAFE outputs are not final safety conclusions and require subsequent medical review; the framework prioritizes review rather than replacing it. Several limitations bear directly on the results. The method presumes high-quality input data, and no missing-data handling strategy is incorporated—the case-study comparisons across different trials also inherit heterogeneity between trial populations that the analysis cannot fully exclude. Power loss relative to direct methods is inherent to the two-finding requirement, and the choice of K1(i)=⋃jH1(i,j)7 (or adaptive rules such as rounding K1(i)=⋃jH1(i,j)8) is left to operating-characteristic evaluation rather than resolved by theory. Under strong positive within-SA correlation, the default BH layer can be mildly anti-conservative, necessitating BY. Finally, the equivalence to partial conjunction testing means more powerful combining rules (e.g., shifted Simes) could improve power under additional assumptions, but these are not developed here. Open questions include extension to ongoing-trial monitoring, post-marketing surveillance, and integration of composite benefit–risk scores.
Conclusion
SAFE provides a statistically disciplined mechanism for requiring substantial, internally corroborated evidence before flagging a safety area, with FDR control across areas and FWER-type control within them under arbitrary dependence. Simulations confirm nominal error control across dependency regimes, and two real-data applications illustrate its ability to filter single-extreme-observation findings that direct multiplicity-adjusted methods would report. Its main costs—reduced power and sensitivity to data quality—are explicit trade-offs rather than defects, and the framework's modularity permits both deeper hierarchies over MedDRA levels and use as a component within broader safety evaluation infrastructures.