Safeterm Trial-Safety App Overview
- The Safeterm Trial-Safety App is a clinical trial safety analytics platform that integrates transformer-based MedDRA encoding with automated signal detection and PRO instrument optimization.
- It utilizes unsupervised semantic clustering, spectral analysis, and advanced visualizations to balance patient burden with comprehensive adverse event coverage.
- The modular design ensures seamless integration with clinical workflows, enabling reproducible safety assessments and query generation for enhanced regulatory review.
The Safeterm Trial-Safety App is a web- and API-based platform for clinical trial safety analytics and patient-reported outcome (PRO) instrument optimization. Safeterm leverages a high-dimensional transformer-based embedding model to encode MedDRA Preferred Terms (PTs), integrating historical adverse event (AE) data, semantic mapping, clustering, utility-driven selection, and advanced visualizations to support automated signal detection, PRO-CTCAE design, and knowledge-based review. This approach streamlines patient burden–signal coverage trade-offs, enables unsupervised or reproducible MedDRA query generation, and enriches trial data interpretation for sponsors and regulatory professionals (Vandenhende et al., 7 Dec 2025, Vandenhende et al., 8 Dec 2025, Vandenhende et al., 8 Dec 2025, Vandenhende et al., 24 Nov 2025).
1. System Architecture and Data Flow
The Safeterm Trial-Safety App operates through modular backend and frontend components:
- Frontend (Web Client): Users interact via a React/TypeScript interface, inputting historical AE profiles as MedDRA PT lists (with optional incidence counts). The app provides ranked PRO-CTCAE candidate tables, interactive plots (2D projections, leverage vs. rank), and CSV/Excel export (Vandenhende et al., 7 Dec 2025).
- Backend API: Implemented using Python (FastAPI/Flask), it exposes RESTful endpoints (e.g.,
/select_profor PRO-CTCAE selection, AMQ endpoints for MedDRA queries) that orchestrate mapping, embedding, scoring, clustering, and spectral selection pipelines. - Data Stores: SQL/NoSQL databases hold MedDRA dictionaries, mapping tables linking PRO-CTCAE items to PTs, and the Safeterm embedding model (PyTorch, d=300).
- Outputs: Structured JSON returns candidate term rankings (relevance, utility, diversity, leverage), recommended cut-offs (
k_opt), and scores, with direct export and browser-based visualization capabilities.
This architecture supports seamless integration with EDC/pharmacovigilance workflows and enables interactive data-driven refinement for safety monitoring, PRO selection, and query generation.
2. MedDRA Mapping and Semantic Embedding
PRO-CTCAE to MedDRA Mapping: Each PRO-CTCAE symptom (≈124 plain-language items) is manually mapped by expert terminologists to one or two MedDRA PTs, resolving lexical ambiguity via LLTs; this preserves the original PRO intent while providing semantic linkage (Vandenhende et al., 7 Dec 2025).
Safeterm Embedding Model: All MedDRA PTs are encoded in a transformer-based model, trained on large biomedical corpora and MedDRA hierarchy, yielding normalized vectors . This embedding space forms the basis for all semantic computations (cosine similarity, clustering, diversity scoring).
Semantic Similarity: For two normalized vectors , , similarity is ; broader relationships (clinical, mechanistic, linguistic) are captured beyond strict MedDRA hierarchy (Vandenhende et al., 24 Nov 2025).
3. Relevance, Utility, and Diversity Ranking
Relevance Scoring:
- Redundancy among PRO items: , .
- Relevance to AE history: , .
- Raw relevance: .
- Incidence weighting: , with 0.
Utility Function:
- Saturated relevance: 1 (2, 3).
- Combined utility: 4, 5.
L-kernel for Utility/Diversity: From Determinantal Point Process theory, 6, 7. Diagonal entries capture utility; off-diagonals encode semantic overlap penalties.
4. Spectral Analysis and Orthogonal Symptom Selection
Eigen-Decomposition and Explained Variance:
- 8, where 9, 0 orthonormal eigenvectors.
- Cumulative explained variance: 1.
- Minimal orthogonal set size: 2 (3–4 typical).
Diversity Leverage Score:
- For item 5: 6.
- Items are rank-ordered by leverage, enforcing selection of the top 7 for coverage across all axes.
5. Automated MedDRA Query Generation and Validation
Safeterm incorporates AMQ (Automated Medical Query) features for MedDRA term retrieval (Vandenhende et al., 8 Dec 2025, Vandenhende et al., 8 Dec 2025):
- Workflow: Free-text query or MedDRA PT input 8 embedding 9 cosine similarity computation 0 extreme-value (two-means) clustering 1 knee-point threshold selection 2 ranked PT candidate list.
- Thresholding: Lower thresholds (e.g., 0.50–0.60) maximize recall (≈0.94 for SMQs, ≈0.95 for OCMQs); higher thresholds (0.70–0.90) increase precision (up to 0.89 for SMQs, 0.86 for OCMQs), sacrificing recall.
- Performance: For the optimal F1 threshold (30.70): SMQ recall 0.48/precision 0.45/F1 0.44; OCMQ recall 0.57/precision 0.34/F1 0.37.
- Narrow-term PTs: Require slightly higher similarity thresholds, maintain recall, slightly reduced precision by gold set size.
- Recommendations: Use valid MedDRA PTs as queries, adjust thresholds to match sensitivity/specificity needs, integrate with EDC systems for real-time query generation and review.
6. Visualization, Knowledge Layer, and Clustering
Hidden Medical Knowledge Layer: Safeterm augments MedDRA PTs with high-dimensional embeddings, semantic descriptors, and precomputed pairwise cosine similarities, forming a latent relationship graph (Vandenhende et al., 24 Nov 2025).
Automatic Clustering:
- Trial-observed PT embeddings are reduced (PCA) and clustered via agglomerative or k-means algorithms.
- Cluster identity is decoded via AI translators from embedding centroids; ungrouped PTs (low silhouette scores) are flagged and colored distinctly.
Shrinkage Incidence Ratio (SIR) and Cluster-Level EBGM:
- Expected count: 4; SIR 5 (gamma-Poisson shrinkage).
- Cluster-level aggregation: Precision-weighted mean 6, 7.
Visualization Outputs:
- Semantic Map: 2D PCA/t-SNE projection of PTs, colored by semantic cluster, sized by incidence rate; interactive filtering and tooltip details.
- Expectedness-versus-Disproportionality Plot (EVD): X-axis: expectedness (cosine similarity to disease indication vector), Y-axis: 8. Points colored by cluster, sized by incidence. Outliers (low expectedness, high 9) denote novel safety signals.
7. Empirical Results and Practical Integration
Monte Carlo Simulations (N=100,000): Mean recall 0.70, precision 0.72, F1 0.70 (info threshold 97.5%), stable across signal/noise levels (Vandenhende et al., 7 Dec 2025).
Oncology Case Study (Multiple Myeloma):
- Phase I: Algorithm selected 0 PRO-CTCAE items; all matched AE PTs; 9 exact-matches flagged and excluded for redundancy.
- Phase II: Automated list overlapped with 8 of 15 manual PROs; coverage was comparable (auto 11/16, manual 11/15 retrieved).
- Automated selection provided objective, reproducible design and explicit burden–coverage justification.
Legacy Trials with Semantic Clustering:
- Duchenne Muscular Dystrophy: Liver damage cluster detected (semantic map, cluster-level EBGM); minor hepatotoxicity signals enriched.
- Narcolepsy Dose-Response: Dose-dependent stress cluster SIR rise detected.
- Hodgkin’s Lymphoma: Bone marrow failure cluster differentiated between treatments.
Practical Recommendations:
- Start broad signal detection at moderate thresholds, refine for specificity as needed.
- Leverage semantic clustering and visualization for hypothesis generation and transparent safety review.
- Integrate app endpoints with clinical EDC, pharmacovigilance, and dashboard systems.
Safeterm transforms trial safety workflows by embedding MedDRA PTs in a semantically calibrated hidden space, enabling objective, reproducible PRO selection, rapid and unsupervised term query generation, and advanced clustering-based signal analysis, validated across diverse oncology and neurology trials (Vandenhende et al., 7 Dec 2025, Vandenhende et al., 8 Dec 2025, Vandenhende et al., 8 Dec 2025, Vandenhende et al., 24 Nov 2025).