GenFacts: Dual Frameworks in Genomics & AI
- GenFacts is a term defining two distinct frameworks: a historical, data-driven genomics testing approach using FAB techniques, and a VAE-based method for counterfactual explanations in time series.
- The genomics framework employs probabilistic tensor factorization to distill multi-omics data, enabling enhanced hypothesis testing while preserving frequentist error guarantees.
- The time-series framework integrates class-discriminative variational autoencoders with prototype-based latent optimization to generate plausible, interpretable counterfactual explanations for sequential data.
Searching arXiv for papers on “GenFacts” and closely related uses of the term. “GenFacts” denotes at least two distinct methodological usages in the literature. In one usage, it refers to a genomics testing framework built around “frequentist assisted by Bayes” (FAB), in which historical multi-omics data are distilled by a probabilistic tensor factorization model and then used to sharpen hypothesis tests in smaller specialized studies while preserving null-valid -values (Bryan et al., 2020). In another usage, it denotes a generative framework for counterfactual explanations in multivariate time series, based on a class-discriminative variational autoencoder, prototype-based initialization, and realism-constrained latent optimization (Seifi et al., 25 Sep 2025). The shared theme is transfer of structured prior information into a downstream inference or explanation task, but the two usages address different data modalities, objectives, and validation criteria.
1. Terminological scope and distinct usages
The genomics usage of GenFacts is organized around a two-stage pipeline. Historical genomics data across cell lines, genes, and modalities are represented as a tensor of effects, and a low-rank probabilistic model is fit to distill latent features for cell lines and genes. These distilled features are then used to construct hypothesis-specific prior directions that shift classical -tests into FAB tests, yielding smaller -values when the historical information is relevant and classical behavior when it is not (Bryan et al., 2020).
The time-series usage of GenFacts is a counterfactual-explanation framework for multivariate time series classification. It first trains a class-discriminative variational autoencoder and then generates counterfactuals by optimizing in the latent space with classification, proximity, and realism constraints. In that setting, the method is evaluated on radar gesture data and handwritten letter trajectories, with emphasis on plausibility and human-centered interpretability rather than sparsity alone (Seifi et al., 25 Sep 2025).
This dual usage suggests that “GenFacts” is not yet a single stable technical term across domains. A plausible implication is that the term should be interpreted contextually: in genomics, it designates a history-informed testing procedure; in explainable AI for sequential data, it designates a generative counterfactual method.
2. Genomics GenFacts: distilled historical information for hypothesis testing
In the genomics setting, the problem is to improve power in small or moderately sized hypothesis-driven studies by borrowing information from a large historical corpus without sacrificing frequentist validity. The testing problem is formulated over many hypotheses , with estimators modeled as
and classical two-sided tests based on
where (Bryan et al., 2020).
The historical data are organized as a tensor
where indexes cancer cell lines, 0 indexes genes, and 1 indexes modalities. The low-rank probabilistic model is
2
with latent feature vectors 3 for cell lines and 4 for genes, and modality-specific interaction matrices 5 (Bryan et al., 2020). The design matrix for a new study is then formed as
6
or a subset of its rows corresponding to the tested effects.
The distilled historical information is converted into an approximate prior
7
estimated by an empirical Bayes leave-one-out strategy to preserve independence between the prior-derived shift and the test statistic. The FAB shift is then
8
and the resulting 9-value becomes
0
The critical property is that for any fixed 1, the FAB 2-value is uniform under 3, provided 4 is independent of 5 (Bryan et al., 2020).
This construction makes the method adaptive. When historical information is predictive, the shift behaves like an oracle directional bias and can increase discoveries. When the historical information is weak or irrelevant, 6, and the procedure reduces to the ordinary two-sided 7-test (Bryan et al., 2020).
3. Probabilistic structure and multimodal transfer in the genomics framework
A central feature of the genomics formulation is its multimodal probability model. Rather than fitting separate decompositions for RNA-seq, CRISPR dependency, RNAi dependency, mutation calls, or drug viability profiling, the model shares latent cell-line and gene representations across modalities while allowing modality-specific interactions through 8 (Bryan et al., 2020). This shared-latent design is what supports transfer from broad historical data to a new assay or targeted follow-up study.
The framework also accommodates different observation models. For continuous data it uses the normal model
9
while the details note extensions to a probit model for binary mutation calls and a Tobit model for strictly positive or censored continuous data such as RNA-seq (Bryan et al., 2020). This makes the historical representation explicitly multimodal rather than restricted to a single assay type.
The practical objective is not direct effect estimation from the historical tensor alone, but distillation of reusable features that generalize across modalities and studies. This differs from classical empirical Bayes shrinkage on a single experiment. Here the historical corpus functions as a structured map of cell-line and gene relationships, and the new study uses that map to obtain hypothesis-specific prior directions (Bryan et al., 2020).
The role of the leave-one-out construction is especially important. It ensures that the estimated shift does not recycle the same evidence used in the downstream test, thereby preserving the exact null calibration of the resulting 0-values (Bryan et al., 2020).
4. Empirical behavior of the genomics framework
The genomics paper reports both simulation and real-data results. Under null simulations, based on 10,000 simulated datasets with 1, the FAB 2-values were empirically uniform, and BH-adjusted FAB results achieved target FDR 3, matching classical tests (Bryan et al., 2020). This addresses the main methodological concern that power gains might come at the expense of type I error inflation.
Under non-null simulations, the signal model is
4
The reported behavior is monotone in the relevance of historical information: if 5, history is irrelevant and FAB matches classical testing; as 6 decreases, FAB produces more discoveries; and overall performance interpolates between the classical two-sided test and an oracle one-sided test (Bryan et al., 2020).
The real-data studies span several genomics settings. In CRISPR essentiality screens in acute myeloid leukemia, FAB yielded more discoveries than the standard two-sided test. In drug viability profiling for repurposing, FAB produced as many or more discoveries for most drugs, and for 67% of drugs it produced strictly more discoveries. In mammary cell subpopulation differential expression, FAB gave more discoveries for all tested contrasts. In a serine metabolism and differential dependency example, the historical fit was weak, but the method remained conservative and aligned closely with classical results (Bryan et al., 2020).
A further reported application is a Ras-mutant follow-up study, in which FAB found 30 discoveries versus 23 for classical testing. Additional FAB-only hits included genes such as NRAS, ICMT, ELAVL1, and ELOVL1, described in the details as biologically plausible Ras-related genes (Bryan et al., 2020). The principal significance of these findings is methodological rather than purely biological: they show that the framework can exploit historical relevance when present and avoid degradation when absent.
5. Time-series GenFacts: generative counterfactual explanations
In the multivariate time-series setting, GenFacts addresses counterfactual explanations for a trained classifier 7, where 8. The objective is to generate a counterfactual 9 such that 0, while keeping the sequence close to the original input, plausible with respect to the data manifold, and interpretable to a human observer (Seifi et al., 25 Sep 2025).
The method is explicitly motivated by limitations of existing counterfactual procedures for multivariate time series. The reported failure modes are poor plausibility, ignoring cross-variable and temporal correlations, and limited actionability when optimization is driven primarily by minimal-change or sparsity criteria (Seifi et al., 25 Sep 2025). The paper argues that in applications such as radar gestures, quantities like range, velocity, and angle are not independent, so feature-wise sparse perturbations can violate domain structure.
GenFacts therefore adopts a two-stage generative design. Stage 1 trains a class-discriminative variational autoencoder. Stage 2 freezes the encoder, decoder, and classifier and performs gradient-based optimization in latent space, initialized from a target-class prototype and regularized toward realism (Seifi et al., 25 Sep 2025). The resulting framework is designed to generate class-changing explanations that remain on or near the learned data manifold.
This use of “GenFacts” is conceptually unrelated to the FAB genomics procedure, despite the shared name. One is a hypothesis-testing framework for genomics; the other is an explainability framework for sequential sensor and trajectory data.
6. Architecture, objectives, and evaluation of the time-series framework
The encoder in the time-series model consists of a 2-layer bidirectional LSTM, followed by 1D CNN layers and a fully connected layer producing 1 and 2. The decoder consists of a fully connected layer, transposed CNN layers, and a bidirectional LSTM. The latent variable is sampled from
3
with prior 4 (Seifi et al., 25 Sep 2025).
The VAE is trained with a multi-objective loss
5
where the terms correspond to reconstruction, KL regularization, classification consistency, and contrastive learning (Seifi et al., 25 Sep 2025). The classification-consistency term matches the classifier’s softmax prediction on the reconstruction to that on the original input, while the contrastive term uses 6 to pull together same-class latent vectors and push apart different-class latent vectors.
Counterfactual generation optimizes
7
with a classification loss driving the decoded sample toward the target class, a proximity loss encouraging closeness to the original sample, and a realism loss
8
that keeps the optimized code near the latent prior center (Seifi et al., 25 Sep 2025). A distinctive feature is prototype-based initialization, using the mean latent vector of the 9 nearest neighbors from the target class, which is reported to improve convergence speed, success rate, and plausibility.
The reported implementation uses latent dimension 40, Adam with learning rate 0.002, batch size 32, training up to 600 epochs, KL annealing with a cosine schedule, and linearly annealed classification and contrastive losses over the first half of training. The VAE loss weights are 0, 1, 2, and 3; the counterfactual generation weights are 4, 5, and 6 (Seifi et al., 25 Sep 2025).
Evaluation is performed on a 60 GHz FMCW radar gesture dataset and handwritten letter trajectories. For radar gestures, a GRU-based classifier trained on 6,809 samples with an 80/20 train-validation split achieved 7. For letter trajectories, a GRU classifier trained on 2,858 samples with a 75/25 split and 8 hidden units for 100 epochs achieved 8 (Seifi et al., 25 Sep 2025).
Quantitatively, the main metrics are validity, proximity, and plausibility. Plausibility is defined by inlier status under an ensemble of isolation forests and membership within the 90th percentile of pairwise training distances, using fixed-length descriptors based on statistical moments, temporal dynamics, and correlations (Seifi et al., 25 Sep 2025). On the radar benchmark, the reported table gives GenFacts proximity 9, validity 0, and plausibility 1, outperforming baselines in plausibility and maintaining full validity. The paper summarizes this as a 2 plausibility improvement and the highest interpretability scores in a human study (Seifi et al., 25 Sep 2025).
The human study uses five independent evaluators focused on ambiguous diagonal radar swipes. The reported interpretability scores are 38.6 for CoMTE, 3.4 for TSEvo, 2.1 for CFProto, 9.7 for SPARCE, 5.5 for Multi-SpaCE, and 90.4 for GenFacts (Seifi et al., 25 Sep 2025). The ablation study attributes distinct roles to the three counterfactual loss terms: removing 3 worsens proximity by 4, removing 5 lowers validity by 18.7 percentage points, and removing 6 lowers plausibility by 18.7 percentage points (Seifi et al., 25 Sep 2025).
7. Conceptual significance, limitations, and relation to neighboring methods
Across both usages, GenFacts represents a strategy of structured borrowing of information, but the meaning of “information” differs by domain. In the FAB genomics framework, the borrowed information is historical statistical regularity across cell lines, genes, and modalities, distilled into latent factors and converted into directional shifts for valid hypothesis tests (Bryan et al., 2020). In the time-series framework, the borrowed information is manifold structure and class geometry, learned by a discriminative generative model and used to constrain counterfactual search toward plausible regions of latent space (Seifi et al., 25 Sep 2025).
The genomics framework is adjacent to empirical Bayes multiple testing and multimodal tensor factorization, but its defining feature is the preservation of frequentist guarantees through null-valid FAB 7-values and leave-one-out independence constructions (Bryan et al., 2020). The time-series framework is adjacent to VAE-based counterfactual generation, prototype methods, and realism-constrained latent optimization, but its main emphasis is that plausibility and human-centered interpretability matter more than sparsity alone for actionable time-series explanations (Seifi et al., 25 Sep 2025).
The limitations are domain-specific. In the genomics case, gains depend on the relevance of the historical corpus to the new study, even though the method is designed to revert to classical testing when that relevance is weak (Bryan et al., 2020). In the time-series case, the learned manifold is only as good as the VAE and classifier used to define it, and the reported evaluations are centered on two datasets: radar gestures and handwritten letter trajectories (Seifi et al., 25 Sep 2025).
A likely source of confusion is simple nomenclature. Because the same label is attached to methods in genomics and in time-series explainable AI, readers must disambiguate by citation and context. In current usage, “GenFacts” is best understood not as a single canonical theory but as a name attached to at least two domain-specific frameworks with distinct objectives, mathematical constructions, and empirical criteria (Bryan et al., 2020, Seifi et al., 25 Sep 2025).