Papers
Topics
Authors
Recent
Search
2000 character limit reached

bfact: Boolean & Bayesian Factor Analysis

Updated 10 July 2026
  • bfact is an overloaded term that encompasses Boolean matrix factorisation, Bayesian factor analysis, and Bayes factor methods, as well as B factory physics.
  • In Boolean matrix factorisation, bfact refers to a Python package that employs disjoint factor selection and hybrid combinatorial optimization to decompose binary data effectively.
  • Within Bayesian contexts, bfact describes advanced factor models that integrate robust outlier detection and multiple hypothesis testing using Bayes factors, aiding applications in genomics and econometrics.

to=arxiv_search.search 娱乐开号 彩神争霸苹果"query":"bfact","max_results":10,"sort_by":"relevance"}{"result":"Search results for 'bfact' (max 10):\n\n1. (Visscher et al., 7 Sep 2025) - Hybrid restricted master problem for Boolean matrix factorisation\n Authors: Michael S. S. Law, Benjamin P. C. Tomlinson, Guido Sanguinetti\n Published: 2025-09-07\n Categories: cs.LG, cs.DS, math.OC\n Abstract: We present bfact, a Python package for performing accurate low-rank Boolean matrix factorisation (BMF). bfact uses a hybrid combinatorial optimisation approach based on a priori candidate factors generated from clustering algorithms. It selects the best disjoint factors before performing either a second combinatorial or heuristic algorithm to recover the BMF. We show that bfact does particularly well at estimating the true rank of matrices in simulated settings. In real benchmarks, using a collation of single-cell RNA-sequencing datasets from the Human Lung Cell Atlas, we show that bfact achieves strong signal recovery, with a much lower rank.\n\n2. (Liang et al., 23 Jun 2025) - Bayesian integrative factor analysis methods, with application in nutrition and genomics data\n Authors: Xinyue Liang, Wenyi Wang, Gen Li\n Published: 2025-06-23\n Categories: stat.ME, stat.AP, q-bio.GN\n Abstract: High-dimensional data are crucial in biomedical research. Integrating such data from multiple studies is a critical process that relies on the choice of advanced statistical models, enhancing statistical power, reproducibility, and scientific insight compared to analyzing each study separately. Factor analysis (FA) is a core dimensionality reduction technique that models observed data through a small set of latent factors. Bayesian extensions of FA have recently emerged as powerful tools for multi-study integration, enabling researchers to disentangle shared biological signals from study-specific variability. In this tutorial, we provide a practical and comparative guide to five advanced Bayesian integrative factor models: Perturbed Factor Analysis (PFA), Bayesian Factor Regression with non-local spike-and-slab priors (MOM-SS), Subspace Factor Analysis (SUFA), Bayesian Multi-study Factor Analysis (BMSFA), and Bayesian Combinatorial Multi-study Factor Analysis (Tetris). To contextualize these methods, we also include two benchmark approaches: standard FA applied to pooled data (Stack FA) and FA applied separately to each study (Ind FA). We evaluate all methods through extensive simulations, assessing computational efficiency and accuracy in the estimation of loadings and number of factors. To bridge theory and practice, we present a full analytical workflow, with detailed R code, demonstrating how to apply these models to real-world datasets in nutrition and genomics.\n\n3. (Billio et al., 25 Mar 2025) - Bayesian Outlier Detection for Matrix-variate Models\n Authors: Alessio Bissiri, Luca Casarin, Roberto Golinelli, Francesco Ravazzolo, Fabio Rigat\n Published: 2025-03-25\n Categories: stat.ME, econ.EM, stat.CO\n Abstract: Bayes Factor (BF) is one of the tools used in Bayesian analysis for model selection. The predictive BF finds application in detecting outliers, which are relevant sources of estimation and forecast errors. An efficient framework for outlier detection is provided and purposely designed for large multidimensional datasets. Online detection and analytical tractability guarantee the procedure's efficiency. The proposed sequential Bayesian monitoring extends the univariate setup to a matrix--variate one. Prior perturbation based on power discounting is applied to obtain tractable predictive BFs. This way, computationally intensive procedures used in Bayesian Analysis are not required. The conditions leading to inconclusive responses in outlier identification are derived, and some robust approaches are proposed that exploit the predictive BF's variability to improve the standard discounting method. The effectiveness of the procedure is studied using simulated data. An illustration is provided through applications to relevant benchmark datasets from macroeconomics and finance.\n\n4. (Ma et al., 2024) - Robust bilinear factor analysis based on the matrix-variate t distribution\n Authors: Xuejun Pan, Jiahui Ai, Fuping Huang, Song Xi Chen, Shaobo Jin\n Published: 2024-01-04\n Categories: stat.ME\n Abstract: Factor Analysis based on multivariate t distribution (tfa) is a useful robust tool for extracting common factors on heavy-tailed or contaminated data. However, tfa is only applicable to vector data. When tfa is applied to matrix data, it is common to first vectorize the matrix observations. This introduces two challenges for tfa: (i) the inherent matrix structure of the data is broken, and (ii) robustness may be lost, as vectorized matrix data typically results in a high data dimension, which could easily lead to the breakdown of tfa. To address these issues, starting from the intrinsic matrix structure of matrix data, a novel robust factor analysis model, namely bilinear factor analysis built on the matrix-variate t distribution (tbfa), is proposed in this paper. The novelty is that it is capable to simultaneously extract common factors for both row and column variables of interest on heavy-tailed or contaminated matrix data. Two efficient algorithms for maximum likelihood estimation of tbfa are developed. Closed-form expression for the Fisher information matrix to calculate the accuracy of parameter estimates are derived. Empirical studies are conducted to understand the proposed tbfa model and compare with related competitors. The results demonstrate the superiority and practicality of tbfa. Importantly, tbfa exhibits a significantly higher breakdown point than tfa, making it more suitable for matrix data.\n\n5. (Mittal et al., 2020) - Deep Bayesian Nonparametric Factor Analysis\n Authors: Erik A. Platanios, Michael I. Jordan, et al.\n Published: 2020-11-09\n Categories: stat.ML, cs.LG\n Abstract: We propose a deep generative factor analysis model with beta process prior that can approximate complex non-factorial distributions over the latent codes. We outline a stochastic EM algorithm for scalable inference in a specific instantiation of this model and present some preliminary results.\n\n6. (Wen, 2013) - Robust Bayesian FDR Control using Bayes Factors, with Applications to Multi-tissue eQTL Discovery\n Authors: Xiaoquan Wen\n Published: 2013-11-15\n Categories: stat.AP, q-bio.GN\n Abstract: Motivated by the genomic application of expression quantitative trait loci (eQTL) mapping, we propose a new procedure to perform simultaneous testing of multiple hypotheses using Bayes factors as input test statistics. One of the most significant features of this method is its robustness in controlling the targeted false discovery rate (FDR) even under misspecifications of parametric alternative models. Moreover, the proposed procedure is highly computationally efficient, which is ideal for treating both complex system and big data in genomic applications. We discuss the theoretical properties of the new procedure and demonstrate its power and computational efficiency in applications of single-tissue and multi-tissue eQTL mapping.\n\n7. (Johnson et al., 2022) - Bayes factor functions for reporting outcomes of hypothesis tests\n Authors: Donald B. Johnson, et al.\n Published: 2022-09-30\n Categories: stat.ME, math.ST\n Abstract: Bayes factors represent the ratio of probabilities assigned to data by competing scientific hypotheses. Drawbacks of Bayes factors are their dependence on prior specifications that define null and alternative hypotheses and difficulties encountered in their computation. To address these problems, we define Bayes factor functions (BFF) directly from common test statistics. BFFs depend on a single non-centrality parameter that can be expressed as a function of standardized effect sizes, and plots of BFFs versus effect size provide informative summaries of hypothesis tests that can be easily aggregated across studies. Such summaries eliminate the need for arbitrary P-value thresholds to define ``statistical significance.'' BFFs are available in closed form and can be computed easily from z, t, chi-squared, and F statistics.\n\n8. (Nenova et al., 2013) - An FCA-based Boolean Matrix Factorisation for Collaborative Filtering\n Authors: Y. Y. Yevtushenko, D. M. Kourie, G. J. J. Oosthuizen, S. O. Kuznetsov, R. Kruger\n Published: 2013-10-16\n Categories: cs.IR\n Abstract: We propose a new approach for Collaborative Filtering which is based on Boolean Matrix Factorisation (BMF) and Formal Concept Analysis. In a series of experiments on real data (Movielens dataset) we compare the approach with the SVD- and NMF-based algorithms in terms of Mean Average Error (MAE). One of the experimental consequences is that it is enough to have a binary-scaled rating data to obtain almost the same quality in terms of MAE by BMF than for the SVD-based algorithm in case of non-scaled data.\n\n9. (Bevan, 2012) - B Factories\n Authors: Adrian Bevan, et al.\n Published: 2012-02-03\n Categories: hep-ex, hep-ph\n Abstract: The B factories were constructed with a very specific purpose in mind: to test the Cabibbo-Kobayashi-Maskawa description of quark mixing and CP violation in the Standard Model of particle physics. The goals of testing this part of the Standard Model were achieved, and have been surpassed beyond all expectation. As a result the B factories have revolutionised our understanding of many areas of the Standard Model of particle physics, and also provide a number of stringent limits on possible scenarios of physics beyond the Standard Model. In some cases these limits on new physics effects equal or surpass those achievable at the CERN based Large Hadron Collider."} “bfact” is a context-dependent label rather than a single standardized technical term in arXiv literature. It appears in work on Boolean matrix factorisation, especially as the name of a Python package for low-rank Boolean decompositions; in Bayesian factor-analysis literature, including multi-study, robust, and deep nonparametric models; in Bayes-factor-based statistical inference; and, in an unrelated high-energy-physics usage, in connection with B factories. The term therefore denotes a family of abbreviation-driven usages whose meaning is determined by domain and surrounding notation rather than by a unique definition (Visscher et al., 7 Sep 2025, Liang et al., 23 Jun 2025, Billio et al., 25 Mar 2025, Bevan, 2012).

1. Terminological scope and domain-specific meanings

In current arXiv usage, “bfact” is best treated as an overloaded label. The main technical senses represented in the literature are summarized below.

Usage Core object Representative paper
Boolean matrix factorisation Binary factor models and Boolean products (Nenova et al., 2013, Visscher et al., 7 Sep 2025)
Bayesian factor analysis Latent-factor models with Bayesian priors and inference (Liang et al., 23 Jun 2025, Ma et al., 2024, Mittal et al., 2020)
Bayes-factor methods Model comparison, multiple testing, and reporting of evidence (Billio et al., 25 Mar 2025, Wen, 2013, Johnson et al., 2022)
B Factories Asymmetric-energy e+ee^+e^- facilities for flavor physics (Bevan, 2012)

This distribution of meanings has methodological consequences. In Boolean matrix factorisation, the central object is a binary matrix and the key algebra is logical OR over logical AND. In Bayesian factor analysis, the central object is a latent-factor decomposition with priors over loadings, factor scores, and variance components. In Bayes-factor usage, the core quantity is a ratio of marginal or predictive probabilities. In B-factory physics, the phrase refers to accelerator-detector complexes constructed to test the Cabibbo–Kobayashi–Maskawa description of quark mixing and CP violation.

2. Boolean matrix factorisation in collaborative filtering

A major “bfact” usage concerns Boolean matrix factorisation for recommender systems. In the FCA-based collaborative-filtering formulation, a binary matrix I{0,1}n×mI \in \{0,1\}^{n\times m} is decomposed into a Boolean product of binary matrices A{0,1}n×rA \in \{0,1\}^{n\times r} and B{0,1}r×mB \in \{0,1\}^{r\times m},

(AB)ij=k=1r(AikBkj),(A \odot B)_{ij} = \bigvee_{k=1}^r (A_{ik} \wedge B_{kj}),

with the paper using (PQ)(P \circ Q) for the same semantics. This differs from SVD or NMF, which use real-valued arithmetic and additive aggregation. The Boolean setting yields discrete biclusters rather than continuous latent components, and the decomposition is non-linear under Boolean algebra (Nenova et al., 2013).

The formal-concept layer is supplied by Formal Concept Analysis. A formal context K=(G,M,I)\mathcal{K}=(G,M,I) consists of objects GG, attributes MM, and a binary relation IG×MI \subseteq G \times M. For I{0,1}n×mI \in \{0,1\}^{n\times m}0 and I{0,1}n×mI \in \{0,1\}^{n\times m}1, the derivation operators are

I{0,1}n×mI \in \{0,1\}^{n\times m}2

A formal concept is a pair I{0,1}n×mI \in \{0,1\}^{n\times m}3 such that I{0,1}n×mI \in \{0,1\}^{n\times m}4 and I{0,1}n×mI \in \{0,1\}^{n\times m}5, where I{0,1}n×mI \in \{0,1\}^{n\times m}6 is the extent and I{0,1}n×mI \in \{0,1\}^{n\times m}7 is the intent. Given a subset of concepts I{0,1}n×mI \in \{0,1\}^{n\times m}8, one constructs indicator matrices I{0,1}n×mI \in \{0,1\}^{n\times m}9 and A{0,1}n×rA \in \{0,1\}^{n\times r}0, and A{0,1}n×rA \in \{0,1\}^{n\times r}1 reproduces the covered part of A{0,1}n×rA \in \{0,1\}^{n\times r}2. The theoretical basis is the universality statement that for every binary matrix A{0,1}n×rA \in \{0,1\}^{n\times r}3, there exists A{0,1}n×rA \in \{0,1\}^{n\times r}4 such that A{0,1}n×rA \in \{0,1\}^{n\times r}5, and the optimality statement that if A{0,1}n×rA \in \{0,1\}^{n\times r}6 for binary A{0,1}n×rA \in \{0,1\}^{n\times r}7, then there exists A{0,1}n×rA \in \{0,1\}^{n\times r}8 with A{0,1}n×rA \in \{0,1\}^{n\times r}9 such that B{0,1}r×mB \in \{0,1\}^{r\times m}0.

Algorithmically, the approach uses Algorithm 2 from Belohlavek & Vychodil (2010), which avoids computing the entire concept lattice and instead incrementally discovers closed biclusters that increase coverage. Inputs are a binary-scaled user–item matrix and either a target coverage B{0,1}r×mB \in \{0,1\}^{r\times m}1 or maximal number of factors B{0,1}r×mB \in \{0,1\}^{r\times m}2. The procedure initializes B{0,1}r×mB \in \{0,1\}^{r\times m}3, repeatedly selects a seed associated with uncovered B{0,1}r×mB \in \{0,1\}^{r\times m}4s, computes its closure to form a concept B{0,1}r×mB \in \{0,1\}^{r\times m}5, marks B{0,1}r×mB \in \{0,1\}^{r\times m}6 as covered, and then builds B{0,1}r×mB \in \{0,1\}^{r\times m}7 and B{0,1}r×mB \in \{0,1\}^{r\times m}8. The worst-case time is B{0,1}r×mB \in \{0,1\}^{r\times m}9, with (AB)ij=k=1r(AikBkj),(A \odot B)_{ij} = \bigvee_{k=1}^r (A_{ik} \wedge B_{kj}),0 the number of factors. Model complexity is effectively regularized through coverage targets such as (AB)ij=k=1r(AikBkj),(A \odot B)_{ij} = \bigvee_{k=1}^r (A_{ik} \wedge B_{kj}),1 or (AB)ij=k=1r(AikBkj),(A \odot B)_{ij} = \bigvee_{k=1}^r (A_{ik} \wedge B_{kj}),2.

The recommender itself does not use the Boolean product to generate numeric predictions. Instead, the user–factor matrix (AB)ij=k=1r(AikBkj),(A \odot B)_{ij} = \bigvee_{k=1}^r (A_{ik} \wedge B_{kj}),3 is used to compute cosine similarities, and a standard memory-based (AB)ij=k=1r(AikBkj),(A \odot B)_{ij} = \bigvee_{k=1}^r (A_{ik} \wedge B_{kj}),4NN predictor aggregates neighbors’ numeric ratings: (AB)ij=k=1r(AikBkj),(A \odot B)_{ij} = \bigvee_{k=1}^r (A_{ik} \wedge B_{kj}),5 Ratings (AB)ij=k=1r(AikBkj),(A \odot B)_{ij} = \bigvee_{k=1}^r (A_{ik} \wedge B_{kj}),6 are binarized by thresholds (AB)ij=k=1r(AikBkj),(A \odot B)_{ij} = \bigvee_{k=1}^r (A_{ik} \wedge B_{kj}),7, (AB)ij=k=1r(AikBkj),(A \odot B)_{ij} = \bigvee_{k=1}^r (A_{ik} \wedge B_{kj}),8, (AB)ij=k=1r(AikBkj),(A \odot B)_{ij} = \bigvee_{k=1}^r (A_{ik} \wedge B_{kj}),9, or (PQ)(P \circ Q)0. Evaluation uses

(PQ)(P \circ Q)1

with (PQ)(P \circ Q)2 the 20,000 held-out ratings in the MovieLens split.

On MovieLens-100K, after retaining users with (PQ)(P \circ Q)3 ratings, the matrix has size (PQ)(P \circ Q)4, with 80,000 train and 20,000 test ratings. At (PQ)(P \circ Q)5, the number of factors is (PQ)(P \circ Q)6 for SVD and (PQ)(P \circ Q)7 for BMF. At (PQ)(P \circ Q)8 coverage and numbers of neighbors (PQ)(P \circ Q)9, the reported MAE values are

K=(G,M,I)\mathcal{K}=(G,M,I)0

K=(G,M,I)\mathcal{K}=(G,M,I)1

K=(G,M,I)\mathcal{K}=(G,M,I)2

The reported consequence is that binary-scaled rating data can yield MAE nearly the same as SVD on non-scaled data. Tighter thresholds K=(G,M,I)\mathcal{K}=(G,M,I)3 and K=(G,M,I)\mathcal{K}=(G,M,I)4 tend to improve MAE relative to K=(G,M,I)\mathcal{K}=(G,M,I)5 or K=(G,M,I)\mathcal{K}=(G,M,I)6. When recommending a fixed number of items such as 20, BMF-based recommendations are slightly worse than using the full rating matrix directly, and using only the ratings covered by BMF for similarity computation increases MAE, whereas using K=(G,M,I)\mathcal{K}=(G,M,I)7 yields better results.

3. The Python package “bfact” for low-rank Boolean matrix factorisation

A second, more recent use makes “bfact” the proper name of a Python package for accurate low-rank Boolean matrix factorisation. The package operates on an observed binary matrix K=(G,M,I)\mathcal{K}=(G,M,I)8, with Boolean product

K=(G,M,I)\mathcal{K}=(G,M,I)9

and uses either reconstruction error

GG0

or complexity-based MDL-like objectives. Its defining structural bias is a disjoint column-set factorisation in which

GG1

The package therefore begins by encouraging disjointness and only later relaxes it during refinement (Visscher et al., 7 Sep 2025).

The first stage is a hybrid restricted master problem. Candidate feature-sets are generated by hierarchical clustering on pairwise Hamming distances of features and by Leiden community detection at multiple resolutions on a GG2-NN graph of features. The union of these cluster-derived feature sets forms a candidate pool GG3, typically with GG4. The restricted master problem GG5 then selects up to GG6 factors through a small MILP with factor-selection variables GG7 and feature slack variables GG8. The coverage inequality

GG9

ensures every feature is either assigned to a selected candidate or absorbed by slack, while equality would yield an exact disjoint cover. Delayed column generation was attempted, but the pricing subproblem was found too slow to improve the bound in time; in practice, MM0 alone gives high-quality disjoint scaffolds.

Two second-stage recovery options are provided. In the combinatorial route, MM1 uses the unique observation-membership patterns discovered by MM2, introduces variables MM3, MM4, and MM5, and penalizes false negatives, false positives, and complexity through MM6 and MM7. This stage was tractable for matrices up to MM8 and MM9, but required heavy memory of about IG×MI \subseteq G \times M0 GB. In the heuristic route, the algorithm starts from IG×MI \subseteq G \times M1 and IG×MI \subseteq G \times M2, reassigns features by majority support within each observation group using a threshold grid IG×MI \subseteq G \times M3, evaluates either reconstruction error or PRIMP’s code-table MDL cost IG×MI \subseteq G \times M4, and then greedily prunes redundant factors. For reconstruction error, the pruning tolerance IG×MI \subseteq G \times M5 is chosen close to IG×MI \subseteq G \times M6, with IG×MI \subseteq G \times M7 used on simulations and a real-data rule of thumb

IG×MI \subseteq G \times M8

truncated to three decimals.

Rank estimation is built into the package. The pipeline is run over IG×MI \subseteq G \times M9, with early stopping if the metric fails to improve for I{0,1}n×mI \in \{0,1\}^{n\times m}00 successive rank increments. The exposed pipelines are bfact-recon, bfact-MDL, and bfact-MIP. Candidate generation has worst-case cost I{0,1}n×mI \in \{0,1\}^{n\times m}01 for pairwise Hamming distances; I{0,1}n×mI \in \{0,1\}^{n\times m}02 scales with approximately I{0,1}n×mI \in \{0,1\}^{n\times m}03 in variables and constraints; heuristic refinement is dominated by computing I{0,1}n×mI \in \{0,1\}^{n\times m}04 and scanning thresholds.

Empirically, the three variants achieve similar F1 scores on the ground-truth signal I{0,1}n×mI \in \{0,1\}^{n\times m}05 in simulations and recover ranks accurately. PANDA consistently overestimates rank and has lower F1; PRIMP struggles when I{0,1}n×mI \in \{0,1\}^{n\times m}06 is dense; MDL4BMF estimates rank well but yields lower F1 and is much slower. On UCI Chess, Mushroom, and MovieLens 10M with binarised ratings, the bfact variants match or surpass baselines in reconstruction or F1 while using far fewer factors. On Human Lung Cell Atlas single-cell RNA-sequencing data, comprising 14 datasets with sizes up to I{0,1}n×mI \in \{0,1\}^{n\times m}07 cells I{0,1}n×mI \in \{0,1\}^{n\times m}08 genes and density I{0,1}n×mI \in \{0,1\}^{n\times m}09–I{0,1}n×mI \in \{0,1\}^{n\times m}10, bfact-recon and bfact-MDL recover strong signals at substantially lower ranks than competing methods. The package description therefore associates “bfact” specifically with a disjoint-first, warm-started, hybrid combinatorial optimisation strategy for BMF.

4. Bayesian factor analysis and matrix-valued generalizations

In another major usage, “bfact” refers to Bayesian factor analysis rather than Boolean matrix factorisation. The basic single-study factor model is

I{0,1}n×mI \in \{0,1\}^{n\times m}11

with I{0,1}n×mI \in \{0,1\}^{n\times m}12, loadings I{0,1}n×mI \in \{0,1\}^{n\times m}13, factor scores I{0,1}n×mI \in \{0,1\}^{n\times m}14, and diagonal noise I{0,1}n×mI \in \{0,1\}^{n\times m}15. With I{0,1}n×mI \in \{0,1\}^{n\times m}16 and I{0,1}n×mI \in \{0,1\}^{n\times m}17, the marginal covariance is I{0,1}n×mI \in \{0,1\}^{n\times m}18. The model is non-identifiable up to orthogonal rotations, so identifiability is handled by structural constraints such as lower-triangular I{0,1}n×mI \in \{0,1\}^{n\times m}19 with positive diagonal, heteroscedastic factors, or post hoc alignment by orthogonal Procrustes, varimax, or spectral decomposition. In the multi-study setting, one writes

I{0,1}n×mI \in \{0,1\}^{n\times m}20

and the recent tutorial organizes five advanced Bayesian integrative factor models around this decomposition: Perturbed Factor Analysis (PFA), Bayesian Factor Regression with non-local spike-and-slab priors (MOM-SS), Subspace Factor Analysis (SUFA), Bayesian Multi-study Factor Analysis (BMSFA), and Bayesian Combinatorial Multi-study Factor Analysis (Tetris), with Stack FA and Ind FA as benchmarks. The priors include MGPS shrinkage, Dirichlet–Laplace shrinkage, non-local spike-and-slab slabs, and Indian Buffet Process priors; the inference schemes include Gibbs, EM with coordinate descent, hybrid Gibbs–HMC, and Metropolis-within-Gibbs (Liang et al., 23 Jun 2025).

The tutorial’s comparative picture is explicitly application-driven. In five simulation scenarios, no method is uniformly best. MOM-SS is the fastest because it uses EM. SUFA is efficient when I{0,1}n×mI \in \{0,1\}^{n\times m}21 is small, but runtime grows with I{0,1}n×mI \in \{0,1\}^{n\times m}22. BMSFA is moderate and robust but slower. Tetris and PFA are the slowest at high I{0,1}n×mI \in \{0,1\}^{n\times m}23, while Tetris_fixT considerably reduces runtime. The guidance is correspondingly conditional: unknown or partial sharing suggests Tetris, clearly shared plus study-specific structure suggests BMSFA, study-specific structure inside the shared loading subspace suggests SUFA, and perturbation alignment suggests PFA. In the nutrition application with six studies, I{0,1}n×mI \in \{0,1\}^{n\times m}24, and I{0,1}n×mI \in \{0,1\}^{n\times m}25, typical choices are I{0,1}n×mI \in \{0,1\}^{n\times m}26–I{0,1}n×mI \in \{0,1\}^{n\times m}27 and I{0,1}n×mI \in \{0,1\}^{n\times m}28–I{0,1}n×mI \in \{0,1\}^{n\times m}29, with Tetris giving the lowest MSE but clear overfitting and SUFA and BMSFA balancing error and interpretability. In the genomics application with four studies, I{0,1}n×mI \in \{0,1\}^{n\times m}30, and I{0,1}n×mI \in \{0,1\}^{n\times m}31, SUFA and BMSFA achieved the lowest MSE of about I{0,1}n×mI \in \{0,1\}^{n\times m}32, while Tetris did not finish within 5 days.

A matrix-valued robust extension is the matrix-variate I{0,1}n×mI \in \{0,1\}^{n\times m}33-based bilinear factor analysis model, or I{0,1}n×mI \in \{0,1\}^{n\times m}34-BFA. It preserves row and column structure by writing

I{0,1}n×mI \in \{0,1\}^{n\times m}35

with I{0,1}n×mI \in \{0,1\}^{n\times m}36 and I{0,1}n×mI \in \{0,1\}^{n\times m}37, and models each matrix observation through a matrix-variate I{0,1}n×mI \in \{0,1\}^{n\times m}38 distribution obtained from a scale-mixture representation with I{0,1}n×mI \in \{0,1\}^{n\times m}39. This yields robust sample weights

I{0,1}n×mI \in \{0,1\}^{n\times m}40

and supports ECME, AECM, PX-ECME, and PX-AECM maximum-likelihood algorithms. The model provides a closed-form Fisher information matrix, posterior factor scores

I{0,1}n×mI \in \{0,1\}^{n\times m}41

and a higher empirical breakdown behavior than vectorized I{0,1}n×mI \in \{0,1\}^{n\times m}42-FA, which deteriorates because vectorization inflates the effective dimension I{0,1}n×mI \in \{0,1\}^{n\times m}43 (Ma et al., 2024).

A deeper nonparametric variant places a beta process prior on binary latent features and feeds the binary code I{0,1}n×mI \in \{0,1\}^{n\times m}44 through a nonlinear neural network I{0,1}n×mI \in \{0,1\}^{n\times m}45 to parameterize a continuous latent I{0,1}n×mI \in \{0,1\}^{n\times m}46, with a linear observation model

I{0,1}n×mI \in \{0,1\}^{n\times m}47

In the specific instantiation described, I{0,1}n×mI \in \{0,1\}^{n\times m}48 is Dirichlet-distributed, I{0,1}n×mI \in \{0,1\}^{n\times m}49, and inference uses a stochastic MAP-EM procedure with Beta–Bernoulli conjugacy for I{0,1}n×mI \in \{0,1\}^{n\times m}50, Gaussian conjugacy for I{0,1}n×mI \in \{0,1\}^{n\times m}51, greedy sparse coding for I{0,1}n×mI \in \{0,1\}^{n\times m}52, and ADAM updates for I{0,1}n×mI \in \{0,1\}^{n\times m}53. On MNIST, the reported setup uses a 3-layer network with 100 hidden units, I{0,1}n×mI \in \{0,1\}^{n\times m}54, truncation I{0,1}n×mI \in \{0,1\}^{n\times m}55, I{0,1}n×mI \in \{0,1\}^{n\times m}56, I{0,1}n×mI \in \{0,1\}^{n\times m}57, I{0,1}n×mI \in \{0,1\}^{n\times m}58, I{0,1}n×mI \in \{0,1\}^{n\times m}59, ADAM stepsize I{0,1}n×mI \in \{0,1\}^{n\times m}60, I{0,1}n×mI \in \{0,1\}^{n\times m}61, I{0,1}n×mI \in \{0,1\}^{n\times m}62, batch size I{0,1}n×mI \in \{0,1\}^{n\times m}63, and 10,000 iterations (Mittal et al., 2020).

5. Bayes-factor-based meanings of “bfact”

A distinct family of uses reads “bfact” as Bayes factor rather than factor analysis. The canonical Bayes factor is

I{0,1}n×mI \in \{0,1\}^{n\times m}64

In matrix-variate online monitoring, the predictive Bayes factor at time I{0,1}n×mI \in \{0,1\}^{n\times m}65 is

I{0,1}n×mI \in \{0,1\}^{n\times m}66

where the alternative predictive is generated through power discounting,

I{0,1}n×mI \in \{0,1\}^{n\times m}67

For matrix-normal and matrix-variate I{0,1}n×mI \in \{0,1\}^{n\times m}68 models, the paper derives closed-form predictives, closed-form I{0,1}n×mI \in \{0,1\}^{n\times m}69, an upper bound I{0,1}n×mI \in \{0,1\}^{n\times m}70, a geometric ellipsoid acceptance region, and three variability-aware robust versions: the minimum BF, the integrated BF, and the normalized integrated BF. The sequential algorithm is analytic and requires no MCMC. In simulations with I{0,1}n×mI \in \{0,1\}^{n\times m}71 and I{0,1}n×mI \in \{0,1\}^{n\times m}72, calibrated thresholds produced about I{0,1}n×mI \in \{0,1\}^{n\times m}73 false positives under I{0,1}n×mI \in \{0,1\}^{n\times m}74, while power increased with the outlier magnitude I{0,1}n×mI \in \{0,1\}^{n\times m}75 and the number of contaminated entries. Applications to EU macroeconomics, an international trade network, and a volatility network flag the COVID-19 outbreak, the start of the Ukraine conflict, and abrupt volatility-regime changes as outlying periods (Billio et al., 25 Mar 2025).

Bayes factors also underpin a robust multiple-testing framework for eQTL discovery. In the hierarchical mixture setup with latent indicators I{0,1}n×mI \in \{0,1\}^{n\times m}76, the posterior non-null probability is

I{0,1}n×mI \in \{0,1\}^{n\times m}77

and the robust procedure replaces I{0,1}n×mI \in \{0,1\}^{n\times m}78 with a conservative upper-bound estimator I{0,1}n×mI \in \{0,1\}^{n\times m}79. Two such estimators are provided. The EBF estimator sorts Bayes factors ascending, finds

I{0,1}n×mI \in \{0,1\}^{n\times m}80

and sets I{0,1}n×mI \in \{0,1\}^{n\times m}81. The QBF estimator uses null quantiles I{0,1}n×mI \in \{0,1\}^{n\times m}82 and

I{0,1}n×mI \in \{0,1\}^{n\times m}83

Ranking by the calibrated posterior I{0,1}n×mI \in \{0,1\}^{n\times m}84 is equivalent to ranking by I{0,1}n×mI \in \{0,1\}^{n\times m}85, and the step-up rule rejects the largest set whose average conservative local false discovery rate is at most I{0,1}n×mI \in \{0,1\}^{n\times m}86. In a single-tissue simulation with I{0,1}n×mI \in \{0,1\}^{n\times m}87, 10,000 genes, and I{0,1}n×mI \in \{0,1\}^{n\times m}88–I{0,1}n×mI \in \{0,1\}^{n\times m}89 cis-SNPs per gene, EBF took I{0,1}n×mI \in \{0,1\}^{n\times m}90, QBF with 100 permutations took I{0,1}n×mI \in \{0,1\}^{n\times m}91, BH or Storey with 500 permutations took I{0,1}n×mI \in \{0,1\}^{n\times m}92, and BH or Storey with 5000 permutations took I{0,1}n×mI \in \{0,1\}^{n\times m}93. In the multi-tissue Dimas et al. dataset with 5,012 genes and three tissues, QBF found 1,002 eGenes at I{0,1}n×mI \in \{0,1\}^{n\times m}94 FDR, EBF found 927, Storey on permutation I{0,1}n×mI \in \{0,1\}^{n\times m}95-values of BF found 1,012, and the naive min-I{0,1}n×mI \in \{0,1\}^{n\times m}96 test with Storey found 627 (Wen, 2013).

A reporting-oriented generalization is the Bayes factor function (BFF), defined directly from a test statistic I{0,1}n×mI \in \{0,1\}^{n\times m}97 with noncentrality parameter I{0,1}n×mI \in \{0,1\}^{n\times m}98,

I{0,1}n×mI \in \{0,1\}^{n\times m}99

The framework supplies closed-form or standard-library-evaluable BFFs for A{0,1}n×rA \in \{0,1\}^{n\times r}00, A{0,1}n×rA \in \{0,1\}^{n\times r}01, A{0,1}n×rA \in \{0,1\}^{n\times r}02, and A{0,1}n×rA \in \{0,1\}^{n\times r}03 statistics, and maps A{0,1}n×rA \in \{0,1\}^{n\times r}04 to standardized effect sizes such as Cohen’s A{0,1}n×rA \in \{0,1\}^{n\times r}05 or RMSES. For a one-sample A{0,1}n×rA \in \{0,1\}^{n\times r}06-test,

A{0,1}n×rA \in \{0,1\}^{n\times r}07

For independent studies, aggregation is multiplicative,

A{0,1}n×rA \in \{0,1\}^{n\times r}08

or additive on the log scale. The proposed interpretation is evidential rather than dichotomous: curves of BFF versus effect size replace single-threshold significance reporting and are designed to summarize how evidence varies across plausible effect magnitudes (Johnson et al., 2022).

6. B Factories as an unrelated high-energy-physics usage

In a fully separate literature, “B Factories” refers to the asymmetric-energy A{0,1}n×rA \in \{0,1\}^{n\times r}09 colliders PEP-II at SLAC and KEKB at KEK, together with the BABAR and Belle detectors, built to test the Cabibbo–Kobayashi–Maskawa mechanism of quark mixing and CP violation. The experiments operated near the A{0,1}n×rA \in \{0,1\}^{n\times r}10 resonance to produce coherent A{0,1}n×rA \in \{0,1\}^{n\times r}11 pairs. The asymmetric boosts were A{0,1}n×rA \in \{0,1\}^{n\times r}12 for PEP-II and A{0,1}n×rA \in \{0,1\}^{n\times r}13 for KEKB, enabling time-difference measurements through

A{0,1}n×rA \in \{0,1\}^{n\times r}14

Integrated luminosities at the A{0,1}n×rA \in \{0,1\}^{n\times r}15 were A{0,1}n×rA \in \{0,1\}^{n\times r}16 for BABAR and A{0,1}n×rA \in \{0,1\}^{n\times r}17 for Belle, for a combined A{0,1}n×rA \in \{0,1\}^{n\times r}18 and more than A{0,1}n×rA \in \{0,1\}^{n\times r}19 A{0,1}n×rA \in \{0,1\}^{n\times r}20 pairs (Bevan, 2012).

The flagship observable was time-dependent CP violation in A{0,1}n×rA \in \{0,1\}^{n\times r}21 and related channels. For a CP eigenstate A{0,1}n×rA \in \{0,1\}^{n\times r}22,

A{0,1}n×rA \in \{0,1\}^{n\times r}23

with

A{0,1}n×rA \in \{0,1\}^{n\times r}24

For tree-dominated A{0,1}n×rA \in \{0,1\}^{n\times r}25 modes such as A{0,1}n×rA \in \{0,1\}^{n\times r}26, one has A{0,1}n×rA \in \{0,1\}^{n\times r}27, A{0,1}n×rA \in \{0,1\}^{n\times r}28, and A{0,1}n×rA \in \{0,1\}^{n\times r}29 up to the CP-eigenvalue sign. The combined B-factory average quoted for A{0,1}n×rA \in \{0,1\}^{n\times r}30 is A{0,1}n×rA \in \{0,1\}^{n\times r}31, corresponding to A{0,1}n×rA \in \{0,1\}^{n\times r}32 with sub-degree precision. Additional headline results include A{0,1}n×rA \in \{0,1\}^{n\times r}33, A{0,1}n×rA \in \{0,1\}^{n\times r}34 from A{0,1}n×rA \in \{0,1\}^{n\times r}35, and A{0,1}n×rA \in \{0,1\}^{n\times r}36 from combined GLW, ADS, and GGSZ analyses.

The experiments also constrained loop-level flavor dynamics and possible new-physics contributions. The effective Hamiltonian for A{0,1}n×rA \in \{0,1\}^{n\times r}37 transitions is

A{0,1}n×rA \in \{0,1\}^{n\times r}38

with A{0,1}n×rA \in \{0,1\}^{n\times r}39, A{0,1}n×rA \in \{0,1\}^{n\times r}40, and A{0,1}n×rA \in \{0,1\}^{n\times r}41 among the key operators. The inclusive branching fraction A{0,1}n×rA \in \{0,1\}^{n\times r}42 is reported near A{0,1}n×rA \in \{0,1\}^{n\times r}43–A{0,1}n×rA \in \{0,1\}^{n\times r}44, strongly constraining A{0,1}n×rA \in \{0,1\}^{n\times r}45 and charged-Higgs scenarios. BABAR and Belle also pioneered A{0,1}n×rA \in \{0,1\}^{n\times r}46 and A{0,1}n×rA \in \{0,1\}^{n\times r}47 measurements, contributed to charm-mixing observations, and reshaped heavy-hadron spectroscopy through discoveries such as A{0,1}n×rA \in \{0,1\}^{n\times r}48, A{0,1}n×rA \in \{0,1\}^{n\times r}49, and A{0,1}n×rA \in \{0,1\}^{n\times r}50. In this usage, therefore, “bfact” has no relation to factor models or Bayes factors; it abbreviates a major experimental program in flavor physics.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to bfact.