bfact: Boolean & Bayesian Factor Analysis
- bfact is an overloaded term that encompasses Boolean matrix factorisation, Bayesian factor analysis, and Bayes factor methods, as well as B factory physics.
- In Boolean matrix factorisation, bfact refers to a Python package that employs disjoint factor selection and hybrid combinatorial optimization to decompose binary data effectively.
- Within Bayesian contexts, bfact describes advanced factor models that integrate robust outlier detection and multiple hypothesis testing using Bayes factors, aiding applications in genomics and econometrics.
to=arxiv_search.search 娱乐开号 彩神争霸苹果"query":"bfact","max_results":10,"sort_by":"relevance"}{"result":"Search results for 'bfact' (max 10):\n\n1. (Visscher et al., 7 Sep 2025) - Hybrid restricted master problem for Boolean matrix factorisation\n Authors: Michael S. S. Law, Benjamin P. C. Tomlinson, Guido Sanguinetti\n Published: 2025-09-07\n Categories: cs.LG, cs.DS, math.OC\n Abstract: We present bfact, a Python package for performing accurate low-rank Boolean matrix factorisation (BMF). bfact uses a hybrid combinatorial optimisation approach based on a priori candidate factors generated from clustering algorithms. It selects the best disjoint factors before performing either a second combinatorial or heuristic algorithm to recover the BMF. We show that bfact does particularly well at estimating the true rank of matrices in simulated settings. In real benchmarks, using a collation of single-cell RNA-sequencing datasets from the Human Lung Cell Atlas, we show that bfact achieves strong signal recovery, with a much lower rank.\n\n2. (Liang et al., 23 Jun 2025) - Bayesian integrative factor analysis methods, with application in nutrition and genomics data\n Authors: Xinyue Liang, Wenyi Wang, Gen Li\n Published: 2025-06-23\n Categories: stat.ME, stat.AP, q-bio.GN\n Abstract: High-dimensional data are crucial in biomedical research. Integrating such data from multiple studies is a critical process that relies on the choice of advanced statistical models, enhancing statistical power, reproducibility, and scientific insight compared to analyzing each study separately. Factor analysis (FA) is a core dimensionality reduction technique that models observed data through a small set of latent factors. Bayesian extensions of FA have recently emerged as powerful tools for multi-study integration, enabling researchers to disentangle shared biological signals from study-specific variability. In this tutorial, we provide a practical and comparative guide to five advanced Bayesian integrative factor models: Perturbed Factor Analysis (PFA), Bayesian Factor Regression with non-local spike-and-slab priors (MOM-SS), Subspace Factor Analysis (SUFA), Bayesian Multi-study Factor Analysis (BMSFA), and Bayesian Combinatorial Multi-study Factor Analysis (Tetris). To contextualize these methods, we also include two benchmark approaches: standard FA applied to pooled data (Stack FA) and FA applied separately to each study (Ind FA). We evaluate all methods through extensive simulations, assessing computational efficiency and accuracy in the estimation of loadings and number of factors. To bridge theory and practice, we present a full analytical workflow, with detailed R code, demonstrating how to apply these models to real-world datasets in nutrition and genomics.\n\n3. (Billio et al., 25 Mar 2025) - Bayesian Outlier Detection for Matrix-variate Models\n Authors: Alessio Bissiri, Luca Casarin, Roberto Golinelli, Francesco Ravazzolo, Fabio Rigat\n Published: 2025-03-25\n Categories: stat.ME, econ.EM, stat.CO\n Abstract: Bayes Factor (BF) is one of the tools used in Bayesian analysis for model selection. The predictive BF finds application in detecting outliers, which are relevant sources of estimation and forecast errors. An efficient framework for outlier detection is provided and purposely designed for large multidimensional datasets. Online detection and analytical tractability guarantee the procedure's efficiency. The proposed sequential Bayesian monitoring extends the univariate setup to a matrix--variate one. Prior perturbation based on power discounting is applied to obtain tractable predictive BFs. This way, computationally intensive procedures used in Bayesian Analysis are not required. The conditions leading to inconclusive responses in outlier identification are derived, and some robust approaches are proposed that exploit the predictive BF's variability to improve the standard discounting method. The effectiveness of the procedure is studied using simulated data. An illustration is provided through applications to relevant benchmark datasets from macroeconomics and finance.\n\n4. (Ma et al., 2024) - Robust bilinear factor analysis based on the matrix-variate t distribution\n Authors: Xuejun Pan, Jiahui Ai, Fuping Huang, Song Xi Chen, Shaobo Jin\n Published: 2024-01-04\n Categories: stat.ME\n Abstract: Factor Analysis based on multivariate t distribution (tfa) is a useful robust tool for extracting common factors on heavy-tailed or contaminated data. However, tfa is only applicable to vector data. When tfa is applied to matrix data, it is common to first vectorize the matrix observations. This introduces two challenges for tfa: (i) the inherent matrix structure of the data is broken, and (ii) robustness may be lost, as vectorized matrix data typically results in a high data dimension, which could easily lead to the breakdown of tfa. To address these issues, starting from the intrinsic matrix structure of matrix data, a novel robust factor analysis model, namely bilinear factor analysis built on the matrix-variate t distribution (tbfa), is proposed in this paper. The novelty is that it is capable to simultaneously extract common factors for both row and column variables of interest on heavy-tailed or contaminated matrix data. Two efficient algorithms for maximum likelihood estimation of tbfa are developed. Closed-form expression for the Fisher information matrix to calculate the accuracy of parameter estimates are derived. Empirical studies are conducted to understand the proposed tbfa model and compare with related competitors. The results demonstrate the superiority and practicality of tbfa. Importantly, tbfa exhibits a significantly higher breakdown point than tfa, making it more suitable for matrix data.\n\n5. (Mittal et al., 2020) - Deep Bayesian Nonparametric Factor Analysis\n Authors: Erik A. Platanios, Michael I. Jordan, et al.\n Published: 2020-11-09\n Categories: stat.ML, cs.LG\n Abstract: We propose a deep generative factor analysis model with beta process prior that can approximate complex non-factorial distributions over the latent codes. We outline a stochastic EM algorithm for scalable inference in a specific instantiation of this model and present some preliminary results.\n\n6. (Wen, 2013) - Robust Bayesian FDR Control using Bayes Factors, with Applications to Multi-tissue eQTL Discovery\n Authors: Xiaoquan Wen\n Published: 2013-11-15\n Categories: stat.AP, q-bio.GN\n Abstract: Motivated by the genomic application of expression quantitative trait loci (eQTL) mapping, we propose a new procedure to perform simultaneous testing of multiple hypotheses using Bayes factors as input test statistics. One of the most significant features of this method is its robustness in controlling the targeted false discovery rate (FDR) even under misspecifications of parametric alternative models. Moreover, the proposed procedure is highly computationally efficient, which is ideal for treating both complex system and big data in genomic applications. We discuss the theoretical properties of the new procedure and demonstrate its power and computational efficiency in applications of single-tissue and multi-tissue eQTL mapping.\n\n7. (Johnson et al., 2022) - Bayes factor functions for reporting outcomes of hypothesis tests\n Authors: Donald B. Johnson, et al.\n Published: 2022-09-30\n Categories: stat.ME, math.ST\n Abstract: Bayes factors represent the ratio of probabilities assigned to data by competing scientific hypotheses. Drawbacks of Bayes factors are their dependence on prior specifications that define null and alternative hypotheses and difficulties encountered in their computation. To address these problems, we define Bayes factor functions (BFF) directly from common test statistics. BFFs depend on a single non-centrality parameter that can be expressed as a function of standardized effect sizes, and plots of BFFs versus effect size provide informative summaries of hypothesis tests that can be easily aggregated across studies. Such summaries eliminate the need for arbitrary P-value thresholds to define ``statistical significance.'' BFFs are available in closed form and can be computed easily from z, t, chi-squared, and F statistics.\n\n8. (Nenova et al., 2013) - An FCA-based Boolean Matrix Factorisation for Collaborative Filtering\n Authors: Y. Y. Yevtushenko, D. M. Kourie, G. J. J. Oosthuizen, S. O. Kuznetsov, R. Kruger\n Published: 2013-10-16\n Categories: cs.IR\n Abstract: We propose a new approach for Collaborative Filtering which is based on Boolean Matrix Factorisation (BMF) and Formal Concept Analysis. In a series of experiments on real data (Movielens dataset) we compare the approach with the SVD- and NMF-based algorithms in terms of Mean Average Error (MAE). One of the experimental consequences is that it is enough to have a binary-scaled rating data to obtain almost the same quality in terms of MAE by BMF than for the SVD-based algorithm in case of non-scaled data.\n\n9. (Bevan, 2012) - B Factories\n Authors: Adrian Bevan, et al.\n Published: 2012-02-03\n Categories: hep-ex, hep-ph\n Abstract: The B factories were constructed with a very specific purpose in mind: to test the Cabibbo-Kobayashi-Maskawa description of quark mixing and CP violation in the Standard Model of particle physics. The goals of testing this part of the Standard Model were achieved, and have been surpassed beyond all expectation. As a result the B factories have revolutionised our understanding of many areas of the Standard Model of particle physics, and also provide a number of stringent limits on possible scenarios of physics beyond the Standard Model. In some cases these limits on new physics effects equal or surpass those achievable at the CERN based Large Hadron Collider."} “bfact” is a context-dependent label rather than a single standardized technical term in arXiv literature. It appears in work on Boolean matrix factorisation, especially as the name of a Python package for low-rank Boolean decompositions; in Bayesian factor-analysis literature, including multi-study, robust, and deep nonparametric models; in Bayes-factor-based statistical inference; and, in an unrelated high-energy-physics usage, in connection with B factories. The term therefore denotes a family of abbreviation-driven usages whose meaning is determined by domain and surrounding notation rather than by a unique definition (Visscher et al., 7 Sep 2025, Liang et al., 23 Jun 2025, Billio et al., 25 Mar 2025, Bevan, 2012).
1. Terminological scope and domain-specific meanings
In current arXiv usage, “bfact” is best treated as an overloaded label. The main technical senses represented in the literature are summarized below.
| Usage | Core object | Representative paper |
|---|---|---|
| Boolean matrix factorisation | Binary factor models and Boolean products | (Nenova et al., 2013, Visscher et al., 7 Sep 2025) |
| Bayesian factor analysis | Latent-factor models with Bayesian priors and inference | (Liang et al., 23 Jun 2025, Ma et al., 2024, Mittal et al., 2020) |
| Bayes-factor methods | Model comparison, multiple testing, and reporting of evidence | (Billio et al., 25 Mar 2025, Wen, 2013, Johnson et al., 2022) |
| B Factories | Asymmetric-energy facilities for flavor physics | (Bevan, 2012) |
This distribution of meanings has methodological consequences. In Boolean matrix factorisation, the central object is a binary matrix and the key algebra is logical OR over logical AND. In Bayesian factor analysis, the central object is a latent-factor decomposition with priors over loadings, factor scores, and variance components. In Bayes-factor usage, the core quantity is a ratio of marginal or predictive probabilities. In B-factory physics, the phrase refers to accelerator-detector complexes constructed to test the Cabibbo–Kobayashi–Maskawa description of quark mixing and CP violation.
2. Boolean matrix factorisation in collaborative filtering
A major “bfact” usage concerns Boolean matrix factorisation for recommender systems. In the FCA-based collaborative-filtering formulation, a binary matrix is decomposed into a Boolean product of binary matrices and ,
with the paper using for the same semantics. This differs from SVD or NMF, which use real-valued arithmetic and additive aggregation. The Boolean setting yields discrete biclusters rather than continuous latent components, and the decomposition is non-linear under Boolean algebra (Nenova et al., 2013).
The formal-concept layer is supplied by Formal Concept Analysis. A formal context consists of objects , attributes , and a binary relation . For 0 and 1, the derivation operators are
2
A formal concept is a pair 3 such that 4 and 5, where 6 is the extent and 7 is the intent. Given a subset of concepts 8, one constructs indicator matrices 9 and 0, and 1 reproduces the covered part of 2. The theoretical basis is the universality statement that for every binary matrix 3, there exists 4 such that 5, and the optimality statement that if 6 for binary 7, then there exists 8 with 9 such that 0.
Algorithmically, the approach uses Algorithm 2 from Belohlavek & Vychodil (2010), which avoids computing the entire concept lattice and instead incrementally discovers closed biclusters that increase coverage. Inputs are a binary-scaled user–item matrix and either a target coverage 1 or maximal number of factors 2. The procedure initializes 3, repeatedly selects a seed associated with uncovered 4s, computes its closure to form a concept 5, marks 6 as covered, and then builds 7 and 8. The worst-case time is 9, with 0 the number of factors. Model complexity is effectively regularized through coverage targets such as 1 or 2.
The recommender itself does not use the Boolean product to generate numeric predictions. Instead, the user–factor matrix 3 is used to compute cosine similarities, and a standard memory-based 4NN predictor aggregates neighbors’ numeric ratings: 5 Ratings 6 are binarized by thresholds 7, 8, 9, or 0. Evaluation uses
1
with 2 the 20,000 held-out ratings in the MovieLens split.
On MovieLens-100K, after retaining users with 3 ratings, the matrix has size 4, with 80,000 train and 20,000 test ratings. At 5, the number of factors is 6 for SVD and 7 for BMF. At 8 coverage and numbers of neighbors 9, the reported MAE values are
0
1
2
The reported consequence is that binary-scaled rating data can yield MAE nearly the same as SVD on non-scaled data. Tighter thresholds 3 and 4 tend to improve MAE relative to 5 or 6. When recommending a fixed number of items such as 20, BMF-based recommendations are slightly worse than using the full rating matrix directly, and using only the ratings covered by BMF for similarity computation increases MAE, whereas using 7 yields better results.
3. The Python package “bfact” for low-rank Boolean matrix factorisation
A second, more recent use makes “bfact” the proper name of a Python package for accurate low-rank Boolean matrix factorisation. The package operates on an observed binary matrix 8, with Boolean product
9
and uses either reconstruction error
0
or complexity-based MDL-like objectives. Its defining structural bias is a disjoint column-set factorisation in which
1
The package therefore begins by encouraging disjointness and only later relaxes it during refinement (Visscher et al., 7 Sep 2025).
The first stage is a hybrid restricted master problem. Candidate feature-sets are generated by hierarchical clustering on pairwise Hamming distances of features and by Leiden community detection at multiple resolutions on a 2-NN graph of features. The union of these cluster-derived feature sets forms a candidate pool 3, typically with 4. The restricted master problem 5 then selects up to 6 factors through a small MILP with factor-selection variables 7 and feature slack variables 8. The coverage inequality
9
ensures every feature is either assigned to a selected candidate or absorbed by slack, while equality would yield an exact disjoint cover. Delayed column generation was attempted, but the pricing subproblem was found too slow to improve the bound in time; in practice, 0 alone gives high-quality disjoint scaffolds.
Two second-stage recovery options are provided. In the combinatorial route, 1 uses the unique observation-membership patterns discovered by 2, introduces variables 3, 4, and 5, and penalizes false negatives, false positives, and complexity through 6 and 7. This stage was tractable for matrices up to 8 and 9, but required heavy memory of about 0 GB. In the heuristic route, the algorithm starts from 1 and 2, reassigns features by majority support within each observation group using a threshold grid 3, evaluates either reconstruction error or PRIMP’s code-table MDL cost 4, and then greedily prunes redundant factors. For reconstruction error, the pruning tolerance 5 is chosen close to 6, with 7 used on simulations and a real-data rule of thumb
8
truncated to three decimals.
Rank estimation is built into the package. The pipeline is run over 9, with early stopping if the metric fails to improve for 00 successive rank increments. The exposed pipelines are bfact-recon, bfact-MDL, and bfact-MIP. Candidate generation has worst-case cost 01 for pairwise Hamming distances; 02 scales with approximately 03 in variables and constraints; heuristic refinement is dominated by computing 04 and scanning thresholds.
Empirically, the three variants achieve similar F1 scores on the ground-truth signal 05 in simulations and recover ranks accurately. PANDA consistently overestimates rank and has lower F1; PRIMP struggles when 06 is dense; MDL4BMF estimates rank well but yields lower F1 and is much slower. On UCI Chess, Mushroom, and MovieLens 10M with binarised ratings, the bfact variants match or surpass baselines in reconstruction or F1 while using far fewer factors. On Human Lung Cell Atlas single-cell RNA-sequencing data, comprising 14 datasets with sizes up to 07 cells 08 genes and density 09–10, bfact-recon and bfact-MDL recover strong signals at substantially lower ranks than competing methods. The package description therefore associates “bfact” specifically with a disjoint-first, warm-started, hybrid combinatorial optimisation strategy for BMF.
4. Bayesian factor analysis and matrix-valued generalizations
In another major usage, “bfact” refers to Bayesian factor analysis rather than Boolean matrix factorisation. The basic single-study factor model is
11
with 12, loadings 13, factor scores 14, and diagonal noise 15. With 16 and 17, the marginal covariance is 18. The model is non-identifiable up to orthogonal rotations, so identifiability is handled by structural constraints such as lower-triangular 19 with positive diagonal, heteroscedastic factors, or post hoc alignment by orthogonal Procrustes, varimax, or spectral decomposition. In the multi-study setting, one writes
20
and the recent tutorial organizes five advanced Bayesian integrative factor models around this decomposition: Perturbed Factor Analysis (PFA), Bayesian Factor Regression with non-local spike-and-slab priors (MOM-SS), Subspace Factor Analysis (SUFA), Bayesian Multi-study Factor Analysis (BMSFA), and Bayesian Combinatorial Multi-study Factor Analysis (Tetris), with Stack FA and Ind FA as benchmarks. The priors include MGPS shrinkage, Dirichlet–Laplace shrinkage, non-local spike-and-slab slabs, and Indian Buffet Process priors; the inference schemes include Gibbs, EM with coordinate descent, hybrid Gibbs–HMC, and Metropolis-within-Gibbs (Liang et al., 23 Jun 2025).
The tutorial’s comparative picture is explicitly application-driven. In five simulation scenarios, no method is uniformly best. MOM-SS is the fastest because it uses EM. SUFA is efficient when 21 is small, but runtime grows with 22. BMSFA is moderate and robust but slower. Tetris and PFA are the slowest at high 23, while Tetris_fixT considerably reduces runtime. The guidance is correspondingly conditional: unknown or partial sharing suggests Tetris, clearly shared plus study-specific structure suggests BMSFA, study-specific structure inside the shared loading subspace suggests SUFA, and perturbation alignment suggests PFA. In the nutrition application with six studies, 24, and 25, typical choices are 26–27 and 28–29, with Tetris giving the lowest MSE but clear overfitting and SUFA and BMSFA balancing error and interpretability. In the genomics application with four studies, 30, and 31, SUFA and BMSFA achieved the lowest MSE of about 32, while Tetris did not finish within 5 days.
A matrix-valued robust extension is the matrix-variate 33-based bilinear factor analysis model, or 34-BFA. It preserves row and column structure by writing
35
with 36 and 37, and models each matrix observation through a matrix-variate 38 distribution obtained from a scale-mixture representation with 39. This yields robust sample weights
40
and supports ECME, AECM, PX-ECME, and PX-AECM maximum-likelihood algorithms. The model provides a closed-form Fisher information matrix, posterior factor scores
41
and a higher empirical breakdown behavior than vectorized 42-FA, which deteriorates because vectorization inflates the effective dimension 43 (Ma et al., 2024).
A deeper nonparametric variant places a beta process prior on binary latent features and feeds the binary code 44 through a nonlinear neural network 45 to parameterize a continuous latent 46, with a linear observation model
47
In the specific instantiation described, 48 is Dirichlet-distributed, 49, and inference uses a stochastic MAP-EM procedure with Beta–Bernoulli conjugacy for 50, Gaussian conjugacy for 51, greedy sparse coding for 52, and ADAM updates for 53. On MNIST, the reported setup uses a 3-layer network with 100 hidden units, 54, truncation 55, 56, 57, 58, 59, ADAM stepsize 60, 61, 62, batch size 63, and 10,000 iterations (Mittal et al., 2020).
5. Bayes-factor-based meanings of “bfact”
A distinct family of uses reads “bfact” as Bayes factor rather than factor analysis. The canonical Bayes factor is
64
In matrix-variate online monitoring, the predictive Bayes factor at time 65 is
66
where the alternative predictive is generated through power discounting,
67
For matrix-normal and matrix-variate 68 models, the paper derives closed-form predictives, closed-form 69, an upper bound 70, a geometric ellipsoid acceptance region, and three variability-aware robust versions: the minimum BF, the integrated BF, and the normalized integrated BF. The sequential algorithm is analytic and requires no MCMC. In simulations with 71 and 72, calibrated thresholds produced about 73 false positives under 74, while power increased with the outlier magnitude 75 and the number of contaminated entries. Applications to EU macroeconomics, an international trade network, and a volatility network flag the COVID-19 outbreak, the start of the Ukraine conflict, and abrupt volatility-regime changes as outlying periods (Billio et al., 25 Mar 2025).
Bayes factors also underpin a robust multiple-testing framework for eQTL discovery. In the hierarchical mixture setup with latent indicators 76, the posterior non-null probability is
77
and the robust procedure replaces 78 with a conservative upper-bound estimator 79. Two such estimators are provided. The EBF estimator sorts Bayes factors ascending, finds
80
and sets 81. The QBF estimator uses null quantiles 82 and
83
Ranking by the calibrated posterior 84 is equivalent to ranking by 85, and the step-up rule rejects the largest set whose average conservative local false discovery rate is at most 86. In a single-tissue simulation with 87, 10,000 genes, and 88–89 cis-SNPs per gene, EBF took 90, QBF with 100 permutations took 91, BH or Storey with 500 permutations took 92, and BH or Storey with 5000 permutations took 93. In the multi-tissue Dimas et al. dataset with 5,012 genes and three tissues, QBF found 1,002 eGenes at 94 FDR, EBF found 927, Storey on permutation 95-values of BF found 1,012, and the naive min-96 test with Storey found 627 (Wen, 2013).
A reporting-oriented generalization is the Bayes factor function (BFF), defined directly from a test statistic 97 with noncentrality parameter 98,
99
The framework supplies closed-form or standard-library-evaluable BFFs for 00, 01, 02, and 03 statistics, and maps 04 to standardized effect sizes such as Cohen’s 05 or RMSES. For a one-sample 06-test,
07
For independent studies, aggregation is multiplicative,
08
or additive on the log scale. The proposed interpretation is evidential rather than dichotomous: curves of BFF versus effect size replace single-threshold significance reporting and are designed to summarize how evidence varies across plausible effect magnitudes (Johnson et al., 2022).
6. B Factories as an unrelated high-energy-physics usage
In a fully separate literature, “B Factories” refers to the asymmetric-energy 09 colliders PEP-II at SLAC and KEKB at KEK, together with the BABAR and Belle detectors, built to test the Cabibbo–Kobayashi–Maskawa mechanism of quark mixing and CP violation. The experiments operated near the 10 resonance to produce coherent 11 pairs. The asymmetric boosts were 12 for PEP-II and 13 for KEKB, enabling time-difference measurements through
14
Integrated luminosities at the 15 were 16 for BABAR and 17 for Belle, for a combined 18 and more than 19 20 pairs (Bevan, 2012).
The flagship observable was time-dependent CP violation in 21 and related channels. For a CP eigenstate 22,
23
with
24
For tree-dominated 25 modes such as 26, one has 27, 28, and 29 up to the CP-eigenvalue sign. The combined B-factory average quoted for 30 is 31, corresponding to 32 with sub-degree precision. Additional headline results include 33, 34 from 35, and 36 from combined GLW, ADS, and GGSZ analyses.
The experiments also constrained loop-level flavor dynamics and possible new-physics contributions. The effective Hamiltonian for 37 transitions is
38
with 39, 40, and 41 among the key operators. The inclusive branching fraction 42 is reported near 43–44, strongly constraining 45 and charged-Higgs scenarios. BABAR and Belle also pioneered 46 and 47 measurements, contributed to charm-mixing observations, and reshaped heavy-hadron spectroscopy through discoveries such as 48, 49, and 50. In this usage, therefore, “bfact” has no relation to factor models or Bayes factors; it abbreviates a major experimental program in flavor physics.