Papers
Topics
Authors
Recent
Search
2000 character limit reached

Stochastic Co-Block Model (ScBM)

Updated 9 July 2026
  • ScBM is a probabilistic model that segments bipartite and directed networks into distinct sender and receiver communities using separate latent structures.
  • It leverages singular value decomposition of the expected adjacency matrix to achieve precise co-clustering and community recovery.
  • Extensions relax Bernoulli and degree homogeneity constraints, broadening its application to dynamic networks and distribution-free settings.

Searching arXiv for recent and foundational papers on stochastic co-block models. The stochastic co-block model (ScBM) is a stochastic blockmodel for binary bipartite graphs and, through the directed–bipartite equivalence, for directed unweighted networks. Its defining feature is an asymmetric latent structure: row nodes and column nodes, or equivalently senders and receivers, possess separate community memberships, and edge probabilities depend only on the corresponding row-block and column-block. In its standard formulation, the adjacency matrix has expectation Ω=ρZrPZc′\Omega = \rho Z_r P Z_c', so the expected edge intensity is block-constant over each row-community/column-community pair. This places ScBM as the directed or bipartite analogue of the classical stochastic block model while making singular-value-based co-clustering the natural inferential mechanism (Qing et al., 2021).

1. Formal specification and equivalent representations

In the bipartite formulation, one observes a binary adjacency matrix A∈{0,1}nr×ncA \in \{0,1\}^{n_r \times n_c}, where nrn_r row nodes interact with ncn_c column nodes. Row nodes are partitioned into KrK_r communities and column nodes into KcK_c communities, encoded by one-hot membership matrices

Zr∈Mnr,Kr,Zc∈Mnc,Kc,Z_r \in \mathbb{M}_{n_r,K_r}, \qquad Z_c \in \mathbb{M}_{n_c,K_c},

with exactly one $1$ per row. If row node iri_r belongs to community girg_{i_r} and column node A∈{0,1}nr×ncA \in \{0,1\}^{n_r \times n_c}0 to community A∈{0,1}nr×ncA \in \{0,1\}^{n_r \times n_c}1, then ScBM assumes conditionally independent Bernoulli edges

A∈{0,1}nr×ncA \in \{0,1\}^{n_r \times n_c}2

where A∈{0,1}nr×ncA \in \{0,1\}^{n_r \times n_c}3 is the block probability matrix and A∈{0,1}nr×ncA \in \{0,1\}^{n_r \times n_c}4 is a sparsity parameter. Equivalently,

A∈{0,1}nr×ncA \in \{0,1\}^{n_r \times n_c}5

The directed-network formulation uses a single node set but assigns each node two latent labels: a sender label A∈{0,1}nr×ncA \in \{0,1\}^{n_r \times n_c}6 and a receiver label A∈{0,1}nr×ncA \in \{0,1\}^{n_r \times n_c}7. For a directed adjacency matrix A∈{0,1}nr×ncA \in \{0,1\}^{n_r \times n_c}8 with A∈{0,1}nr×ncA \in \{0,1\}^{n_r \times n_c}9,

nrn_r0

with nrn_r1. This makes ScBM explicitly asymmetric: the number and composition of sender communities may differ from those of receiver communities, and the block matrix need not be square or symmetric (Qing, 20 Aug 2025).

A recurring interpretation is that ScBM is not merely a directed SBM with nuisance asymmetry, but a model in which outgoing and incoming relational roles are structurally distinct. That distinction is the reason singular vectors, rather than eigenvectors of a symmetric matrix, occupy a central place in ScBM theory and algorithms.

2. Core assumptions and structural limitations

The standard ScBM rests on four assumptions. First, edges are binary: nrn_r2. Second, edges are conditionally independent given the row and column memberships. Third, each edge follows a Bernoulli law with mean equal to the block-specific probability. Fourth, there is block homogeneity: all node pairs lying in the same row-block/column-block pair share the same expected connection probability. These assumptions yield a block-structured expectation matrix and an analytically tractable model class (Qing et al., 2021).

The same assumptions also delimit the model’s scope. Because the edge law is Bernoulli, ScBM records only presence or absence. Weighted interactions are collapsed to nrn_r3 whenever present, irrespective of magnitude, sign, or count. The restriction is explicit: ScBM “requires that elements of the adjacency matrix should be generated from the Bernoulli distribution,” which is why it is applied to unweighted bipartite networks. This means that block structure in ratings, intensities, signed ties, or counts is outside the standard ScBM likelihood, even when the expectation still appears block-structured.

Another limitation is degree homogeneity within blocks. In the basic model, nodes within the same block have similar expected degrees because the mean matrix is constant within each block pair. Later degree-corrected constructions relax this by introducing sender- and receiver-specific propensities, but the baseline ScBM itself does not contain such terms.

These restrictions should be understood as modeling choices rather than defects. They isolate a clean asymmetric block structure in nrn_r4, and much of the subsequent literature preserves that expectation-level structure while relaxing Bernoulli and degree-homogeneity assumptions.

3. Spectral geometry, community recovery, and model-order estimation

The inferential backbone of ScBM is the singular-value decomposition of the population mean matrix. Under the bipartite model nrn_r5, the compact SVD

nrn_r6

has a blockwise geometry: nrn_r7 for some nrn_r8, and each row of nrn_r9 is constant within a row community; similarly ncn_c0. Distinct blocks are separated in the singular-vector embedding, so in the population case clustering the rows of ncn_c1 and ncn_c2 exactly recovers the row and column memberships. In the observed network, the practical analogue is truncated SVD of ncn_c3 followed by ncn_c4-means on the row embeddings; in the ideal case ncn_c5, this yields exact recovery up to label permutation (Qing et al., 2021).

This geometry explains why ScBM is naturally paired with directed spectral co-clustering. Sender and receiver communities are inferred from left and right singular spaces, respectively. In degree-corrected variants, raw singular-vector rows are no longer constant within clusters because degree parameters rescale them, but row normalization restores a blockwise directional structure. The normalized rows become constant within clusters, and distinct row clusters are separated by ncn_c6 in the population embedding (Qing et al., 2021).

Recent work addresses a problem that earlier ScBM methods typically treated as fixed input: the joint estimation of the sender and receiver community numbers ncn_c7. A goodness-of-fit framework forms a normalized residual matrix ncn_c8 under a candidate pair ncn_c9 and uses

KrK_r0

as a test statistic. Under the null hypothesis of correct community numbers, the upper bound of KrK_r1 converges to zero; under underfitting, KrK_r2 diverges to infinity. This dichotomy yields a lexicographic sequential testing procedure, DiGoF, and a ratio-based variant, RDiGoF, both proven consistent for recovering the true asymmetric community counts under ScBM (Qing, 20 Aug 2025).

The methodological significance is twofold. First, ScBM inference remains fundamentally spectral even when the goal shifts from label estimation to model selection. Second, the asymmetry of the model is not a minor complication: it changes both the geometry of embeddings and the appropriate goodness-of-fit machinery, which is based on singular values rather than symmetric-matrix eigenvalue theory.

4. Generalizations beyond the Bernoulli model

A major line of development keeps the ScBM expectation structure while releasing its distributional restriction. In the Bipartite Distribution-Free Model (BiDFM), the entries KrK_r3 may follow any distribution KrK_r4 as long as

KrK_r5

Bernoulli ScBM is recovered as the special case in which KrK_r6 is Bernoulli and KrK_r7 is nonnegative. The degree-corrected extension, BiDCDFM, replaces the mean by

KrK_r8

so expected edge weights factor into a block-level term and node-specific sender and receiver propensities. When KrK_r9, KcK_c0, and the edge law is Bernoulli, BiDCDFM reduces to degree-corrected ScBM (Qing et al., 2021).

Model Expected adjacency Edge law
ScBM KcK_c1 Bernoulli
BiDFM KcK_c2 arbitrary KcK_c3
BiDCDFM KcK_c4 arbitrary KcK_c5

These generalizations preserve the central ScBM idea: the community structure resides in the block structure of the expectation, not in a specific parametric law for edge generation. The same paper develops BiSC and nBiSC, spectral algorithms with theoretical guarantees for consistent estimation of node labels under the distribution-free and degree-corrected models. Simulations and empirical examples cover Bernoulli, Normal, and signed KcK_c6 settings, and the methods are applied to weighted directed networks including political blogs, monks’ ratings in “Crisis in a Cloister,” Dutch college friendship, highschool friendship networks, and a Facebook-like message-count network (Qing et al., 2021).

A plausible implication is that ScBM is best viewed as a core expectation-level template. Its later extensions do not discard the co-block concept; they retain it while broadening the permissible edge distributions and heterogeneity structure.

5. Dynamic and structured extensions

ScBM has also been embedded into dynamic multivariate time-series models. In a recent framework for VAR-type systems, a degree-corrected ScBM is imposed on transition matrices through

KcK_c7

where KcK_c8 and KcK_c9 are sending and receiving membership matrices, Zr∈Mnr,Kr,Zc∈Mnc,Kc,Z_r \in \mathbb{M}_{n_r,K_r}, \qquad Z_c \in \mathbb{M}_{n_c,K_c},0 is the block interaction matrix, and Zr∈Mnr,Kr,Zc∈Mnc,Kc,Z_r \in \mathbb{M}_{n_r,K_r}, \qquad Z_c \in \mathbb{M}_{n_c,K_c},1 are degree-correction matrices. This separates who tends to send shocks from who tends to receive shocks and summarizes directional dependence in a low-dimensional form (Kim et al., 14 Apr 2026).

The same work develops dynamic directed spectral co-clustering with eigenvector smoothing to estimate latent community paths over time or across dependence horizons. In periodic VAR (PVAR), the model tracks season-specific sending and receiving communities; in generalized VHAR, it tracks horizon-specific communities across short, medium, and long dependence scales. The smoothing procedure, based on PisCES-style projector updates, is designed to stabilize season- or horizon-specific singular-vector embeddings and reveal how directional groups split, merge, or persist.

The theoretical contribution is non-asymptotic misclassification bounds that combine time-series estimation error and ScBM graph randomness. Applications to U.S. nonfarm payrolls distinguish a recurrent business-centered core from more mobile, seasonally sensitive sectors. In global stock volatilities, the framework reveals a compact U.S.-centered long-horizon block, a Europe-heavy developed core, and a more dynamic short-horizon reallocation of peripheral and bridge markets (Kim et al., 14 Apr 2026).

This development extends the original static ScBM perspective in a specific direction: co-block structure is treated not as a single network partition, but as an ordered sequence of asymmetric partitions indexed by season or horizon. The underlying object remains recognizably ScBM—separate sending and receiving communities linked by a block matrix—but its interpretation shifts from static mesoscale structure to latent community paths.

6. Relation to latent block models, Bayesian formulations, and acronym collisions

In the bipartite literature, ScBM is closely aligned with the latent block model (LBM). An LBM places latent groups on both row and column node sets and assumes that edges are independent conditional on the pair of latent groups: Zr∈Mnr,Kr,Zc∈Mnc,Kc,Z_r \in \mathbb{M}_{n_r,K_r}, \qquad Z_c \in \mathbb{M}_{n_c,K_c},2 In this sense, the LBM is exactly a stochastic co-clustering model, and the paper “Blockmodels: A R-package for estimating in Latent Block Model and Stochastic Block Model, with various probability functions, with or without covariates” explicitly presents the LBM as the bipartite counterpart of SBM and mathematically what a stochastic co-block model means. That framework supports Bernoulli, Gaussian, and Poisson edge models, with or without covariates, and estimates memberships via variational EM with automatic group-number selection through the ICL criterion (Leger, 2016).

A Bayesian counterpart arises from collapsed blockmodel methodology. The collapsed SBM integrates out mixing proportions and block parameters, yielding a posterior over the number of clusters and cluster memberships. The same conjugate collapsing logic extends to bipartite co-block structure by introducing separate row and column memberships and a posterior over Zr∈Mnr,Kr,Zc∈Mnc,Kc,Z_r \in \mathbb{M}_{n_r,K_r}, \qquad Z_c \in \mathbb{M}_{n_c,K_c},3. This suggests a fully discrete allocation-sampler approach to ScBM inference, including direct estimation of row and column block numbers without transdimensional MCMC (McDaid et al., 2012).

A further point of clarification concerns terminology. The acronym “SCBM” is not unique. In “Generative models for two-ground-truth partitions in networks,” SCBM denotes the stochastic cross-block model, a different generative construction in which a network is an SBM over cross-partitions representing joint membership in two planted partitions, with cross-block probabilities defined by normalized products of two block matrices. That model is a benchmark for coexistence of multiple ground-truth partitions and should not be conflated with the stochastic co-block model discussed here (Mangold et al., 2023).

Taken together, these neighboring formulations show that ScBM has both a narrow and a broad meaning. In the narrow sense, it is the Bernoulli bipartite or directed co-block model with expectation Zr∈Mnr,Kr,Zc∈Mnc,Kc,Z_r \in \mathbb{M}_{n_r,K_r}, \qquad Z_c \in \mathbb{M}_{n_c,K_c},4. In the broader sense used across related literatures, it refers to probabilistic co-clustering models with distinct row and column latent structures, often extended to weighted edges, covariates, degree correction, Bayesian collapsing, and dynamic directional dependence.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Stochastic Co-Block Model (ScBM).