---
title: Stochastic Co-Block Model (ScBM)
url: https://www.emergentmind.com/topics/stochastic-co-block-model-scbm
type: topic
---

# Stochastic Co-Block Model (ScBM)

Searching arXiv for recent and foundational papers on stochastic co-block models.
The stochastic co-block model (ScBM) is a stochastic blockmodel for binary bipartite graphs and, through the directed–bipartite equivalence, for directed unweighted networks. Its defining feature is an asymmetric latent structure: row nodes and column nodes, or equivalently senders and receivers, possess separate community memberships, and edge probabilities depend only on the corresponding row-block and column-block. In its standard formulation, the adjacency matrix has expectation \( \Omega = \rho Z_r P Z_c' \), so the expected edge intensity is block-constant over each row-community/column-community pair. This places ScBM as the directed or bipartite analogue of the classical stochastic block model while making singular-value-based co-clustering the natural inferential mechanism [2109.10319].

## 1. Formal specification and equivalent representations

In the bipartite formulation, one observes a binary adjacency matrix \(A \in \{0,1\}^{n_r \times n_c}\), where \(n_r\) row nodes interact with \(n_c\) column nodes. Row nodes are partitioned into \(K_r\) communities and column nodes into \(K_c\) communities, encoded by one-hot membership matrices
\[
Z_r \in \mathbb{M}_{n_r,K_r}, \qquad Z_c \in \mathbb{M}_{n_c,K_c},
\]
with exactly one \(1\) per row. If row node \(i_r\) belongs to community \(g_{i_r}\) and column node \(j_c\) to community \(g_{j_c}\), then ScBM assumes conditionally independent Bernoulli edges
\[
A(i_r,j_c)\mid (g_{i_r},g_{j_c}) \sim \mathrm{Bernoulli}\big(\rho\, P(g_{i_r},g_{j_c})\big),
\]
where \(P \in [0,1]^{K_r\times K_c}\) is the block probability matrix and \(\rho>0\) is a sparsity parameter. Equivalently,
\[
\mathbb{E}[A] = \Omega := \rho Z_r P Z_c'.
\]

The directed-network formulation uses a single node set but assigns each node two latent labels: a sender label \(g^s(i)\in\{1,\dots,K_s\}\) and a receiver label \(g^r(i)\in\{1,\dots,K_r\}\). For a directed adjacency matrix \(A\in\{0,1\}^{n\times n}\) with \(A_{ii}=0\),
\[
\mathbb{P}\big(A(i,j)=1\mid g^s,g^r,B\big)=B\big(g^s(i),g^r(j)\big), \qquad i\neq j,
\]
with \(B\in[0,1]^{K_s\times K_r}\). This makes ScBM explicitly asymmetric: the number and composition of sender communities may differ from those of receiver communities, and the block matrix need not be square or symmetric [2508.14816].

A recurring interpretation is that ScBM is not merely a directed SBM with nuisance asymmetry, but a model in which outgoing and incoming relational roles are structurally distinct. That distinction is the reason singular vectors, rather than eigenvectors of a symmetric matrix, occupy a central place in ScBM theory and algorithms.

## 2. Core assumptions and structural limitations

The standard ScBM rests on four assumptions. First, edges are binary: \(A(i_r,j_c)\in\{0,1\}\). Second, edges are conditionally independent given the row and column memberships. Third, each edge follows a Bernoulli law with mean equal to the block-specific probability. Fourth, there is block homogeneity: all node pairs lying in the same row-block/column-block pair share the same expected connection probability. These assumptions yield a block-structured expectation matrix and an analytically tractable model class [2109.10319].

The same assumptions also delimit the model’s scope. Because the edge law is Bernoulli, ScBM records only presence or absence. Weighted interactions are collapsed to \(1\) whenever present, irrespective of magnitude, sign, or count. The restriction is explicit: ScBM “requires that elements of the adjacency matrix should be generated from the Bernoulli distribution,” which is why it is applied to unweighted bipartite networks. This means that block structure in ratings, intensities, signed ties, or counts is outside the standard ScBM likelihood, even when the expectation still appears block-structured.

Another limitation is degree homogeneity within blocks. In the basic model, nodes within the same block have similar expected degrees because the mean matrix is constant within each block pair. Later degree-corrected constructions relax this by introducing sender- and receiver-specific propensities, but the baseline ScBM itself does not contain such terms.

These restrictions should be understood as modeling choices rather than defects. They isolate a clean asymmetric block structure in \( \mathbb{E}[A] \), and much of the subsequent literature preserves that expectation-level structure while relaxing Bernoulli and degree-homogeneity assumptions.

## 3. Spectral geometry, community recovery, and model-order estimation

The inferential backbone of ScBM is the singular-value decomposition of the population mean matrix. Under the bipartite model \( \Omega=\rho Z_r P Z_c' \), the compact SVD
\[
\Omega = U_r \Lambda U_c'
\]
has a blockwise geometry: \(U_r = Z_r X_r\) for some \(X_r\), and each row of \(U_r\) is constant within a row community; similarly \(U_c=Z_cX_c\). Distinct blocks are separated in the singular-vector embedding, so in the population case clustering the rows of \(U_r\) and \(U_c\) exactly recovers the row and column memberships. In the observed network, the practical analogue is truncated SVD of \(A\) followed by \(k\)-means on the row embeddings; in the ideal case \(A=\Omega\), this yields exact recovery up to label permutation [2109.10319].

This geometry explains why ScBM is naturally paired with directed spectral co-clustering. Sender and receiver communities are inferred from left and right singular spaces, respectively. In degree-corrected variants, raw singular-vector rows are no longer constant within clusters because degree parameters rescale them, but row normalization restores a blockwise directional structure. The normalized rows become constant within clusters, and distinct row clusters are separated by \( \sqrt{2} \) in the population embedding [2109.10319].

Recent work addresses a problem that earlier ScBM methods typically treated as fixed input: the joint estimation of the sender and receiver community numbers \( (K_s,K_r) \). A goodness-of-fit framework forms a normalized residual matrix \( \hat R \) under a candidate pair \( (K_{s0},K_{r0}) \) and uses
\[
\hat T_n = \sigma_1(\hat R)-2
\]
as a test statistic. Under the null hypothesis of correct community numbers, the upper bound of \( \hat T_n \) converges to zero; under underfitting, \( \hat T_n \) diverges to infinity. This dichotomy yields a lexicographic sequential testing procedure, DiGoF, and a ratio-based variant, RDiGoF, both proven consistent for recovering the true asymmetric community counts under ScBM [2508.14816].

The methodological significance is twofold. First, ScBM inference remains fundamentally spectral even when the goal shifts from label estimation to model selection. Second, the asymmetry of the model is not a minor complication: it changes both the geometry of embeddings and the appropriate goodness-of-fit machinery, which is based on singular values rather than symmetric-matrix eigenvalue theory.

## 4. Generalizations beyond the Bernoulli model

A major line of development keeps the ScBM expectation structure while releasing its distributional restriction. In the Bipartite Distribution-Free Model (BiDFM), the entries \(A(i_r,j_c)\) may follow any distribution \( \mathcal{F} \) as long as
\[
\mathbb{E}[A] = \rho Z_r P Z_c'.
\]
Bernoulli ScBM is recovered as the special case in which \( \mathcal{F} \) is Bernoulli and \(P\) is nonnegative. The degree-corrected extension, BiDCDFM, replaces the mean by
\[
\Omega := \Theta_r Z_r P Z_c' \Theta_c,
\]
so expected edge weights factor into a block-level term and node-specific sender and receiver propensities. When \( \Theta_r=\sqrt{\rho}\,I \), \( \Theta_c=\sqrt{\rho}\,I \), and the edge law is Bernoulli, BiDCDFM reduces to degree-corrected ScBM [2109.10319].

| Model | Expected adjacency | Edge law |
|---|---|---|
| ScBM | \( \rho Z_r P Z_c' \) | Bernoulli |
| BiDFM | \( \rho Z_r P Z_c' \) | arbitrary \( \mathcal{F} \) |
| BiDCDFM | \( \Theta_r Z_r P Z_c' \Theta_c \) | arbitrary \( \mathcal{F} \) |

These generalizations preserve the central ScBM idea: the community structure resides in the block structure of the expectation, not in a specific parametric law for edge generation. The same paper develops BiSC and nBiSC, spectral algorithms with theoretical guarantees for consistent estimation of node labels under the distribution-free and degree-corrected models. Simulations and empirical examples cover Bernoulli, Normal, and signed \( \pm 1 \) settings, and the methods are applied to weighted directed networks including political blogs, monks’ ratings in “Crisis in a Cloister,” Dutch college friendship, highschool friendship networks, and a Facebook-like message-count network [2109.10319].

A plausible implication is that ScBM is best viewed as a core expectation-level template. Its later extensions do not discard the co-block concept; they retain it while broadening the permissible edge distributions and heterogeneity structure.

## 5. Dynamic and structured extensions

ScBM has also been embedded into dynamic multivariate time-series models. In a recent framework for VAR-type systems, a degree-corrected ScBM is imposed on transition matrices through
\[
\mathcal{A} := \mathbb{E}[A] = \mu\, \Theta^y Y B Z' \Theta^z,
\]
where \(Y\) and \(Z\) are sending and receiving membership matrices, \(B\) is the block interaction matrix, and \( \Theta^y,\Theta^z \) are degree-correction matrices. This separates who tends to send shocks from who tends to receive shocks and summarizes directional dependence in a low-dimensional form [2604.12563].

The same work develops dynamic directed spectral co-clustering with eigenvector smoothing to estimate latent community paths over time or across dependence horizons. In periodic VAR (PVAR), the model tracks season-specific sending and receiving communities; in generalized VHAR, it tracks horizon-specific communities across short, medium, and long dependence scales. The smoothing procedure, based on PisCES-style projector updates, is designed to stabilize season- or horizon-specific singular-vector embeddings and reveal how directional groups split, merge, or persist.

The theoretical contribution is non-asymptotic misclassification bounds that combine time-series estimation error and ScBM graph randomness. Applications to U.S. nonfarm payrolls distinguish a recurrent business-centered core from more mobile, seasonally sensitive sectors. In global stock volatilities, the framework reveals a compact U.S.-centered long-horizon block, a Europe-heavy developed core, and a more dynamic short-horizon reallocation of peripheral and bridge markets [2604.12563].

This development extends the original static ScBM perspective in a specific direction: co-block structure is treated not as a single network partition, but as an ordered sequence of asymmetric partitions indexed by season or horizon. The underlying object remains recognizably ScBM—separate sending and receiving communities linked by a block matrix—but its interpretation shifts from static mesoscale structure to latent community paths.

## 6. Relation to latent block models, Bayesian formulations, and acronym collisions

In the bipartite literature, ScBM is closely aligned with the latent block model (LBM). An LBM places latent groups on both row and column node sets and assumes that edges are independent conditional on the pair of latent groups:
\[
X_{ij}\mid \big(Z^{(1)}_{iq} Z^{(2)}_{jl}=1\big) \overset{\mathrm{ind}}{\sim} F_{ql}^{Y_{ij}}.
\]
In this sense, the LBM is exactly a stochastic co-clustering model, and the paper “Blockmodels: A R-package for estimating in Latent Block Model and Stochastic Block Model, with various probability functions, with or without covariates” explicitly presents the LBM as the bipartite counterpart of SBM and mathematically what a stochastic co-block model means. That framework supports Bernoulli, Gaussian, and Poisson edge models, with or without covariates, and estimates memberships via variational EM with automatic group-number selection through the ICL criterion [1602.07587].

A Bayesian counterpart arises from collapsed blockmodel methodology. The collapsed SBM integrates out mixing proportions and block parameters, yielding a posterior over the number of clusters and cluster memberships. The same conjugate collapsing logic extends to bipartite co-block structure by introducing separate row and column memberships and a posterior over \( (K_U,K_V,z^{(U)},z^{(V)}) \). This suggests a fully discrete allocation-sampler approach to ScBM inference, including direct estimation of row and column block numbers without transdimensional MCMC [1203.3083].

A further point of clarification concerns terminology. The acronym “SCBM” is not unique. In “Generative models for two-ground-truth partitions in networks,” SCBM denotes the stochastic cross-block model, a different generative construction in which a network is an SBM over cross-partitions representing joint membership in two planted partitions, with cross-block probabilities defined by normalized products of two block matrices. That model is a benchmark for coexistence of multiple ground-truth partitions and should not be conflated with the stochastic co-block model discussed here [2302.02787].

Taken together, these neighboring formulations show that ScBM has both a narrow and a broad meaning. In the narrow sense, it is the Bernoulli bipartite or directed co-block model with expectation \( \rho Z_r P Z_c' \). In the broader sense used across related literatures, it refers to probabilistic co-clustering models with distinct row and column latent structures, often extended to weighted edges, covariates, degree correction, Bayesian collapsing, and dynamic directional dependence.

Source: https://www.emergentmind.com/topics/stochastic-co-block-model-scbm