Papers
Topics
Authors
Recent
Search
2000 character limit reached

MCBP: Multi-Domain Technical Constructs

Updated 10 July 2026
  • MCBP is an overloaded acronym used in atomic physics, geometric machine learning, graphical-model inference, and LLM accelerator design, with each domain employing distinct computational workflows.
  • In atomic physics, the multi-configuration Breit–Pauli method models dielectronic recombination through configuration-interaction and perturbation theory to achieve plasma rate accuracies.
  • In machine learning and inference, MCBP encompasses mean curvature boundary detection, MCMC assisted by belief propagation, and memory-compute co-design for LLM acceleration with optimized hardware.

MCBP is an overloaded acronym that denotes several unrelated technical constructs in current research literature. In the sources considered here, it refers to the multi-configuration Breit–Pauli formalism in atomic collision theory, Mean Curvature Boundary Points in geometric machine learning, MCMC assisted by Belief Propagation in graphical-model inference, and a Memory-Compute co-design exploiting Bit-slice sparsity and Repetitiveness for large-language-model inference acceleration. The shared acronym does not imply conceptual continuity across these domains; each usage has its own mathematical objects, computational workflow, and validation regime (Novotný et al., 2012, Levada, 5 May 2026, Ahn et al., 2016, Wang et al., 12 Sep 2025).

1. Disambiguation and research domains

In the cited literature, MCBP appears in four principal senses.

Expansion Domain Representative source
Multi-configuration Breit–Pauli Atomic physics; dielectronic recombination (Novotný et al., 2012)
Mean Curvature Boundary Points Geometric machine learning; unsupervised learning (Levada, 5 May 2026)
MCMC assisted by Belief Propagation Probabilistic inference in graphical models (Ahn et al., 2016)
Memory-Compute co-design exploiting Bit-slice sparsity and Repetitiveness LLM accelerator architecture (Wang et al., 12 Sep 2025)

The ambiguity is especially relevant because two of these senses are currently active in high-dimensional data analysis: Mean Curvature Boundary Points as an unsupervised-learning method, and a later acceleration paper that explicitly targets the curvature computation used by that method (Levada, 5 May 2026, Levada, 4 Jun 2026). By contrast, the atomic-physics sense predates them and is embedded in the AUTOSTRUCTURE/IPIRDW tradition for recombination-rate calculations (Nikolić et al., 2010, Kaur et al., 2018).

2. MCBP as multi-configuration Breit–Pauli theory

In atomic physics, MCBP is an atomic-structure and collision methodology in which the non-relativistic Hamiltonian is supplemented by Breit–Pauli operators, bound and autoionizing levels are represented by configuration-interaction expansions, and scattering and radiative transitions are computed in lowest-order perturbation theory (Kaur et al., 2018). In intermediate coupling, the Hamiltonian is written as

HBP=HNR+HMV+HD+HSO+HSS+HSOO,H_{\rm BP}=H_{\rm NR}+H_{\rm MV}+H_{\rm D}+H_{\rm SO}+H_{\rm SS}+H_{\rm SOO},

with mass-velocity, Darwin, spin–orbit, spin–spin, and spin–other-orbit contributions explicitly included (Nikolić et al., 2010). The corresponding ionic states are expanded as

Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,

and the coefficients are obtained by diagonalizing the Breit–Pauli Hamiltonian in a CSF basis (Nikolić et al., 2010, Kaur et al., 2018).

The method is used extensively for dielectronic recombination (DR). In the independent-processes, isolated-resonance, distorted-wave approximation, the DR cross section is represented as a sum over isolated resonances whose widths are determined by autoionization and radiative rates. For argon-like ions, the partial DR cross section is written in Lorentzian form and Maxwellian averaging gives

αDR(T)≈(8πme)1/2(kBT)−3/2∫0∞E σDR(E)e−E/kBT dE,\alpha^{\rm DR}(T)\approx \Bigl(\tfrac{8}{\pi m_e}\Bigr)^{1/2}(k_BT)^{-3/2} \int_0^\infty E\,\sigma^{\rm DR}(E)e^{-E/k_BT}\,dE,

with resonance sums carried explicitly to n=1000n=1000 and ℓ=10\ell=10, followed by a hydrogenic top-up (Nikolić et al., 2010). For the silicon isoelectronic sequence, the partial DR rate from initial ii to final ff is expressed as

αif(T)=(4πa02IHkBT)3/2∑dωd2ωie−Ec/kBTAd→iaAd→fr∑hAd→hr+∑mAd→ma,\alpha_{if}(T)=\Bigl(\frac{4\pi a_0^2 I_H}{k_BT}\Bigr)^{3/2} \sum_d \frac{\omega_d}{2\omega_i}e^{-E_c/k_BT} \frac{A^a_{d\to i}A^r_{d\to f}} {\sum_h A^r_{d\to h}+\sum_m A^a_{d\to m}},

with both Δnc=0\Delta n_c=0 and Δnc=1\Delta n_c=1 core excitations included (Kaur et al., 2018).

AUTOSTRUCTURE is the principal implementation platform in the cited work. In the Fe XII study, all structure, autoionization, and radiative data generation was done with AUTOSTRUCTURE, using scaled Thomas–Fermi–Dirac–Amaldi potentials, CI+BP Hamiltonians with several thousand CSFs, and continuum functions in analytic Coulomb form (Novotný et al., 2012). The same paper compared merged-beams recombination rate coefficients measured at the TSR heavy-ion storage ring with MCBP theory and found significant differences in resonance energies and strengths at the MBRRC level, including Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,0 ranging from Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,1 to Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,2 across different energy intervals (Novotný et al., 2012). Yet the Maxwellian-averaged plasma rate coefficient was much more robust: the MCBP PRRC agreed with experiment to within Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,3 at photoionized-plasma temperatures and within the Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,4 experimental uncertainty at collisionally ionized-plasma temperatures (Novotný et al., 2012).

A recurring conclusion in the atomic-physics literature is that accurate threshold positions and CI completeness are decisive at low temperature. For argon-like ions, the lowest Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,5 thresholds control the strong low-Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,6 DR enhancement, and the reported MCBP rates at Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,7 K are larger by factors of Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,8–Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,9 than widely used empirical formulas (Nikolić et al., 2010). For silicon-like ions, older recommended fits miss low-αDR(T)≈(8πme)1/2(kBT)−3/2∫0∞E σDR(E)e−E/kBT dE,\alpha^{\rm DR}(T)\approx \Bigl(\tfrac{8}{\pi m_e}\Bigr)^{1/2}(k_BT)^{-3/2} \int_0^\infty E\,\sigma^{\rm DR}(E)e^{-E/k_BT}\,dE,0 fine-structure DR, whereas MCBP with full fine-structure mixing and αDR(T)≈(8πme)1/2(kBT)−3/2∫0∞E σDR(E)e−E/kBT dE,\alpha^{\rm DR}(T)\approx \Bigl(\tfrac{8}{\pi m_e}\Bigr)^{1/2}(k_BT)^{-3/2} \int_0^\infty E\,\sigma^{\rm DR}(E)e^{-E/k_BT}\,dE,1 coverage yields data intended for generalized collisional-radiative modelling frameworks and OPEN-ADAS archiving (Kaur et al., 2018).

3. MCBP as Mean Curvature Boundary Points

In geometric machine learning, MCBP denotes Mean Curvature Boundary Points, a boundary-detection framework for unsupervised learning that uses local mean curvature as a descriptor of boundary structure (Levada, 5 May 2026). The method starts from the differential-geometric observation that the shape operator

αDR(T)≈(8πme)1/2(kBT)−3/2∫0∞E σDR(E)e−E/kBT dE,\alpha^{\rm DR}(T)\approx \Bigl(\tfrac{8}{\pi m_e}\Bigr)^{1/2}(k_BT)^{-3/2} \int_0^\infty E\,\sigma^{\rm DR}(E)e^{-E/k_BT}\,dE,2

encodes second-order bending, and that the mean curvature can be written as

αDR(T)≈(8πme)1/2(kBT)−3/2∫0∞E σDR(E)e−E/kBT dE,\alpha^{\rm DR}(T)\approx \Bigl(\tfrac{8}{\pi m_e}\Bigr)^{1/2}(k_BT)^{-3/2} \int_0^\infty E\,\sigma^{\rm DR}(E)e^{-E/k_BT}\,dE,3

For a point cloud αDR(T)≈(8πme)1/2(kBT)−3/2∫0∞E σDR(E)e−E/kBT dE,\alpha^{\rm DR}(T)\approx \Bigl(\tfrac{8}{\pi m_e}\Bigr)^{1/2}(k_BT)^{-3/2} \int_0^\infty E\,\sigma^{\rm DR}(E)e^{-E/k_BT}\,dE,4, explicit manifold parametrization is avoided by estimating a local patch αDR(T)≈(8πme)1/2(kBT)−3/2∫0∞E σDR(E)e−E/kBT dE,\alpha^{\rm DR}(T)\approx \Bigl(\tfrac{8}{\pi m_e}\Bigr)^{1/2}(k_BT)^{-3/2} \int_0^\infty E\,\sigma^{\rm DR}(E)e^{-E/k_BT}\,dE,5 from αDR(T)≈(8πme)1/2(kBT)−3/2∫0∞E σDR(E)e−E/kBT dE,\alpha^{\rm DR}(T)\approx \Bigl(\tfrac{8}{\pi m_e}\Bigr)^{1/2}(k_BT)^{-3/2} \int_0^\infty E\,\sigma^{\rm DR}(E)e^{-E/k_BT}\,dE,6-nearest neighbors, computing the empirical covariance

αDR(T)≈(8πme)1/2(kBT)−3/2∫0∞E σDR(E)e−E/kBT dE,\alpha^{\rm DR}(T)\approx \Bigl(\tfrac{8}{\pi m_e}\Bigr)^{1/2}(k_BT)^{-3/2} \int_0^\infty E\,\sigma^{\rm DR}(E)e^{-E/k_BT}\,dE,7

approximating the first fundamental form by αDR(T)≈(8πme)1/2(kBT)−3/2∫0∞E σDR(E)e−E/kBT dE,\alpha^{\rm DR}(T)\approx \Bigl(\tfrac{8}{\pi m_e}\Bigr)^{1/2}(k_BT)^{-3/2} \int_0^\infty E\,\sigma^{\rm DR}(E)e^{-E/k_BT}\,dE,8, and forming a discrete proxy αDR(T)≈(8πme)1/2(kBT)−3/2∫0∞E σDR(E)e−E/kBT dE,\alpha^{\rm DR}(T)\approx \Bigl(\tfrac{8}{\pi m_e}\Bigr)^{1/2}(k_BT)^{-3/2} \int_0^\infty E\,\sigma^{\rm DR}(E)e^{-E/k_BT}\,dE,9 for the second fundamental form from a local quadratic fit (Levada, 5 May 2026).

The resulting shape-operator approximation is

n=1000n=10000

and the raw mean-curvature score is

n=1000n=10001

These scores are normalized to

n=1000n=10002

after which boundary points are extracted through percentile thresholding: n=1000n=10003 The authors describe this as an adaptive percentile-based thresholding scheme enabling multiscale boundary extraction without ad hoc density parameters (Levada, 5 May 2026).

The geometric interpretation is broader than classical contour detection. High-curvature regions are said to correspond to transitions between clusters, geometric irregularities, and low-density interfaces; the same curvature field therefore yields a unified interpretation of boundary points, outliers, and transition points (Levada, 5 May 2026). The method also induces a curvature-driven decomposition into a smooth set

n=1000n=10004

and a boundary set

n=1000n=10005

with the smooth subset functioning as a non-linear geometric filter for downstream clustering (Levada, 5 May 2026).

The reported complexity is dominated by ambient dimension. Building the n=1000n=10006-nearest-neighbor graph costs n=1000n=10007, while per-point covariance estimation, eigendecomposition, and local quadratic fitting lead to a total complexity of n=1000n=10008; a prior PCA to n=1000n=10009 reduces this to ℓ=10\ell=100 (Levada, 5 May 2026). Empirically, synthetic experiments on Gaussian blob, two-blobs, anisotropic ellipses, and two-moons datasets showed curvature peaks at outer contours and inter-cluster interfaces, and experiments on 25 real datasets reported average clustering improvements after filtering out the top 25% curvature points: Silhouette Coefficient increased by ℓ=10\ell=101, Calinski–Harabasz by ℓ=10\ell=102, and Davies–Bouldin decreased by ℓ=10\ell=103 (Levada, 5 May 2026).

4. Scalable curvature computation for MCBP

A later paper addresses the principal computational bottleneck of the Mean Curvature Boundary Points pipeline: local mean-curvature computation in high ambient dimension (Levada, 4 Jun 2026). The original estimator constructs, for each neighborhood, a covariance matrix

â„“=10\ell=104

computes its eigendecomposition â„“=10\ell=105, and then forms a feature matrix â„“=10\ell=106 with â„“=10\ell=107 columns from squared and cross products of eigenvectors (Levada, 4 Jun 2026). The discrete shape-operator estimator is

â„“=10\ell=108

and the mean-curvature estimate is

â„“=10\ell=109

Forming ii0 costs ii1 per point, which the paper identifies as the dominant obstacle to using MCBP on data with more than a few dozen features (Levada, 4 Jun 2026).

The first acceleration is an exact algebraic identity: ii2 where ii3 is the columnwise element-wise square of ii4. This removes explicit construction of ii5 from the trace computation and yields

ii6

with ii7 (Levada, 4 Jun 2026). After eigendecomposition, the trace computation drops to ii8.

The second acceleration exploits the fact that ii9 has rank at most ff0. Replacing full eigendecomposition with a truncated SVD of the centered ff1 data matrix reduces the cost to ff2, and an analytical approximation for the null-space contribution is derived from the expected outer product of null-space eigenvectors under the Haar measure (Levada, 4 Jun 2026). The resulting total complexity is

ff3

For fixed small ff4, the dominant term is ff5 (Levada, 4 Jun 2026).

The empirical consequences are substantial. On 40 OpenML datasets with ff6 from 4 to 279, the combined exact/fast implementation achieved median Spearman ff7, Chatterjee ff8 after normalization, and median normalized MAE below ff9. Median wall-clock time was αif(T)=(4πa02IHkBT)3/2∑dωd2ωie−Ec/kBTAd→iaAd→fr∑hAd→hr+∑mAd→ma,\alpha_{if}(T)=\Bigl(\frac{4\pi a_0^2 I_H}{k_BT}\Bigr)^{3/2} \sum_d \frac{\omega_d}{2\omega_i}e^{-E_c/k_BT} \frac{A^a_{d\to i}A^r_{d\to f}} {\sum_h A^r_{d\to h}+\sum_m A^a_{d\to m}},0 s versus αif(T)=(4πa02IHkBT)3/2∑dωd2ωie−Ec/kBTAd→iaAd→fr∑hAd→hr+∑mAd→ma,\alpha_{if}(T)=\Bigl(\frac{4\pi a_0^2 I_H}{k_BT}\Bigr)^{3/2} \sum_d \frac{\omega_d}{2\omega_i}e^{-E_c/k_BT} \frac{A^a_{d\to i}A^r_{d\to f}} {\sum_h A^r_{d\to h}+\sum_m A^a_{d\to m}},1 s for the original on low dimensions, while for αif(T)=(4πa02IHkBT)3/2∑dωd2ωie−Ec/kBTAd→iaAd→fr∑hAd→hr+∑mAd→ma,\alpha_{if}(T)=\Bigl(\frac{4\pi a_0^2 I_H}{k_BT}\Bigr)^{3/2} \sum_d \frac{\omega_d}{2\omega_i}e^{-E_c/k_BT} \frac{A^a_{d\to i}A^r_{d\to f}} {\sum_h A^r_{d\to h}+\sum_m A^a_{d\to m}},2–αif(T)=(4πa02IHkBT)3/2∑dωd2ωie−Ec/kBTAd→iaAd→fr∑hAd→hr+∑mAd→ma,\alpha_{if}(T)=\Bigl(\frac{4\pi a_0^2 I_H}{k_BT}\Bigr)^{3/2} \sum_d \frac{\omega_d}{2\omega_i}e^{-E_c/k_BT} \frac{A^a_{d\to i}A^r_{d\to f}} {\sum_h A^r_{d\to h}+\sum_m A^a_{d\to m}},3 the reported speedups were αif(T)=(4πa02IHkBT)3/2∑dωd2ωie−Ec/kBTAd→iaAd→fr∑hAd→hr+∑mAd→ma,\alpha_{if}(T)=\Bigl(\frac{4\pi a_0^2 I_H}{k_BT}\Bigr)^{3/2} \sum_d \frac{\omega_d}{2\omega_i}e^{-E_c/k_BT} \frac{A^a_{d\to i}A^r_{d\to f}} {\sum_h A^r_{d\to h}+\sum_m A^a_{d\to m}},4–αif(T)=(4πa02IHkBT)3/2∑dωd2ωie−Ec/kBTAd→iaAd→fr∑hAd→hr+∑mAd→ma,\alpha_{if}(T)=\Bigl(\frac{4\pi a_0^2 I_H}{k_BT}\Bigr)^{3/2} \sum_d \frac{\omega_d}{2\omega_i}e^{-E_c/k_BT} \frac{A^a_{d\to i}A^r_{d\to f}} {\sum_h A^r_{d\to h}+\sum_m A^a_{d\to m}},5; specific examples included USPS (αif(T)=(4πa02IHkBT)3/2∑dωd2ωie−Ec/kBTAd→iaAd→fr∑hAd→hr+∑mAd→ma,\alpha_{if}(T)=\Bigl(\frac{4\pi a_0^2 I_H}{k_BT}\Bigr)^{3/2} \sum_d \frac{\omega_d}{2\omega_i}e^{-E_c/k_BT} \frac{A^a_{d\to i}A^r_{d\to f}} {\sum_h A^r_{d\to h}+\sum_m A^a_{d\to m}},6) from αif(T)=(4πa02IHkBT)3/2∑dωd2ωie−Ec/kBTAd→iaAd→fr∑hAd→hr+∑mAd→ma,\alpha_{if}(T)=\Bigl(\frac{4\pi a_0^2 I_H}{k_BT}\Bigr)^{3/2} \sum_d \frac{\omega_d}{2\omega_i}e^{-E_c/k_BT} \frac{A^a_{d\to i}A^r_{d\to f}} {\sum_h A^r_{d\to h}+\sum_m A^a_{d\to m}},7 s to αif(T)=(4πa02IHkBT)3/2∑dωd2ωie−Ec/kBTAd→iaAd→fr∑hAd→hr+∑mAd→ma,\alpha_{if}(T)=\Bigl(\frac{4\pi a_0^2 I_H}{k_BT}\Bigr)^{3/2} \sum_d \frac{\omega_d}{2\omega_i}e^{-E_c/k_BT} \frac{A^a_{d\to i}A^r_{d\to f}} {\sum_h A^r_{d\to h}+\sum_m A^a_{d\to m}},8 s and Arrhythmia (αif(T)=(4πa02IHkBT)3/2∑dωd2ωie−Ec/kBTAd→iaAd→fr∑hAd→hr+∑mAd→ma,\alpha_{if}(T)=\Bigl(\frac{4\pi a_0^2 I_H}{k_BT}\Bigr)^{3/2} \sum_d \frac{\omega_d}{2\omega_i}e^{-E_c/k_BT} \frac{A^a_{d\to i}A^r_{d\to f}} {\sum_h A^r_{d\to h}+\sum_m A^a_{d\to m}},9) from Δnc=0\Delta n_c=00 s to Δnc=0\Delta n_c=01 s (Levada, 4 Jun 2026). Exact mode is algebraically identical, with errors of order Δnc=0\Delta n_c=02, whereas the fast mode introduces Δnc=0\Delta n_c=03 bias that is described as negligible for Δnc=0\Delta n_c=04 and rank-order preserving (Levada, 4 Jun 2026).

5. MCBP as MCMC assisted by Belief Propagation

In graphical-model inference, MCBP stands for MCMC assisted by Belief Propagation (Ahn et al., 2016). The framework addresses the contrast between BP, which is typically fast but approximate on loopy graphs, and MCMC, which is asymptotically exact but may mix exponentially slowly. The formal starting point is loop calculus for a pairwise binary Markov random field

Δnc=0\Delta n_c=05

together with the identity

Δnc=0\Delta n_c=06

where Δnc=0\Delta n_c=07 is the set of generalized loops and Δnc=0\Delta n_c=08 is the Bethe approximation produced by BP (Ahn et al., 2016).

A central truncation keeps only 2-regular loops,

Δnc=0\Delta n_c=09

leading to the 2-loop series

Δnc=1\Delta n_c=10

For planar pairwise binary graphical models this truncated sum is computable in polynomial time by Pfaffian methods; the paper’s contribution is to approximate it in general graphs by MCMC (Ahn et al., 2016). The proposed sampler uses the worm algorithm on an enlarged state space containing both 2-regular loops and subgraphs with exactly two odd-degree vertices, plus a rejection scheme that turns endpoint states into samples from the target distribution over 2-regular loops (Ahn et al., 2016).

For the full loop series, the paper introduces a rejection-free chain based on a cycle basis and a fixed path set. Any generalized loop is decomposed as an XOR over elements of Δnc=1\Delta n_c=11, and the Markov chain proposes Δnc=1\Delta n_c=12 for a uniformly selected basis element Δnc=1\Delta n_c=13, accepting with probability Δnc=1\Delta n_c=14 whenever Δnc=1\Delta n_c=15 remains a generalized loop (Ahn et al., 2016). Both truncated and full-series estimators are embedded in a simulated-annealing schedule Δnc=1\Delta n_c=16 to estimate partition-function ratios stage by stage.

The theoretical guarantees focus on polynomial-time approximation under explicit assumptions. The worm chain is shown to mix in time

Δnc=1\Delta n_c=17

and the paper states that relative-error estimation of the truncated or full loop series is polynomial in Δnc=1\Delta n_c=18, Δnc=1\Delta n_c=19, Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,00, and Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,01, provided Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,02 and the sign imbalance remain at least inverse-polynomial (Ahn et al., 2016). Empirically, on Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,03 grid Ising models with and without random external field, and on the hard-core model, the MCBP variants reduced the Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,04 error in Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,05 relative to BP and standard Gibbs-MCMC, often by an order of magnitude, while incurring comparable chain-run cost (Ahn et al., 2016).

6. MCBP as a memory-compute efficient LLM accelerator

In computer architecture, MCBP denotes a Memory-Compute co-design exploiting Bit-slice sparsity and Repetitiveness for decoder-only LLM inference (Wang et al., 12 Sep 2025). The architecture is organized around three bottlenecks—GEMM computation, weight loading, and KV-cache loading—and seeks to optimize all three simultaneously by operating at the bit-slice level rather than the value level (Wang et al., 12 Sep 2025). Quantized 8-bit weights are stored in off-chip HBM in a bit-slice–first layout, decoded into on-chip SRAM, processed by BRCR units for GEMM reduction, and combined with BGPP units that progressively prune KV accesses during attention (Wang et al., 12 Sep 2025).

The first mechanism, BS-Repetitiveness-Enabled Computation Reduction (BRCR), exploits repetition patterns among columns of an Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,06 bit-slice block. A naive bit-plane GEMV requires

Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,07

where Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,08 is average bit-sparsity. BRCR rewrites a group multiply as

Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,09

so that the addition count per bit-plane becomes

Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,10

and over Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,11 bit-planes

Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,12

(Wang et al., 12 Sep 2025).

The second mechanism, BS-Sparsity-Enabled Two-State Coding (BSTC), compresses each bit-slice group columnwise into either a zero-block or a non-zero block with a one-bit state tag and, for nonzero entries, the raw Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,13 bits (Wang et al., 12 Sep 2025). The paper reports that on LLaMA-7B, high-order bit-slices with Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,14 exhibit Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,15, yielding compression ratios of approximately Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,16–Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,17, and that overall weight traffic is reduced by Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,18 on average in decoding (Wang et al., 12 Sep 2025). The third mechanism, Bit-Grained Progressive Prediction (BGPP), performs progressive top-Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,19 attention filtering from MSB to LSB using thresholds

Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,20

with empirical radius Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,21. Candidates whose partial score plus the maximum remaining-bit contribution cannot exceed the threshold are pruned early; the reported effect is up to Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,22 reduction in KV-cache accesses and Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,23 reduction in attention computation with negligible quality loss below Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,24 for Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,25 around Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,26–Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,27 (Wang et al., 12 Sep 2025).

The accelerator instantiates these mechanisms with HBM2 channels, dedicated SRAMs, PE clusters for BRCR, lightweight BSTC codecs, and BGPP units running asynchronously with the main pipeline (Wang et al., 12 Sep 2025). On 26 benchmarks, the standard design achieved Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,28 speedup over Nvidia A100 INT8 TensorRT and Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,29 GOPS/W, while the aggressive design achieved Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,30 speedup and Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,31 GOPS/W; the abstract additionally reports Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,32 higher energy efficiency than the A100 and energy savings of Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,33, Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,34, and Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,35 relative to SpAtten, FACT, and SOFA, respectively (Wang et al., 12 Sep 2025). Area and power figures at 28 nm and 1 GHz are given as Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,36 and Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,37 W, excluding the note that the HBM interface accounts for about Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,38 of power (Wang et al., 12 Sep 2025). The limitations explicitly identified are dependence of BGPP on Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,39, the fixed group size Ψi=∑kck(i)Φk,\Psi_i=\sum_k c_k^{(i)}\Phi_k,40, BSTC’s one-bit overhead for each non-zero block, and the need to extend the co-design to mixed precision, INT4/FP8, and broader system-level co-optimization (Wang et al., 12 Sep 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to MCBP.