---
title: 'WCFE: Weight Clustering for CNN Extraction'
url: https://www.emergentmind.com/topics/weight-clustering-feature-extraction-wcfe
type: topic
---

# WCFE: Weight Clustering for CNN Extraction

Weight Clustering Feature Extraction (WCFE) denotes a weight-clustering-based CNN feature extractor used as the front-end of HDnn-family on-device learning accelerators, notably FSL-HDnn and Clo-HDnn. In this usage, WCFE is a hardware-efficient convolution engine that reduces compute, memory, and data movement by clustering similar weights into shared representative values, storing weights as low-bit indices rather than full-precision values, accumulating input activations by weight cluster before multiplication, and reusing clustered activation patterns across filters or channels. Within these systems, WCFE provides the feature vector consumed by a downstream Hyperdimensional Computing (HDC) module; the feature extractor is kept frozen during adaptation, while learning is performed in the classifier stage [2409.10918], [2507.17953], [2512.11826].

## 1. Definition and architectural position

In FSL-HDnn, WCFE is the CNN-based feature extraction module of an end-to-end few-shot learning pipeline whose two major parts are feature extraction and HDC-based classification or learning. The feature extractor uses “per-filter weight clustering and pattern sharing across filters,” and it remains frozen during few-shot learning, so on-chip adaptation updates only the HDC classifier rather than CNN weights [2409.10918].

In Clo-HDnn, WCFE is likewise the CNN-like feature extraction front-end, but it is embedded in a continual on-device learning architecture composed of a WCFE module, an HD module for encoding, training, and inference, and a global FIFO connecting them. Clo-HDnn further supports dual-mode operation: for simple datasets, WCFE can be bypassed and raw inputs can be sent directly toward the HD pipeline, whereas more complex datasets use WCFE before HDC [2507.17953].

The 2025 FSL-HDnn formulation makes the same division of labor explicit. WCFE is described there as a parameter-efficient feature extractor that keeps the CNN feature extractor frozen after ImageNet pretraining, applies weight clustering to pretrained CNN weights, and uses the clustered representation to reduce memory footprint, data movement, and MAC operations while preserving feature quality for the downstream classifier [2512.11826].

## 2. Weight clustering and convolutional computation

The core WCFE mechanism is weight clustering. Similar weights are clustered into the same average value, so a conventional convolution filter weight tensor is approximated by a small set of shared representative values. In the FSL-HDnn description, prior studies are cited as indicating that up to 16 unique weights per filter can preserve accuracy comparable to unclustered feature extraction; this corresponds to 4-bit indices, with storage split into cluster representative values and a 4-bit index identifying the cluster membership of each original weight [2409.10918].

Operationally, WCFE exploits repeated weight indices. Input activations associated with the same weight index are first accumulated and are then multiplied by the corresponding centroid or codebook value. The later FSL-HDnn formulation states this explicitly in clustered form: activations with the same index are accumulated, then \(N\) accumulated inputs are multiplied with \(N\) codebook weights, and the final output is formed by summing those results. It also gives the corresponding operation-count change: the original convolution cost is
\[
2 \times K^2 - 1
\]
whereas the clustered version becomes
\[
K^2 + N - 1.
\]
This is the direct computational rationale for fewer multiplications and reduced redundant accumulation [2512.11826].

A second optimization is pattern reuse across filters. In the 2024 FSL-HDnn description, the clustering pattern is shared across filters for different channels so that accumulated input pixels can be reused by the filters for many output channels. In Clo-HDnn, the same idea is described as grouping, accumulating, and multiplying once for inputs sharing the same weight value during convolution. This suggests that WCFE is not merely weight compression; it is a reuse-aware execution scheme whose efficiency depends on both shared values and shared clustering patterns [2409.10918], [2507.17953].

## 3. Hardware organization and PE-level execution

WCFE is realized as a specialized convolution accelerator. Across both HDnn implementations, the compute core is a \(4 \times 16\) PE array, corresponding to 64 PEs. In the 2024 FSL-HDnn description, PEs on the same row share one input pixel bus, and PEs on the same column share one index/weight bus, enabling reuse of input data across rows and reuse of index or weight information across columns [2409.10918].

The 2025 FSL-HDnn paper gives the storage organization in more detail. The FE module contains an 8-bank, 128KB activation memory with double buffering, a 16-bank, 36KB index memory for weight indices, and a 16-bank, 4KB weight memory for BF16 codebooks. The PE array uses a codebook-stationary dataflow: codebook weights remain local, while indices and activations are streamed or broadcast to PEs. Each PE column corresponds to one output channel, and the four rows operate on different activations from consecutive rows in the input feature map so that four output pixels in consecutive rows are computed in parallel [2512.11826].

At PE granularity, the design is optimized for \(3 \times 3\) convolution kernels. The 2024 description states that each PE contains four register files, with three RFs accumulating input activations from three separate positions of the sliding convolution window and one RF used for the multiplication stage with actual weight values; the timing diagram overlaps accumulation in three RFs with multiplication in the fourth RF. The 2025 description gives the corresponding microarchitecture as four RFs, four adders, and one MAC unit, again with a time-overlapped schedule in which three RFs process three horizontally adjacent output pixels in parallel while the remaining RF performs MAC on a previous output pixel [2409.10918], [2512.11826].

Clo-HDnn preserves the same broad structure. Its WCFE uses a dedicated \(4 \times 16\) PE array with 4 RFs and 1 MAC per PE, and the paper reports operation over 50–250 MHz at 0.7–1.2 V. The stated purpose of this organization is to minimize idle time and improve efficiency [2507.17953].

## 4. Interface with HDC and end-to-end learning modes

WCFE is upstream of HDC rather than a learning engine in its own right. In FSL-HDnn, WCFE produces an \(F\)-dimensional feature vector, which is then encoded into a \(D\)-dimensional hypervector with \(D \gg F\). The 2025 formulation writes the HDC encoding as
\[
\mathbf{h} = \text{Encode}(\mathbf{x}) = \mathbf{B} \cdot \mathbf{x},
\]
where \(\mathbf{x} \in \mathbb{R}^F\) and \(\mathbf{B} \in \{-1,+1\}^{D \times F}\). Training aggregates encoded hypervectors into class hypervectors,
\[
\mathbf{C}_j = \sum_{i=1}^{k} \mathbf{h}_i^j,
\]
and inference is performed by nearest-class comparison,
\[
\underset{j}{\arg\min}~\text{Distance}(\mathbf{q}, \mathbf{C}_j).
\]
Only the HDC classifier is updated during few-shot learning; the CNN feature extractor remains frozen [2512.11826].

The same upstream role appears in Clo-HDnn. WCFE extracts features, after which the HD module performs Kronecker HD encoding, training or inference, and progressive search. The paper states that HDC class hypervectors are updated by adding or subtracting the query hypervector depending on inference correctness. WCFE is therefore the compressed feature generator that prepares data for gradient-free continual learning rather than the module that stores learned knowledge [2507.17953].

The later FSL-HDnn design extends this interface with two system-level optimizations that explicitly involve WCFE. First, early exit with branch feature extraction allows intermediate outputs of each ResNet-18 CONV block to be extracted and classified early; with \(E_s=2, E_c=2\), the paper reports skipping about 20–25% of layers with \(<1\%\) accuracy loss and about 32% inference latency or energy reduction. Second, batched single-pass training groups multiple samples per class so that FE extraction, HDC aggregation, and class-HV formation improve weight reuse and PE utilization; the reported effect is 18% to 32% per-image latency and energy savings over non-batched training [2512.11826].

## 5. Quantitative characteristics and reported implementations

The published HDnn papers report WCFE as both a compression method and a measured silicon subsystem. The following summary consolidates the WCFE-specific design points and outcomes explicitly reported in those works.

| System | WCFE configuration | Reported WCFE-related outcomes |
|---|---|---|
| FSL-HDnn [2409.10918] | Weight clustering with up to 16 unique weights per filter; 4-bit indices; 64 PEs in a \(4 \times 16\) array; optimized for \(3 \times 3\) kernels | 3.7× reduction in operations and 4.4× reduction in parameters on VGG16; 5.7 TOPS/W for feature extraction; feature-extractor power of 27 mW at 0.9 V; 2.6× higher peak TOPS/W than state-of-the-art CNN accelerators |
| Clo-HDnn [2507.17953] | Post-training weight clustering and pattern reuse; \(4 \times 16\) PE array; 4 RFs and 1 MAC per PE; 50–250 MHz; 0.7–1.2 V | Parameters reduced by 1.9×; CONV computations reduced by 2.1×; 1.44–4.66 TFLOPS/W for WCFE; WCFE accounts for 94.2% of total energy consumption and 87.7% of total latency; negligible accuracy drop compared to the FP baseline |
| FSL-HDnn [2512.11826] | \(Ch_{\text{sub}} = 64\); K-means clustering; BF16 codebooks; 8-bank 128KB activation memory; 16-bank 36KB index memory; 16-bank 4KB weight memory; codebook-stationary \(4 \times 16\) PE array | 1.8× memory saving and 2.1× computational reduction compared to INT8 baseline; full-chip training energy of 6 mJ/image; training throughput of 28 images/s; end-to-end training time of 1.7 s for a 10-way 5-shot task; 2× to 20.9× end-to-end training latency reduction over prior ODL chips |

Taken together, these figures indicate a consistent design objective: WCFE reduces the feedforward CNN cost sufficiently that HDC-based single-pass learning can become the dominant learning abstraction. A plausible implication is that the primary architectural value of WCFE lies less in isolated convolution speedup than in enabling an end-to-end on-device learning pipeline whose front-end remains practical under edge power and memory constraints.

## 6. Scope, tradeoffs, and related meanings

A common misconception is to equate WCFE with the learning mechanism of HDnn systems. The cited accelerator papers state the opposite: WCFE is the feature extractor, while HDC performs the few-shot or continual learning update. In Clo-HDnn, the distinction is strong enough that WCFE can be bypassed entirely for simpler datasets, which also reflects a central tradeoff: even after clustering and reuse optimizations, WCFE remains the dominant energy and latency component when it is active [2507.17953].

Another important scope condition is that WCFE assumes clustered weights are acceptable for the target workload. In the FSL-HDnn and Clo-HDnn descriptions, the benefit depends on repeated or similar trained convolution weights, compact index representations, and the ability to accumulate activations before multiplication. This suggests that WCFE is most naturally understood as a hardware-aware compressed inference method for frozen CNN front-ends rather than as a generic learning rule.

Outside accelerator design, the acronym can be read more generically. A survey of K-Means-based feature weighting describes feature weighting as a generalization of feature selection and notes that such methods are best seen as weighting or re-scaling existing features rather than discovering new transformed features [1601.03483]. The WISE framework goes further in a mixed-type tabular setting by unifying representation, feature weighting, clustering, and interpretation, and by treating multiple learned feature-weight vectors as semantic views that drive clustering and explanation [2604.05857]. These works are conceptually related only at the level of weighted feature relevance; they do not define the CNN accelerator WCFE used in FSL-HDnn and Clo-HDnn.

In the narrow hardware sense established by the HDnn literature, WCFE is therefore best characterized as a frozen, clustered, reuse-aware CNN feature extractor that compresses weights into codebooks and indices, restructures convolution around index-based accumulation and codebook multiplication, and supplies the compact feature embeddings required by downstream HDC classifiers for on-device few-shot and continual learning.

Source: https://www.emergentmind.com/topics/weight-clustering-feature-extraction-wcfe