Papers
Topics
Authors
Recent
Search
2000 character limit reached

WCFE: Weight Clustering for CNN Extraction

Updated 7 July 2026
  • WCFE is a hardware-aware CNN feature extractor that clusters similar weights into low-bit indices to reduce computation and memory without sacrificing feature quality.
  • It operates as a frozen front-end module, feeding compact features to an adaptable HDC classifier for efficient few-shot and continual learning.
  • The approach achieves notable savings, such as up to 3.7× reduction in operations and 4.4× parameter reduction, optimizing energy and latency for on-device learning.

Weight Clustering Feature Extraction (WCFE) denotes a weight-clustering-based CNN feature extractor used as the front-end of HDnn-family on-device learning accelerators, notably FSL-HDnn and Clo-HDnn. In this usage, WCFE is a hardware-efficient convolution engine that reduces compute, memory, and data movement by clustering similar weights into shared representative values, storing weights as low-bit indices rather than full-precision values, accumulating input activations by weight cluster before multiplication, and reusing clustered activation patterns across filters or channels. Within these systems, WCFE provides the feature vector consumed by a downstream Hyperdimensional Computing (HDC) module; the feature extractor is kept frozen during adaptation, while learning is performed in the classifier stage (Yang et al., 2024, Song et al., 23 Jul 2025, Xu et al., 2 Dec 2025).

1. Definition and architectural position

In FSL-HDnn, WCFE is the CNN-based feature extraction module of an end-to-end few-shot learning pipeline whose two major parts are feature extraction and HDC-based classification or learning. The feature extractor uses “per-filter weight clustering and pattern sharing across filters,” and it remains frozen during few-shot learning, so on-chip adaptation updates only the HDC classifier rather than CNN weights (Yang et al., 2024).

In Clo-HDnn, WCFE is likewise the CNN-like feature extraction front-end, but it is embedded in a continual on-device learning architecture composed of a WCFE module, an HD module for encoding, training, and inference, and a global FIFO connecting them. Clo-HDnn further supports dual-mode operation: for simple datasets, WCFE can be bypassed and raw inputs can be sent directly toward the HD pipeline, whereas more complex datasets use WCFE before HDC (Song et al., 23 Jul 2025).

The 2025 FSL-HDnn formulation makes the same division of labor explicit. WCFE is described there as a parameter-efficient feature extractor that keeps the CNN feature extractor frozen after ImageNet pretraining, applies weight clustering to pretrained CNN weights, and uses the clustered representation to reduce memory footprint, data movement, and MAC operations while preserving feature quality for the downstream classifier (Xu et al., 2 Dec 2025).

2. Weight clustering and convolutional computation

The core WCFE mechanism is weight clustering. Similar weights are clustered into the same average value, so a conventional convolution filter weight tensor is approximated by a small set of shared representative values. In the FSL-HDnn description, prior studies are cited as indicating that up to 16 unique weights per filter can preserve accuracy comparable to unclustered feature extraction; this corresponds to 4-bit indices, with storage split into cluster representative values and a 4-bit index identifying the cluster membership of each original weight (Yang et al., 2024).

Operationally, WCFE exploits repeated weight indices. Input activations associated with the same weight index are first accumulated and are then multiplied by the corresponding centroid or codebook value. The later FSL-HDnn formulation states this explicitly in clustered form: activations with the same index are accumulated, then NN accumulated inputs are multiplied with NN codebook weights, and the final output is formed by summing those results. It also gives the corresponding operation-count change: the original convolution cost is

2×K2−12 \times K^2 - 1

whereas the clustered version becomes

K2+N−1.K^2 + N - 1.

This is the direct computational rationale for fewer multiplications and reduced redundant accumulation (Xu et al., 2 Dec 2025).

A second optimization is pattern reuse across filters. In the 2024 FSL-HDnn description, the clustering pattern is shared across filters for different channels so that accumulated input pixels can be reused by the filters for many output channels. In Clo-HDnn, the same idea is described as grouping, accumulating, and multiplying once for inputs sharing the same weight value during convolution. This suggests that WCFE is not merely weight compression; it is a reuse-aware execution scheme whose efficiency depends on both shared values and shared clustering patterns (Yang et al., 2024, Song et al., 23 Jul 2025).

3. Hardware organization and PE-level execution

WCFE is realized as a specialized convolution accelerator. Across both HDnn implementations, the compute core is a 4×164 \times 16 PE array, corresponding to 64 PEs. In the 2024 FSL-HDnn description, PEs on the same row share one input pixel bus, and PEs on the same column share one index/weight bus, enabling reuse of input data across rows and reuse of index or weight information across columns (Yang et al., 2024).

The 2025 FSL-HDnn paper gives the storage organization in more detail. The FE module contains an 8-bank, 128KB activation memory with double buffering, a 16-bank, 36KB index memory for weight indices, and a 16-bank, 4KB weight memory for BF16 codebooks. The PE array uses a codebook-stationary dataflow: codebook weights remain local, while indices and activations are streamed or broadcast to PEs. Each PE column corresponds to one output channel, and the four rows operate on different activations from consecutive rows in the input feature map so that four output pixels in consecutive rows are computed in parallel (Xu et al., 2 Dec 2025).

At PE granularity, the design is optimized for 3×33 \times 3 convolution kernels. The 2024 description states that each PE contains four register files, with three RFs accumulating input activations from three separate positions of the sliding convolution window and one RF used for the multiplication stage with actual weight values; the timing diagram overlaps accumulation in three RFs with multiplication in the fourth RF. The 2025 description gives the corresponding microarchitecture as four RFs, four adders, and one MAC unit, again with a time-overlapped schedule in which three RFs process three horizontally adjacent output pixels in parallel while the remaining RF performs MAC on a previous output pixel (Yang et al., 2024, Xu et al., 2 Dec 2025).

Clo-HDnn preserves the same broad structure. Its WCFE uses a dedicated 4×164 \times 16 PE array with 4 RFs and 1 MAC per PE, and the paper reports operation over 50–250 MHz at 0.7–1.2 V. The stated purpose of this organization is to minimize idle time and improve efficiency (Song et al., 23 Jul 2025).

4. Interface with HDC and end-to-end learning modes

WCFE is upstream of HDC rather than a learning engine in its own right. In FSL-HDnn, WCFE produces an FF-dimensional feature vector, which is then encoded into a DD-dimensional hypervector with D≫FD \gg F. The 2025 formulation writes the HDC encoding as

NN0

where NN1 and NN2. Training aggregates encoded hypervectors into class hypervectors,

NN3

and inference is performed by nearest-class comparison,

NN4

Only the HDC classifier is updated during few-shot learning; the CNN feature extractor remains frozen (Xu et al., 2 Dec 2025).

The same upstream role appears in Clo-HDnn. WCFE extracts features, after which the HD module performs Kronecker HD encoding, training or inference, and progressive search. The paper states that HDC class hypervectors are updated by adding or subtracting the query hypervector depending on inference correctness. WCFE is therefore the compressed feature generator that prepares data for gradient-free continual learning rather than the module that stores learned knowledge (Song et al., 23 Jul 2025).

The later FSL-HDnn design extends this interface with two system-level optimizations that explicitly involve WCFE. First, early exit with branch feature extraction allows intermediate outputs of each ResNet-18 CONV block to be extracted and classified early; with NN5, the paper reports skipping about 20–25% of layers with NN6 accuracy loss and about 32% inference latency or energy reduction. Second, batched single-pass training groups multiple samples per class so that FE extraction, HDC aggregation, and class-HV formation improve weight reuse and PE utilization; the reported effect is 18% to 32% per-image latency and energy savings over non-batched training (Xu et al., 2 Dec 2025).

5. Quantitative characteristics and reported implementations

The published HDnn papers report WCFE as both a compression method and a measured silicon subsystem. The following summary consolidates the WCFE-specific design points and outcomes explicitly reported in those works.

System WCFE configuration Reported WCFE-related outcomes
FSL-HDnn (Yang et al., 2024) Weight clustering with up to 16 unique weights per filter; 4-bit indices; 64 PEs in a NN7 array; optimized for NN8 kernels 3.7× reduction in operations and 4.4× reduction in parameters on VGG16; 5.7 TOPS/W for feature extraction; feature-extractor power of 27 mW at 0.9 V; 2.6× higher peak TOPS/W than state-of-the-art CNN accelerators
Clo-HDnn (Song et al., 23 Jul 2025) Post-training weight clustering and pattern reuse; NN9 PE array; 4 RFs and 1 MAC per PE; 50–250 MHz; 0.7–1.2 V Parameters reduced by 1.9×; CONV computations reduced by 2.1×; 1.44–4.66 TFLOPS/W for WCFE; WCFE accounts for 94.2% of total energy consumption and 87.7% of total latency; negligible accuracy drop compared to the FP baseline
FSL-HDnn (Xu et al., 2 Dec 2025) 2×K2−12 \times K^2 - 10; K-means clustering; BF16 codebooks; 8-bank 128KB activation memory; 16-bank 36KB index memory; 16-bank 4KB weight memory; codebook-stationary 2×K2−12 \times K^2 - 11 PE array 1.8× memory saving and 2.1× computational reduction compared to INT8 baseline; full-chip training energy of 6 mJ/image; training throughput of 28 images/s; end-to-end training time of 1.7 s for a 10-way 5-shot task; 2× to 20.9× end-to-end training latency reduction over prior ODL chips

Taken together, these figures indicate a consistent design objective: WCFE reduces the feedforward CNN cost sufficiently that HDC-based single-pass learning can become the dominant learning abstraction. A plausible implication is that the primary architectural value of WCFE lies less in isolated convolution speedup than in enabling an end-to-end on-device learning pipeline whose front-end remains practical under edge power and memory constraints.

A common misconception is to equate WCFE with the learning mechanism of HDnn systems. The cited accelerator papers state the opposite: WCFE is the feature extractor, while HDC performs the few-shot or continual learning update. In Clo-HDnn, the distinction is strong enough that WCFE can be bypassed entirely for simpler datasets, which also reflects a central tradeoff: even after clustering and reuse optimizations, WCFE remains the dominant energy and latency component when it is active (Song et al., 23 Jul 2025).

Another important scope condition is that WCFE assumes clustered weights are acceptable for the target workload. In the FSL-HDnn and Clo-HDnn descriptions, the benefit depends on repeated or similar trained convolution weights, compact index representations, and the ability to accumulate activations before multiplication. This suggests that WCFE is most naturally understood as a hardware-aware compressed inference method for frozen CNN front-ends rather than as a generic learning rule.

Outside accelerator design, the acronym can be read more generically. A survey of K-Means-based feature weighting describes feature weighting as a generalization of feature selection and notes that such methods are best seen as weighting or re-scaling existing features rather than discovering new transformed features (Amorim, 2015). The WISE framework goes further in a mixed-type tabular setting by unifying representation, feature weighting, clustering, and interpretation, and by treating multiple learned feature-weight vectors as semantic views that drive clustering and explanation (Li et al., 7 Apr 2026). These works are conceptually related only at the level of weighted feature relevance; they do not define the CNN accelerator WCFE used in FSL-HDnn and Clo-HDnn.

In the narrow hardware sense established by the HDnn literature, WCFE is therefore best characterized as a frozen, clustered, reuse-aware CNN feature extractor that compresses weights into codebooks and indices, restructures convolution around index-based accumulation and codebook multiplication, and supplies the compact feature embeddings required by downstream HDC classifiers for on-device few-shot and continual learning.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Weight Clustering Feature Extraction (WCFE).