Clo-HDnn: Continual Learning Accelerator
- Clo-HDnn is a hardware accelerator that combines CNN-like weight clustering and hyperdimensional computing, providing a compact solution for continual on-device learning.
- The chip replaces gradient-based backpropagation with gradient-free HDC updates, using simple hypervector additions and subtractions to minimize energy and memory overhead.
- Its dual-mode architecture, integrating selective feature extraction with a bypass path, achieves significant efficiency gains (e.g., 4.66 TFLOPS/W) for diverse edge applications.
Clo-HDnn is a hardware accelerator chip for continual on-device learning that integrates a compact CNN-like Weight Clustering Feature Extraction (WCFE) engine with a hyperdimensional computing (HDC) module built around a Kronecker HD Encoder and progressive search (Song et al., 23 Jul 2025). It is designed for edge devices that must learn from streaming or sequential data under tight energy, memory, and area constraints without cloud retraining, and it replaces gradient-based backpropagation with gradient-free HDC updates that store learned knowledge as class hypervectors (CHVs) (Song et al., 23 Jul 2025).
1. Problem setting and design rationale
Clo-HDnn targets emerging continual learning (CL) workloads in which data arrive incrementally and the deployed system must update its model in situ. The design is motivated by three constraints identified in on-device learning (ODL): the high cost of gradient-based training on-device, catastrophic forgetting in continual learning, and the hardware overheads of naïve HDC implementations, especially large encoders, exhaustive similarity search, and large associative memories for high-dimensional class representations (Song et al., 23 Jul 2025).
The accelerator addresses these constraints by combining compressed feature extraction with gradient-free HDC. In this formulation, prior knowledge is stored in CHVs rather than in a backpropagation-trained parameter stack. Continual updates are therefore reduced to simple additions and subtractions on hypervectors, eliminating gradient buffers and stochastic-gradient-descent-style optimization in the deployed device (Song et al., 23 Jul 2025). This makes Clo-HDnn particularly suited to resource-constrained edge scenarios in which adaptation must occur continuously and locally.
A central architectural premise is that HDC can act as the continual-learning substrate while WCFE supplies task-dependent features only when necessary. The paper therefore presents Clo-HDnn as a full-chip, end-to-end continual on-device learning accelerator integrating WCFE and HDC, rather than as an HDC classifier in isolation (Song et al., 23 Jul 2025).
2. Macro-architecture and operating modes
Clo-HDnn consists of three main blocks: a WCFE engine, an HD module, and a global FIFO / CDC interconnect that also enables dual-mode operation (Song et al., 23 Jul 2025).
| Block | Function | Notable implementation detail |
|---|---|---|
| WCFE engine | CNN-like feature extraction with weight clustering | 4×16 PE array |
| HD module | Kronecker encoding, HDC inference, and CHV update | XOR-tree search and CHV cache |
| Global FIFO / CDC | Interconnect, buffering, and dual-mode routing | Supports bypass and normal dataflow |
The WCFE engine is used for complex datasets such as CIFAR-100, where raw data first undergo feature extraction before entering the HD module. In this normal mode, WCFE outputs compact feature vectors that are then encoded into query hypervectors (QHVs) by the Kronecker HD Encoder and classified by the HDC subsystem (Song et al., 23 Jul 2025).
The same chip can also operate in bypass mode. For simpler datasets such as ISOLET and UCIHAR, raw or preprocessed features are sent directly to the HD module, bypassing WCFE entirely (Song et al., 23 Jul 2025). This dual-mode arrangement is significant because the chip-level measurements show that on CIFAR-100, WCFE contributes 94.2% of total energy and 87.7% of latency, making feature extraction the dominant cost on complex vision tasks (Song et al., 23 Jul 2025). The bypass path is therefore not an auxiliary feature but a primary mechanism for matching the compute path to the structure of the dataset.
The overall architecture also includes a control core with a custom ISA, which gives the chip a degree of programmability beyond fixed-function datapaths (Song et al., 23 Jul 2025). That combination of programmable control, mode switching, and modular HDC training/inference is central to the chip’s positioning as an end-to-end CL accelerator rather than a single-kernel accelerator.
3. Hyperdimensional continual learning pipeline
In Clo-HDnn, information is represented as high-dimensional vectors, and hypervector elements are stored as INT8 values in the CHV cache (Song et al., 23 Jul 2025). Each class is associated with a class hypervector , while each input sample is encoded into a query hypervector . The CHVs function as class prototypes in hyperdimensional space (Song et al., 23 Jul 2025).
Inference proceeds by generating a QHV from either WCFE features or bypassed raw features, then comparing that QHV against the stored CHVs. The HD Search module uses XOR-tree-based distance computation; the text specifies that the 64-b MSBs of each CHV are fetched and compared with the segmented QHV using the XOR tree (Song et al., 23 Jul 2025). In the binary case, this corresponds to a Hamming-style distance: The predicted class is the best match under the distance or similarity measure implemented in hardware (Song et al., 23 Jul 2025).
Training is single-pass and gradient-free. For each labeled example, the system first performs inference to obtain a prediction and then updates the relevant CHVs by adding or subtracting the QHV according to correctness (Song et al., 23 Jul 2025). The reinforcement rule is
and when prediction is incorrect, the true and predicted classes are adjusted as
Because the update is localized to CHVs, the system does not backpropagate through a deep network and does not require gradient memory (Song et al., 23 Jul 2025).
This update rule also supports incremental class addition. A new class can be introduced by allocating a new CHV slot in the CHV cache, initializing it, and then updating it as labeled examples arrive (Song et al., 23 Jul 2025). The paper states that catastrophic forgetting is mitigated because prior knowledge remains encoded in existing CHVs and new data add information rather than overwriting a shared weight tensor (Song et al., 23 Jul 2025). Within the paper’s scope, this is the principal continual-learning argument for HDC as deployed in Clo-HDnn.
4. Encoder, feature compression, and progressive search
Encoding is identified as the main bottleneck in HDC. Clo-HDnn addresses that bottleneck with a Kronecker HD Encoder that uses a Kronecker-product-structured weight matrix in place of a full dense projection matrix (Song et al., 23 Jul 2025). Conceptually, if is the input feature vector and the encoder matrix, the encoder computes
but is represented through a Kronecker structure,
0
so the hardware can implement the encoding by two-stage reshaping and block matrix multiplications over smaller base matrices (Song et al., 23 Jul 2025).
The resulting savings are substantial in the paper’s measurements: the Kronecker HD Encoder achieves a 43× speedup and 1376× memory capacity savings over existing HD encoders (Song et al., 23 Jul 2025). Its microarchitecture includes an 8-bank 1KB weight buffer, 256-bit weight fetch per cycle, feature streaming into 32 parallel 8-to-1 adder trees, and binary-INT encoding that replaces multiplications with additions because the weights use limited discrete levels (Song et al., 23 Jul 2025).
Feature extraction efficiency is handled separately by WCFE. WCFE begins from a conventional CNN pretrained offline; after training, weights with similar values are clustered into a codebook, and each original weight is replaced by a codebook index (Song et al., 23 Jul 2025). This reduces the number of unique parameters by 1.9× and allows pattern reuse in convolution, where input positions sharing a weight index are accumulated before a shared multiplication (Song et al., 23 Jul 2025). The resulting effect is a 2.1× reduction in convolution computations (Song et al., 23 Jul 2025). The WCFE hardware implements this with a 4×16 processing-element array; each PE contains 4 register files for local accumulation and 1 MAC unit (Song et al., 23 Jul 2025).
A third efficiency mechanism is progressive search. Instead of always computing and comparing full-length QHVs, Clo-HDnn first generates only a partial QHV segment, fetches the corresponding partial CHV segments, and computes partial distances (Song et al., 23 Jul 2025). It then evaluates the margin between the best and second-best classes; if the margin exceeds a preset confidence threshold, the search terminates early. Otherwise, additional QHV and CHV segments are generated and compared (Song et al., 23 Jul 2025). This reduces encoding operations, CHV-cache bandwidth, and search cost, and the paper reports up to 61% complexity reduction with negligible accuracy loss (Song et al., 23 Jul 2025).
Taken together, WCFE, the Kronecker HD Encoder, and progressive search partition the efficiency problem into three distinct domains: front-end feature generation, HDC encoding, and associative search. That decomposition is one of the defining characteristics of the design.
5. Fabrication, measurements, and reported performance
Clo-HDnn was fabricated in 40 nm CMOS with an area of 14.4 mm1 (Song et al., 23 Jul 2025). The chip includes the WCFE subsystem, Kronecker encoder, HDC search engine, HDC training engine, CHV cache, global FIFO / CDC interface, and the custom-ISA control core (Song et al., 23 Jul 2025). Both the WCFE and HDC classifier operate over 50–250 MHz and 0.7–1.2 V (Song et al., 23 Jul 2025).
The paper defines energy efficiency as
2
Under this definition, WCFE achieves 1.44–4.66 TFLOPS/W, while the HDC classifier achieves 1.29–3.78 TOPS/W (Song et al., 23 Jul 2025). The peak figures highlighted in the title and abstract are 4.66 TFLOPS/W for feature extraction and 3.78 TOPS/W for the classifier (Song et al., 23 Jul 2025). Relative to state-of-the-art ODL accelerators, the paper reports 7.77× higher energy efficiency for feature extraction and 4.85× higher energy efficiency for the classifier (Song et al., 23 Jul 2025).
Evaluation is reported on three benchmarks. CIFAR-100 is used in normal mode as a complex dataset requiring WCFE; ISOLET and UCIHAR are used in bypass mode as simpler tasks with lower-dimensional or pre-extracted features (Song et al., 23 Jul 2025). According to the paper, Clo-HDnn shows negligible accuracy drop compared to a floating-point baseline for both the bypass-mode and normal-mode settings, indicating that WCFE compression, Kronecker encoding, and progressive search do not significantly damage classification performance (Song et al., 23 Jul 2025).
The paper further reports Clo-HDnn as the first chip to support end-to-end CL for HDC tasks, combining feature extraction with HD-based classification in a fabricated accelerator (Song et al., 23 Jul 2025). Within the measured workload mix, WCFE dominates cost on complex vision tasks, while HDC provides the low-cost adaptation substrate. The empirical significance of the design is therefore not only its peak efficiency numbers but also its asymmetric compute profile: expensive feature extraction is used selectively, while learning updates remain inexpensive.
6. Scope, limitations, and nomenclature
The reported scope of Clo-HDnn is classification under continual-learning scenarios, specifically images, spoken-letter features, and smartphone sensor data (Song et al., 23 Jul 2025). The paper does not provide explicit support or results for complex sequence models such as RNNs or Transformers, nor for regression or structured prediction (Song et al., 23 Jul 2025). It also assumes fixed-dimensional hypervectors with INT8 elements, and it notes that CHV capacity may limit scaling to very large numbers of classes or tasks given fixed on-chip memory (Song et al., 23 Jul 2025). A further design assumption is that the structured Kronecker projection remains sufficiently discriminative for the evaluated tasks, even though it is less flexible than a fully random projection matrix (Song et al., 23 Jul 2025).
These constraints delimit the accelerator’s intended use cases. The paper positions Clo-HDnn for IoT nodes, embedded sensors, wearables, mobile or robotic platforms, and industrial or environmental sensors that must continuously adapt without cloud connectivity (Song et al., 23 Jul 2025). That application profile follows directly from the chip’s gradient-free update path, compact memory footprint, dual-mode dataflow, and measured energy efficiency.
The name also benefits from disambiguation. In other arXiv literature, “HDNN” can denote a highway deep neural network for small-footprint acoustic modeling (Lu et al., 2016) or a hybrid analog-digital deep neural network for mmWave massive-MIMO systems (Morsali et al., 2021). Clo-HDnn, by contrast, refers specifically to a continual on-device learning accelerator that combines WCFE with HDC, a Kronecker HD Encoder, and progressive search in a fabricated 40 nm CMOS chip (Song et al., 23 Jul 2025).