---
title: 'Clo-HDnn: Continual Learning Accelerator'
url: https://www.emergentmind.com/topics/clo-hdnn
type: topic
---

# Clo-HDnn: Continual Learning Accelerator

Clo-HDnn is a hardware accelerator chip for continual on-device learning that integrates a compact CNN-like Weight Clustering Feature Extraction (WCFE) engine with a hyperdimensional computing (HDC) module built around a Kronecker HD Encoder and progressive search [2507.17953]. It is designed for edge devices that must learn from streaming or sequential data under tight energy, memory, and area constraints without cloud retraining, and it replaces gradient-based backpropagation with gradient-free HDC updates that store learned knowledge as class hypervectors (CHVs) [2507.17953].

## 1. Problem setting and design rationale

Clo-HDnn targets emerging continual learning (CL) workloads in which data arrive incrementally and the deployed system must update its model in situ. The design is motivated by three constraints identified in on-device learning (ODL): the high cost of gradient-based training on-device, catastrophic forgetting in continual learning, and the hardware overheads of naïve HDC implementations, especially large encoders, exhaustive similarity search, and large associative memories for high-dimensional class representations [2507.17953].

The accelerator addresses these constraints by combining compressed feature extraction with gradient-free HDC. In this formulation, prior knowledge is stored in CHVs rather than in a backpropagation-trained parameter stack. Continual updates are therefore reduced to simple additions and subtractions on hypervectors, eliminating gradient buffers and stochastic-gradient-descent-style optimization in the deployed device [2507.17953]. This makes Clo-HDnn particularly suited to resource-constrained edge scenarios in which adaptation must occur continuously and locally.

A central architectural premise is that HDC can act as the continual-learning substrate while WCFE supplies task-dependent features only when necessary. The paper therefore presents Clo-HDnn as a full-chip, end-to-end continual on-device learning accelerator integrating WCFE and HDC, rather than as an HDC classifier in isolation [2507.17953].

## 2. Macro-architecture and operating modes

Clo-HDnn consists of three main blocks: a WCFE engine, an HD module, and a global FIFO / CDC interconnect that also enables dual-mode operation [2507.17953].

| Block | Function | Notable implementation detail |
|---|---|---|
| WCFE engine | CNN-like feature extraction with weight clustering | 4×16 PE array |
| HD module | Kronecker encoding, HDC inference, and CHV update | XOR-tree search and CHV cache |
| Global FIFO / CDC | Interconnect, buffering, and dual-mode routing | Supports bypass and normal dataflow |

The WCFE engine is used for complex datasets such as CIFAR-100, where raw data first undergo feature extraction before entering the HD module. In this normal mode, WCFE outputs compact feature vectors that are then encoded into query hypervectors (QHVs) by the Kronecker HD Encoder and classified by the HDC subsystem [2507.17953].

The same chip can also operate in bypass mode. For simpler datasets such as ISOLET and UCIHAR, raw or preprocessed features are sent directly to the HD module, bypassing WCFE entirely [2507.17953]. This dual-mode arrangement is significant because the chip-level measurements show that on CIFAR-100, WCFE contributes 94.2% of total energy and 87.7% of latency, making feature extraction the dominant cost on complex vision tasks [2507.17953]. The bypass path is therefore not an auxiliary feature but a primary mechanism for matching the compute path to the structure of the dataset.

The overall architecture also includes a control core with a custom ISA, which gives the chip a degree of programmability beyond fixed-function datapaths [2507.17953]. That combination of programmable control, mode switching, and modular HDC training/inference is central to the chip’s positioning as an end-to-end CL accelerator rather than a single-kernel accelerator.

## 3. Hyperdimensional continual learning pipeline

In Clo-HDnn, information is represented as high-dimensional vectors, and hypervector elements are stored as INT8 values in the CHV cache [2507.17953]. Each class \(c\) is associated with a class hypervector \(\mathbf{H}_c\), while each input sample is encoded into a query hypervector \(\mathbf{q}\). The CHVs function as class prototypes in hyperdimensional space [2507.17953].

Inference proceeds by generating a QHV from either WCFE features or bypassed raw features, then comparing that QHV against the stored CHVs. The HD Search module uses XOR-tree-based distance computation; the text specifies that the 64-b MSBs of each CHV are fetched and compared with the segmented QHV using the XOR tree [2507.17953]. In the binary case, this corresponds to a Hamming-style distance:
\[
d_H(\mathbf{h}_1,\mathbf{h}_2)=\sum_{i=1}^{D} (\mathbf{h}_1[i]\oplus \mathbf{h}_2[i]).
\]
The predicted class is the best match under the distance or similarity measure implemented in hardware [2507.17953].

Training is single-pass and gradient-free. For each labeled example, the system first performs inference to obtain a prediction and then updates the relevant CHVs by adding or subtracting the QHV according to correctness [2507.17953]. The reinforcement rule is
\[
\mathbf{H}_c \leftarrow \mathbf{H}_c + \mathbf{q},
\]
and when prediction is incorrect, the true and predicted classes are adjusted as
\[
\mathbf{H}_{c_{\text{true}}} \leftarrow \mathbf{H}_{c_{\text{true}}} + \mathbf{q}, \qquad
\mathbf{H}_{c_{\text{pred}}} \leftarrow \mathbf{H}_{c_{\text{pred}}} - \mathbf{q}.
\]
Because the update is localized to CHVs, the system does not backpropagate through a deep network and does not require gradient memory [2507.17953].

This update rule also supports incremental class addition. A new class can be introduced by allocating a new CHV slot in the CHV cache, initializing it, and then updating it as labeled examples arrive [2507.17953]. The paper states that catastrophic forgetting is mitigated because prior knowledge remains encoded in existing CHVs and new data add information rather than overwriting a shared weight tensor [2507.17953]. Within the paper’s scope, this is the principal continual-learning argument for HDC as deployed in Clo-HDnn.

## 4. Encoder, feature compression, and progressive search

Encoding is identified as the main bottleneck in HDC. Clo-HDnn addresses that bottleneck with a Kronecker HD Encoder that uses a Kronecker-product-structured weight matrix in place of a full dense projection matrix [2507.17953]. Conceptually, if \(\mathbf{x}\) is the input feature vector and \(\mathbf{W}\) the encoder matrix, the encoder computes
\[
\mathbf{q}=\mathbf{W}\mathbf{x},
\]
but \(\mathbf{W}\) is represented through a Kronecker structure,
\[
\mathbf{W}\approx \mathbf{A}\otimes \mathbf{B}\otimes \cdots,
\]
so the hardware can implement the encoding by two-stage reshaping and block matrix multiplications over smaller base matrices [2507.17953].

The resulting savings are substantial in the paper’s measurements: the Kronecker HD Encoder achieves a 43× speedup and 1376× memory capacity savings over existing HD encoders [2507.17953]. Its microarchitecture includes an 8-bank 1KB weight buffer, 256-bit weight fetch per cycle, feature streaming into 32 parallel 8-to-1 adder trees, and binary-INT encoding that replaces multiplications with additions because the weights use limited discrete levels [2507.17953].

Feature extraction efficiency is handled separately by WCFE. WCFE begins from a conventional CNN pretrained offline; after training, weights with similar values are clustered into a codebook, and each original weight is replaced by a codebook index [2507.17953]. This reduces the number of unique parameters by 1.9× and allows pattern reuse in convolution, where input positions sharing a weight index are accumulated before a shared multiplication [2507.17953]. The resulting effect is a 2.1× reduction in convolution computations [2507.17953]. The WCFE hardware implements this with a 4×16 processing-element array; each PE contains 4 register files for local accumulation and 1 MAC unit [2507.17953].

A third efficiency mechanism is progressive search. Instead of always computing and comparing full-length QHVs, Clo-HDnn first generates only a partial QHV segment, fetches the corresponding partial CHV segments, and computes partial distances [2507.17953]. It then evaluates the margin between the best and second-best classes; if the margin exceeds a preset confidence threshold, the search terminates early. Otherwise, additional QHV and CHV segments are generated and compared [2507.17953]. This reduces encoding operations, CHV-cache bandwidth, and search cost, and the paper reports up to 61% complexity reduction with negligible accuracy loss [2507.17953].

Taken together, WCFE, the Kronecker HD Encoder, and progressive search partition the efficiency problem into three distinct domains: front-end feature generation, HDC encoding, and associative search. That decomposition is one of the defining characteristics of the design.

## 5. Fabrication, measurements, and reported performance

Clo-HDnn was fabricated in 40 nm CMOS with an area of 14.4 mm\(^2\) [2507.17953]. The chip includes the WCFE subsystem, Kronecker encoder, HDC search engine, HDC training engine, CHV cache, global FIFO / CDC interface, and the custom-ISA control core [2507.17953]. Both the WCFE and HDC classifier operate over 50–250 MHz and 0.7–1.2 V [2507.17953].

The paper defines energy efficiency as
\[
\text{Energy efficiency}=\frac{\text{Throughput (operations/s)}}{\text{Power (W)}}.
\]
Under this definition, WCFE achieves 1.44–4.66 TFLOPS/W, while the HDC classifier achieves 1.29–3.78 TOPS/W [2507.17953]. The peak figures highlighted in the title and abstract are 4.66 TFLOPS/W for feature extraction and 3.78 TOPS/W for the classifier [2507.17953]. Relative to state-of-the-art ODL accelerators, the paper reports 7.77× higher energy efficiency for feature extraction and 4.85× higher energy efficiency for the classifier [2507.17953].

Evaluation is reported on three benchmarks. CIFAR-100 is used in normal mode as a complex dataset requiring WCFE; ISOLET and UCIHAR are used in bypass mode as simpler tasks with lower-dimensional or pre-extracted features [2507.17953]. According to the paper, Clo-HDnn shows negligible accuracy drop compared to a floating-point baseline for both the bypass-mode and normal-mode settings, indicating that WCFE compression, Kronecker encoding, and progressive search do not significantly damage classification performance [2507.17953].

The paper further reports Clo-HDnn as the first chip to support end-to-end CL for HDC tasks, combining feature extraction with HD-based classification in a fabricated accelerator [2507.17953]. Within the measured workload mix, WCFE dominates cost on complex vision tasks, while HDC provides the low-cost adaptation substrate. The empirical significance of the design is therefore not only its peak efficiency numbers but also its asymmetric compute profile: expensive feature extraction is used selectively, while learning updates remain inexpensive.

## 6. Scope, limitations, and nomenclature

The reported scope of Clo-HDnn is classification under continual-learning scenarios, specifically images, spoken-letter features, and smartphone sensor data [2507.17953]. The paper does not provide explicit support or results for complex sequence models such as RNNs or Transformers, nor for regression or structured prediction [2507.17953]. It also assumes fixed-dimensional hypervectors with INT8 elements, and it notes that CHV capacity may limit scaling to very large numbers of classes or tasks given fixed on-chip memory [2507.17953]. A further design assumption is that the structured Kronecker projection remains sufficiently discriminative for the evaluated tasks, even though it is less flexible than a fully random projection matrix [2507.17953].

These constraints delimit the accelerator’s intended use cases. The paper positions Clo-HDnn for IoT nodes, embedded sensors, wearables, mobile or robotic platforms, and industrial or environmental sensors that must continuously adapt without cloud connectivity [2507.17953]. That application profile follows directly from the chip’s gradient-free update path, compact memory footprint, dual-mode dataflow, and measured energy efficiency.

The name also benefits from disambiguation. In other arXiv literature, “HDNN” can denote a highway deep neural network for small-footprint acoustic modeling [1608.00892] or a hybrid analog-digital deep neural network for mmWave massive-MIMO systems [2107.14704]. Clo-HDnn, by contrast, refers specifically to a continual on-device learning accelerator that combines WCFE with HDC, a Kronecker HD Encoder, and progressive search in a fabricated 40 nm CMOS chip [2507.17953].

Source: https://www.emergentmind.com/topics/clo-hdnn