Papers
Topics
Authors
Recent
Search
2000 character limit reached

HD Spam Filtering Accelerator

Updated 30 December 2025
  • HD Spam Filtering Accelerator is a system that uses hyperdimensional n-gram encoding to represent spam-related textual features efficiently.
  • It employs random bipolar vectors with cyclic permutations to capture order and frequency, offering significant speed and memory improvements over traditional methods.
  • Empirical results show competitive filtering accuracy alongside substantial computational and memory savings, making it ideal for large-scale spam detection.

N-gram HD (Hyperdimensional) Encoders are methods for representing the distributional statistics of contiguous character or token sequences (n-grams) in fixed-length, high-dimensional vectors. Such encoders leverage principles from hyperdimensional computing and vector-symbolic architectures to yield compact representations that capture order, frequency, and context, providing scalable alternatives to traditional n-gram histograms. Recent advances, exemplified by HyperEmbed and ZEN 2.0, demonstrate that n-gram HD encoders support efficient memory and computation trade-offs while achieving competitive or state-of-the-art results across domains, languages, and tasks (Alonso et al., 2020, Song et al., 2021).

1. Foundations of Hyperdimensional N-gram Encoding

Hyperdimensional computing (HDC) leverages random vectors in very high-dimensional spaces (D=103D=10^3–10510^5) where base vectors are generated i.i.d. and possess near-orthogonality properties. In n-gram encoding, each symbol S∈ΣS\in\Sigma is assigned a random bipolar vector vS∈{+1,−1}Dv_S\in\{+1,-1\}^D stored in an item memory. Positional information is encoded via cyclic permutations ρj(vS)\rho^j(v_S) where jj indexes the position within an n-gram. Binding is performed as element-wise multiplication:

vw=⨀j=1nρj(vSj),w=S1S2⋯Snv_w = \bigodot_{j=1}^{n} \rho^j(v_{S_j}),\quad w=S_1S_2\cdots S_n

All observed n-grams in a document are bundled via weighted summation, resulting in a single D-dimensional vector VV:

V=∑w∈Σnc(w) vwV = \sum_{w\in\Sigma^n} c(w)\,v_w

Normalization (e.g., l2l_2) is typical prior to downstream task application (Alonso et al., 2020).

2. Algorithmic Implementations and Pseudocode

The standard sliding-window algorithm processes text to construct HD representations for all character n-grams. For document 10510^50, alphabet 10510^51, n-gram length 10510^52, and HD dimension 10510^53:

  • Initialize: For each 10510^54, generate random 10510^55.
  • Process: For each position 10510^56, bind and permute vectors for the n-gram 10510^57.
  • Bundle: Accumulate the resulting 10510^58 into 10510^59.
  • Normalize: S∈ΣS\in\Sigma0.

Time complexity for document encoding is S∈ΣS\in\Sigma1 (Alonso et al., 2020):

VV8

3. Trade-offs: Dimensionality, Fidelity, and Resource Efficiency

Classical n-gram histograms require S∈ΣS\in\Sigma2 counters for alphabet size S∈ΣS\in\Sigma3. HD encoders collapse all n-gram statistics into a single S∈ΣS\in\Sigma4-dimensional vector—S∈ΣS\in\Sigma5 is independent of S∈ΣS\in\Sigma6. Increasing S∈ΣS\in\Sigma7 improves representational fidelity and classification accuracy, while reducing S∈ΣS\in\Sigma8 offers substantial savings in memory and computation.

Empirically, accuracy (micro-averaged S∈ΣS\in\Sigma9) grows with vS∈{+1,−1}Dv_S\in\{+1,-1\}^D0 until a saturation point (vS∈{+1,−1}Dv_S\in\{+1,-1\}^D1): vS∈{+1,−1}Dv_S\in\{+1,-1\}^D2512 for small corpora, vS∈{+1,−1}Dv_S\in\{+1,-1\}^D34096 for large. Beyond vS∈{+1,−1}Dv_S\in\{+1,-1\}^D4, additional gains are negligible or negative. Speed-up and memory reduction scale approximately as vS∈{+1,−1}Dv_S\in\{+1,-1\}^D5 (Alonso et al., 2020). In practical terms, for AskUbuntu with MLP (vS∈{+1,−1}Dv_S\in\{+1,-1\}^D6, vS∈{+1,−1}Dv_S\in\{+1,-1\}^D7), vS∈{+1,−1}Dv_S\in\{+1,-1\}^D8 vs. baseline vS∈{+1,−1}Dv_S\in\{+1,-1\}^D9, with ρj(vS)\rho^j(v_S)0 faster training, ρj(vS)\rho^j(v_S)1 faster testing, and ρj(vS)\rho^j(v_S)2 smaller memory footprint.

4. Integration in Deep Architectures: ZEN 2.0

ZEN 2.0 advances n-gram HD encoding by marrying external n-gram features with Transformer-based models. Extraction proceeds via PMI thresholds on large corpora (Chinese, Arabic), yielding lexicons (ρj(vS)\rho^j(v_S)3 = 261K for Chinese, 194K for Arabic). For each n-gram ρj(vS)\rho^j(v_S)4, a learnable embedding ρj(vS)\rho^j(v_S)5 is provided and contextualized via a 6-layer Transformer (no positional encoding). At every main encoder layer, token states ρj(vS)\rho^j(v_S)6 are fused by summation with weighted n-gram representations:

ρj(vS)\rho^j(v_S)7

where ρj(vS)\rho^j(v_S)8 is proportional to ρj(vS)\rho^j(v_S)9.

Pre-training objectives mirror BERT (MLM, NSP), with optional whole n-gram masking (WNM). No additional n-gram prediction head is required; the n-gram encoder is learned via the masked language modeling loss. Relative positional encoding and fusion architecture apply across domains and languages (Song et al., 2021).

5. Experimental Results and Empirical Performance

HyperEmbed was validated on three small intent-classification corpora (Chatbot, AskUbuntu, WebApplication) and a large-scale news corpus (20NewsGroups). HD embeddings (jj0–4, jj1–jj2) were supplied to multiple classifiers: Ridge, KNN, MLP, PA, RF, LSVC, SGD, NC, BNB. Benchmarked metrics include jj3, training/test time, and memory.

  • As jj4 increases, F1 approaches baseline attained with traditional n-gram features.
  • For 20NewsGroups (jj5, jj6–3): up to jj7 of baseline jj8 retained, with jj9–vw=⨀j=1nρj(vSj),w=S1S2⋯Snv_w = \bigodot_{j=1}^{n} \rho^j(v_{S_j}),\quad w=S_1S_2\cdots S_n0 speed-up and vw=⨀j=1nρj(vSj),w=S1S2⋯Snv_w = \bigodot_{j=1}^{n} \rho^j(v_{S_j}),\quad w=S_1S_2\cdots S_n1 memory reduction.
  • Linear classifiers (Ridge, MLP, PA, SGD, LSVC) show best trade-offs, while KNN and tree-based classifiers typically lose accuracy due to the distributed nature of the HD embeddings (Alonso et al., 2020).

ZEN 2.0 training on 8.4B-token Chinese and 7.3B-token Arabic datasets established new state-of-the-art performance over existing BERT and AraBERT baselines across 10 Chinese and multiple Arabic tasks. Task-specific gains in vw=⨀j=1nρj(vSj),w=S1S2⋯Snv_w = \bigodot_{j=1}^{n} \rho^j(v_{S_j}),\quad w=S_1S_2\cdots S_n2 and accuracy range from vw=⨀j=1nρj(vSj),w=S1S2⋯Snv_w = \bigodot_{j=1}^{n} \rho^j(v_{S_j}),\quad w=S_1S_2\cdots S_n3 to vw=⨀j=1nρj(vSj),w=S1S2⋯Snv_w = \bigodot_{j=1}^{n} \rho^j(v_{S_j}),\quad w=S_1S_2\cdots S_n4 points (Song et al., 2021).

Model Corpora Size n-gram Vocab F1/Acc Gain
HyperEmbed Small/Large – ca. vw=⨀j=1nρj(vSj),w=S1S2⋯Snv_w = \bigodot_{j=1}^{n} \rho^j(v_{S_j}),\quad w=S_1S_2\cdots S_n5 baseline, substantial efficiency (see above)
ZEN 2.0(L) 8.4B/7.3B 261K/194K vw=⨀j=1nρj(vSj),w=S1S2⋯Snv_w = \bigodot_{j=1}^{n} \rho^j(v_{S_j}),\quad w=S_1S_2\cdots S_n6–vw=⨀j=1nρj(vSj),w=S1S2⋯Snv_w = \bigodot_{j=1}^{n} \rho^j(v_{S_j}),\quad w=S_1S_2\cdots S_n7 over BERT/AraBERT

6. Practical Guidelines for Deployment and Tuning

Optimal configuration depends on corpus scale and resource constraints.

  • For small corpora, select vw=⨀j=1nρj(vSj),w=S1S2⋯Snv_w = \bigodot_{j=1}^{n} \rho^j(v_{S_j}),\quad w=S_1S_2\cdots S_n8–vw=⨀j=1nρj(vSj),w=S1S2⋯Snv_w = \bigodot_{j=1}^{n} \rho^j(v_{S_j}),\quad w=S_1S_2\cdots S_n9; for large, VV0–VV1.
  • Sweep VV2 from VV3 to VV4, choose the smallest VV5 achieving at least VV6–VV7 of baseline performance.
  • Prefer linear models and shallow MLP for HD representations.
  • Consider binarization of the HD vectors and classifier for maximal speed and memory efficiency in extreme settings.
  • Domain and language adaptation in ZEN 2.0 is architecture-neutral; PMI/frequency thresholds and n-gram lexicons are tuned per language, supporting broad coverage without rearchitecting (Alonso et al., 2020, Song et al., 2021).

7. Context, Significance, and Directions

N-gram HD encoders combine scalable representational capacity with memory and computational efficiency, enabling large-vocabulary or long-span n-gram features to be collapsed into manageable fixed-length vectors. Their use in modern NLP architectures, particularly in ZEN 2.0, demonstrates concrete accuracy gains and efficiency for diverse languages and domains.

This suggests the feasibility of high-performance, resource-conscious NLP pipeline designs, and opens potential for further research intersecting distributed representations, symbolic reasoning, and large-scale neural architectures. Continued investigation may address optimal permutation and binding schemes, fusion strategies, and generalization of HD encoding to non-linguistic sequence domains.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to HD Spam Filtering Accelerator.