Papers
Topics
Authors
Recent
Search
2000 character limit reached

N-gram HD Encoders & Transformer Fusion

Updated 30 December 2025
  • N-gram HD encoders are computational frameworks that represent n-gram statistics and contextual spans as fixed-length, high-dimensional vectors using hyperdimensional computing principles.
  • They leverage binding and bundling operations to encode n-gram sequences, enabling efficient integration into both classic classifiers and modern Transformer-based models.
  • Empirical results show these encoders offer significant trade-offs with marked improvements in speed and memory usage while maintaining competitive accuracy across diverse NLP tasks.

N-gram HD (High-Dimensional) Encoders are computational frameworks for representing n-gram statistics and contextual span information from text as fixed-length, high-dimensional vectors. These approaches synthesize ideas from hyperdimensional computing (HDC) and neural encoding architectures to produce distributed, resource-efficient representations suitable for both classic classifiers and modern Transformer-based models. Two principal lines—hyperdimensional binding/bundling schemes (Alonso et al., 2020) and n-gram Transformer fusion architectures (Song et al., 2021)—define current methodologies.

1. Hyperdimensional Computing Principles for N-gram Encoding

Hyperdimensional (HD) or Vector-Symbolic Architectures encode symbols, sequences, and sets by associating each with a high-dimensional vector, typically with dimension D∈[103,105]D \in [10^3, 10^5]. Each base vector vS∈{±1}Dv_S \in \{\pm1\}^D is drawn i.i.d. and nearly orthogonal. This property supports distributed operations:

  • Binding (element-wise multiplication, ⊙\odot): Used to encode sequence or order, crucial for n-grams.
  • Bundling (element-wise addition, ++): Aggregates multiple hypervectors (e.g., n-gram counts) in the same DD-dimensional space.

For character n-gram encoding: each character is mapped to a random bipolar vector, permuted by a fixed operator ρ\rho to encode positional information, and bound together to form n-gram representations. These are aggregated across a text to form a single DD-dimensional summary vector. The process is defined as:

vw=⨀j=1nρj(vSj),V=∑wc(w) vw,V^=V∥V∥2v_w = \bigodot_{j=1}^n \rho^j(v_{S_j}), \qquad V = \sum_{w} c(w)\,v_w, \qquad \widehat{V} = \frac{V}{\|V\|_2}

where c(w)c(w) is the n-gram count in a document, and ρ\rho denotes a fixed cyclic permutation of vector coordinates.

2. Algorithmic Workflow and Pseudocode

A standard workflow for N-gram HD encoding (Alonso et al., 2020) consists of the following steps:

  1. Initialization: Assign each symbol vS∈{±1}Dv_S \in \{\pm1\}^D0 in the alphabet vS∈{±1}Dv_S \in \{\pm1\}^D1 a random bipolar base vector vS∈{±1}Dv_S \in \{\pm1\}^D2.
  2. N-gram Extraction: Use a sliding window over the text vS∈{±1}Dv_S \in \{\pm1\}^D3 to enumerate all overlapping character n-grams.
  3. Binding: For each n-gram vS∈{±1}Dv_S \in \{\pm1\}^D4, bind the permuted symbol hypervectors as above.
  4. Bundling: Accumulate the bound n-gram hypervectors into the sum vS∈{±1}Dv_S \in \{\pm1\}^D5.
  5. Normalization: Obtain vS∈{±1}Dv_S \in \{\pm1\}^D6 by vS∈{±1}Dv_S \in \{\pm1\}^D7 normalization.
  6. Classifier Input: The normalized HD vector is input to standard classifiers.

Pseudocode for the encoder (as per (Alonso et al., 2020)):

vw=⨀j=1nρj(vSj),V=∑wc(w) vw,V^=V∥V∥2v_w = \bigodot_{j=1}^n \rho^j(v_{S_j}), \qquad V = \sum_{w} c(w)\,v_w, \qquad \widehat{V} = \frac{V}{\|V\|_2}6

The computational complexity is vS∈{±1}Dv_S \in \{\pm1\}^D8.

3. Trade-offs: Dimensionality, Context Size, and Resource Efficiency

A key advantage of N-gram HD encoders is the decoupling of output dimensionality vS∈{±1}Dv_S \in \{\pm1\}^D9 from the combinatorial explosion in ⊙\odot0 (traditional n-gram models require ⊙\odot1 counters; HD encoding always yields a ⊙\odot2-dimensional vector). Selection of ⊙\odot3 governs fidelity and resource consumption:

  • Larger ⊙\odot4: Higher fidelity to true n-gram histograms, increased classification accuracy, with increased memory and compute cost.
  • Smaller ⊙\odot5: Substantial memory and throughput savings, potential loss in accuracy if under-parameterized.

Empirical results indicate that F⊙\odot6 scores increase rapidly for small ⊙\odot7 (e.g., ⊙\odot8 offers marked improvement), saturating at dataset-dependent ⊙\odot9. For instance, ++0 for small corpora, ++1 for large (Alonso et al., 2020). Resource improvements scale proportionally: train/test speedups and memory reductions of ++2 to ++3 over dense n-gram representations are common at minimal accuracy loss.

4. Integration into Transformer Architectures: N-gram Fusion Techniques

In neural text encoders such as ZEN 2.0 (Song et al., 2021), n-gram high-dimensional embeddings are constructed via unsupervised statistical extraction and Transformer-based contextualization:

  • N-gram Extraction: Spans of length ++4–++5 are selected by pointwise mutual information (PMI) and frequency thresholds; for Chinese, PMI++6, ++7 yields ++8; for Arabic, PMI++9, DD0, DD1.
  • HD Embedding & Encoding: A learnable lookup table DD2 (with DD3 or DD4), followed by a six-layer Transformer encoder for contextualization.

Fusion with character/token-level Transformer states is executed layer-wise. Let DD5 denote the token state at layer DD6; overlapping n-grams DD7 contribute weighted contextual vectors. Fusion is done by: DD8 where DD9; no gating or concatenation is used. N-gram signals augment token-level representations directly. The n-gram encoder operates as a parallel Transformer.

5. Empirical Evaluation and Performance Metrics

HyperEmbed (Alonso et al., 2020) evaluated N-gram HD encoders on three small (Chatbot, AskUbuntu, WebApplication) and one large (20NewsGroups) corpus:

  • Baselines: Conventional character n-gram counts (ρ\rho0200,000 dimensions).
  • HD Embeddings: ρ\rho1–4, ρ\rho2–ρ\rho3.
  • Classifiers: Ridge, KNN, MLP, PA, RF, LSVC, SGD, NC, BNB.

Key results:

  • For AskUbuntu (MLP, ρ\rho4, ρ\rho5): Fρ\rho6 = 0.91 vs. baseline 0.92, ρ\rho7 faster training, ρ\rho8 faster test, ρ\rho9 memory reduction.
  • For 20NewsGroups (DD0, DD1–3): Most classifiers maintained DD290% of baseline FDD3 with DD4–DD5 speedup and DD6 memory reduction.

Linear classifiers and shallow MLPs generally outperformed local or tree-based models, which lost accuracy due to the distributed representation's smoothing effects.

ZEN 2.0 (Song et al., 2021) demonstrates consistent state-of-the-art improvements across a battery of Chinese and Arabic NLP tasks (e.g., MSR-CWS FDD7 = 98.66, CMRC2018 FDD8 = 89.92), outperforming prior benchmarks typically by 0.1–2.0 points in absolute metric terms.

6. Practical Guidelines and Adaptation Considerations

Recommendations for practitioners (Alonso et al., 2020, Song et al., 2021):

  • Select DD9–4 in small corpora; vw=⨀j=1nρj(vSj),V=∑wc(w) vw,V^=V∥V∥2v_w = \bigodot_{j=1}^n \rho^j(v_{S_j}), \qquad V = \sum_{w} c(w)\,v_w, \qquad \widehat{V} = \frac{V}{\|V\|_2}0–3 in large.
  • Sweep vw=⨀j=1nρj(vSj),V=∑wc(w) vw,V^=V∥V∥2v_w = \bigodot_{j=1}^n \rho^j(v_{S_j}), \qquad V = \sum_{w} c(w)\,v_w, \qquad \widehat{V} = \frac{V}{\|V\|_2}1 from vw=⨀j=1nρj(vSj),V=∑wc(w) vw,V^=V∥V∥2v_w = \bigodot_{j=1}^n \rho^j(v_{S_j}), \qquad V = \sum_{w} c(w)\,v_w, \qquad \widehat{V} = \frac{V}{\|V\|_2}2 to vw=⨀j=1nρj(vSj),V=∑wc(w) vw,V^=V∥V∥2v_w = \bigodot_{j=1}^n \rho^j(v_{S_j}), \qquad V = \sum_{w} c(w)\,v_w, \qquad \widehat{V} = \frac{V}{\|V\|_2}3; select minimal vw=⨀j=1nρj(vSj),V=∑wc(w) vw,V^=V∥V∥2v_w = \bigodot_{j=1}^n \rho^j(v_{S_j}), \qquad V = \sum_{w} c(w)\,v_w, \qquad \widehat{V} = \frac{V}{\|V\|_2}4 achieving vw=⨀j=1nρj(vSj),V=∑wc(w) vw,V^=V∥V∥2v_w = \bigodot_{j=1}^n \rho^j(v_{S_j}), \qquad V = \sum_{w} c(w)\,v_w, \qquad \widehat{V} = \frac{V}{\|V\|_2}5–98% of baseline accuracy.
  • Prefer linear and shallow neural classifiers for HD vectors.
  • For resource-constrained environments, binarize encodings and classifier weights.
  • ZEN 2.0 architecture adapts to multiple languages (Chinese, Arabic) and domains via threshold tuning and separate n-gram vocabularies, without structural changes.

7. Significance and Applications

N-gram HD encoders address scaling and efficiency challenges in embedding n-gram statistics for NLP tasks. By leveraging high-dimensionality and distributed encoding, they provide concise, accurate representations with dramatic resource savings. They integrate seamlessly into classic ML pipelines and Transformer-based neural architectures, enabling robust, end-to-end modeling of contextual spans. This framework offers practical trade-offs between memory, speed, and accuracy, and demonstrates superiority in multilingual, multi-domain applications, verified experimentally on several production-scale corpora (Alonso et al., 2020, Song et al., 2021).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to N-gram HD Encoders.