Papers
Topics
Authors
Recent
Search
2000 character limit reached

Attention-Free TreeFFN Architectures

Updated 13 May 2026
  • The paper presents TreeFFN architectures that replace dense attention with structured, tree-like feedforward propagation, achieving linear or near-linear complexity.
  • Empirical evaluations reveal models like TreeGPT, Wave-Attractor-Tree, and Vision KAN outperform traditional attention networks in accuracy and parameter efficiency.
  • These architectures leverage deterministic tree topologies and RBF-based decompositions to offer scalable, parallelizable solutions for structured reasoning, sequence modeling, and vision.

Attention-free TreeFFN architectures are a class of neural network models that eliminate explicit attention mechanisms in favor of structured, feedforward propagation organized around tree-like or hierarchical topologies. These architectures span a range of formulations, including bidirectional neighbor-only propagation, hierarchical binary-tree reduction, and theoretically grounded feedforward decompositions based on the Kolmogorov–Arnold theorem. Across modalities such as structured reasoning, sequence modeling, and vision, attention-free TreeFFN designs aim to achieve linear or near-linear complexity, improved scalability, and competitive or superior task performance as compared to standard attention-based networks.

1. Structural Principles and Variants

Attention-free TreeFFN architectures diverge fundamentally from Transformer-style models by eliminating the requirement for pairwise interaction matrices and global token-token affinity calculations. The principal design variants include:

  • Bidirectional local TreeFFNs: Exemplified by TreeGPT, these architectures employ sparse, neighbor-to-neighbor message passing in both left-to-right and right-to-left directions, with updates propagated via structured, parallel TreeFFN blocks. No global aggregation or attention map is used (Li, 6 Sep 2025).
  • Hierarchical tree reduction: The Wave-Attractor-Tree model utilizes a fixed, balanced binary-tree hierarchy. Successive GLU-based merge operations produce higher-level representations via strictly parallel bottom-up composition, with tree depth scaling as O(log⁡n)\mathcal O(\log n) (Berezkin, 28 Feb 2026).
  • Analytic TreeFFNs via Kolmogorov–Arnold decompositions: In Vision KAN, the architecture directly encodes a two-layer univariate function decomposition over each patch or token, parameterized by trainable radial basis function (RBF) expansions for maximum expressive power without explicit token mixing (Yang et al., 29 Jan 2026).

The defining commonality is the reliance on deterministic or learned tree topologies that induce sparse and interpretable computational graphs, replacing the dense, data-dependent connectivity of self-attention.

2. Mathematical Formulation and Update Mechanisms

Bidirectional TreeFFN (TreeGPT)

Each TreeFFN block operates over local adjacency defined by ordered input sequences:

  • For H(t)∈RN×dH^{(t)}\in\mathbb{R}^{N\times d}, the encoder and decoder updates are

hi(t+1)=f(Wshi(t)+Wnhj(i)(t)+b)h_i^{(t+1)} = f(W_s h_i^{(t)} + W_n h_{j(i)}^{(t)} + b)

where j(i)=i+1j(i)=i+1 (encoder) or j(i)=i−1j(i)=i-1 (decoder), and ff is a pointwise nonlinearity (e.g., GELU). Each pass updates all nodes in parallel along local edges, followed by residual summation (Li, 6 Sep 2025).

Hierarchical Binary Tree Reduction (Wave-Attractor-Tree)

The reduction proceeds through levels, merging pairs (ℓ,r)∈Rd×Rd(\ell, r)\in\mathbb{R}^d\times\mathbb{R}^d via:

  • Stack inputs: x=[ℓ; r]∈R2dx = [\ell;\ r]\in\mathbb{R}^{2d}
  • Compute projections:

v=Wvalx g=σ(Wgatex) ρ=σ(Wresx)\begin{aligned} v &= W_{\mathrm{val}} x\ g &= \sigma(W_{\mathrm{gate}} x)\ \rho &= \sigma(W_{\mathrm{res}} x) \end{aligned}

  • Merge:

m=RMSNorm(v⊙g),merged=ρ⊙m + (1−ρ)⊙ℓ+r2m = \mathrm{RMSNorm}(v \odot g),\qquad \text{merged} = \rho \odot m\ +\ (1-\rho)\odot\tfrac{\ell+r}{2}

With all merges at a level performed in parallel, tree depth is H(t)∈RN×dH^{(t)}\in\mathbb{R}^{N\times d}0 and total merges H(t)∈RN×dH^{(t)}\in\mathbb{R}^{N\times d}1 (Berezkin, 28 Feb 2026).

Analytic TreeFFN (Vision KAN)

Leveraging the Kolmogorov–Arnold theorem, a two-layer univariate decomposition is instantiated as:

  • H(t)∈RN×dH^{(t)}\in\mathbb{R}^{N\times d}2
  • Each H(t)∈RN×dH^{(t)}\in\mathbb{R}^{N\times d}3 and H(t)∈RN×dH^{(t)}\in\mathbb{R}^{N\times d}4 is parameterized by an RBF expansion,

H(t)∈RN×dH^{(t)}\in\mathbb{R}^{N\times d}5

and similarly for H(t)∈RN×dH^{(t)}\in\mathbb{R}^{N\times d}6. In practice, Vision KAN merges this into a patch-wise nonlinear token mixing layer, plus additional axis-wise separable mixing and low-rank global projections (Yang et al., 29 Jan 2026).

3. Parallelism and Computational Complexity

TreeFFN architectures are designed for full data-parallel execution:

  • Bidirectional neighbor passes (TreeGPT): All local updates performed in parallel via masked convolutions or banded matrix multiplications, scaling as H(t)∈RN×dH^{(t)}\in\mathbb{R}^{N\times d}7 per block. No global softmax nor H(t)∈RN×dH^{(t)}\in\mathbb{R}^{N\times d}8 intermediate states are needed (Li, 6 Sep 2025).
  • Hierarchical tree reduction (Wave-Attractor-Tree): Each tree level requires a single parallel merge across all sibling pairs. Depth is H(t)∈RN×dH^{(t)}\in\mathbb{R}^{N\times d}9, total work is hi(t+1)=f(Wshi(t)+Wnhj(i)(t)+b)h_i^{(t+1)} = f(W_s h_i^{(t)} + W_n h_{j(i)}^{(t)} + b)0, and memory is hi(t+1)=f(Wshi(t)+Wnhj(i)(t)+b)h_i^{(t+1)} = f(W_s h_i^{(t)} + W_n h_{j(i)}^{(t)} + b)1. In contrast, Transformer attention is hi(t+1)=f(Wshi(t)+Wnhj(i)(t)+b)h_i^{(t+1)} = f(W_s h_i^{(t)} + W_n h_{j(i)}^{(t)} + b)2 (Berezkin, 28 Feb 2026).
  • KAN-based token mixers (Vision KAN): Each path (patch-wise, axis-wise, low-rank global) is parallelized. Overall complexity is linear in number of tokens: hi(t+1)=f(Wshi(t)+Wnhj(i)(t)+b)h_i^{(t+1)} = f(W_s h_i^{(t)} + W_n h_{j(i)}^{(t)} + b)3, where hi(t+1)=f(Wshi(t)+Wnhj(i)(t)+b)h_i^{(t+1)} = f(W_s h_i^{(t)} + W_n h_{j(i)}^{(t)} + b)4 is the number of RBF bases, hi(t+1)=f(Wshi(t)+Wnhj(i)(t)+b)h_i^{(t+1)} = f(W_s h_i^{(t)} + W_n h_{j(i)}^{(t)} + b)5 the patch flattening dimension, hi(t+1)=f(Wshi(t)+Wnhj(i)(t)+b)h_i^{(t+1)} = f(W_s h_i^{(t)} + W_n h_{j(i)}^{(t)} + b)6 convolution width, and hi(t+1)=f(Wshi(t)+Wnhj(i)(t)+b)h_i^{(t+1)} = f(W_s h_i^{(t)} + W_n h_{j(i)}^{(t)} + b)7 global projection rank (Yang et al., 29 Jan 2026).

These properties yield significant improvements in efficiency and scalability relative to attention-based architectures, especially on long-sequence or high-resolution data.

4. Empirical Results and Benchmark Performance

Empirical evaluations demonstrate the efficacy of attention-free TreeFFNs in multiple domains:

  • Structured reasoning (TreeGPT on ARC Prize 2025): TreeGPT achieves 99% validation accuracy and 100% token-level accuracy on selected samples, using only 3.16M parameters and converging within approximately 1500 training steps. This outperforms billion-parameter transformer baselines by over an order of magnitude in both accuracy (by 83–98 percentage points) and parameter efficiency (Li, 6 Sep 2025).
  • Long-range structural dependencies (Wave-Attractor-Tree): In language modeling (TinyShakespeare, sequence length 512), WAT achieves higher next-character accuracies (up to +11% over transformers) and orders-of-magnitude faster convergence. On bracket balance classification, WAT yields 75% accuracy versus 57% for transformers (18 percentage point improvement), with 10× less time per epoch (Berezkin, 28 Feb 2026).
  • Vision classification (Vision KAN on ImageNet-1K): The ViK-Small model reaches 76.5% top-1 accuracy with 13.5M parameters and 1.6 GFLOPs, matching or exceeding architectures such as ResMLP-S12 and DeiT-Tiny. Ablations confirm the importance of all attention-free token-mixing paths (RBF bases, separable mixing, low-rank global mapping) (Yang et al., 29 Jan 2026).

These results support the assertion that attention-free TreeFFNs can provide state-of-the-art performance in both structured and perceptual tasks, at substantially reduced computational cost.

5. Practical Considerations, Inductive Biases, and Limitations

Attention-free TreeFFNs instantiate strong inductive biases aligned with specific problem domains:

  • Hierarchical compositionality: Tree reduction and KAN-style decompositions encode hierarchical dependencies (e.g., syntax trees, nested structures) directly in the network structure, favoring domains with latent or explicit tree-like data.
  • Scalability: Linear or near-linear scaling in token/patch count enables applicability to long sequences or high-dimensional data.
  • Parameter efficiency: Parameter counts are typically independent of sequence length, because local weights or shared merge operators are reused across positions and levels.

Limitations are observed:

  • Generalization beyond trees: Models tailored to tree or grid structures (e.g., TreeGPT) may not generalize immediately to unstructured natural language or free-form sequential data (Li, 6 Sep 2025).
  • Dependency range: Lack of global connections in neighborhood-only designs may hinder learning of long-range dependencies unless additional mechanisms (more layers, explicit global paths) are introduced.
  • Empirical breadth: Thus far, systematic validation outside specialized benchmarks and visual classification tasks remains limited, and general-purpose replacement of attention mechanisms is not yet established (Li, 6 Sep 2025, Yang et al., 29 Jan 2026).

6. Theoretical Foundations and Future Directions

The theoretical basis for TreeFFN architectures encompasses:

  • Kolmogorov–Arnold representation: Vision KAN directly implements a constructive version of the theorem, with RBF parameterizations delivering universal function approximation in feedforward form (Yang et al., 29 Jan 2026).
  • Learned and fixed tree topologies: Both fixed (Wave-Attractor-Tree’s binary hierarchy (Berezkin, 28 Feb 2026)) and data-driven (TreeGPT’s neighbor graphs (Li, 6 Sep 2025)) can be used, with potential to learn richer adjacency relations.
  • Hybridization opportunities: Open questions include integrating lightweight global tokens, learning adjacency for arbitrarily structured data, and extending TreeFFN principles to multi-modal or graph-structured domains (Li, 6 Sep 2025).

A plausible implication is that as scalability demands grow and architectural interpretability becomes critical, attention-free TreeFFNs based on function-theoretic and compositional principles will form an important class of highly efficient neural models.

7. Comparison Table of Core Attention-Free TreeFFN Architectures

Architecture Domain Core Mechanism Reference
TreeGPT Structured reasoning Bidirectional neighbor TreeFFN (Li, 6 Sep 2025)
Wave-Attractor-Tree Sequence modeling Hierarchical binary GLU composition (Berezkin, 28 Feb 2026)
Vision KAN (ViK) Vision (ImageNet-1K) Patch-wise RBFKAN (K–A decomposition) (Yang et al., 29 Jan 2026)

This summary highlights the diversity of mechanisms applicable across different data modalities and how TreeFFN paradigms offer flexible and efficient alternatives to attention-centric architectures.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Attention-Free TreeFFN Architectures.