---
title: Attention-Free TreeFFN Architectures
url: https://www.emergentmind.com/topics/attention-free-treeffn-architectures
type: topic
---

# Attention-Free TreeFFN Architectures

Attention-free TreeFFN architectures are a class of neural network models that eliminate explicit attention mechanisms in favor of structured, feedforward propagation organized around tree-like or hierarchical topologies. These architectures span a range of formulations, including bidirectional neighbor-only propagation, hierarchical binary-tree reduction, and theoretically grounded feedforward decompositions based on the Kolmogorov–Arnold theorem. Across modalities such as structured reasoning, sequence modeling, and vision, attention-free TreeFFN designs aim to achieve linear or near-linear complexity, improved scalability, and competitive or superior task performance as compared to standard attention-based networks.

## 1. Structural Principles and Variants

Attention-free TreeFFN architectures diverge fundamentally from Transformer-style models by eliminating the requirement for pairwise interaction matrices and global token-token affinity calculations. The principal design variants include:

- **Bidirectional local TreeFFNs:** Exemplified by TreeGPT, these architectures employ sparse, neighbor-to-neighbor message passing in both left-to-right and right-to-left directions, with updates propagated via structured, parallel TreeFFN blocks. No global aggregation or attention map is used [2509.05550].
- **Hierarchical tree reduction:** The Wave-Attractor-Tree model utilizes a fixed, balanced binary-tree hierarchy. Successive GLU-based merge operations produce higher-level representations via strictly parallel bottom-up composition, with tree depth scaling as $\mathcal O(\log n)$ [2603.00812].
- **Analytic TreeFFNs via Kolmogorov–Arnold decompositions:** In Vision KAN, the architecture directly encodes a two-layer univariate function decomposition over each patch or token, parameterized by trainable radial basis function (RBF) expansions for maximum expressive power without explicit token mixing [2601.21541].

The defining commonality is the reliance on deterministic or learned tree topologies that induce sparse and interpretable computational graphs, replacing the dense, data-dependent connectivity of self-attention.

## 2. Mathematical Formulation and Update Mechanisms

### Bidirectional TreeFFN (TreeGPT)
Each TreeFFN block operates over local adjacency defined by ordered input sequences:
- For $H^{(t)}\in\mathbb{R}^{N\times d}$, the encoder and decoder updates are
  $$
  h_i^{(t+1)} = f(W_s h_i^{(t)} + W_n h_{j(i)}^{(t)} + b)
  $$
  where $j(i)=i+1$ (encoder) or $j(i)=i-1$ (decoder), and $f$ is a pointwise nonlinearity (e.g., GELU). Each pass updates all nodes in parallel along local edges, followed by residual summation [2509.05550].

### Hierarchical Binary Tree Reduction (Wave-Attractor-Tree)
The reduction proceeds through levels, merging pairs $(\ell, r)\in\mathbb{R}^d\times\mathbb{R}^d$ via:
- Stack inputs: $x = [\ell;\ r]\in\mathbb{R}^{2d}$
- Compute projections:
  $$
  \begin{aligned}
    v &= W_{\mathrm{val}} x\\
    g &= \sigma(W_{\mathrm{gate}} x)\\
    \rho &= \sigma(W_{\mathrm{res}} x)
  \end{aligned}
  $$
- Merge:
  $$
  m = \mathrm{RMSNorm}(v \odot g),\qquad \text{merged} = \rho \odot m\ +\ (1-\rho)\odot\tfrac{\ell+r}{2}
  $$
  With all merges at a level performed in parallel, tree depth is $\lceil\log_2 n\rceil$ and total merges $n-1$ [2603.00812].

### Analytic TreeFFN (Vision KAN)
Leveraging the Kolmogorov–Arnold theorem, a two-layer univariate decomposition is instantiated as:
- $f(x_1,\dots,x_n) = \sum_{q=0}^{2n} g_q\left(\sum_{p=1}^n h_{q,p}(x_p)\right)$
- Each $h_{q,p}$ and $g_q$ is parameterized by an RBF expansion,
  $$
  h_{q,p}(t) = \sum_{j=1}^M \alpha_{q,p,j} \exp\left(-\frac{(t-\mu_{q,p,j})^2}{2\sigma_{q,p,j}^2}\right)
  $$
  and similarly for $g_q$. In practice, Vision KAN merges this into a patch-wise nonlinear token mixing layer, plus additional axis-wise separable mixing and low-rank global projections [2601.21541].

## 3. Parallelism and Computational Complexity

TreeFFN architectures are designed for full data-parallel execution:
- **Bidirectional neighbor passes (TreeGPT):** All local updates performed in parallel via masked convolutions or banded matrix multiplications, scaling as $\mathcal{O}(N d^2)$ per block. No global softmax nor $N^2$ intermediate states are needed [2509.05550].
- **Hierarchical tree reduction (Wave-Attractor-Tree):** Each tree level requires a single parallel merge across all sibling pairs. Depth is $\mathcal O(\log n)$, total work is $\mathcal O(n d^2)$, and memory is $\mathcal O(n d)$. In contrast, Transformer attention is $\mathcal O(n^2 d)$ [2603.00812].
- **KAN-based token mixers (Vision KAN):** Each path (patch-wise, axis-wise, low-rank global) is parallelized. Overall complexity is linear in number of tokens: $\mathcal O(N C (MF + k + r))$, where $M$ is the number of RBF bases, $F$ the patch flattening dimension, $k$ convolution width, and $r$ global projection rank [2601.21541].

These properties yield significant improvements in efficiency and scalability relative to attention-based architectures, especially on long-sequence or high-resolution data.

## 4. Empirical Results and Benchmark Performance

Empirical evaluations demonstrate the efficacy of attention-free TreeFFNs in multiple domains:

- **Structured reasoning (TreeGPT on ARC Prize 2025):** TreeGPT achieves 99% validation accuracy and 100% token-level accuracy on selected samples, using only 3.16M parameters and converging within approximately 1500 training steps. This outperforms billion-parameter transformer baselines by over an order of magnitude in both accuracy (by 83–98 percentage points) and parameter efficiency [2509.05550].
- **Long-range structural dependencies (Wave-Attractor-Tree):** In language modeling (TinyShakespeare, sequence length 512), WAT achieves higher next-character accuracies (up to +11% over transformers) and orders-of-magnitude faster convergence. On bracket balance classification, WAT yields 75% accuracy versus 57% for transformers (18 percentage point improvement), with 10× less time per epoch [2603.00812].
- **Vision classification (Vision KAN on ImageNet-1K):** The ViK-Small model reaches 76.5% top-1 accuracy with 13.5M parameters and 1.6 GFLOPs, matching or exceeding architectures such as ResMLP-S12 and DeiT-Tiny. Ablations confirm the importance of all attention-free token-mixing paths (RBF bases, separable mixing, low-rank global mapping) [2601.21541].

These results support the assertion that attention-free TreeFFNs can provide state-of-the-art performance in both structured and perceptual tasks, at substantially reduced computational cost.

## 5. Practical Considerations, Inductive Biases, and Limitations

Attention-free TreeFFNs instantiate strong inductive biases aligned with specific problem domains:
- **Hierarchical compositionality:** Tree reduction and KAN-style decompositions encode hierarchical dependencies (e.g., syntax trees, nested structures) directly in the network structure, favoring domains with latent or explicit tree-like data.
- **Scalability:** Linear or near-linear scaling in token/patch count enables applicability to long sequences or high-dimensional data.
- **Parameter efficiency:** Parameter counts are typically independent of sequence length, because local weights or shared merge operators are reused across positions and levels.

Limitations are observed:
- **Generalization beyond trees:** Models tailored to tree or grid structures (e.g., TreeGPT) may not generalize immediately to unstructured natural language or free-form sequential data [2509.05550].
- **Dependency range:** Lack of global connections in neighborhood-only designs may hinder learning of long-range dependencies unless additional mechanisms (more layers, explicit global paths) are introduced.
- **Empirical breadth:** Thus far, systematic validation outside specialized benchmarks and visual classification tasks remains limited, and general-purpose replacement of attention mechanisms is not yet established [2509.05550][2601.21541].

## 6. Theoretical Foundations and Future Directions

The theoretical basis for TreeFFN architectures encompasses:
- **Kolmogorov–Arnold representation:** Vision KAN directly implements a constructive version of the theorem, with RBF parameterizations delivering universal function approximation in feedforward form [2601.21541].
- **Learned and fixed tree topologies:** Both fixed (Wave-Attractor-Tree’s binary hierarchy [2603.00812]) and data-driven (TreeGPT’s neighbor graphs [2509.05550]) can be used, with potential to learn richer adjacency relations.
- **Hybridization opportunities:** Open questions include integrating lightweight global tokens, learning adjacency for arbitrarily structured data, and extending TreeFFN principles to multi-modal or graph-structured domains [2509.05550].

A plausible implication is that as scalability demands grow and architectural interpretability becomes critical, attention-free TreeFFNs based on function-theoretic and compositional principles will form an important class of highly efficient neural models.

## 7. Comparison Table of Core Attention-Free TreeFFN Architectures

| Architecture              | Domain                | Core Mechanism                          | Reference   |
|---------------------------|-----------------------|-----------------------------------------|-------------|
| TreeGPT                   | Structured reasoning  | Bidirectional neighbor TreeFFN          | [2509.05550]|
| Wave-Attractor-Tree       | Sequence modeling     | Hierarchical binary GLU composition     | [2603.00812]|
| Vision KAN (ViK)          | Vision (ImageNet-1K)  | Patch-wise RBFKAN (K–A decomposition)   | [2601.21541]|

This summary highlights the diversity of mechanisms applicable across different data modalities and how TreeFFN paradigms offer flexible and efficient alternatives to attention-centric architectures.

Source: https://www.emergentmind.com/topics/attention-free-treeffn-architectures