---
title: TreeFFN Encoder-Decoder Architecture
url: https://www.emergentmind.com/topics/treeffn-encoder-decoder
type: topic
---

# TreeFFN Encoder-Decoder Architecture

TreeFFN Encoder-Decoder denotes a class of tree-structured or tree-inspired encoder-decoder systems in which encoding and decoding are carried out without standard transformer self-attention, but the term is used most explicitly in TreeGPT to describe a pure TreeFFN encoder-decoder for structured reasoning tasks [2509.05550]. In that formulation, an Encoder TreeFFN processes left-to-right dependencies and a Decoder TreeFFN processes right-to-left patterns through adjacent connections executed in parallel, with the stated aim of eliminating attention computation while maintaining sequence modeling capabilities [2509.05550]. Related antecedents in the literature include soft decision-tree autoencoders, in which both encoder and decoder are differentiable trees trained by stochastic gradient descent, and forest-based auto-encoders, in which encoding is a vector of leaf indices and decoding reconstructs inputs from the intersection of tree path constraints [1409.7461] [1709.09018].

## 1. Definition and directional structure

In TreeGPT, the TreeFFN Encoder-Decoder is defined as an attention-free neural architecture based on two directional components: an Encoder TreeFFN with only left-to-right adjacent edges and a Decoder TreeFFN with only right-to-left adjacent edges [2509.05550]. The encoder edge set is

$$
E_{\mathrm{enc}} = \{(i, i+1) : i = 0 \ldots N-2\},
$$

and the decoder edge set is

$$
E_{\mathrm{dec}} = \{(i, i-1) : i = N-1 \ldots 1\}.
$$

The design principle is explicitly local. The encoder propagates information from position $i$ to $i+1$ in parallel across the full sequence, the decoder propagates information from position $i$ to $i-1$ in parallel, and both layers run simultaneously with no sequential “for $i$” loops and exchange only neighbor messages [2509.05550]. The architecture is therefore not a tree in the sense of a classical decision tree; it is a graph-style message-passing construction simplified to 1-hop neighbor communication, with no global attention [2509.05550].

A plausible implication is that the term “TreeFFN” in this setting names the structure of localized feed-forward propagation rather than a branching decision-tree topology. That distinction is important when comparing TreeGPT with earlier tree-based encoder-decoder models.

## 2. Message computation and bidirectional update equations

The TreeGPT formulation specifies message computation per edge through a projected edge representation and an MLP:

$$
e_{ij}^{\mathrm{proj}} = \mathrm{Linear}(e_{ij}),
$$

$$
m_{ij} = \mathrm{MLP}([h_i; h_j; e_{ij}^{\mathrm{proj}}]).
$$

An optional gating mechanism is also defined:

$$
g_{ij} = \sigma(\mathrm{Gate}(h_i, h_j)),
\qquad
\mathrm{agg}_i = \sum_{j \in N(i)} g_{ij}\cdot m_{ij}.
$$

Let $H^{(t)} \in \mathbb{R}^{N \times d}$ denote node features at iteration $t$. The encoder update is

$$
h_{\mathrm{enc}}^{(t+1)} = \mathrm{TreeFFN}_{L \rightarrow R}(H^{(t)}, E_{\mathrm{enc}}),
$$

the decoder update is

$$
h_{\mathrm{dec}}^{(t+1)} =
\mathrm{TreeFFN}_{R \leftarrow L}(H^{(t)} + h_{\mathrm{enc}}^{(t+1)}, E_{\mathrm{dec}}),
$$

and the residual combination is

$$
H^{(t+1)} = H^{(t)} + h_{\mathrm{dec}}^{(t+1)}.
$$

The paper also gives a single-step bidirectional view:

$$
H_{\mathrm{final}} =
H_{\mathrm{input}}
+ \mathrm{TreeFFN}_{L \rightarrow R}(H_{\mathrm{input}}, E_{\mathrm{enc}})
+ \mathrm{TreeFFN}_{R \leftarrow L}(H_{\mathrm{input}}, E_{\mathrm{dec}}).
$$

This formulation makes the encoder-decoder coupling explicit: the decoder is applied to $H^{(t)} + h_{\mathrm{enc}}^{(t+1)}$, so the right-to-left pass is conditioned on the output of the left-to-right pass [2509.05550]. This suggests that bidirectionality is achieved not through attention heads but through residual composition of two directional TreeFFN operators.

## 3. Architectural specification and computational profile

The reported TreeGPT instantiation has approximately 3.16 million parameters, hidden dimension $d=256$, and 2 stacked “bidirectional TreeFFN” layers [2509.05550]. Each stacked layer contains one $L \rightarrow R$ TreeFFN and one $R \leftarrow L$ TreeFFN, and within each component 2 internal TreeFFN iterations $(T=2)$ are applied [2509.05550]. Residual connections wrap each TreeFFN component to preserve gradient flow.

The architecture is described as computationally efficient because neighbor-to-neighbor TreeFFN has complexity $O(N \cdot d^2)$ per layer, in contrast to $O(N^2 \cdot d)$ for self-attention, and it eliminates all quadratic cost of attention matrices [2509.05550]. The empirical ablation timings reported for convergence are 563 s for edge-projection only, 579 s for edge-proj + gating, and 894 s for the full TreeFFN baseline with attention off; under the same setup, small transformer baselines typically exceed 1,200 s [2509.05550].

These complexity claims are specific to the reported architecture and setup. They characterize a design space in which performance is sought by constraining communication to adjacent edges rather than by constructing dense pairwise attention maps.

## 4. Training regime and reported ARC-AGI-2 results

TreeGPT is evaluated on the ARC Prize 2025 dataset, described as grid-based visual puzzles requiring abstract rule induction [2509.05550]. Optimization uses AdamW with a cosine learning rate schedule, and the reported convergence behavior is that validation accuracy reached 99% within 1,500 gradient steps while training loss decays smoothly under the cosine schedule [2509.05550].

The final performance figures given are 99% full-task validation accuracy and 100% token-level accuracy on selected held-out examples [2509.05550]. The abstract states that the model achieves 99% validation accuracy using 3.16M parameters, converges within 1500 training steps, and demonstrates 100% token-level accuracy on selected evaluation samples [2509.05550]. The implementation is described as trainable on mid-range hardware, and complete reproducibility scripts are said to be provided [2509.05550].

The summary also reports a parameter-efficiency comparison of 3.16 M versus 1.5 B in small CoT models, alongside 99% versus approximately 1–2% accuracy on ARC-AGI-2 [2509.05550]. Within the scope of the reported experiments, this positions TreeFFN Encoder-Decoder as a specialized structured-reasoning architecture rather than a generic language-model substitute. The paper itself frames the findings as preliminary and states that further investigation across diverse tasks and datasets would be valuable for establishing broader applicability [2509.05550].

## 5. Relation to earlier tree-based encoder-decoder paradigms

Before TreeGPT, encoder-decoder systems based on trees were already present in substantially different forms. In "Autoencoder Trees" [1409.7461], the encoder and decoder are soft binary decision trees. The encoder tree $t_e$ maps $x \in \mathbb{R}^d$ to a lower-dimensional code $z \in \mathbb{R}^k$, and the decoder tree $t_d$ maps $z$ back to a reconstruction $\hat{x} \in \mathbb{R}^d$ [1409.7461]. At each internal node $m$, routing is defined by a soft multivariate sigmoid split,

$$
g_m(x) = \frac{1}{1+\exp(-w_m^T x)},
$$

and the response at an internal node is a convex combination of its children:

$$
y_m(x) = g_m(x)\, y_{\ell}(x) + (1-g_m(x))\, y_r(x).
$$

Training minimizes mean squared reconstruction error and uses online SGD with diagonal AdaGrad; trees begin at depth 2 and are grown incrementally every 40 epochs up to final depth 5 or 6, with $L_2$ regularization on split weights and leaf vectors [1409.7461]. Empirically, the model is reported to outperform perceptron autoencoders on MNIST at $k=2$, to be competitive at $k=10$, and to consistently beat AE perceptrons on 20 NewsGroups; increasing tree depth from 5 to 6 yields larger gains than increasing $k$ [1409.7461]. In this lineage, the encoder-decoder interpretation is literal: both maps are tree-valued functions.

A second antecedent is "AutoEncoder by Forest" [1709.09018], which proposes EncoderForest, or eForest, as a tree ensemble based auto-encoder. Here the encoder $f:\mathbb{R}^d \to \mathbb{Z}^T$ maps an input to a vector of leaf indices, one per tree, and the decoder reconstructs by intersecting the path constraints implied by those leaf indices [1709.09018]. The resulting equivalence class is called the Maximal-Compatible-Rule (MCR), and reconstruction selects a representative point from that intersection, such as interval midpoints for continuous features [1709.09018]. Unlike Autoencoder Trees, eForest does not backpropagate through encoder and decoder; trees are grown either by completely random splits in the unsupervised variant or by Random Forest procedures in the supervised variant, and reconstruction quality is then measured by mean squared error [1709.09018]. Training complexity is reported as $O(T \cdot n \cdot \log n \cdot d)$ for balanced trees, and the implementation on a 68-core CPU is said to be 100× faster than a CNN-AE on GPU, although decoding is empirically slower than DNN autoencoders [1709.09018].

Taken together, these earlier systems show that “tree encoder-decoder” has referred to at least three distinct mechanisms: differentiable soft decision trees, forest path-code reconstruction, and TreeGPT’s bidirectional neighbor-to-neighbor TreeFFN propagation. A common misconception is to treat these as interchangeable. They share an encoder-decoder template, but their computational primitives, training procedures, and representational assumptions are materially different.

## 6. Scope, limitations, and prospective extensions

The reported TreeGPT formulation has explicit scope conditions. Its current limitations are that it assumes explicit tree or AST structures, does not generalize directly to unstructured text or very large trees, may incur growing computational cost for extremely deep or wide trees despite being linear in sequence length, and has been evaluated only on ARC-AGI-2, leaving multi-task generality untested [2509.05550].

The suggested future directions are automatic tree inference from raw sequences when no AST is available, extension to multi-modal reasoning with code, text, and images, adaptation of TreeFFN blocks to graph-structured data beyond chains such as general DAGs and knowledge graphs, and advanced optimizations including mixed-precision, gradient checkpointing, structured sparsity, and model compression [2509.05550]. Additional application domains proposed in the summary are natural language parsing, molecular graphs, and knowledge-base reasoning [2509.05550].

These qualifications delimit the meaning of the reported results. The ARC-AGI-2 performance demonstrates that an attention-free TreeFFN Encoder-Decoder can be highly effective on a structured reasoning benchmark under the stated setup, but the same source explicitly stops short of claiming general replacement of attention-based architectures [2509.05550]. A plausible implication is that the main significance of TreeFFN Encoder-Decoder lies in identifying a regime where extremely local neighbor passages and residual bidirectional composition are sufficient for high-performance reasoning, rather than in establishing a universal architecture for all sequence or multimodal tasks.

Source: https://www.emergentmind.com/topics/treeffn-encoder-decoder