Papers
Topics
Authors
Recent
Search
2000 character limit reached

Symbol-Equivariant Recurrent Reasoning Models

Updated 4 March 2026
  • SE-RRMs are neural architectures that explicitly enforce symbol permutation equivariance to improve robust reasoning performance.
  • They reduce computational complexity by eliminating extensive symbol-permutation data augmentation, enhancing scalability and generalization.
  • Empirical evaluations on Sudoku, ARC-AGI, and Maze tasks demonstrate SE-RRM's superior efficiency and accuracy over traditional recurrent reasoning models.

Symbol-Equivariant Recurrent Reasoning Models (SE-RRMs) are neural architectures specifically designed for structured reasoning tasks, such as Sudoku and ARC-AGI, where symmetries over symbols (digits, colors, etc.) provide crucial inductive bias. Unlike earlier Recurrent Reasoning Models (RRMs)—notably Hierarchical Reasoning Model (HRM) and Tiny Recursive Model (TRM)—which enforce permutation symmetry only implicitly via extensive data augmentation, SE-RRMs achieve permutation equivariance at the architectural level. This is realized through symbol-equivariant layers, ensuring the model's outputs are invariant under any relabeling of input symbols. As a result, SE-RRMs produce identical solutions for all permutations of the symbol set and exhibit enhanced robustness, data-efficiency, and generalization across task scales and symbol sets (Freinschlag et al., 2 Mar 2026).

1. Permutation Equivariance: Formalism and Implementation

Permutation equivariance is defined as the commutativity of a function f:X→Yf : X \rightarrow Y with the action of a permutation group GG on XX and YY, i.e., f(g⋅X)=g⋅f(X)f(g \cdot X) = g \cdot f(X) for every g∈Gg \in G. In the context of SE-RRMs, two distinct symmetry groups are considered: SIS_I (permutations over positions/cells) and SKS_K (permutations over KK symbols or colors). Input data are encoded as three-way tensors X∈RD×I×KX \in \mathbb{R}^{D \times I \times K}, with symbol permutations GG0 acting as GG1 and position permutations GG2 as GG3.

SE-RRMs guarantee symbol-equivariance by designing every model layer—attention, MLP, normalization, residual connections—to commute with the action of GG4. Specifically, for any SE-RRM block mapping GG5 and for any GG6, the relation GG7 holds. The output is an GG8 tensor, where permuting the symbol dimension corresponds exactly to relabeling.

Architecturally, equivariant linear maps GG9 satisfy XX0. These are constructed in practice via weight-sharing and explicit attention operations over the symbol axis.

2. Model Architecture and Computational Details

Let XX1 denote the number of positions (cells), XX2 the number of symbols, and XX3 the feature dimension. Inputs XX4 (with XX5) are embedded into XX6. The recurrent hidden state at time XX7 is XX8.

The model operates as a fixed-point iteration for XX9, with each block YY0 composed of YY1 layers. The layer operations per block are:

  • Positional self-attention YY2 along positions, treating each (feature, symbol) slice as a sequence.
  • Symbol self-attention YY3 along symbols, shared across positions.
  • A pointwise MLP YY4 (SwiGLU) per (i, c).
  • RMS normalization applied over the feature dimension.

Operations per layer YY5 are: YY6 with the block output YY7. The output projection is a linear map YY8, shared across all YY9, yielding f(g⋅X)=g⋅f(X)f(g \cdot X) = g \cdot f(X)0 logits, followed by a row-wise softmax for class probabilities.

The architectural design, where all layers (attention, MLP, normalization) commute with permutations in symbol axis, is central to enforcing exact equivariance.

3. Training Regime and Objective Function

Deep supervision is applied at each of f(g⋅X)=g⋅f(X)f(g \cdot X) = g \cdot f(X)1 unrolled steps; for each iteration f(g⋅X)=g⋅f(X)f(g \cdot X) = g \cdot f(X)2, a cross-entropy loss

f(g⋅X)=g⋅f(X)f(g \cdot X) = g \cdot f(X)3

is computed with f(g⋅X)=g⋅f(X)f(g \cdot X) = g \cdot f(X)4 as the target symbol at position f(g⋅X)=g⋅f(X)f(g \cdot X) = g \cdot f(X)5. At each step, gradients are backpropagated through the current f(g⋅X)=g⋅f(X)f(g \cdot X) = g \cdot f(X)6 only, detaching f(g⋅X)=g⋅f(X)f(g \cdot X) = g \cdot f(X)7 to stabilize training. A random halting scheme, with halt probability f(g⋅X)=g⋅f(X)f(g \cdot X) = g \cdot f(X)8 at each step except the last, serves to replace a Q-learning halting policy and reduces compute.

Optimization uses AdamW with weight decay and a warmup plus constant or cosine learning rate scheduling. In typical Sudoku experiments, hyperparameters include learning rate f(g⋅X)=g⋅f(X)f(g \cdot X) = g \cdot f(X)9, weight decay 1, batch size g∈Gg \in G0, g∈Gg \in G1 deep supervision steps, feature dimension g∈Gg \in G2, and a 2 million parameter model.

Crucially, SE-RRM’s S_K-equivalence eliminates the need for symbol-permutation data augmentation: only spatial augmentations are required (e.g., dihedral symmetries in ARC-AGI), reducing augmentation needs by two orders of magnitude relative to HRM/TRM.

4. Empirical Evaluation and Benchmark Results

SE-RRM performance is evaluated across structured reasoning tasks, with primary comparisons against HRM and TRM. All results reported are as in (Freinschlag et al., 2 Mar 2026).

A. Sudoku

  • Training: 1,000 base 9×9 puzzles × 1,000 symbol-permutation augmentations.
  • Test: 422,786 9×9 puzzles; zero-shot generalization on 4×4, 16×16, and 25×25 puzzles.
Model 4×4 FSR (GPA) 9×9 FSR (GPA) 16×16 GPA 25×25 GPA
HRM 0% (29%) 63.5% (86.1%) -- --
TRM 0% (46%) 71.9% (89.8%) -- --
SE-RRM 95.5% (99.2%) 93.7% (97.6%) 51.9% 31.5%

SE-RRM dramatically outperforms prior RRMs, including near-perfect generalization to 4×4 and >50% accuracy on larger unseen 16×16 and 25×25 grids, despite training solely on 9×9.

Test-time scaling with increased steps g∈Gg \in G3 demonstrates improved solution rates, e.g., 93.7% FSR at g∈Gg \in G4, rising to 98.8% at g∈Gg \in G5.

B. ARC-AGI

Model ARC-AGI-1 pass@2 ARC-AGI-2 pass@2
HRM 40.3% 5.0%
TRM 44.6% 7.8%
SE-RRM 45.3% 7.1%

SE-RRM matches or slightly surpasses prior results with only 8 dihedral augmentations per puzzle, compared to ~1,000 symbol-permutation augmentations required for HRM/TRM.

C. Maze

  • Dataset: 1,000 train/test 30×30 mazes (path length ≥110); four distinct symbols (not treated equivariantly).
  • Metric: Fully solved rate (FSR).
Model Maze FSR
HRM 74.5%
TRM 85.3%
SE-RRM 88.8%

SE-RRM achieves the highest FSR, even on tasks where symbol-equivariance is explicitly broken by distinct embeddings.

5. Training Workflow and Pseudocode

A typical SE-RRM training step proceeds as follows:

g∈Gg \in G7

This approach, particularly the detachment at each block, ensures stable and efficient training.

6. Analysis and Broader Implications

SE-RRM’s explicit architectural enforcement of symbol-permutation equivariance yields significant benefits:

  • Elimination of the need for g∈Gg \in G6 symbol-permutation augmentations, greatly reducing training sample requirements.
  • Robust out-of-distribution generalization, with models trained on 9×9 Sudoku achieving competitive or superior results on 4×4, 16×16, and 25×25 puzzles.
  • Improved scalability, with competitive or improved performance over HRM and TRM, often with fewer parameters (2 million) and far fewer augmentations.
  • Sample efficiency, with only spatial augmentations required for most tasks.
  • At inference, solution accuracy improves monotonically with the number of recurrent steps, enabling a test-time computation-accuracy tradeoff.

A plausible implication is that permutation-equivariant architectures, such as SE-RRM, are crucial for robust symbolic reasoning in domains with inherent symmetries, and their explicit symmetry handling addresses core generalization limitations of previous methods (Freinschlag et al., 2 Mar 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Symbol-Equivariant Recurrent Reasoning Models (SE-RRMs).