Papers
Topics
Authors
Recent
Search
2000 character limit reached

When to Re-Plan: Subgoal Persistence in Hierarchical Latent Reasoning

Published 2 Jun 2026 in cs.AI | (2606.03741v1)

Abstract: Long-horizon reasoning requires a system to commit to medium-horizon intent without becoming rigid: re-plan too often and computation never coheres into multi-step structure; commit too long and the plan goes stale. We study this stability-adaptivity tradeoff in the latent reasoning setting, where multi-step computation occurs inside hidden state rather than externalized token traces. We extend the Hierarchical Reasoning Model (HRM) with a feudal-style manager-worker interface: a slow high-level module periodically emits a normalized directional subgoal that persists for P low-level steps, biasing the worker's hidden-state updates and supplying an intrinsic cosine alignment loss. On ARC and ConceptARC, we find that subgoal persistence -- not subgoal injection alone -- is the central knob: moderate periods P in [3, 6] consistently outperform both very frequent (P=1) and very long horizons, with a clear minimum LM loss at P=3 (1.544 vs. 1.674 at P=1, 1.640 baseline; replicated over 5 seeds at mean 1.595, std 0.045). The intrinsic alignment weight lambda shows a complementary narrow optimum (lambda approximately 0.05). A controlled ablation at past-sweet-spot lambda isolates learned directional structure -- not architectural capacity or auxiliary loss alone -- as the source of interference when the alignment signal exceeds its optimum. Together these findings implicate a design principle for compositional planning in latent reasoning systems: medium-horizon intent must be coherent across enough computational steps for compositional structure to form.

Authors (1)

Summary

  • The paper identifies subgoal persistence as the key design variable in hierarchical latent reasoning, with manager periods of P=3–6 outperforming no-subgoal and P=1 settings on ARC-AGI.
  • A sweep finds the best performance at P=3, reaching mean language-model loss of 1.595 across five seeds, while overly frequent revision destabilizes computation and longer commitments degrade only gradually.
  • The method works best as a lightweight planning prior, with alignment weight λ≈0.05; stronger weighting harms performance, and ablations show that learned directional structure—not added capacity—causes this interference.

Motivation and problem statement

Latent reasoning architectures such as the Hierarchical Reasoning Model (HRM) perform multi-step computation inside hidden state rather than as externalized token traces, gaining speed and compactness at the cost of the implicit commitment structure that token-level chain-of-thought provides for free. In an externalized trace, each emitted token constrains subsequent deliberation; in a latent reasoner, nothing forces medium-horizon intent to persist across micro-steps. The paper asks a precise question: how long should a latent reasoner commit to an intent before revising it? The answer is operationalized through subgoal persistence — the number of low-level steps PP that a manager-emitted directional subgoal remains in force — and the resulting stability–adaptivity tradeoff is characterized empirically on ARC-AGI and ConceptARC.

The framing imports the commitment-duration lens from feudal reinforcement learning and the options framework, where commitment duration is already recognized as a first-class design choice. The transfer is nontrivial: the worker's "actions" are hidden-state updates rather than environment interactions, and the cost of a stale plan is representational rather than behavioral.

Method: Subgoal-Augmented HRM

The method extends HRM with a feudal manager–worker interface. Every PP low-level micro-steps, the high-level state zHz^H is projected through a learned matrix WgW_g and ℓ2\ell_2-normalized to produce a directional subgoal gkg_k. The authors argue explicitly for direction rather than target-state goals: a target would be brittle under nonstationary recurrent dynamics, whereas direction encodes "where to move" without overspecifying execution. A learned scalar gate αk=σ(wα⊤ztkH)\alpha_k = \sigma(w_\alpha^\top z^H_{t_k}) modulates commitment strength.

During its persistence window, the active goal enters the worker's update as an additive bias via a projection VLV_L, optionally also into the high-level update. Injection alone does not guarantee progress along the issued direction, so a cosine alignment loss rewards net low-level displacement over the window:

Lalign(k)=1−cos⁡(ΔzkL, gk),ΔzkL=ztk+PL−ztkL\mathcal{L}^{(k)}_{\text{align}} = 1 - \cos(\Delta z^L_k,\, g_k), \qquad \Delta z^L_k = z^L_{t_k+P} - z^L_{t_k}

added to the HRM objective (task loss plus ACT halting losses) with weight λ\lambda. Because HRM detaches state between segments, alignment gradients do not cross segment boundaries. The backbone follows the standard HRM configuration: hidden size 512, 4+4 transformer layers, PP0 high-level update period, ACT with up to 16 internal steps.

Persistence is the central knob

Sweeping PP1 at fixed PP2 yields three findings that together constitute the paper's strongest claim:

Manager period PP3 LM loss
1 1.674 (worst; above baseline)
2 1.638
3 1.544 (best)
4 1.564
5 1.590
6 1.564
8 1.568
no-subgoal baseline 1.640

First, persistence is necessary, not optional: at PP4 the full subgoal infrastructure underperforms the no-subgoal baseline (PP5 vs.\ PP6). This is presented as the cleanest evidence that the persistence mechanism, not the injection mechanism, carries the benefit — a subgoal overwritten every step provides no temporal coherence and effectively injects noise. Second, the transition from PP7 to PP8 produces the largest single-step improvement in the sweep (PP9 relative to baseline), suggesting a minimum coherence horizon below which compositional structure cannot form. Third, decay beyond the optimum is gradual: loss stays within zHz^H0 out to zHz^H1. The asymmetry supports the interpretation that staleness is a soft failure mode while absence of commitment is not.

Replication across 5 seeds at the best configurations gives mean LM loss zHz^H2 (std zHz^H3) at zHz^H4 and zHz^H5 (std zHz^H6) at zHz^H7; seed variability is small relative to the zHz^H8-vs-zHz^H9 gap, ruling out a single-seed artifact. One caveat the authors flag: cross-task validation on ConceptARC-mini shows only a marginal improvement (WgW_g0 vs.\ WgW_g1 baseline, roughly WgW_g2), which they themselves downgrade to a directionally consistent secondary observation rather than independent benchmark evidence.

Narrow optimum in the alignment weight

Fixing WgW_g3 and sweeping WgW_g4 reveals a narrow optimum at WgW_g5: LM loss WgW_g6 there, versus WgW_g7 at WgW_g8, WgW_g9 at ℓ2\ell_20 (essentially matching baseline), and degradation below baseline for ℓ2\ell_21. Token-level accuracy tracks the same optimum (ℓ2\ell_22 at ℓ2\ell_23). The interpretation is that the alignment signal functions best as a lightweight planning prior rather than a dominant geometric constraint; when it competes with the task gradient on comparable terms, it reduces the worker's representational flexibility. Removing the ℓ2\ell_24 normalization of ℓ2\ell_25 or the commitment gate degrades performance, confirming the importance of directional semantics and the soft-prior role.

Past-sweet-spot ablation isolates learned directional structure

A three-cell controlled ablation at ℓ2\ell_26, ℓ2\ell_27 (batch size 64 on a single L4 GPU, num_aug=100) distinguishes architectural capacity, auxiliary loss, and learned directional content:

Cell Train LM loss Δ vs. baseline
A_full (learned directions) 1.327 +0.100
B_baseline (vanilla HRM) 1.227 —
E_random (random unit directions) 1.230 +0.003

Two contrasts are decisive. Random directions with identical architecture and auxiliary loss match baseline within ℓ2\ell_28 — an order of magnitude below the main study's seed-level std of ℓ2\ell_29 — so neither added capacity nor the intrinsic term alone causes harm. The full mechanism's gkg_k0 gap over E_random therefore attributes past-sweet-spot interference specifically to learned directional content: when over-weighted, real directional structure captures representational capacity otherwise used for the task objective. As the authors put it, a mechanism that does nothing past its optimum cannot do anything at it either — the fact that learned directions help at gkg_k1 and hurt at gkg_k2 is two sides of the same representational work.

A methodological caveat deserves emphasis: absolute loss values are not comparable between the main study (CPU, batch size 768, num_aug=1000) and the ablation (GPU, batch size 64, num_aug=100); the same nominal configuration yields gkg_k3 versus gkg_k4. All claims rest on within-regime contrasts, and the ablation uses a single fixed seed evaluated at one optimization step.

Limitations and open questions

The paper is candid about scope. Evaluation is restricted to ARC-style grid puzzles, and the held-out ConceptARC-mini gain is too small to serve as generalization evidence. The ablation's single-seed design is bounded by main-study variance but not directly replicated. Most substantively, the paper offers no representation-level analysis of how persistent subgoals shape worker hidden-state geometry — the analysis that would directly probe the claimed compositional substrate — because late-training checkpoints were not retained; this is left explicitly open. Whether the identified sweet spot gkg_k5 transfers beyond this task family and backbone scale also remains untested.

Conclusion

This work establishes subgoal persistence, not subgoal injection, as the operative variable in feudal-style latent reasoning: moderate manager periods (gkg_k6) yield consistent gains, gkg_k7 actively harms performance relative to no subgoals, and staleness degrades far more gracefully than instability. Combined with the narrow gkg_k8 optimum and the ablation attributing past-sweet-spot interference to learned directional structure, the results support a concrete design principle: medium-horizon intent must remain coherent across enough computational steps for compositional structure to form, while remaining soft enough not to override task optimization.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 2 tweets with 13 likes about this paper.