Papers
Topics
Authors
Recent
Search
2000 character limit reached

Epistemic Uncertainty for Test-Time Discovery

Published 11 May 2026 in cs.LG and cs.AI | (2605.11328v1)

Abstract: Automated scientific discovery using LLMs relies on identifying genuinely novel solutions. Standard reinforcement learning penalizes high-variance mutations, which leads the policy to prioritize familiar patterns. As a result, the maximum reward plateaus even as the average reward increases. Overcoming this limitation requires a signal that distinguishes unexplored regions from intrinsically difficult problems. This necessitates measuring disagreement across independently adapted weight hypotheses rather than relying on a single network's confidence. UG-TTT addresses this challenge by maintaining a small ensemble of low-rank adapters over a frozen base model. The per-token disagreement, quantified as the mutual information between ensemble predictions and weight hypotheses, isolates epistemic uncertainty and identifies positions where insufficient coverage leads to adapter divergence rather than intrinsic problem difficulty. This measure is incorporated as an exploration bonus into the policy gradient, directing the policy toward positions where persistent adapter disagreement signals low training coverage, the same frontier where genuine discovery is possible. A nuclear norm regularizer ensures the adapters remain distinct from one another, thereby preserving the exploration signal throughout training. Across four scientific discovery benchmarks, UG-TTT increases the maximum reward on three tasks, maintains substantially higher solution diversity, and an ablation study confirms that the regularizer is essential for sustaining this behavior.

Summary

  • The paper introduces Uncertainty-Guided Test-Time Training (UG-TTT), combining BALD-style token uncertainty, temperature-coupled exploration bonuses, and nuclear-norm adapter regularization to prevent diversity collapse.
  • UG-TTT improves maximum reward on three of four benchmarks, retains 1.1–1.7 bits of solution-family entropy, reaches the CP26 ceiling with 384 versus 25,600 baseline rollouts, and finds AC1 targets about 16 times faster.
  • The nuclear-norm regularizer is essential: without it, adapter subspaces align within two epochs, mutual information falls by 3–4 orders of magnitude, and uncertainty-guided exploration effectively disappears.

Motivation: diversity collapse as a structural failure

The paper addresses a specific pathology in LLM-driven scientific discovery systems such as AlphaEvolve and TTT-Discover (Novikov et al., 16 Jun 2025, Yuksekgonul et al., 22 Jan 2026): while average reward rises under standard policy-gradient optimization, the maximum reward — the quantity that matters for genuine discovery — plateaus prematurely. The authors argue this is not an incidental search failure but a structural consequence of the objective itself. Expected-reward maximization penalizes high-variance mutations, filtering out uncertain but potentially high-payoff directions and driving mode collapse onto "safe" solutions from the pretraining distribution. Existing remedies (entropy bonuses, difficulty-aware certainty weighting) derive their signal from a single network, which conflates aleatoric uncertainty with epistemic uncertainty; only weight-space disagreement can separate them.

The proposed framework, Uncertainty-Guided Test-Time Training (UG-TTT), extends TTT-Discover with three coupled components: (i) an ensemble of KK LoRA adapters over a frozen base model whose per-token disagreement yields a BALD-style mutual information signal; (ii) an exploration bonus injected into the policy gradient via a shaped advantage whose gain is coupled to the adaptive entropic temperature β(s)\beta(s); and (iii) a nuclear-norm regularizer on stacked adapter down-projections that provably enforces subspace orthogonality across adapters (2605.11328).

Method

Token-level epistemic uncertainty. Each adapter Θ(k)=Θbase+B(k)A(k)\Theta^{(k)} = \Theta_{\text{base}} + B^{(k)}A^{(k)} is read as a sample from a variational posterior over the shared frozen base. At each decoding position tt, the epistemic component is isolated by the BALD decomposition:

MIt=H(pˉt)−1K∑kH(pθk(⋅∣q,o<t)),MI_t = H(\bar p_t) - \frac{1}{K}\sum_k H(p_{\theta_k}(\cdot \mid q, o_{<t})),

where pˉt\bar p_t is the mixture predictive. This quantity is non-negative, vanishes when members agree, and grows when they place mass on incompatible continuations — precisely the operational signature of unresolved positions. Because most tokens in code rollouts are syntactic boilerplate on which adapters trivially agree, MItMI_t is reduced to a scalar UiU_i via the top-7%7\% mean over positions, a threshold motivated by prior findings that high-entropy minority tokens carry the RL signal (Wang et al., 2 Jun 2025) and that entropy-triggered mechanisms fire on roughly 4–12% of decoding steps.

Shaped advantage. Within each rollout group, UiU_i is standardized and clipped to β(s)\beta(s)0 (provably inactive for group sizes β(s)\beta(s)1, Proposition 1), then added to the leave-one-out advantage as β(s)\beta(s)2. Two design choices matter. The prefactor β(s)\beta(s)3 rescales the bonus to the magnitude of the reward signal, damping exploration in near-uniform groups. More importantly, the coefficient is coupled to the entropic temperature, β(s)\beta(s)4: since β(s)\beta(s)5 grows monotonically as within-group rewards equalize — the regime in which the policy contracts onto a single template — exploration pressure intensifies exactly at collapse onset rather than following a fixed schedule. The authors are explicit that the bonus is intentionally biased toward high-MI rollouts; Proposition 1 bounds only the residual scale perturbation, not estimator unbiasedness.

Subspace-diversity regularizer. The central technical obstacle is that jointly trained adapters share data and objectives and converge to identical weights, driving mutual information to zero — a degeneracy the authors observe empirically within two to three epochs. Their remedy maximizes the nuclear norm of the stacked down-projections β(s)\beta(s)6. Proposition 2 establishes that, under equal adapter-wise Frobenius norms and β(s)\beta(s)7, the global maximizer consists of blocks with mutually orthogonal row spaces and equal singular values β(s)\beta(s)8 — exactly the geometric condition keeping the BALD decomposition non-trivial. Regularization applies only to down-projections β(s)\beta(s)9; joint regularization of up-projections oscillates, consistent with prior observations (Zhai et al., 2023). Proposition 3 confirms that UG-TTT reduces to the TTT-Discover objective when Θ(k)=Θbase+B(k)A(k)\Theta^{(k)} = \Theta_{\text{base}} + B^{(k)}A^{(k)}0 or when adapters are tied and Θ(k)=Θbase+B(k)A(k)\Theta^{(k)} = \Theta_{\text{base}} + B^{(k)}A^{(k)}1.

Experimental results

All runs use Qwen3-8B with Θ(k)=Θbase+B(k)A(k)\Theta^{(k)} = \Theta_{\text{base}} + B^{(k)}A^{(k)}2 rank-16 adapters on four TTT-Discover benchmarks (AC1, AC2, CP26, Erdős), six epochs on a single RTX Pro 6000 (~32 h per run).

Problem Baseline Θ(k)=Θbase+B(k)A(k)\Theta^{(k)} = \Theta_{\text{base}} + B^{(k)}A^{(k)}3 UG-TTT Θ(k)=Θbase+B(k)A(k)\Theta^{(k)} = \Theta_{\text{base}} + B^{(k)}A^{(k)}4 Θ(k)=Θbase+B(k)A(k)\Theta^{(k)} = \Theta_{\text{base}} + B^{(k)}A^{(k)}5 Baseline Θ(k)=Θbase+B(k)A(k)\Theta^{(k)} = \Theta_{\text{base}} + B^{(k)}A^{(k)}6 UG-TTT Θ(k)=Θbase+B(k)A(k)\Theta^{(k)} = \Theta_{\text{base}} + B^{(k)}A^{(k)}7
AC1 0.6381 0.6406 +0.0025 0.70 1.71
AC2 0.8532 0.8563 +0.0031 0.50 1.12
CP26 2.6302 2.6359 +0.0058 0.70 1.23
Erdős 2.6174 2.6167 −0.0007 0.67 1.62

UG-TTT improves Θ(k)=Θbase+B(k)A(k)\Theta^{(k)} = \Theta_{\text{base}} + B^{(k)}A^{(k)}8 on three of four problems and matches the fourth to within 0.03% of the bound, while retaining 1.1–1.7 bits of solution-family entropy where the baseline collapses below 0.71 bits — a cross-task mean entropy gap of Θ(k)=Θbase+B(k)A(k)\Theta^{(k)} = \Theta_{\text{base}} + B^{(k)}A^{(k)}9. Absolute deltas understate the effect's position on the discovery curve: on AC1, UG-TTT crosses tt0 at step 12 versus step 189 for the baseline (~16× faster); on CP26 it matches the published TTT-Discover ceiling using 384 rollouts versus 25,600 for the baseline on the same model. On AC2, the gain comes entirely from the scipy.differential_evolution family, which yields zero correct rollouts under the unregularized ensemble. On Erdős, where the scalar bound ties, UG-TTT logs 13 new-best events versus 7 (+86%) and nearly doubles discretization resolution.

The mechanism claim is supported directly: surviving families are those producing winning solutions (e.g., the Cauchy–Schwarz target tt1 appears in 18/30 best AC1 rollouts versus 1/30 for the baseline), and every such motif lies inside a family the baseline extinguishes by the final epoch. Preserved diversity is thus a precondition for discovery, not merely a side effect.

The ablation identifies the nuclear-norm regularizer as load-bearing. Without NNM, adapter row spaces align within two epochs, ensemble MI falls to ~tt2 (three to four orders of magnitude below the regularized run), the MI bonus silently zeroes out, and the run reverts to ordinary RL on tied adapters, stalling at tt3. The full method reaches its ceiling on 25% fewer training tokens (tt4 vs. tt5), aided by streaming MI early-stop, which fires on 50.0%/55.6% of AC1/CP26 rollouts.

Limitations and open questions

The evaluation covers one base model (Qwen3-8B), tt6, a single GPU configuration, and a single seed across four benchmarks, so entropy gains and streaming behavior may not transfer. Streaming early-stop required paired empirical calibration: it reduced tt7 by 0.0148 on AC2 and 0.0017 on Erdős, and no length–reward diagnostic predicts this asymmetry in advance — the calibration remains empirical rather than principled. The three coupled hyperparameters (tt8, tt9, MIt=H(pˉt)−1K∑kH(pθk(⋅∣q,o<t)),MI_t = H(\bar p_t) - \frac{1}{K}\sum_k H(p_{\theta_k}(\cdot \mid q, o_{<t})),0) and the fixed top-7% token threshold are heuristic. The ensemble incurs roughly MIt=H(pˉt)−1K∑kH(pθk(⋅∣q,o<t)),MI_t = H(\bar p_t) - \frac{1}{K}\sum_k H(p_{\theta_k}(\cdot \mid q, o_{<t})),1-fold parameter overhead plus a chunked MIt=H(pˉt)−1K∑kH(pθk(⋅∣q,o<t)),MI_t = H(\bar p_t) - \frac{1}{K}\sum_k H(p_{\theta_k}(\cdot \mid q, o_{<t})),2-way scoring pass, only partially offset by streaming. Finally, the method inherits deterministic program-checker rewards from TTT-Discover; porting to learned or noisy verifiers would require recalibrating the MI bonus, and the authors note a dual-use risk: rewarding verifier-uncertain regions under a loosely specified verifier would actively seek loopholes.

Conclusion

UG-TTT reframes premature reward plateauing in test-time discovery as a structural consequence of single-model, scalar-reward optimization, and addresses it by making ensemble weight-space disagreement — a Bayesian epistemic signal — part of the training objective itself. The nuclear-norm regularizer is essential: without it, the uncertainty signal collapses within two epochs and the method degenerates to the baseline. The strongest quantitative claims are the sustained 3–4 orders of magnitude higher ensemble mutual information, the 1.1–1.7 bits of retained family entropy against baseline collapse, and the CP26 state-of-the-art match at roughly 1.7% of the baseline rollout budget. The main open question left by the paper is whether the epistemic-exploration mechanism transfers beyond deterministic verifiers and beyond the single-model, single-seed configuration evaluated here.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 1 like about this paper.