Evolving in Thought Space: Training a Small Model at Test Time Unlocks Better Discoveries

This presentation explores Guidance-TTT, a novel test-time reinforcement learning framework that separates strategic decision-making from code implementation. Instead of adapting the large model that generates complete programs, the system trains only a compact guidance model to propose high-level modifications while a powerful frozen executor translates those strategic ideas into working code. Across four challenging discovery tasks—from combinatorial optimization to GPU kernel design—this hierarchical approach achieves state-of-the-art results while dramatically reducing computational cost, demonstrating that learning in thought space can be more effective than learning in solution space.
Script
Most program discovery systems face a costly dilemma: the model smart enough to implement a solution reliably is often too expensive to retrain every time you search for a better approach. Guidance TTT solves this by splitting the job, adapting only a small model that learns which changes to propose while a frozen powerhouse translates those ideas into working code.
The system maintains an archive of tested programs, each linked to its parent and the strategic modification that created it. When verifier scores come back, only the compact guidance policy gets updated. The expensive executor that writes the actual code never sees a gradient, yet the whole system learns which high level changes are worth pursuing.
Guidance TTT was tested on four structurally different tasks: polyomino packing, where pieces must tile efficiently; lasso regularization paths, requiring sparse solutions across penalty levels; an agricultural planning challenge balancing investment and harvest timing; and GPU kernel optimization for matrix multiplication. These aren't toy problems—they demand real algorithmic insight, not just parameter tweaking.
On polyomino packing, Guidance TTT reached a score of 91.89, exceeding even a web-enabled four-agent baseline at 89.40. The ablations tell the real story: freezing the guidance policy drops performance to 84.85, while adapting the solution model directly collapses to 41.63. Learning strategic modifications outperforms learning complete rewrites, even when the rewrite model has more expressive power.
The method isn't perfect. Execution errors create ambiguous feedback: a great strategic idea can get zero reward if the executor fails to implement it correctly. And archive search can get stuck refining a narrow family of solutions rather than exploring fundamentally different approaches. After reaching a peak, dozens of valid local improvements sometimes fail to push the global frontier forward.
Guidance TTT proves that where you place learning matters as much as how much you learn. By evolving strategic intent rather than executable artifacts, a small model can steer discovery more effectively than a large one adapting its entire solution. To dive deeper into the paper, explore the full methods, or create your own video summary, visit EmergentMind.com and see how test time adaptation is reshaping program search.