Principled Thoughts for Latent Recursive LLM Systems
Abstract: LLMs can reason in continuous space instead of decoded text, by recurring on their own hidden states or by passing those states between agents, while training supervises only the Cross-Entropy (CE) of the final decoded answer and does not constrain the thought. Theoretical and empirical analyses establish and confirm four failures of CE-only training that lead to a lower probability of the correct answer such as collapsing thoughts across distinct questions and retaining irrelevant information. We introduce REST (REpresentation-Supervised Thoughts), a training objective that turns four properties of a valid thought representation (causality, minimality, separability, and stability) into differentiable losses added to CE. We instantiate it in latent single-agent and multi-agent systems, without architectural changes or added parameters at inference. Across 7 benchmarks spanning mathematics, science, medicine, and code generation, with the same training data, compute, and latent budget, REST increases accuracy over CE-only training across agent settings and model sizes by up to 7.5 percentage points and convergence on a final answer by 30\%. Furthermore, REST thoughts encode more of what is required to achieve the correct answer, and decoding them better recovers the intended output of the agent, which makes latent communication easier to interpret. Project Website: https://fard-lab.github.io/REST
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper studies how artificial intelligence systems can think without writing every thought as text.
Usually, a LLM, such as ChatGPT, reasons by generating words one after another. For example, to solve a math problem, it might write:
First, I calculate... Then, I substitute... Therefore, the answer is...
The paper explores a different method. Instead of turning every thought into words, the model keeps its reasoning inside its hidden states—large groups of numbers inside the model. These hidden states are called latent thoughts.
The researchers argue that simply training the model to get the final answer right is not enough. The hidden thoughts might be messy, contain irrelevant information, or become too similar for different questions.
To solve this problem, they introduce a training method called REST, short for REpresentation-Supervised Thoughts.
2. What questions does the research ask?
The paper focuses on several main questions:
- Can a model’s hidden thoughts be trained to contain the right information?
- What goes wrong when hidden thoughts are trained only to produce a correct final answer?
- Can models be encouraged to:
- keep useful information,
- remove unnecessary information,
- represent different questions differently, and
- show uncertainty when several answers are possible?
- Does this improve performance on mathematics, science, medicine, and programming tasks?
- Does REST work both when:
- one model thinks repeatedly by itself, and
- several models pass hidden thoughts to one another?
The researchers describe these goals using four properties:
| Property | Simple meaning |
|---|---|
| Causality | The thought should contain information that helps produce the answer. |
| Minimality | The thought should leave out information that is not needed. |
| Separability | Different questions should create different thoughts. |
| Stability | The thought should represent uncertainty and possible answers, not just one lucky answer. |
A useful analogy is a student making notes before an exam. Good notes should contain the important facts, leave out unrelated details, look different for different topics, and show when the student is unsure.
3. How did the researchers conduct the study?
Latent reasoning
An LLM is made of many layers that process numbers called hidden states. These numbers contain information about the input and the model’s current reasoning.
In a normal system, the model turns its reasoning into text. In the systems studied here, the model can instead send its hidden states directly into another step of reasoning. This is similar to passing a student’s private notes to another student without reading the notes aloud.
The paper tests two types of systems:
- Single-agent system: One model repeatedly uses its own hidden thoughts to continue reasoning.
- Multi-agent system: Several models work together:
- a planner suggests an approach,
- a refiner improves it,
- a solver gives the final answer.
The models themselves are kept frozen, meaning their main abilities are not changed. The researchers train only the connections that transfer hidden thoughts from one reasoning step or agent to another.
The usual training method
The normal method uses cross-entropy loss, often shortened to CE. This is a score showing how surprised the model is by the correct answer.
For example, if the correct answer is “42” and the model gives “42” a high probability, the CE score is low. If the model gives it a very low probability, the score is high.
CE only checks the final answer. It does not carefully check whether the hidden thought was useful or sensible.
The REST method
REST keeps the usual CE score but adds four extra checks, one for each desired property:
- Causality check: Does the hidden thought help the next model behave as if it had received the producer’s written explanation?
- Minimality check: Does the thought focus on the producer’s useful output instead of copying the original question or irrelevant details?
- Separability check: Are thoughts for different problems far enough apart to avoid confusion?
- Stability check: Does the thought show how uncertain the producer was?
These checks are converted into mathematical penalties called losses. During training, the system tries to reduce both the final-answer error and these thought-related errors.
The researchers tested REST on seven types of benchmarks, including:
- mathematics,
- difficult science questions,
- medical questions, and
- computer programming.
They compared REST with CE-only training and with other methods that try to make hidden thoughts resemble written reasoning.
4. What did the researchers find?
REST improved accuracy
REST generally produced more correct answers than CE-only training.
Across the experiments, the average improvement was approximately:
- 3.3 percentage points for single-agent systems,
- 3.5 percentage points for multi-agent systems.
The strongest settings improved accuracy by as much as:
- 6.5 points in a single-agent system,
- 7.5 points in a multi-agent system.
These improvements occurred while using the same training data, similar computing resources, and the same amount of space for latent thoughts.
CE-only thoughts had important problems
The researchers found that training only on the final answer caused hidden thoughts to behave poorly.
They often:
- became too similar for unrelated questions,
- included too much information from the original prompt,
- failed to preserve the useful information from the producer’s answer, and
- did not properly represent uncertainty.
This is like asking students only whether they got the final answer correct. Some students might guess correctly, copy irrelevant notes, or use confusing shortcuts. Their final answer may be right, but their notes would not be reliable or useful to another student.
REST preserved more useful information
REST thoughts contained more of the producer’s actual plan or answer and less irrelevant information from the original input.
When the researchers replaced a model’s normal written explanation with its hidden thought, REST preserved more of the model’s performance than CE-only training did. This suggests that REST thoughts carried more useful information.
REST helped models reach final answers
REST sometimes used more output tokens—about 15.4% more on average. However, the researchers argue that these extra tokens were not simply wasted. The REST systems were more likely to finish and clearly produce a final answer.
The rate of reaching a final answer increased from about:
- 73% with CE-only training
- to 95% with REST
So, REST may make the model think for longer, but it also makes the model more likely to complete the task successfully.
Results were generally consistent across system sizes
REST helped both smaller and larger models. Larger models benefited especially from additional rounds of hidden reasoning, while smaller models often gained most of their improvement after one round.
The different REST properties were not equally useful in every experiment. For example, minimality was especially helpful in some single-agent settings, while combinations of properties worked particularly well in multi-agent systems.
5. Why is this research important?
This paper suggests that training an AI only to produce the correct final answer is like grading a student only on the last line of a test. The student may get the answer right for the wrong reasons, or their work may be impossible for someone else to understand or use.
REST provides a way to train the model’s hidden reasoning itself. It encourages the model to create thoughts that are:
- useful,
- focused,
- different for different problems, and
- able to represent uncertainty.
If this approach continues to work, it could make latent-reasoning AI systems more accurate and dependable. It might also make communication between several AI agents more effective because each agent would receive a clearer and more useful internal message.
However, the method has a cost: REST can require more computation or produce longer final responses. Also, the experiments do not prove that the hidden thoughts are fully understandable to humans. They show that the thoughts contain more useful information and behave better according to the researchers’ tests.
Overall, the paper’s main message is that AI should not be trained only to give the right answer. Its internal reasoning should also be organized and supervised so that it carries the right information to the next step.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- Limited model and system diversity: The evaluation uses only two size regimes and a small set of Qwen, Llama, and Gemma models; it remains unclear whether REST generalizes to substantially larger models, other architectures, multimodal models, or models trained specifically for reasoning.
- Incomplete baseline coverage: Comparisons are primarily against CE-only training, CODI, SIM-CoT, and a text baseline. REST is not compared with broader alternatives such as reinforcement learning, preference optimization, verifier-guided training, latent-state contrastive learning, or larger latent-step budgets.
- Unclear contribution of each loss component: Although single-property ablations are reported, the study does not fully disentangle interactions among causality, minimality, separability, and stability, especially when their gradients conflict or when one property dominates optimization.
- No systematic analysis of loss interference: The paper does not quantify whether optimizing one property degrades another, nor does it study gradient conflicts, loss scaling, or adaptive weighting among the four auxiliary objectives.
- Hyperparameter robustness is incompletely established: The paper reports selected values and some weight sweeps, but the sensitivity to , , temperature , the number of negative examples , pooling architecture, and probe capacity is not systematically characterized.
- Potential train–inference mismatch: Training uses teacher-forced producer outputs, whereas inference generates latent trajectories without access to the ground-truth or producer text. The extent to which REST remains effective under longer free-running trajectories and accumulated generation errors is unresolved.
- Stability is only approximated through entropy: The stability loss matches the producer’s mean token-level predictive entropy along one sampled or teacher-forced output. This does not establish that the thought encodes the full conditional output distribution, alternative solution strategies, calibration, or uncertainty over complete sequences.
- Validity of the entropy proxy is uncertain: The paper does not show that matching average predictive entropy improves downstream uncertainty estimation or preserves the producer’s distribution over complete answers. Entropy can be identical for distributions with very different support and semantics.
- Minimality may depend on an imperfect causal interpretation: The minimality objective uses consumer-side cross-entropies to infer information about producer inputs and outputs, but it is not demonstrated that reducing recoverability of the producer input corresponds to removing genuinely irrelevant information rather than useful context.
- Separability may encourage arbitrary dispersion: The separability loss penalizes cosine similarity with preceding training examples, regardless of whether those examples are semantically related. The study does not test whether it separates genuinely distinct tasks while preserving similarity among equivalent or paraphrased questions.
- Separability evaluation is largely geometric: PCA visualizations and pooled-vector cosine distances do not establish that representations are semantically separable in a way that improves compositional reasoning or prevents harmful collisions under distribution shift.
- Attention pooling introduces an additional learned representation: Separability and stability are measured through learned pooling and probing modules. It is unclear whether the observed properties hold in the full thought sequence independently of these probes or whether the probes learn to extract information that is not usable by the consumer.
- No causal intervention tests for representation properties: The analysis primarily correlates REST-trained representations with accuracy, decoding, clustering, and answer rates. Controlled interventions—such as replacing, perturbing, shuffling, or selectively masking thought dimensions—are needed to establish that the targeted properties causally produce the gains.
- Interpretability claims remain limited: Better decoding of thoughts and higher recovery of oracle-text performance do not demonstrate that thoughts are human-interpretable, faithful to the model’s actual computation, or suitable for reliable monitoring.
- Semantic equivalence of decoded thoughts is not rigorously evaluated: The analysis compares decoded content with producer inputs and outputs, but it does not specify robust semantic similarity measures or determine whether decoded thoughts preserve reasoning steps, conclusions, or merely correlated lexical information.
- Accuracy gains may partly reflect increased computation: REST often increases token generation, particularly in multi-agent and scaled settings. The comparison does not fully normalize wall-clock time, FLOPs, energy, memory, or total inference cost, so the gains cannot yet be attributed solely to better representations.
- No compute-normalized comparison across latent budgets: The study matches the latent budget but does not evaluate whether CE-only systems with more latent steps, more recursion rounds, or more sampling can achieve comparable accuracy at similar computational cost.
- Best-round reporting may introduce selection bias: Each system is reported at the recursion round that performs best, while the practical procedure for selecting that round is not specified. This may overstate performance if the optimal round is chosen using evaluation results.
- Small or unstable benchmark subsets limit statistical confidence: Several reported benchmarks, especially AIME and GPQA, contain relatively few evaluation examples. The paper provides averages over three training seeds but does not report confidence intervals, significance tests, per-example variance, or sensitivity to benchmark composition.
- Benchmark contamination and temporal generalization are not addressed: The paper does not establish that the base models or training data are free from contamination involving the evaluation benchmarks, including the newer AIME and LiveCodeBench versions.
- Training-data scale and composition are narrow: Training relies on Sequential-Math derived from s1K and m1K data. It remains unknown whether REST works with noisy, weakly supervised, multilingual, adversarial, or substantially larger datasets.
- Domain generalization is incomplete: The benchmarks cover mathematics, science, medicine, and code generation, but do not test dialogue, planning, retrieval, long-context reasoning, factual knowledge under uncertainty, safety-critical decisions, or multimodal tasks.
- Robustness to distribution shift is unexplored: The paper does not evaluate out-of-distribution questions, paraphrases, adversarial prompts, domain shifts, or changes in prompt formatting and role assignments.
- Robustness to noisy or incorrect producer outputs is unknown: Since later agents rely on earlier latent thoughts, the effect of incorrect, uncertain, contradictory, or deliberately misleading producer reasoning on REST-trained systems remains unexamined.
- Error propagation across recursion is insufficiently characterized: The paper reports aggregate performance by round but does not identify how REST changes the types, frequency, and persistence of errors as thoughts circulate through multiple agents and rounds.
- Scalability to larger numbers of agents and rounds is unresolved: Experiments use either a single self-looping agent or a three-agent planner–refiner–solver chain. The computational and optimization behavior of REST with more agents, heterogeneous links, asynchronous communication, or many recursion rounds is unknown.
- Outer-link capacity is not systematically studied: Only the outer link is trained and the inner link is frozen. It remains unclear whether the benefits depend on the capacity, initialization, architecture, or trainability of , and whether jointly training both links would improve or destabilize performance.
- Frozen base models constrain the conclusions: Because all base LLM parameters and inner links remain frozen, the results do not show whether REST remains effective when the underlying models are jointly fine-tuned or when the latent interface is learned end-to-end.
- Transferability across model pairs is unclear: The study uses fixed heterogeneous model assignments. It does not test whether a thought trained for one producer–consumer pair transfers to another consumer, model family, tokenizer, hidden size, or prompt template.
- Theoretical guarantees rely on strong assumptions: The claimed lower bounds and divergence arguments do not establish that optimizing the proposed surrogate losses guarantees improved accuracy in realistic, nonstationary, multi-round systems. The assumptions behind the theoretical results are not stress-tested empirically.
- The causality loss may overconstrain useful transformations: Matching the consumer’s predictive distribution under textual and latent transfers could discourage beneficial compression, abstraction, or alternative reasoning representations. The trade-off between faithful transfer and useful latent transformation is not investigated.
- The relationship between representation quality and final answer correctness is not fully established: REST improves average benchmark accuracy, but the paper does not determine which representation metrics best predict per-example success or whether the four properties are necessary and sufficient for reliable reasoning.
- Failure cases are not analyzed in detail: The paper does not provide systematic qualitative or quantitative analyses of examples where REST decreases accuracy, particularly the scaled separability and stability configurations that reduce performance relative to CE-only.
- Calibration and uncertainty quality are untested: Although stability is motivated by producer uncertainty, the study does not evaluate answer calibration, selective prediction, abstention, confidence accuracy, or uncertainty propagation across agents.
- Security and privacy implications are unexplored: Latent thoughts may transmit sensitive prompt information or conceal undesirable content from ordinary text inspection. The paper does not assess information leakage, prompt memorization, malicious latent communication, or auditability.
- Reproducibility is difficult to assess from the provided text: The paper refers to appendices for datasets, hyperparameters, implementation details, and results, but the supplied manuscript does not provide enough information to independently reproduce the training setup, loss normalization, sampling procedures, or statistical analyses.
- The conclusion and implications are incomplete: The manuscript ends during the “Implications” subsection, leaving the authors’ discussion of limitations, deployment risks, practical recommendations, and broader consequences unresolved.
Practical Applications
Immediate Applications
- Software and LLM infrastructure: improved latent-reasoning fine-tuning
- Integrate REST-style auxiliary losses into existing latent-recursive LLM pipelines by adding causality, minimality, separability, and stability terms to the usual final-answer cross-entropy objective.
- This is immediately feasible because the method trains only the outer representation-transfer link and does not require architectural changes, additional inference parameters, or changes to the base LLM.
- Potential tools: a PyTorch loss module, a Hugging Face Trainer callback, or a monitoring dashboard that reports thought collision, information retention, output-distribution stability, and oracle-text substitution loss.
- Expected value: higher answer accuracy and more reliable convergence in latent single-agent systems and planner–refiner–solver pipelines.
- Dependencies and assumptions: access to hidden states, differentiable transfer links, suitable training examples with producer inputs and outputs, and sufficient compute for auxiliary forward passes. The reported gains are based on open-weight models and benchmark-style data; production gains require domain-specific validation.
- Latent multi-agent orchestration for complex technical questions
- Use REST to train heterogeneous agents that communicate through hidden representations rather than long textual intermediate messages.
- Sectors: scientific research, enterprise knowledge systems, technical support, legal document analysis, and engineering design.
- A practical workflow could assign separate agents to planning, critique, refinement, and final synthesis while using REST to reduce irrelevant context transfer and preserve useful producer information.
- Expected value: improved cross-agent communication, fewer semantic collisions between unrelated tasks, and better retention of plans across multiple recursive rounds.
- Dependencies and assumptions: agents must have compatible or learnable representation-transfer links; privacy and audit requirements may favor textual fallbacks because latent messages are harder to inspect directly.
- Single-model latent reasoning for lower-cost inference
- Apply the single-agent self-loop configuration to models that need additional reasoning depth without generating a full chain of thought in text.
- Sectors: customer-service automation, coding assistants, educational tutoring, and personal productivity tools.
- Minimality is particularly relevant when the model repeatedly receives the same question: it can encourage latent states to store solution-relevant information rather than duplicating the prompt.
- Expected value: improved accuracy without necessarily increasing decoded reasoning length; in the Light single-agent configuration, minimality improved accuracy while reducing total output tokens.
- Dependencies and assumptions: the inner latent-step link must already function reliably, and inference latency must be evaluated at the hidden-state level rather than only by token count.
- Medical decision-support and scientific question answering
- Fine-tune latent reasoning systems for medical, scientific, and biomedical question-answering workflows using REST as a representation-quality objective.
- The paper reports improvements on MedQA and GPQA-Diamond, suggesting possible use in literature triage, clinical knowledge retrieval, diagnostic hypothesis generation, and scientific assistants.
- Potential workflow: a planner identifies relevant mechanisms, a refiner checks the reasoning, and a solver produces a cited or structured answer.
- Dependencies and assumptions: benchmark accuracy does not establish clinical safety. Deployment would require expert review, calibrated uncertainty, retrieval grounding, data privacy controls, and prospective evaluation. Stability should not be interpreted as a substitute for clinically meaningful confidence estimates.
- Code-generation assistants and automated program repair
- Apply REST-trained latent agents to code generation, debugging, unit-test creation, and repository-level planning.
- A latent planner could generate an implementation strategy, a refiner could identify edge cases, and a solver could produce code and tests.
- Expected value: better preservation of implementation plans across agents and potentially fewer failures caused by irrelevant prompt information.
- Dependencies and assumptions: code correctness must be measured through compilation, tests, security analysis, and repository-level benchmarks rather than answer accuracy alone. The reported token increases in some multi-agent configurations may offset computational savings.
- Representation-level monitoring and debugging of LLM systems
- Use the four REST properties as diagnostic metrics even when REST is not used for training.
- Teams can test whether latent thoughts:
- preserve answer-relevant information under text-to-latent substitution (causality);
- suppress duplicated or irrelevant prompt content (minimality);
- remain distinguishable for semantically different inputs (separability); and
- encode uncertainty or alternative candidate solutions (stability).
- Potential products: latent-state regression tests, representation drift monitors, and pre-deployment safety audits for recurrent or multi-agent LLM services.
- Dependencies and assumptions: the metrics require access to internal representations and may be sensitive to pooling choices, similarity measures, and probe design.
- Academic training and benchmarking for latent reasoning
- Use REST as a reproducible baseline for research on latent chain-of-thought, recurrent LLMs, hidden-state communication, and neural representation learning.
- The method provides a common experimental framework for comparing final-answer supervision with representation-level supervision across model sizes, domains, recursion depths, and agent topologies.
- Potential outputs: benchmark suites that separately evaluate answer accuracy, thought separability, information preservation, uncertainty encoding, convergence rate, and inference cost.
- Dependencies and assumptions: the paper’s reported results use particular model families, datasets, loss weights, and latent budgets. Independent replication is needed before treating the gains as universal.
- Policy and governance evaluations for opaque AI systems
- Regulators and internal governance teams can incorporate REST-style tests into audits of systems that reason or communicate internally through hidden states.
- For example, separability tests may identify representation collapse, while causality tests can measure whether an internal message actually preserves the information claimed by an upstream agent.
- Potential use: evidence for model cards, deployment reviews, algorithmic impact assessments, and documentation of latent-agent communication behavior.
- Dependencies and assumptions: the approach evaluates functional properties, not semantic interpretability or intentionality. Organizations would need standardized access protocols for hidden states and safeguards against exposing proprietary model internals.
Long-Term Applications
- Scalable latent-agent architectures for autonomous research and engineering
- Develop larger recursive systems in which specialized agents exchange compact latent plans over many rounds.
- The results suggest that scaled agents can benefit more from additional recursion, while smaller systems may obtain most of their benefit from a single round. This could support autonomous literature review, experiment planning, simulation design, and engineering optimization.
- Potential products: research copilots that maintain latent project state, multi-agent design systems, and automated experiment-planning platforms.
- Dependencies and assumptions: performance must continue to scale with recursion depth; error accumulation, representation-transfer compatibility, memory use, and verification remain unresolved. Long-horizon autonomous action would also require robust external grounding and human approval.
- Latent communication protocols between independently trained models
- Establish interoperable hidden-state communication standards so that models from different vendors or training runs can exchange plans without translating them into public text.
- REST’s causality and stability objectives could serve as training criteria for such protocols.
- Sectors: robotics, cloud AI services, edge computing, and enterprise software integration.
- Potential benefits: lower communication bandwidth, reduced exposure of sensitive intermediate text, and faster coordination between specialized models.
- Dependencies and assumptions: hidden representations are generally model-specific. Interoperability would require standardized latent spaces, alignment adapters, security controls, and defenses against malformed or adversarial latent messages.
- Robotics and embodied multi-agent coordination
- Adapt REST to systems in which perception, planning, control, and monitoring agents exchange continuous internal states.
- A planner could communicate task-relevant state to a controller while minimality suppresses irrelevant sensor history, separability distinguishes different tasks, and stability preserves alternative action possibilities.
- Potential applications: warehouse robots, autonomous vehicles, surgical robotics, and household robots.
- Dependencies and assumptions: real-world robotics requires strict latency, temporal consistency, safety guarantees, and robustness to distribution shift. The paper evaluates language benchmarks, so transfer to sensorimotor representations is an open research problem.
- Adaptive compute and energy-efficient inference
- Use stability and convergence signals to allocate more latent recursion only when the model represents substantial uncertainty, while terminating early for simple or confident cases.
- This could produce adaptive-compute systems for cloud inference, mobile assistants, and energy-constrained edge devices.
- Potential benefits: lower average energy consumption, reduced latency, and better allocation of computation across easy and difficult inputs.
- Dependencies and assumptions: the learned entropy probe must correlate with actual answer uncertainty. The paper also reports that REST can increase decoded tokens in some settings, so net energy savings are not established and would require hardware-level measurement.
- High-assurance decision-support systems with latent uncertainty tracking
- Extend the stability term to represent multiple plausible solutions rather than only a single sampled answer.
- In finance, healthcare, climate modeling, and policy analysis, this could support systems that preserve alternative scenarios for downstream agents instead of prematurely committing to one plan.
- Potential tools: scenario-ranking engines, uncertainty-aware forecasting assistants, and decision-support systems that expose candidate distributions to human reviewers.
- Dependencies and assumptions: predictive entropy is only a proxy for epistemic and aleatoric uncertainty. Reliable deployment would require calibration, distributional evaluation, abstention policies, and domain-specific uncertainty representations.
- Privacy-preserving internal collaboration
- Investigate whether latent communication can reduce the need to expose sensitive intermediate text between agents or services.
- For example, a medical or financial system might transmit a task-relevant representation rather than a full document or reasoning trace.
- Dependencies and assumptions: latent vectors are not automatically private or non-invertible. They may leak personal, proprietary, or regulated information and would require inversion testing, access controls, encryption, differential privacy, and formal leakage analysis.
- Curricula and educational tutoring systems with compact internal reasoning
- Build tutoring systems that use latent planning and refinement internally while presenting only age-appropriate explanations to learners.
- REST could help separate internal solution planning from irrelevant student-profile or prompt information and maintain multiple pedagogical strategies before selecting one.
- Potential products: adaptive math tutors, programming coaches, and teacher-assistance platforms.
- Dependencies and assumptions: educational use requires explanation faithfulness and pedagogical validity. A latent thought that improves answer accuracy may not produce an accurate or understandable explanation, so additional supervision connecting latent states to student-facing feedback is necessary.
- Formal verification and controllable latent reasoning
- Combine REST with proof assistants, program verifiers, retrieval systems, or symbolic solvers so that causality is evaluated against verifiable intermediate artifacts rather than only natural-language answers.
- This could support mathematically reliable theorem proving, safety-critical code synthesis, and auditable planning.
- Dependencies and assumptions: REST alone does not guarantee correctness, interpretability, or faithfulness. Long-term systems would need explicit verification objectives, adversarial evaluation, formal guarantees where possible, and mechanisms for recovering from incorrect or unstable latent states.
Glossary
- Affine map: A function formed by applying a linear transformation followed by a translation. “where is an attention pooling over the positions of with an affine map from to .”
- Auto-regressive transformer: A transformer model that generates each token conditionally on preceding tokens. “Let denote an auto-regressive transformer model with vocabulary and hidden size .”
- Auxiliary loss: An additional optimization objective used alongside the primary loss. “REST against auxiliary-loss baselines, Light vs Scaled.”
- Causality: The property that a representation preserves information necessary to reproduce a model’s output. “Causality requires that hold the required information about the text generated by a producer agent when given to a consumer agent without altering what a consumer agent would have predicted.”
- Chain-of-Thought (CoT) prompting: Prompting that encourages a LLM to produce intermediate reasoning steps. “Chain-of-Thought (CoT) prompting enables LLMs to solve complex problems by generating each intermediate reasoning step as explicit text.”
- Cosine similarity: A measure of the angular similarity between two vectors. “let denote the cosine similarity.”
- Cross-Entropy (CE): A loss function measuring the difference between a target probability distribution and a model’s predicted distribution. “The main objective is usually Cross-Entropy (CE) of the final decoded answer.”
- Decoder: A component or process that converts an internal representation into a sequence of output tokens. “It decodes through the consumer's vocabulary, scores each candidate refined plan by the decoded tokens it receives, and reports the exponentiated entropy of the resulting posterior over candidates.”
- Differentiable loss: An optimization objective whose value has derivatives with respect to model parameters. “REST translates each of these four properties into a differentiable loss term.”
- Effective superposition: A measure of how many alternative reasoning paths a representation can support simultaneously. “Higher indicates higher effective superposition.”
- Entropy: A measure of uncertainty in a probability distribution. “Estimating the entropy on the other hand can track the property without bias.”
- Forward pass: The computation of a model’s outputs from its inputs. “Given a token sequence with input embeddings , a forward pass yields last-layer hidden states.”
- Gradient: The derivative of a loss with respect to model parameters, used to update those parameters during training. “The property term is averaged over each transfer, therefore receiving a gradient from its own term in addition to the final CE.”
- Hidden state: An internal vector representation produced by a neural network layer. “Within a single model, this is realized by training the model to consume its own hidden states as the next reasoning step instead of a token embedding.”
- Hyperparameter: A training setting selected externally rather than learned directly by the model. “Each term is weighted by a hyperparameter , added on top of the existing CE loss.”
- Inference: The process of using a trained model to generate predictions. “During inference no is available in advance, so the producer starts from its prompt alone and takes latent steps.”
- Latent space: A continuous vector space in which a model represents information internally. “A growing line of work reasons directly in the continuous latent space of LLMs rather than through decoded text.”
- Latent thought: An internal vector sequence used for reasoning instead of explicit text. “A latent thought of a producer to a consumer composes both links as .”
- Minimality: The property that a representation discards information irrelevant to producing its output. “Minimality requires that should remove irrelevant information that was present in the input of the producer agent while maintaining relevant information relative to its output.”
- Monte Carlo sampling: Estimating a distribution or quantity by drawing repeated random samples. “Sampling that distribution would require many outputs at every stage of the system.”
- Next-token distribution: The probability distribution over possible tokens predicted for the next position in a sequence. “A next-token distribution is lower case (, ) while the sequence is the matching upper case (, ).”
- Oracle text: A reference text assumed to provide ideal or fully informative supervision. “Empirically, CE-only does not reach the accuracy of the oracle text from the producer agent.”
- Outer link: A learned transformation that maps one agent’s internal representation into another agent’s input space. “An outer link maps the output of into the input space of another agent.”
- Posterior: A probability distribution representing updated beliefs after incorporating evidence. “It decodes through the consumer's vocabulary, scores each candidate refined plan by the decoded tokens it receives, and reports the exponentiated entropy of the resulting posterior over candidates.”
- Predictive entropy: The uncertainty of a model’s predicted distribution at a particular generation step. “let denote the producer's mean predictive entropy along .”
- Recursion: Repeatedly applying a computational process to its own output or state. “Once the solver finishes, its thought returns to the planner and the chain repeats for a further round.”
- Separability: The property that representations of semantically distinct inputs remain distinguishable. “Separability requires that two thought representations for semantically distinct outputs should be distinguishable or separable to represent that they encode semantically distinct information.”
- Shannon entropy: An information-theoretic measure of uncertainty in a discrete probability distribution. “let denote Shannon entropy.”
- Softmax: A function that converts a vector of scores into a probability distribution. “the next-token distribution over .”
- Stability: The property that a representation captures a distribution of possible outputs rather than only one sampled output. “Stability requires encoding the output distribution rather than one sequence.”
- Teacher forcing: Training a sequence model by providing the ground-truth previous tokens as inputs. “During training, is the producer's ground-truth output, and one teacher-forced pass over the producer's prompt concatenated with yields at 's positions.”
- Temperature: A parameter controlling the concentration or randomness of a probability distribution. “where is a scalar for temperature and denotes the parameters of the attention pooling.”
- Thought collision: The situation in which distinct inputs produce identical or nearly identical internal representations. “If training resulted in two different examples having colliding thoughts, then a consumer agent must answer both with nearly the same distribution even if the two target texts are different.”
- Token embedding: A vector representation assigned to an individual token. “Within a single model, this is realized by training the model to consume its own hidden states as the next reasoning step instead of a token embedding.”
- Transfer: The passage of a textual or latent representation from one agent or processing stage to another. “Every handoff of this kind is one transfer of \cref{def:transfer}.”
- Teacher-forced target: The known target sequence supplied to a model during supervised sequence generation. “where is the teacher-forced target answer for the consumer.”
- Underconstrained representation: A representation for which the training objective fails to impose sufficient structural requirements. “Therefore, only minimizing CE would make the thought representation underconstrained.”






