Papers
Topics
Authors
Recent
Search
2000 character limit reached

World Model Science: Self-Organized Criticality, Weak Chaos, and Metastable Belief Dynamics in Long-Horizon LLM Agents

Published 12 Jul 2026 in cs.AI | (2609.17419v1)

Abstract: Long-horizon LLM agents must maintain task state across extended sequences of observations, actions, tool calls, and intermediate beliefs. We study these trajectories through three dynamical views: self-organized criticality, weak chaos, and metastable belief dynamics. Our framework aligns agent-implied states with benchmark-grounded states and measures stress accumulation, error avalanches, temporal dependence, local--global mismatch, bounded divergence, belief-basin transitions, and finite-size scaling under explicit null models. Across 22 experiments spanning controlled puzzles, tool use, embodied tasks, multi-hop retrieval, general-assistant reasoning, and Game of Life, we find that locally valid actions can persist after global state fidelity fails, stress can trigger abrupt collapse, error sequences exhibit long memory, dependency depth changes the propagation regime, and larger horizons support larger avalanches. At the same time, divergence remains bounded, belief states show metastable rather than fully chaotic behavior, and stronger claims of universal power laws, critical points, or shared intervention optima are not supported. These results suggest a science of agent world models based on trajectory-level dynamical diagnostics rather than terminal reward alone.

Authors (2)

Summary

  • The paper suggests analyzing the trajectories of long-horizon LLM agents through a finite dynamical system framework that detects stress-sensitive collapses and structural conditionings without imposing classical SOC requirements.
  • The analysis assumes that each LLM agent experiences propagating stress even though in this distributed context it does not claim SOC phenomena anything more than the claim of the criticality dynamics in the given settings.
  • Flow chart experiments show that these stress-sensitive behavior signifiers in LLM agents perform weaker than the classical SOC.

Research objective and conceptual framework

The paper proposes a trajectory-level framework for analyzing long-horizon LLM agents as finite dynamical systems. Its central claim is deliberately narrower than the physical claim suggested by self-organized criticality (SOC): agent trajectories can exhibit measurable signatures of correlated, stress-sensitive, structurally propagated collapse without thereby establishing physical SOC, a universal critical point, or universal power-law behavior. The framework is intended to distinguish these signatures from independent per-step errors and from ordinary persistence caused by Markovian dependence (2609.17419).

The proposed analogy treats unresolved uncertainty, contradictions, retrieval conflicts, unverified assumptions, and tool-error debt as accumulated stress. A threshold-crossing event corresponds to degradation in the agent’s benchmark-grounded world state, while an avalanche is a temporally clustered sequence of state errors. This framing extends established diagnostics from SOC and crackling-noise systems—event-size distributions, temporal dependence, finite-size scaling, and invariance under surface perturbations—without assuming that LLM agents instantiate the underlying physical mechanisms.

Figure 1

Figure 1: SOC-inspired diagnostic framing in which stress accumulation, threshold crossing, and cascade release are used as diagnostic signatures of correlated agent collapse.

The measured object is not a hidden activation trajectory. Instead, the authors define a substrate-specific state extractor over logged interactions. The resulting state vector contains progress, belief, constraints, uncertainty, risk or tool debt, memory or retrieval context, and plan intention. A corresponding gold-state extractor is obtained from simulator state, database state, supporting facts, audit logs, or benchmark annotations. World-state fidelity is then compared with local action validity. The distinction is essential: an action may be syntactically valid, immediately admissible, and environmentally executable even when the agent’s latent task state has already diverged.

The paper formalizes this distinction through the local-global gap, defined as local validity minus global world-state fidelity. A proposition establishes that local-validity traces cannot, in general, identify global fidelity: two histories can generate identical sequences of locally admissible actions while differing in hidden constraints or task state. This is not merely a measurement preference but an identifiability limitation. Evaluations based only on action validity therefore cannot recover the state trajectory relevant to long-horizon correctness.

Measurement protocol and diagnostic criteria

The empirical program comprises 22 experiments over StatefulPuzzle-SOC, τ\tau-bench Retail and Airline, GAIA Level 1, ALFWorld, HotpotQA-RAG, and Game of Life. The experiments use a fixed model interface, temperature zero, and fixed seeds. The authors freeze state extractors, coordinate distances, fidelity weights, and stress weights before outcome analysis. This precommitment is intended to constrain researcher degrees of freedom, although the validity of each extractor remains substrate-dependent.

The diagnostic pipeline maps interaction traces into state estimates, stress components, thresholded error events, and statistical summaries. It compares observed trajectories with several null models: independent Bernoulli errors, Markov persistence, shuffled spectra, task-difficulty predictors, and surface-perturbation controls.

Figure 2

Figure 2: Measurement pipeline separating local action validity from global world-state fidelity and comparing observed collapse with independent-error and persistence nulls.

The paper defines finite world-model SOC through four jointly required properties:

  1. Temporal dependence: thresholded state errors are correlated beyond an independent-error null.
  2. Stress response: collapse probability or magnitude increases with accumulated or externally injected stress.
  3. Finite-size scaling: cascade scale depends on a finite system-size variable such as trajectory horizon.
  4. Structural conditioning: propagation depends on dependency depth or task-graph topology.

This definition intentionally excludes stronger requirements often associated with SOC. A universal exponent, exact critical point, asymptotic scale-free behavior, and shared intervention optimum are not necessary. The resulting construct is therefore an operational finite-system diagnostic rather than a claim that LLM agents reproduce classical SOC mechanisms.

Figure 3

Figure 3: Cross-substrate evidence matrix distinguishing supporting, partial, boundary, alternative, capacity-limited, and untested diagnostic results.

Stress-sensitive collapse and the local-global mismatch

The strongest causal result comes from the controlled StatefulPuzzle-SOC intervention. Exogenously injected stress is varied before outcome measurement, using horizon 64 and dependency depth one. Stress predicts collapse with an AUROC of 0.979, and the first nonzero stress condition moves the model from a measurable zero-stress stability floor to near-deterministic collapse. Because stress is manipulated rather than inferred from the same error signal used to define collapse, this result provides more than a correlation between internal difficulty and failure.

Figure 4

Figure 4: Controlled stress causes a sharp transition from mostly stable behavior to near-deterministic collapse over horizon 64.

The implication is that small perturbations can have strongly nonlinear consequences when the agent is operating near a state-dependent failure boundary. However, the result is established in a controlled synthetic substrate; it does not by itself demonstrate that naturally occurring stress in benchmark environments has comparable causal leverage.

Benchmark traces provide a complementary result. In τ\tau-bench Airline, information-only tasks show little local-global mismatch, whereas more complex task classes produce gaps between approximately 0.62 and 0.68. GAIA intermediate-conclusion steps produce the largest reported gap, ΔLG=0.857\Delta^{LG}=0.857. These actions remain locally coherent even after the evidence state has collapsed.

Figure 5

Figure 5: Local action validity can remain high after global evidence or task-state fidelity has substantially degraded.

This result directly challenges evaluations that use executable tool calls or immediate admissibility as proxies for reliable reasoning. A locally valid action is not evidence that the agent’s belief state, constraints, or evidence set remain correct. The measurement does, however, depend on the fidelity and granularity of the benchmark-specific gold-state extractor.

The surface-invariance experiments reinforce a structural interpretation. Paraphrases, renamings, distractors, order changes, and stylistic transformations preserve macro avalanche statistics when they preserve the underlying task graph. In StatefulPuzzle, six surface variants maintain comparable avalanche size, collapse rate, and collapse timing; reported KS distances remain below the corresponding same-distribution null threshold. Similar qualitative stability appears in HotpotQA and Retail.

Figure 6

Figure 6: Macro collapse statistics remain comparatively stable under prompt-surface changes that preserve task structure.

The result suggests that the measured dynamics are not reducible to superficial wording. It does not establish universality: the relevant invariance is conditional on preserving the task graph, and the paper later finds that intervention regimes and geometric signatures vary substantially across substrates.

Memory, bounded divergence, and metastability

Temporal analyses reject a white-noise interpretation of several error streams. StatefulPuzzle remains in a long-memory regime across tested horizons, with spectral error exponents in the range αe∈[1.38,1.52]\alpha_e \in [1.38,1.52] and valid-horizon DFA estimates of approximately 1.34–1.38. The authors explicitly discount short-series DFA artifacts by treating DFA as reliable only at sufficiently long horizons and using spectral estimates as the primary cross-horizon diagnostic.

HotpotQA supplies a mechanism-sensitive comparison. Full-context retrieval nearly decorrelates the error stream, whereas narrow top-two retrieval shifts errors toward flicker-like persistence. Thus, temporal dependence is not presented as an invariant property of the model alone; it changes with the information channel and retrieval regime.

Figure 7

Figure 7: Error persistence depends on horizon and information access, with narrow retrieval producing stronger temporal dependence in HotpotQA.

The dependency-depth experiment is framed as weak chaos rather than unbounded chaos. With a fixed initial discrepancy, exponential fits are preferred at depth one. At depth two, the mean AIC difference changes sign, and the fraction of paired trajectories favoring power-law fits increases from 0.133 at depth one to 0.667–0.700 at depths two through eight. Mean divergence rises from 0.120 at depth one to approximately 1.88–1.95 at larger depths, while saturation occurs within the finite state space.

Figure 8

Figure 8: Recursive dependency depth changes the shape of bounded divergence, with an AIC crossover at depth two rather than unbounded chaotic growth.

The authors correctly avoid interpreting these fits as estimates of a positive Lyapunov exponent. The power-law preference describes the shape of finite-system propagation before saturation. The result supports a depth-dependent propagation transition, but it does not establish deterministic chaos in the dynamical-systems sense.

The metastability analysis further limits the interpretation. ALFWorld agents escape wrong belief basins rapidly, but they also under-exploit success-associated basins. The measured escape probability from wrong basins is 0.97, compared with 0.82 from success-associated basins. This pattern is inconsistent with a simple account in which agents become rigidly trapped in incorrect metastable states. It is better characterized as shallow basin structure with excessive exploration or insufficient exploitation. Consequently, the paper’s title-level reference to metastable belief dynamics denotes partial, substrate-specific evidence rather than robust wrong-basin lock-in.

Geometry, finite-size effects, and capability boundaries

The geometry analyses test whether error clusters reflect the topology of evidence or task dependencies. HotpotQA, with fixed two-hop evidence topology, produces relatively stable fractal dimensions between 0.835 and 0.904, with a spread of 0.069. ALFWorld exhibits considerably larger topology-conditioned variation: reported values are 0.639 for long-chain tasks, 1.042 for container tasks, and 1.50 for multi-room tasks.

Figure 9

Figure 9: Error geometry is stable under fixed evidence topology but varies with embodied task-graph structure.

The implication is that fractal dimension should not be treated as a universal scalar characteristic of an agent. It is an estimand conditioned on the substrate’s graph structure, state representation, and error-cluster construction. The variation across ALFWorld task classes supports structural conditioning, but it also makes cross-benchmark numerical comparisons difficult without stronger normalization.

Finite-size scaling is most clearly demonstrated by varying trajectory horizon in StatefulPuzzle. The maximum avalanche size increases monotonically from 7 to 490, with every adjacent-horizon comparison remaining significant after false-discovery correction; the reported corrected tests are below 10−810^{-8}. This establishes a horizon-dependent cutoff: longer trajectories permit larger cascades, while the finite task constrains the maximum event size.

The Game-of-Life experiment demonstrates why capacity controls are indispensable. For grid sizes from 4 through 64, the measured fractal dimension increases from 1.09 to 1.49. At grid sizes of at least 96, however, there are no valid generated trajectories. The absence of a signature at those sizes cannot be interpreted as evidence against critical-like dynamics, because the model has left the measurable operating regime.

Figure 10

Figure 10: Avalanche size grows with horizon in StatefulPuzzle, while Game of Life exposes a separate capacity boundary.

Figure 11

Figure 11: A missing diagnostic signature is interpretable only when the model can generate valid trajectories on the evaluated substrate.

This capability-matched-support condition is one of the paper’s most important methodological points. A benchmark can fail to reveal a dynamical signature either because the signature is absent or because the model cannot produce valid trajectories. These cases are observationally distinct only if validity is measured independently.

Limits on stronger SOC interpretations

The paper reports several results that qualify rather than strengthen the universal interpretation. In Retail, natural early stress predicts an independent reward-error label with only modest performance: AUROC 0.616. After accounting for task difficulty, the incremental cross-validated contribution of stress is small.

Figure 12

Figure 12: Natural stress is modestly predictive of independent Retail reward errors and adds little beyond task difficulty.

This result separates controlled stress sensitivity from observational precursor value. The controlled intervention demonstrates causal responsiveness in StatefulPuzzle; the Retail analysis does not show that naturally extracted stress is a strong or generally useful early-warning signal.

Prompt-regime clustering provides similarly limited evidence for discrete universality classes. Macro statistics remain stable across variants, with KS distances of approximately 0.03–0.17, but six interpretable clusters are not the statistically preferred partition. Silhouette analysis favors two clusters, with a reported silhouette score of 0.321, and bootstrap ARI indicates continuous rather than sharply separated regimes.

Figure 13

Figure 13: Stable macro statistics coexist with continuous prompt-regime structure rather than sharply discrete universality classes.

Intervention optima are also substrate-specific. Retail performs best under no intervention or high verification, with a reported best score of 0.50, whereas ALFWorld favors exploration-heavy or memory-heavy regimes, with a best score of 0.433.

Figure 14

Figure 14: Verification and exploration have different operating trade-offs in Retail and ALFWorld.

This contradicts the idea of a single near-critical intervention balance transferable across environments. The measured control policy must be conditioned on the substrate’s error ecology, task topology, and state observability.

Several open questions remain within the paper’s own scope. The validity of the conclusions depends on substrate-specific state maps and distance functions, whose construction may affect fidelity, stress, and avalanche estimates. Persistence models explain a substantial portion of Retail burstiness, so the incremental evidence for critical organization beyond correlated dependence is limited. Pure power laws are rejected in favor of truncated or lognormal tails, and no unique critical point is identified. Finally, the experiments use one model interface, fixed decoding settings, and a finite set of benchmarks; whether the same diagnostics preserve their interpretation across model families and independently designed state extractors remains unresolved.

Conclusion

The paper develops a finite, measurement-oriented account of long-horizon agent collapse. Its most consistent evidence concerns stress-sensitive failure, local-global state divergence, long-memory errors, dependency-conditioned propagation, topology-dependent error geometry, and horizon-dependent avalanche cutoffs. The strongest numerical findings are the controlled-stress AUROC of 0.979, the GAIA local-global gap of 0.857, the depth-two divergence transition, and avalanche growth from 7 to 490 with horizon.

The results support the use of SOC-inspired diagnostics for world-model reliability, but not the stronger claim that LLM agents possess universal critical dynamics. Natural stress has modest predictive value, persistence explains part of the observed clustering, belief basins are shallow rather than rigidly trapping, regime structure is continuous, and intervention optima are substrate-specific. The paper’s principal contribution is therefore methodological: long-horizon agent evaluation should measure intermediate world-state fidelity and trajectory dynamics rather than relying solely on terminal reward or stepwise action validity (2609.17419).

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

Explain it Like I'm 14

1. What is this paper about?

This paper studies how AI agents powered by LLMs, behave during long tasks.

An LLM agent might need to:

  • remember information,
  • use tools,
  • follow rules,
  • search for evidence,
  • make plans,
  • and complete many steps in a row.

The researchers ask whether small mistakes can build up and suddenly cause a much larger failure. They compare this behavior to systems in nature, such as a sandpile. A sandpile may look stable while grains slowly pile up, but one small extra grain can cause a large avalanche.

The paper calls this idea self-organized criticality, or SOC. The authors do not claim that AI systems are exactly like sandpiles. Instead, they use similar measurements to study whether AI mistakes are connected, clustered, and affected by the length and structure of a task.

2. What questions did the researchers ask?

The paper focuses on several main questions:

  1. Do small problems build up before an AI agent suddenly fails? For example, can uncertainty, contradictions, or unverified assumptions accumulate until the agent collapses?
  2. Can an agent take locally correct actions while its overall understanding is already wrong? An action may be allowed at one moment but still be based on a mistaken picture of the whole task.
  3. Are mistakes independent, or do they come in connected groups? The researchers compare isolated errors with “avalanches,” meaning bursts of related errors.
  4. Does the structure of a task affect how mistakes spread? A task with many connected steps may allow one mistake to affect many later decisions.
  5. Does a longer task allow larger failures?
  6. Do these behaviors look like complete chaos? The researchers examine whether errors grow without limits or remain bounded by the task and the agent’s abilities.

3. How did they study the problem?

Tracking the agent’s hidden task state

The researchers examined the logs of an agent’s actions, observations, and tool results. They then estimated the agent’s current “world state,” including:

  • how much progress it had made,
  • what it believed,
  • which rules or constraints it needed to follow,
  • how uncertain it was,
  • how much risk or tool-error debt had accumulated,
  • what information it remembered,
  • and what it planned to do next.

They also created a gold state: the state that was actually correct according to the environment, database, or task record.

This allowed them to compare:

  • Global fidelity: Is the agent’s overall understanding correct?
  • Local validity: Is its current action allowed and immediately sensible?

An everyday analogy is a student solving a complicated puzzle. The student may write a legal next move, but if they misunderstood an earlier clue, their overall solution may already be going wrong.

Measuring stress

The researchers created a stress score based on things such as:

  • unresolved questions,
  • contradictions,
  • conflicting pieces of information,
  • assumptions that had not been checked,
  • and previous tool errors.

They then tested whether higher stress was connected to later collapse.

Measuring “avalanches”

An avalanche was a period when the agent’s understanding was much more wrong than usual. The researchers measured:

  • how many steps the avalanche lasted,
  • how large it was,
  • and how much total error it contained.

Comparing against simpler explanations

The researchers used several null models. A null model is a simple comparison system that shows what would happen if the interesting effect were not present.

For example, they compared the agent’s errors with:

  • independent random errors,
  • errors that only depend on the immediately previous step,
  • and errors explained by task difficulty.

This helped them determine whether the agent’s failures were truly clustered or merely the result of making mistakes at a steady rate.

Testing many tasks

The study used 22 experiments across different environments, including:

  • tool-use tasks in retail and airline settings,
  • general assistant questions,
  • embodied tasks where the agent navigates rooms and objects,
  • multi-step information retrieval,
  • controlled puzzles,
  • and the computer simulation Game of Life.

The researchers also changed task length, task dependency depth, prompts, and the amount of available information.

4. What did they find?

Stress can lead to sudden collapse

In the controlled puzzle experiments, adding even a small amount of artificial stress made the agent much more likely to fail. The stress measure predicted collapse very well, with an AUROC of 0.979.

In simple terms, the agent could appear to work normally until hidden problems pushed it close to a breaking point. Then a small additional difficulty could cause a sharp failure.

Local actions can remain correct while the bigger picture is wrong

One of the most important findings is that an action can be locally valid even when the agent’s overall understanding has already become incorrect.

For example, an agent might:

  • make a grammatically correct tool call,
  • follow the immediate rules,
  • or produce a reasonable-looking sentence,

while using the wrong facts or an outdated plan.

In one GAIA assistant task, the difference between local action validity and global state correctness reached 0.857, showing a large mismatch.

This means that checking only whether each action is executable is not enough. Evaluators also need to check whether the agent still understands the entire task correctly.

Errors have memory

The errors were not always random and independent. In several experiments, mistakes tended to remain connected over time.

This is similar to a student making one incorrect assumption and then using it again and again. The first mistake creates conditions for later mistakes.

The amount of memory depended on the information available. For example, narrow retrieval in HotpotQA caused more persistent errors, while giving the agent fuller context reduced this effect.

Task structure affects how mistakes spread

When tasks had deeper chains of dependencies, an early mistake could spread in a different way. The researchers found a change in error behavior around dependency depth two.

However, this did not prove that the systems were completely chaotic. The divergence stayed bounded because the tasks and state spaces were finite. The errors could grow, but they did not increase forever.

Longer tasks allow larger avalanches

In the controlled puzzle environment, the largest avalanche increased as the allowed task length increased. The maximum avalanche cutoff grew from 7 steps in shorter settings to 490 in longer ones.

This suggests that long tasks give mistakes more opportunities to interact and spread.

Error patterns depend on task geometry

The shape of error clusters depended on the structure of the task.

For example:

  • information-retrieval tasks had relatively stable error patterns when their evidence structure stayed the same;
  • embodied navigation tasks showed different patterns for long chains, containers, and multiple rooms.

This suggests that there is no single error pattern shared by every kind of AI task.

Prompt wording mattered less than task structure

Changing the wording, names, order, or style of prompts often did not greatly change the overall collapse pattern when the underlying task remained the same.

This suggests that the task’s structure may matter more than its surface wording.

The researchers did not find proof of universal laws

The authors are careful about what their results do not show. They did not find strong evidence for:

  • one universal power-law pattern,
  • a single critical point shared by all AI tasks,
  • one universal class of agent behavior,
  • or one intervention strategy that works best everywhere.

For instance, high verification helped in some retail tasks, while more exploration and memory helped in some embodied tasks.

5. Why are these findings important?

Most AI evaluations focus on the final answer: did the agent succeed or fail?

This paper argues that this is not enough. An agent may already be losing track of the task long before its final answer becomes obviously wrong.

The research suggests that AI evaluations should also record:

  • what the agent believes at each step,
  • whether its beliefs match the real task state,
  • how uncertainty and contradictions build up,
  • whether errors occur in bursts,
  • and how task length and structure affect failure.

This could help developers detect problems earlier and design better safeguards. For example, an agent might be required to:

  • check its beliefs after important actions,
  • verify facts before using them repeatedly,
  • stop when contradictions become too large,
  • or use different strategies for different types of tasks.

Conclusion

The paper presents a new way to study failures in long-running AI agents. Its main idea is that AI mistakes may behave less like separate random accidents and more like connected chains or avalanches.

The strongest evidence shows that:

  • hidden stress can build up,
  • small problems can trigger sudden collapse,
  • locally valid actions can hide a globally incorrect understanding,
  • mistakes can persist over time,
  • task structure affects how errors spread,
  • and longer tasks can produce larger failures.

At the same time, the paper does not claim that all AI systems follow one universal law or that they are literally chaotic or critical physical systems. Instead, it offers a useful set of tools for examining how an AI agent’s understanding changes throughout a task, rather than judging it only by its final answer.

Knowledge Gaps

Conocimiento faltante, limitaciones y preguntas abiertas

  • Validez de los extractores de estado: No se informa con suficiente detalle cómo se construyen ni validan los extractores ϕs\phi_s y ϕs⋆\phi_s^\star, sus métricas ds,kd_{s,k} y sus pesos; futuros trabajos deberían medir su fiabilidad interevaluador, sensibilidad a errores de anotación y validez frente a evaluaciones humanas independientes.
  • Dependencia de decisiones de medición: Los resultados pueden variar sustancialmente con los pesos de fidelidad, pesos de estrés, umbral τe\tau_e y definición de avalancha; falta un análisis sistemático de sensibilidad y robustez frente a distintas parametrizaciones.
  • Coordenadas no observables del estado: Cuando una dimensión del estado no está disponible se fija su distancia a cero, lo que puede ocultar errores reales; se necesita cuantificar cuánto cambian las conclusiones al imputar, eliminar o estimar explícitamente esas dimensiones.
  • Muestra y potencia estadística: El texto no reporta de forma completa el número de trayectorias por condición, intervalos de confianza, tamaños de efecto ni análisis de potencia; esto dificulta evaluar la estabilidad de los resultados, especialmente en experimentos de colas, DFA, fractalidad y agrupamiento.
  • Generalización entre modelos: Los experimentos usan una única configuración de interfaz, temperatura y semilla, pero no establecen si los patrones se mantienen entre familias de LLM, escalas de modelo, proveedores, ventanas de contexto o políticas de decodificación.
  • Efecto de la aleatoriedad de generación: La temperatura cero y la semilla fija impiden estudiar cómo la variabilidad estocástica afecta la formación de avalanchas, la memoria temporal y la divergencia entre trayectorias.
  • Separación entre capacidad del modelo y dinámica del entorno: Los resultados mezclan propiedades del agente con restricciones de los benchmarks, formatos de salida, herramientas, límites de contexto y reglas del entorno; faltan diseños factoriales que separen sistemáticamente estos factores.
  • Causalidad del estrés: El experimento de StatefulPuzzle muestra que el estrés inyectado predice colapso, pero no determina qué componentes —incertidumbre, contradicciones, conflictos de recuperación, supuestos no verificados o deuda de herramientas— son causalmente responsables ni cómo interactúan.
  • Validez ecológica del estrés controlado: No se demuestra que el estrés inyectado reproduzca la distribución, composición o evolución del estrés que aparece en tareas reales; sería necesario comparar perturbaciones sintéticas con episodios naturales de error.
  • Precursores naturales débiles: En Retail, el estrés aporta una mejora predictiva modesta sobre dificultad, pero no se identifica qué señales adicionales permitirían anticipar el colapso ni si la predicción se mantiene fuera de muestra y en otros dominios.
  • Ambigüedad entre persistencia y avalanchas: Los modelos de persistencia explican parte de la estructura de los tamaños de avalancha; queda sin resolver qué mecanismo adicional distingue una dinámica de avalanchas genuina de errores autocorrelacionados, dificultad serial o dependencia en las etiquetas.
  • No identificación de un mecanismo generativo: Las asociaciones entre estrés, memoria, topología y colapso no muestran cómo se generan internamente las transiciones; faltan modelos causales o mecanísticos que conecten memoria, planificación, recuperación y actualización de creencias con los errores observados.
  • Estatus de la analogía con SOC: La definición de “SOC finito” es operacional y más amplia que la noción física clásica; no se establece qué predicciones adicionales diferencian este marco de modelos alternativos de acumulación de errores, procesos de renovación, modelos ocultos de Markov o fallos por saturación de contexto.
  • Ausencia de un punto crítico identificable: El estudio no determina si existe un parámetro de control cuyo ajuste produzca una transición crítica reproducible, ni si las señales observadas se maximizan cerca de un régimen crítico en vez de reflejar simplemente degradación progresiva.
  • Escalamiento finito limitado al horizonte: El crecimiento del tamaño máximo de avalancha con el horizonte no prueba una ley de escalamiento universal; faltan colapsos de tamaño finito, exponentes de escalamiento, múltiples variables de tamaño y pruebas en horizontes y tareas independientes.
  • Sensibilidad al umbral de error: Las avalanchas dependen de τe\tau_e; no se muestra si los resultados de agrupamiento, duración y escalamiento sobreviven a umbrales absolutos, relativos, adaptativos o basados en cuantiles.
  • Interpretación de la memoria temporal: Los espectros y el DFA pueden verse afectados por series cortas, no estacionariedad, truncamiento por finalización de tareas y mezcla de trayectorias con distinta duración; falta una evaluación más amplia con métodos robustos para procesos no estacionarios y datos faltantes.
  • Dirección causal entre memoria y recuperación: El resultado de HotpotQA sugiere que el ancho de recuperación modifica la dependencia temporal, pero no aclara si la recuperación estrecha causa memoria de errores, si ambos dependen de dificultad, o si el agente cambia su estrategia en respuesta a señales no observadas.
  • Robustez de la dimensión fractal: Las estimaciones de DfD_f se presentan para pocos tipos de grafos y pueden depender de la representación espacial, resolución, tamaño de muestra y algoritmo de box-counting; falta comparar métodos geométricos y validar la interpretabilidad causal de esta medida.
  • Estructura de dependencia insuficientemente aislada: El cambio asociado con la profundidad de dependencia podría confundirse con longitud de trayectoria, dificultad, número de restricciones o carga de memoria; se requieren intervenciones que varíen cada propiedad manteniendo las demás constantes.
  • No equivalencia entre grafos benchmark y procesos cognitivos: Las topologías de evidencia o tareas usadas pueden ser representaciones diseñadas por los investigadores y no necesariamente corresponder a la estructura que el agente utiliza internamente; falta medir o inferir el grafo efectivo de dependencias del agente.
  • Limitaciones de la comparación local-global: La brecha ΔLG\Delta^{LG} depende de cómo se puntúan la validez local y la fidelidad global; no se establece si esta métrica predice de forma consistente fallos posteriores, recuperación espontánea o daño irreversible.
  • Predicción de recuperabilidad: El estudio identifica divergencia de estado, pero no analiza qué errores pueden corregirse, cuánto cuesta la recuperación ni qué señales tempranas distinguen una desviación reversible de un colapso terminal.
  • Intervenciones no evaluadas de forma causal: Las comparaciones entre verificación, exploración y memoria sugieren políticas específicas por sustrato, pero no aíslan los mecanismos de cada intervención ni miden sus costes computacionales, latencia, uso de herramientas o efectos sobre la utilidad.
  • Falta de optimización adaptativa: No se estudia si una política que ajusta dinámicamente verificación, exploración o recuperación según el estrés y la fidelidad supera a las políticas estáticas evaluadas.
  • Generalización de la invariancia de superficie: La estabilidad ante paráfrasis, renombrado y cambios de estilo se prueba en un conjunto limitado de variantes; queda abierta su robustez ante cambios lingüísticos, traducción, formatos, instrucciones adversariales o transformaciones que preserven semántica pero alteren la distribución de tokens.
  • Regímenes macro continuos: El agrupamiento no encuentra clases universales nítidas, pero no se determina si los regímenes continuos dependen de la representación elegida, del algoritmo de clustering o de variables latentes no medidas.
  • Cobertura limitada de tareas y modelos de interacción: Aunque se incluyen varios benchmarks, faltan tareas de software, navegación web, interacción multiagente, negociación y entornos con cambios dinámicos para establecer si los hallazgos se aplican más allá de los siete sustratos considerados.
  • Validez del control Game of Life: El control natural está limitado por la capacidad del modelo para generar estados válidos; no se separa completamente la dinámica del sistema celular de los errores de serialización, seguimiento de formato o incapacidad de planificación del agente.
  • Capacidad como confusor no cuantificado: Registrar una frontera de capacidad no basta para determinar cómo la capacidad afecta las estimaciones; sería necesario variar sistemáticamente el tamaño del modelo, contexto, herramientas y representación antes de comparar firmas dinámicas.
  • Reproducibilidad incompleta de los resultados: Aunque se proporciona un repositorio, el texto no especifica de manera suficiente las versiones de modelos, prompts completos, filtros de datos, criterios de exclusión, seeds, procedimientos de anotación y configuraciones exactas para reproducir cada una de las 22 pruebas.
  • Riesgo de dependencia entre pruebas: Las 22 pruebas pueden reutilizar trayectorias, variantes o métricas relacionadas; falta aclarar la unidad estadística efectiva y controlar la dependencia entre análisis más allá de la corrección de Benjamini–Hochberg.
  • Relación con el rendimiento final: El trabajo muestra que la fidelidad intermedia aporta información adicional, pero no cuantifica cuánto mejora la predicción de recompensa, seguridad, coste o éxito final frente a métricas convencionales.
  • Aplicación a seguridad y operación real: No se evalúa si los diagnósticos permiten activar salvaguardas antes de una acción peligrosa, reducir daños en herramientas reales o mejorar la monitorización de agentes desplegados.
  • Interpretación de las “creencias” del agente: El estado medido es una reconstrucción externa basada en trazas, no necesariamente la creencia interna del LLM; queda abierta la relación entre la fidelidad observada, activaciones internas, memoria de contexto y representaciones latentes.
  • Persistencia temporal de las conclusiones: No se sabe si las firmas cambian después de fine-tuning, entrenamiento con feedback, uso de memoria externa, reflexión, recuperación aumentada o aprendizaje durante la interacción.
  • Comparación con baselines de agentes más simples: Faltan controles con políticas heurísticas, agentes sin memoria, agentes con memoria perfecta, planificadores simbólicos y modelos de error explícitos para determinar qué firmas son específicas de LLM y cuáles emergen en cualquier sistema secuencial.

Practical Applications

Immediate Applications

The paper’s most deployable contribution is a trajectory-level monitoring framework that separates local action validity from global world-state fidelity. These applications can be implemented with existing agent logs, benchmark state extractors, and the released code, provided the relevant state variables are observable.

  • Long-horizon LLM agent observability and reliability dashboards — Software, enterprise automation
    • Instrument agents to log, at every step, progress, beliefs, constraints, uncertainty, memory context, plan state, tool debt, and action validity.
    • Track the local–global gap, ΔLG=L−F\Delta^{LG}=L-F, to identify cases where an action is syntactically valid or executable but inconsistent with the actual task state.
    • A monitoring product could display:
    • unresolved uncertainties and contradictions;
    • retrieval conflicts and unverified assumptions;
    • tool-call failures and accumulated tool debt;
    • error-episode size and duration;
    • horizon-dependent collapse risk.
    • Actionability: Existing agent traces and tool logs are sufficient for a first implementation.
    • Dependencies: Requires a reliable, substrate-specific gold-state extractor and frozen evaluation weights. If global state cannot be reconstructed, the system can monitor stress proxies but cannot claim world-state fidelity.
  • Runtime detection of silent agent-state drift — Software, customer service, workflow automation
    • Add a “state-fidelity gate” before consequential actions such as account changes, bookings, refunds, database mutations, or sending external communications.
    • Pause or escalate an agent when local validity remains high but global fidelity falls, as in the paper’s GAIA and airline examples.
    • A practical workflow is:
    • 1. estimate current state fidelity;
    • 2. compare it with the latest trusted checkpoint;
    • 3. require re-verification when the discrepancy exceeds a threshold;
    • 4. resume only after retrieving evidence or confirming constraints.
    • Dependencies: Thresholds must be calibrated to the application. The paper shows that a single universal intervention threshold is unlikely to work across substrates.
  • Pre-action verification for high-impact tool calls — Finance, healthcare administration, travel, e-commerce
    • Use the stress variables UU, KK, RR, VV, and BB as a checklist before mutating actions:
    • unresolved uncertainty;
    • contradictions;
    • retrieval disagreement;
    • unverified assumptions;
    • accumulated tool errors.
    • Require confirmation, independent retrieval, or human approval when stress is elevated.
    • Potential products include an agent middleware layer that classifies actions as read-only, reversible, or mutating and applies progressively stronger verification.
    • Dependencies: The method is most suitable where constraints and environment state are auditable. It should not be treated as a guarantee of correctness; the paper finds that natural stress is only a modest predictor of failure in Retail.
  • Process-level evaluation for LLM agents — Academia, model development, benchmarking
    • Extend evaluations beyond terminal reward, exact-match answers, and tool-call syntax.
    • Report:
    • world-state fidelity over time;
    • local–global mismatch;
    • avalanche size and duration;
    • temporal error dependence;
    • dependency-depth effects;
    • horizon-dependent error cutoffs.
    • Compare observed traces against independent Bernoulli, Markov-persistence, shuffled-spectrum, and task-difficulty null models.
    • This can reveal whether two models with similar success rates differ in their failure dynamics and recoverability.
    • Dependencies: Benchmarks need intermediate annotations, simulator state, supporting evidence, or audit logs. Results are only meaningful on capability-matched trajectories.
  • Improved debugging and incident analysis for agent systems — Software engineering and MLOps
    • Use error avalanches to distinguish isolated mistakes from cascades caused by an earlier state-tracking failure.
    • Error geometry and dependency-depth analysis can identify whether failures originate in:
    • retrieval topology;
    • recursive planning;
    • tool chains;
    • embodied navigation;
    • long memory contexts.
    • This supports targeted remediation, such as improving retrieval width for multi-hop QA or inserting checkpoints at dependency boundaries.
    • Dependencies: The task graph must be explicitly represented or reconstructed. Fractal or geometric statistics should be used comparatively within a task family, not as universal measures.
  • Retrieval and memory-channel tuning — Search, knowledge management, RAG systems
    • Use the paper’s finding that narrow top-2 retrieval produces more persistent error behavior than fuller context to evaluate retrieval policies.
    • Compare candidate retrieval workflows by measuring whether errors are:
    • isolated or clustered;
    • persistent across steps;
    • concentrated around evidence conflicts;
    • amplified by missing context.
    • A practical RAG workflow could dynamically widen retrieval or trigger evidence reconciliation when temporal dependence increases.
    • Dependencies: More context may increase cost, latency, or distraction. The paper does not establish that wider retrieval is always optimal; the effect is task- and substrate-dependent.
  • Human escalation policies for agentic customer and operational workflows — Industry and public services
    • Replace simple “tool call failed” rules with escalation criteria based on trajectory state:
    • high stress;
    • repeated contradictions;
    • increasing avalanche duration;
    • large local–global gap;
    • repeated recovery failure.
    • This is especially relevant for airline rebooking, retail support, insurance intake, scheduling, and administrative assistants.
    • Dependencies: Escalation policies require calibrated false-positive and false-negative costs. Observed stress can correlate with task difficulty, so it should not be used as the sole escalation signal.
  • Educational tools for teaching reliable reasoning and debugging — Education
    • Build learner-facing or developer-facing tools that visualize how small early discrepancies propagate through a multi-step task.
    • Students can compare:
    • locally valid versus globally correct steps;
    • independent errors versus temporally clustered errors;
    • effects of additional verification or wider evidence retrieval;
    • consequences of increasing dependency depth.
    • Dependencies: Educational versions require interpretable state representations and should avoid presenting the SOC analogy as evidence that LLMs possess physical criticality.
  • Policy and audit standards for high-risk AI agents — Regulation and governance
    • Establish minimum logging requirements for deployed agents, including intermediate state, tool calls, evidence provenance, constraint checks, and recovery events.
    • Require evaluations across multiple horizons, task topologies, and dependency depths rather than only aggregate success rates.
    • Use the local–global gap as an audit criterion for systems that can make consequential changes.
    • Dependencies: Policies must define sector-specific observables and privacy-preserving logging. The paper does not provide evidence for universal safety thresholds or a universal “critical point.”

Long-Term Applications

The following applications are plausible extensions of the findings but require larger datasets, better state extraction, causal validation, or integration with adaptive agent architectures.

  • Adaptive runtime control of agent verification and exploration — Software, robotics, autonomous systems
    • Develop controllers that adjust verification, exploration, memory retrieval, and replanning based on measured stress and error dynamics.
    • For example:
    • increase verification when contradictions and tool debt accumulate;
    • widen retrieval when error persistence rises;
    • encourage exploration when an embodied agent is trapped in a narrow or stale belief basin;
    • restart or roll back when a cascade begins.
    • Dependencies: The paper explicitly finds substrate-specific intervention optima: Retail favors no intervention or high verification, whereas ALFWorld benefits more from exploration or memory. Adaptive policies therefore require online calibration rather than a universal rule.
  • Checkpointing, rollback, and state repair for autonomous agents — Robotics, software agents, industrial automation
    • Use avalanche onset and local–global divergence to trigger automatic rollback to the last trusted world-state checkpoint.
    • In robotics, this could restore a verified map, inventory, or subgoal state. In software agents, it could revert uncommitted edits or database transactions.
    • Dependencies: Requires reversible actions, transactional environments, reliable state snapshots, and recovery policies that do not themselves amplify the cascade.
  • World-model-aware agent architectures — AI research
    • Integrate explicit belief-state tracking, contradiction ledgers, uncertainty budgets, and task-graph memory into agent architectures.
    • Rather than relying only on chain-of-thought or conversation history, agents could maintain structured state variables corresponding to the paper’s q,b,c,u,r,m,pq,b,c,u,r,m,p representation.
    • Training objectives could penalize:
    • rising state error;
    • unnecessary local–global gaps;
    • persistent error spectra;
    • failure to recover after a contradiction.
    • Dependencies: The measured state is benchmark-grounded and not the model’s private activation state. Better external state tracking may improve observability without necessarily changing the underlying reasoning capability.
  • Dependency-aware planning and task decomposition — Software agents, robotics, project management
    • Use task-graph topology and dependency depth to identify plans likely to amplify early mistakes.
    • Planners could prefer:
    • shallower dependency structures;
    • explicit validation at recursive boundaries;
    • independent subplans where possible;
    • evidence checkpoints before high-dependency decisions.
    • Dependencies: Graph structure must accurately represent semantic dependencies, not merely the order of actions. The observed bounded divergence transition should not be interpreted as unbounded chaos or used to infer universal scaling laws.
  • Safety certification for long-horizon autonomous systems — Healthcare, finance, transportation, public-sector automation
    • Create certification protocols requiring an agent to demonstrate bounded error propagation under controlled perturbations, increasing horizons, and altered dependency depth.
    • Certification could include:
    • stress-injection tests;
    • local–global fidelity tests;
    • recovery after injected contradictions;
    • robustness to paraphrase and surface changes;
    • capability-boundary testing.
    • Dependencies: Certification requires sector-specific gold states and representative operational traces. Benchmark behavior may not transfer directly to real-world distribution shifts.
  • Multi-agent coordination and cascade containment — Distributed AI and robotics
    • Extend the framework from single-agent traces to networks of agents exchanging beliefs, evidence, and tool results.
    • Potential applications include identifying whether one agent’s stale belief creates a cascade across:
    • software development teams of agents;
    • supply-chain planning systems;
    • robot fleets;
    • clinical decision-support pipelines.
    • Network-level diagnostics could measure which communication links or shared memories propagate errors most strongly.
    • Dependencies: Multi-agent attribution is substantially harder than single-agent attribution. Shared failures may arise from common prompts, models, tools, or data rather than inter-agent propagation.
  • Benchmark suites for dynamical reliability and agent stress testing — Academia and industry evaluation
    • Develop standardized benchmarks that vary one structural factor at a time:
    • horizon;
    • dependency depth;
    • task-graph topology;
    • retrieval width;
    • injected contradictions;
    • tool-error debt;
    • reversible versus irreversible actions.
    • Such benchmarks would complement WebArena, SWE-bench, tool-use evaluations, and embodied tasks by measuring how agents fail rather than only whether they succeed.
    • Dependencies: Benchmark designers must avoid circular labels, short-horizon artifacts, and invalid-generation regimes. As shown by the Game of Life experiments, missing signatures may indicate model incapacity rather than absence of dynamics.
  • Predictive maintenance for agentic workflows — Enterprise operations
    • Treat persistent increases in stress, error autocorrelation, or avalanche duration as early-warning indicators of impending workflow failure.
    • Organizations could monitor populations of agent runs and schedule model updates, prompt revisions, retrieval-index repairs, or tool maintenance when failure dynamics deteriorate.
    • Dependencies: The paper finds only modest observational predictive value for natural stress after controlling for task difficulty. Longitudinal validation is needed before using these measures for operational forecasting.
  • Scientific study of agent belief basins and metastability — Cognitive modeling and AI theory
    • Investigate whether agents occupy recurring belief basins from which they can recover, become trapped, or transition abruptly after small perturbations.
    • This could inform memory design, belief revision, and uncertainty calibration in long-horizon reasoning systems.
    • Dependencies: The paper reports metastable rather than fully chaotic behavior, but it does not establish a universal basin structure. Identifying basins requires robust state representations and repeated perturbation experiments.
  • Energy- and cost-aware agent orchestration — Cloud computing and sustainable AI
    • Use dynamical diagnostics to decide when additional verification is cost-effective and when an agent should be terminated, restarted, or escalated.
    • Avoiding long error avalanches could reduce wasted tool calls, retrieval operations, model inference, and human review.
    • Dependencies: Cost savings depend on the accuracy and latency of monitoring. Because intervention optima differ by task, optimization must jointly consider reliability, compute cost, and user impact.

Overall, the paper supports practical monitoring, evaluation, and intervention tools now, but it does not justify claims of universal criticality, universal power laws, or a single optimal control strategy. The feasibility of nearly all applications depends on obtaining trustworthy intermediate-state annotations and validating them across the specific task topology, horizon, and capability regime in which an agent will operate.

Glossary

  • Agent-implied state: A representation of the task state inferred from an agent’s interaction history rather than directly observed from the environment. “The measured object is not a hidden activation state; it is a benchmark-grounded state vector extracted from logs”
  • Avalanche: A temporally connected episode of unusually large errors or system activity following accumulated stress. “For threshold τe\tau_e, define the avalanche set, size, weighted size, and duration”
  • Benjamini–Hochberg false-discovery control: A procedure for limiting the expected proportion of false positives among statistically significant results. “confirmatory p-values are corrected with Benjamini-Hochberg false-discovery control at q=0.05q=0.05”
  • Belief basin: A region of belief-state space in which the system tends to remain temporarily stable. “Belief basins”
  • Capability boundary: A limit beyond which a model cannot produce valid outputs needed for meaningful evaluation. “Game of Life exposes the boundary condition: once grid size exceeds the model's valid-generation regime”
  • Cellular automaton: A discrete computational system in which cells update according to local transition rules. “cellular automata near phase transitions”
  • Chaos: Sensitive, potentially unbounded dependence of system behavior on initial conditions. “a depth-dependent change in bounded propagation, not a positive Lyapunov exponent or unbounded chaos”
  • Crackling noise: Irregular bursts of activity produced by systems responding to gradual external changes. “Related intuitions appear in crackling-noise systems”
  • Critical point: A parameter value at which a system undergoes a qualitative phase transition and often exhibits scale-free behavior. “do not imply a universal critical point”
  • DFA (detrended fluctuation analysis): A method for estimating long-range correlations in nonstationary time series. “spectral and valid-horizon DFA tests”
  • Dynamic range: The span of input intensities over which a system responds effectively. “Neural criticality adapts related measurements to cascades and dynamic range near phase boundaries”
  • Epistemic calibration: The degree to which an agent’s confidence accurately reflects the correctness or uncertainty of its beliefs. “epistemic-calibration studies”
  • Exogenous perturbation: An externally introduced change that is not generated by the system being studied. “a small exogenous stress perturbation can move a capability-matched agent”
  • Finite-size scaling: The analysis of how system behavior changes as the finite size of the system varies. “finite-size scaling: cascade scale depends on a finite system-size variable, such as horizon”
  • Flicker-like behavior: Temporal behavior associated with a power spectrum approximately proportional to $1/f$, indicating persistent fluctuations. “narrow retrieval moves errors toward flicker-like behavior”
  • Fractal dimension: A measure of the complexity or scaling structure of an object or pattern across spatial or graph-based scales. “In HotpotQA, fixed two-hop evidence graphs yield stable fractal dimensions”
  • Graph topology: The structural arrangement of nodes and connections in a graph, independent of geometric layout. “Error clusters follow evidence and task graphs”
  • Heavy-tailed distribution: A probability distribution in which extreme values occur more frequently than under light-tailed distributions such as the Gaussian. “heavy-tailed event sizes”
  • Horizon degradation: The deterioration of task performance as the number of sequential steps increases. “Long-horizon studies measure task completion, compounding error, behavioral drift, and horizon degradation”
  • Identification limitation: A restriction on what can be inferred from the available observations or measurements. “Proposition~\ref{prop:lg_insuff} gives a minimal identification limitation”
  • Independent Bernoulli error process: A model in which each step independently produces an error with a fixed probability. “The observed avalanches are too clustered for a matched independent Bernoulli error process”
  • Local–global gap: The difference between whether an individual action is valid and whether the overall task state remains correct. “The local-global gap is”
  • Long-memory process: A stochastic process in which observations remain statistically dependent across long time intervals. “StatefulPuzzle remains in a long-memory regime after DFA correction”
  • Lyapunov exponent: A quantity measuring the average exponential rate at which nearby system trajectories diverge. “not a positive Lyapunov exponent or unbounded chaos”
  • Markov persistence: Dependence in which the probability of the next state is determined by a limited history, commonly the immediately preceding state. “a Markov-persistence null”
  • Metastable: Temporarily stable but capable of transitioning to another state under perturbation. “belief states show metastable rather than fully chaotic behavior”
  • Multi-hop retrieval: Retrieval that requires gathering and connecting information through multiple reasoning or evidence steps. “HotpotQA-style multi-hop reasoning”
  • Null model: A baseline statistical model representing the absence of the effect or structure being tested. “Matched nulls turn the SOC analogy into finite estimands”
  • Partially observable process: A sequential decision process in which the underlying state cannot be directly observed in full. “We model an agent benchmark as a finite-horizon partially observable process”
  • Power law: A relationship in which the frequency or magnitude of an event scales as a power of another variable. “Tail universality, a unique power-law exponent, and a shared intervention optimum are not required”
  • Propagation regime: A characteristic pattern governing how disturbances or errors spread through a system. “dependency depth changes the propagation regime”
  • Self-organized criticality (SOC): A property of some systems that naturally evolve toward a critical state in which small disturbances can produce cascades of many sizes. “This paper asks whether a similar measurement language can be operationalized as finite world-model SOC”
  • Spectral exponent: An exponent describing how the power of a time series varies across frequencies. “spectral exponents stay in the long-memory regime across horizons”
  • State-space divergence: The separation between two evolving states or trajectories over time. “Dependency-depth divergence transition”
  • Structural conditioning: Dependence of a statistical outcome on the structure of a task, such as its dependency graph. “structural conditioning: propagation statistics depend on task dependency depth or task-graph topology”
  • Substrate: The specific environment, benchmark, or task domain in which a process is measured. “For each substrate ss, let Hs\mathcal H_s be the finite trace space”
  • Temporal dependence: Statistical dependence between observations at different times. “temporal dependence: thresholded state errors are correlated beyond an independent-error null”
  • Universality class: A category of systems sharing common large-scale statistical behavior despite differences in microscopic details. “The stronger universality-class interpretation receives partial evidence”
  • White noise: A random signal whose values are uncorrelated across time and whose power is distributed uniformly across frequencies. “We test whether error streams are compatible with white or nearly white independent noise”
  • World-state fidelity: The degree to which an agent’s represented task state matches the benchmark-grounded or true state. “World-state fidelity records state errors that final success, local syntax checks, and executable tool-call validity do not observe”

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.