Papers
Topics
Authors
Recent
Search
2000 character limit reached

State-Centric Decision Process

Published 12 May 2026 in cs.AI | (2605.12755v1)

Abstract: Language environments such as web browsers, code terminals, and interactive simulations emit raw text rather than states, and provide none of the runtime structure that MDP analysis requires. No explicit state space, no observation-to-state mapping, no certified transitions, and no termination criterion. We introduce the State-Centric Decision Process (SDP), a runtime framework that constructs these missing inputs by having the agent build them, predicate by predicate, as it acts. At each step the agent commits to a natural-language predicate describing how the world should look, takes an action to make it true, and checks the observation against it. Predicates that pass become certified states, and the resulting trajectory carries the four objects language environments do not provide, namely a task-induced state space, an observation-to-state mapping, certified transitions, and a termination criterion. We evaluate SDP on five benchmarks spanning planning, scientific exploration, web reasoning, and multi-hop question answering. SDP achieves the best training-free results on all five, with the advantage widening as the horizon grows. The certified trajectories additionally support analyses unavailable to reactive agents, including per-predicate credit assignment, failure localization, partial-progress measurement, and modular operator replacement.

Summary

  • The paper introduces SDP, a runtime framework where agents propose predicate-based state plans, realize them through actions, validate observations, and replan after failures to close the specification gap in language environments.
  • SDP achieves leading training-free results across five benchmarks, including 61.7 TravelPlanner pass rate, 59.16 ScienceWorld performance, and 58.3/67.2 EM/F1 on HotpotQA, with larger gains on longer-horizon tasks.
  • Validation is the strongest contributor in ablations, while cascades improve performance under tight budgets and replanning helps recoverable environments, although validator errors, retrieval limits, and higher LLM costs remain open challenges.

The specification gap in language environments

The paper's starting point is a formal observation about language agents: the environments they inhabit—web browsers, code terminals, interactive simulators, retrieval pipelines—emit raw text and supply none of the objects that Markov Decision Process (MDP) analysis presupposes. The authors enumerate four missing inputs: a state space SS, an observation-to-state mapping ϕ:H→S\phi : \mathcal{H} \to S, certified transitions (s,a,s′)(s, a, s'), and a termination criterion. Their argument for why no environment-side fix exists is goal-dependence of useful abstractions. Since two histories that a summarization policy can safely identify must be separated by a checkout policy, any single abstraction fine enough to be Markov across all goals collapses toward identity on the history space, while any coarser one is goal-specific. The correct object is therefore a family {ϕg}\{\phi_g\} indexed by goals, which the MDP formalism provides no mechanism to select at runtime.

Two standard responses are dismissed on explicit grounds. The POMDP relaxation presupposes precisely the SS and TT whose existence is in question, so it does not resolve the specification problem. Operating directly on raw histories—the dominant reactive-agent approach—works empirically but recovers no analytic object over which transitions, value functions, or progress can be defined. The paper frames this as a specification gap rather than a sample-complexity problem: more data cannot help when there is no target object to approximate.

The State-Centric Decision Process

SDP supplies the four missing inputs by having the agent construct them at runtime. An SDP is a tuple over a predicate space Σ\Sigma with four operators: Propose maps the current state and goal predicate to a next target predicate; Realize selects an action intended to make that target hold; Validate consumes the resulting observation (it is the sole interface to raw outputs) and returns an integer k≥0k \geq 0 counting how many consecutive head predicates the observation satisfies; Replan replaces the plan tail after repeated failures exceed a budget bb. Predicates that pass validation become certified states, and the trajectory accumulates transitions of the form (st,at,st+k)(s_t, a_t, s_{t+k}).

Two design choices carry most of the weight. First, the decision variable is reorganized from actions to states: Propose solves an outer optimization over predicate chains in ϕ:H→S\phi : \mathcal{H} \to S0 before execution begins, while Realize solves an inner per-step problem over actions conditioned on the committed target. Both stages are realized through prompted LLM calls rather than explicit optimization—a concession the authors state plainly, deferring learned operators to future work. Second, execution failures and planning failures are corrected on separate timescales: a failed action retries against the same predicate without touching the plan, while only budget exhaustion at a target triggers Replan, which locally repairs the suffix rather than restarting from ϕ:H→S\phi : \mathcal{H} \to S1.

A notable mechanism is the cascade: because Validate returns an integer rather than a binary verdict, one action can certify several consecutive predicates at once (e.g., a single script computing both an athlete's pace and a derived total distance). Cascades generalize single-step MDP transitions to ϕ:H→S\phi : \mathcal{H} \to S2 and act as a budget multiplier under tight attempt limits.

The framework's structural guarantee is a Markov property (Proposition 1): conditional on the environment's response depending on history only through the current certified state and plan tail, the next certified state depends only on ϕ:H→S\phi : \mathcal{H} \to S3. This holds by construction—Propose and Realize read only local inputs—and rests on a mild conditional independence assumption about the environment rather than an empirical claim about its full dynamics. The authors are careful to note that supplying the MDP inputs makes downstream methods well-posed on the artifact; it does not constitute application of any particular downstream method.

Empirical results

SDP achieves the best training-free results across five benchmarks spanning planning, web reasoning, scientific exploration, and multi-hop QA, with the advantage widening as horizon grows.

Benchmark Headline result Comparison
TravelPlanner 61.7 final pass rate (GPT-4o); 97.4/93.8 hard-constraint Micro/Macro +14.8/+19.4 over ATLAS despite ATLAS using Gemini-2.5-Pro
AssistantBench 31.8 accuracy, 14.9 EM Best overall; Easy-tier 92.8% exceeds next best by >10 points
ScienceWorld 59.16 overall +11.3 over Plan-and-Act; Long tasks 50.41 (+15.6)
HotpotQA 58.3 EM / 67.2 F1 Best among compared methods
MuSiQue 41.4 EM / 51.9 F1 Largest margin, consistent with deeper chains

On TravelPlanner, the hard-constraint advantage traces directly to predicate structure: each constraint becomes a separately certified predicate, catching violations where they arise. Reproduction runs showed 12–18% malformed outputs and 20–30% budget overflow for plan-as-text baselines; SDP eliminates both because Realize selects among pre-filtered feasible options. On ScienceWorld, predicates compress plans substantially—one "boil water" task used 7 predicates subsuming a 36-step oracle sequence—and the Long-task lead confirms the horizon-scaling claim. On multi-hop QA, Validate blocks hallucinated intermediate findings before propagation, decisive in HotpotQA's distractor setting.

The paper concedes where the structure does not win. On AssistantBench Hard tier, Magentic-One leads because multi-page browsing sessions exceed SDP's search-and-scrape interface. On MuSiQue, roughly 76% of failures trace to a retrieval bottleneck—Validate correctly rejects unsupported findings, but with BM25 alone there is no better paragraph to try—rather than planning or validation errors. These are honest boundary conditions on the headline claims.

Trajectory anatomy and ablations

Because every rejection attaches to a specific predicate, certified trajectories support analyses unavailable to reactive agents. Cascade rates range from 0% (TravelPlanner) to 37% (ScienceWorld), and replan utility tracks environment recoverability—ScienceWorld retains full success through one replan while TravelPlanner declines steadily, since replanning cannot rescue infeasible option sets. Failed runs are fractions rather than binary outcomes: failed TravelPlanner runs certify 44% of their plan, failed MuSiQue runs certify 60–64% of hops, with longer chains failing farther in. Validator calibration is auditable: goal certification agrees with correct answers 79% (HotpotQA) and 60% (MuSiQue) of the time, while forced-finalization runs show sharply lower precision (41% and 19%), indicating the certification signal carries information beyond parametric guessing.

Ablations isolate each mechanism. Removing Validate produces the largest drop on four of five benchmarks (e.g., ScienceWorld 59.2 → 15.7); TravelPlanner is the exception because its environment structurally pre-filters candidates, partially substituting for Validate. Replan's contribution scales with recoverability (strongest on ScienceWorld, weakest on AssistantBench). Cascade matters mainly under tight budgets (ScienceWorld 59.2 → 32.0). One methodological caveat deserves emphasis: ablation estimates are computed by replaying recorded trajectories under counterfactual rules with rescaling factors, not by re-execution. The authors acknowledge these estimates assume removing a mechanism does not alter subsequent behavior, making them optimistic bounds on true effects—valid for ordering mechanisms, less so for exact magnitudes.

Limitations and open questions

The limitations are stated directly. Validate can produce false positives on superficially matching observations, so the certified-state guarantee is only as strong as validator soundness, quantified through calibration. Propose bounds plan quality by the LLM's decomposition ability, and Replan cannot recover consistently unreachable targets. Natural-language predicates restrict expressible conditions, particularly continuous quantities. SDP also uses more LLM calls per environment step than reactive baselines, and the cost–benefit tradeoff is not systematically characterized. The optimization view of Eqs. (outer/inner) remains aspirational: both stages are prompted LLM calls, and whether learning a state generator for Propose, a policy over ϕ:H→S\phi : \mathcal{H} \to S4 for Realize, or offline RL on certified tuples ϕ:H→S\phi : \mathcal{H} \to S5 outperforms prompting is left open. Whether validator soundness can be improved systematically—using the identifiable subset of false-certification runs the anatomy analysis surfaces—is likewise unresolved.

Conclusion

SDP reframes the missing MDP structure of language environments as a specification problem the agent itself can solve: commit to falsifiable natural-language predicates before acting, certify observations against them, and accumulate a trajectory carrying states, mappings, certified transitions, termination, and credit. The empirical case is strong and horizon-consistent, and the diagnostic analyses demonstrate that the artifact's value extends beyond task scores. The framework's principal open questions concern replacing prompted operators with learned ones and strengthening validator soundness, both made well-posed by the interfaces the paper defines.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.