Decision-theoretic integration of clarification value and cost for long-horizon agents

Develop a decision-theoretic framework that integrates the value of information obtained from clarifying questions with the costs of user interaction—specifically interruption, latency, and trust costs—to optimally determine when a long-horizon workflow agent should ask for clarification versus proceed autonomously.

Background

The paper studies agents executing long-horizon workflows where human clarification has nontrivial costs. It separately analyzes the value of information gained through clarification and the costs of asking questions, which include interruption, latency, and trust impacts.

While LHAW provides a framework to generate and evaluate underspecified tasks and measures performance gains from clarification, it does not yet unify these gains with explicit costs in a formal decision policy. The authors explicitly defer creating such a unified decision-theoretic framework.

References

Toward this goal, we study the value of information and the cost of corresponding questions separately, with their integration into a decision-theoretic framework left for future work.

LHAW: Controllable Underspecification for Long-Horizon Tasks  (2602.10525 - Pu et al., 11 Feb 2026) in Section 2.1 (Long-Horizon Workflows)

This paper focuses on single-turn clarification; extending CLAIM to multi-turn interaction requires explicit dialogue-state tracking, history-dependent uncertainty estimation, and planning over future turns, which we leave for future work.

A workflow can therefore be technically valid while no longer representing exactly what the user asked for. The present release should consequently be viewed as an inspectable, validation-gated route for constructing a bounded class of collider-simulation workflows requiring expert approval, not as an autonomous replacement for scientific judgement or an operating-system sandbox. The most important next step is to make the interaction with the user more reliable when the request is incomplete or inconsistent, particularly through explicit clarification and stronger preservation of scientific intent. Within this scope, the results support combining LLMs with deterministic grounding and workflow construction, while also showing that reliable preservation of scientific intent remains the principal unresolved boundary.

HEPLocalAgent 1.0: Running Collider Simulations from Plain-Language Requests on Your Own Computer  (2608.28244 - Singh et al., 28 Aug 2026) in Section 6, Conclusion, p. 16

Pedagogical correspondence is partly interpretive, however, so when and what educators should confirm remains an empirical question rather than a fixed checkpoint schedule.

Beyond Prompt-to-App: Accountable Translation in Teacher-Facing Agentic Authoring  (2609.04679 - Kadir et al., 4 Sep 2026) in Section 6.2, “Agentic authoring requires scoped status claims”

In particular, it remains unclear whether an agent actually recovers formulation-critical missing requirements and whether it knows when enough information has been obtained to proceed.

Ask Before You Optimize: Dynamic Pre-Formulation Clarification for Interactive Optimization  (2609.05258 - Ge et al., 4 Sep 2026) in Section 1, Introduction