Generalization of Signal-Based Sampling to Real Users and Broader Domains

Determine whether the advantages of the lightweight, deterministic signal-based trajectory triage and sampling framework over random and heuristic sampling observed on τ-bench generalize to a broader range of application domains and to interactions with real user populations rather than LLM-simulated users.

Background

The paper proposes a lightweight, signal-based framework that computes deterministic interaction and execution signals from agent trajectories to triage and prioritize trajectories for human review without model calls.

Empirical results on τ-bench (airline and retail domains with LLM-simulated users) show higher developer-informativeness rates and annotation efficiency compared to random and heuristic baselines, but the authors note that whether these advantages carry over to other domains and real user populations is unresolved.

References

While these domains exercise all signal categories in the taxonomy, whether the observed advantages generalize to a broader range of domains and to real user populations remains an open question.

Signals: Trajectory Sampling and Triage for Agentic Interactions  (2604.00356 - Chen et al., 1 Apr 2026) in Limitations (Section 5)

Closing this gap properly --- designing a signal that recovers more of the 11--12\% headroom at a charged probe cost --- is left to follow-up work.

Sampling Luck Masquerades as Allocation Gain: Auditing Test-Time Budget Allocation for Neural Combinatorial Optimization  (2608.13087 - Bae, 13 Aug 2026) in Section VI, Limitations, item 4; related discussion in Section IV-F

The tau-bench domain is the third open problem.

CatchBench: When Can an Agent Failure Be Caught?  (2608.22808 - Zhao, 24 Aug 2026) in Section 6, Discussion and Limitations, paragraph “Measured Headroom Locates Three Open Problems”; see also Section 4.2, subsection “LIVE: How Early Is Failure Visible?” and Appendix, subsection “Additional LIVE Prefix Results,” paragraph “Tau-Bench Remains Below the Warning Bar.”

First, our experiments use a ReAct-based agent framework, so it remains unclear whether the findings extend to other agent architectures and interaction policies.

Monitoring Web Agents Without Internal Signals: Observable Trajectories and Key-Step Supervision  (2609.02057 - Pan et al., 2 Sep 2026) in Limitations section