Sensitivity of benchmark results to simulated-user design

Quantify how THINKINGBOX-BENCH results change when varying the simulated user’s backbone model, disclosure behavior, or cooperativeness.

Background

THINKINGBOX-BENCH holds a single simulated user fixed across all evaluated agents: a GPT-5.4-mini deployment that follows a constrained, cooperative, information-bounded interaction policy. This design improves comparability between agents but limits the benchmark’s claims because the simulator does not model misremembered facts, shifting objectives, uncooperative behavior, or user-initiated actions.

The paper explicitly identifies sensitivity to simulator design as unresolved. Establishing how benchmark scores and conclusions vary across different simulator backbones, disclosure policies, and levels of cooperativeness would clarify the robustness and external validity of the reported agent evaluations.

References

Quantifying sensitivity to the simulator is left to future work, by varying its backbone model, its disclosure behavior, or its cooperativeness.

One Success Isn't Reliability: Thinkingbox, a Sandbox and Benchmark for Agents in Stateful Business Workflows  (2608.19741 - Li et al., 20 Aug 2026) in Appendix A, subsection “Simulator and judge dependence”