Determine whether Qwen would answer from decoy content under valid tool access

Determine whether Qwen3.6-27B answers from a decoy datasheet when document tools remain callable, rather than treating its exclusion from the wrong-content probe as evidence of either success or failure.

Background

The wrong-content and null-tool probe arms were intended to test failures that preserve the tool-call surface while corrupting the retrieved document content. Qwen3.6-27B experienced frequent token-exhaustion and engine errors under forced tool choice in these arms, so the authors exclude its results rather than interpreting them as behavioral evidence. Its response to decoy content therefore remains unresolved.

References

Those arms would measure the forcing rather than the model, so we report them on Claude Sonnet~4.6 and GPT-5.1 only, and the 100-run denominator in Section~\ref{sec:res-silent} is two arms $\times$ two models $\times$ 25 claims. The exclusion is a coverage limit, not a result: nothing here says Qwen would or would not answer from a decoy.

Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction  (2608.28439 - Ye et al., 28 Aug 2026) in Appendix, Section “Probe Arms and Their Coverage,” paragraph “Why Qwen is excluded from two of the three” (Section~\ref{sec:appendix-probes})