Generalization to Agents That Share Their Thought Process

Determine whether the findings on covert indirect prompt-injection success and the ICoA attack in ReAct-style tool-using agents apply to agents that share their thought process with the user.

Background

The paper evaluates covert and overt indirect prompt-injection outcomes primarily in ReAct-style agents whose intermediate reasoning, tool calls, and observations remain hidden while only the final response is shown to the user. Its central mechanism depends on this separation between internal actions and user-visible output: an injected action can remain covert when the agent returns to the original user task before producing its final response. The authors explicitly leave unresolved whether these results extend to agents that expose their reasoning or thought process to users, since such disclosure could substantially change the user's ability to detect the injected action.

References

Second, our tests focus only on ReAct-style loops. We do not know if the results apply to agents that share their thought process with the user.

Will the User Ever Know? Covert Indirect Prompt Injection on Tool-Using LLM Agents  (2608.30362 - Lee et al., 31 Aug 2026) in Limitations section