Detecting Concealed Indirect Prompt Injections via CoT or Activation Monitoring
Investigate whether monitoring complete chain-of-thought traces or probing internal neural activations using activation probes can detect concealed indirect prompt injection compromises that are not evident from user-facing outputs in the Indirect Prompt Injection Arena scenarios.
References
Since our evaluation focused on user-facing outputs, an open question is whether monitoring full CoT traces or internal representations via activation probes~\citep{kramar2026building} could detect concealed attacks that evade output-level scrutiny.
Whether internal-activation anchors (mechanistic interpretability, activation monitoring) could detect it is an open question this study does not address; our claim is bounded to output-layer statistical detection, not to verification in general.