Evaluate ResidualMux in interactive agentic workloads

Determine whether the ResidualMux covert channel persists under interactive agentic dialogue, tool use, and long-context production workloads beyond the controlled 50-turn dialogue-history protocol evaluated in the paper.

Background

The paper evaluates autoregressive generation and repeated reinjection over up to 50 conversational turns using a controlled dialogue-history protocol. These experiments indicate that KV-cache reuse and increasing dialogue history do not eliminate the channel under the tested conditions.

The authors do not establish whether the same behavior holds in more complex agentic settings involving tool calls, interactive state transitions, and production-scale long contexts. These workloads are explicitly identified as outside the measured evaluation.

References

Second, our autoregressive evaluation (\S\ref{sec:eval:autoregressive}) covers repeated reinjection over 50 turns under a controlled dialogue-history protocol. Interactive agentic dialogue, tool use, and long-context production workloads remain unmeasured.

— Your Model Is Leaking: Covert Information Transfer through LLM Residual Streams  (2609.27996 - Li et al., 23 Sep 2026) in Appendix, Section 'Scope and Limitations' (\S\ref{app:limitations})