Persistence of answer-side triggers across longer conversations

Determine whether the answer-side backdoor trigger persists and remains effective over longer conversational gaps and intervening turns in multi-turn large language model dialogues.

Background

The paper evaluates its answer-side backdoor primarily in two-turn dialogues, where the trigger-bearing assistant response immediately precedes the harmful query. Although the formal framework permits the trigger to occur at an earlier point in the dialogue history, the authors identify the persistence of the backdoor across longer gaps and intervening interactions as an unresolved empirical issue. Resolving this question would establish whether the attack generalizes beyond the short conversational setting used in the main experiments.

References

Although our formulation allows the trigger to appear at an earlier turn in the dialogue history, its persistence over longer conversational gaps and intervening turns remains to be systematically evaluated.

— The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language Models  (2610.07723 - Zhang et al., 6 Oct 2026) in Limitations