Determine the real-world prevalence of unresolved pending plans

Determine the prevalence of at-risk pending plans—user-stated plans that are abandoned or remain unresolved—in real-world conversational-assistant interactions, in order to establish the real-world performance of speculation-contamination safeguards.

Background

The evaluation uses constructed plan-to-outcome datasets because existing public conversational-memory datasets contain too few cases in which speculative plans can be tracked through abandonment or non-resolution. Consequently, the reported contamination rates do not establish how frequently this failure occurs in deployed conversational assistants.

The paper explicitly identifies the prevalence of such pending plans as unknown. Longitudinal conversational corpora are mentioned as a possible source for future validation, but the prevalence itself is not measured in the study.

References

First, the unknown prevalence of at-risk pending plans means our constructed datasets do not establish real-world performance (Section~\ref{sec:datasets}).

— AgentMemGate: Addressing Speculation Contamination in Conversational Assistant Memory  (2610.07707 - Sharma et al., 6 Oct 2026) in Section “Discussion, limitations, and conclusion,” paragraph beginning “Five limitations bound generalization”

Third, we do not test whether staging harms answers about plans.

— AgentMemGate: Addressing Speculation Contamination in Conversational Assistant Memory  (2610.07707 - Sharma et al., 6 Oct 2026) in Section “Discussion, limitations, and conclusion,” paragraph beginning “Five limitations bound generalization”