Determine whether stronger literal matching increases ordinary-QA degradation
Determine whether the stronger literal matching of the Qwen3-Embedding-8B backbone causes the larger loss in ordinary single-hop question answering observed when the Madeleine association term is added.
References
The loss comes from direct single-hop questions, where the association term pulls the ranking away from the literally best-matching turn towards related situations; for the earlier v2-only 4B retriever, 91 questions lost their evidence from the top 10 and 62 gained it. The extra v2d source, which removes this cost on 4B, does not remove it on 8B; we conjecture that the stronger literal matching of the 8B backbone makes the cost larger.
— Madeleine: Learning Involuntary Recall for Conversational Memory from Simulated Lives
(2610.01118 - Shi, 1 Oct 2026) in Appendix, Section “Full Results,” paragraph beginning “Ordinary LoCoMo QA”