Long-context scaling of quantization-amplified proactive interference
Determine whether the proactive-interference penalty caused by bitsandbytes INT4 quantization strengthens or attenuates when the PI-LLM retrieval paradigm is extended from the few-hundred-token prompts tested in the study to contexts of 8,000 or more tokens.
References
Our prompts are also short-to-medium (at most a few hundred tokens), so whether the effect strengthens or attenuates in the 8k+-token long-context regime is an open, practically important question -- one that prior work on quantized long-context retrieval \citep{mekala2025quantization} suggests could go either way and would be worth testing directly with our paradigm.
— Compress and Forget: bitsandbytes Quantization Amplifies Proactive Interference in LLMs
(2608.18578 - Shahrabi-Farahani et al., 19 Aug 2026) in Section 6, “Limitations and Future Work”