Characterize the role of real prior model-generated context
Determine whether the quantity of real, model-generated prior context in a session affects how qwen3:8b, llama3.1:8b, and mistral:7b respond to rare tool-call failures, independently of the immediate pre-failure token pattern.
References
Whether the sheer quantity of real prior self-generated context (as opposed to its content) affects how a model responds to a failure is a question this design does not resolve, because it was not built to resolve it; it was built to make the rate-of-failure and the immediate pre-failure context comparable, and it succeeds at that narrower goal.
All figures in Section~\ref{sec:results} were taken before the cap was applied. The per-model and per-surface rates should therefore be read as a lower bound with unknown per-model bias, since large-context models were penalised more than small ones. Re-reading the corpus under a fixed cap is the first item of future work.