Establish whether evaluation outputs fed back into model weights
Establish whether any information from the evaluation episode flowed back into the language models’ weights.
References
Evaluation versus training is not decidable: a multi-rollout structure is evidenced and fits both; the fictitious date carries no behavioural signal ($\rho=-0.135$, $p=0.126$); the agents call themselves benchmark and evaluator subjects (35 and 20 names), never reward or training (zero). Whether anything flowed back into weights is in no available source.
— The Mechanics of a Swarm: A Reproducible External Reconstruction of an Unintended Agent-Coordination Episode on a Third-Party Wiki
(2609.12748 - Lütje, 11 Sep 2026) in Section 3.7, “What the corpus documents about grading: no correctness feedback”