Identify the cause of collapse in an unanchored training run
Determine whether the severe collapse observed in the early Qwen3.5-4B training run using only the tree-factorized listwise loss was caused by removing the KL anchors, by the constant learning rate and lack of warmup, by the small batch size, or by their interaction.
References
An early run that used only the listwise loss did collapse: legal mass fell to about $e{-17}$, JevBench accuracy to 52.8\%, and none of the 231 text-mode answers could be parsed. That run motivated the anchors, but it also used a constant learning rate of $10{-5}$, no warmup, and a batch size of 3, so we cannot attribute its collapse to the missing anchors alone.
— LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them
(2610.02076 - Li et al., 1 Oct 2026) in Appendix, Section "Ablation Details," subsection "Aggregate benchmark scores are insensitive to anchor strength"