Identify the cause of collapse in an unanchored training run

Determine whether the severe collapse observed in the early Qwen3.5-4B training run using only the tree-factorized listwise loss was caused by removing the KL anchors, by the constant learning rate and lack of warmup, by the small batch size, or by their interaction.

Background

The LLM2Jev fine-tuning objective combines a tree-factorized listwise loss with KL anchors intended to preserve legal-token probability mass and general generation behavior. An early run using only the listwise loss experienced extreme degradation: legal probability mass fell substantially, JevBench accuracy dropped, and no text-mode answers could be parsed. However, that run also differed from the later ablations in several optimization settings, including learning rate, warmup, and batch size. Because these factors were confounded, the paper explicitly leaves the causal source of the collapse unresolved.

References

An early run that used only the listwise loss did collapse: legal mass fell to about $e{-17}$, JevBench accuracy to 52.8\%, and none of the 231 text-mode answers could be parsed. That run motivated the anchors, but it also used a constant learning rate of $10{-5}$, no warmup, and a batch size of 3, so we cannot attribute its collapse to the missing anchors alone.

— LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them  (2610.02076 - Li et al., 1 Oct 2026) in Appendix, Section "Ablation Details," subsection "Aggregate benchmark scores are insensitive to anchor strength"