Determine the cause of option-10 displacement

Determine the cause of the observed shift away from option 10 toward option 1 on ten-option MMLU-Pro questions after fine-tuning Qwen3.5-4B with the LLM2Jev objective, particularly under stronger KL-anchor weights.

Background

The LLM2Jev numeric-suffix interface causes options 1 and 10 to share their first token under digit-level tokenizers. In the MMLU-Pro evaluation, fine-tuning shifts the model's predictions away from option 10, with the effect becoming stronger as the KL-anchor weight increases. For the 4B model, many displaced predictions move to option 1, and most of the net MMLU-Pro losses at the strongest anchor setting occur on these questions. The paper reports the empirical phenomenon but does not establish whether it arises from tokenizer-induced competition, optimization dynamics, the regularization scheme, or another mechanism.

References

At 4B with $\lambda=1$, the displaced answers go to option 1 (178 picks vs.\ 123), and 15 of the model's 17 net MMLU-Pro losses occur on these questions. We have not identified the cause.

— LLM2Jev: LLMs Are Already Jev-Style Decision Models -- When and How to Fine-Tune Them  (2610.02076 - Li et al., 1 Oct 2026) in Appendix, Section "Details of the Data Comparisons," subsection "Option 1 versus option 10" (Appendix option10)