Scaling behavior of domain-shift degradation

Determine whether the domain-shift degradation patterns observed for the Qwen2.5-0.5B and 1.5B financial named entity recognition models attenuate at much larger model scales.

Background

The study evaluates relatively small generative models—Qwen2.5-0.5B and Qwen2.5-1.5B—trained on approximately 1,000 sentences. Although the 1.5B scale check qualitatively reproduces the smaller model’s findings, it uses only a single random seed and therefore does not establish a robust scaling trend.

The unresolved issue is whether increasing model scale reduces the degradation in extraction quality and confidence reliability caused by shifts from SEC filings to financial news and general-topic social media. Resolving this question would clarify whether the reported robustness limitations are primarily consequences of small model capacity or persist in substantially larger models.

References

The 1.5B scale check rests on a single seed, so its agreement with the 0.5B results is qualitative rather than a demonstrated replication; whether the shift-degradation patterns attenuate at much larger scale is open.

Reliable Financial Named Entity Recognition under Domain Shift  (2608.19558 - Zheng et al., 20 Aug 2026) in Section 7, “Limitations and Ethical Considerations,” paragraph “Scale”