Relative causes of code-retrieval performance decline

Determine the relative contribution of backbone capability retention and code-specific supervision coverage to the decline in CoFree-4B's code-retrieval performance, particularly on the LeetCode and Pony subsets of BRIGHT.

Background

The paper reports that CoFree-4B performs worse than its Qwen3-Embedding-4B backbone on the LeetCode and Pony subsets of BRIGHT. Contrastive-only fine-tuning on the same query–document pairs also reduces performance on these datasets, suggesting that the decline may result from continued training altering code-retrieval capabilities learned by the backbone. The authors further note that limited coverage of task-relevant code reasoning in the retained supervision could contribute, but they do not establish how much each factor matters.

References

The relative contribution of each factor to the decline remains to be established.

— Reasoning Quality Matters: Combating Reasoning Collapse in LLM-based Embedding Learning  (2609.20563 - Gong et al., 17 Sep 2026) in Appendix, Section 'Discussion of Code Retrieval' (label: app:code_retrieval_discussion)