Characterize the Effect of Cross-Editing on Edit Correctness

Characterize how cross-editing affects a language model’s ability to produce correct code edits that pass the collected tests, including whether editing code generated by another model systematically changes edit correctness relative to self-editing.

Background

The main analysis finds a consistent self-editing versus cross-editing pattern for edit size, but the corresponding effect on correctness is not established. In the reported edit-success matrix, self-editing does not consistently outperform cross-editing: three of the five models perform best on Olmo 3 implementations, and all models perform worst on Claude Haiku implementations. The authors explicitly leave the relationship between cross-editing and correct-edit performance unresolved for future work.

References

How cross-editing affects a model's ability to make correct edits is an interesting question that we leave to future work.

CROCODIL: Cross-Model Code Editing with LLMs  (2609.03894 - Zhong et al., 3 Sep 2026) in Appendix, Section “Edit Success Across Cross-editing”