Develop training-time methods to reproduce the external-comment benefit

Develop training-time methods that reliably reproduce the pass@1 improvement produced by correct externally prefixed comment blocks through self-elicited comments or plans generated by the recipient model.

Background

The study finds that correct comments supplied by stronger source models substantially improve weaker recipient models' code-generation pass@1, whereas prompting the recipient to generate its own comments, explicit self-plans, or two-stage plans does not reliably reproduce that improvement. Across the tested prompting configurations, the best self-elicitation result recovers only about 24% of the external-comment lift, leaving unresolved how recipient models could be trained to generate solution content that is both correct and useful for their subsequent code generation. The authors identify training-time methods as a likely direction but do not develop or evaluate such a method in the paper.

References

Despite testing ten prompt variants plus explicit self-planning and two-stage pipelines (Section~\ref{sec:self-elicit}), we did not find a prompting method that reliably reproduces the external-comment lift through self-elicitation; closing this gap likely requires training-time methods, which we leave to future work.

— Talking to Itself While Coding: What Makes Comments Help Code Generation?  (2609.09242 - Pan et al., 8 Sep 2026) in Section 'Limitations'