Establish the causal effect of interface repair on multi-turn RL learning

Establish whether repairing function-call parsing within verl’s own AgentLoop, while holding the remaining training setup fixed, enables Qwen2.5-Coder to learn multi-turn repair through reinforcement learning.

Background

The observed function-calling training arm produced no parser-accepted or executed tool calls, whereas ReAct supplied a working interaction channel. Because the available comparison changed both protocol and interface, it could not isolate whether interface censoring itself prevents learning multi-turn repair. The paper identifies a repaired-function-calling RL comparison at 7B with parameter-efficient tuning as the decisive unresolved experiment.

References

Because this arm changes protocol and interface together it cannot separate a working channel is insufficient'' fromReAct in particular is insufficient''; it is 75 steps on one seed at 1.5B against 150-step baselines, and the step-75 evaluation was lost when the checkpoint write exhausted the disk after training completed.

Interface-Induced Trajectory Censoring  (2609.03966 - Wang, 3 Sep 2026) in Section 3.8, subsection “Why no repaired-FC arm exists yet, and what a working channel did not buy”; Section 4, subsection “What would change our conclusions”