Establish the causal effect of interface repair on multi-turn RL learning
Establish whether repairing function-call parsing within verl’s own AgentLoop, while holding the remaining training setup fixed, enables Qwen2.5-Coder to learn multi-turn repair through reinforcement learning.
References
Because this arm changes protocol and interface together it cannot separate a working channel is insufficient'' fromReAct in particular is insufficient''; it is 75 steps on one seed at 1.5B against 150-step baselines, and the step-75 evaluation was lost when the checkpoint write exhausted the disk after training completed.
— Interface-Induced Trajectory Censoring
(2609.03966 - Wang, 3 Sep 2026) in Section 3.8, subsection “Why no repaired-FC arm exists yet, and what a working channel did not buy”; Section 4, subsection “What would change our conclusions”