Optimal RL recipe for agentic reasoning
Develop the optimal reinforcement learning training recipe for agentic reasoning in large language model agents that integrate external tools, specifying the algorithmic components and settings that yield the best performance and stability.
References
Despite rapid progress in GRPO-based variants, the optimal RL recipe for agentic reasoning remains unclear.
Effective reward design, credit assignment, and optimization stability remain open challenges, particularly under joint multimodal interaction and agentic execution.
Finally, while GRPO shows advantages over PPO in our setting, improving the stability and effectiveness of RL methods for structured generation remains an open challenge.
How to model and assemble these calls into training samples remains an open question.
It remains unclear whether the same routing paradigm generalizes to substantially different settings such as agent planning or open-ended instruction following.