Learn meta-skills that co-evolve with reasoning policies through reinforcement learning

Develop reinforcement-learning methods for learning meta-skills that co-evolve with the reasoning policy, rather than relying on predefined meta-skill programs or fixed workflows.

Background

The paper identifies a limitation in existing meta-skill-driven skill-optimization methods: they generally rely on predefined prompts, rubrics, specifications, or multi-agent workflows to revise individual skills, while the reasoning agent is typically held fixed. These fixed update rules can become misaligned with an evolving reasoning policy.

The unresolved problem is to learn the meta-skill itself through reinforcement learning so that its skill-update policy can adapt jointly with the reasoning policy. The proposed CoSkill framework is presented as an approach addressing this previously open direction.

References

Learning meta-skills that co-evolve with the reasoning policy through RL therefore remains open.

CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution  (2609.04865 - Feng et al., 4 Sep 2026) in Section 1, Introduction