Scaling DARS annotators and policies

Determine whether larger annotators improve supervision quality for Dependency-Aware Reward Shaping and whether larger policies benefit similarly from Dependency-Aware Reward Shaping.

Background

Dependency-Aware Reward Shaping relies on an annotator to identify predicate verification, invalidation, and repair events, while the policy uses the resulting dependency-aware step rewards. The paper evaluates a distilled Qwen3-8B annotator and policies ranging from 1.5B to 8B parameters, but it does not systematically study how scaling either component affects supervision quality or policy performance.

The authors explicitly identify both directions of scaling as unresolved: whether increasing annotator size improves the quality of structural supervision, and whether increasing policy size produces comparable benefits from DARS. These questions remain open beyond the reported experiments.

References

First, we do not systematically explore scaling: whether larger annotators improve supervision quality or larger policies benefit similarly from DARS remains an open question.

— Dependency-Aware Reward Shaping for Agentic Reinforcement Learning  (2610.01207 - Chen et al., 1 Oct 2026) in Section 6, Conclusion and Limitations