Effectiveness of substantially stronger analyzers

Determine whether a substantially stronger analyzer model could write more accurate instructions for Privileged Self-Practice training than the Qwen3.6-27B analyzer used in the experiments.

Background

Privileged Self-Practice relies on a frozen analyzer model to convert failed student rollouts and reference trajectories into task-specific instructions. The experiments use Qwen3.6-27B as the analyzer because invoking a substantially larger or frontier model throughout training would incur high computational and token costs.

The paper reports strong results with the selected analyzer but does not establish whether a substantially stronger analyzer would generate more accurate instructions or improve training outcomes. The question is explicitly identified as unresolved in the conclusion.

References

Second, we did not make our analyzer as a larger or frontier models because the analyzer is invoked throughout training. Whether a substantially stronger analyzer could write more accurate instructions is left open.

— From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents  (2609.29051 - Su et al., 24 Sep 2026) in Section Conclusion