Extend SCOPE to larger backbones and additional problem families

Replicate the SCOPE results on larger language-model backbones and transfer the SCOPE recipe to additional mathematical problem families beyond the constructively generated induction, inequality, and number-theory domains.

Background

The paper’s main experiments use SmolLM2-135M, with only a preliminary three-configuration comparison involving a 1B backbone. The reported gains are attributed primarily to restricting free generation and exposing current symbolic-engine state, but the extent to which these conclusions persist at larger model scales and across broader mathematical domains has not been established.

The authors therefore identify replication on larger backbones and transfer to additional problem families as future work rather than as questions resolved by the present experiments. This is distinct from the paper’s completed validation on the public Lean-Workbook subset, whose admissible domain remains limited.

References

Replication on larger backbones and transfer of the recipe to more problem families remain future work.

— SCOPE: Certified Theorem Proving with a Language Model as the Policy Planner  (2610.08319 - Zhou et al., 6 Oct 2026) in Section XII, Discussion

One negative result is reported in full. Merging the corpus of the existing routing families into training, so that routing problems enter the model's capability domain, fails the zero-regression gate in all three mixing ratios (176, 184, and 172 out of 218, all below 191), with degradation not monotone in the ratio; we register this as a sensitivity of corpus mixing to the training trajectory, and model-level inclusion of the routing domain is left to future methodological work.

— SCOPE: Certified Theorem Proving with a Language Model as the Policy Planner  (2610.08319 - Zhou et al., 6 Oct 2026) in Section X, Domain Extension: Validation on a Public Problem Library