Assess reward-weight sensitivity at larger problem sizes

Determine whether changing the expert-iteration reward coefficients produces a difference in the learned coarsening policy at problem sizes larger than the tested values of N=10 and N=20.

Background

The paper evaluates alternative weightings of the expert-iteration reward by re-scoring cached candidate coarsenings. At N=10 and N=20, the resulting policies are effectively indistinguishable across the tested configurations, suggesting that the reward coefficients primarily encode a priority ordering in this regime.

The authors explicitly qualify this result as limited to the smaller tested problem sizes: they do not determine whether reward-weight changes would affect policy performance at larger N. Establishing this would clarify whether the apparent insensitivity of the learned coarsener is robust when feasibility becomes more difficult and the fixed-depth policy degrades.

References

This was measured at $N{=}10$ and $N{=}20$, where both configurations lie close to the feasibility ceiling; whether a difference emerges at larger $N$ was not tested.

GNN-Guided Graph Coarsening and Adaptive QUBO Penalties for the Capacitated Vehicle Routing Problem with Time Windows on a Quantum Annealer  (2609.04593 - Rezk et al., 4 Sep 2026) in Section “Robustness: architecture and reward sensitivity,” paragraph “Reward coefficients”