Limits of size generalization for GALLOP

Determine whether the size generalization observed for GALLOP policies trained on smaller linear programs extends to substantially larger linear programs than those evaluated in the study.

Background

GALLOP uses a dimension-agnostic reinforcement-learning policy and, in the reported experiments, generalizes from smaller training instances to LPs up to 400 times larger, including Transport LPs with 10.24 million variables. However, the maximum evaluated problem sizes were constrained by the available GPU memory.

The authors identify the unresolved question of whether the observed generalization behavior continues at substantially larger scales. Addressing it requires experiments on hardware with more GPU memory and evaluation beyond the problem sizes considered in the paper.

References

Experiments with more GPU memory would allow us to test whether the observed size generalization extends to substantially larger LPs.

— Reinforcement Learning to Accelerate Primal-Dual Hybrid Gradient for Linear Programming  (2610.01546 - Sul et al., 1 Oct 2026) in Section “Conclusion and Limitations,” subsection “Limitations”