Effectiveness of diverse beam search

Determine whether diverse beam search provides a benefit for post-training robotic world models that learn from alternative futures.

Background

The paper analyzes the reward component of MemSPO under a fixed candidate set and derives a pairwise reward-gradient identity together with a magnitude bound under candidate collapse. However, the analysis does not establish that the search procedure itself generates better candidates or guarantees monotonic improvement, because search changes the candidate distribution and the resulting update is not an unbiased on-policy gradient.

The unresolved issue concerns whether the diversity induced by diverse beam search translates into a substantive learning advantage for robotic world-model post-training, rather than merely producing greater candidate diversity or dispersion. The paper evaluates this empirically, but the theoretical benefit remains unestablished.

References

The benefit of diverse beam search remains an empirical question.

— FutureWorlds: Learning Robotic World Models from Alternative Futures  (2610.01019 - Wu et al., 1 Oct 2026) in Appendix A, Scope, p. 15