Necessity of Test-Time Future Imagination in World Action Models
Determine whether explicit future generation of visual observations during inference is necessary to achieve strong action performance in World Action Models, or whether the primary gains arise from the video prediction objective used during training.
References
More fundamentally, it remains unclear whether explicit future imagination is actually necessary for strong action performance.
Two candidate mechanisms, not mutually exclusive: \emph{(1) Amortized test-time compute:} lookahead generation runs the backbone's forward dynamics at a horizon the action pass never explicitly computes, materializing an implicit forecast into an explicit, reusable conditioning signal. (2) A training-time scaffold: offline lookahead supervision factorizes the demonstrated behavior into where to go and \emph{how to get there}, and the test-time lookahead merely keeps the input distribution matched to that factorization. Our evidence does not yet separate the two --- the flat dose curve is consistent with both.