Transfer of frequency-conditioned action generation to other architectures

Investigate whether frequency-conditioned action generation transfers to other architectures, including world-action models.

Background

FreqFM is introduced as a frequency-conditioned Flow Matching framework for vision-language-action models, using frequency-dependent action statistics to shape the source distribution, adapt the training objective, and constrain classifier-free guidance. The experiments evaluate the method with the π₀ and π₀.₅ VLA backbones, but do not establish whether the approach generalizes beyond these architectures.

The authors explicitly identify transfer to other architectures, including world-action models, as a future-work question. This remains unresolved in the paper because no experiments or theoretical results address such transfer.

References

Future work could also test whether frequency-conditioned action generation transfers to other architectures, including world-action models.

Frequency-Conditioned Flow Matching for Vision-Language-Action Models  (2609.10405 - Niu et al., 9 Sep 2026) in Future work, Conclusion