Assess inference-cost reduction through prefix caching and step distillation

Determine whether conditioning the FLUX.2 prefix on a fixed noise level enables prefix caching and whether step distillation reduces the inference cost of PatchWAM by decreasing the number of denoising solver steps.

Background

PatchWAM recomputes the full conditioning prefix at every solver step because FLUX.2 modulates prefix tokens with the current noise level, and it also updates future-frame and action tokens jointly. These design choices make inference substantially slower than the dual-expert control.

The paper identifies two possible mitigation strategies: training with a prefix conditioned on a fixed noise level so that the prefix can be cached, and applying step distillation to reduce the number of solver steps. Neither strategy is evaluated in the reported work.

References

A model trained with its prefix conditioned on a fixed noise level could cache the prefix as the control does, and step distillation \citep{akbari2026flashwam} could reduce the number of solver steps; we have tested neither.

— An Action Is Worth One Patch: Unified World-Action Modeling with PatchWAM  (2609.25961 - Wang et al., 22 Sep 2026) in Section 5.1, “Inference Cost”