Assess inference-cost reduction through prefix caching and step distillation
Determine whether conditioning the FLUX.2 prefix on a fixed noise level enables prefix caching and whether step distillation reduces the inference cost of PatchWAM by decreasing the number of denoising solver steps.
References
A model trained with its prefix conditioned on a fixed noise level could cache the prefix as the control does, and step distillation \citep{akbari2026flashwam} could reduce the number of solver steps; we have tested neither.
— An Action Is Worth One Patch: Unified World-Action Modeling with PatchWAM
(2609.25961 - Wang et al., 22 Sep 2026) in Section 5.1, “Inference Cost”