Papers
Topics
Authors
Recent
Search
2000 character limit reached

Modular MeanFlow: Unified One-Step Modeling

Updated 26 November 2025
  • Modular MeanFlow (MMF) is a framework that efficiently generates high-quality data samples in one step via time-averaged velocity regression.
  • It introduces a tunable gradient modulation mechanism with a curriculum warmup to balance training stability and model expressiveness.
  • Empirical results show state-of-the-art performance in image synthesis, low-data regimes, and out-of-distribution scenarios.

Modular MeanFlow (MMF) is a unifying framework for stable and scalable one-step generative modeling, developed to efficiently generate high-quality data samples via direct mapping in a single function evaluation. MMF generalizes and interpolates between flow-matching and consistency-based models by introducing a principled family of regression losses built upon time-averaged velocity fields. Central to its design are a differential identity linking instantaneous and averaged velocities, a tunable gradient modulation mechanism, and a curriculum-style warmup schedule for training stability and expressiveness. Empirically, MMF achieves state-of-the-art performance across image synthesis, low-data, out-of-distribution (OOD), and trajectory modeling tasks, while circumventing the computational burden of higher-order derivatives (You et al., 24 Aug 2025).

1. Theoretical Framework

MMF builds on the continuous-time generative model defined by the ordinary differential equation (ODE):

dxtdt=v(xt,t),x1∼pprior,    x0∼pdata\frac{dx_t}{dt} = v(x_t, t), \qquad x_1 \sim p_\text{prior}, \;\; x_0 \sim p_\text{data}

where v(xt,t)v(x_t, t) denotes the instantaneous velocity field parameterizing the mapping from ppriorp_\text{prior} (usually a tractable distribution) to pdatap_\text{data}. MMF introduces the time-averaged velocity field over the interval [r,t][r, t]:

u(xt,r,t):=1t−r∫rtv(xτ,τ) dτu(x_t, r, t) := \frac{1}{t - r} \int_{r}^{t} v(x_\tau, \tau)\,d\tau

With Lipschitz assumptions on vv, the averaged velocity recovers the instantaneous field as t→rt \to r:

lim⁡t→ru(xt,r,t)=v(xr,r)\lim_{t \to r} u(x_t, r, t) = v(x_r, r)

A key identity underpins MMF:

v(xt,t)=u(xt,r,t)+(t−r)ddtu(xt,r,t)v(x_t, t) = u(x_t, r, t) + (t - r) \frac{d}{dt} u(x_t, r, t)

where v(xt,t)v(x_t, t)0. This relation enables the regression of averaged velocities and their time derivatives to approximate the model's functional path, decoupling expressiveness from the risk of instability intrinsic to higher-order supervision.

2. Modular Loss Construction and Gradient Modulation

MMF defines a spectrum of regression losses parametrized both by velocity averaging interval and a scalar v(xt,t)v(x_t, t)1 controlling gradient flow. The “full” regression loss is given by:

v(xt,t)v(x_t, t)2

To trade off training stability and functional expressiveness, a partial stop-gradient operator is introduced:

v(xt,t)v(x_t, t)3

The MMF loss then generalizes as:

v(xt,t)v(x_t, t)4

Here, v(xt,t)v(x_t, t)5 fully propagates gradients (maximum expressiveness but possible instability), v(xt,t)v(x_t, t)6 detaches Jacobian-vector products (maximum stability), and intermediate v(xt,t)v(x_t, t)7 governs a stability-expressiveness continuum. Explicitly blocking gradient flow through higher-order terms prevents gradient explosions and training oscillations.

3. Curriculum Warmup and Training Protocol

MMF adopts a curriculum for v(xt,t)v(x_t, t)8:

v(xt,t)v(x_t, t)9

In early training (ppriorp_\text{prior}0), MMF behaves as a consistency or flow-matching model, yielding high stability by restricting second-order signal propagation. As ppriorp_\text{prior}1, expressive gradients are introduced, allowing the model to capture richer curvature and achieve lower asymptotic loss. Empirically, this curriculum schedule yields both rapid convergence and low variance in training.

A typical MMF training protocol:

[r,t][r, t]3

Sampling proceeds in one step: ppriorp_\text{prior}2.

4. Connections to Prior Methods

MMF subsumes prior consistency and flow-matching models as special cases:

  • Consistency Models: Fixing ppriorp_\text{prior}3 and ppriorp_\text{prior}4 recovers the fixed-time consistency loss ppriorp_\text{prior}5.
  • Flow Matching: The instantaneous limit ppriorp_\text{prior}6, ppriorp_\text{prior}7, and ppriorp_\text{prior}8 recovers the flow-matching loss ppriorp_\text{prior}9.
  • Gradient Efficiency: Applying stop-gradient to the Jacobian-vector term ensures that backward computation never traverses pdatap_\text{data}0, eliminating pdatap_\text{data}1 cost and Hessian-vector product overhead.

This unification allows MMF to inherit the interpretability and theoretical properties of both frameworks, while providing a tunable control for interpolation between them.

5. Empirical Results and Model Analysis

MMF's empirical evaluation focuses on image synthesis, robustness, and path modeling:

Model FID (↓) 1-step MSE (↓) LPIPS (↓) Inference Time (s) (↓)
MeanFlow (full) 3.91 0.087 0.132 0.031
MeanFlow (stop-grad) 4.27 0.095 0.156 0.024
MMF (λ=0) 4.19 0.093 0.148 0.023
MMF (λ=0.5) 3.78 0.084 0.120 0.026
MMF (λ=1) 3.62 0.080 0.109 0.034
MMF (curriculum) 3.41 0.076 0.097 0.025

On CIFAR-10 and ImageNet-64, curriculum MMF achieves the lowest FID, lowest 1-step MSE, and the highest diversity (LPIPS), matching or exceeding the efficiency of prior mean flow and consistency baselines. Few-shot and OOD experiments demonstrate that curriculum MMF retains low FID even with as little as 1% of CIFAR-10 data and achieves 10–20% lower FID in OOD settings (SVHN, STL-10, CIFAR-C) compared to baselines. In ODE-fitting and 2D control tasks, curriculum MMF yields smooth, accurate paths, outperforming noisy full-gradient and oversmoothed stop-grad alternatives.

Path deviation is formalized as:

pdatap_\text{data}2

Curriculum MMF achieves the lowest pdatap_\text{data}3, supporting latent interpolation smoothness.

6. Ablations and Practicalities

Extensive ablations reveal that:

  • Varying pdatap_\text{data}4: pdatap_\text{data}5 yields maximum stability but higher FID (underfitting curvature). pdatap_\text{data}6 is most expressive but unstable (loss oscillations). pdatap_\text{data}7 provides some smoothing but with late-stage variance. Curriculum pdatap_\text{data}8 combines low early variance and best final performance.
  • Curriculum Horizon: Short warmup (small pdatap_\text{data}9) induces early instability; long warmup is too conservative with slower convergence. Optimal [r,t][r, t]0–[r,t][r, t]1 of total steps.
  • Compute: Forward-mode autodiff for the Jacobian-vector-product yields ~15% overhead; with stopgrad, no backward is needed through this term.

The standard MMF implementation utilizes a UNet architecture with sinusoidal time embeddings, Adam optimizer (learning rate [r,t][r, t]2, batch size 128, cosine decay), and a curriculum warmup over 100k steps.

7. Significance, Limitations, and Outlook

MMF provides a theoretically grounded, computationally efficient, and practically robust approach for one-step generative modeling. By enabling a tunable spectrum between expressiveness and stability—mediated by gradient modulation and curriculum scheduling—it addresses the instability and inefficiency intrinsic to prior higher-order methods. MMF’s empirical results demonstrate high generalization under low-data and out-of-distribution regimes and applicability beyond image synthesis to trajectory modeling. A plausible implication is that MMF may be extensible to other domains requiring stable, one-shot sampling of complex data distributions via learnable ODE flows (You et al., 24 Aug 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Modular MeanFlow (MMF).