Papers
Topics
Authors
Recent
Search
2000 character limit reached

Separable GRPO-Based Training (SepGRPO)

Updated 5 January 2026
  • SepGRPO is a multimodal RL methodology that alternates policy optimization to jointly align MLLMs and DiTs, enabling effective Chain-of-Thought reasoning.
  • It employs clipped surrogate objectives, KL penalties, and module-specific rewards to decouple training processes and ensure stable, scalable updates.
  • Empirical benchmarks from GenEval, WISE, and RISEBench demonstrate significant improvements in reasoning generation and visual rendering performance.

The separable GRPO-based training paradigm (SepGRPO) is a reinforcement learning methodology designed for joint alignment of Multimodal LLMs (MLLMs) and Diffusion Transformers (DiTs) within the ThinkGen framework. It employs an alternating policy-gradient approach via Group-Relative Policy Optimization (GRPO), utilizing clipped surrogate objectives with KL penalties and module-specific rewards. SepGRPO enables effective Chain-of-Thought (CoT) reasoning for general visual generation tasks, supporting flexible multi-scenario training while maintaining strict decoupling between the instruction-generating MLLM and the image-generating DiT modules (Jiao et al., 29 Dec 2025).

1. Mathematical Foundation and Policy Representation

SepGRPO formalizes joint RL as alternated optimization over two parameterized policies:

  • MLLM Policy πM(o ∣ q;θM)\pi_M(o\,|\,q; \theta_M):

Inputs qq (captions, reference image plus edit instructions) yield autoregressive “think-phase” token sequences o=(o1,…,oT)o=(o_1,\ldots,o_T) under parameters θM\theta_M.

  • DiT Policy πD(x0:T ∣ z, c; θD)\pi_D(x_{0:T}\,|\,z,\,c;\,\theta_D):

Generates denoising trajectories x0:Tx_{0:T}; latent noise z∼N(0,I)z\sim \mathcal{N}(0, I); conditional input cc extracted from post-</think> hidden states by VGI-Refine; parameters θD\theta_D.

Rewards are module-specific:

  • RM(o):=Rrule(DiTgenerate(o))R_M(o) := R_{\text{rule}}(\text{DiT}_\text{generate}(o)): MLLM tokens are fed to a frozen DiT, resulting images scored by scenario-based models (GenEval, HPSv3, OCR word-accuracy, SigLIP2, NED).
  • qq0: DiT rollouts scored directly when MLLM is frozen.

Advantage estimation employs group-relative normalization: qq1 GRPO applies clipped probability ratios qq2 at token-step granularity: qq3 The surrogate objective is: qq4

Parameter updates are blockwise decoupled: qq5

2. Alternating Training Algorithm and Data Flow

SepGRPO alternates between module-specific RL epochs, described as:

θM\theta_M6

Data flow distinctions:

  • MLLM-GRPO: User inputs prompt autoregressive MLLM token generation up to <\think>; VGI-Refine extracts hidden states; “Prepadding States” are prepended; DiT (frozen) receives these as instructions for sampling images; rewards are computed; only qq6 updated.
  • DiT-GRPO: MLLM (frozen) provides CoT rollout and VGI-Refine output; DiT samples multiple denoising trajectories with fixed qq7; rewards computed by rule models; only qq8 updated.

3. Training Regimen and Hyperparameter Specification

Rollout counts:

  • qq9 (MLLM-GRPO): 8
  • o=(o1,…,oT)o=(o_1,\ldots,o_T)0 (DiT-GRPO): 24

Clipping and KL-Penalty:

  • o=(o1,…,oT)o=(o_1,\ldots,o_T)1 (clip): 0.2
  • o=(o1,…,oT)o=(o_1,\ldots,o_T)2 (KL-penalty): o=(o1,…,oT)o=(o_1,\ldots,o_T)3

Learning rates:

  • o=(o1,…,oT)o=(o_1,\ldots,o_T)4 (MLLM-GRPO): o=(o1,…,oT)o=(o_1,\ldots,o_T)5 (lower than supervised)
  • o=(o1,…,oT)o=(o_1,\ldots,o_T)6 (DiT-GRPO): o=(o1,…,oT)o=(o_1,\ldots,o_T)7

Batch structure:

  • MLLM batches: 32 prompts × o=(o1,…,oT)o=(o_1,\ldots,o_T)8
  • DiT batches: 16 instructions × o=(o1,…,oT)o=(o_1,\ldots,o_T)9
  • Optimizer: AdamW (no weight decay, gradient clip 1.0)

Datasets:

  • MLLM-GRPO: 5K GenEval, 10K reasoning, 3K rendering, 3K editing, 3K reflection
  • DiT-GRPO: Simple-Scene (GenEval), Text-Rendering (CVTG), θM\theta_M0K samples

Sampling efficiency:

  • Denoising steps reduced to 20 per sample at θM\theta_M1 px
  • CFG θM\theta_M2 for initial θM\theta_M3 steps

4. Theoretical Analysis and Convergence

Blockwise coordinate ascent in SepGRPO, under standard smoothness and bounded-reward assumptions, ensures monotonic improvement in module-level reward objectives. KL-penalized surrogate objectives are maximized alternately: θM\theta_M4 The sequence of updates converges to a stationary point of

θM\theta_M5

as per classical blockwise coordinate ascent theory [Nocedal & Wright, 2006].

Advantages of the separable paradigm:

  1. Module-local rewards: Each policy is optimized using tailored reward signals without cross-module gradient entanglement.
  2. Lower memory footprint: Only a single module’s rollout graph is instantiated during its respective update.
  3. Variance reduction: Sampling distributions remain module-specific (token generation for MLLM; latent denoising for DiT), resulting in more stable gradient estimates.

A plausible implication is that such strict separation mitigates delayed reward propagation and credit assignment difficulties, common in end-to-end multimodal RL pipelines.

5. Empirical Evaluation and Benchmark Performance

Ablation results across GenEval, WISE, CVTG, and RISEBench demonstrate the quantitative benefits of SepGRPO. Summary table:

Task w/o RL +MLLM-GRPO* +DiT-GRPO*
GenEval 0.88 0.86 0.89
WISE 0.55 0.76 0.76
CVTG 0.75 0.79 0.84
RISEBench 3.6 13.0 13.0

(* denotes CoT reasoning enabled)

Notable metrics:

  • WISE (reasoning generation): +MLLM-GRPO achieves 0.76 vs. 0.55 (w/o RL), a 21% improvement.
  • RISEBench (reasoning editing): SepGRPO CoT yields 13.0 average vs. 3.6 (w/o CoT), a gain of 9.4.
  • CVTG text rendering accuracy: Stage 3: 0.75; +MLLM-GRPO: 0.79; +DiT-GRPO: 0.84.

This suggests consistent advantages in reasoning, text rendering, semantic alignment, and editing over supervised-only or non-CoT RL training.

6. Context and Research Significance

SepGRPO represents a principled approach for multimodal generative model alignment under RL, addressing limitations of scenario-specific or coupled update strategies. Its blockwise policy optimization enhances the generalizability of CoT reasoning for both instruction and image generation, validated across multiple diverse datasets and benchmarks.

A plausible implication is that such a modular RL training regime can facilitate scalable adaptation to novel generation scenarios, lower computational overhead during training, and enable more interpretable module-specific improvements (Jiao et al., 29 Dec 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Separable GRPO-based Training Paradigm (SepGRPO).