---
title: 'RGBX-Next: G-Buffer Video Rendering'
url: https://www.emergentmind.com/topics/rgbx-next
type: topic
---

# RGBX-Next: G-Buffer Video Rendering

RGBX-Next is a unified generative framework for forward and inverse rendering from screen-space G-buffers. It fine-tunes the Wan 2.1 14B video diffusion-transformer (DiT) to estimate intrinsic scene representations from images and videos and to synthesize realistic RGB images and videos from complete or partial G-buffer inputs. Its modality set includes albedo, surface normals, depth, material properties, diffuse irradiance, and albedo-free direct lighting. The framework treats G-buffers as an interface between explicit rendering and generative modeling: they provide spatially explicit control while the diffusion prior supplies appearance detail, lighting effects, and plausible completion of unspecified information [2608.13929].

## 1. Problem formulation and rendering paradigm

Traditional rendering maps an explicitly specified scene to an image,

$$
I=\mathcal{R}(G),
$$

where $G$ includes geometry, materials, camera parameters, and lighting. RGBX-Next instead uses a learned renderer conditioned on a collection of G-buffer modalities $X$:

$$
X\longrightarrow \mathrm{RGB}.
$$

The framework also addresses the inverse problem,

$$
\mathrm{RGB}\longrightarrow X,
$$

in which images or video frames are decomposed into plausible scene-property buffers. The forward model represents conditional synthesis as $\hat I\sim p_\theta(I\mid X)$, whereas the inverse model estimates $\hat X\sim p_\theta(X\mid I)$.

Forward rendering is intrinsically underdetermined because G-buffers may omit complete geometry, hidden surfaces, environment lighting, and transport information. Inverse rendering is likewise ill-posed: multiple combinations of geometry, materials, illumination, and camera can explain the same RGB observation. RGBX-Next therefore produces plausible generative solutions rather than guaranteeing physically unique reconstructions.

The framework occupies an intermediate position between physically based rendering and unconstrained diffusion generation. Classical renderers offer deterministic behavior, temporal stability, and explicit scene control, but require detailed assets and computationally expensive illumination. Diffusion models provide learned priors for realistic appearance but conventionally offer less precise per-pixel control. RGBX-Next combines these properties by using G-buffers as structured conditions and a large video DiT as a learned renderer.

## 2. G-buffer modalities and lighting representations

The notation $X$ denotes a collection of intrinsic or rendering-related modalities rather than a single auxiliary channel. The principal buffers are:

| Symbol | Representation | Function |
|---|---|---|
| $a$ | Albedo | Diffuse material color |
| $n$ | Surface normal | Camera-coordinate surface orientation |
| $z$ | Depth | Per-pixel geometric distance |
| $m$ | Material properties | Roughness, metallicity, and transparency |
| $E$ | Diffuse irradiance | Integrated incoming illumination |
| $d$ | Direct lighting | Albedo-free direct illumination |

Albedo represents diffuse material color. Surface normals are expressed in camera coordinates. Depth is a per-pixel geometric buffer. Material properties include roughness, metallicity, and transparency. Diffuse irradiance $E$ represents illumination integrated over incoming directions, while $d$ is constructed to retain direct-lighting information without multiplying it by diffuse or specular albedo.

The lighting representations are intended to provide controllable alternatives to a full environment map or image-based lighting representation. Diffuse irradiance is useful for controlling broad diffuse illumination, whereas albedo-free direct lighting is comparatively orthogonal to material color and can provide a direct-lighting control signal.

RGBX-Next does not require every buffer to be available. A forward model may condition on albedo alone, depth alone, albedo plus irradiance, or other modality combinations. The general input-output configuration is represented by subsets of a full modality set $\mathcal{M}$:

$$
M_{\mathrm{in}},M_{\mathrm{out}}\subseteq\mathcal{M},
$$

where clean G-buffers occupy $M_{\mathrm{in}}$ and noisy modalities to be generated occupy $M_{\mathrm{out}}$.

## 3. Video DiT architecture and latent interface

RGBX-Next is built by fine-tuning Wan 2.1 14B. The pretrained VAE encoder and decoder remain frozen, while all DiT parameters are fine-tuned. Each modality $Y^m$ is independently encoded into the shared video latent space:

$$
x^m=E_{\mathrm{VAE}}(Y^m).
$$

The latent has spatial resolution reduced by $8\times$, temporal resolution reduced by approximately $4\times$, and 16 channels. For an input with $F$ RGB frames, the latent-frame count is

$$
F'=\frac{F+3}{4},
$$

so $4k+1$ RGB frames become $k+1$ latent frames.

Latents are patchified into tokens using a $2\times2$ spatial block. The DiT applies transformer self-attention, text cross-attention, adaptive LayerNorm timestep conditioning, and 3D rotary positional embeddings. For a token at spatial location $(i,j)$ and latent-frame index $k$, query and key positional modulation is:

$$
Q^m_{i,j,k}=R_{i,j,k}W_Q u^m_{i,j,k},
\qquad
K^m_{i,j,k}=R_{i,j,k}W_K u^m_{i,j,k}.
$$

The 3D positional representation distinguishes spatial and temporal locations, allowing the model to process video motion and temporal consistency.

A central architectural operation is frame-wise concatenation. Clean input modalities occupy some latent-frame slots, while noisy output modalities occupy others:

$$
x_{\mathrm{fw}}(t)=
\left\{x_0^m:m\in M_{\mathrm{in}}\right\}
\cup
\left\{x_t^m:m\in M_{\mathrm{out}}\right\}.
$$

The transformer processes the resulting multimodal token sequence jointly. Clean input tokens provide context and do not receive a denoising loss; output tokens are denoised and supervised.

For multiple clean G-buffer inputs in $X\rightarrow\mathrm{RGB}$ rendering, RGBX-Next uses X-patchify. The clean latent inputs are concatenated along the channel dimension and mapped to conditioning tokens by a learned patchification operator. Noisy output modalities remain separate token streams. Packing only clean inputs avoids optimization difficulties associated with mixing multiple noisy output modalities while preserving the pretrained output structure.

## 4. Flow matching, token typing, and conditional denoising

The Wan model uses flow matching. A clean latent $x_0^m$ is linearly interpolated with Gaussian noise:

$$
x_t^m=(1-t)x_0^m+t\epsilon,
\qquad
t\in[0,1],
\qquad
\epsilon\sim\mathcal{N}(0,I).
$$

The target velocity is $\epsilon-x_0^m$, and the DiT predicts

$$
v_\theta=f_\theta(x_t,t,p).
$$

The standard flow-matching objective is

$$
\mathcal{L}_{\mathrm{FM}}
=
\mathbb{E}_{m,t,\epsilon}
\left[
\left\|
f_\theta(x_t^m,t,p)-(\epsilon-x_0^m)
\right\|_2^2
\right].
$$

For forward and inverse rendering, the loss is restricted to generated modalities:

$$
\mathcal{L}_{\mathrm{fw}}
=
\mathbb{E}
\left[
\sum_{m\in M_{\mathrm{out}}}
\left\|
f_\theta(x_{\mathrm{fw}}(t),t,p)^m
-
(\epsilon-x_0^m)
\right\|_2^2
\right].
$$

RGBX-Next introduces QK type embeddings. Each token receives a type $\tau=(r,m)$, where $r$ identifies its role as an input or output and $m$ identifies its modality. Learned type-dependent offsets are added directly to queries and keys:

$$
Q_\ell=R_\ell W_Q u_\ell+e^Q_{\tau_\ell},
\qquad
K_\ell=R_\ell W_K u_\ell+e^K_{\tau_\ell}.
$$

This explicitly distinguishes, for example, clean RGB input tokens from noisy albedo output tokens. The model also assigns different timesteps to input and output tokens:

$$
\tilde t_\ell=
\begin{cases}
0,&r(\tau_\ell)=\mathrm{in},\\
t,&r(\tau_\ell)=\mathrm{out}.
\end{cases}
$$

Clean input tokens therefore behave as fixed, noise-free latents, while output tokens receive the current diffusion timestep. Ablations indicate that QK type embeddings and clean input tokens improve convergence and output quality.

Classifier-free guidance is used for realistic rendering. With positive and negative prompt velocities $v^+$ and $v^-$,

$$
v_{\mathrm{CFG}}
=
v^+
+
w(v^+-v^-).
$$

Positive prompts such as “photograph, real video, high-resolution, high-quality” and negative prompts containing “synthetic, CG, rendered, video game, unrealistic” steer the model toward photographic appearance. Reversing the prompt meanings produces more synthetic or stylized results while generally retaining G-buffer structure.

## 5. Forward and inverse rendering workflows

### RGB-to-$X$ inverse rendering

The inverse model receives RGB frames as clean inputs and generates selected G-buffers. The output may initially be selected by a text prompt such as “albedo” or “normal”; the improved model uses explicit modality and role embeddings.

Training begins with paired synthetic sequences containing aligned RGB images and G-buffers. The internal synthetic dataset contains 900 path-traced sequences of approximately 100 frames covering indoor and outdoor scenes. Public datasets such as Hypersim and InteriorVerse are also relevant sources.

The RGB-to-$X$ model can estimate albedo, normals, depth, materials, diffuse irradiance, and direct lighting. Estimating diffuse irradiance from real video is reported as a capability not previously shown in this framework. Because inverse rendering is ambiguous, the predicted buffers are plausible decompositions and may not be physically exact.

### $X$-to-RGB forward rendering

The forward model receives one or more clean G-buffer modalities and generates RGB. Specialized models use albedo, normal, depth, albedo plus irradiance, or albedo plus direct lighting. A generalized model accepts albedo, normal, irradiance, and material inputs with modality dropout.

Forward rendering is trained on realistic paired data constructed indirectly. Real videos are collected from Pexels, G-buffers are estimated using RGB-to-$X$, and the estimated buffers are paired with the original RGB videos. Each video is captioned with Qwen2.5-VL, including whether it appears photographic or rendered. The dataset contains approximately 6,622 videos, with 371 held out for testing. Training uses 17-frame clips at $1280\times720$; still-image fine-tuning uses $1920\times1088$.

The use of real RGB targets reduces the synthetic or game-like appearance that can arise when forward rendering is trained solely on synthetic RGB/G-buffer pairs. However, errors in RGB-to-$X$ estimates propagate into the paired conditions.

### Variable-strength and partial conditioning

RGBX-Next trains with modality dropout, spatial masks, and Gaussian blur. For modality $m$, the condition can be transformed as

$$
\tilde x^m(n)
=
\delta^{m,n}M\odot B_{\sigma_m}(x^m),
$$

where $\delta^{m,n}$ gates the modality, $M$ is a spatial mask, and $B_{\sigma_m}$ applies Gaussian blur. A modality can remain active for only an initial fraction of denoising iterations:

$$
\delta^{m,n}
=
\begin{cases}
1,&n\leq cN,\\
0,&n>cN.
\end{cases}
$$

These operations support full modality dropout, partial or spatial conditioning, and coarse structural guidance. During training, modalities are dropped with probability 50%, masked spatially, or blurred with $\sigma\in[1,10]$. The model learns to interpret incomplete G-buffers as weak or absent guidance rather than as literal black or blurred scene content.

The framework also uses SAM 2 to segment regions for spatial control. Albedo is randomly removed from selected segments, encouraging detailed generation in unspecified regions rather than flat outputs.

## 6. Streaming, evaluation, and computational characteristics

RGBX-Next extends fixed-length video generation through chunked streaming. A streaming model receives a stitching frame from the previous chunk, new input frames, and generates new output frames. The reported experiments use $1/16$-streaming models with 16 new frames per chunk.

Teacher forcing uses ground-truth stitching frames during training, but inference uses generated stitching frames, producing a train-test mismatch that can cause drift and blur. Stitching-frame augmentation introduces brightness shifts up to $\pm2\%$, Gaussian blur or unsharp masking with $\sigma\leq1.5$, and VAE decode-encode round trips. Long-context tokens append a clean reference-frame token set to preserve color, lighting, materials, and global appearance.

Self forcing trains on the model’s own previous predictions. RGBX-Next uses a hybrid schedule consisting of augmented teacher forcing, followed by equal-probability teacher forcing and self forcing with a $10\times$ learning-rate reduction. The final RGB-to-$X$ streaming model is reported to remain stable for approximately 1,000 frames, although drift can still occur in very long sequences.

Training uses AdamW, a global batch size of 128 across 32 A100 80 GB GPUs, and learning rates of $10^{-5}$ for image/video fine-tuning and $10^{-6}$ for streaming fine-tuning. Sampling uses 20 diffusion steps. The reported X-to-RGB inference cost for 17 frames is approximately 400 seconds on one A100 80 GB GPU with CFG equal to 3. RGB-to-$X$ uses CFG equal to 1. The models are therefore teacher-quality systems rather than real-time renderers.

On the Hypersim test set, RGBX-Next reports the following intrinsic-decomposition results:

| Metric | Albedo | Normal | Irradiance | Depth |
|---|---:|---:|---:|---:|
| PSNR | 20.17 | 21.22 | 25.19 | 29.47 |
| SSIM | 0.8219 | 0.7638 | 0.8589 | 0.9032 |
| LPIPS | 0.1420 | 0.1986 | 0.3535 | 0.1166 |

It is reported as best across all three metrics for albedo and depth and best on two of three metrics for normal and irradiance among the stated comparisons. Qualitatively, it produces flatter albedo with less residual shading, more plausible sky-region depth, and temporally coherent irradiance, normals, depth, and materials.

On a held-out real RGB/G-buffer test set, forward-rendering FID is:

| Method | FID |
|---|---:|
| Channel-wise concatenation | 88.6039 |
| VACE | 62.2883 |
| RGB$\leftrightarrow X$ | 62.3884 |
| RGBX-Next | **45.3871** |

The reported results indicate improved realism, lighting, texture synthesis, and adherence to supplied G-buffers. Video outputs are temporally coherent because the underlying model is a video DiT rather than an image model applied independently to each frame.

## 7. Limitations, significance, and future directions

RGBX-Next does not guarantee physical correctness. It can generate plausible but physically inconsistent materials, lighting, or geometry; inverse-rendered G-buffers are not uniquely correct scene reconstructions. Diffuse irradiance and direct lighting improve control but do not fully represent global illumination or general relighting. Real-data supervision depends on estimated G-buffers, so inverse-rendering errors can propagate into forward-rendering training.

The framework remains computationally expensive. It is no faster than the original Wan 2.1 14B model, and a 17-frame X-to-RGB segment requires approximately 400 seconds on one A100 with CFG. No diffusion distillation method such as DMD or DMD2 is applied. Streaming reduces fixed-context limitations but does not eliminate long-term drift, and the current system uses a single reference frame rather than richer 3D or video-level memory.

The unified RGB$\times X$ model is a proof of concept for changing input and output modality configurations through token embeddings. A complete treatment of arbitrary modality combinations remains future work. Additional challenges include improved physical consistency, richer long-term memory, more precise object-level and material-level partial control, faster inference, broader data, and improved lighting representations.

RGBX-Next’s principal methodological contribution is a general recipe for repurposing a pretrained DiT as a multimodal forward and inverse renderer. The recipe combines frame-wise concatenation of clean inputs and noisy outputs, QK type embeddings, timestep-zero clean tokens, X-patchify for multiple G-buffer inputs, modality dropout, spatial masking, blur-based conditioning, realistic-video supervision, and self-forced streaming. Its significance lies less in introducing a new diffusion objective than in redefining the latent-frame interface so that a video diffusion model can operate as a typed, spatially aligned renderer for heterogeneous G-buffer modalities.

Source: https://www.emergentmind.com/topics/rgbx-next