Papers
Topics
Authors
Recent
Search
2000 character limit reached

Neural Six-Way Lightmaps

Updated 2 July 2026
  • The paper introduces a neural network framework that infers six directional lightmaps from coarse screen-space guiding maps, enabling real-time volumetric rendering of smoke.
  • The approach integrates a U-Net architecture with specialized channel adapters to efficiently capture directional lighting cues and dynamic media variations.
  • Results show a PSNR of 40.7 dB with a 3.9 ms lighting overhead, outperforming traditional methods while maintaining interactive performance.

Neural six-way lightmaps are a data-driven framework for real-time rendering of participating media, such as smoke, that synthesizes directional volumetric lighting with fidelity approaching path tracing while retaining the efficiency, interactivity, and integration of precomputed lightmap techniques. The approach leverages a neural network to infer six directional lightmaps from a coarse, view-aligned screen-space representation derived via approximate ray marching in the volume, enabling dynamic effects that support arbitrary camera movement, lighting changes, and real-time media simulation in interactive applications (Li et al., 4 Apr 2026).

1. Volume Rendering Foundations and Six-way Lightmaps

Volumetric participating media are governed by the volume rendering equation (VRE). For a point x\mathbf{x} and outgoing direction ωo\omega_o, the VRE is

L(x,ωo)=∫0dT(x ⁣↔ ⁣y(s))[σa(y(s)) Le(y(s),ωo)+σs(y(s)) Ls(y(s),ωo)] ds+T(x ⁣↔ ⁣y(d)) Lo(y(d),ωo)L(\mathbf{x},\omega_o) = \int_{0}^{d} T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(s)\bigr) \bigl[ \sigma_a(\mathbf{y}(s))\,L_e(\mathbf{y}(s),\omega_o) + \sigma_s(\mathbf{y}(s))\,L_s(\mathbf{y}(s),\omega_o) \bigr]\, ds + T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(d)\bigr)\, L_o(\mathbf{y}(d),\omega_o)

with σa\sigma_a, σs\sigma_s as absorption and scattering coefficients, TT transmittance, LeL_e emission, and LsL_s in-scattered radiance (see data for full definitions). Direct Monte Carlo evaluation of the VRE is costly due to the need for many samples per pixel.

Conventional six-way lightmaps approximate volumetric effects by precomputing six 3D lightmaps aligned with the Cartesian axes (“front/back”, “left/right”, “up/down”), which are blended in real-time proportional to local lighting direction. This approach is widely adopted for games but is limited to pre-simulated sequences and static view assumptions, yielding inflexible results and suboptimal integration with dynamic, user-driven interactions.

2. Screen-space Guiding Map Generation

To enable a neural predictor to condition on dynamic volumetric geometry and view, a screen-space guiding map G(x)∈R3\mathcal{G}(x)\in\mathbb{R}^3 is computed per pixel xx through coarse, forward ray marching along the camera view direction ωo\omega_o0. Key components calculated at each ray-marching step are:

  • Transmittance:

ωo\omega_o1

  • Approximate in-scattered radiance: A weighted sum for primary axes (“front,” “top,” “bottom”) with large steps ωo\omega_o2 and threshold-based accept/reject sampling.
  • Layer Depth Estimation: Depth ωo\omega_o3, indicating entry into the dense phase of the volume.

The resulting guiding map stacks ωo\omega_o4, representing coarse 1st-order scattering, transmission, and silhouette structure.

3. Neural Network Architecture

Given the guiding map ωo\omega_o5, a neural network models the mapping

ωo\omega_o6

where each ωo\omega_o7 approximates the light integral along cartesian directions. The architecture consists of:

  • Input: ωo\omega_o8 guiding map.
  • Backbone: U-Net encoder–decoder, 4 down/up stages, 2×(Conv–ReLU–BatchNorm) per stage; bottleneck at ωo\omega_o9.
  • Channel Adapters: Four lateral heads (3 NAFBlocks each; 1×1 then 3×3 convolutions with SwiGLU) group outputs by axis pairs.
  • Output: L(x,ωo)=∫0dT(x ⁣↔ ⁣y(s))[σa(y(s)) Le(y(s),ωo)+σs(y(s)) Ls(y(s),ωo)] ds+T(x ⁣↔ ⁣y(d)) Lo(y(d),ωo)L(\mathbf{x},\omega_o) = \int_{0}^{d} T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(s)\bigr) \bigl[ \sigma_a(\mathbf{y}(s))\,L_e(\mathbf{y}(s),\omega_o) + \sigma_s(\mathbf{y}(s))\,L_s(\mathbf{y}(s),\omega_o) \bigr]\, ds + T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(d)\bigr)\, L_o(\mathbf{y}(d),\omega_o)0 tensor of six-way lightmaps.

Channel adapters enable the architecture to efficiently model correlations among the grouped axes and output features relevant to per-direction lighting synthesis.

4. Training Methodology and Loss Formulation

Training data is generated from 14 smoke sequences (400³ grid, 200 frames each). For each frame and 9 camera azimuths, high-fidelity ground-truth lightmaps L(x,ωo)=∫0dT(x ⁣↔ ⁣y(s))[σa(y(s)) Le(y(s),ωo)+σs(y(s)) Ls(y(s),ωo)] ds+T(x ⁣↔ ⁣y(d)) Lo(y(d),ωo)L(\mathbf{x},\omega_o) = \int_{0}^{d} T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(s)\bigr) \bigl[ \sigma_a(\mathbf{y}(s))\,L_e(\mathbf{y}(s),\omega_o) + \sigma_s(\mathbf{y}(s))\,L_s(\mathbf{y}(s),\omega_o) \bigr]\, ds + T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(d)\bigr)\, L_o(\mathbf{y}(d),\omega_o)1 and guiding maps L(x,ωo)=∫0dT(x ⁣↔ ⁣y(s))[σa(y(s)) Le(y(s),ωo)+σs(y(s)) Ls(y(s),ωo)] ds+T(x ⁣↔ ⁣y(d)) Lo(y(d),ωo)L(\mathbf{x},\omega_o) = \int_{0}^{d} T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(s)\bigr) \bigl[ \sigma_a(\mathbf{y}(s))\,L_e(\mathbf{y}(s),\omega_o) + \sigma_s(\mathbf{y}(s))\,L_s(\mathbf{y}(s),\omega_o) \bigr]\, ds + T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(d)\bigr)\, L_o(\mathbf{y}(d),\omega_o)2 are rendered in Houdini Karma using Monte Carlo (512 spp, single bounce). The resulting dataset comprises 25,200 examples.

The network is trained to minimize a compound loss: L(x,ωo)=∫0dT(x ⁣↔ ⁣y(s))[σa(y(s)) Le(y(s),ωo)+σs(y(s)) Ls(y(s),ωo)] ds+T(x ⁣↔ ⁣y(d)) Lo(y(d),ωo)L(\mathbf{x},\omega_o) = \int_{0}^{d} T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(s)\bigr) \bigl[ \sigma_a(\mathbf{y}(s))\,L_e(\mathbf{y}(s),\omega_o) + \sigma_s(\mathbf{y}(s))\,L_s(\mathbf{y}(s),\omega_o) \bigr]\, ds + T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(d)\bigr)\, L_o(\mathbf{y}(d),\omega_o)3 where L(x,ωo)=∫0dT(x ⁣↔ ⁣y(s))[σa(y(s)) Le(y(s),ωo)+σs(y(s)) Ls(y(s),ωo)] ds+T(x ⁣↔ ⁣y(d)) Lo(y(d),ωo)L(\mathbf{x},\omega_o) = \int_{0}^{d} T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(s)\bigr) \bigl[ \sigma_a(\mathbf{y}(s))\,L_e(\mathbf{y}(s),\omega_o) + \sigma_s(\mathbf{y}(s))\,L_s(\mathbf{y}(s),\omega_o) \bigr]\, ds + T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(d)\bigr)\, L_o(\mathbf{y}(d),\omega_o)4 denotes VGG feature extraction for perceptual similarity and the FlowNet term promotes temporal coherence. Coefficients are L(x,ωo)=∫0dT(x ⁣↔ ⁣y(s))[σa(y(s)) Le(y(s),ωo)+σs(y(s)) Ls(y(s),ωo)] ds+T(x ⁣↔ ⁣y(d)) Lo(y(d),ωo)L(\mathbf{x},\omega_o) = \int_{0}^{d} T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(s)\bigr) \bigl[ \sigma_a(\mathbf{y}(s))\,L_e(\mathbf{y}(s),\omega_o) + \sigma_s(\mathbf{y}(s))\,L_s(\mathbf{y}(s),\omega_o) \bigr]\, ds + T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(d)\bigr)\, L_o(\mathbf{y}(d),\omega_o)5, L(x,ωo)=∫0dT(x ⁣↔ ⁣y(s))[σa(y(s)) Le(y(s),ωo)+σs(y(s)) Ls(y(s),ωo)] ds+T(x ⁣↔ ⁣y(d)) Lo(y(d),ωo)L(\mathbf{x},\omega_o) = \int_{0}^{d} T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(s)\bigr) \bigl[ \sigma_a(\mathbf{y}(s))\,L_e(\mathbf{y}(s),\omega_o) + \sigma_s(\mathbf{y}(s))\,L_s(\mathbf{y}(s),\omega_o) \bigr]\, ds + T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(d)\bigr)\, L_o(\mathbf{y}(d),\omega_o)6. Optimization uses Adam (lr L(x,ωo)=∫0dT(x ⁣↔ ⁣y(s))[σa(y(s)) Le(y(s),ωo)+σs(y(s)) Ls(y(s),ωo)] ds+T(x ⁣↔ ⁣y(d)) Lo(y(d),ωo)L(\mathbf{x},\omega_o) = \int_{0}^{d} T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(s)\bigr) \bigl[ \sigma_a(\mathbf{y}(s))\,L_e(\mathbf{y}(s),\omega_o) + \sigma_s(\mathbf{y}(s))\,L_s(\mathbf{y}(s),\omega_o) \bigr]\, ds + T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(d)\bigr)\, L_o(\mathbf{y}(d),\omega_o)7, batch 12, 200 epochs, L(x,ωo)=∫0dT(x ⁣↔ ⁣y(s))[σa(y(s)) Le(y(s),ωo)+σs(y(s)) Ls(y(s),ωo)] ds+T(x ⁣↔ ⁣y(d)) Lo(y(d),ωo)L(\mathbf{x},\omega_o) = \int_{0}^{d} T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(s)\bigr) \bigl[ \sigma_a(\mathbf{y}(s))\,L_e(\mathbf{y}(s),\omega_o) + \sigma_s(\mathbf{y}(s))\,L_s(\mathbf{y}(s),\omega_o) \bigr]\, ds + T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(d)\bigr)\, L_o(\mathbf{y}(d),\omega_o)860 hours on a single GPU).

5. Pipeline Integration and Rendering Process

At runtime, the pipeline executes the following stages:

  1. Simulation: 3D LBM velocity (4 ms) and density advection (11.5 ms).
  2. Guiding Map Generation: Coarse ray march at L(x,ωo)=∫0dT(x ⁣↔ ⁣y(s))[σa(y(s)) Le(y(s),ωo)+σs(y(s)) Ls(y(s),ωo)] ds+T(x ⁣↔ ⁣y(d)) Lo(y(d),ωo)L(\mathbf{x},\omega_o) = \int_{0}^{d} T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(s)\bigr) \bigl[ \sigma_a(\mathbf{y}(s))\,L_e(\mathbf{y}(s),\omega_o) + \sigma_s(\mathbf{y}(s))\,L_s(\mathbf{y}(s),\omega_o) \bigr]\, ds + T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(d)\bigr)\, L_o(\mathbf{y}(d),\omega_o)9 (2 ms).
  3. Neural Inference: TensorRT-accelerated U-Net evaluation (1.2 ms).
  4. Shading: Camera-facing billboard samples σa\sigma_a0 at each fragment (0.3 ms).
  5. Shadowing: Screen-space depth test compares guiding map depth σa\sigma_a1 to obstacle shadow map in light space (0.8 ms).

The resulting pixel color under arbitrary light direction σa\sigma_a2 is computed as: σa\sigma_a3 where σa\sigma_a4 are the directional lightmaps, σa\sigma_a5 is predicted transparency, and σa\sigma_a6 is an optional emission component.

6. Benchmarks: Performance, Fidelity, and Comparisons

Comparisons to existing methods demonstrate the efficiency and visual quality of neural six-way lightmaps:

Method Real-time PSNR (↑) Time (ms ↓) Memory
Traditional 6-way [Müller] yes 27.5 dB 0.5 2×512² × RGB
Volumetric ReSTIR (1 spp) no 26.3 dB 10.4 –
ReSTIR + denoise no 34.8 dB 12.6 –
MRPNN no 36.1 dB 184.0 –
Ours (neural 6-way) yes 40.7 dB 3.9 512² × 6 ch

Key metrics:

  • Full pipeline runs at σa\sigma_a720 ms per frame (50 FPS), with 3.9 ms attributed to lighting.
  • Visual fidelity measured by PSNR: +4 dB over denoised ReSTIR, +4.6 dB over traditional lightmaps.
  • Memory footprint: 6-channel lightmap per pixel.

7. Limitations and Future Research Directions

Limitations include:

  • Shadow singularities: Screen-space depth test yields under-shadowing beneath complex geometry due to absence of complete volumetric shadow propagation.
  • Distribution shift: For smoke densities scaled by σa\sigma_a8 or σa\sigma_a9, PSNR degrades by σs\sigma_s0 dB but plausibility is preserved.

Identified avenues for future research:

  • Integration of higher-order scattering via learned residual corrections.
  • Generalization to area light sources or environment lighting.
  • Improved dynamic geometry shadowing through volumetric carving in light space.
  • Extension to other participating media (clouds, fire with anisotropic phase functions) (Li et al., 4 Apr 2026).

Neural six-way lightmaps thus provide an efficient and scalable solution for realistic, interactive volumetric rendering in dynamic virtual environments, bridging the gap between traditional precomputed lightmaps and computationally intensive path-traced methods.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Neural Six-Way Lightmaps.