---
title: Neural Six-Way Lightmaps
url: https://www.emergentmind.com/topics/neural-six-way-lightmaps-cf290ead-db18-482a-80e8-e46f8d48e922
type: topic
---

# Neural Six-Way Lightmaps

Neural six-way lightmaps are a data-driven framework for real-time rendering of participating media, such as smoke, that synthesizes directional volumetric lighting with fidelity approaching path tracing while retaining the efficiency, interactivity, and integration of precomputed lightmap techniques. The approach leverages a neural network to infer six directional lightmaps from a coarse, view-aligned screen-space representation derived via approximate ray marching in the volume, enabling dynamic effects that support arbitrary camera movement, lighting changes, and real-time media simulation in interactive applications [2604.03748].

## 1. Volume Rendering Foundations and Six-way Lightmaps

Volumetric participating media are governed by the volume rendering equation (VRE). For a point $\mathbf{x}$ and outgoing direction $\omega_o$, the VRE is
\[
L(\mathbf{x},\omega_o) = \int_{0}^{d} T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(s)\bigr) \bigl[ \sigma_a(\mathbf{y}(s))\,L_e(\mathbf{y}(s),\omega_o) + \sigma_s(\mathbf{y}(s))\,L_s(\mathbf{y}(s),\omega_o) \bigr]\, ds + T\bigl(\mathbf{x}\!\leftrightarrow\!\mathbf{y}(d)\bigr)\, L_o(\mathbf{y}(d),\omega_o)
\]
with $\sigma_a$, $\sigma_s$ as absorption and scattering coefficients, $T$ transmittance, $L_e$ emission, and $L_s$ in-scattered radiance (see data for full definitions). Direct Monte Carlo evaluation of the VRE is costly due to the need for many samples per pixel. 

Conventional six-way lightmaps approximate volumetric effects by precomputing six 3D lightmaps aligned with the Cartesian axes (“front/back”, “left/right”, “up/down”), which are blended in real-time proportional to local lighting direction. This approach is widely adopted for games but is limited to pre-simulated sequences and static view assumptions, yielding inflexible results and suboptimal integration with dynamic, user-driven interactions.

## 2. Screen-space Guiding Map Generation

To enable a neural predictor to condition on dynamic volumetric geometry and view, a screen-space guiding map $\mathcal{G}(x)\in\mathbb{R}^3$ is computed per pixel $x$ through coarse, forward ray marching along the camera view direction $\omega_v$. Key components calculated at each ray-marching step are:

- **Transmittance:** 
  \[
  T_n = T_{n-1}\,\exp\!\bigl(-\sigma_t(\mathbf{x}_n)\,\Delta t\bigr),\quad T_0=1
  \]
- **Approximate in-scattered radiance:** A weighted sum for primary axes (“front,” “top,” “bottom”) with large steps $\Delta t$ and threshold-based accept/reject sampling.
- **Layer Depth Estimation:** Depth $D=\min_{n}\{ n\Delta t:\sigma_s(\mathbf{x}_n)>\tau \}$, indicating entry into the dense phase of the volume.

The resulting guiding map stacks $\left(\tilde L_{\text{scatt}},T,D\right)$, representing coarse 1st-order scattering, transmission, and silhouette structure.

## 3. Neural Network Architecture

Given the guiding map $\mathcal{G}$, a neural network models the mapping 

\[
\{\tilde L_{\rm scatt},T,D\} \xrightarrow{f_\theta} \left(\hat L_x^+,\hat L_x^-,\hat L_y^+,\hat L_y^-,\hat L_z^+,\hat L_z^-\right)
\]
where each $\hat L_i$ approximates the light integral along cartesian directions. The architecture consists of:

- **Input:** $512\times512\times3$ guiding map.
- **Backbone:** U-Net encoder–decoder, 4 down/up stages, 2×(Conv–ReLU–BatchNorm) per stage; bottleneck at $32\times32\times512$.
- **Channel Adapters:** Four lateral heads (3 NAFBlocks each; 1×1 then 3×3 convolutions with SwiGLU) group outputs by axis pairs.
- **Output:** $512\times512\times6$ tensor of six-way lightmaps.

Channel adapters enable the architecture to efficiently model correlations among the grouped axes and output features relevant to per-direction lighting synthesis.

## 4. Training Methodology and Loss Formulation

Training data is generated from 14 smoke sequences (400³ grid, 200 frames each). For each frame and 9 camera azimuths, high-fidelity ground-truth lightmaps $\mathcal{L}$ and guiding maps $\mathcal{G}$ are rendered in Houdini Karma using Monte Carlo (512 spp, single bounce). The resulting dataset comprises 25,200 examples.

The network is trained to minimize a compound loss:
\[
\mathcal{L}(\theta)   =   \|\hat{\mathcal{L}}-\mathcal{L}\|_2^2
+   \lambda_{\rm perc}\, \|\phi(\hat{\mathcal{L}})-\phi(\mathcal{L})\|_1 
+   \lambda_{\rm flow}\, \|\mathrm{FlowNet}(\hat{\mathcal{L}}^{t-1},\hat{\mathcal{L}}^t) -\mathrm{FlowNet}(\mathcal{L}^{t-1},\mathcal{L}^t)\|_1
\]
where $\phi$ denotes VGG feature extraction for perceptual similarity and the FlowNet term promotes temporal coherence. Coefficients are $\lambda_{\rm perc} = 0.1$, $\lambda_{\rm flow} = 0.1$. Optimization uses Adam (lr $10^{-3}$, batch 12, 200 epochs, $\sim$60 hours on a single GPU).

## 5. Pipeline Integration and Rendering Process

At runtime, the pipeline executes the following stages:

1. **Simulation:** 3D LBM velocity (4 ms) and density advection (11.5 ms).
2. **Guiding Map Generation:** Coarse ray march at $512^2$ (2 ms).
3. **Neural Inference:** TensorRT-accelerated U-Net evaluation (1.2 ms).
4. **Shading:** Camera-facing billboard samples $\hat{\mathcal{L}}(x)$ at each fragment (0.3 ms).
5. **Shadowing:** Screen-space depth test compares guiding map depth $D(x)$ to obstacle shadow map in light space (0.8 ms).

The resulting pixel color under arbitrary light direction $\omega_l$ is computed as:
\[
C(x) = \sum_{p \in \{x, y, z\}} |\omega_{l, p}|\, \hat L_{p}^{\mathrm{sign}(\omega_{l,p})}(x) + T(x)\, L_{\rm em}(x)
\]
where $\hat L_{p}^{\pm}$ are the directional lightmaps, $T$ is predicted transparency, and $L_{\rm em}$ is an optional emission component.

## 6. Benchmarks: Performance, Fidelity, and Comparisons

Comparisons to existing methods demonstrate the efficiency and visual quality of neural six-way lightmaps:

| Method                     | Real-time | PSNR (↑) | Time (ms ↓) | Memory          |
|----------------------------|:---------:|----------|-------------|-----------------|
| Traditional 6-way [Müller] |  yes      | 27.5 dB  | 0.5         | 2×512² × RGB    |
| Volumetric ReSTIR (1 spp)  |  no       | 26.3 dB  | 10.4        | –               |
| ReSTIR + denoise           |  no       | 34.8 dB  | 12.6        | –               |
| MRPNN                      |  no       | 36.1 dB  | 184.0       | –               |
| Ours (neural 6-way)        |  yes      | 40.7 dB  | 3.9         | 512² × 6 ch     |

Key metrics:
- Full pipeline runs at $\approx$20 ms per frame (50 FPS), with 3.9 ms attributed to lighting.
- Visual fidelity measured by PSNR: +4 dB over denoised ReSTIR, +4.6 dB over traditional lightmaps.
- Memory footprint: 6-channel lightmap per pixel.

## 7. Limitations and Future Research Directions

Limitations include:
- **Shadow singularities:** Screen-space depth test yields under-shadowing beneath complex geometry due to absence of complete volumetric shadow propagation.
- **Distribution shift:** For smoke densities scaled by $0.5 \times$ or $2 \times$, PSNR degrades by $2-5$ dB but plausibility is preserved.

Identified avenues for future research:
- Integration of higher-order scattering via learned residual corrections.
- Generalization to area light sources or environment lighting.
- Improved dynamic geometry shadowing through volumetric carving in light space.
- Extension to other participating media (clouds, fire with anisotropic phase functions) [2604.03748]. 

Neural six-way lightmaps thus provide an efficient and scalable solution for realistic, interactive volumetric rendering in dynamic virtual environments, bridging the gap between traditional precomputed lightmaps and computationally intensive path-traced methods.

Source: https://www.emergentmind.com/topics/neural-six-way-lightmaps-cf290ead-db18-482a-80e8-e46f8d48e922