---
title: 'CL3AN: Ambient Lighting Normalization Benchmark'
url: https://www.emergentmind.com/topics/cl3an
type: topic
---

# CL3AN: Ambient Lighting Normalization Benchmark

Searching arXiv for recent papers on CL3AN and related ambient lighting normalization benchmarks.
arXiv search query: CL3AN ambient lighting normalization benchmark colored light sources.
CL3AN is a benchmark for color ambient lighting normalization under arbitrary, multi-colored illumination. It was introduced as the first large-scale, high-resolution dataset of its kind for restoring images captured under multiple colored light sources to ambient-normalized references, and it was subsequently used as the evaluation suite for CANDLE, a DINOv3-guided restoration model designed for the same regime [2508.02168][2604.02785]. The benchmark targets cases in which illumination-induced chromatic bias dominates the image formation process, including severe chromatic shifts, local color spill, specular highlight saturation, and material-dependent reflectance. In that sense, CL3AN occupies a distinct position relative to white-balance, low-light enhancement, and conventional image restoration benchmarks, because the central difficulty is disentangling illumination color from object-intrinsic appearance rather than merely correcting exposure or denoising.

## 1. Task definition and problem setting

CL3AN is organized around the task of color ambient lighting normalization: mapping an image captured under one or more RGB-tinted directional lights to a reference image acquired under uniform ambient illumination. The motivating premise is that practical illumination is inherently complex, involving colored light sources, occlusions, and diverse material interactions that produce intricate reflectance and shading effects, whereas existing methods often assume a single light source or uniform, white-balanced lighting [2508.02168].

The benchmark emphasizes conditions under which conventional geometric and low-level priors are insufficient. The reported failure factors include multi-colored directional sources, strong chromatic shifts and local color spill, specular highlights and saturation that often clip the brightest channel, and material-dependent reflectance spanning metallic, glossy, matte, conductive, dielectric, and transparent surfaces [2604.02785]. This suggests that CL3AN is intended not simply as a color-correction dataset, but as a stress test for models that must recover object-intrinsic color when the observed RGB is heavily biased by illumination.

A further distinction is that the benchmark supports both colored-light and white-light settings. In the original dataset description, each scene yields three registered images: a colored direct-light image, a white direct-light image, and an ambient reference. In the later CANDLE evaluation, the color-light input \(I\) and ambient ground truth \(Y_{GT}\) define the principal restoration pair for the color-light track [2508.02168][2604.02785].

## 2. Dataset composition, capture pipeline, and annotation structure

CL3AN contains 4,535 samples following the official split of 3,667 training triplets, 437 validation triplets, and 431 test triplets. The original dataset paper also specifies 105 distinct, cluttered tabletop scenes, with the 10 validation and 10 test scenes held out at the scene level [2508.02168]. The scenes span a wide range of materials and were designed to preserve difficult reflectance phenomena rather than normalize them away.

Each scene is acquired as a registered triplet:

| Component | Description |
|---|---|
| Colored direct lighting | Multi-RGB directional lights |
| White direct lighting | White-aligned directional lights |
| Ambient reference | Uniform diffuse white light |

Capture is performed with a Canon R6 MKII at full 24 MP RAW and demosaiced RGB. The camera configuration is specified as exposure \(1/60\) s, \(f/11\), ISO 100, white balance 6400 K, and 50 mm focal length. The ambient system uses five white softboxes at \(6400\ \text{K} \pm 5\%\) and 100% intensity, arranged around the scene. The direct system uses up to three programmable RGB fixtures with intensity range 30%–100%, hue range \(0^\circ\)–\(360^\circ\), and freely varied positions and orientations subject to at least one active light [2508.02168].

Ground-truth ambient references are acquired under the diffuse system with geometry-based filtering to suppress residual shadows. Pixel-perfect registration is guaranteed by a locked tripod and remote release. Calibration is performed via a 24-patch color checker under full white direct lighting to align the two lighting setups. The dataset also logs per-scene object lists, camera parameters, and precise light settings, including hue, intensity, and number of active lights [2508.02168].

Two resolution conventions appear in the literature. The source captures are approximately \(6000\times 4000\) (\(\approx 24\) MP), while the original RLN² benchmarking protocol uses \(1920\times 1440\) downsampled images. The later CANDLE protocol resizes images to \(1024\times 768\) for both training and inference [2508.02168][2604.02785]. Reported numbers are therefore protocol-specific.

## 3. Evaluation suite and formal metrics

CL3AN is paired with a restoration-oriented evaluation suite centered on PSNR, SSIM, and LPIPS, with FID additionally used for challenge evaluation. In the CANDLE formulation, the primary metrics are defined as follows [2604.02785].

Peak Signal-to-Noise Ratio:
$$
\mathrm{PSNR}(x,y)=10\cdot \log_{10}\!\left(\frac{L^2}{\mathrm{MSE}(x,y)}\right),
$$
where
$$
\mathrm{MSE}(x,y)=\frac{1}{N}\sum_i \lVert x_i-y_i\rVert^2
$$
and \(L\) is the maximum pixel value, with \(L=1\) after normalization.

Structural SIMilarity:
$$
\mathrm{SSIM}(x,y)=
\frac{(2\mu_x\mu_y+C_1)\cdot(2\sigma_{xy}+C_2)}
{(\mu_x^2+\mu_y^2+C_1)\cdot(\sigma_x^2+\sigma_y^2+C_2)},
$$
with \(\mu\) and \(\sigma\) denoting local means and variances, and \(C_1,C_2\) small stabilizing constants.

Learned Perceptual Image Patch Similarity:
$$
\mathrm{LPIPS}(x,y)=\mathrm{average}_p\ \lVert \phi(x)_p-\phi(y)_p\rVert_2
$$
over deep-network feature maps \(\phi\), as defined in Zhang et al. 2018.

Fréchet Inception Distance:
$$
\mathrm{FID}(P_r,P_g)=\lVert \mu_r-\mu_g\rVert^2 + \mathrm{Tr}\!\left(\Sigma_r+\Sigma_g-2(\Sigma_r\Sigma_g)^{1/2}\right),
$$
where \((\mu_r,\Sigma_r)\) and \((\mu_g,\Sigma_g)\) are the mean and covariance of pretrained Inception-v3 features for real and generated images.

The benchmark is also associated with concrete preprocessing protocols. In the CANDLE experiments, all images are resized to \(1024\times 768\), training uses random \(512\times 512\) crops with progressive growth up to \(768\times 576\), and augmentations include random horizontal and vertical flips together with random \(90^\circ\) rotations [2604.02785]. In RLN², benchmarking is carried out at \(1920\times 1440\), and training uses progressive patch training with standard geometric augmentations [2508.02168].

## 4. Retinex-based modeling on CL3AN: the RLN² framework

The first dedicated methodology built around CL3AN is RLN², introduced in “After the Party: Navigating the Mapping From Color to Ambient Lighting” [2508.02168]. RLN² is motivated by the observation that leading approaches on the benchmark produce illumination inconsistencies, texture leakage, and color distortion because they cannot precisely disentangle illumination from reflectance.

The model adopts a Retinex decomposition under paired images:
$$
\begin{aligned}
I(x,y) &= L(x,y)\cdot R(x,y),\\
I'(x,y) &= L'(x,y)\cdot R'(x,y),
\end{aligned}
$$
where \(L,L'\) denote per-pixel illumination and \(R,R'\) denote per-pixel reflectance. Restoration is then formulated through residuals \(\overline{L},\overline{R}\) such that
$$
\begin{aligned}
L(x,y) &= L'(x,y)+\overline{L}(x,y),\\
R(x,y) &= R'(x,y)+\overline{R}(x,y).
\end{aligned}
$$

A central design choice is explicit chromaticity and luminance guidance derived from HSV. The Value channel \(V\) guides the luminance branch, while Hue–Saturation \((H,S)\) guide the reflectance branch. RLN² uses two parallel streams, an \(L\)-branch for luminance residuals and an \(R\)-branch for reflectance residuals. Its encoder combines RGB downsampling with Haar DWT to extract low- and high-frequency maps, and its inner refinement blocks employ Cross-Domain Feature Fusion Attention (CDFFA) to inject HSV guidance [2508.02168].

At the decoder, RLN² uses high-frequency skip connections via inverse DWT, cross-attention fusion with a pretrained ConvNeXt feature extractor, channel attention, and final residual predictors \(\overline{L}\) and \(\overline{R}\). The reported model variants are RLN²-S, RLN²-Sf, RLN²-L, and RLN²-Lf, with RLN²-Lf defined as the large variant plus the frequency branch and reported at approximately 22.72 G MACs for a \(128\times 128\) patch [2508.02168].

Training uses a single-term reconstruction loss,
$$
\mathcal{L}_{\ell_1}=\left\lVert \widehat I - I_{\mathrm{ambient}}\right\rVert_1,
$$
with Adam, learning rate \(2\times 10^{-4}\), a cosine scheduler with two cycles, gradient clipping at 0.01, and progressive patch training on up to three NVIDIA L40 (48 GB) cards under PyTorch/CUDA 12.6 [2508.02168].

## 5. DINOv3-guided modeling on CL3AN: the CANDLE framework

CANDLE, introduced in “CANDLE: Illumination-Invariant Semantic Priors for Color Ambient Lighting Normalization,” is the second major method centered on CL3AN and is explicitly motivated by a representation-level observation: DINOv3 self-supervised features remain highly consistent between colored-light inputs and ambient-lit ground truth [2604.02785]. CANDLE uses this consistency as an illumination-robust semantic prior.

The core restoration network builds on the PromptNorm encoder–decoder. Four layers—6, 12, 18, and 24—of a frozen DINOv3 ViT-L/16 are extracted and injected into successive encoder stages through DINO Omni-layer Guidance (D.O.G.). The decoder then applies a bifurcated color–frequency refinement design, denoted BFACG + SFFB, to suppress decoder-side chromatic collapse and detail contamination [2604.02785].

The training protocol is staged. First, the full CANDLE model is trained under a progressive patch-size curriculum \(256\rightarrow 512\rightarrow 768\). Second, the main backbone is frozen and a lightweight NAFNet refiner is trained on the Stage-1 outputs. Third, the backbone and refiner are jointly fine-tuned end-to-end with the combined loss. Additional challenge-specific procedures include retrieval-based finetuning via DINO-CLIP similarity and the use of extra synthetic color-light data from Ambient6K, approximately 300 images [2604.02785].

Optimization uses Adam with weight decay \(=0\) and \(\beta=(0.9,0.999)\). The learning-rate schedule proceeds as \(1\times 10^{-4}\rightarrow 5\times 10^{-5}\rightarrow 2\times 10^{-5}\), with cosine annealing in later fine-tuning. Within the paper’s interpretation, D.O.G. provides global semantic consistency, so that object identity and material boundaries remain correctly colored even in strongly tinted or highlight-saturated regions, while BFACG + SFFB suppress residual chromatic collapse in specular areas and block contamination from illumination-biased encoder features [2604.02785].

## 6. Benchmark results, qualitative behavior, and nomenclature

Quantitatively, RLN² established the initial dedicated baseline on CL3AN at \(1920\times 1440\). On the test set, the unprocessed input scores 10.84 dB PSNR, 0.447 SSIM, and 0.518 LPIPS; IFBlend reaches 20.37 dB, 0.720, and 0.228; and RLN²-Lf reaches 20.523 dB, 0.746, and 0.208. On AMBIENT6K, RLN²-Lf similarly improves over IFBlend while using 22.72 G MACs versus IFBlend’s 26.01 G, approximately 15% lower [2508.02168].

Under the later \(1024\times 768\) CL3AN protocol used by CANDLE, the comparison includes both general restoration networks and earlier ambient-lighting normalization methods:

| Method | PSNR / SSIM / LPIPS | MACs |
|---|---|---|
| NAFNet | 18.55 / 0.6121 / 0.3804 | 4.04 G |
| SFNet | 16.78 / 0.5489 / 0.6183 | 30.59 G |
| Uformer | 17.79 / 0.6815 / 0.3437 | 20.89 G |
| Restormer | 18.24 / 0.6570 / 0.3541 | 35.31 G |
| HINet | 17.73 / 0.6035 / 0.3887 | 39.46 G |
| IFBlend | 19.42 / 0.7581 / 0.2505 | 24.65 G |
| Retinexformer | 18.82 / 0.6943 / 0.3167 | 3.96 G |
| RLN2-Lf | 19.85 / 0.7436 / 0.2569 | 21.93 G |
| PromptNorm | 19.22 / 0.7485 / 0.2590 | 20.81 G |
| CANDLE | 21.07 / 0.7788 / 0.2325 | 23.52 G |

On this protocol, CANDLE reports a 21.07 dB PSNR, which is a \(+1.22\) dB gain over the strongest prior method, RLN2-Lf. In the NTIRE 2026 ambient-lighting challenge setting, CANDLE additionally reports FID 88.20 on the color-lighting track, corresponding to third-place fidelity ranking, and 49.84 on the white-lighting track, which is the lowest FID overall; the paper also states that the method achieved 2nd place in fidelity on the White Lighting track [2604.02785].

The qualitative observations reported across the two papers are consistent. General restoration backbones such as NAFNet and Restormer often produce washed-out, desaturated results under strong chromatic shifts because they lack a mechanism to disentangle illumination color from object reflectance. IFBlend can recover detail through frequency-domain fusion, but it cannot reliably infer true material color when the RGB input is heavily biased. PromptNorm’s surface-normal guidance improves shape and shading, yet still conflates local color shifts with reflectance under multi-colored lighting. RLN² is described as correcting color-bleeding and shadow spill while preserving fine texture in highlights and deep shadows, and CANDLE is described as largely overcoming incomplete correction in the hardest highlight regions and material-dependent color inconsistency through illumination-invariant DINOv3 features [2508.02168][2604.02785].

A recurrent point of confusion is nomenclature. In image restoration, CL3AN denotes the ambient-lighting benchmark described above. A separate 2026 graph-learning paper uses the similar acronym CL³AN-GNN for “Curriculum-Guided Feature Learning and Three-Stage Attention Network” in imbalanced node classification; that usage is unrelated to the ambient-lighting dataset and restoration benchmark [2602.03808].

Source: https://www.emergentmind.com/topics/cl3an