- The paper presents WildRelight, the first dataset designed for real-world single-image relighting using precise spatial/temporal alignment and HDR environment maps.
- The proposed physics-guided adaptation framework leverages Diffusion Posterior Sampling and sampling-aware temporal test-time adaptation to achieve near-supervised performance without retraining.
- Experimental results reveal significant domain shifts in synthetic-trained models, underscoring the need for real-world illumination signals in inverse rendering.
WildRelight: Real-World Benchmark and Physics-Guided Domain Adaptation for Single-Image Relighting
Motivation and Domain Gap Analysis
Single-image relighting—the manipulation of illumination in a solitary photograph—underpins diverse applications in computational photography, AR, and cinematic content creation. Contemporary models, particularly latent diffusion-based generative architectures, achieve high photorealism on synthetic datasets by decomposing images into albedo, geometry, and illumination components, enabling accurate rerendering under arbitrary lights. However, these synthetic benchmarks are insufficient proxies for the physical complexity of natural scenes. They systematically omit atmospheric scattering, indirect light, and the non-ideal spectral and materials variability inherent in real-world environments. This induces severe domain shifts; existing inverse rendering and relighting models generalize poorly beyond such datasets.
WildRelight is introduced to directly address this gap: it is distinguished as the first dataset designed for in-the-wild, real-world single-image relighting evaluation, prioritizing strict spatial/temporal pixel alignment and radiometric accuracy. High-resolution outdoor images are paired with co-located, physically faithful HDR environment maps, capturing illumination at multiple times of day per scene, thus encompassing complex environmental variability. This setup enables rigorous benchmarking and exposes the limitations of current synthetic-trained models in real deployment.

Figure 1: Example image/environment map pairs from WildRelight, illustrating strict spatial and radiometric alignment across temporal lighting variations.
Dataset Design and Acquisition Protocol
WildRelight comprises 30 scenes with 5–7 illumination variants each, sampled under the immutable sun's trajectory, demanding hours-long monitoring rather than rapid active lighting. To guarantee scene/envmap correspondence, a dual-camera system (Sony A7 for scenes, Insta360 Pro 2 for envmaps) is meticulously co-located, aligning optical centers via nodal point calibration and minimizing temporal delays (<1 minute in most cases).
Color calibration is performed across both sensors using X-Rite ColorChecker targets, ensuring radiometric consistency. All imagery is captured and stored in 16-bit linear RAW, bypassing non-linear CRF effects and enabling direct HDR synthesis through exposure merging, maximizing shadow and highlight retention.
Dynamic scene elements (e.g., foliage, clouds) are rigorously masked via manually annotated binary regions, rather than algorithmic warping, preserving ground truth photometric integrity and allowing for selective evaluation excluding non-static areas.

Figure 2: Challenging scenario examples from WildRelight, including high-complexity glass, foliage, and reflective surfaces, each temporally sampled for illumination evolution.
Physics-Guided Adaptation and Test-Time Inference Framework
WildRelight's spatiotemporally aligned design enables a new paradigm: domain adaptation via self-supervised constraints. The reference framework integrates:
- Diffusion Posterior Sampling (DPS): Physical regularization of G-buffer predictions occurs through differentiable rendering, enforcing consistency between rendered outputs (using the Cook–Torrance model) and observed images under ground-truth HDR illumination. This guides the diffusion trajectory during inference without retraining, constraining generative priors to the actual scene structure.
- Sampling-Aware Temporal Test-Time Adaptation (TTA): Temporal self-supervision is achieved through a leave-one-out protocol—N−1 captured illuminations serve as adaptation anchors, with the model optimized for the held-out lighting. Only lightweight LoRA modules within the diffusion UNet are updated, aligning model statistics to scene-specific light transport characteristics while mitigating overfitting.
Together, these mechanisms transform the sim-to-real adaptation challenge into an instance-specific, physically grounded task, leveraging WildRelight's temporal structure.
Experimental Benchmarking and Strong Results
Zero-shot evaluation of pretrained SOTA models—DiffusionRenderer, RGB↔X, Materialist—demonstrates significant domain degradation (<16 dB PSNR for synthetic-trained methods), inability to reproduce high-frequency shadowing, and clear failures under complex outdoor illumination. Materialist's optimization-centric pipeline fares better with known environment maps but struggles with complex materials and geometric artifacts.
Supervised global finetuning of DiffusionRenderer on WildRelight's training scenes yields a substantial improvement (23.28 dB → 25.95 dB PSNR), validating the dataset's critical bridging role and empirically quantifying the domain gap.
Qualitative evaluations further highlight the limitations of synthetic priors: diffusion-based approaches post-finetuning produce outputs more faithful to real world brightness and shadowing, while optimization-based baselines often falter under geometric complexity or diverse foliage.

Figure 3: Qualitative comparison on relighting: zero-shot models fail on brightness/structure; finetuning enables accurate reproduction of GT illumination.
The physics-guided DPS and temporal TTA framework achieves near-supervised performance at inference time (25.04 dB PSNR, 0.345 LPIPS), approaching global finetuning results without retraining. Ablation analysis reveals that TTA alone can overfit photometric intensities at the expense of perceptual realism, while DPS regularizes adaptation, preserving physical plausibility. The synergistic combination outperforms both naïve adaptation and unconstrained optimization, substantiating the efficacy of scene-intrinsic constraints.

Figure 4: Visualization of ablation study: DPS regularizes shadow realism, TTA improves alignment with temporal lighting, combined adaptation yields best fidelity.
Practical and Theoretical Implications
WildRelight establishes a new standard for single-image relighting evaluation, offering uniquely aligned, high-fidelity supervision for robust domain transfer. Its temporal structure enables self-supervised, physics-constrained adaptation, dramatically reducing dependence on expensive synthetic training or manual annotation.
Empirically, the dataset exposes the necessity for real-world domain signals: synthetic-only models are fundamentally inadequate for deployment. The physics-guided adaptation framework demonstrates that leveraging temporal illumination evolution—in the form of real environment maps—enables tractable instance-specific alignment, suggesting feasible pathways for robust inverse rendering in unconstrained environments.
Practically, WildRelight will catalyze physically grounded relighting research beyond controlled laboratory or synthetic datasets, facilitating benchmarking and adaptation of generative models, optimization-based approaches, and hybrid techniques. The dataset's rigorous protocol also paves the way for future advances in accurate outdoor light synthesis, environment-aware augmentation, and AR/VR integration.
Future Directions
Progress will depend on further modeling dynamic elements, advanced scene-based adaptation (e.g., leveraging video data or spatial context), and optimizing computational efficiency in test-time adaptation. Extending the protocol to indoor/outdoor transitions, more sophisticated masking, or integrating direct sensor data (e.g., LiDAR) could broaden its applicability. Ultimately, WildRelight frames the challenge of real-world relighting as a physically tractable, data-rich, and self-supervised problem, propelling the next generation of robust inverse rendering algorithms.

Figure 5: Additional samples from WildRelight, illustrating diversity and complexity of scenes and illumination conditions.
Conclusion
WildRelight, through its precise alignment and temporal sampling, provides an actionable bridge across the sim-to-real gap in single-image relighting. The integration of physics-guided inference and temporal self-supervised adaptation achieves near-supervised performance without retraining, substantiating both the dataset's benchmark value and the feasibility of robust domain transfer. WildRelight is poised to advance physically grounded relighting research and real-world deployment, enabling models to effectively navigate the unconstrained complexity of natural illumination and materials.