Papers
Topics
Authors
Recent
Search
2000 character limit reached

LightCrafter: PBR-Conditioned Video Diffusion Refinement for Controllable and Consistent Relighting

Published 9 Jul 2026 in cs.CV and cs.GR | (2607.08016v2)

Abstract: Video relighting requires balancing long-form temporal consistency with a physically grounded understanding of light transport, which depends on accurate estimation of intrinsic scene properties such as materials, geometry, and illumination. Existing methods follow two paradigms: (1) reconstruct a video's photometric properties via inverse rendering and relight them to a target illumination via forward rendering, using physically-based rendering (PBR) or a neural renderer; these suffer from noisy reconstructions and struggle with hard-to-model effects such as global illumination. (2) Frame the task as generative video-to-video translation conditioned on relighting targets (a target environment map or text); this limits relighting control and temporal stability, since diffusion models struggle to translate long-form videos, and is constrained by the availability of input/relit training pairs. We propose LightCrafter, a hybrid pipeline that reformulates video relighting as video translation of a proxy video: rather than translating the input video directly to the target, we translate a PBR rendering of the input under the target illumination to the final target. This bakes illumination targets into the PBR proxy, removing the need to teach the diffusion model illumination concepts like environment maps, and enables more intricate lighting control while naturally providing long-form temporal consistency. We show PBR renders alone already outperform some prior art but struggle with effects like global illumination; to capture these, we leverage photometric priors in video generation models by post-training CogVideoX on synthetic video pairs and real-world unpaired videos. We outperform prior state-of-the-art on existing real-world relighting benchmarks and contribute a synthetic benchmark for further analysis. We will release our dataset, benchmark, metrics, and code.

Summary

  • The paper introduces a pipeline that first renders physically-based proxy videos and then refines them with a diffusion model for artifact correction.
  • It achieves high temporal consistency and controlled illumination, evidenced by improved PSNR, SSIM, and LPIPS metrics on both synthetic and real-world data.
  • The method also facilitates downstream scene editing by allowing manipulation of lighting, materials, and camera trajectories without retraining.

LightCrafter: PBR-Conditioned Video Diffusion Refinement for Controllable and Consistent Relighting

Motivation and Context

Video relighting demands temporally consistent and physically plausible scene modification, contingent on accurate estimation of geometry, materials, and illumination. Prior art either pursues explicit inverse rendering followed by physically-based or neural forward rendering, or trains generative video-to-video models conditioned on semantic or environment map signals. Both paradigms exhibit deficiencies: explicit pipelines propagate reconstruction artifacts and cannot model effects such as global illumination, while generative diffusions lack precise lighting control and struggle with long-form consistency.

LightCrafter reframes the task as PBR-proxy video translation: rather than directly translating input to relit output, the pipeline renders a physically-based proxy video under the target illumination, then employs a video diffusion refiner to photorealistically harmonize artifacts. This design leverages explicit illumination control and temporal determinism inherent to PBR rendering, while exploiting photometric refinement capabilities of diffusion models. The pipeline's key insight is that PBR proxies encode most relighting structure and temporal coherence, which allows the diffusion model to focus exclusively on artifact removal.

Figure 1

Figure 1: LightCrafter formulates controllable video relighting as PBR-proxy rendering refinement, combining physically-based rendering for temporal and lighting control with a diffusion model for photorealistic artifact correction.

Methodological Overview

The system operates in three phases:

  1. Inverse Rendering: Scene intrinsics are recovered via state-of-the-art estimators, including DiffusionRenderer for photometric properties, DiffusionLight for illumination, and MegaSAM for geometry. Explicit world coordinate recovery enables accurate shadow computation and illumination alignment across viewpoints.
  2. Forward PBR Rendering: The scene is rendered under the target illumination using a physically-designed renderer implementing a Cook-Torrance material model with GGX distribution and mesh-based shadowing. The rendering directly exposes explicit controls for lighting, camera, materials, and object placement.
  3. Diffusion Refinement: A CogVideoX-5B video diffusion model, modified for video-to-video translation, is fine-tuned to convert noisy PBR proxies into photorealistic relit outputs. The model operates on temporally-aligned latents, augmented with conditioning noise to robustify against proxy artifacts. Temporal tiling with overlap-fused prediction enables relighting for arbitrarily long sequences while maintaining global consistency.

Figure 2

Figure 2: LightCrafter recovers scene state, renders a PBR video for explicit relighting control, and refines the proxy via diffusion for photorealistic, temporally coherent output.

Training Data Curation

Artifact-matched training is achieved by constructing both synthetic and real-world supervision:

  • Synthetic Pairs: 3,000 pairs are synthesized by rendering compositional scenes under diverse lighting and camera motion, passing rendered outputs through the complete inverse rendering stack to produce PBR proxies with realistic artifacts.
  • Real-world Pseudo-Pairs: 1,000 DL3DV video chunks are inverse-rendered and their source illumination optimized such that a PBR proxy matches the input. The proxy/input pairs teach artifact correction on real-world appearance, compensating for the absence of ground-truth relit data.

This dual-source curation exposes the diffusion model to in-the-wild artifact distributions at training time, enabling robust generalization and effective artifact correction during inference.

Figure 3

Figure 3: Ablation reveals the criticality of combining synthetic and real-world supervision—training on both ensures accurate and faithful relighting.

Figure 4 further illustrates curation pipelines, emphasizing alignment and preservation of realistic reconstruction errors.

Figure 4

Figure 4: Data curation pipeline for synthetic and real-world video pairs, ensuring artifact-matched training supervision.

Experimental Evaluation

LightCrafter is validated on synthetic paired relighting, MIT Multi-Illumination (image relighting), and DL3DV pseudo-pair benchmarks.

  • On synthetic videos, LightCrafter achieves PSNR of 24.16, SSIM of 0.7366, and LPIPS of 0.2096, outperforming prior baselines by substantial margins.
  • On real-world data, LightCrafter yields PSNR of 19.91 and SSIM of 0.7450, with strong temporal consistency (T-CLIP 0.9892, Warp-SSIM 0.9598).

Contrary to previous claims that generative models generalize best for video relighting, the PBR proxy alone (without diffusion) frequently surpasses several end-to-end baselines, substantiating advances in monocular inverse rendering. The full LightCrafter pipeline achieves not only superior photorealism but also precise lighting control and coherence.

Figure 5

Figure 5: LightCrafter relights synthetic and real-world videos with higher fidelity, temporal consistency, and faithful lighting, compared to current baselines.

Long-form experimental results demonstrate LightCrafter's ability to maintain consistent shadows, highlights, and global appearance across arbitrarily long sequences, sidestepping temporal drift observed in chunked generative diffusion approaches.

Figure 6

Figure 6: LightCrafter delivers stable, temporally coherent relighting across long sequences, while generative baselines exhibit shadow and color inconsistencies.

Scene Editing Capabilities

Explicit controllability of the PBR proxy enables downstream editing applications—including lighting manipulation, material and object insertion, and camera trajectory change—without retraining the diffusion refiner. Editing intrinsics or lighting and re-rendering yields modified PBR proxies, which are then harmonized by the same diffusion model.

Figure 7

Figure 7: Downstream scene-editing applications are enabled, such as modifying light sources, materials, or inserting virtual objects.

Theoretical and Practical Implications

LightCrafter establishes a hybrid paradigm combining explicit physically-based rendering control and photometric learning-based refinement, thereby reconciling the tradeoff between deterministic structure and generative flexibility. The approach dispenses with the necessity of training diffusion models to learn explicit illumination semantics, since PBR proxies encode lighting directly. The artifact-matched training pipeline provides a template for bridging synthetic and real-world data domains, critical for forward-looking generalization in complex video tasks.

These design choices forecast several implications:

  • Efficient Relighting at Scale: By decoupling illumination control from generative refinement, large-scale relighting applications (e.g., entertainment, AR/VR, digital preservation) become tractable, with user-facing control interfaces for arbitrary lighting, materials, and objects.
  • Robustness and Consistency: The explicit pipeline mitigates temporal consistency issues ubiquitous in diffusion models, and the artifact-matched supervision paradigm can be generalized to other domains where forward rendering is feasible.
  • Future Directions: Co-optimization of inverse rendering and refinement may further reduce error propagation. Integration of stronger priors (multi-view geometry, semantic segmentation) may address persistent failure modes in challenging scenes with specular, transparent, or thin structures.

Limitations

LightCrafter is contingent on the quality of inverse-rendered intrinsics and their susceptibility to failure cases—specular surfaces, poor parallax, motion blur, or thin geometry. The diffusion refiner can only correct artifacts within its training distribution; excessive errors in geometry or lighting estimation may exceed its capability. Further methodological advances are required to address robustness in diverse real-world scenarios.

Conclusion

LightCrafter introduces a practically controllable, temporally stable video relighting system that synergizes physically-based rendering for explicit illumination control with a diffusion-based artifact correction mechanism. Strong quantitative results and qualitative coherence, alongside scene editing flexibility, establish the system as a robust foundation for relighting and downstream video editing applications. The artifact-matched curation pipeline ensures robust generalization and paves the way for future integration of explicit and generative paradigms in complex, temporally extended video tasks (2607.08016).

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 6 tweets with 34 likes about this paper.