Determine the optimal spatial conditioning modality for thermal generation

Determine which spatial modality—RGB imagery, edge maps, depth maps, segmentation maps, or combinations of these—best serves thermal image generation within spatially conditioned diffusion models.

Background

Spatial conditioning is introduced to provide diffusion models with scene layout information that text prompts alone do not specify. The paper discusses several possible spatial signals, including edges, depth, segmentation, and RGB images, and notes that prior thermal-generation studies have evaluated spatial conditioning only in narrow scenarios.

The unresolved issue is which spatial representation provides the most useful structural information for thermal image synthesis. The paper evaluates these alternatives experimentally using the Text2Thermal framework, but the broader question is explicitly identified as remaining open rather than definitively settled.

References

Within the thermal domain, spatial conditioning has been applied to closing the synthetic-to-real gap in infrared segmentation and to object-centric synthesis, but evaluation in both cases is limited to narrow scenarios, and which spatial modality best serves thermal generation remains open.

Text2Thermal: Physics-Aware Thermal Image Synthesis from Textual Priors  (2609.03585 - Qazi et al., 3 Sep 2026) in Section 2.3, “Spatial control for diffusion models” (subsection of Related work)