Disentangling Vision-Encoder and Graph-Backbone Contributions

Determine how much origin-destination generation performance is attributable to the quality of the spatial vision encoder versus the choice of generative graph-diffusion backbone, by evaluating alternative backbones alongside the fixed WeDAN architecture.

Background

The experiments hold the WeDAN GraphTransformer diffusion architecture fixed while varying only the satellite vision encoder. This design isolates encoder differences, but it leaves unresolved whether the observed transfer failures and performance gaps arise from the encoders, from WeDAN's architectural limitations, or from their interaction. The paper specifically notes that alternative score-based graph-diffusion and variational flow-matching backbones may offer different scalability and generalization properties, while the contribution of backbone choice remains untested.

References

Disentangling encoder quality from backbone choice remains an open question for future work.

Do Satellites See Commuters? A Critical Benchmark of Vision Foundation Models  (2609.00661 - Iqbal et al., 1 Sep 2026) in Section 6, subsection “Backbone Architectural Constraints and Alternatives”