Evaluate layout quality under underspecified prompts

Develop evaluation methods for layout quality in identity-preserving group-image generation when prompts are underspecified and multiple spatial arrangements may be equally valid, rather than measuring quality against a single reference layout.

Background

The paper evaluates layout from two perspectives: whether faces are planned correctly and whether the predicted arrangement preserves relative ground-truth relationships, such as left–right ordering. The authors note that this evaluation paradigm is problematic when a prompt does not uniquely specify the scene layout, because several arrangements may be equally appropriate. In that setting, comparison with one reference image cannot reliably determine layout quality, and the paper leaves the development of a solution unresolved.

References

When the prompt is underspecified, however, which is common in real-world use, many layouts may be equally valid and layout quality becomes difficult to measure against a single reference. We leave this to future work.

WithEveryone: Unified Planning and Identity Grounding for Group Image Generation  (2608.20336 - Xu et al., 20 Aug 2026) in Section Limitations, Future Work, and Responsible Use, paragraph “Layout is hard to evaluate”