Perceptual Filtering and Downstream World-Model Utility

Determine how aesthetic and luminance-based perceptual filtering of synthetic Unreal Engine video data affects the dynamics learned by action-conditioned world models, including whether aggressive filtering discards useful structural diversity and whether permissive filtering retains samples that interfere with training.

Background

The production pipeline curates rendered samples using perceptual proxies, primarily aesthetic and luminance scores, to remove severely dark, overexposed, visually empty, or poorly rendered outputs. However, perceptual attractiveness may not correspond to the usefulness of a scene for learning environment dynamics, since visually unattractive scenes can still contain valuable geometry, motion, or interaction structure.

The paper states that the current thresholds were selected empirically and have not been calibrated against downstream world-model performance. It therefore identifies as unresolved the relationship between perceptual filtering decisions and the quality of learned action-conditioned dynamics, noting the competing risks of discarding structurally useful environments and retaining low-quality samples that could impair training.

References

An open question is therefore how perceptual filtering affects learned dynamics: overly aggressive filtering may discard useful structural diversity, whereas permissive filtering may retain low-quality samples that interfere with training. Establishing this relationship requires controlled downstream experiments and remains outside the scope of the present production-focused report.

Building Pretraining Data for World Models: An Unreal Engine-Based Pipeline for Action-Conditioned Video Generation  (2609.03557 - Wang et al., 3 Sep 2026) in Section 7, Limitations and Discussion, paragraph “Perceptual Quality and Downstream Utility”