Explain the source of explicit-conditioning benefits for visual memory
Determine whether the stronger visual-memory performance of 4D foundation models that explicitly project input observations onto target views before inpainting is caused by explicit memory conditioning simplifying invisible-segment generation, given that these models are predominantly trained on videos where target objects remain continuously visible.
References
We conjecture this is because most 4D models are predominantly trained on videos where target objects remain continuously visible.
Architectures that condition on explicit geometric structure partially compensate for this by encoding object persistence directly into the input, but it remains unclear whether such architectural advantages would persist under a memory-aware training paradigm.