Parallelized Spatiotemporal Binding
Abstract: While modern best practices advocate for scalable architectures that support long-range interactions, object-centric models are yet to fully embrace these architectures. In particular, existing object-centric models for handling sequential inputs, due to their reliance on RNN-based implementation, show poor stability and capacity and are slow to train on long sequences. We introduce Parallelizable Spatiotemporal Binder or PSB, the first temporally-parallelizable slot learning architecture for sequential inputs. Unlike conventional RNN-based approaches, PSB produces object-centric representations, known as slots, for all time-steps in parallel. This is achieved by refining the initial slots across all time-steps through a fixed number of layers equipped with causal attention. By capitalizing on the parallelism induced by our architecture, the proposed model exhibits a significant boost in efficiency. In experiments, we test PSB extensively as an encoder within an auto-encoding framework paired with a wide variety of decoder options. Compared to the state-of-the-art, our architecture demonstrates stable training on longer sequences, achieves parallelization that results in a 60% increase in training speed, and yields performance that is on par with or better on unsupervised 2D and 3D object-centric scene decomposition and understanding.
- Object-centric image generation with factored depths, locations, and appearances. arXiv preprint arXiv:2004.00642, 2020.
- Monet: Unsupervised scene decomposition and representation. arXiv preprint arXiv:1901.11390, 2019.
- Inferno: Inferring object-centric 3d scene representations without supervision. In ICLR2022 Workshop on the Elements of Reasoning: Objects, Structure and Causality, 2022.
- Object representations as fixed points: Training iterative refinement algorithms with implicit differentiation. arXiv preprint arXiv:2207.00787, 2022.
- Dilated recurrent neural networks. Advances in neural information processing systems, 30, 2017.
- Roots: Object-centric representation and rendering of 3d scenes. J. Mach. Learn. Res., 22:259:1–259:36, 2020.
- Infogan: Interpretable representation learning by information maximizing generative adversarial nets. In Advances in neural information processing systems, pp. 2172–2180, 2016.
- Exploiting spatial invariance for scalable unsupervised object tracking. arXiv preprint arXiv:1911.09033, 2019.
- Learning 3d object-oriented world models from unlabeled videos. In Workshop on Object-Oriented Learning at ICML, 2020.
- Dang-Nhu, R. Evaluating disentanglement of structured representations. arXiv preprint arXiv:2101.04041, 2021.
- Language modeling with gated convolutional networks. In International conference on machine learning, pp. 933–941. PMLR, 2017.
- An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020.
- Unsupervised discovery of 3d physical objects from video. ArXiv, 2021.
- SAVi++: Towards end-to-end object-centric learning from real-world videos. In Advances in Neural Information Processing Systems, 2022.
- Genesis: Generative scene inference and sampling with object-centric latent representations, 2019.
- Genesis-v2: Inferring unordered object representations without iterative refinement. arXiv preprint arXiv:2104.09958, 2021.
- Masked autoencoders as spatiotemporal learners. ArXiv, abs/2205.09113, 2022.
- Latent bottlenecked attentive neural processes. arXiv preprint arXiv:2211.08458, 2022.
- Connectionism and cognitive architecture: A critical analysis. Cognition, 28(1-2):3–71, 1988.
- Unsupervised learning of temporal abstractions with slot-based transformers. arXiv preprint arXiv:2203.13573, 2022.
- Neural expectation maximization. In Advances in Neural Information Processing Systems, pp. 6691–6701, 2017.
- Multi-object representation learning with iterative variational inference. In International Conference on Machine Learning, pp. 2424–2433. PMLR, 2019.
- On the binding problem in artificial neural networks. arXiv preprint arXiv:2012.05208, 2020.
- Kubric: a scalable dataset generator. 2022.
- Efficiently modeling long sequences with structured state spaces. ArXiv, abs/2111.00396, 2021.
- Streetsurf: Extending multi-view implicit surface reconstruction to street views. ArXiv, abs/2306.04988, 2023.
- Deep residual learning for image recognition. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778, 2016.
- Unsupervised object-centric video generation and decomposition in 3d. Advances in Neural Information Processing Systems, 33:3106–3117, 2020.
- Gaussian error linear units (gelus). arXiv: Learning, 2016.
- beta-vae: Learning basic visual concepts with a constrained variational framework. In ICLR, 2017.
- Axial attention in multidimensional transformers. arXiv preprint arXiv:1912.12180, 2019.
- Dorsal: Diffusion for object-centric representations of scenes ${et al.}$. 2023.
- Perceiver: General perception with iterative attention. In International conference on machine learning, pp. 4651–4664. PMLR, 2021.
- Improving object-centric learning with query optimization. In The Eleventh International Conference on Learning Representations, 2023.
- Scalor: Generative world models with scalable object representations. In International Conference on Learning Representations, 2019.
- Object-centric slot diffusion. NeurIPS, 2023.
- Simone: View-invariant, temporally-abstracted object representations via unsupervised video decomposition. arXiv preprint arXiv:2106.03849, 2021.
- Conditional Object-Centric Learning from Video. In International Conference on Learning Representations (ICLR), 2022.
- Sequential attend, infer, repeat: Generative modelling of moving objects. In Neural Information Processing Systems, 2018.
- Nerf-vae: A geometry aware 3d scene generative model. In International Conference on Machine Learning, pp. 5742–5752. PMLR, 2021.
- Consistent generative query networks. arXiv preprint arXiv:1807.02033, 2018.
- Building machines that learn and think like people. Behavioral and brain sciences, 40:e253, 2017.
- Set transformer. 2018.
- Learning object-centric representations of multi-object scenes from multiple views. Advances in Neural Information Processing Systems, 33:5656–5666, 2020.
- Object-centric representation learning with generative spatial-temporal factorization. Advances in Neural Information Processing Systems, 34:10772–10783, 2021.
- Independently recurrent neural network (indrnn): Building a longer and deeper rnn. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 5457–5466, 2018.
- Improving generative imagination in object-centric world models. In International Conference on Machine Learning, pp. 6140–6149. PMLR, 2020a.
- Space: Unsupervised object-oriented scene representation via spatial attention and decomposition. ArXiv, abs/2001.02407, 2020b.
- Object-centric learning with slot attention. Advances in Neural Information Processing Systems, 33:11525–11538, 2020.
- Decoupled weight decay regularization. arXiv preprint arXiv:1711.05101, 2017.
- Complex-valued autoencoders for object discovery. arXiv preprint arXiv:2204.02075, 2022.
- Rotating features for object discovery. arXiv preprint arXiv:2306.00600, 2023.
- Nerf: Representing scenes as neural radiance fields for view synthesis. Commun. ACM, 65:99–106, 2020.
- Blockgan: Learning 3d object-aware scene representations from unlabelled images. Advances in neural information processing systems, 33:6767–6778, 2020.
- Giraffe: Representing scenes as compositional generative neural feature fields. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11453–11464, 2021.
- Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499, 2016.
- On the difficulty of training recurrent neural networks. In International conference on machine learning, pp. 1310–1318. Pmlr, 2013.
- Exploring the limits of transfer learning with a unified text-to-text transformer. The Journal of Machine Learning Research, 21(1):5485–5551, 2020.
- Object scene representation transformer. Advances in Neural Information Processing Systems, 35:9512–9524, 2022.
- Rust: Latent neural scene representations from unposed imagery. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 17297–17306, 2023.
- Scene representation transformer: Geometry-free novel view synthesis through set-latent scene representations. 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 6219–6228, 2021.
- Dyst: Towards dynamic neural scene representations on real-world videos. arXiv preprint arXiv:2310.06020, 2023.
- Seeing 3d objects in a single image via self-supervised static-dynamic disentanglement. arXiv preprint arXiv:2207.11232, 2022.
- Sequential neural processes. Advances in Neural Information Processing Systems, 32, 2019.
- Illiterate dall-e learns to compose. In International Conference on Learning Representations, 2021.
- Simple unsupervised object-centric learning for complex and naturalistic videos. Advances in Neural Information Processing Systems, 35:18181–18196, 2022.
- Core knowledge. Developmental science, 10(1):89–96, 2007.
- Contrastive training of complex-valued autoencoders for object discovery. arXiv preprint arXiv:2305.15001, 2023.
- Decomposing 3d scenes into objects via unsupervised volume segmentation. ArXiv, abs/2104.01148, 2021.
- Learning longer-term dependencies in rnns with auxiliary losses. In International Conference on Machine Learning, pp. 4965–4974. PMLR, 2018.
- Suds: Scalable urban dynamic scenes. 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 12375–12385, 2023.
- Conditional image generation with pixelcnn decoders. Advances in neural information processing systems, 29, 2016.
- Investigating object compositionality in generative adversarial networks. Neural Networks, 130:309–325, 2020.
- Attention is all you need. In Neural Information Processing Systems, 2017.
- Entity abstraction in visual model-based reinforcement learning. arXiv preprint arXiv:1910.12827, 2019.
- Towards causal generative scene models via competition of experts. arXiv preprint arXiv:2004.12906, 2020.
- D2nerf: Self-supervised decoupling of dynamic and static objects from a monocular video. arxiv preprint [2022-07-07], 2022.
- Inverted-attention transformers can learn object representations: Insights from slot attention. 2023a.
- Slotdiffusion: Object-centric generative modeling with diffusion models. NeurIPS, 2023b.
- Emernerf: Emergent spatial-temporal scene decomposition via self-supervision. ArXiv, abs/2311.02077, 2023a.
- Unisim: A neural closed-loop sensor simulator. In CVPR, 2023b.
- Robustifying sequential neural processes. In International Conference on Machine Learning, pp. 10861–10870. PMLR, 2020.
- Unsupervised discovery of object radiance fields. ArXiv, abs/2107.07905, 2021.
- Object-centric learning for real-world videos by predicting temporal feature similarities. arXiv preprint arXiv:2306.04829, 2023.
- Robust and controllable object-centric learning through energy-based models. arXiv preprint arXiv:2210.05519, 2022.
Paper Prompts
Sign up for free to create and run prompts on this paper.