Recurrent Complex-Weighted Autoencoders for Unsupervised Object Discovery
Abstract: Current state-of-the-art synchrony-based models encode object bindings with complex-valued activations and compute with real-valued weights in feedforward architectures. We argue for the computational advantages of a recurrent architecture with complex-valued weights. We propose a fully convolutional autoencoder, SynCx, that performs iterative constraint satisfaction: at each iteration, a hidden layer bottleneck encodes statistically regular configurations of features in particular phase relationships; over iterations, local constraints propagate and the model converges to a globally consistent configuration of phase assignments. Binding is achieved simply by the matrix-vector product operation between complex-valued weights and activations, without the need for additional mechanisms that have been incorporated into current synchrony-based models. SynCx outperforms or is strongly competitive with current models for unsupervised object discovery. SynCx also avoids certain systematic grouping errors of current models, such as the inability to separate similarly colored objects without additional supervision.
- Max Wertheimer. Untersuchungen zur lehre von der gestalt. ii. Psychologische Forschung, 4(1):301–350, 1923.
- Kurt Koffka. Principles of gestalt psychology. Philosophy and Scientific Method, 32(8), 1935.
- Wolfgang Köhler. Gestalt psychology. Psychologische Forschung, 31(1), 1967.
- Neural networks trained on natural scenes exhibit gestalt closure. Computational Brain & Behavior, 4:251–263, 2021.
- Experience-dependent perceptual grouping and object-based attention. Journal of Experimental Psychology: Human Perception and Performance, 28(1):202, 2002.
- Elizabeth S. Spelke. Principles of object perception. Cognitive Science, 14(1):29–56, 1990.
- Core knowledge. Developmental science, 10(1):89–96, 2007.
- The discovery of structural form. Proceedings of the National Academy of Sciences, 105(31):10687–10692, 2008.
- Simulation as an engine of physical scene understanding. Proceedings of the National Academy of Sciences, 110(45):18327–18332, 2013.
- On the binding problem in artificial neural networks. Preprint arXiv:2012.05208, 2020.
- Attention over learned object embeddings enables complex visual reasoning. In Proc. Advances in Neural Information Processing Systems (NeurIPS), 2021.
- Structured agents for physical construction. In Proc. Int. Conf. on Machine Learning (ICML), 2019.
- Learning to generalize with object-centric agents in the open world survival game crafter. IEEE Transactions on Games, pages 1–20, 2023.
- An investigation into pre-training object-centric representations for reinforcement learning. In Proc. Int. Conf. on Machine Learning (ICML), 2023.
- Learning dexterous grasping with object-centric visual affordances. In 2021 IEEE international conference on robotics and automation (ICRA), 2021.
- Apex: Unsupervised, object-centric scene segmentation and tracking for robot manipulation. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2021.
- Self-supervised visual reinforcement learning with object-centric representations. In Int. Conf. on Learning Representations (ICLR), 2021.
- Anne Treisman. The binding problem. Current opinion in neurobiology, 6(2):171–178, 1996.
- Adina L Roskies. The binding problem. Neuron, 24(1):7–9, 1999.
- Learning to Segment Images Using Dynamic Feature Binding. Neural Computation, 4(5):650–665, 1992.
- Michael C Mozer. A principle for unsupervised hierarchical decomposition of visual scenes. In Proc. Advances in Neural Information Processing Systems (NIPS), 1998.
- Neuronal synchrony in complex-valued deep networks. In Int. Conf. on Learning Representations (ICLR), 2014.
- Complex-valued autoencoders for object discovery. Transactions on Machine Learning Research, 2022.
- Contrastive training of complex-valued autoencoders for object discovery. In Proc. Advances in Neural Information Processing Systems (NeurIPS), 2023.
- Rotating features for object discovery. In Proc. Advances in Neural Information Processing Systems (NeurIPS), 2023.
- Geoffrey Hinton. How to Represent Part-Whole Hierarchies in a Neural Network. Neural Computation, 35(3):413–452, 2023.
- Binding dynamics in rotating features. ArXiv, abs/2402.05627, 2024.
- David Waltz. Generating semantic descriptions from drawings of scenes with shadows. Technical report, Massachusetts Institute of Technology, MIT MAC-TR 271, 1972.
- Deep complex networks. In Int. Conf. on Learning Representations (ICLR), 2018.
- Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2015.
- Multi-object datasets. https://github.com/deepmind/multi-object-datasets/, 2019.
- Multi-object representation learning with iterative variational inference. In Proc. Int. Conf. on Machine Learning (ICML), 2019.
- Object-centric learning with slot attention. In Proc. Advances in Neural Information Processing Systems (NeurIPS), 2020.
- Efficient iterative amortized inference for learning symmetric and disentangled multi-object representations. In Proc. Int. Conf. on Machine Learning (ICML), 2021.
- Adam: A method for stochastic optimization. In Int. Conf. on Learning Representations (ICLR), 2015.
- William M Rand. Objective criteria for the evaluation of clustering methods. Journal of the American Statistical association, 66(336):846–850, 1971.
- Comparing partitions. Journal of classification, 2:193–218, 1985.
- Laurens van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of Machine Learning Research, 9(86):2579–2605, 2008.
- Attend, infer, repeat: Fast scene understanding with generative models. In Proc. Advances in Neural Information Processing Systems (NIPS), 2016.
- Sequential attend, infer, repeat: Generative modelling of moving objects. In Proc. Advances in Neural Information Processing Systems (NeurIPS), 2018.
- Spatially invariant unsupervised object detection with convolutional neural networks. In Proc. AAAI Conf. on Artificial Intelligence, 2019.
- Hierarchical relational inference. In Proc. AAAI Conf. on Artificial Intelligence, 2021.
- SPACE: unsupervised object-oriented scene representation via spatial attention and decomposition. In Int. Conf. on Learning Representations (ICLR), 2020.
- Matrix capsules with EM routing. In Int. Conf. on Learning Representations (ICLR), 2018.
- Tagger: Deep unsupervised perceptual grouping. In Proc. Advances in Neural Information Processing Systems (NIPS), 2016.
- Neural expectation maximization. In Proc. Advances in Neural Information Processing Systems (NIPS), 2017.
- Inverted-attention transformers can learn object representations: Insights from slot attention. In UniReps: the First Workshop on Unifying Representations in Neural Models, NeurIPS, 2023.
- Jürgen Schmidhuber. Learning to control fast-weight memories: An alternative to recurrent nets. Neural Computation, 4(1):131–139, 1992.
- Monet: Unsupervised scene decomposition and representation. arXiv preprint arXiv:1901.11390, 2019.
- Genesis: Generative scene inference and sampling with object-centric latent representations. In Int. Conf. on Learning Representations (ICLR), 2020.
- Neural systematic binder. In Int. Conf. on Learning Representations (ICLR), 2022a.
- Conditional object-centric learning from video. In Int. Conf. on Learning Representations (ICLR), 2022.
- SAVi++: Towards end-to-end object-centric learning from real-world videos. In Proc. Advances in Neural Information Processing Systems (NeurIPS), 2022.
- Simple unsupervised object-centric learning for complex and naturalistic videos. In Proc. Advances in Neural Information Processing Systems (NeurIPS), 2022b.
- Object-centric learning for real-world videos by predicting temporal feature similarities. In Proc. Advances in Neural Information Processing Systems (NeurIPS), 2023.
- Learning object-centric representations of multi-object scenes from multiple views. In Proc. Advances in Neural Information Processing Systems (NeurIPS), 2020.
- Decomposing 3d scenes into objects via unsupervised volume segmentation. arXiv:2104.01148, 2021.
- Object Scene Representation Transformer. In Proc. Advances in Neural Information Processing Systems (NeurIPS), 2022.
- Audioslots: A slot-centric generative model for audio separation. In IEEE International Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSPW), 2023.
- Unsupervised musical object discovery from audio. In Workshop on Machine Learning for Audio, NeurIPS, 2023.
- Unsupervised Learning of Temporal Abstractions With Slot-Based Transformers. Neural Computation, 35(4):593–626, 2023.
- Christoph von der Malsburg and Werner Schneider. A neural cocktail-party processor. Biological cybernetics, 54(1):29–40, 1986.
- Christoph von der Malsburg and Joachim Buhmann. Sensory segmentation with coupled neural oscillators. Biological cybernetics, 67(3):233–242, 1992.
- DeLiang Wang. The time dimension for scene analysis. IEEE Transactions on Neural Networks, 16(6):1401–1426, 2005.
- Unsupervised segmentation with dynamical units. IEEE Transactions on Neural Networks, 19(1):168–182, 2008.
- Christoph von der Malsburg. Binding in models of perception and brain function. Current opinion in neurobiology, 5(4):520–526, 1995.
- Visual feature integration and the temporal correlation hypothesis. Annual review of neuroscience, 18(1):555–586, 1995.
- Wolf Singer. Neuronal synchrony: A versatile code for the definition of relations? Neuron, 24(1):49–65, 1999.
- Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830, 2011.
Paper Prompts
Sign up for free to create and run prompts on this paper.