Papers
Topics
Authors
Recent
Search
2000 character limit reached

Recurrent Complex-Weighted Autoencoders for Unsupervised Object Discovery

Published 27 May 2024 in cs.LG and cs.NE | (2405.17283v3)

Abstract: Current state-of-the-art synchrony-based models encode object bindings with complex-valued activations and compute with real-valued weights in feedforward architectures. We argue for the computational advantages of a recurrent architecture with complex-valued weights. We propose a fully convolutional autoencoder, SynCx, that performs iterative constraint satisfaction: at each iteration, a hidden layer bottleneck encodes statistically regular configurations of features in particular phase relationships; over iterations, local constraints propagate and the model converges to a globally consistent configuration of phase assignments. Binding is achieved simply by the matrix-vector product operation between complex-valued weights and activations, without the need for additional mechanisms that have been incorporated into current synchrony-based models. SynCx outperforms or is strongly competitive with current models for unsupervised object discovery. SynCx also avoids certain systematic grouping errors of current models, such as the inability to separate similarly colored objects without additional supervision.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (69)
  1. Max Wertheimer. Untersuchungen zur lehre von der gestalt. ii. Psychologische Forschung, 4(1):301–350, 1923.
  2. Kurt Koffka. Principles of gestalt psychology. Philosophy and Scientific Method, 32(8), 1935.
  3. Wolfgang Köhler. Gestalt psychology. Psychologische Forschung, 31(1), 1967.
  4. Neural networks trained on natural scenes exhibit gestalt closure. Computational Brain & Behavior, 4:251–263, 2021.
  5. Experience-dependent perceptual grouping and object-based attention. Journal of Experimental Psychology: Human Perception and Performance, 28(1):202, 2002.
  6. Elizabeth S. Spelke. Principles of object perception. Cognitive Science, 14(1):29–56, 1990.
  7. Core knowledge. Developmental science, 10(1):89–96, 2007.
  8. The discovery of structural form. Proceedings of the National Academy of Sciences, 105(31):10687–10692, 2008.
  9. Simulation as an engine of physical scene understanding. Proceedings of the National Academy of Sciences, 110(45):18327–18332, 2013.
  10. On the binding problem in artificial neural networks. Preprint arXiv:2012.05208, 2020.
  11. Attention over learned object embeddings enables complex visual reasoning. In Proc. Advances in Neural Information Processing Systems (NeurIPS), 2021.
  12. Structured agents for physical construction. In Proc. Int. Conf. on Machine Learning (ICML), 2019.
  13. Learning to generalize with object-centric agents in the open world survival game crafter. IEEE Transactions on Games, pages 1–20, 2023.
  14. An investigation into pre-training object-centric representations for reinforcement learning. In Proc. Int. Conf. on Machine Learning (ICML), 2023.
  15. Learning dexterous grasping with object-centric visual affordances. In 2021 IEEE international conference on robotics and automation (ICRA), 2021.
  16. Apex: Unsupervised, object-centric scene segmentation and tracking for robot manipulation. In IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2021.
  17. Self-supervised visual reinforcement learning with object-centric representations. In Int. Conf. on Learning Representations (ICLR), 2021.
  18. Anne Treisman. The binding problem. Current opinion in neurobiology, 6(2):171–178, 1996.
  19. Adina L Roskies. The binding problem. Neuron, 24(1):7–9, 1999.
  20. Learning to Segment Images Using Dynamic Feature Binding. Neural Computation, 4(5):650–665, 1992.
  21. Michael C Mozer. A principle for unsupervised hierarchical decomposition of visual scenes. In Proc. Advances in Neural Information Processing Systems (NIPS), 1998.
  22. Neuronal synchrony in complex-valued deep networks. In Int. Conf. on Learning Representations (ICLR), 2014.
  23. Complex-valued autoencoders for object discovery. Transactions on Machine Learning Research, 2022.
  24. Contrastive training of complex-valued autoencoders for object discovery. In Proc. Advances in Neural Information Processing Systems (NeurIPS), 2023.
  25. Rotating features for object discovery. In Proc. Advances in Neural Information Processing Systems (NeurIPS), 2023.
  26. Geoffrey Hinton. How to Represent Part-Whole Hierarchies in a Neural Network. Neural Computation, 35(3):413–452, 2023.
  27. Binding dynamics in rotating features. ArXiv, abs/2402.05627, 2024.
  28. David Waltz. Generating semantic descriptions from drawings of scenes with shadows. Technical report, Massachusetts Institute of Technology, MIT MAC-TR 271, 1972.
  29. Deep complex networks. In Int. Conf. on Learning Representations (ICLR), 2018.
  30. Delving deep into rectifiers: Surpassing human-level performance on imagenet classification. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2015.
  31. Multi-object datasets. https://github.com/deepmind/multi-object-datasets/, 2019.
  32. Multi-object representation learning with iterative variational inference. In Proc. Int. Conf. on Machine Learning (ICML), 2019.
  33. Object-centric learning with slot attention. In Proc. Advances in Neural Information Processing Systems (NeurIPS), 2020.
  34. Efficient iterative amortized inference for learning symmetric and disentangled multi-object representations. In Proc. Int. Conf. on Machine Learning (ICML), 2021.
  35. Adam: A method for stochastic optimization. In Int. Conf. on Learning Representations (ICLR), 2015.
  36. William M Rand. Objective criteria for the evaluation of clustering methods. Journal of the American Statistical association, 66(336):846–850, 1971.
  37. Comparing partitions. Journal of classification, 2:193–218, 1985.
  38. Laurens van der Maaten and Geoffrey Hinton. Visualizing data using t-sne. Journal of Machine Learning Research, 9(86):2579–2605, 2008.
  39. Attend, infer, repeat: Fast scene understanding with generative models. In Proc. Advances in Neural Information Processing Systems (NIPS), 2016.
  40. Sequential attend, infer, repeat: Generative modelling of moving objects. In Proc. Advances in Neural Information Processing Systems (NeurIPS), 2018.
  41. Spatially invariant unsupervised object detection with convolutional neural networks. In Proc. AAAI Conf. on Artificial Intelligence, 2019.
  42. Hierarchical relational inference. In Proc. AAAI Conf. on Artificial Intelligence, 2021.
  43. SPACE: unsupervised object-oriented scene representation via spatial attention and decomposition. In Int. Conf. on Learning Representations (ICLR), 2020.
  44. Matrix capsules with EM routing. In Int. Conf. on Learning Representations (ICLR), 2018.
  45. Tagger: Deep unsupervised perceptual grouping. In Proc. Advances in Neural Information Processing Systems (NIPS), 2016.
  46. Neural expectation maximization. In Proc. Advances in Neural Information Processing Systems (NIPS), 2017.
  47. Inverted-attention transformers can learn object representations: Insights from slot attention. In UniReps: the First Workshop on Unifying Representations in Neural Models, NeurIPS, 2023.
  48. Jürgen Schmidhuber. Learning to control fast-weight memories: An alternative to recurrent nets. Neural Computation, 4(1):131–139, 1992.
  49. Monet: Unsupervised scene decomposition and representation. arXiv preprint arXiv:1901.11390, 2019.
  50. Genesis: Generative scene inference and sampling with object-centric latent representations. In Int. Conf. on Learning Representations (ICLR), 2020.
  51. Neural systematic binder. In Int. Conf. on Learning Representations (ICLR), 2022a.
  52. Conditional object-centric learning from video. In Int. Conf. on Learning Representations (ICLR), 2022.
  53. SAVi++: Towards end-to-end object-centric learning from real-world videos. In Proc. Advances in Neural Information Processing Systems (NeurIPS), 2022.
  54. Simple unsupervised object-centric learning for complex and naturalistic videos. In Proc. Advances in Neural Information Processing Systems (NeurIPS), 2022b.
  55. Object-centric learning for real-world videos by predicting temporal feature similarities. In Proc. Advances in Neural Information Processing Systems (NeurIPS), 2023.
  56. Learning object-centric representations of multi-object scenes from multiple views. In Proc. Advances in Neural Information Processing Systems (NeurIPS), 2020.
  57. Decomposing 3d scenes into objects via unsupervised volume segmentation. arXiv:2104.01148, 2021.
  58. Object Scene Representation Transformer. In Proc. Advances in Neural Information Processing Systems (NeurIPS), 2022.
  59. Audioslots: A slot-centric generative model for audio separation. In IEEE International Conference on Acoustics, Speech, and Signal Processing Workshops (ICASSPW), 2023.
  60. Unsupervised musical object discovery from audio. In Workshop on Machine Learning for Audio, NeurIPS, 2023.
  61. Unsupervised Learning of Temporal Abstractions With Slot-Based Transformers. Neural Computation, 35(4):593–626, 2023.
  62. Christoph von der Malsburg and Werner Schneider. A neural cocktail-party processor. Biological cybernetics, 54(1):29–40, 1986.
  63. Christoph von der Malsburg and Joachim Buhmann. Sensory segmentation with coupled neural oscillators. Biological cybernetics, 67(3):233–242, 1992.
  64. DeLiang Wang. The time dimension for scene analysis. IEEE Transactions on Neural Networks, 16(6):1401–1426, 2005.
  65. Unsupervised segmentation with dynamical units. IEEE Transactions on Neural Networks, 19(1):168–182, 2008.
  66. Christoph von der Malsburg. Binding in models of perception and brain function. Current opinion in neurobiology, 5(4):520–526, 1995.
  67. Visual feature integration and the temporal correlation hypothesis. Annual review of neuroscience, 18(1):555–586, 1995.
  68. Wolf Singer. Neuronal synchrony: A versatile code for the definition of relations? Neuron, 24(1):49–65, 1999.
  69. Scikit-learn: Machine learning in Python. Journal of Machine Learning Research, 12:2825–2830, 2011.
Citations (1)

Summary

  • The paper introduces SynCx, a recurrent complex-weighted autoencoder that uses complex-valued weights and iterative phase propagation to achieve unsupervised object binding.
  • Empirical evaluations show SynCx achieving superior object grouping performance (ARI 0.89 on Tetrominoes) and competitive results on other datasets without relying on external inductive biases.
  • This approach simplifies entity segmentation and offers a more principled method for unsupervised object discovery, aligning with neurobiological theories and proving robust in challenging scenarios.

Recurrent Complex-Weighted Autoencoders for Unsupervised Object Discovery

The paper presents a novel approach for unsupervised object discovery through the employment of a recurrent architecture with complex-valued weights. The model, termed SynCx, is designed as a fully convolutional autoencoder that leverages iterative constraint satisfaction, contrasting with the existing synchrony-based models. Current state-of-the-art models utilize real-valued weights and an array of architectural features such as gating mechanisms to achieve object binding, which are shown to possess limitations in certain grouping scenarios and datasets. This paper argues that these extras are unnecessary, and instead, complex-valued weights can natively handle binding through the inherent properties of complex numbers, such as phase relationships.

Key Contributions and Methodology

The principal contribution of this research is the SynCx model, which employs complex-weighted matrix operations to achieve binding between objects. Significant features of SynCx include:

  • Complex-Valued Weights: Previous models treated complex-valued activations but operated with real-valued weights. SynCx instead utilizes complex-valued weights entirely, which can handle constructive and destructive interferences devoid of additional mechanisms like χ-binding or cosine binding.
  • Iterative Phase Propagation: A distinctive aspect of SynCx is its recurrent nature, where the model iteratively updates the phases of the activation maps, propagating soft constraints throughout. This aligns phases to achieve binding over iterations, akin to forming global interpretations of objects from local features.
  • Avoidance of Inductive Biases: The SynCx model does not rely on contrastive training or pre-trained features, factors required by many contemporary models to avoid grouping errors on datasets with similarly colored objects.

Empirical Findings

SynCx has been evaluated on a set of diverse datasets (Tetrominoes, dSprites, and CLEVR) to quantify its object grouping efficacy. Notably, SynCx achieved superior performance on the Tetrominoes dataset, with an ARI of 0.89, indicating its reliability even when objects are similarly colored, a known challenge for other models like RF. While on dSprites and CLEVR, SynCx performed competitively with state-of-the-art models, it demonstrated the significant merit of its unsupervised binding ability without reliance on additional supervision.

Implications and Future Directions

The implications of adopting complex-valued weights for unsupervised learning are substantial. It removes the need for manual feature engineering and complicated network structures, providing a simpler, more principled approach to entity segmentation. The findings suggest potential applications in contexts where color is not a reliable cue and where unsupervised feature discovery remains paramount. Additionally, this research aligns with the temporal correlation hypothesis, strengthening the linkage between computational models and neurobiological phenomena.

Future research could explore augmentations to SynCx, such as integrating spatial priors and enhancing the clustering mechanism to account for phase value variances and outliers, enhancing the robustness of object discovery mechanisms. Additionally, employing a temporal difference loss to refine iterative phase adjustment further represents an intriguing future direction.

Overall, the work substantiates the theoretical considerations of complex-weighted neural structures and positions SynCx as a robust baseline for further explorations in synchrony and binding implementations within unsupervised neural networks.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 7 tweets with 152 likes about this paper.