Papers
Topics
Authors
Recent
Search
2000 character limit reached

On the weight dynamics of learning networks

Published 30 Apr 2024 in cs.LG and nlin.CD | (2405.00743v1)

Abstract: Neural networks have become a widely adopted tool for tackling a variety of problems in machine learning and artificial intelligence. In this contribution we use the mathematical framework of local stability analysis to gain a deeper understanding of the learning dynamics of feed forward neural networks. Therefore, we derive equations for the tangent operator of the learning dynamics of three-layer networks learning regression tasks. The results are valid for an arbitrary numbers of nodes and arbitrary choices of activation functions. Applying the results to a network learning a regression task, we investigate numerically, how stability indicators relate to the final training-loss. Although the specific results vary with different choices of initial conditions and activation functions, we demonstrate that it is possible to predict the final training loss, by monitoring finite-time Lyapunov exponents or covariant Lyapunov vectors during the training process.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (16)
  1. B. Ghani and S. Hallerberg, Applied Sciences 11, 9226 (2021).
  2. B. Warner and M. Misra, The american statistician 50, 284 (1996).
  3. E. Haber and L. Ruthotto, Inverse problems 34, 014004 (2017).
  4. H. Huang, Statistical mechanics of neural networks (Springer, 2021).
  5. E. Weinan, Communications in Mathematics and Statistics 1, 1 (2017).
  6. G.-H. Liu and E. A. Theodorou, arXiv preprint arXiv:1908.10920  (2019).
  7. E. Gelenbe, Neural computation 2, 239 (1990).
  8. Y. Fang and T. G. Kincaid, IEEE Transactions on Neural Networks 7, 996 (1996).
  9. C. L. Wolfe and R. M. Samelson, Tellus A 59, 355 (2007).
  10. P. V. Kuptsov and U. Parlitz, Journal of nonlinear science 22, 727 (2012).
  11. K. Fukushima, IEEE Transactions on Systems Science and Cybernetics 5, 322 (1969).
  12. D. Hendrycks and K. Gimpel, arXiv preprint arXiv:1606.08415  (2016).
  13. X. Glorot and Y. Bengio, in Proceedings of the thirteenth international conference on artificial intelligence and statistics (JMLR Workshop and Conference Proceedings, 2010) pp. 249–256.
  14. S. K. Kumar, arXiv preprint arXiv:1704.08863  (2017).
  15. C. W. Granger, Journal of econometrics 39, 199 (1988).
  16. M. Gori and A. Tesi, IEEE Transactions on Pattern Analysis and Machine Intelligence 14, 76 (1992).

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 3 tweets with 0 likes about this paper.