SGD method for entropy error function with smoothing l0 regularization for neural networks
Abstract: The entropy error function has been widely used in neural networks. Nevertheless, the network training based on this error function generally leads to a slow convergence rate, and can easily be trapped in a local minimum or even with the incorrect saturation problem in practice. In fact, there are many results based on entropy error function in neural network and its applications. However, the theory of such an algorithm and its convergence have not been fully studied so far. To tackle the issue, we propose a novel entropy function with smoothing l0 regularization for feed-forward neural networks. Using real-world datasets, we performed an empirical evaluation to demonstrate that the newly conceived algorithm allows us to substantially improve the prediction performance of the considered neural networks. More importantly, the experimental results also show that our proposed function brings in more precise classifications, compared to well-founded baselines. Our work is novel as it enables neural networks to learn effectively, producing more accurate predictions compared to state-of-the-art algorithms. In this respect, we expect that the algorithm will contribute to existing studies in the field, advancing research in Machine Learning and Deep Learning.
- IEEE Geoscience and Remote Sensing Letters 17(6), 1087–1091 (2020). URL https://doi.org/10.1109/LGRS.2019.2937872
- IEEE Transactions on Information Theory 51(12), 4203–4215 (2005). DOI 10.1109/TIT.2005.858979
- Appl. Intell. 53(5), 5732–5749 (2023). DOI 10.1007/s10489-022-03539-8. URL https://doi.org/10.1007/s10489-022-03539-8
- Appl. Math. Comput. 311(C), 22–28 (2017). DOI 10.1016/j.amc.2017.05.010. URL https://doi.org/10.1016/j.amc.2017.05.010
- In: M. Kearns, S. Solla, D. Cohn (eds.) Advances in Neural Information Processing Systems, vol. 11. MIT Press (1999). URL https://proceedings.neurips.cc/paper/1998/file/a14ac55a4f27472c5d894ec1c3c743d2-Paper.pdf
- Neurocomputing 316, 262–269 (2018). DOI https://doi.org/10.1016/j.neucom.2018.07.075. URL https://www.sciencedirect.com/science/article/pii/S0925231218309111
- J. Sci. Comput. 87(1), 31 (2021). DOIÂ 10.1007/S10915-021-01443-W. URL https://doi.org/10.1007/s10915-021-01443-w
- Neurocomputing 128, 128–135 (2014). DOI https://doi.org/10.1016/j.neucom.2013.01.057. URL https://www.sciencedirect.com/science/article/pii/S0925231213007339
- Springer Series in Statistics. Springer New York Inc., New York, NY, USA (2001)
- Ishikawa, M.: Structural learning with forgetting. Neural Networks 9(3), 509–521 (1996). DOI 10.1016/0893-6080(96)83696-3. URL https://doi.org/10.1016/0893-6080(96)83696-3
- IEEE Transactions on Circuits and Systems II: Analog and Digital Signal Processing 39(7), 453–474 (1992). DOI 10.1109/82.160170
- Neurocomputing 314, 109–119 (2018). DOI 10.1016/j.neucom.2018.06.046. URL https://doi.org/10.1016/j.neucom.2018.06.046
- Information Sciences 588, 196–213 (2022). DOI https://doi.org/10.1016/j.ins.2021.12.065. URL https://www.sciencedirect.com/science/article/pii/S0020025521012871
- Neural Comput. Appl. 32(4), 1037–1050 (2020). DOI 10.1007/s00521-018-3933-z. URL https://doi.org/10.1007/s00521-018-3933-z
- Neurocomputing 138, 229–237 (2014). DOI 10.1016/j.neucom.2014.01.041. URL https://doi.org/10.1016/j.neucom.2014.01.041
- In: 7th International Conference on Learning Representations, ICLR 2019, New Orleans, LA, USA, May 6-9, 2019. OpenReview.net (2019). URL https://openreview.net/forum?id=Bkg6RiCqY7
- IEEE Transactions on Signal Processing 64(21), 5657–5671 (2016). DOI 10.1109/TSP.2016.2585096
- Neural Networks 6(6), 845–853 (1993). DOI 10.1016/S0893-6080(05)80129-7. URL https://doi.org/10.1016/S0893-6080(05)80129-7
- Information Sciences 492, 29–39 (2019). DOI https://doi.org/10.1016/j.ins.2019.04.012. URL https://www.sciencedirect.com/science/article/pii/S0020025519303135
- Nesterov, Y.: Introductory lectures on convex optimization : a basic course / Yurii Nesterov. Mathematics and its applications ; v. 564. Kluwer Academic Publishers, Boston (2004 - 2004)
- Oh, S.H.: Letters: Error back-propagation algorithm for classification of imbalanced data. Neurocomput. 74(6), 1058–1061 (2011). DOI 10.1016/j.neucom.2010.11.024. URL https://doi.org/10.1016/j.neucom.2010.11.024
- In: J. Ortega, W. Rheinboldt (eds.) Iterative Solution of Nonlinear Equations in Several Variables, pp. 1–6. Academic Press (1970). DOI https://doi.org/10.1016/B978-0-12-528550-6.50008-9. URL https://www.sciencedirect.com/science/article/pii/B9780125285506500089
- J. Intell. Manuf. 17(3), 285–299 (2006). DOI 10.1007/s10845-005-0005-x. URL https://doi.org/10.1007/s10845-005-0005-x
- Neurocomputing 410, 1–11 (2020). DOI https://doi.org/10.1016/j.neucom.2020.05.066. URL https://www.sciencedirect.com/science/article/pii/S0925231220309115
- Sharma, A.: Guided parallelized stochastic gradient descent for delay compensation. Applied Soft Computing 102, 107084 (2021). DOIÂ https://doi.org/10.1016/j.asoc.2021.107084. URL https://www.sciencedirect.com/science/article/pii/S1568494621000077
- Reliab. Eng. Syst. Saf. 230, 108920 (2023). DOIÂ 10.1016/J.RESS.2022.108920. URL https://doi.org/10.1016/j.ress.2022.108920
- Energy 284, 128677 (2023). DOIÂ https://doi.org/10.1016/j.energy.2023.128677. URL https://www.sciencedirect.com/science/article/pii/S0360544223020716
- Journal of Inverse and Ill-Posed Problems 21 (2013). DOIÂ 10.1515/jip-2012-0030
- Williams, P.M.: Bayesian regularization and pruning using a laplace prior. Neural Comput. 7(1), 117–143 (1995). DOI 10.1162/neco.1995.7.1.117. URL https://doi.org/10.1162/neco.1995.7.1.117
- Information Sciences 576, 173–186 (2021). DOI https://doi.org/10.1016/j.ins.2021.06.038. URL https://www.sciencedirect.com/science/article/pii/S0020025521006290
- Neural Processing Letters 52(3), 2687–2695 (2020). DOI 10.1007/s11063-020-10374-w. URL https://doi.org/10.1007/s11063-020-10374-w
- NEURAL NETWORKS 139, 17–23 (2021). DOI 10.1016/j.neunet.2021.02.011. URL http://dx.doi.org/10.1016/j.neunet.2021.02.011
- Mobile Networks and Applications 25(6), 2434–2446 (2020). DOI 10.1007/s11036-020-01587-3. URL https://doi.org/10.1007/s11036-020-01587-3
- IEEE Transactions on Neural Networks and Learning Systems pp. 1–15 (2023). DOI 10.1109/TNNLS.2023.3329525
- IEEE Transactions on Systems, Man, and Cybernetics: Systems 53(12), 7852–7863 (2023). DOI 10.1109/TSMC.2023.3300318
- DOIÂ 10.3389/fnins.2022.850932
- Neurocomputing 542, 126240 (2023). DOIÂ https://doi.org/10.1016/j.neucom.2023.126240. URL https://www.sciencedirect.com/science/article/pii/S0925231223003636
- Entropy 24(4) (2022). DOIÂ 10.3390/e24040455. URL https://www.mdpi.com/1099-4300/24/4/455
- IEEE Transactions on Cognitive and Developmental Systems pp. 1–13 (2023). DOI 10.1109/TCDS.2023.3329532
- Neural Comput. Appl. 26(2), 383–390 (2015). DOI 10.1007/s00521-014-1730-x. URL https://doi.org/10.1007/s00521-014-1730-x
Paper Prompts
Sign up for free to create and run prompts on this paper using GPT-5.
Top Community Prompts
Collections
Sign up for free to add this paper to one or more collections.