2000 character limit reached
Memorizing Gaussians with no over-parameterizaion via gradient decent on neural networks
Published 28 Mar 2020 in cs.LG and stat.ML | (2003.12895v1)
Abstract: We prove that a single step of gradient decent over depth two network, with hidden neurons, starting from orthogonal initialization, can memorize independent and randomly labeled Gaussians in . The result is valid for a large class of activation functions, which includes the absolute value.
Paper Prompts
Sign up for free to create and run prompts on this paper.