Towards traversing and understanding the solution space of neural networks

Complex problems often admit many distinct solutions. For example, stabilizing a nonlinear dynamical system (e.g., keeping an inverted pendulum upright) can be achieved with different control algorithms. Likewise, over-parametrized neural networks trained on the same task can achieve the same input-output mapping through different mechanisms, suggesting that training many networks could chart a problem’s solution space. However, in practice, standard neural network optimizers (e.g., stochastic gradient descent) bias training towards a narrow region of the solution space, which becomes further restricted as network size increases. In this talk, I will discuss recent progress towards systematically traversing the solution space of neural networks. First, by viewing recurrent neural networks (RNNs) as dynamical systems, I will show how a regularizer based on Koopman operator theory (which describes nonlinear dynamics via linear dynamics on observable functions) enables RNN training to be “pushed” away from previously found solutions. This allows us to discover networks that solve the same task with vastly different dynamics. Surprisingly, we find these distinct solutions are in a region of the loss landscape that is connected by low-loss paths. Motivated by this observation, I will present the Hessian null space continuation (HNC), a method we develop which allows for efficient exploration of the flat directions of the loss landscape (i.e., null space of the Hessian). I will show that the HNC allows us to generate diverse solutions in neural networks with thousands of parameters (RNNs) to tens of millions of parameters (vision Transformers). I will end by discussing a specific example (stabilizing the inverted pendulum) where the HNC allows us to not only identify different solutions, but also gain new insight into how the solution space is organized.

 

About the speaker:

Will Redman is an Assistant Professor at Johns Hopkins University in the Electrical and Computer Engineering Department and a member of the Data Science and AI Institute. He was previously a senior research scientist at Johns Hopkins Applied Physics Lab, a PhD student at UC Santa Barbara in the Interdepartmental Dynamical Neuroscience program, and an undergraduate student at NYU in mathematics and physics. His research interests are in the dynamics of learning and the learning of dynamics.

Date
Location
Amos Eaton 216
Speaker: Will Redman from Johns Hopkins University
Back to top