A Function-Space Theory of Learning Dynamics in Deep Neural Networks
Summary
Deep neural networks can display regular macroscopic behavior even though their parameter dynamics are highly nonlinear and occur in very high-dimensional spaces. This paper develops a statistical-mechanical description of learning directly in function space, treating parameter configurations as microscopic realizations and functions together with their dynamical operators as macroscopic variables. For mean-squared loss, it shows that exact error dynamics are governed by the learning operator M = JJ*, where J is the relevant Jacobian. The framework combines the dynamical Boltzmann weight of conditional stochastic dynamics with a parameter-space density of states whose local curvature is represented by a statistical operator B. After integrating over local fluctuations, the fluctuation contribution is expressed as Phi_fluc(M;B) = (sigma_xi^2/2) log det(M^-1+B) plus a constant. At fixed spectra, this contribution is rotationally stationary when M and B commute, is minimized when large eigenvalues of M pair with small eigenvalues of B, and produces a local restoring force against rotational mismatch. For ReLU-type function spaces under mild stable statistical conditions, the paper derives B = sigma_xi^2 L* K L. Because L measures coarse-grained second-order structure, the low-B sector corresponds, up to bounded anisotropy in K, to directions with low structural curvature. The result implies a preference for faster relaxation along smooth, data-adaptive directions and identifies function space as a natural macroscopic level for studying stable collective organization during learning.