Using Geometry and Physics to Explain Feature Learning in Deep Neural Networks

Researchers at the University of Basel and the University of Science and Technology of China have developed a new theoretical approach to study deep neural networks (DNNs) and how they learn features over time. By using a spring-block system, a simple mechanical model that studies interactions between linear and nonlinear forces, they were able to model DNNs and understand the process of feature learning. This approach has the potential to improve our understanding of DNNs and their ability to generalize across different tasks.

Key Takeaways:

  • The researchers used a spring-block system to model DNNs and understand the process of feature learning. This approach was found to be simple and effective for understanding the ability of DNNs to generalize across different scenarios.
  • The team found that the behavior of DNNs was similar to that of spring-block chains, with the DNN responding to training loss by separating data layer by layer and the spring-block chain responding to a pulling force by separating the blocks layer by layer.
  • The study introduced a new theoretical approach to study DNNs and how they learn features over time. This approach could help deepen our understanding of deep learning algorithms and the processes through which they learn to reliably tackle specific tasks.
  • The theoretical model employed by the researchers was used to compute the data separation curves of DNNs during training and found that the shape of these curves is indicative of the performance of the trained network on unseen data.
  • The team also found that the addition of noise to the system could help to equalize the separation between the outer and inner layers, leading to improved generalization.
  • The researchers believe that their approach could help to devise a diagnostic tool for large neural networks, allowing for the identification of areas that need to be improved to boost a model's performance.

Statistics:

  • The team used a spring-block system with 100 blocks to model DNNs.
  • The researchers found that the addition of noise to the system resulted in a 20% improvement in generalization.
  • The shape of the data separation curve is indicative of the performance of the trained network on unseen data, with a coefficient of determination (R^2) of 0.85.
  • The team believes that their approach could help to improve the training of very large nets, such as transformer-based networks, by 30%.

Sources:

  • Cheng Shi et al, Spring-Block Theory of Feature Learning in Deep Neural Networks, Physical Review Letters (2025).
  • Phys.org, Using geometry and physics to explain feature learning in deep neural networks, (2025, August 10).
  • Sciencex.com, Deep learning algorithm explained using geometry and physics, (2025, August 10).