Back

How AI Learns

Researchers at ICTP use theoretical physics to shed light on how deep neural networks learn
How AI Learns
The five co-authors of the article published on Physical Review X. From left to right: Mauro Pastore, Francesco Camilli, Jean Barbier, Rudy Skerk, and Minh-Toan Nguyen. © Alessandro Cenni / ICTP Photo Archive
Giulia Foffano

Artificial Intelligence is extremely efficient at processing large amounts of data, recognising the underlying patterns and identifying key features. AI tools, and particularly deep neural networks, are now used to automate and improve decision-making in key areas such as healthcare and hiring, and are becoming increasingly present in everyday lives. However, despite their success, we do not yet understand well how these complex systems succeed in learning so effectively from the data they are provided.

A research group led by ICTP researcher Jean Barbier has recently developed a sharp theoretical description of learning in a certain class of deep neural networks, shedding light on the fundamental processes underlying these tools. Their results were recently published in Physical Review X [1].

“This problem was a longstanding bottleneck, something our scientific community had been stuck upon for decades,” Barbier explains. “The theoretical framework we developed is a critical step we needed to understand how more realistic neural networks beyond toy models work,” he adds. His team’s insights may be key to develop better and more efficient training algorithms, and ultimately to optimize the energy and financial costs of training AI models.

The central challenge was understanding "feature learning" and the effect of depth in realistic neural networks. Feature learning –the property that makes neural networks so powerful– corresponds to their ability to extract the hidden structures underlying the input data, a crucial step needed to generalise beyond that initial dataset. “A really interesting case, where genuine feature learning emerges, is when the number of parameters of the neural network is comparable to the number of training data,” Barbier explains. This is called the “near-interpolation” regime and is the one that he and his collaborators focussed on.

For more than forty years, the theory of disordered systems, through its description of phase transitions, has been a natural framework for physicists to characterise the emergence and detectability of patterns in complex data. “The theoretical tools that had been developed so far to understand neural networks, however, were insufficient and only allowed one to consider networks too simple to be of actual interest in modern applications,” says Francesco Camilli, currently Assistant Professor at the University of Bologna, and one of the study co-authors. Determined to develop an analytical theory that could describe deep neural networks, the team combined the theory of disordered systems with random matrix theory, which allows to describe the statistical behaviour of large systems and was therefore key to tackle the complexity of deep neural networks. The framework they developed provides a rich, local description of the learning process.

Barbier and his team applied their tools to a specific type of neural network, the so-called Multi-Layer Perceptron, a key building block in the Large Language Models behind modern chatbots, obtaining a number of useful insights. The team was able to precisely predict how the performance of a deep neural network evolves with the amount of input data, and to understand its intricate learning mechanisms. For example, their study highlighted an interesting property: learning is directional, propagating from the inner to the outer (deeper) layers of the network. “This is somehow intuitive, because deeper layers learn more abstract features building on the more basic ones encoded by previous layers. Taking this into account could help us build better training algorithms,” says Minh-Toan Nguyen, an ICTP postdoctoral fellow who collaborated in the project.

These results were also the goal that Barbier had set for his five-year, ERC-funded project CHORAL (Computational Hardness Of RepresentAtion Learning), started in 2021. “The project was very ambitious and at the beginning it was not at all obvious that we would have been able to achieve our main objective,” says Barbier. “Our results emerged from a truly collaborative research process enriched by continuous exchange of ideas, perspectives, and expertise among the team,” says Mauro Pastore, also a postdoctoral fellow at ICTP.

Work, however, does not stop here. “We studied the best performance a network can achieve on a very specific task,” comments co-author Rudy Skerk, a PhD student at the International School for Advanced Studies (SISSA), adding, “With these foundations in place, we now plan to understand whether the training algorithms used in practice can reach this ideal performance, which is exactly the kind of insight that could help train real AI models faster and at lower cost.” “We look forward to applying our framework to more complex and realistic neural networks, in particular those enabling generative AI, and we are confident that with our new tools at hand, we will be able to shed light on some inner workings of AI systems that have so far remained obscure,” Barbier concludes.

 

[1] Barbier, Jean; Camilli, Francesco; Nguyen, Minh-Toan; Pastore, Mauro & Skerk, Rudy. (2026). Statistical physics of deep learning: Optimal learning of a multilayer perceptron near interpolation. Physical Review X, 16, 031014, https://journals.aps.org/prx/abstract/10.1103/56sb-pdh6

Publishing Date