Coarse-Graining Hidden Representations: Unsupervised Neuron Selection via Mapping Entropy Margherita Mele,1, 2 Andrea Castagna,1 Roberto Menichetti,1, 2 Raffaello Potestio,1, 2, ∗ and Alessandro Ingrosso3, † 1
arXiv:2609.05126v1 [cs.LG] 4 Sep 2026
2
Physics Department, University of Trento, via Sommarive, 14 I-38123 Trento, Italy INFN-TIFPA, Trento Institute for Fundamental Physics and Applications, I-38123 Trento, Italy 3 Donders Centre for Neuroscience, Radboud University, Nijmegen, The Netherlands (Dated: September 7, 2026)
Overparameterized neural networks carry far more hidden units than a task nominally requires, raising the question of which neurons are essential and whether that distinction is legible in the representation itself, without labels or gradients. We cast neuron selection as the problem of coarsegraining the hidden layer by retaining a subset of its neurons, and score each putative selection by the mapping entropy (ME). This quantity measures the loss of discriminatory power inherent in discarding part of the network neurons, and the selection that minimises the ME is taken as particularly informative. This criterion is fully unsupervised, in that it depends only on hidden-activation statistics. In teacher–student networks, ME optimisation recovers the minimal teacher-consistent representation and retains extra units in proportion to the hidden layer’s residual variability; in a non-linear Gaussian process task, it selects coherent functional-class mappings whose preferred class shifts across training. On this task and on translation-augmented MNIST, ME-selected subnetworks outperform random subsets of equal size, most clearly under strong compression — linking configurational distinguishability to predictive performance.
I.
INTRODUCTION
Modern neural networks have achieved remarkable empirical success, even though the principles through which learning, generalisation, and internal organisation emerge are not yet fully understood. These networks are generally highly overparameterized, in the sense that in order to achieve a high level of generalisation the system has to entail far more internal degrees of freedom than they are nominally required to encode the desired target input-output relation [1–4]. The specific regime in which a network operates strongly shapes its resulting representations. In the infinite-width limit, training dynamics linearise around initialisation and the network is described by a fixed kernel, the neural tangent kernel [5, 6]; in this lazy regime the internal features barely move, and the hidden representation is essentially inherited from the random initial weights [7]. The complementary feature-learning or rich regime is the one in which hidden units reorganise substantially during training, developing structured, data-adapted representations [8–10]. Widely regarded as a central ingredient of the optimisation dynamics and empirical success of neural networks, overparameterization raises fundamental questions regarding the nature of feature learning in deep networks. How is task-relevant computation distributed across the hidden units? Can we identify specific neurons that are essential to preserve predictive performance and others that, instead, play a redundant or secondary role in a learned representation? These questions also bear on two distinct but closely related concerns in deep learning: efficiency and interpretability. ∗ [email protected] † [email protected]
Efficiency pertains to identifying the neurons that carry out most of the task-relevant computation, which is central to pruning and model compression—reducing memory and compute while preserving predictive performance [11–15]. Strategies to reduce network complexity while retaining accuracy are essential for deployment in resource-constrained settings, and they bear on transferability, since a useful compressed representation should remain adaptable under fine-tuning on downstream tasks [16, 17]. Compression is especially valuable when labelled data are scarce, where one wishes to keep only the most robust and transferable features while shedding unnecessary capacity. The same motivations drive the interpretability agenda. Here one asks whether individual hidden units have identifiable functional roles, how strongly information is localised versus distributed, and whether dominant features can be disentangled within the learned representation [18–20]. This question has been brought into sharp focus by recent work in mechanistic interpretability, which seeks to reverse-engineer trained networks (transformers in particular) into human-understandable computational units. This literature has shown that individual neurons are frequently polysemantic, responding to superpositions of unrelated features, and has proposed that networks pack more features than they have neurons through superposition [21]; sparse dictionary methods such as sparse autoencoders have since been used to extract more monosemantic, interpretable directions from the residual stream of large language models [22, 23], and circuit-level analyses have attempted to trace how such features compose into algorithms [24, 25]. Classic vision-side counterparts include network dissection and the study of single directions important for generalisation [18, 19]. From this standpoint, neuron selection is not merely a compression device but a probe of the organ-
2 isation that training has induced—and a natural question is whether the statistical structure of the representation, on its own, already encodes which units are functionally important. A large body of work has attacked the efficiency question through pruning, sparsification, and regularisation. Parameters or whole neurons are ranked by criteria such as weight magnitude, local sensitivity, saliency, or estimated contribution to the loss, and the least important are removed [11–13, 26, 27]. Dropout regularisation has separately shown that networks remain effective when substantial fractions of hidden units are randomly suppressed during training, evidencing considerable redundancy in learned representations [28]. Importantly, this redundancy is highly non-uniform: some subnetworks or subsets of units preserve function far better than others, as emphasised by work on winning tickets, neuron importance, and the behavioural effects of compression [14, 19, 20, 29]. What remains less clear is whether these unequal contributions can be inferred directly from the internal statistical organisation of the hidden representation, rather than from supervised importance scores. In this work we address these questions from an information-theoretic perspective, framing the selection of relevant features in a representation as a coarsegraining problem. Rather than assigning neuron importance through their immediate effect on the supervised loss, we treat the hidden layer as a configuration space generated by the network over the dataset, and define a reduced representation by retaining only a subset of hidden neurons. Such reduction can thus be interpreted as a coarsening of the space of hidden configurations, where the full representation is mapped onto a lowerdimensional description. The central question is then whether the statistical structure of these hidden configurations, by itself, is sufficient to identify informative subsets. To quantify the effect of such reductions we employ the mapping entropy optimization workflow, or MEOW [30– 34], an information-theoretic approach that aims at minimising the mapping entropy (ME) [35–38], that is a measure of the loss of distinguishability between data representations induced by a given coarse-graining. Originally developed to identify maximally informative coarsegrained representations of biomolecular systems, MEOW has recently been applied to study neural dynamics [32] and the structure of parameter spaces in supervised learning problems [39]. Here, we apply MEOW to the hidden activation patterns of trained feed-forward networks. More specifically, in this setting, coarse-graining amounts to a decimation of the hidden layer, and minimising the ME allows one to identify the subset of neurons that best preserves the informational content of the representation. Neuron selection is thereby recast as an unsupervised search for reduced hidden representations that retain the relevant configurational information. The MEOW approach provides a completely unsuper-
vised criterion for the selection of relevant subsets of units in a hidden layer. We systematically analyse how the predictive performance of reduced networks is related to the contribution of selected neurons in preserving input information. We start our analysis in controlled settings using synthetic datasets. In a teacher-student (TS) scenario, the structure of the hidden representation is controlled by construction and the alignment between hidden units and teacher features can be monitored explicitly [4, 40]. This setting lets us ask precisely which neurons ME minimisation selects and how that selection evolves during training. We additionally employ a non-linear Gaussian process (NLGP) model, a synthetic but richer scenario in which hidden units spontaneously split into distinct functional classes [41]; here the goal is to assess whether ME can detect this heterogeneity directly from the hidden representations. Finally, the functional significance of the selected subsets is evaluated through pruning experiments on both NLGP and MNIST trained networks, comparing the predictive performance of the reduced networks against random baselines [42, 43]. II.
METHODS
The hidden layer of a feed-forward neural network is taken as the microscopic representation of the system. A coarse-grained (CG) description is then defined by selecting a subset of ncg < K neurons from a hidden layer of width K. Accordingly, a decimation mapping M specifies which neurons are retained and therefore determines the reduced representation. In particular, given an input pattern xµ , the network produces a hidden postactivation vector hµ = (hµ1 , . . . , hµK ). Its components are binarized by taking their sign, yielding the binary hidden configuration: ϕµ = (sµ1 , . . . , sµK ),
sµi = sign(hµi ) ∈ {−1, +1}. (1)
Although binarization discards potentially important magnitude information, such conservative choice allows us to reconstruct the empirical distribution without arbitrary binning and ensures that any detected structure is highly robust. The collection of the ϕµ configurations over the dataset defines the ensemble of microscopic states considered in the analysis. Since the same binary configuration can occur for different input patterns, each microscopic state can be associated with an empirical probability p(ϕ), given by its frequency in the dataset. For each ϕµ , the mapping M retains only selected neurons and, thus, induces a reduced configuration Φµ = M(ϕµ ). As a consequence, distinct microscopic configurations may be mapped onto the same reduced state and thus become indistinguishable after coarse-graining. To compare different mappings, it is therefore necessary to quantify how much of the statistical information contained in the original ensemble is lost through this reduction.
3 This loss is measured by the mapping entropy [31– 35, 37]. For a given mapping, the probability of a reduced configuration is obtained by summing the probabilities of all microscopic configurations mapped onto it. From this reduced description, a back-mapped distribution p̄M (ϕ) is constructed by redistributing the probability of each reduced state uniformly among the microscopic configurations compatible with it. The ME is then defined as the Kullback–Leibler divergence between the original microscopic distribution and the distribution reconstructed from the reduced representation, Smap (M) =
X
p(ϕ) log
ϕ
p(ϕ) . p̄M (ϕ)
M = arg
min
M∈Mncg
Smap (M).
(3)
The optimisation is carried out with different strategies depending on the size of the hidden layer and, correspondingly, on the number of admissible mappings. When the mapping space is sufficiently small, the minimum is determined through an exhaustive exploration of all configurations. For larger hidden layers, where such a complete enumeration becomes computationally prohibitive, the mappings minimizing the ME are searched through a stochastic simulated annealing, performed using the EXCOGITO software package implementing the MEOW protocol [31, 32].
A.
M M 1 X 1 X µ yTµ (x) = √ ϕ(xµ · Bi ) = √ ti , M i=1 M i=1
(4)
√ (2)
Accordingly, a smaller Smap indicates that the reduced representation preserves more of the statistical information contained in the original hidden-layer ensemble. At fixed reduced size ncg , the optimal subset is identified by minimising the ME over the set Mncg of all mappings that retain exactly ncg neurons: ⋆
Accordingly, the student has K = nM hidden neurons, indexed by pairs (i, α), where i ∈ {1, . . . , M } labels the associated teacher neuron and α ∈ {1, . . . , n} labels the replica within that group. Input patterns are sampled independently from the standard Gaussian distribution, xµ ∼ N (0, IN ). The teacher and student outputs are defined by
Teacher-Student
As an initial controlled setting, we consider a onehidden-layer network within the TS framework [44, 45]. Both teacher and student take inputs in RN and have hidden layers of size M and K > M , respectively. This yields a controlled over-realised TS setting, in which the student has more hidden units than the teacher, while the teacher provides a natural reference structure [4, 46]. Both networks are modelled as soft committee machines with fixed second-layer weights under the standard TS normalisation, and training acts only on the student firstlayer weights [4, 40]. Two instances of this setting are considered. The first is an analytically constructed replicated configuration, in which the student hidden units are organised into groups of replicas associated with the teacher units. The second consists of trained student networks obtained by online learning on teacher-generated examples. The teacher first-layer weights are taken to be an orN thonormal set {Bi }M and Bi · Bj = δij . i=1 , with Bi ∈ R
ySµ (x) =
√ M M n n M XX M XX µ µ ϕ(x · Ji,α ) = s , K i=1 α=1 K i=1 α=1 i,α (5)
where the activation function is x . ϕ(x) = erf √ 2
(6)
In the controlled setting, student first-layer weights are generated directly from the teacher weights through a tunable mismatch parameter η ∈ [0, 1]. Specifically, p Ji,α = 1 − η 2 Bi + η vi,α , (7) where vi,α is a unit vector drawn uniformly in the subspace orthogonal to Bi , so that vi,α · Bi = 0. By construction, each student vector remains normalised and has overlap p Ji,α · Bj = 1 − η 2 δij , (8) with the teacher weights. The parameter η therefore controls the degree of alignment: η = 0 corresponds to exact replication of the teacher representation, while η = 1 corresponds to a fully orthogonal representation within the complementary subspace. Alongside the controlled construction, trained student networks were also analysed. The student was trained by an online learning procedure on examples generated by the teacher, with only the first-layer weights updated and the second-layer weights kept fixed at their prescribed normalised values. As noted above, training naturally drives the student towards the replicated organisation that is enforced analytically in the in silico construction. To apply the same (i, α) labelling to trained networks, each student neuron is assigned to the teacher class with which it has the largest overlap at the end of training. With this convention, an effective mismatch parameter is defined as s 2 eff ηi,α (t) = 1 − max Ji,α (t) · Bj , (9) j
and its network average as ηeff (t) =
1 X eff η (t). K i,α i,α
(10)
4 The quantity ηeff (t) provides a time-dependent measure of the student-teacher alignment, enabling a direct comparison between trained networks and the controlled in silico reference. For each value of η (or ηeff ), the ME analysis was performed over subsets of student neurons. Since the replicated construction associates each student neuron with a well-defined teacher class, the selected mappings can be characterised not only by their ME value but also by their composition. For a mapping M, let ni (M) denote the number of selected student neurons associated with teacher neuron i, and define pi (M) =
ni (M) , ncg
i = {1, . . . , M },
(11)
PM with i=1 pi (M) = 1. The balance of the mapping is then quantified by ∆(M) =
1 − maxi pi (M) . 1 − 1/M
(12)
By construction, ∆ = 0 when all selected neurons belong to the same class, corresponding to a maximally redundant mapping, whereas ∆ = 1 when the selected subset is perfectly balanced across all M classes. The quantity ∆ was used throughout the analysis as a complementary descriptor of the structure of the selected mappings, and its definition applies without modification to both the in silico and trained settings. We investigated two TS configurations, both with input dimension N = 100. In the first, the teacher and student hidden layers have sizes M = 2 and K = 10, respectively, and the analysis is based on 3 × 104 examples; this system is small enough to allow an exhaustive exploration of the mapping space. In the second, M = 5 and K = 25, and the analysis is based on 105 examples; in this case, the larger mapping space makes it necessary to minimise the ME by simulated annealing rather than by exhaustive search. Apart from this difference in numerical treatment, the model construction and the definitions of η, ηeff , and ∆ are identical in the two settings.
±
where zξ ∈ RN is a zero-mean Gaussian vector with covariance ± ± |i − j|2 ξ± Cij = ⟨ziξ zjξ ⟩ = exp − ± 2 , (14) (ξ ) with periodic boundary conditions imposed on the input coordinate index. The two classes are therefore distin+ guished by their respective correlation lengths √ ξ and − ξ . The non-linear function ψ(z) = erf(z/ 2) introduces non-Gaussian statistics controlled by the parameter γ, while the normalisation factor Z(γ) is chosen such ± that Var(xξ ) = 1. Throughout the analysis, the values N = 50, ξ + = 7.5, ξ − = 2, and γ = 15 were employed. Both the training and test sets contain a total of P = αN patterns, with α = 300, equally divided between the two classes. The labels assigned to the two classes are y = +1 and y = −1, respectively to configurations with correlation length ξ + and ξ − . The network is a one-hidden-layer perceptron with N input units, K = 30 hidden neurons, and a single linear output unit. The hidden-layer activation function is x ϕ(x) = erf √ , (15) 2 √ and the output weight vector is fixed to wiout = 1/ K for all i, with zero output bias; only the first-layer weights and biases are trained. Training was performed by stochastic gradient descent on the mean squared error loss with L2 regularisation coefficient 0.1, learning rate 0.1, and batch size 300. Networks were trained for up to 3 × 104 epochs. For the analysis of trained networks, hidden neurons were partitioned into two empirical groups according to the structure of their incoming weight vectors. Neurons with weight vectors concentrated on a restricted subset of input coordinates were classified as localised, whereas neurons with extended alternating-sign profiles were classified as oscillatory [41]. The two groups were quantitatively identified by means of the inverse participation ratio (IPR), defined for each hidden neuron i as PN IPR(wi ) = P
4 j=1 wij
N 2 j=1 wij
B.
Non-Linear Gaussian Process
As a second model, we considered a synthetic binary classification task based on a NLGP [41]. Input patterns are N -dimensional vectors, where each component xi , with i = 1, . . . , N , plays the role of a pixel intensity. Two classes are generated, corresponding to Gaussian fields with correlation lengths ξ + and ξ − , respectively. The patterns belonging to each class are generated as ±
ξ±
x
ψ(γ zξ ) = , Z(γ)
(13)
2 ,
(16)
which ranges between 1/N , attained when all weights are equal in magnitude, and 1, attained when only a single weight is non-zero. Localised neurons therefore exhibit systematically larger IPR values than oscillatory ones. The two groups also develop distinct bias values during training, with oscillatory neurons associated with positive bias and localised neurons with negative one, thus providing a complementary indicator for their identification and tracking. The ME analysis was performed at multiple checkpoints during training. At each epoch, the hidden representations were computed by passing all training patterns
5 through the network and binarising the continuous activations via the sign function. The ME minimisation was then carried out independently at each checkpoint and for each value of the number of neurons retained ncg , using multiple independent stochastic minimisation runs. Additionally, the composition of the selected mappings can be also characterised in terms of the fraction of localised neurons floc (M) among the ncg retained units, floc (M) =
C.
nloc (M) . ncg
(17)
structed by combining the independently sampled spectra obtained at fixed nloc . This reconstruction is possible because each simulation explores the full range of effective energies accessible under the corresponding compositional constraint and because, for fixed ncg and nloc , the total number of admissible mappings is known exactly. These combinatorial counts provide the appropriate normalisation for each restricted DoS, thereby allowing the different contributions to be assembled into the full twodimensional DoS. The same analysis was performed at different training epochs in order to characterise the evolution of the mapping space structure during learning.
Wang-Landau Sampling of the Mapping Space
In addition to the direct minimisation of the ME, the structure of the mapping space of the NLGP system was characterised via Wang–Landau (WL) sampling in the 1/t formulation [39, 47, 48]. While ME minimisation identifies mappings with minimal information loss, it does not provide information on the global organisation of the mapping space. WL sampling complements this analysis by estimating the density of states (DoS), i.e., the number of configurations associated with a given value of an effective energy. The DoS is computed by adaptively biasing the Monte Carlo sampling to promote a broad exploration of the effective energy range while iteratively refining the current DoS estimate. The magnitude of these updates is controlled by a refinement parameter, which is progressively reduced during the simulation according to the 1/t scheme. In the present application, a configuration corresponds to a mapping M, and the ME, Smap (M), is taken as its effective energy. Sampling was terminated when the refinement parameter reached 10−5 . At fixed retained neuron size ncg and fixed number of localised neurons nloc , the mapping space is defined as the set of all subsets of ncg hidden neurons containing exactly nloc localised units or, analogously, ncg − nloc oscillatory units. The WL analysis was performed at fixed ncg = 15. This choice is particularly relevant for the present system, as the hidden layer consists of 30 neurons, evenly divided into 15 localised and 15 oscillatory units, making ncg = 15 a natural CG size for the analysis. For each admissible value of nloc , an independent WL simulation was carried out on the corresponding constrained mapping space. Trial moves were defined by exchanging one (or more) retained neuron(s) with one (or more) discarded neuron(s) belonging to the same class, thereby preserving both ncg and nloc throughout the sampling dynamics. Each simulation was continued until the refinement parameter reached 10−5 . The quantity of interest is the joint DoS: g(Smap , nloc ) ,
(18)
or, equivalently, its representation in terms of the fraction of localised neurons floc = nnloc . The joint DoS was reconcg
D.
Performance Evaluation of Reduced Networks
To test whether the mappings selected by ME minimisation also identify effective reduced representations, we construct pruned networks by retaining only the hidden neurons selected by the mapping. For each ncg value, the performance of the selected subset is compared with that obtained from random subsets of the same size. In the NLGP case, pruning is applied to the hidden layer of the trained classifier. After selecting the retained neurons, the output bias is re-optimised while keeping the hidden-layer parameters fixed. As a second classification benchmark, we consider a feed-forward neural network with a single hidden layer of K = 30 neurons and erf activation function. The task is binary classification on the MNIST dataset [42, 43], restricted to digits 1 and 7. To make the problem translationally invariant and more challenging, the dataset is augmented by random translations of the input images. The network is trained by stochastic gradient descent, with only the first-layer parameters optimised. Reduced networks are then obtained by pruning the hidden layer according to the selected mapping, and their test accuracy is compared with that obtained from random subsets of the same size.
III. A.
RESULTS
Structure of the Mappings Selected by Mapping Entropy Minimisation
In this section, we analyse the properties of the subsets selected by ME minimisation in controlled settings where the organisation of the hidden representation is already understood. To this end, we consider two complementary scenarios: the overparameterised TS regression model, where hidden units can be interpreted through their alignment with the teacher’s units, and the NLGP classification problem, where neurons organise into distinct functional classes. In both cases, this well-defined structure, known a priori, makes it possible to interpret the subsets selected by ME.
6 1.
Teacher-Student system
We begin with the controlled TS setting, where the structure of the student hidden layer is fixed by construction and therefore provides a direct reference for interpreting the subsets selected by ME minimisation. The relevant observable is the balance index ∆, defined in Eq. 12, whose values are shown in FIG. 1e,f as a function of the mismatch parameter η for different numbers of retained neurons ncg . By construction, ∆ = 1 identifies a mapping that distributes the retained neurons as evenly as possible across teacher groups, whereas smaller values correspond to progressively more concentrated selections. ∆ values close to 1 are therefore desirable, because they indicate that the selected subset preserves the full teacher structure rather than over-representing only a restricted part of it. We start from the exact replicated limit, η = 0, which provides the natural reference case for the discussion. Here, the ME-minimising mapping is fully balanced (∆ = 1) at the smallest number of neurons compatible with the teacher structure: ncg = 2 for the M2-K10 system and ncg = 5 for the M5-K25 system (FIG. 1.e,f). Thus, when the student exactly replicates the teacher representation, the entropy minimum coincides with the smallest subset that still covers all teacher modes. This is a non-trivial result, because the ME is computed solely from the statistics of the hidden configurations and has no explicit access to the teacher labels; nevertheless, in the replicated limit it selects one representative degree of freedom for each teacher mode. As soon as the mismatch becomes finite, deviations from this correspondence begin to emerge. In both systems, the curves in FIG. 1.e,f show that the smallest number of neurons for which the selected mapping is balanced increases monotonically with η. In the M2-K10 case, the mapping selected at the minimal size ncg = 2 is balanced only at η = 0, while already at η ≃ 0.1 it has collapsed to ∆ = 0. By contrast, the first ncg value that remains balanced over an extended mismatch range is ncg = 4, which keeps ∆ = 1 up to η ≈ 0.2. A similar trend is visible in the M5-K25 case, but on a broader scale: the minimal balanced mapping occurs at ncg = 5 only at η = 0, while for finite mismatch the balanced solution is shifted to substantially larger numbers of retained neurons. For instance, curves with ncg ≈ 10 remain at ∆ ≃ 0.875, showing that they still over-represent some teacher groups, whereas the first fully balanced solution appears only for ncg = 15. This systematic shift has a clear interpretation. The ME does not probe functional equivalence with the teacher, but the statistical distinguishability of the hidden units. At η = 0, different replicas associated with the same teacher unit are statistically indistinguishable, so one representative per group is sufficient to fully reproduce the whole network function. At finite η, each student neuron acquires an additional orthogonal component and replicas are no longer equivalent at the level
of hidden-layer statistics. As the variability of the hidden representation increases, the overall minimum of this profile shifts towards larger ncg values, because more neurons must be retained to resolve the additional variability. In this sense, the increase of the optimal ncg is not simply a failure to recover the teacher-minimal representation; rather, it quantifies how far the hidden representation is from the exact replicated limit. It is then natural to ask whether the same scenario survives in trained networks, where the variability is not imposed homogeneously by construction but generated by the learning dynamics. Panels FIG. 1.b–d show that training indeed drives the student towards the teacher representation. In the M2-K10 example, the overlaps with the two teacher directions increase from values close to zero at the beginning of training to values close to one at late times, while the orthogonal components are progressively suppressed (FIG. 1.b). Over the same interval, the generalisation error decreases by several orders of magnitude (FIG. 1.c), and the effective mismatch ηeff drops from approximately 1 to a final value ηfinal ≃ 0.04 (FIG. 1.d). The trained student therefore approaches, but does not exactly reach, the replicated limit. The ME analysis of the trained networks is fully consistent with the controlled construction. In FIG. 1.e,f, the square markers corresponding to the trained networks fall on the same branches defined by the in silico curves when plotted at the corresponding value of ηeff . Quantitatively, the trained networks no longer recover a balanced mapping at the strictly minimal size ncg = M : for M2-K10 the smallest balanced mapping is found at ncg = 4, and for M5-K25 at ncg = 15. These are precisely the same ncg values selected by the controlled curves at mismatch values of order η ≃ 0.04, showing that the trained networks behave as weakly mismatched replicas of the teacher representation. Taken together, these results identify a clear selection principle in the TS setting. In the exact replicated limit, ME minimisation recovers the minimal teacher-consistent hidden representation. As soon as replicas are no longer statistically equivalent, the ncg value for which ∆ = 1 increases, and the amount by which it increases provides a direct measure of the residual variability present in the hidden layer. The agreement between the analytically controlled construction and the trained networks shows that this interpretation remains valid beyond the idealised setting and continues to hold in the presence of heterogeneous, learning-induced correlations.
2.
NLGP System
A richer and less constrained setting is provided by the NLGP classification problem, where the hidden representation is structured, but not through explicit replication. The full set of results is shown in FIG. 2. As already known from previous studies [41], after training the hidden neurons separate into two classes, illustrated
7
Teacher network
N M
…
B1 B2
erf
B1 ⋅ B2 = 0
Student network
N
K J1,1
…
J1,2 J2,1
erf
J2,2 Ji,α ⋅ Bj =
1 − η 2 δij
FIG. 1. TS system and ME analysis of hidden-layer representations. (a) Schematic representation of the teacher and student networks. The teacher has M hidden units with orthogonal first-layer weights Bi , whilepthe student has K > M hidden units labelled by (i, α). In the controlled construction, the student weights satisfy Ji,α ·Bj = 1 − η 2 δij . (b) Training dynamics of the overlaps Ji,α ·Bj in the M2-K10 case, shown as a function of the training epoch. Blue and green curves denote the overlaps with B1 and B2 , respectively, while grey curves denote the components orthogonal to the teacher directions. (c) Generalization error as a function of the training epoch for the same system. (d) Effective mismatch parameter ηeff as a function of the training epoch. The dashed horizontal line indicates the final value ηfinal = 0.04. (e, f ) Balance ∆ of the ME-minimising mapping as a function of the mismatch parameter η, for the M2-K10 and M5-K25 systems, respectively. Different colours identify different values of the retained CG dimensionality ncg , as indicated by the colour bar. Circular markers denote the controlled in silico construction, and square markers denote the trained networks, plotted at the corresponding effective mismatch value ηeff . Note the different scale of the ∆ axis in the two plots.
in FIG. 2.b: localised neurons, whose weights are concentrated on a restricted subset of input coordinates, and oscillatory neurons, whose weights display an extended alternating-sign profile. Here this prior structural information is used as a reference frame to analyse the mappings selected by the ME minimisation. Before turning to the ME analysis, it is important to establish that the two neuron classes correspond to distinct yet individually informative representations of the classification task. This is shown in FIG. 2.c, where the network output is evaluated after retaining only one class of hidden neurons at a time. In both cases, the outputs associated with the two input classes remain clearly separated, showing that either population alone is sufficient to preserve discriminative information. The difference lies in the form of the representation: localised neurons and oscillatory neurons induce distinct output distributions, reflecting two qualitatively different internal encodings of the same classification rule. FIG. 2.c therefore shows that the two populations should not be interpreted
as playing complementary roles that become meaningful only when combined; rather, they define alternative representational modes, each of which can by itself support the separation of the input classes. The ME analysis shows that the optimal mappings do not combine these two representations arbitrarily. In FIG. 2.d (left subpanel) we report the average fraction of localised neurons among the optimal mappings for different training epochs. Early in the training phase the neurons selected in the mapping are almost invariably localised irrespective of the subset size ncg . This picture is confirmed by the selection probability pin that a given neuron is included in an optimal mapping, reported in the right subpanel of FIG. 2.d: at epoch 100, in fact, the optimisation repeatedly selects the same band of localised neurons with pin ≈ 1, whereas oscillatory neurons are almost never included. As training proceeds, this preference changes in a strongly size-dependent way. Around epoch 300, the average localised fraction starts to decrease for intermedi-
8
FIG. 2. NLGP system and ME analysis of hidden-layer representations. (a) Schematic of the classification task and network architecture. Input patterns from the two classes, characterised by different correlation lengths, are shown alongside the onehidden-layer network with K = 30 neurons. (b) Representative examples of hidden neurons after training: a localised neuron (left), with weights concentrated on a subset of input coordinates, and an oscillatory neuron (right), with extended alternatingsign structure. (c) Distribution of the network output evaluated using only localised neurons (green) or only oscillatory neurons (blue), compared across the two classes (hatched histograms). (d) Left: average fraction of localised neurons in the ME-minimising mapping as a function of the number of retained neurons ncg , for different training epochs (colour-coded). Right: selection probability pin of each neuron, shown as a function of ncg , for two representative epochs (100 and 1000). Neurons are ordered according to their bias value. (e) ME Smap evaluated for mappings composed exclusively of localised neurons (green) or oscillatory neurons (blue), as a function of training epoch. (f ) DoS of the mapping space as a function of Smap and the number of localised neurons, for different training epochs. Overlaid points indicate average optimisation trajectories.
ate and large ncg values, signalling the appearance of a competing oscillatory solution. By late training, the transition is sharp: at epoch 1000, the optimal mappings are still almost fully localised for very small subsets, but the localised fraction drops below 1/2 already around ncg ≃ 6 − 7, and becomes essentially zero for ncg ≳ 10. The selection probability map at epoch 1000 shows that this crossover does not arise from a gradual mixing of the two populations. Rather, for small ncg the selected neurons come almost exclusively from the localised group, whereas for larger ncg the selected mappings are almost entirely oscillatory. This interpretation is corroborated by the direct com-
parison shown in FIG. 2.e at ncg = 15, where the retained subset coincides with the full set of oscillatory neurons or, alternatively, with the full set of localised neurons. At epoch 100, purely localised mappings have a substantially lower ME than purely oscillatory ones, with Smap ≈ 0.26 for the localised case against Smap ≈ 0.31 for the oscillatory one. Around epoch 200, the two values become comparable, Smap ≃ 0.32, marking the point at which the two representational modes compete on equal footing. At later epochs the ordering is reversed: the value of the ME of selections containing exclusively oscillatory neurons remains approximately flat in the interval 0.31−0.33, whereas the one including only localised neu-
9 rons increases steadily up to ≈ 0.38 by epoch 1000. The growing gap quantitatively explains why, at late times, oscillatory mappings dominate the optimal set at large ncg . The WL reconstruction of the full mapping space provides a more global view of this transition. In FIG. 2.f, the DoS is shown as a function of Smap and the number nloc of localised neurons retained in mappings of fixed size ncg = 15. A first important point is that the ME minima are generally not found in mixed mappings containing substantial contributions from both functional classes; rather, they lie close to the edges of the space, where the retained subset is dominated by one class or the other. At epoch 100, the low-Smap basin is found at large nloc , at the fully localised edge, and the optimisation trajectories converge to this region. At epoch 300, the landscape becomes effectively bimodal: two competing basins are visible, and the optimisation trajectories split between them. By epoch 1000, the low-Smap basin has moved to nloc = 0, i.e., to purely oscillatory mappings, while the localised side remains only as a higher-Smap sector of the mapping space. Therefore, the change observed in FIG. 2.d is not a finite sampling artefact of the annealing dynamics, but it reflects a true restructuring of the ME landscape. Overall, the NLGP results show that ME minimisation does not favour sparse or mixed subsets. Instead, it selects coherent mappings associated with one of the two pre-existing representational classes, and the preferred class depends both on the stage of training and on the number of neurons retained. At early epochs, localised neurons provide the most informative reduced description. During training, an oscillatory representation progressively becomes competitive and eventually dominates for sufficiently large ncg . The ME therefore acts as a probe of the internal organisation of the hidden layer, resolving which representational mode provides the statistically more significant CG description.
B.
Performance of the Reduced Network
Having characterised the structure of the mappings selected by ME minimisation in controlled settings, we now turn to their functional validation. The goal is to determine whether these subsets, beyond exhibiting a clear and non-trivial organisation in the benchmark cases discussed above, also preserve the predictive performance of the original network after pruning. Since the MEOW criterion relies solely on the statistics of the hidden activation patterns, with no direct access to the task or to the output labels, this constitutes a stringent test of its effectiveness as an unsupervised pruning criterion. We begin with the NLGP classifier, where the performance of reduced networks can be monitored at different stages of training and for several values of retained neurons ncg (FIG. 3). For each ncg value, the accuracy obtained from the MEOW-selected subset is compared with
FIG. 3. NLGP reduced-network performance at different training epochs. Classification accuracy as a function of the number of retained neurons ncg . For each ncg , the green markers denote the accuracy of the reduced network obtained from the ME-selected subset, the orange markers the accuracies obtained from random subsets of the same size, and the dashed grey line the accuracy of the full network. After pruning, the output bias is re-optimised while keeping the hidden-layer parameters fixed. Note that the y axes do not start from zero.
that of random subsets of the same cardinality, as well as with the accuracy of the full network. After pruning, the output bias is re-optimised while keeping the hiddenlayer parameters fixed. This is a minimal readjustment, introduced to compensate for the shift of the output distribution induced by the selected hidden subset. Its use is motivated looking at FIG. 2.c: localised and oscillatory populations define distinct, yet individually viable, representations of the classification rule, but they generally induce different offsets in the output distribution. As a consequence, pruning changes the effective decision threshold of the classifier. If the zero threshold used for the full network were kept unchanged, all reduced selections would yield an accuracy close to 0.5, with one of the two classes being almost entirely misclassified. Reoptimising the output bias therefore removes this trivial source of performance loss and allows one to compare different subsets on the basis of the information retained in the hidden representation. From an inspection of FIG. 3, a clear trend emerges. For all epochs considered, the ME-selected subsets typically achieve an accuracy above the bulk of the random distribution, showing that the mappings identified
10
FIG. 4. Reduced-network performance for the MNIST binary classification task (1 vs 7) with data augmentation by random translations. Classification accuracy as a function of the number of retained neurons ncg . For each ncg , the green markers denote the accuracy of the reduced network obtained from the ME-selected subset, the orange markers the accuracies obtained from random subsets of the same size, and the dashed grey line the accuracy of the full network. Note that the y axes do not start from zero.
through the activation statistics also retain a comparatively informative representation for the task. This advantage is most visible in the strongly compressed regime. For instance, at ncg = 8 and ncg = 12, the optimised subset lies systematically above the typical random choice at all epochs. The same behaviour persists at ncg = 15 and 18, although the separation becomes progressively smaller as the number of retained neurons increases. By contrast, at ncg = 22 the MEOW selection and the random choice of retained neurons provide nearly indistinguishable outcomes, as expected when most of the hidden layer is already retained. Overall, the NLGP results indicate that ME minimisation provides a non-trivial and functionally meaningful guide for pruning, especially when the compression is so strong (i.e. when up to half of the hidden layer neurons are culled) that the specific choice of subset matters. The same analysis can then be extended to a more realistic classification problem, namely the binary MNIST task with data augmentation by random translations, using a network with K = 30 hidden-neurons. In this case, shown in FIG. 4, we report the reduced-network performance only at the end of training, since the behaviour is qualitatively stable across epochs. Also here the accuracy obtained from the MEOW-selected subset is compared with the distribution generated by random subsets of equal size. The overall picture is consistent with the NLGP benchmark, but with a clearer dependence on the number of retained neurons. Excluding the extreme case ncg = 2, where all reduced networks perform poorly, the ME-selected mappings systematically lie in the upper part of the random distribution for small and intermedi-
ate ncg values. In particular, for 3 ≲ ncg ≲ 12, the selected subset typically yields an accuracy well above the bulk of the random configurations and often close to their upper tail. This indicates that, in the strongly pruned regime, the ME criterion is able to identify hidden subsets that preserve most effectively the task-relevant information. For larger ncg values, however, the distinction progressively weakens: the random and optimised subsets become compatible within the observed spread. This is consistent with the fact that, when the reduction is mild, the space of admissible subsets becomes much less heterogeneous and the benefit of a structured selection is correspondingly reduced. Taken together, these results show that the mappings selected by ME minimisation are not only structurally meaningful, but also functionally relevant. The selection criterion does not identify the optimal reduced network in an absolute sense, nor should this be expected, since the approach is blind to the specific task and relies only on the hidden activation statistics; nevertheless, in a broad range of regimes it provides an effective proxy for pruning. Its advantage is most evident when the compression is sufficiently strong that the choice of the retained subset is non-trivial, while it naturally diminishes as ncg approaches the size of the full hidden layer. Because the MEOW protocol is strictly unsupervised, we omit comparisons with supervised pruning measures (e.g., weight magnitude, saliency, or gradient scores) that exploit information ME deliberately ignores. This isolates the capacity of activation statistics alone to identify critical units, leaving benchmarks against supervised criteria for future work.
IV.
CONCLUSIONS
In this work, we have investigated the space of coarsegrainings of neural networks through the lens of the MEOW approach, with the aim of identifying reduced representations that are both structurally meaningful and functionally effective. Our information-theoretic framework yields a purely unsupervised method to select informative subsets of neurons in a network, highlight functional heterogeneity within hidden layers and identify the evolution of learned representations throughout training. Minimizing the mapping entropy thus presents itself as a simple strategy to identify relevant features of the internal representation. The proposed strategy was put at test in multiple and diverse contexts. In controlled TS settings, the selected neurons are strongly aligned with the informative directions defined by the teacher. In the NLGP classification task, where the hidden layer spontaneously organises into two functionally distinct populations, the method does not select generic high-variance units: it selects coherent mappings drawn from a single representational class, and the preferred class changes over the course of training. Crucially, reduced networks obtained by retaining the
11 neurons selected through ME minimisation consistently outperform those constructed from random subsets of the same size. This establishes a concrete link between information-theoretic optimality and predictive performance, showing that the preservation of configurational distinguishability translates into better-preserved predictive performance relative to random pruning: at equal compression, ME-selected subsets degrade markedly less than random ones. These findings suggest that the ME captures structural information that is intrinsic to the data representation, rather than being tied to a specific task or loss function. In this sense, the MEOW strategy represents a principled pruning criterion that is complementary to standard approaches based on weight magnitude, sensitivity analysis, or gradient-based importance measures [11, 26]. Unlike these methods, which rely on supervised signals, the MEOW approach operates directly on the statistical structure of the hidden configurations, making it applicable in a broader range of settings. Several concrete directions stem from this work, which we detail below. a. Scaling and the lazy-to-rich transition. An interesting, relevant avenue would be the development of a simple representation-level diagnostic that is sensitive to the crossover from lazy to rich behaviour. In fact, a natural next step for our work is to scale the approach to substantially larger and wider networks and to use it as a diagnostic for the crossover between the feature-learning and lazy (kernel/NTK) regimes [5, 7, 9]. Because the ME is sensitive to whether hidden units are statistically distinguishable or interchangeable, we expect that the ME landscape, and in particular the balance and composition of ME-optimal subsets, should behave qualitatively differently in the two regimes: lazy-regime representations, inherited from initialisation and only weakly reorganised, should exhibit a flatter, more degenerate selection landscape than feature-learning ones. Systematically varying width, initialisation scale, and learning rate would allow the ME to be assessed as a representation-level order parameter for locating this transition. b. Iterative pruning for deep networks. The present study prunes a single hidden layer in one shot. A promising extension of our protocol is an iterative ME-based pruning scheme for deep architectures, in which units are removed in several rounds, each time re-estimating the hidden configuration statistics and re-minimising the ME on the surviving representation, in the spirit of iterative magnitude pruning and the lottery-ticket procedure [14, 29]. Interleaving decimation with brief fine-tuning between rounds could compensate for the shift of the internal representation and allow much stronger compres-
sion than single-shot selection. c. Other layers and architectures. Our framework applies to any layer that produces a well-defined configuration ensemble, so it can be carried beyond single fullyconnected layers to convolutional feature maps and, most interestingly, to the internal representations of attentionbased models and transformers. On the methodological side, such extensions call for more general forms of coarsegraining, going beyond the simple selection of individual neurons employed in this work. Applying ME-based coarse-graining to the residual stream, attention heads, or MLP neurons of transformers would connect this line of work to mechanistic interpretability, where superposition and polysemanticity make principled, unsupervised notions of unit importance particularly valuable [21, 22]; the ME minimisation could offer a statistics-driven criterion for head or feature selection that is complementary to dictionary-learning approaches. d. Robustness and adversarial vulnerability. Since the MEOW approach identifies the subset of units that carry the bulk of the representation’s distinguishability, it also singles out the degrees of freedom on which the network’s behaviour most sharply depends. This suggests using the ME as a lens on robustness: whether ME-critical neurons are disproportionately implicated in a network’s sensitivity to adversarial perturbations and distribution shift, and whether ME-guided pruning or regularisation of these units alters the trade-off between accuracy and robustness. Such a study would tie the geometry of the hidden representation to the security properties of trained models.
[1] I. Goodfellow, Y. Bengio, and A. Courville, Deep Learning. Cambridge, MA: MIT Press, 2016.
[2] L. Zdeborová, “Understanding deep learning is also a job for physicists,” Nature Physics, vol. 16, no. 6, pp. 602– 604, 2020.
ACKNOWLEDGMENTS
This work was conducted in the spirit of the Slow Science Manifesto, advocating for collaborative and sustainable research (slow-science.com).
AUTHOR CONTRIBUTIONS
RP and AI proposed the study; RP, AI and MM conceived the work plan and proposed the method; AC and MM carried out the simulations; MM carried out the preliminary data analyses. All authors contributed to the analysis and interpretation of the data. MM drafted the manuscript. All authors reviewed the results, contributed to writing the paper, and approved the final version of the manuscript.
12 [3] P. Mehta, M. Bukov, C.-H. Wang, A. G. R. Day, C. Richardson, C. K. Fisher, and D. J. Schwab, “A highbias, low-variance introduction to machine learning for physicists,” Physics Reports, vol. 810, pp. 1–124, 2019. [4] A. Engel and C. van den Broeck, Statistical Mechanics of Learning. Cambridge: Cambridge University Press, 2001. [5] A. Jacot, F. Gabriel, and C. Hongler, “Neural tangent kernel: Convergence and generalization in neural networks,” Advances in neural information processing systems, vol. 31, 2018. [6] J. Lee, L. Xiao, S. S. Schoenholz, Y. Bahri, R. Novak, J. Sohl-Dickstein, and J. Pennington, “Wide neural networks of any depth evolve as linear models under gradient descent,” 2019. [7] L. Chizat, E. Oyallon, and F. Bach, “On lazy training in differentiable programming,” Advances in neural information processing systems, vol. 32, 2019. [8] S. Mei, A. Montanari, and P.-M. Nguyen, “A mean field view of the landscape of two-layer neural networks,” Proceedings of the National Academy of Sciences, vol. 115, no. 33, pp. E7665–E7671, 2018. [9] M. Geiger, S. Spigler, A. Jacot, and M. Wyart, “Disentangling feature and lazy training in deep neural networks,” Journal of Statistical Mechanics: Theory and Experiment, vol. 2020, no. 11, p. 113301, 2020. [10] G. Yang and E. J. Hu, “Tensor programs iv: Feature learning in infinite-width neural networks,” in International Conference on Machine Learning, pp. 11727– 11737, PMLR, 2021. [11] Y. LeCun, I. Kanter, and S. A. Solla, “Second order properties of error surfaces: Learning time and generalization,” in Advances in Neural Information Processing Systems 3, pp. 918–924, 1990. [12] B. Hassibi, D. G. Stork, and G. J. Wolff, “Optimal brain surgeon and general network pruning,” in Proceedings of the IEEE International Conference on Neural Networks, vol. 1, pp. 293–299, IEEE, 1993. [13] S. Han, H. Mao, and W. J. Dally, “Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding,” in International Conference on Learning Representations (ICLR), 2016. [14] D. Blalock, J. J. Gonzalez Ortiz, J. Frankle, and J. Guttag, “What is the state of neural network pruning?,” Proceedings of Machine Learning and Systems, vol. 2, pp. 129–146, 2020. [15] G. E. Hinton, O. Vinyals, and J. Dean, “Distilling the knowledge in a neural network,” CoRR, vol. abs/1503.02531, 2015. [16] J. Yosinski, J. Clune, Y. Bengio, and H. Lipson, “How transferable are features in deep neural networks?,” in Advances in Neural Information Processing Systems 27, pp. 3320–3328, 2014. [17] E. Iofinova, A. Peste, M. Kurtz, and D. Alistarh, “How well do sparse imagenet models transfer?,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 12266–12276, 2022. [18] D. Bau, B. Zhou, A. Khosla, A. Oliva, and A. Torralba, “Network dissection: Quantifying interpretability of deep visual representations,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 3319–3327, 2017. [19] A. S. Morcos, D. G. T. Barrett, N. C. Rabinowitz, and M. M. Botvinick, “On the importance of single direc-
tions for generalization,” in International Conference on Learning Representations (ICLR), 2018. [20] S. Hooker, D. Erhan, P.-J. Kindermans, and B. Kim, “A benchmark for interpretability methods in deep neural networks,” in Advances in Neural Information Processing Systems 32, pp. 9737–9748, 2019. [21] N. Elhage, T. Hume, C. Olsson, N. Schiefer, T. Henighan, S. Kravec, Z. Hatfield-Dodds, R. Lasenby, D. Drain, C. Chen, et al., “Toy models of superposition,” arXiv preprint arXiv:2209.10652, 2022. [22] T. Bricken, A. Templeton, J. Batson, et al., “Towards monosemanticity: Decomposing language models with dictionary learning,” Transformer Circuits Thread, 2023. [23] H. Cunningham, A. Ewart, L. Riggs, R. Huben, and L. Sharkey, “Sparse autoencoders find highly interpretable features in language models,” arXiv preprint arXiv:2309.08600, 2023. [24] C. Olsson, N. Elhage, N. Nanda, N. Joseph, N. DasSarma, T. Henighan, B. Mann, A. Askell, Y. Bai, A. Chen, et al., “In-context learning and induction heads,” arXiv preprint arXiv:2209.11895, 2022. [25] K. Wang, A. Variengien, A. Conmy, B. Shlegeris, and J. Steinhardt, “Interpretability in the wild: a circuit for indirect object identification in gpt-2 small,” arXiv preprint arXiv:2211.00593, 2022. [26] D. Molchanov, A. Ashukha, and D. Vetrov, “Variational dropout sparsifies deep neural networks,” in Proceedings of the 34th International Conference on Machine Learning, vol. 70 of Proceedings of Machine Learning Research, pp. 2498–2507, PMLR, 2017. [27] H. Cheng, M. Zhang, and J. Q. Shi, “A survey on deep neural network pruning: Taxonomy, comparison, analysis, and recommendations,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 12, pp. 10558–10578, 2024. [28] N. Srivastava, G. E. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,” Journal of Machine Learning Research, vol. 15, no. 56, pp. 1929–1958, 2014. [29] J. Frankle and M. Carbin, “The lottery ticket hypothesis: Finding sparse, trainable neural networks,” in International Conference on Learning Representations (ICLR), 2019. [30] M. Giulini, R. Menichetti, M. S. Shell, and R. Potestio, “An information-theory-based approach for optimal model reduction of biomolecules,” Journal of Chemical Theory and Computation, vol. 16, no. 11, pp. 6795–6813, 2020. [31] M. Giulini, R. Fiorentini, L. Tubiana, R. Potestio, and R. Menichetti, “EXCOGITO, an extensible coarsegraining toolbox for the investigation of biomolecules by means of low-resolution representations,” Journal of Chemical Information and Modeling, vol. 64, no. 12, pp. 4912–4927, 2024. [32] R. Aldrigo, R. Menichetti, and R. Potestio, “Lowresolution descriptions of model neural activity reveal hidden features and underlying system properties,” Physical Review E, vol. 111, no. 4, p. 044315, 2025. [33] M. Rigoli, R. Potestio, and R. Menichetti, “A multiscale analysis of the czra transcription repressor highlights the allosteric changes induced by metal ion binding,” The Journal of Physical Chemistry B, vol. 129, no. 2, pp. 611– 625, 2025.
13 [34] A. Guadagnin Pattaro, R. Menichetti, and R. Potestio, “Detailed insight into the chignolin folding process from maximally informative low-resolution representations of its isocommittor hypersurfaces,” Journal of Chemical Theory and Computation, 08 2026. [35] A. Chaimovich and M. S. Shell, “Relative entropy as a universal metric for multiscale errors,” Physical Review E—Statistical, Nonlinear, and Soft Matter Physics, vol. 81, no. 6, p. 060104, 2010. [36] J. F. Rudzinski and W. Noid, “Coarse-graining entropy, forces, and structures,” The Journal of Chemical Physics, vol. 135, no. 21, p. 214101, 2011. [37] K. M. Kidder and W. Noid, “Analysis of mapping atomic models to coarse-grained resolution,” The Journal of Chemical Physics, vol. 161, no. 13, 2024. [38] S. Hummerich, T. Bereau, and U. Köthe, “Split-flows: Measure transport and information loss across molecular resolutions,” 2026. [39] M. Mele, R. Menichetti, A. Ingrosso, and R. Potestio, “Density of states in neural networks: an in-depth exploration of learning in parameter space,” Transactions on Machine Learning Research, 2025. [40] D. Saad and S. A. Solla, “On-line learning in soft committee machines,” Physical Review E, vol. 52, no. 4, pp. 4225–4243, 1995. [41] A. Ingrosso and S. Goldt, “Data-driven emergence of convolutional structure in neural networks,” Proceedings of the National Academy of Sciences, vol. 119, no. 40, p. e2201854119, 2022.
[42] Y. LeCun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based learning applied to document recognition,” Proceedings of the IEEE, vol. 86, no. 11, pp. 2278– 2324, 1998. [43] L. Deng, “The MNIST database of handwritten digit images for machine learning research [best of the web],” IEEE Signal Processing Magazine, vol. 29, pp. 141–142, Nov. 2012. [44] S. Goldt, M. S. Advani, A. M. Saxe, F. Krzakala, and L. Zdeborová, “Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup,” in Advances in Neural Information Processing Systems 32, pp. 6979–6989, 2019. [45] S. Goldt, M. S. Advani, A. M. Saxe, F. Krzakala, and L. Zdeborová, “Dynamics of stochastic gradient descent for two-layer neural networks in the teacher-student setup,” Journal of Statistical Mechanics: Theory and Experiment, no. 12, p. 124010, 2020. [46] T. L. H. Watkin, A. Rau, and M. Biehl, “The statistical mechanics of learning a rule,” Reviews of Modern Physics, vol. 65, no. 2, pp. 499–556, 1993. [47] F. Wang and D. P. Landau, “Efficient, multiplerange random walk algorithm to calculate the density of states,” Physical Review Letters, vol. 86, no. 10, pp. 2050–2053, 2001. [48] R. E. Belardinelli and V. D. Pereyra, “Wang-landau algorithm: A theoretical analysis of the saturation of the error,” The Journal of Chemical Physics, vol. 127, no. 18, p. 184105, 2007.