Conceptio › Archive › arXiv CS
arXiv CSopen access

Geometric Mean Pooling for Equal-Weight Multiplicative Coarse-Graining

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

arXiv:2609.21876v1 [cs.LG] 18 Sep 2026

Geometric Mean Pooling for Equal-Weight Multiplicative Coarse-Graining

Ang-Kun Wu Department of Physics and Astronomy, University of Tennessee, Knoxville Knoxville, Tennessee 37996, USA [email protected] Fangdi Wen Department of Physics and Astronomy, Rutgers University New Brunswick, New Jersey 08901, USA

Jingtao Zhang Google Mountain View, CA 94043, USA

Abstract As an alternative to the additive and extremal biases of average and max pooling, we introduce Geometric Mean Pooling (GMP), a signed pooling operator that combines the product of feature signs with the geometric mean of feature magnitudes. Motivated by local-to-global composition in quantum many-body physics, GMP retains both joint sign information and a characteristic multiplicative scale without introducing learnable pooling parameters. We show that non-overlapping hierarchical GMP preserves the corresponding global multiplicative statistic and evaluate it on synthetic sequence tasks, iterative coarse-graining, image classification, and molecular lipophilicity regression. On the synthetic tasks, GMP recovers product-based signals more accurately than average and max pooling and maintains predictive performance under the tested levels of multiplicative input noise. On image and molecular data, however, its effectiveness depends on the representation, target parameterization, and placement of local and global pooling. These results position GMP as a complementary, regime-dependent inductive bias for tasks in which equal-weight multiplicative composition is plausible, rather than as a universal replacement for standard pooling operators.

1

Introduction

Pooling is a fundamental form of coarse-graining in convolutional neural networks (CNNs): it reduces spatial or sequence resolution while retaining features useful for downstream prediction. Classical CNN architectures established local max pooling as a source of spatial selectivity and translation tolerance [28, 27], while average pooling provides a complementary aggregation that preserves local mean responses. Global average pooling later became a standard way to replace fully connected layers with a spatially aggregated representation, reducing parameter count and encouraging correspondence between feature maps and class-level predictions [30, 18]. Modern surveys and learnable pooling methods have extended this basic design space to generalized means, learned norms, attention-based aggregation, and higher-order feature interactions [41, 16, 8]. These operators are effective when the target is controlled by additive statistics, salient responses, or learned weighted combinations of local features. They are less naturally matched to signals encoded in the joint product of many comparable features. Multiplicative interactions provide a distinct alternative to additive aggregation. Product units and multiplicative neural networks were introduced to represent interactions that cannot be expressed efficiently by purely additive units [10, 32], and multiplicative gates remain central to recurrent and Preprint.

attention-based architectures [21, 43]. Similar product-like structure appears in scientific settings in which a global quantity is assembled from local factors, including molecular properties, variational wavefunctions, and statistical-mechanical order parameters. In the present work, we focus on the QN more specific equal-weight Pcase in which the target depends on i=1 xi , or, for positive magnitudes, equivalently on the sum i log |xi |. The individual feature marginals may then be identical across classes even though their joint products differ. In such a setting, average and max pooling can discard the relevant dependence structure, whereas a multiplicative pooling operator can preserve the product of feature signs and the average of their log-magnitudes. Theoretically, quantum many-body states provide additional physical motivation for studying multiplicative composition. In the non-interacting limit, single-particle eigenmodes provide independent degrees of freedom, and a many-body occupation state is generated by applying creation operators to the vacuum; for example, |Ψk1 ,k2 ⟩ = a†k1 a†k2 |0⟩, where a†k creates a particle in eigenstate k. Thus, the occupation-number basis forms a product basis in Fock space, whereas interactions mix these configurations and generate superpositions and entanglement [12]. For a fermionic occupation state built from single-particle orbitals, the coordinate-space wavefunction is a Slater determinant: an antisymmetrized sum of products of one-particle orbitals [39, 9, 13]. In neural-network quantum-state representations, this determinant is commonly retained as an explicit antisymmetric component, while neural networks parameterize the orbitals and correlation factors[39, 9, 34, 20]. Valence-bond-solid states provide a related example of structured multiplicative composition, with amplitudes organized from local singlet-bond factors and a superposition over compatible bond configurations [2, 1, 50]. With the observations mentioned above, GMP, as a parameter-free operator that averages feature logmagnitudes while retaining sign parity, provides a multiplicative complement to average and max pooling. We study global GMP, which summarizes full feature maps or sequences, and local GMP, which preserves blockwise multiplicative summaries for subsequent processing. GMP outperforms average and max pooling applied directly to inputs on synthetic signed-product classification tasks; image classification and molecular regression examine its dependence on learned representations, pooling placement, and target parameterization. Together, these studies connect the preservation of multiplicative observables across scales to pooling choice as a structural prior. Contributions (1) A signed multiplicative pooling primitive. We define a pooling operator that combines sign parity with an equal-weight geometric mean of magnitudes for local and global aggregation, and distinguish it from unsigned geometric pooling and learned arithmetic modules. (2) A coarse-graining analysis. We establish hierarchical consistency for complete, equal-sized, non-overlapping groups with a shared clamp, and characterize how input signs, magnitudes, and activation choices affect the operator. (3) Controlled evaluation of applicability. We organize synthetic tasks around the statistics preserved by GMP, and use image and molecular experiments to examine its dependence on representation, activation, and pooling placement. Matched comparisons with average and max pooling isolate the role of the pooling primitive.

2

Related Work

Generalized means and pooling. Pooling methods include fixed average and max operators, learned norms, generalized means, and higher-order feature summaries [41, 16, 8]. Generalized Mean pooling (GeM), used in image retrieval [35, 29], is defined for positive inputs by !1/p n 1X p GeMp (x) = x . n i=1 i Its limit as p → 0 is the geometric mean. Weighted geometric pooling [46], alpha-integration pooling [11], attention-aware GeM [15], and groupwise GeM [24] further explore this family. GenAgg P [25] provides a directly relevant framework based on generalized f -means, f −1 (n−1 i f (xi )), with additional parameters extending the aggregation family. Its geometricmagnitude construction uses log-magnitudes. GMP builds on this established connection between geometric magnitudes and averaging in transformed coordinates. Its magnitude component inherits the equal-sized hierarchical consistency of quasi-arithmetic means. GMP fixes the magnitude trans2

form and adds the product of input signs; for strictly positive inputs, GMP clamps the inputs from below, averages their logarithms, and exponentiates the result. Learned multiplication and set representations. Neural Arithmetic Units include the Neural Multiplication Unit (NMU), which learns products over subsets of inputs [31]. Neural Power Units (NPU) learn power functions with a treatment of negative inputs through complex arithmetic [19]. These methods address learnable arithmetic structure, whereas GMP applies a prescribed parity rule and equal log-magnitude weights. GMP’s sign-parity convention provides a specific real-valued composition rule alongside these learned power functions. Deep Sets [55] studies invariant representations built from learned elementwise transformations and sum aggregation. This separates two complementary design choices: selecting the aggregation statistic and learning the representation supplied to it. Our experiments focus on the first choice through matched pooling substitutions. Geometry and physical coarse-graining. Geometric pooling on graphs and meshes concerns the geometry of the domain [53, 33, 4, 38]; geometric network renormalization likewise studies structural coarse-graining [56]. Here “geometric” refers to a mean of feature magnitudes. Connections between neural pooling, renormalization, and tensor networks motivate a block-coarse-graining interpretation [44, 22, 47, 17, 40, 48]. We use this interpretation to study which statistics are preserved as local summaries are composed across scales.

3

Methodology

3.1

The Geometric Mean Pooling Primitive

We define the signed geometric mean pooling (GMP) operator over a window W = {x1 , . . . , xk } as ! ! k k Y  1X log max(|xi |, ε) , (1) GMP(W ) = sign(xi ) exp k i=1 i=1 where ε > 0 is a lower bound on magnitudes inside the logarithm and sign(0) = 0. The implementation defaults are 10−6 in 1D and 10−12 in 2D; the hierarchical statement below assumes the same value at every level. The first factor records sign parity for nonzero inputs, and the second computes a clamped geometric magnitude. Log-space evaluation avoids forming the full magnitude product. Appendix A characterizes gradients and behavior near zero. For nonzero inputs with an inactive clamp, GMP is related to the product decomposition ! ! k k k Y Y X xi = sign(xi ) exp log |xi | . i=1

i=1

i=1

Thus, GMP preserves the joint sign and an equal-weight multiplicative scale. For known k and an inactive clamp, the product is recovered as sign(GMP)|GMP|k . The normalization gives GMP a characteristic magnitude scale; reconstruction of the full product additionally uses the window size. In the global regime, the window is the entire sequence, so k = N , and GMP reduces to a single equal-weight multiplicative summary of the whole input. In the local regime, GMP is applied blockwise with non-overlapping windows of size k and stride k:   N GMPlocal (x) = [GMP(W1 ), . . . , GMP(Wm )] , m= . (2) k Each complete block is replaced by its signed multiplicative summary. If k does not divide N , this definition omits the trailing entries; the hierarchical identity applies only when the blocks cover all inputs. For 2D inputs, GMP extends directly to spatial pooling windows. Given a window Wu,v of height r and width s, the 2D operator is ! ! r Y s r s Y  1 XX GMP2D (Wu,v ) = sign(xh,w ) exp log max(|xh,w |, ε) , (3) rs w=1 w=1 h=1

h=1

and is applied independently to each channel unless otherwise specified. 3

Renormalization-group perspective and algebraic contrast From a machine-learning perspective, a pooling layer replaces each local group of activations with a lower-resolution feature. Repeated application therefore composes local summaries into progressively coarser representations, which can be formalized as a coarse-graining map: each block is replaced by a single summary variable. In the language of Kadanoff block-spin RG [22, 47], a lattice of microscopic variables {xi }N i=1 is partitioned into non-overlapping blocks Bj of size b, and each block is replaced by a single block variable x′j = Rb (Bj ). Repeated application generates a flow R R R R x(0) −−−b→ x(1) −−−b→ x(2) −−−b→ · · · −−−b→ x(n) , Nn = N/bn . An order parameter is a fixed point of this flow if its value is preserved across scales; in other words, repeated blockwise coarse-graining yields the same statistic as a single global summary. This is the natural RG language for distinguishing pooling operators. For Ravg = b (B) P average pooling, the block map is the additive block-spin transform 1 max x , and for max pooling it is the extremal order-statistic map R (B) = max x i i∈B i. b i∈B b These are natural RG maps for linear or extremal observables, but they are not fixed points for equal-weight multiplicative   GMP defines the block map Q Pstatistics. By contrast, 1 sign(x ) exp RGMP (B) = log max(|x |, ε) . Under the change of variables i i b i∈B  i∈B b ui = log max(|xi |, ε) , the magnitude part of GMP becomes an arithmetic average in logmagnitude space: 1X log |RGMP (B)| = ui . b b i∈B

Hence GMP is algebraically conjugate to additive averaging in log-space, and repeated nonoverlapping block pooling preserves the global multiplicative statistic exactly (proof in Appendix B). In this precise sense, GMP is a fixed-point RG map for equal-weight multiplicative order parameters. When the target depends on a product of many features, average and max pooling may discard information needed to recover the target. In Case A of the experiments below, GMP directly retains the sign parity that determines the class label, whereas the tested average- and max-pooling classifiers perform near chance in the independent-input setting. 3.2

Local and Global Placement

Global GMP summarizes a full sequence or feature map and is structurally aligned with a global equal-weight multiplicative statistic. Local GMP preserves blockwise statistics for downstream processing and is aligned with tasks in which the relevant factors occur within those blocks. These two placements let the pooling scale reflect the spatial or sequential organization of the signal. The pooling inputs determine the role of the sign component. With strictly positive inputs, GMP equals the geometric mean of clamped inputs. With ReLU inputs, any zero in a window forces its signed-GMP output to zero, and the implemented backward pass gives zero derivatives for all inputs in that window. Thus, activation choice affects both the information available to pooling and its optimization behavior.

4

Experiments

The synthetic tasks examine prescribed sign and magnitude statistics; the application experiments examine pooling within learned representations. 4.1

Synthetic Sequence Classification

Case A (global signed-product classification). For each example, we draw a clean sequence QN i.i.d. xi ∼ N (0, 1) of length N and set the label y = 1{ i=1 xi < 0}, with each xi being independent identically distributed (i.i.d.). Equivalently, the two classes differ only in the parity of negative entries: class 0 has an even number of negative entries and class 1 an odd number. For this independent-input setting (N > 1), the class prior is balanced and every individual coordinate has the same marginal distribution in both classes. We compare global max, average, and GMP, each 4

0.8

Test accuracy

1.0 Geometric Max Average

0.6

0.8

0.4 1.0

0.8

0.8

0.6 0.4 0.2

Geometric Max Average

0.6

0.4 1.0

Test R 2

Test R 2

Test accuracy

1.0

0.6 0.4 0.2

30

60

90 120 Sequence length L

150

180

2

4

6 Unit cell M

8

10

Figure 1: Sequence-length and cell-size sweeps for Case B classification (top) and R2 regression (bottom). Left: varying sequence length with cell size 3. Right: varying cell size with sequence length 120. Clean classification inputs are standard Gaussian; clean regression inputs are lognormal with standard Gaussian logarithms. Targets are computed before adding Gaussian input noise with standard deviation 0.05. Error bars denote the standard deviation across ten random seeds. followed by a linear classifier: Poolglobal → FC(1, 2), where FC denotes a fully connected layer mapping. No convolution, normalization, or pointwise activation is used, so that the pooling operator acts directly on the variables carrying the product structure. Across the reported single-seed sweeps over sequence length N ∈ {16, 32, 64, 128} and adjacent-site correlation ρ ∈ {0, 0.3, 0.5, 0.7, 0.9}, GMP achieves 100% accuracy. Average and max pooling remain near chance for independent inputs and both reach 60.4% accuracy at ρ = 0.9. Sweep details and full results appear in Appendix C, Table 3. Case B (local signed-product-sum classification). To test local multiplicative coarse-graining, i.i.d. we partition a clean sequence xi ∼ N (0, 1) into non-overlapping cells Cj of size k. Each cell contributes the signed geometric mean     Y X 1 gj =  sign(xi ) exp log(max{|xi |, ε}) , k i∈Cj

i∈Cj

PN/k

and the sample-level statistic is S = j=1 gj , y = 1{S < 0}. Thus the label depends on a sum of local, equal-weight multiplicative statistics. The classifier receives a noisy observation x̃i = xi + ηi , where ηi ∼ N (0, 0.052 ) independently; labels are always computed from the clean sequence. For every pooling primitive, we use the matched architecture Poollocal (k, stride = k) → Flatten → FC(N/k, 2). The local windows are therefore aligned exactly with the generative cells. As in Case A, there are no convolutional layers, normalizations, or activations before pooling. This deliberately minimal setting isolates the information preserved by the pooling primitive: on clean inputs with matching numerical conventions, GMP produces the local statistics gj directly, after which the linear classifier can learn their sum. Noisy observations test how accurately these local summaries retain the clean class signal. The classification panels of Fig. 1 show a consistent advantage for GMP over average and max pooling in the aligned settings. Accuracy remains high across the sequence-length sweep and declines as the cell size grows. The two sweeps illustrate how the operator’s multiplicative prior interacts with the number and size of local groups under input noise. Impact of Sign Information. To quantify the effect of sign information on classification accuracy, we perform an ablation study comparing signed and unsigned geometric-mean pooling (Table 1). We first embed each input sequence into a higher-dimensional space to increase model expressivity, using the architecture Conv1d(1, dembed , k = 1) → Pool → FC(dembed , 2), 5

with dembed = 32. We evaluate both Case A and Case B under the corresponding global and local pooling settings. The only difference between the pooling variants is whether the geometric mean is computed with or without the sign-parity term. Removing the sign-parity term reduces accuracy to approximately chance level in both tasks, demonstrating its importance in these settings. Table 1: Ablation of the sign-parity term in geometric mean pooling. Results are reported as mean ± standard deviation across 10 random seeds. The sequence length is 120, and the cell size is 3 for Case B. Setting

Accuracy 0.9330 ± 0.0659 0.5107 ± 0.0165 0.9856 ± 0.0047 0.5072 ± 0.0136

Case A, signed Case A, unsigned Case B, signed Case B, unsigned

4.2

Synthetic Magnitude Regression and Iterative Pooling

The synthetic regression tasks use MSE loss and a linear output layer. Lognormal clean inputs provide a controlled setting for studying multiplicative magnitudes. Local aggregation, learned embeddings, and input perturbations then probe how this structure is carried into the prediction. R1 (Global geometric-mean regression).P The clean task draws xi ∼ lognormal(0, 1) independently and sets y = exp N −1 i log xi . The model uses a learned pointwise Conv1d(1, dembed , k = 1), followed by variable global pooling and a linear regression head, with dembed = 1 for R1. The geometric mean of the raw inputs provides the exact target statistic when the clamp is inactive. The learned affine embedding and regression head allow the model to adapt this summary during training. For N ∈ {16, 32, 64, 128} in the reported single-seed experiment, GMP achieves R2 = 0.993– 1.000, compared with 0.609–0.638 for average pooling and 0.019–0.137 for max pooling. These results support the correspondence between geometric aggregation and the multiplicative target; average pooling also retains predictive information through correlations between arithmetic and geometric magnitudes. Full results appear in Appendix C, Table 4. R2 (Local cell log-mean regression): We use i.i.d. lognormal features, xi ∼ lognormal(0, 1), with cell size c ∈ {2, 3, 4, 5, 6, 8, 10} and sequence length N ∈ {18, 36, 54, 72, 90,108, 126, 144, 162,180} (fixed c = 3) or N = 120 (varying c). Each cell conP P 1 tributes sj = exp 1c i∈cellj log |xi | , and the target y = ncell j sj is computed from the clean sequence. The architecture is local pooling (kernel and stride c), flattening, and FC(ncell , 32) → ReLU → FC(32, 1). Pooling is selected from GMP, average, and max and acts directly on the noisy sequence, without a preceding convolutional embedding. The sweeps appear in Fig. 1 (bottom panels). Observations are perturbed as x̃i = xi + ηi , with ηi ∼ N (0, 0.052 ). R3 (Noise robustness): Clean lognormal features are corrupted by multiplicative noise applied to 2 every feature, followed by additive noise: x̃ = x · exp(ϵmul ) + ϵadd , ϵmul ∼ N (0, σmul ), ϵadd ∼ 2 N (0, σadd ). The target is the geometric mean of the clean features. The model uses the same architecture as R1, but with dembed = 32 to increase the capacity of the learned representation before global pooling. Figure 2 compares responses to multiplicative and additive perturbations. Multiplicative noise shifts raw log-magnitudes additively, directly matching the coordinates used by GMP. Additive noise probes a different perturbation and can also change input signs. The sweep examines these two effects within the same prediction task. R4 (Iterative preservation of a multiplicative statistic). Clean lognormal sequences of length L = 256 are perturbed once by multiplying a randomly selected fraction p = 25% of entries by exp(η), with η ∼ N (0, 1). Repeated pooling uses kernel and stride 2, reducing 256 entries to one in eight steps. At every step, the evaluator takes a global geometric-mean readout of the remaining entries for all three pooling methods and compares it with the clean global geometric mean using MAE. 6

0.2

0.95

0.95

0.94

0.93

0.91

0.76

0.98

0.97

0.97

0.96

0.94

0.81

0.99

0.99

0.98

0.95

0.95

0.88

0.0

0.01

0.02

0.05

0.1

0.2

Additive noise add

0.19

0.41

0.45

0.45

0.45

0.45

0.45

0.54

0.58

0.58

0.59

0.59

0.59

0.56

0.61

0.61

0.61

0.61

0.61

0.58

0.57

0.57

0.57

0.56

0.56

0.0

0.01

0.02

0.05

0.1

0.2

Additive noise add

0.00

0.00

0.04

0.04

0.04

0.04

0.03

0.02

0.08

0.06

0.06

0.06

0.05

0.04

0.6

0.10

0.09

0.08

0.07

0.06

0.05

0.4

0.10

0.10

0.10

0.07

0.06

0.05

0.10

0.10

0.09

0.08

0.07

0.06

0.0

0.01

0.02

0.05

0.1

0.2

0.8

R2

0.70

0.19

2.0

0.74

0.19

0.00

1.0

0.75

0.19

0.00

0.5

0.77

0.19

0.00

0.2

0.76

0.18

1.0

0.01

Multiplicative noise mul

0.5

0.77

0.02

0.1

0.42

0.02

0.0

0.42

2.0

0.45

0.02

1.0

0.47

0.02

0.5

0.46

0.02

0.2

1.0

0.43

Max Pooling

0.02

Multiplicative noise mul

0.13

0.1

0.14

0.0

2.0

0.15

Multiplicative noise mul

0.19

0.1

Average Pooling

0.16

0.0

Geometric Pooling 0.10

Additive noise add

0.2 0.0

Mean Absolute Error (MAE)

Figure 2: R3: global regression under multiplicative and additive input noise, x̃ = x exp(ϵmul ) + ϵadd , with log x ∼ N (0, 1) and independent Gaussian perturbations of standard deviations σmul and σadd . Over the tested noise ranges, GMP maintains higher predictive performance than max pooling, including at low noise levels. Geometric Pooling Average Pooling Max Pooling

101

100

10 1 1

2

3

4

5

RG Step (Iteration)

6

7

8

Figure 3: R4: preservation of a multiplicative observable across pooling levels. MAE is evaluated against the clean geometric-mean target (σ = 1, p = 25%, L = 256), using a global geometric readout after each local pooling step for all methods. GMP maintains an approximately constant readout error across the hierarchy. Figure 3 illustrates the connection between hierarchical consistency and preservation of a multiplicative observable. The GMP curve remains approximately constant across successive pooling steps, whereas average and max pooling change the common geometric readout. This behavior reflects the statistic preserved by each block transformation: GMP composes geometric summaries, average pooling composes arithmetic means, and max pooling composes extrema. 4.3

Image Classification and Activation Compatibility

We evaluate GMP on MNIST [28], Fashion-MNIST [52], and CIFAR-10 [26] using convolutional classifiers with matched channel dimensions, optimizers, and training procedures. The experiments isolate three different placements of the variable pooling primitive: global aggregation, repeated local spatial reduction, and synchronized local-plus-global pooling. For each configuration, we compare GMP, max pooling, and average pooling. The three configurations specify where the variable pooling operator is used: Configuration Local pooling Global fixed max Local variable Synchronized variable

Global pooling variable fixed max same as local

The variable operator is selected from {GMP, Max, Avg}. Thus, the Global configuration isolates the final spatial summary, the Local configuration tests repeated local coarse-graining, and the Synchronized configuration applies the same pooling operator at both levels. All backbones use 7

MNIST

Test Accuracy

1.0

Fashion-MNIST

1.0

1.0

0.9

0.9

0.9

0.8

0.8

0.8

0.7

0.7

0.7

0.6

0.6

0.6

0.5

0.5

0.5

0.4

0.4 Global Pooling

Local Pooling

Sync Pooling

CIFAR-10 Geo

Max

Avg

0.4 Global Pooling

Local Pooling

Sync Pooling

Global Pooling

Local Pooling

Sync Pooling

Figure 4: Validation accuracy on MNIST, Fashion-MNIST, and CIFAR-10 across ten random seeds. Each panel reports the three pooling primitives (GMP, max, and average) under three architectures: variable global pooling with fixed intermediate max pooling, variable local pooling with fixed global max pooling, and synchronized local-plus-global pooling. Local GMP is generally more competitive than global GMP on MNIST and Fashion-MNIST, whereas the gap between GMP and the gap between GMP and the average- and max-pooling baselines increases on CIFAR-10. Error bars denote standard deviation across 10 seeds. All models are trained with cross-entropy loss and Adam [23] using batch size 128 and learning rate 10−3 .

convolution–ReLU blocks, with matched architectures and training procedures within each dataset. Full architecture details appear in Appendix D. Figure 4 shows that local GMP remains competitive with max and average pooling on MNIST and Fashion-MNIST, whereas global GMP performs less consistently across the tested settings. The performance gap is larger on CIFAR-10. The sensitivity of signed GMP to exact zeros after ReLU offers a possible explanation for the weaker global results. The results therefore support a regime-specific interpretation of GMP: it is a useful structural prior when multiplicative composition is plausible, but it is not generally interchangeable with additive pooling for generic image classification. 4.4

Lipophilicity Regression

We evaluate on the Lipophilicity regression task from the MoleculeNet molecular benchmark suite [51], specifically version 1 (V1), curated from ChEMBL with 4,200 small organic molecules. Molecules are partitioned by a Bemis–Murcko scaffold split [3], with approximately 80% of scaffolds assigned to training and the remaining scaffolds assigned to testing, so the test set contains molecular scaffolds absent from training. The dataset label is a log-scale lipophilicity value, y = log D, the experimental octanol/water distribution coefficient at pH 7.4 (range approximately [−1.5, 4.5]). To restore the true multiplicative structure of the molecular property, we transform the regression target as z = exp(y) before optimization and evaluation, so the reported RMSE and R2 are measured on the z scale (range approximately [0.2, 90]) rather than the log D scale. Each molecule is represented by a 2,048-bit Morgan fingerprint (Morgan FP) with radius 2 (ECFP4) [36]. Because the raw fingerprint is fixed, high-dimensional, and sparse, our 2D CNN maps it through a fully connected projection layer, R2048 → R1024 followed by ReLU, giving the network learnable capacity to re-express the fingerprint into a denser representation rather than operating directly on raw binary bits. The projected vector is reshaped into a single-channel 32×32 latent feature map, to which a softplus activation is applied before three convolutional blocks, each consisting of a 3 × 3 convolution (channel widths 1 → 32 → 64 → 128), the same activation, and a 2 × 2 local pooling operator with stride 2 (reducing the spatial resolution 32 × 32 → 16 × 16 → 8 × 8 → 4 × 4). The resulting 128-channel feature map is reduced by a global pooling operator to a 128-dimensional vector, which is passed to a two-layer regression head, FC(128, 32) → ReLU → FC(32, 1). For each experiment, the local and global pooling operators are independently selected from geometric, max, and average pooling. The learned 32 × 32 layout supplies latent neighborhoods for convolutional processing, with locality defined by the representation. We compare the 2D CNN against four baselines trained on the same scaffold split, summarized in Table 2: Morgan fingerprint features with XGBoost [7], a Random Forest regressor [5], and 8

Lipophilicity Regression (Morgan CNN 2D)

0.400

Global: Geo

Global: Max

Global: Avg

0.375 0.350

Test R 2

0.325 0.300 0.275 0.250 0.225 0.200

Local: Geo

Local: Max

Local: Avg

Figure 5: Test R2 on the auxiliary target z = exp(y) for the Morgan-fingerprint 2D CNN under all combinations of local and global pooling operators. Each molecule is represented by a radius-2, 2,048-bit Morgan fingerprint, projected to a learned 32 × 32 latent feature map, and processed by a three-block convolutional hierarchy. Values and error bars denote the mean and standard deviation over ten random seeds.

a fully connected network, as well as a learned SMILES [45] token embedding passed through the same fully connected head. The three Morgan-fingerprint baselines achieve similar performance and outperform the tested SMILES-embedding baseline, suggesting that Morgan fingerprints provide a more effective representation in this experimental setting. Table 2: Lipophilicity regression performance on the original logarithmic label scale, using a scaffold split. Baseline Morgan FP + XGBoost Morgan FP + Random Forest Morgan FP + FC SMILES Embedding + FC

RMSE 0.8870 0.8718 0.8752 1.1596

R2 0.4631 0.4813 0.4773 0.0824

Fig. 5 shows the R2 performance of the Morgan FP 2D CNN under all combinations of local and global pooling operators. Global GMP outperforms max and average pooling, while local GMP becomes detrimental when combined with global max pooling. The best pooling configuration achieves R2 = 0.3545 on the auxiliary target z. Because the baselines report results on the log scale of z, those values provide an indirect comparison with this experiment. The results overall highlight the importance of matching the pooling strategy to the underlying algebraic structure of the target property. As shown in Appendix E, training the same architecture on the original label scale changes the ranking, with max pooling achieving the strongest performance.

5

Conclusion

We introduced GMP, a parameter-free pooling operator that combines sign parity with equal-weight geometric magnitudes. It preserves the corresponding global statistic under complete, equal-sized, non-overlapping hierarchical pooling with a shared clamp and no intervening transformations. Synthetic experiments demonstrate its utility for prescribed multiplicative signals, while image and molecular experiments show that its effectiveness depends on representation, activation, pooling placement, and target parameterization. GMP therefore provides a structural prior for sign-sensitive or equal-weight multiplicative tasks. For positive targets, predicting their logarithms with additive aggregation offers an alternative. Extending this framework to complex-valued representations [42], developing learnable hybrid operators, and testing them on physical [49, 6, 34], chemical [14, 37, 54], and other scientific datasets are important directions for future work. 9

Acknowledgments and Disclosure of Funding The work at UT Knoxville was primarily supported by the National Science Foundation Materials Research Science and Engineering Center program through the UT Knoxville Center for Advanced Materials and Manufacturing (DMR-2309083). Computations were performed using the University of Tennessee Infrastructure for Scientific Applications and Advanced Computing (ISAAC) computational resources. All experiments used an NVIDIA RTX A6000 GPU. Reproducibility statement The results presented in this paper are readily reproducible based on the descriptions, algorithms, and code provided in the main text, appendix. The codes are publicly available at https://github. com/angkun-research/GeometricMeanPooling.

References [1] Ian Affleck, Tom Kennedy, Elliott H. Lieb, and Hal Tasaki. Rigorous results on valencebond ground states in antiferromagnets. Physical Review Letters, 59(7):799–802, 1987. doi: 10.1103/PhysRevLett.59.799. [2] Philip W. Anderson. Resonating valence bonds: A new kind of insulator. Materials Research Bulletin, 8(2):153–160, 1973. doi: 10.1016/0025-5408(73)90167-0. [3] Guy W. Bemis and Mark A. Murcko. The properties of known drugs. 1. molecular frameworks. Journal of Medicinal Chemistry, 39(15):2887–2893, 1996. doi: 10.1021/jm9602928. [4] Filippo Maria Bianchi, Carlo Abate, Ivan Marisca, et al. Torch geometric pool: the pytorch library for pooling in graph neural networks. arXiv preprint arXiv:2512.12642, 2025. [5] Leo Breiman. Random forests. Machine Learning, 45(1):5–32, 2001. doi: 10.1023/A: 1010933404324. [6] Giuseppe Carleo and Matthias Troyer. Solving the quantum many-body problem with artificial neural networks. Science, 355(6325):602–606, 2017. doi: 10.1126/science.aag2302. [7] Tianqi Chen and Carlos Guestrin. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, pages 785–794, 2016. doi: 10.1145/2939672.2939785. [8] Yin Cui, Feng Zhou, Jiang Wang, Xiao Liu, Yuanqing Lin, and Serge Belongie. Kernel pooling for convolutional neural networks. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 2921–2930, 2017. [9] Paul A. M. Dirac. On the theory of quantum mechanics. Proceedings of the Royal Society A, 112(762):661–677, 1926. doi: 10.1098/rspa.1926.0133. [10] Richard Durbin and David E. Rumelhart. Product units: A computationally powerful extension to backpropagation networks. Neural Computation, 1(1):133–142, 1989. doi: 10.1162/neco. 1989.1.1.133. [11] Hayoung Eom and Heeyoul Choi. Alpha-integration pooling for convolutional neural networks. arXiv preprint arXiv:1811.03436, 2018. [12] Alexander L. Fetter and John Dirk Walecka. Quantum Theory of Many-Particle Systems. McGraw-Hill, New York, 1971. [13] W. M. C. Foulkes, L. Mitas, R. J. Needs, and G. Rajagopal. Quantum monte carlo simulations of solids. Rev. Mod. Phys., 73:33–83, Jan 2001. doi: 10.1103/RevModPhys.73.33. URL https://link.aps.org/doi/10.1103/RevModPhys.73.33. [14] Justin Gilmer, Samuel S. Schoenholz, Patrick F. Riley, Oriol Vinyals, and George E. Dahl. Neural message passing for quantum chemistry. Proceedings of the 34th International Conference on Machine Learning, 70:1263–1272, 2017. 10

[15] Yinzheng Gu, Chuanpeng Li, and Jinbin Xie. Attention-aware generalized mean pooling for image retrieval. arXiv preprint arXiv:1811.00202, 2018. [16] Caglar Gulcehre, Kyunghyun Cho, Razvan Pascanu, and Yoshua Bengio. Learned-norm pooling for deep feedforward and recurrent neural networks. arXiv preprint arXiv:1311.1780, November 2013. Comments: ECML/PKDD 2014. [17] Gareth Hallam and Peter Whitfield. Compact neural networks based on the multiscale entanglement renormalization ansatz (mera). In Proceedings of the 25th Australian Pattern Recognition Conference (BMVC), page Paper 103, 2018. [18] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 770–778, 2016. doi: 10.1109/CVPR.2016.90. [19] Niklas Heim, Tomas Pevny, and Vasek Smidl. Neural power units. In Advances in Neural Information Processing Systems, volume 33, 2020. URL https://papers.nips.cc/paper_ files/paper/2020/hash/48e59000d7dfcf6c1d96ce4a603ed738-Abstract.html. [20] Jan Hermann, Zachary Schätzle, and Frank Noé. Ab initio quantum chemistry with neural-network wavefunctions. Nature Chemistry, 12:891–897, 2020. doi: 10.1038/ s41557-020-0544-y. [21] Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory. Neural Computation, 9 (8):1735–1780, 1997. doi: 10.1162/neco.1997.9.8.1735. [22] L. P. Kadanoff. Scaling laws for ising phase transitions or how everything goes wrong when the dimension is two. Physics, 1(5):263–272, 1966. doi: 10.1103/Physics.1.263. [23] Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In International Conference on Learning Representations, 2015. [24] Byungsoo Ko, Han-Gyu Kim, Byeongho Heo, Sangdoo Yun, Sanghyuk Chun, Geonmo Gu, and Wonjae Kim. Group generalized mean pooling for vision transformer. arXiv preprint arXiv:2212.04114, 2022. [25] Ryan Kortvelesy, Steven Morad, and Amanda Prorok. Generalised f-mean aggregation for graph neural networks. In Advances in Neural Information Processing Systems, volume 36, 2023. URL https://papers.nips.cc/paper_files/paper/2023/hash/ 6c78ae0c1140902bf3a430b1725bcc4e-Abstract-Conference.html. [26] Alex Krizhevsky. Learning multiple layers of features from tiny images. Technical report, University of Toronto, 2009. [27] Alex Krizhevsky, Ilya Sutskever, and Geoffrey E. Hinton. Imagenet classification with deep convolutional neural networks. In Advances in Neural Information Processing Systems, volume 25, 2012. [28] Yann LeCun, Leon Bottou, Yoshua Bengio, and Patrick Haffner. Gradient-based learning applied to document recognition. Proceedings of the IEEE, 86(11):2278–2324, 1998. doi: 10.1109/5.726791. [29] Teng Li, Z. Meng, Bingbing Ni, Jianbing Shen, and Meng Wang. Robust geometric ℓp -norm feature pooling for image classification and action recognition. Image and Vision Computing, 55:1–12, 2016. [30] Min Lin, Qiang Chen, and Shuicheng Yan. Network in network. In International Conference on Learning Representations, 2014. [31] Andreas Madsen and Alexander Rosenberg Johansen. Neural arithmetic units. In International Conference on Learning Representations, 2020. URL https://arxiv.org/abs/ 2001.05016. [32] Risto Miikkulainen. Subsymbolic Natural Language Processing: An Integrated Model of Scripts, Lexicon, and Memory. MIT Press, 1996. 11

[33] Francesco Milano, Antonio Loquercio, Antoni Rosinol, Davide Scaramuzza, and Luca Carlone. Primal-dual mesh convolutional neural networks. In Advances in Neural Information Processing Systems, volume 33, 2020. [34] David Pfau, James S. Spencer, James Matthews, and W. M. C. Foulkes. Ab initio solution of the many-electron schrödinger equation with deep neural networks. Physical Review Research, 2(3):033429, 2020. doi: 10.1103/PhysRevResearch.2.033429. [35] Filip Radenovi’c, Giorgos Tolias, and Ondřej Chum. Fine-tuning cnn image retrieval with no human annotation. arXiv preprint arXiv:1711.02512, 2018. [36] David Rogers, Matthew Brown, and Matthias Hahn. Extended-connectivity fingerprints. Journal of Chemical Information and Modeling, 50(5):742–754, 2010. doi: 10.1021/ci100050t. [37] Kristof T. Schütt, Anders Arbab-Birkirkjær, et al. Schnet: A continuous-filter convolutional neural network for modeling quantum interactions. Advances in Neural Information Processing Systems, 30, 2017. [38] Si-Baek Seong, Chongwon Pae, and Hae-Jeong Park. Geometric convolutional neural network for analyzing surface-based neuroimaging data. Frontiers in Neuroinformatics, 12:42, 2018. doi: 10.3389/fninf.2018.00042. [39] John C. Slater. The theory of complex spectra. Physical Review, 34(10):1293–1322, 1929. doi: 10.1103/PhysRev.34.1293. [40] Mengdan Tang and Reinhard Klette. Learning scale-variant and scale-invariant features for deep image classification. arXiv preprint arXiv:1602.01255, February 2016. [41] Zhi Tao, Xiaoyu Chen, Huiling Li, Xinyu Yang, Yuncan Liu, and Xiaomin Zhang. Pooling operations in deep learning: From “invariable” to “variable”. BioMed Research International, 2022:4067581, 2022. doi: 10.1155/2022/4067581. [42] Chiheb Trabelsi, Olexa Bilaniuk, Ying Zhang, Dmit cork Serdyuk, Sandeep Subramanian, João Felipe Santos, Soroush Mehri, Negar Rostamzadeh, Yoshua Bengio, and Christopher J. Pal. Deep complex networks. International Conference on Learning Representations, 2018. [43] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, volume 30, 2017. [44] Banching Wang. Pooling as renormalization: The renormalization group theory of deep learning interpreted, 2017. [45] David Weininger. Smiles, a chemical language and information system. 1. introduction to methodology and encoding rules. Journal of Chemical Information and Computer Sciences, 28(1):31–36, 1988. doi: 10.1021/ci00057a005. [46] Chaoqun Weng, Hongxing Wang, and Junsong Yuan. Learning weighted geometric pooling for image classification. In 2013 IEEE International Conference on Image Processing, pages 3713–3717. IEEE, 2013. [47] K. G. Wilson. Renormalization group and critical phenomena. i. renormalization group and the kadanoff scaling picture. Physical Review B, 4(9):3174–3183, 1971. doi: 10.1103/PhysRevB. 4.3174. [48] Ang-Kun Wu, Benedikt Kloss, Wladislaw Krinitsin, Matthew T. Fishman, J. H. Pixley, and E. M. Stoudenmire. Disentangling interacting systems with fermionic gaussian circuits: Application to quantum impurity models. Phys. Rev. B, 111:035119, Jan 2025. doi: 10.1103/PhysRevB.111.035119. URL https://link.aps.org/doi/10.1103/PhysRevB. 111.035119. [49] Ang-Kun Wu, Louis Primeau, Jingtao Zhang, Kai Sun, Yang Zhang, and Shi-Zeng Lin. Modeling quantum geometry for fractional chern insulators with unsupervised learning. npj Comput. Mater., May 2026. doi: 10.1038/s41524-026-02155-1. URL https://www.nature.com/ articles/s41524-026-02155-1. 12

[50] Ang-Kun Wu, Louis Primeau, Yixin Zhang, Jingtao Zhang, Adrian Del Maestro, and Yang Zhang. Compact spin-charge separated neural quantum states for valence-bond states. arXiv preprint arXiv:2606.17045, June 2026. [51] Zhenqin Wu, Bharath Ramsundar, Evan N Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S Pappu, Karl Leswing, and Vijay Pande. Moleculenet: a benchmark for molecular machine learning. Chemical science, 9(2):513–530, 2018. doi: 10.48550/arXiv.1703.00564. [52] Han Xiao, Kashif Rasul, and Roland Vollgraf. Fashion-mnist: A novel image dataset for benchmarking machine learning algorithms. arXiv preprint arXiv:1708.07747, 2017. [53] Yutong Xu and Zhi Liu. Geometric pooling: Maintaining more representative information for graph node classification. Neurocomputing, 579:127439, 2024. doi: 10.1016/j.neucom.2024. 127439. [54] Kevin Yang, Kyle Swanson, Wengong Jin, Connor Coley, E. S. Jensen, et al. Analyzing learned molecular representations for property prediction. Journal of Chemical Information and Modeling, 59(8):3370–3388, 2019. doi: 10.1021/acs.jcim.9b00237. [55] Manzil Zaheer, Satwik Kottur, Siamak Ravanbakhsh, Barnabas Poczos, Ruslan Salakhutdinov, and Alexander J. Smola. Deep sets. In Advances in Neural Information Processing Systems, volume 30, 2017. URL https://papers.nips.cc/paper_files/paper/2017/ hash/f22e4747da1aa27e363d86d40ff442fe-Abstract.html. [56] Muhua Zheng, Simon L. Evans, Gilles Humair, Michael Small, and Tianhou Zhou. Geometric renormalization of weighted networks. Communications Physics, 7(1):97, March 2024. doi: 10.1038/s42005-024-01589-7.

13

A

Implementation details and pseudocode

GMP operates independently on each batch item and channel. In local mode, it reduces each nonoverlapping (or strided) window; in global mode, the window is the full remaining spatial domain. For nonzero inputs, the signed output retains the parity of negative values, while the magnitude is evaluated in log-space to avoid forming the full product. Algorithm 1 Signed Geometric Mean Pooling (GMP) Require: Input tensor X of shape (B, C, L) in 1D or (B, C, H, W ) in 2D; window size K; stride S; magnitude lower bound ε Ensure: Pooled tensor Y 1: if global pooling is selected then 2: Let W be the full sequence (1D) or full spatial map (2D) 3: else 4: Extract sliding windows {Wj } of size K and stride S ▷ S = K by default 5: end if 6: for each Q batch item b, channel c, and window Wj do 7: s ← x∈Wj sign(x)  P 8: ℓ ← |W1j | x∈Wj log max(|x|, ε) 9: Yb,c,j ← s · exp(ℓ) 10: end for 11: return Y For 1D local pooling, Wj contains k sequence entries and the output has shape (B, C, Lout ), where Lout = 1 + ⌊(L − k)/s⌋. For 2D local pooling, each window contains kh kw spatial entries and the output has shape (B, C, Hout , Wout ). Global pooling returns (B, C, 1) in 1D and (B, C, 1, 1) in 2D. The implementation uses torch.sign, whose autograd derivative is zero, and a clamped logmagnitude computation. For a window of nonzero inputs, away from |xj | = ε, its derivative is  ∂Gε Gε (x)/(kxj ), |xj | > ε, = 0, 0 < |xj | < ε. ∂xj The sign is locally constant within each orthant; learning there proceeds through magnitudes. Small inputs above the clamp can amplify derivatives relative to the output scale, whereas inputs below it have zero magnitude derivatives. If any input equals zero, the sign product vanishes and the implemented backward pass yields zero derivatives for every input in that window, including nonzero entries. Clamping also makes the signed operator discontinuous at zero. With k − 1 inputs equal to 1 and 0 < ε < 1, the remaining input t gives  sign(t) ε1/k , 0 < |t| < ε, Gε (t, 1, . . . , 1) = 0, t = 0. Thus arbitrarily small sign changes can produce a finite output jump. ReLU removes negative signs but introduces exact zeros; Softplus gives positive values in exact arithmetic and removes the negative-parity mechanism. A.1

Image-classification architectures

B

Proof: hierarchical consistency of signed GMP

Let x1 , . . . , xN be the inputs and fix ε > 0. Partition the indices into m non-overlapping blocks of size k, so that N = mk. Define ! ! Y  1 X Gε (B) = sign(xi ) exp log max(|xi |, ε) . |B| i∈B

i∈B

14

If any block contains a zero, then its sign product is zero and its GMP output is zero. Consequently, the sign product of the second-level GMP is also zero, which agrees with the direct GMP because the original input contains a zero. It remains to consider the case in which all inputs are nonzero. For each block Bj , every magnitude max(|xi |, ε) is at least ε, and hence |Gε (Bj )| ≥ ε. Therefore, the outer clamp is inactive:  max |Gε (Bj )|, ε = |Gε (Bj )|. Thus,  Gε Gε (B1 ), . . . , Gε (Bm )     m m Y X X Y  1 1 log max(|xi |, ε)  = sign(xi ) exp  m k j=1 j=1 i∈Bj i∈Bj ! " # N N Y  1 X log max(|xi |, ε) = sign(xi ) exp N i=1 i=1 = Gε (x1 , . . . , xN ). Hence the identity holds in exact arithmetic for the clamped definition, including zero-valued inputs, when equal-sized, non-overlapping blocks cover every input, the same ε is used at both levels, and sign(0) = 0. It extends inductively to such hierarchies without intervening feature transformations. Floating-point implementations may differ by rounding.

C

Additional results for global synthetic tasks

This section reports the complete results for Case A and R1, which test global pooling on prescribed multiplicative targets. Both tables report single-seed experiments using seed 42. C.1

Case A: Global signed-product classification

Case A predicts the sign parity of a sequence using global pooling followed by a linear classifier, without a preceding convolution, normalization, or pointwise activation. We vary the sequence length over N ∈ {16, 32, 64, 128} with independent standard Gaussian inputs. A separate sweep fixes N = 32 and varies the adjacent-site correlation over ρ ∈ {0, 0.3, 0.5, 0.7, 0.9}. Correlated sequences are generated recursively as p x1 ∼ N (0, 1), xi = ρxi−1 + 1 − ρ2 zi , where zi ∼ N (0, 1) are independent innovations. The label is determined by the sign of the resulting sequence product. Table 3 shows that GMP achieves perfect accuracy in all reported settings. Average and max pooling remain near chance for independent inputs, while both reach 0.6040 accuracy at ρ = 0.9. C.2

R1: Global geometric-mean regression

R1 predicts the geometric mean of independent lognormal(0, 1) inputs. The model applies a learned pointwise convolution with dembed = 1, global pooling, and a linear regression head. Inputs are noiseless, and the sequence lengths are N ∈ {16, 32, 64, 128}. Table 4 reports both R2 and MSE. GMP achieves R2 between 0.993 and 1.000, whereas average pooling achieves 0.609–0.638 and max pooling achieves 0.019–0.137. This experiment tests geometric aggregation on an explicitly matched target. 15

Table 3: Case A classification accuracy with seed 42. The correlation sweep fixes N = 32; the sequence-length sweep uses independent inputs (ρ = 0). Sweep

Value

Max acc

Average acc

GMP acc

ρ-sweep (seed=42, seq_len=32)

0.0 0.3 0.5 0.7 0.9

0.4680 0.5080 0.5000 0.5130 0.6040

0.5390 0.4860 0.4900 0.4830 0.6040

1.0000 1.0000 1.0000 1.0000 1.0000

Seq_len sweep (seed=42, ρ = 0)

16 32 64 128

0.5190 0.4680 0.4860 0.5050

0.5190 0.5390 0.4710 0.5050

1.0000 1.0000 1.0000 1.0000

Table 4: R1 global geometric-mean regression on noiseless lognormal inputs with seed 42. Results are reported as R2 and MSE; MSE values are scaled by 103 . N 16 32 64 128

D

Max 0.137 0.081 0.071 0.019

R2 Avg 0.609 0.620 0.631 0.638

GMP 1.000 1.000 0.993 1.000

MSE (×10−3 ) Max Avg GMP 60.7 27.5 0.0 29.3 12.1 0.0 14.6 5.8 0.0 7.9 2.9 0.0

Image-classification architectures

CNN architectures. We use convolutional classifiers in which the pooling operators are the only experimental variables. For MNIST and Fashion-MNIST, the backbone contains three convolution– ReLU–pooling blocks with 32 channels, followed by a fourth convolution–ReLU block. Each convolution uses a 3 × 3 kernel with padding one. Local pooling uses a non-overlapping 2 × 2 window with stride two, reducing the spatial resolution from 28 × 28 to 3 × 3. The resulting 32-channel representation is globally pooled and passed to a linear classifier with ten outputs. For CIFAR-10, we use the same pooling configurations with a wider backbone. The convolutional channel widths are 3 → 48 → 64 → 96 → 128 → 128. The first four convolutional blocks include 2 × 2 stride-two pooling and reduce the spatial resolution from 32 × 32 to 2 × 2; the final convolutional block performs feature refinement without further downsampling. The resulting 128dimensional representation is globally pooled and mapped to ten class logits by a linear layer.

E

Lipophilicity regression on the original label scale

Figure 6 presents the CNN pooling comparison on the original log-label scale. Max pooling gives the strongest performance in this setting, while the auxiliary exponential-target experiment favors a different configuration. This shift illustrates how target parameterization and its associated error weighting interact with the choice of pooling. In both settings, Softplus features allow the geometricmagnitude component to be examined within the same convolutional architecture.

16

Lipophilicity Regression (Morgan CNN 2D)

0.600

Global: Geo

Global: Max

Global: Avg

0.575 0.550

Test R 2

0.525 0.500 0.475 0.450 0.425 0.400

Local: Geo

Local: Max

Local: Avg

Figure 6: Test R2 for lipophilicity regression on the original label scale using CNNs with different combinations of local and global pooling. Error bars denote the standard deviation across ten random seeds. Max pooling achieves the highest R2 in this setting.

17

Record · ID 1006855 · SHA-256 afa7069af4364bfb
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.