ConceptioArchivearXiv CS
arXiv CSopen access

Probing Chemical Language Models: Effects of Pre-training and Fine-tuning

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

Probing Chemical Language Models: Effects of Pre-training and Fine-tuning Anna Karnysheva1,2 , Dietrich Klakow1,2,3∗ , Ji-Ung Lee1 * , 1 RTG Neuroexplicit Models 2 Spoken Language Systems, 3 PharmaScienceHub (PSH) Saarland University [email protected]

arXiv:2607.02140v1 [cs.LG] 2 Jul 2026

Abstract

been explored for MPP tasks, however, they still frequently fall behind feature-based models (Dias et al., 2023; Xia et al., 2023; Sadeghi et al., 2024). Moreover, they often exhibit poor out-ofdistribution generalization (Tossou et al., 2024) and are not evaluated on regression tasks which make up a substantial portion of MPP tasks (Liu et al., 2025). While a few works have tried to establish a better understanding of the shortcomings of DNNs by probing their representation for individual molecular substructures, they are often limited to a small set of molecular substructures and graphbased models which are often trained for individual MPP tasks (Akhondzadeh et al., 2023; Wang* et al., 2023; Volkov et al., 2022). In this work, we focus on CLMs trained on linearized molecular representations (i.e., SMILES, Weininger 1988) which have been frequently used (Singh et al., 2026), but not well studied, especially regarding whether they learn to capture molecular substructures during pre-training (RQ1) and how fine-tuning on chemical downstream tasks affects these representations (RQ2). Our systematic study based on a new probing dataset comprising 78 molecular substructures evaluated across eight pre-trained (PT) and six randomly initialized (RI) models reveals that:

Chemical language models (CLMs) are trained with linearized representations such as SMILES, yet it remains unclear which chemically meaningful substructures they encode. To foster a better understanding of CLMs, we conduct a systematic study and probe for 78 molecular substructures across eight pre-trained and six randomly initialized models. We furthermore study how fine-tuning on chemical downstream tasks affects the learned representations of molecular substructures. Our results show that pre-training generally improves molecular structure awareness of CLMs, particularly in the upper layers. Moreover, randomly initialized models already encode ring structures well in the first layer. Our analysis on two chemical downstream tasks further reveals that, interestingly, fine-tuning affects task-relevant molecular substructures more than others, indicating that the changes in the representations follow chemical theory.1

1

Introduction

Drug discovery is an inherently expensive process that includes labor- and time-intensive steps such as designing molecules that are both effective against a target disease and can be safely administered to humans (Jia et al., 2020). One important task in this process is molecular property prediction (MPP, Shen and Nicolaou 2019), i.e., to reliably predict properties such as the lipophilicity, solubility, permeability, bioactivity, or toxicity of a molecule.2 Many deep neural network (DNN) architectures—including general-purpose large language models (LLMs, Zhao et al. 2023), GNNs (Scarselli et al., 2009), graph transformers (Yun et al., 2019), and sequence-based chemical language models (CLMs, Wang et al. 2019)—have

• Pre-training improves molecular structure awareness towards the upper layers. Also, all molecular substructures exhibit a change larger than ±1% in at least one model. • RI models already encode ring structures well, but not other substructures. • Some molecular substructures are unlearned in all models during pre-training. Studying how fine-tuning on lipophilicity and solubility prediction affects the representations of molecular substructures reveals that:

* Co senior authors. 1 2

Code and data will be released under open source licenses. We provide an introduction into chemistry in Section A.

1

Decoder-only models While early works utilize decoder-only models to generate new molecules (Xue et al., 2021; Bagal et al., 2022; Wang et al., 2023), later studies consider them for other tasks such as MPP, CRP, and MO (He et al., 2022; Mazuz et al., 2023). Others have even devised novel text-centric tasks such as molecule captioning (Edwards et al., 2021) or augmented tasks with textual instructions (Edwards et al., 2022; Wang et al., 2023; Liu et al., 2023b; Christofidellis et al., 2023; Fang et al., 2024; Lin et al., 2026).

• Pre-training increases the robustness of molecular substructures during fine-tuning. • Fine-tuning affects the representations less compared to pre-training with changes occuring more frequently in the upper layers. • Molecular substructures that are theoretically more relevant for lipophilicity and solubility prediction are more affected by fine-tuning. Finally, we showcase how probing can be used to identify molecular substructures on which models have not been sufficiently trained, and how to mitigate this by further pre-training them on molecules that include these molecular substructures.

2

Related Work

2.1

Molecular Representation Learning

Limitations Despite all efforts to utilize LLMs in chemistry (Wang et al., 2023; Ye et al., 2025; Xian et al., 2025), recent works found that both chemical and general-purpose LLMs struggle to understand molecular structure (Jang et al., 2025; Ganeeva et al., 2024) or are outperformed by simple baselines (Guo et al., 2023).3 For instance, Xian et al. (2025) show that GNNs outperform general-purpose LLMs on classification as well as regression tasks such as lipophilicity, which constitute a large portion of MPP tasks (Liu et al., 2025).

Molecules can be represented in various ways in order to be processed by models or hand-crafted algorithms. The choice of representation directly affects what architectures are suited; e.g., representing molecules as graphs enables the use of different kinds of GNNs. To train language models, we use linearized molecule representations such as SMILES (Weininger, 1988). For chemistry, works have trained encoder-decoder, encoder-only, and decoder-only models.

2.2

Probing

Probing is widely used in NLP to investigate the extent to which LMs capture linguistic knowledge, such as syntactic (Jawahar et al., 2019; Tenney et al., 2019b; Liu et al., 2019; Hou and Sachan, 2021) or semantic information (Tenney et al., 2019b). Typically, probing involves training a classifier for a specific probing task (e.g., partof-speech tagging) using hidden representations extracted from a pre-trained LM (Belinkov, 2021).

Encoder-decoder models Works have utilized encoder-decoder models in tasks that mirror their sequence-to-sequence nature (Schwaller et al., 2019; Irwin et al., 2022; Lu and Zhang, 2022). Example tasks are chemical reaction prediction (CRP), i.e., predicting the outputs for a given set of inputs (Fooshee et al., 2018) and molecular optimization (MO), where an input molecule is altered to achieve desired properties (He et al., 2021).

Probing for linguistic knowledge Various works localize linguistic knowledge in pre-trained LMs (particularly BERT), attributing syntax and semantics to different layers (Peters et al., 2018; Tenney et al., 2019a; Jawahar et al., 2019; Hewitt and Manning, 2019). Others investigate the effect of fine-tuning, finding that changes are more centered around upper layers (Mosbach et al., 2020; Zhou and Srikumar, 2022) and are task-dependent (Merchant et al., 2020). In general, works have found that many of these changes are less pronounced than during pre-training and vary from task to task.

Encoder-only models Encoder-only models (here, referred to as CLMs) have primarily been trained for MPP. To improve task performance, works have utilized different linear molecular representations (Krenn et al., 2020; Yüksel et al., 2023; Leon et al., 2024), domain-specific auxiliary training objectives (Fabian et al., 2020; Ahmad et al., 2022; Wu et al., 2022; Li and Jiang, 2021; Park et al., 2024b,a), different tokenization schemes (Chithrananda et al., 2020; Ahmad et al., 2022; Leon et al., 2024), positional encodings (Ross et al., 2022; Liu et al., 2023a) and attention mechanisms (Ross et al., 2022).

Probing for chemical knowledge Few works explore the encoding of molecular substructures in models. Prior works focus on graph-based models, finding that GTs generally encode ten molec3

2

Our experiments in Section J support these findings.

ular substructures better than message-passing GNNs; and that the molecular substructures are already encoded well by random initializations (Akhondzadeh et al., 2023). Others study differences between pre-training and fine-tuning of GNNs on MPP tasks and report a positive correlation between probing and MPP performance (Wang* et al., 2023). Only Payne et al. (2020) and Fender et al. (2025) investigate CLMs either for visualization or only for two molecular substructures. Finally, some works have shown that even image-based models capture chemical and biological knowledge in their representations (Alampara et al., 2025; Naghdloo et al., 2025). In summary, it is not well understood whether CLMs learn molecular substructures well and if this follows any chemical theory. With this work, we make a first attempt to address this gap by conducting systematic probing experiments with CLMs across 78 molecular substructures.

3

ing tasks for molecular substructures that appear fewer than 200 times in either the train or test set. Our final probing dataset comprises a diverse set of 78 molecular substructures—from functional groups such as amides or phenols to different types of ring structures (cf. Section C.2).

4

Experimental Setup

We investigate our research questions across eight pre-trained CLMs, which we first probe for the presence of molecular substructures, comparing them against their randomly initialized counterparts. We then study the effects of fine-tuning on two well-studied MPP tasks (lipophilicity and solubility prediction). This allows us to compare the changes of a model’s molecular substructure representation against existing chemical knowledge. Pre-trained CLMs We focus on models pretrained on small molecules, particularly encoderonly CLMs trained with the masked language modeling (MLM) objective. These models can consider both left and right context, in contrast to decoderonly models trained on causal language modeling. Chemberta Chithrananda et al. (2020) release multiple six-layer models based on RoBERTa (Liu et al., 2020). We use both publicly available models, chemberta-base and chemberta. Chemberta-2 In subsequent work, Ahmad et al. (2022) release models pre-trained on different numbers of molecules (chemberta-2-5M, chemberta-2-10M, and chemberta-2-77M). Chemberta-3 Most recently, Singh et al. (2026) released their training framework along with a 12-layer model trained on 100M molecules (chemberta3). Molformer Ross et al. (2022) train a model with 12 layers using linear attention and rotary positional embeddings. We use the publicly available model trained on 100M molecules (molformer). Roberta-zinc-480m Heyer (2023) release a 14layer RoBERTa-based model trained on 480M molecules (roberta-zinc-480m). Except for molformer (which also uses PubChem, Kim et al. 2018), all models are trained on molecules from the ZINC dataset (Irwin and Shoichet, 2005). A more detailed description of all models is provided in Section B.

Probing Dataset Creation

Our goal is to curate a probing dataset that captures a wide range of molecular substructures and at the same time allows us to conduct meaningful analysis. Each probing task is formulated as a binary classification problem predicting the presence or absence of a molecular substructure (e.g., functional group, ring, etc.) in a molecule. The probing dataset is derived from PCQM4Mv2, a publicly available dataset designed for predicting the HOMO-LUMO energy gap (Hu et al., 2021). Preprocessing We first discard all molecules with invalid SMILES strings or those that lead to processing errors (cf. Section C.1). The remaining molecules—represented as SMILES strings— are then canonicalized and annotated with binary labels. Preprocessing results in an initial set of 101 unique molecular substructures. All processing except for binarization is performed using RDKit (Landrum et al., 2024). Data sampling and cleaning The PCQM4Mv2 dataset comprises ≈ 3.7 million molecules—too many to conduct extensive probing experiments. Hence, we create three subsets by uniformly randomly sampling 100k and 20k instances from the preprocessed train and validation splits of PCQM4Mv2, respectively. Note that sampling from the original splits prevents any leakage between train and test sets. Finally, we discard prob-

Probing setup For each probing task, we train a linear classifier on the CLS token representation of each encoder layer (cf. Section D.1). We evaluate 3

probing performance using the macro-averaged F1 score to account for class imbalances in the test splits. We compare each pre-trained model (PT) against its randomly initialized (RI) counterpart (note, chemberta2 has one shared RI model) and a majority class prediction baseline (maj). We further downsample instances in the training set to account for class imbalances. Preliminary experiments show that this substantially improves probing performance. For the remainder of the paper, all reported results refer to the performance on the downsampled probing tasks. All dataset statistics are provided in Section C.3.

5

form exceptionally well at identifying ring structures compared to all other groups ( and x). Moreover, RI models encode ring structures well already at the first encoder layer, suggesting that these surface-level patterns are easy for the models to extract directly from the input, even without pre-training. We also find that the benefit of pretraining diminishes for ring structures compared to that of other molecular substructures (i.e., the gap between x and is much smaller compared to the gap between x and ). Different types of rings A closer analysis of different types of ring structures (i.e., aliphatic, aromatic, and saturated rings) reveals that particularly aromatic and aliphatic are already well encoded in random initializations, with pre-training slightly reducing probing performance in the upper layers (cf. Figure 5) for all models except for chemberta-3 (see Section I). The high performance on aromatic rings might stem from a distinct surface-level pattern. When SMILES strings are canonicalized, atoms in aromatic rings (see Figure 4 for an example) are represented with lowercase letters. This contrast to other molecular substructures (which consist of upper-cased atoms) results in a strong signal for the model.

The Effect of Pre-Training (RQ1)

We study the effect of pre-training (RQ1) using our probing dataset (§3) and analyzing the results with increasing levels of granularity. We first compare the probing performance of pre-trained (PT) and randomly initialized (RI) models (§5.1), then with respect to different ring structures (§5.2), and finally, for individual molecular substructures (§5.3). 5.1

Probing Results

Figure 1 shows the macro-averaged F1 score of all eight PT and six RI models averaged across all 78 probing tasks and three datasetes each. In addition, we show the average performance of the majority class prediction baseline (maj). On average, models learn to better encode molecular substructures during pre-training: most pre-trained models (x) exhibit substantial improvements in performance relative to their randomly initialized counterparts ( ) and the majority classifier (–). We also observe that for most PT models (excluding chemberta and chemberta-base), the probing performance is higher in the upper layers (0.5– 1.0). In particular, we see that chemberta-2-10M, molformer and chemberta-2-5M benefit the most from pre-training. In contrast, we observe negligible improvement for chemberta or even small drops for chemberta-base. Most notably, we find that in the most recent model (chemberta-3), probing performance deteriorates substantially in the lower layers. We further investigate this phenomenon in Section I by further pre-training the model on different datasets. 5.2

5.3

Individual Molecular Substructures

Finally, we investigate if there are molecular substructures that undergo changes consistently across all models. For visualization, we focus on the model with the most pronounced changes (molformer) and provide the rest in Section E. We observe three patterns shown in Figure 2. PT > RI First, we find that pre-training generally leads to a better encoding of most molecular substructures in upper layers, as reflected by the higher density of the red shade in the middle and upper layers. In particular, all models exhibit substantial improvement on carboxylic acids (COO, COO2, Al_COO, Ar_COO), aromatic hydroxy groups (Ar_OH), phenol groups (phenol, phenol_nonorthobound), amides, and ketones (ketone, ketone_Topliss). Furthermore, all models improve on carbonyls (C_O_noCOO), aldehyde and imide, although to a lesser degree.

Ring Structures

PT < RI Second, some molecular substructures consistently exhibit lower performance after pretraining. In particular, probing performance decreases for aromatic nitrogens (Ar_N) and certain

We further analyze the representations of rings and other molecular substructures, finding that the representations of both RI ( ) and PT (x) models per4

chemberta

1.0

chemberta-base chemberta-2-10M chemberta-2-5M chemberta-2-77M

chemberta-3

molformer

roberta-zinc-480m

0.9

Macro F1

0.8 0.7 0.6 0.5 0.4 PT (rings) RI (rings) 1.0 0.0 0.5

0.3 0.2

0.0

0.5

1.0 0.0

0.5

maj (rings) PT (other) 1.0 0.0 0.5

RI (other) maj (other) 1.0 0.0

PT (all) RI (all) 0.5 1.0 0.0

maj (all) 0.5

1.0 0.0

0.5

1.0 0.0

0.5

1.0

Relative Layer Depth

Molecular Substructure

Figure 1: Average probing performance (macro-averaged F1 score ↑) of pre-trained (PT), randomly initialized (RI), and the majority class prediction (maj) models on molecular substructures. We report average performance on 12 ring types (avg rings), all other 66 substructures (avg other) and all 78 substructures (avg all). NHOH NO AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings HAcceptors HDonors Heteroatoms RotatableBonds SaturatedCarbocycles SaturatedHeterocycles SaturatedRings Ring Al_COO Al_OH Al_OH_noTert ArN Ar_COO Ar_N Ar_NH Ar_OH COO COO2

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

-0.21 ±0.07 0.80 ±0.95 1.80 ±0.24 2.86 ±0.21 0.51 ±0.08 0.20 ±0.08 -2.23 ±0.19 -0.03 ±0.03 -0.82 ±0.70 -0.21 ±0.15 0.30 ±1.23 1.75 ±0.17 1.45 ±0.46 3.46 ±0.57 3.06 ±0.18 0.03 ±0.01 2.40 ±0.04 -0.10 ±0.35 0.29 ±0.26 2.89 ±0.32 2.55 ±0.40 -0.70 ±0.20 -9.65 ±0.67 2.64 ±0.64 2.33 ±0.36 2.12 ±0.22

1.07 ±0.26 -2.37 ±0.45 2.29 ±0.27 3.47 ±0.05 0.37 ±0.16 0.11 ±0.16 -0.35 ±0.07 -0.11 ±0.03 -3.67 ±0.13 1.31 ±0.09 -0.04 ±1.15 4.35 ±0.41 3.05 ±0.46 5.16 ±0.57 3.71 ±0.37 0.01 ±0.01 4.99 ±0.40 2.98 ±0.08 3.15 ±0.26 4.92 ±0.23 3.38 ±0.61 -0.46 ±0.03 -1.32 ±0.61 5.10 ±0.19 4.70 ±0.34 4.82 ±0.27

0.99 ±0.15 -3.86 ±0.65 2.39 ±0.46 3.66 ±0.16 0.37 ±0.14 0.05 ±0.09 -0.45 ±0.08 -0.13 ±0.05 -3.80 ±0.48 0.86 ±0.10 -1.48 ±0.89 5.91 ±0.37 3.65 ±0.44 5.32 ±0.46 4.40 ±0.17 0.03 ±0.03 8.23 ±0.12 3.65 ±0.22 3.75 ±0.05 6.96 ±0.88 5.43 ±0.19 -0.70 ±0.08 -5.40 ±0.73 6.20 ±0.89 7.95 ±0.26 7.64 ±0.49

1.73 ±0.40 -4.97 ±0.98 3.18 ±0.12 3.78 ±0.06 0.16 ±0.07 -0.12 ±0.04 -0.79 ±0.14 -0.15 ±0.04 -4.64 ±0.36 1.85 ±0.19 -1.20 ±1.12 7.20 ±0.25 4.87 ±0.38 5.47 ±0.57 4.28 ±0.37 -0.01 ±0.04 11.14 ±0.25 5.80 ±0.15 5.49 ±0.17 9.64 ±0.47 6.30 ±0.55 -0.95 ±0.10 -3.65 ±0.16 7.87 ±0.64 11.43 ±0.60 11.12 ±0.38

2.70 ±0.07 -6.71 ±1.17 3.29 ±0.15 4.22 ±0.21 0.08 ±0.09 -0.44 ±0.12 -1.10 ±0.08 -0.19 ±0.08 -7.25 ±1.30 2.70 ±0.32 -3.50 ±1.29 7.57 ±0.17 4.51 ±0.38 6.63 ±0.06 4.51 ±0.19 -0.00 ±0.01 14.10 ±0.48 7.28 ±0.33 6.88 ±0.53 15.50 ±0.26 7.85 ±0.70 -1.24 ±0.18 -4.41 ±0.47 11.54 ±0.41 14.70 ±0.44 14.34 ±0.41

2.70 ±0.22 -8.01 ±0.39 3.31 ±0.27 4.45 ±0.09 0.01 ±0.15 -0.46 ±0.18 -1.84 ±0.18 -0.25 ±0.03 -9.05 ±0.20 3.01 ±0.38 -4.16 ±0.79 8.11 ±0.67 4.96 ±0.65 6.62 ±0.29 4.68 ±0.41 -0.10 ±0.06 16.23 ±0.34 7.97 ±0.49 7.40 ±0.67 15.53 ±0.07 9.47 ±0.49 -2.01 ±0.06 -5.29 ±0.44 11.16 ±0.66 17.46 ±0.63 17.84 ±0.66

4.62 ±0.14 -9.77 ±1.18 4.28 ±0.44 4.44 ±0.12 -0.10 ±0.09 -0.34 ±0.16 -2.13 ±0.04 -0.44 ±0.09 -9.49 ±0.69 4.76 ±0.35 -5.24 ±0.29 8.41 ±0.39 6.11 ±0.44 7.31 ±0.07 4.87 ±0.45 -0.17 ±0.01 19.64 ±0.73 11.08 ±0.47 9.99 ±0.13 16.12 ±0.57 10.19 ±0.61 -2.36 ±0.12 -5.27 ±0.23 12.17 ±0.78 20.59 ±0.49 20.44 ±0.87

6.17 ±0.25 -9.66 ±1.19 5.05 ±0.22 4.62 ±0.43 -0.15 ±0.04 -0.14 ±0.24 -2.72 ±0.13 -0.65 ±0.04 -10.22 ±1.53 6.28 ±0.29 -6.47 ±0.76 9.48 ±0.88 6.90 ±0.31 7.57 ±0.62 4.90 ±0.28 -0.25 ±0.05 21.94 ±0.74 12.71 ±0.17 11.75 ±0.23 16.97 ±0.42 12.78 ±0.13 -3.13 ±0.05 -5.91 ±1.05 13.11 ±0.14 24.56 ±0.98 24.47 ±0.97

8.65 ±0.50 -10.73 ±0.95 5.37 ±0.24 4.57 ±0.37 -0.47 ±0.07 -0.34 ±0.29 -3.17 ±0.07 -0.90 ±0.08 -11.56 ±1.49 9.37 ±0.51 -6.46 ±1.15 9.97 ±1.09 7.86 ±0.24 7.54 ±0.47 4.44 ±0.26 -0.40 ±0.02 22.22 ±0.77 14.58 ±0.13 13.80 ±0.07 19.51 ±0.20 13.72 ±0.77 -3.69 ±0.14 -7.15 ±0.09 14.19 ±0.44 25.20 ±0.84 24.75 ±0.33

10.98 ±0.49 -10.37 ±1.21 5.97 ±0.38 4.11 ±0.34 -0.75 ±0.12 -0.54 ±0.15 -3.55 ±0.17 -1.02 ±0.10 -11.25 ±0.40 11.38 ±0.51 -6.44 ±0.49 10.26 ±0.31 8.24 ±0.21 6.85 ±0.22 3.68 ±0.27 -0.66 ±0.06 25.10 ±0.69 17.08 ±0.21 16.58 ±0.40 21.49 ±0.52 16.81 ±0.97 -4.34 ±0.38 -9.15 ±0.09 16.59 ±0.63 28.36 ±0.47 28.27 ±0.28

12.13 ±0.41 -11.52 ±0.75 6.08 ±0.29 4.49 ±0.30 -0.84 ±0.09 -1.02 ±0.33 -3.97 ±0.12 -1.36 ±0.09 -12.20 ±1.33 12.67 ±0.59 -6.55 ±1.00 10.27 ±0.50 8.61 ±0.35 7.26 ±0.68 3.83 ±0.25 -1.11 ±0.16 28.08 ±0.37 19.21 ±0.28 18.73 ±0.20 25.67 ±0.95 18.40 ±1.90 -4.84 ±0.14 -13.08 ±0.91 20.34 ±0.19 31.15 ±0.24 31.08 ±0.43

13.09 ±0.61 -9.88 ±0.83 6.94 ±0.41 5.13 ±0.15 -1.51 ±0.02 -0.29 ±0.17 -3.70 ±0.17 -1.67 ±0.05 -10.30 ±0.73 13.82 ±0.54 -5.03 ±0.50 10.73 ±0.87 9.55 ±0.17 8.60 ±0.56 3.64 ±0.32 -1.70 ±0.27 29.00 ±0.41 20.54 ±0.21 19.45 ±0.39 26.64 ±0.73 20.40 ±1.97 -4.58 ±0.18 -12.03 ±0.82 21.96 ±1.15 32.19 ±0.31 31.49 ±0.30

0

1

2

3

4

5

6

7

8

9

10

11

12

C_O C_O_noCOO C_S Imine NH0 NH1 NH2 N_O Ndealkylation1 Ndealkylation2 Nhpyrrole SH aldehyde alkyl_halide allylic_oxid amide amidine aniline aryl_methyl benzene bicyclic ester ether furan guanido halogen

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

0.91 ±0.06 1.54 ±0.16 -4.52 ±0.96 0.85 ±0.40 -0.75 ±0.14 -0.79 ±0.21 0.73 ±0.34 -2.60 ±0.86 1.91 ±0.66 3.44 ±0.86 -9.65 ±0.67 -2.41 ±0.14 0.62 ±0.40 -4.92 ±0.77 -0.79 ±0.07 3.27 ±0.66 0.05 ±0.48 3.68 ±0.53 4.33 ±1.17 0.23 ±0.09 5.04 ±0.24 1.91 ±0.44 1.38 ±0.38 -25.87 ±0.45 -1.46 ±0.73 -6.55 ±0.23

2.75 ±0.16 2.72 ±0.11 -1.01 ±0.60 4.36 ±0.26 0.79 ±0.16 -0.14 ±0.33 1.30 ±0.65 0.57 ±0.58 2.80 ±0.09 5.42 ±0.69 -1.32 ±0.61 -1.57 ±0.59 1.80 ±0.12 -1.13 ±0.34 0.86 ±0.43 6.09 ±0.50 0.44 ±0.59 2.69 ±0.32 5.32 ±0.37 0.15 ±0.16 8.34 ±0.18 3.64 ±0.57 2.56 ±0.46 -4.58 ±0.90 -0.99 ±0.86 -0.65 ±0.12

3.10 ±0.09 3.14 ±0.19 -1.96 ±0.28 4.92 ±0.31 0.63 ±0.19 -0.45 ±0.25 2.34 ±0.38 1.28 ±0.08 3.96 ±0.36 4.83 ±0.37 -5.40 ±0.73 -2.65 ±0.32 2.67 ±0.28 0.23 ±0.06 1.71 ±0.38 6.95 ±0.24 0.77 ±0.47 3.82 ±0.36 8.07 ±0.40 0.04 ±0.09 9.67 ±0.34 4.96 ±0.36 4.24 ±0.31 -5.46 ±0.73 -0.79 ±0.72 -1.19 ±0.18

0

1

2

3

4.38 ±0.11 4.80 ±0.32 -1.04 ±0.25 4.44 ±0.15 1.05 ±0.08 0.25 ±0.43 5.04 ±0.48 1.23 ±0.48 7.36 ±0.46 6.06 ±0.14 -3.65 ±0.16 -0.76 ±0.94 4.77 ±0.51 2.77 ±0.11 3.00 ±0.15 9.10 ±0.33 2.07 ±0.60 4.35 ±0.47 10.72 ±0.24 -0.18 ±0.05 9.45 ±0.19 7.43 ±0.60 6.78 ±0.39 -6.31 ±0.34 -0.22 ±0.96 -1.09 ±0.25

6.19 ±0.21 7.79 ±0.43 0.19 ±0.25 5.43 ±0.34 1.28 ±0.30 0.60 ±0.11 8.34 ±0.10 2.04 ±0.36 9.67 ±0.24 7.94 ±0.51 -4.41 ±0.47 0.62 ±0.47 6.29 ±0.28 3.01 ±0.28 3.41 ±0.39 10.05 ±0.23 2.20 ±0.46 5.48 ±0.47 12.21 ±0.28 -0.46 ±0.07 9.43 ±0.31 7.29 ±0.19 7.90 ±0.49 -8.12 ±0.62 -0.42 ±0.82 -2.40 ±0.18

7.69 ±0.20 9.82 ±0.31 -0.59 ±0.22 5.72 ±0.03 0.81 ±0.17 1.21 ±0.25 9.72 ±0.17 3.00 ±0.42 10.82 ±0.36 8.37 ±0.73 -5.29 ±0.44 0.11 ±0.74 7.24 ±0.36 2.59 ±0.39 4.27 ±0.30 11.62 ±0.42 1.81 ±0.52 5.71 ±0.73 13.04 ±0.19 -0.51 ±0.17 9.31 ±0.24 9.40 ±0.12 8.58 ±0.46 -8.55 ±0.64 -0.47 ±0.63 -3.66 ±0.16

8.70 ±0.05 10.99 ±0.34 -1.09 ±0.41 7.58 ±0.19 1.07 ±0.17 2.74 ±0.32 11.74 ±0.41 4.22 ±0.11 12.46 ±0.70 9.07 ±0.75 -5.27 ±0.23 -0.43 ±0.53 9.89 ±0.60 3.56 ±0.36 4.98 ±0.82 12.26 ±0.05 3.15 ±0.73 5.70 ±0.33 13.07 ±0.42 -0.37 ±0.17 9.34 ±0.18 11.52 ±0.73 9.58 ±0.22 -8.77 ±0.81 -0.69 ±0.31 -3.86 ±0.15

9.21 ±0.09 11.79 ±0.36 -0.75 ±0.86 8.98 ±0.33 0.92 ±0.06 3.26 ±0.46 14.05 ±0.23 4.41 ±0.30 14.36 ±0.97 9.57 ±0.92 -5.91 ±1.05 -0.12 ±0.53 13.71 ±0.68 4.78 ±0.23 5.88 ±0.34 13.10 ±0.43 3.40 ±0.75 5.21 ±0.17 13.28 ±0.34 -0.09 ±0.24 9.30 ±0.22 12.34 ±0.36 11.27 ±0.48 -7.41 ±0.83 -1.00 ±0.24 -4.65 ±0.13

9.28 ±0.13 12.57 ±0.10 0.66 ±1.06 9.61 ±0.49 1.59 ±0.24 3.63 ±0.46 16.42 ±0.12 4.85 ±0.74 15.23 ±0.66 9.15 ±0.54 -7.15 ±0.09 0.37 ±1.78 16.43 ±0.71 3.02 ±0.18 6.49 ±0.39 11.87 ±0.34 3.65 ±0.85 4.94 ±0.44 12.79 ±0.49 -0.36 ±0.26 9.02 ±0.32 13.51 ±0.68 11.81 ±0.74 -9.89 ±1.55 -1.58 ±0.67 -5.73 ±0.09

9.45 ±0.33 13.24 ±0.16 2.97 ±0.98 11.53 ±0.16 3.00 ±0.17 6.32 ±0.54 20.23 ±0.25 4.90 ±0.77 13.89 ±0.58 9.32 ±0.41 -9.15 ±0.09 3.92 ±0.75 19.22 ±0.42 2.45 ±0.53 6.46 ±0.38 13.02 ±0.20 3.86 ±0.67 6.14 ±0.12 12.57 ±0.10 -0.49 ±0.11 8.60 ±0.25 15.37 ±0.37 12.23 ±0.60 -10.29 ±1.02 -1.31 ±0.89 -7.09 ±0.27

10.24 ±0.26 14.66 ±0.24 6.54 ±1.41 11.52 ±0.38 2.96 ±0.11 6.56 ±0.73 22.50 ±0.30 4.49 ±0.44 15.65 ±0.44 9.59 ±0.96 -13.08 ±0.91 6.44 ±1.52 21.86 ±0.56 3.03 ±0.23 6.68 ±0.30 14.29 ±0.54 4.32 ±1.25 5.41 ±0.06 13.78 ±0.17 -0.94 ±0.31 8.52 ±0.28 18.07 ±0.47 12.19 ±0.69 -11.48 ±0.34 -0.58 ±1.16 -5.17 ±0.23

11.00 ±0.46 15.31 ±0.19 7.61 ±1.03 15.81 ±0.25 3.88 ±0.11 6.88 ±0.55 22.95 ±0.26 5.23 ±0.34 15.84 ±0.12 10.76 ±0.57 -12.03 ±0.82 6.17 ±1.08 25.17 ±1.39 3.11 ±0.21 7.99 ±0.32 15.42 ±0.40 7.82 ±0.47 9.46 ±0.26 13.85 ±0.63 -0.25 ±0.08 8.29 ±0.16 19.38 ±0.44 12.76 ±0.33 -9.65 ±0.34 0.05 ±1.07 -6.70 ±0.26

4

5

6

7

8

9

10

11

12

Layer

hdrzine hdrzone imidazole imide ketone ketone_Topliss lactone methoxy morpholine nitrile nitro nitro_arom nitro_arom_nonortho oxime para_hydroxylation phenol phenol_noOrthoHbond piperdine piperzine priamide pyridine sulfide thiazole thiophene unbrch_alkane urea

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

-1.09 ±0.29 -1.85 ±0.08 -1.11 ±0.42 -0.44 ±0.44 -1.21 ±0.32 -0.60 ±0.28 2.46 ±0.38 4.90 ±0.37 2.69 ±0.52 -11.23 ±1.02 -24.41 ±1.25 -25.39 ±1.08 -21.27 ±0.60 -3.59 ±0.29 1.76 ±0.22 1.65 ±0.19 1.16 ±0.09 2.25 ±0.25 1.05 ±0.18 0.03 ±0.57 1.18 ±0.53 -1.34 ±0.53 -18.88 ±0.43 -28.10 ±0.72 3.58 ±0.46 -1.40 ±0.77

-0.95 ±0.19 -0.26 ±0.39 -0.00 ±0.19 3.54 ±0.09 0.46 ±0.32 0.90 ±0.17 3.35 ±0.50 6.59 ±0.40 4.12 ±0.52 -0.12 ±0.21 -2.46 ±0.44 -3.01 ±1.99 -4.61 ±1.61 0.66 ±0.59 4.49 ±0.27 3.55 ±0.14 3.54 ±0.26 3.50 ±0.59 2.48 ±0.65 2.59 ±0.86 1.80 ±0.51 -0.88 ±0.20 -2.72 ±0.39 -3.79 ±0.76 3.57 ±0.27 -0.14 ±0.97

-1.34 ±0.32 -0.65 ±1.25 -0.73 ±0.35 3.70 ±1.30 0.96 ±0.49 1.41 ±0.38 3.87 ±0.47 9.53 ±0.67 3.89 ±0.93 -0.79 ±0.31 -2.30 ±0.53 -2.57 ±0.32 -2.85 ±0.46 1.00 ±0.39 6.37 ±0.21 4.68 ±0.14 4.42 ±0.09 4.03 ±0.23 2.12 ±0.70 2.96 ±0.20 3.59 ±0.14 -1.40 ±0.44 -3.80 ±0.19 -5.01 ±0.92 4.63 ±0.40 -0.19 ±0.41

-1.00 ±0.63 0.56 ±0.74 -0.65 ±0.29 4.25 ±0.74 3.04 ±0.35 2.82 ±0.35 3.99 ±0.39 17.33 ±0.11 3.43 ±0.96 -0.10 ±0.44 -5.26 ±0.44 -4.37 ±0.33 -4.94 ±2.32 0.95 ±0.25 6.15 ±0.31 5.93 ±0.19 5.51 ±0.54 4.56 ±0.16 3.23 ±0.67 4.90 ±0.18 5.81 ±0.60 -0.93 ±0.28 -4.21 ±0.51 -5.46 ±0.37 6.94 ±0.88 0.34 ±0.40

0.30 ±0.58 0.90 ±0.56 -0.87 ±0.91 4.63 ±0.79 5.27 ±0.24 4.90 ±0.30 3.74 ±0.17 19.13 ±0.12 3.57 ±1.40 -0.65 ±0.35 -5.79 ±1.16 -6.31 ±1.66 -5.17 ±1.74 1.87 ±0.58 7.32 ±0.68 8.91 ±0.22 8.66 ±0.24 4.93 ±0.59 4.09 ±0.72 6.56 ±0.63 5.84 ±0.64 -0.82 ±0.12 -4.61 ±0.73 -5.75 ±0.81 9.87 ±0.36 0.58 ±0.58

2.20 ±0.66 2.06 ±0.98 -1.07 ±0.55 5.16 ±0.39 8.05 ±0.40 7.86 ±0.22 3.72 ±0.21 22.73 ±0.22 4.15 ±0.37 -0.27 ±0.33 -6.26 ±1.68 -5.65 ±0.69 -4.24 ±1.12 2.27 ±0.33 8.61 ±0.78 8.31 ±0.67 8.15 ±0.86 5.17 ±0.36 5.76 ±1.04 7.83 ±0.40 5.40 ±0.42 -0.91 ±0.19 -6.63 ±0.38 -9.44 ±0.96 11.04 ±0.65 0.46 ±0.30

2.79 ±0.82 2.96 ±0.86 0.39 ±0.62 5.90 ±0.31 10.50 ±0.46 10.12 ±0.04 3.84 ±0.29 23.70 ±0.12 4.69 ±1.36 0.81 ±0.13 -5.16 ±1.51 -4.74 ±0.37 -4.07 ±0.95 3.70 ±0.55 9.01 ±0.31 9.02 ±1.17 8.85 ±0.56 5.44 ±0.27 6.23 ±1.60 11.74 ±1.17 7.14 ±0.74 -0.84 ±0.45 -6.04 ±1.38 -10.52 ±1.04 11.58 ±0.69 1.95 ±0.78

2.43 ±0.63 3.52 ±0.90 1.53 ±0.40 8.18 ±0.38 12.19 ±0.41 11.31 ±0.24 4.34 ±0.10 28.52 ±0.35 4.69 ±0.63 2.21 ±0.08 -5.34 ±1.34 -4.18 ±0.23 -3.97 ±1.78 5.37 ±0.61 9.80 ±0.46 10.30 ±0.95 10.04 ±0.93 5.41 ±0.67 5.29 ±1.69 12.79 ±0.83 9.60 ±0.45 -0.81 ±0.38 -6.72 ±0.51 -11.66 ±0.37 13.09 ±0.57 1.43 ±0.59

2.56 ±0.33 3.26 ±0.46 1.92 ±1.06 6.38 ±0.48 14.14 ±0.90 13.53 ±0.50 3.92 ±0.24 28.82 ±0.30 5.06 ±1.27 0.57 ±0.53 -5.91 ±0.29 -6.07 ±0.25 -5.29 ±1.03 6.92 ±1.49 9.89 ±0.27 11.68 ±0.87 12.09 ±0.92 5.33 ±0.54 5.06 ±1.62 15.29 ±0.45 9.47 ±0.51 -1.36 ±0.45 -7.16 ±0.60 -12.64 ±0.90 14.01 ±0.58 0.52 ±1.17

3.19 ±0.42 5.02 ±0.81 2.19 ±0.33 7.60 ±1.15 17.09 ±0.81 16.62 ±0.35 3.76 ±0.61 29.10 ±0.30 6.08 ±0.52 1.26 ±0.43 -5.79 ±0.19 -5.97 ±0.50 -5.70 ±0.98 8.83 ±1.01 9.99 ±0.21 15.05 ±1.02 15.49 ±0.23 5.41 ±0.74 5.75 ±1.41 17.51 ±2.77 10.11 ±0.56 0.47 ±0.05 -6.62 ±0.71 -11.85 ±0.49 13.84 ±1.29 0.94 ±1.12

3.24 ±0.42 5.07 ±0.69 2.73 ±0.30 6.37 ±0.38 19.95 ±0.96 19.03 ±0.34 4.84 ±0.59 30.64 ±0.44 6.44 ±0.15 3.00 ±0.20 -7.44 ±0.93 -6.81 ±0.65 -7.91 ±1.79 9.06 ±1.04 9.20 ±0.49 20.12 ±0.62 20.18 ±0.07 5.41 ±0.61 6.68 ±2.17 19.01 ±0.71 11.80 ±0.49 -0.48 ±0.30 -8.49 ±1.14 -12.23 ±0.61 13.52 ±0.99 0.63 ±0.66

3.38 ±0.95 6.83 ±1.77 3.25 ±0.54 8.98 ±0.53 25.11 ±0.99 23.32 ±1.02 6.32 ±0.54 30.24 ±0.26 4.75 ±0.29 2.16 ±0.51 -5.14 ±1.15 -4.77 ±0.46 -5.11 ±2.26 12.14 ±0.58 10.80 ±0.39 20.96 ±1.07 20.87 ±1.11 5.99 ±0.74 6.21 ±1.86 17.55 ±1.43 15.43 ±0.57 -1.62 ±0.30 -8.91 ±1.37 -13.27 ±1.21 13.65 ±0.97 2.66 ±0.31

0

1

2

3

4

5

6

7

8

9

10 11 12

30

20

10

0

10

20

Figure 2: Relative difference in probing performance (% macro-averaged F1) between RI and PT molformer on 78 molecular substructures (cf. Table 3). Each cell denotes the relative difference for a specific probing task (y-axis) and layer (x-axis). Red indicates an increase in performance after pretraining, while blue denotes a decrease.

heterocycles such as thiazole, thiophene and furan across all models. Likewise, some ring substructures such as AromaticHeterocycles, AromaticCarbocycles and benzene show slight performance degradation in upper layers. The decrease in performance on AromaticHeterocycles may potentially reflect the substantial drops in performance of thiazole, thiophene and furan— all aromatic heterocycles. Finally, most models (except for chemberta-2-5M/10M) show evidence of unlearning halogens in upper layers while preserving a better encoding in the lower layers.

CLMs, considerably changing the encoding of many molecular substructures. We further observe that RI models already encode aromatic and aliphatic rings very well performing on-par with or better than PT models. Interestingly, unlearning of molecular substructures largely varies between models, however, a few molecular substructures are consistently unlearned during pre-training. We conjecture this might stem from a disparity in the pre-training data and conduct further pretraining experiments for five molecular substructures and across different models (§7). Finally, we find that all molecular substructures exhibit a change larger than ±1% in at least one model.

PT ≈ RI Third, we observe that only a handful molecular substructures undergo small amounts of change (±1%), resulting in an almost uniform distribution of information across layers. However, this behavior is not consistent across models. 5.4

6

The Effect of Fine-Tuning (RQ2)

Our probing experiments have shown how pretraining reconfigures the information encoded in representations and that RI models already encode ring structures well. Next, we investigate changes of RI and PT models during fine-tuning on two

Discussion

Overall, our results suggest that pre-training improves the molecular structure awareness of 5

6.2

well-studied tasks in chemistry, which allows us to contextualize our findings within chemical theory. 6.1

Experimental Setup

For fine-tuning, we replace the classification head of the CLM with either a linear regression layer or a two-layer MLP and minimize the mean squared error loss. Following Wu et al. (2018), we use the root mean squared error (RMSE) as our evaluation metric for both tasks. Since all models were pre-trained on canonicalized SMILES strings, we canonicalize the input SMILES accordingly.

Chemistry Background

We focus on predicting the lipophilicity and aqueous solubility (in short, solubility) of molecules. Here, we provide brief task descriptions and introduce important molecular substructures that affect the lipophilicity and solubility of a molecule; and refer to Section A for more details.

Dataset Both datasets are sampled from the MoleculeNet benchmark (Wu et al., 2018) and consist of 4,200 (lipophilicity) and 1,127 (solubility) molecules. We use the train–validation–test splits (80/10/10) provided by Ross et al. (2022).Detailed dataset statistics and analysis for both tasks are provided in Section F.

Lipophilicity Lipophilicity refers to the ability of a chemical compound to dissolve in fatlike solvents (lipids, fats, oils; Morak-Młodawska et al. 2023). It is an important physicochemical property of molecules which correlates with the (oral) absorption, (tissue) distribution, metabolism, excretion, and toxcicity (ADMET) properties of drugs (Mannhold et al., 2009), essential in determining how a candidate drug will interact with the human body (Waring, 2009). The goal of lipophilicity prediction is to estimate the octanol/water distribution coefficient (logD) of a specific molecule.

Hyperparameters We perform hyperparameter tuning separately for both tasks, considering different batch sizes and learning rates. All pre-trained and randomly initialized models are trained for up to 10–20 epochs. We deploy early stopping with a patience of 2 and use AdamW as our optimizer. We report all hyperparameters in Section F.

Aqueous solubility Aqueous solubility refers to the ability of a molecule to dissolve in water. For drug development, predicting the solubility of a molecule is equally important as predicting the lipophilicity as it also affects their biovailability and ADMET profiles (Llompart et al., 2024; Klopman et al., 1992). The goal of solubility prediction is to estimate the log solubility (logS) of a specific molecule in water. While solubility is closely related to lipophilicity, it is also dependent on other factors such as the melting point of a molecule (Hill and Young, 2010).

Baselines As baselines, we evaluate multiple traditionally used models, namely, linear regression models (LR), support vector machines (SVM), and gradient boosted trees (XGB). For each model, we evaluate four algorithms to extract molecule representation vectors, also known as fingerprints, provided by RDKit (Landrum et al., 2024). Finally, we evaluate two large language models (LLMs): Llama-3.2-3B-Instruct (Grattafiori et al., 2024) and gpt-oss-20B (OpenAI, 2025) with additional chemical knowledge that is important for the respective downstream task. We provide detailed hyperparameters and experimental results for all baselines in appendices J.1 and J.2.

Important molecular substructures Chemical literature distinguishes between two groups of molecular substructures that are known to affect lipophilicity and solubility (Harrold et al., 2023). First, hydrophilic substructures such as carboxylic acids substantially decrease a molecule’s lipophilicity while increasing its solubility. Second, lipophilic substructures such as aromatic rings increase a molecule’s lipophilicity while decreasing its solubility. We follow this classification of molecular substructures in our analysis and put all other molecular substructures that do not substantially affect lipophilicity into a third group (other). We provide a list of all molecular substructures along with their group in Section F.

Probing dataset adjustment In order to conduct meaningful analyses, we accommodate changes to the probing dataset introduced in §3 that consider dataset-specific properties of the respective downstream task. More specifically, we discard any molecular substructure which appears fewer than ten times in either the training or test split of the task-specific dataset; effectively removing outliers from our analysis. This results in probing 60 and 39 molecular substructures for lipophilicity and solubility prediction, respectively. 6

6.3

Downstream Task Results

a model before and after fine-tuning. This is done for both RI and PT models to understand potential differences in their behavior during fine-tuning. In our analysis, we first inspect lipophilic and hydrophilic molecular substructures and then inspect individual molecular substructures. Due to a lack of space, we focus our analysis in the main paper on lipophilicity prediction and provide the results and analysis for solubility prediction in Section G.

Table 1 shows the results of all randomly initialized (RI) and pre-trained models (PT) as well as the best performing model using fingerprints (SVM) and LLM (gpt-oss-20b) for both downstream tasks. We further include the results of the graph-based models (mol ) that were reported by Ross et al. (2022) who use the same data splits. Overall, we observe that PT models consistently outperform RI ones on both tasks with molformer consistently performing best, highlighting the benefit of pre-training CLMs. We further find that the chemberta-2 models, differing only in the pre-training datasets, exhibit differences of 0.073 RMSE on lipophilicity (0.046 on solubility), suggesting that pre-training data plays a major role for downstream task performance. Moreover, the chemberta-2-77M model is often outperformed by its smaller counterparts. This indicates that data quality may play a more important role than data quantity. Finally, consistent with prior findings, we observe that fingerprint-based models perform rather well (Dias et al., 2023; Xia et al., 2023); and that LLMs perform even worse than the mean predictor (Zhao et al., 2023). Model RI

Lipo PT

ESOL RI PT

GCmol A-FPmol MPNNmol

-

0.655 0.578 0.719

-

0.970 0.503 0.580

mean predictor SVM(c=16) + ATFP gpt-oss-20b

-

1.013 0.640 2.532

-

2.057 0.830 8.964

molformer roberta-zinc-480m chemberta-base chemberta chemberta-2-5M chemberta-2-10M chemberta-2-77M chemberta-3

0.832 0.788 0.785 0.779 0.850 " " 1.026

0.565 0.580 0.663 0.675 0.664 0.591 0.632 0.637

0.808 0.878 0.832 0.822 0.872 " " 0.960

0.587 0.746 0.739 0.693 0.682 0.724 0.728 0.757

Group analysis Figure 3 shows heatmaps for eight pre-trained (left) and six randomly initialized (right) models split into hydrophilic (top), lipophilic (middle), and other (bottom) groups. Similar to §5, red indicates an increase in probing performance while blue indicates a decrease. We observe that the groups that are important for lipophilicity prediction (hydrophilic and lipophilic) undergo larger changes than the other group. This indicates that molecular substructure learning follows chemical theory. Interestingly, we find that fine-tuning has a noticeably smaller effect than pre-training on the molecular substructure representations; and that the changes are mostly concentrated in upper layers which corroborates prior observations in the NLP literature (Mosbach et al., 2020; Merchant et al., 2020; Fayyaz et al., 2021). In contrast, RI models behave differently, as fine-tuning appears to mostly negatively affect the encoding of substructures in upper layers. Individual analysis A detailed analysis of individual molecular substructures reveals that the effect of fine-tuning varies across models. Furthermore, even among molecular substructures of the same group (lipophilic, hydrophilic, other), the magnitude may vary (we provide detailed heatmaps in Section G). Nevertheless, there are multiple substructures for which probing performance increases after fine-tuning on lipophilicity prediction. These are primarily carboxylic acids (COO, COO2, Al_COO, and Ar_COO4 ). This is consistent with observations made by chemists suggesting that carboxylic acids contribute most negatively to the logD value and are therefore highly indicative (Landry and Crawford, 2020). Interestingly, chemberta-2 models consistently improve upon halogens after fine-tuning (we study this closer in §7). Again, we do not observe any consistent trends for RI models, except for a degradation of AromaticHeterocycles and Ar_NH.

Table 1: Test performance (RMSE, ↓) for lipophilicity (Lipo) and solubility (ESOL) prediction (Wu et al., 2018). Besides all CLMs, we also include results of the best performing fingerprint-based model (SVM) and LLM (gpt-oss-20b). mol denotes results of graphbased models reported by Ross et al. (2022).

6.4

Probing Results

We conduct probing experiments similar to §5 but with the difference that we now compare the layerwise representations of molecular substructures in

4

7

With the exception of chemberta2-77M.

chemberta-base

0.00 ±0.00

0.55 ±0.17

1.03 ±0.34

0.43 ±0.29

0.25 ±0.31

0.94 ±0.28

1.56 ±0.43

lipophilic (14)

0.00 ±0.00

-0.01 ±0.14

0.40 ±0.09

0.01 ±0.14

-0.11 ±0.20

0.60 ±0.31

0.27 ±0.31

other (20)

0.00 ±0.00

0.48 ±0.11

0.97 ±0.11

0.21 ±0.25

0.09 ±0.35

-0.04 ±0.34

0.01 ±0.42

0

1

2

3

4

5

6

chemberta-2-10M

hydrophilic (26)

0.00 ±0.00

-2.76 ±0.27

-4.37 ±0.47

-1.01 ±0.85

lipophilic (14)

0.00 ±0.00

-2.24 ±0.38

-1.57 ±0.32

-0.17 ±0.29

other (20)

0.00 ±0.00

-2.51 ±0.11

-3.08 ±0.19

-2.54 ±0.22

0

1

2

3

hydrophilic (26)

0.00 ±0.00

0.58 ±0.13

0.99 ±0.30

0.34 ±0.36

lipophilic (14)

0.00 ±0.00

-0.60 ±0.11

1.03 ±0.23

3.05 ±0.35

other (20)

0.00 ±0.00

-0.36 ±0.15

-0.29 ±0.21

-0.17 ±0.22

0

1

2

3

chemberta-2-77M

molformer

0.00 0.02 0.06 -0.08 -0.47 -0.30 -0.51 -0.70 -0.65 -0.15 0.54 -0.16 -0.49 hydrophilic (26) ±0.00 ±0.33 ±0.26 ±0.19 ±0.20 ±0.41 ±0.34 ±0.34 ±0.45 ±0.26 ±0.67 ±0.32 ±0.58 0.00 0.01 0.05 -0.28 -0.60 -0.56 -0.71 -0.72 -0.29 -0.14 0.30 0.29 -0.11 lipophilic (14) ±0.00 ±0.27 ±0.39 ±0.26 ±0.43 ±0.24 ±0.18 ±0.42 ±0.33 ±0.25 ±0.42 ±0.55 ±0.29 0.00 -0.23 -0.10 -0.02 -0.47 -0.48 -0.64 -0.78 -0.74 -0.74 -0.26 -0.64 -1.51 other (20) ±0.00 ±0.24 ±0.23 ±0.22 ±0.28 ±0.25 ±0.28 ±0.26 ±0.40 ±0.30 ±0.22 ±0.30 ±0.27

0

1

2

3

4

5

6

7

8

9

10 11 12

1 0

0.00 ±0.00

-1.31 ±0.25

-1.38 ±0.27

-3.64 ±0.33

-1.99 ±0.33

-0.61 ±0.42

-0.01 ±0.26

0.00 ±0.00

-1.99 ±0.18

-4.26 ±0.21

-4.47 ±0.18

-3.10 ±0.29

-3.29 ±0.32

-2.73 ±0.31

0.00 ±0.00

-1.18 ±0.26

-1.99 ±0.38

-3.74 ±0.31

-3.07 ±0.37

-2.84 ±0.31

-2.32 ±0.26

0

1

2

3

4

5

6

0 2 4

chemberta-2-5M

0.03 ±0.27

-0.57 ±0.22

-1.10 ±0.29

-1.50 ±0.20

-1.95 ±0.29

-2.59 ±0.44

-2.94 ±0.41

4

other (20)

0.00 ±0.00

-0.18 ±0.21

0.15 ±0.12

-0.24 ±0.22

-0.54 ±0.22

-1.01 ±0.23

-1.51 ±0.23

0

1

2

3

4

5

6

chemberta-2-10M

6.10 ±0.80

5.0

hydrophilic (26)

0.00 ±0.00

-2.35 ±0.25

-0.78 ±0.31

-1.30 ±0.29

2.48 ±0.20

3.14 ±0.31

2.5

lipophilic (14)

0.00 ±0.00

-4.25 ±0.31

-4.86 ±0.24

-5.06 ±0.37

0.00 ±0.00

-1.16 ±0.13

0.45 ±0.27

0.68 ±0.16

0.0

other (20)

0.00 ±0.00

-5.28 ±0.17

-4.90 ±0.18

-5.03 ±0.20

0

1

2

3

0

1

2

3

hydrophilic (26)

0.00 ±0.00

-2.35 ±0.25

-0.78 ±0.31

-1.30 ±0.29

lipophilic (14)

0.00 ±0.00

-4.25 ±0.31

-4.86 ±0.24

-5.06 ±0.37

other (20)

0.00 ±0.00

-5.28 ±0.17

-4.90 ±0.18

-5.03 ±0.20

0

1

2

3

chemberta-3

1

2

3

4

5

6

7

8

roberta-zinc-480m

9

1

2

Relative Layer Depth

3

4

5

6

7

8

2 0

10 11 12

0.00 -0.14 -0.03 -0.14 0.08 -0.16 -0.29 -0.32 0.02 0.90 0.91 0.99 0.51 0.13 -0.08 ±0.00 ±0.05 ±0.09 ±0.10 ±0.11 ±0.10 ±0.12 ±0.11 ±0.21 ±0.23 ±0.25 ±0.35 ±0.23 ±0.30 ±0.36

0

0.41 ±0.19

-1.23 ±0.31

2.24 ±0.40

0.00 -0.02 -0.01 -0.14 -0.12 -0.23 -0.28 0.01 0.52 1.08 1.02 1.19 0.97 0.54 0.74 ±0.00 ±0.07 ±0.07 ±0.10 ±0.10 ±0.13 ±0.14 ±0.23 ±0.22 ±0.26 ±0.17 ±0.30 ±0.31 ±0.21 ±0.15

1

0.95 ±0.19

-1.10 ±0.20

-0.10 ±0.16

0.00 -0.17 -0.06 -0.08 0.29 0.00 0.26 0.47 0.90 2.42 3.02 2.93 2.66 1.61 1.66 ±0.00 ±0.05 ±0.08 ±0.09 ±0.11 ±0.10 ±0.11 ±0.27 ±0.21 ±0.35 ±0.29 ±0.28 ±0.41 ±0.31 ±0.39

0

0.77 ±0.20

0.00 ±0.00

-0.73 ±0.18

0.00 -0.05 0.04 0.60 -0.92 -0.75 -1.52 -1.08 -1.31 -1.17 -1.08 -1.04 0.61 ±0.00 ±0.19 ±0.15 ±0.23 ±0.19 ±0.17 ±0.28 ±0.30 ±0.53 ±0.35 ±0.40 ±0.40 ±0.40

0

0.00 ±0.00

lipophilic (14)

0.00 ±0.00

0.00 0.04 0.09 0.07 -0.36 -0.53 0.62 0.49 -0.44 -0.66 -0.78 -1.44 -0.18 ±0.00 ±0.11 ±0.13 ±0.14 ±0.14 ±0.14 ±0.56 ±0.57 ±0.22 ±0.35 ±0.58 ±0.31 ±0.32

0

chemberta-base

hydrophilic (26) 2

0.00 ±0.00

0.00 -0.11 0.27 0.55 -0.49 -0.52 -1.15 -1.00 -0.76 -1.12 -0.34 0.73 2.94 ±0.00 ±0.14 ±0.14 ±0.14 ±0.20 ±0.18 ±0.23 ±0.37 ±0.50 ±0.39 ±0.48 ±0.30 ±0.57

2

chemberta

0

Category

Category

chemberta hydrophilic (26)

9 10 11 12 13 14

chemberta-2-77M

molformer

0.00 -0.16 0.33 0.32 0.10 -0.24 -0.54 -0.85 -1.35 -1.88 -2.60 -3.39 -4.41 hydrophilic (26) ±0.00 ±0.16 ±0.20 ±0.21 ±0.22 ±0.22 ±0.23 ±0.20 ±0.19 ±0.32 ±0.29 ±0.30 ±0.26

2

0.00 -0.09 -0.09 -0.07 -0.14 -0.18 -0.42 -0.62 -0.96 -1.51 -2.40 -3.17 -4.47 lipophilic (14) ±0.00 ±0.13 ±0.18 ±0.11 ±0.21 ±0.23 ±0.29 ±0.34 ±0.39 ±0.36 ±0.39 ±0.21 ±0.20 0.00 -0.37 0.20 0.25 0.16 0.20 0.10 0.01 -0.26 -0.66 -1.28 -1.92 -2.89 other (20) ±0.00 ±0.26 ±0.16 ±0.16 ±0.19 ±0.20 ±0.18 ±0.20 ±0.21 ±0.20 ±0.22 ±0.29 ±0.23

0

0

1

2

3

4

5

6

7

8

9

10 11 12

0 2

0.00 ±0.00

-5.40 ±0.32

-6.97 ±0.33

-8.40 ±0.29

-9.34 ±0.26

-11.78 ±0.35

-16.92 ±0.29

0.00 ±0.00

-8.12 ±0.26

-10.57 ±0.21

-11.81 ±0.16

-12.62 ±0.13

-14.39 ±0.18

-19.36 ±0.22

0.00 ±0.00

-6.84 ±0.23

-8.00 ±0.30

-9.58 ±0.14

-10.41 ±0.22

-12.76 ±0.27

-18.64 ±0.30

0

1

2

3

4

5

6

0 2 4 0 2 4

-2.35 ±0.25

-0.78 ±0.31

-1.30 ±0.29

0.00 ±0.00

-4.25 ±0.31

-4.86 ±0.24

-5.06 ±0.37

0.00 ±0.00

-5.28 ±0.17

-4.90 ±0.18

-5.03 ±0.20

0

1

2

3

chemberta-3

0.00 -2.97 -2.74 -3.24 -4.12 -5.16 -6.20 -7.65 -9.30 -11.89 -14.76 -18.01 -22.10 ±0.00 ±0.24 ±0.26 ±0.24 ±0.32 ±0.34 ±0.32 ±0.29 ±0.37 ±0.36 ±0.39 ±0.72 ±1.24

4

0 2 4 0 10 20

0.00 -3.85 -3.88 -4.33 -5.21 -6.06 -6.98 -8.53 -10.24 -13.17 -16.71 -20.75 -26.26 ±0.00 ±0.26 ±0.19 ±0.22 ±0.25 ±0.24 ±0.34 ±0.26 ±0.33 ±0.33 ±0.31 ±0.47 ±0.50

1

2

3

4

5

6

7

8

roberta-zinc-480m

9

10 11 12

0.00 0.79 1.04 0.96 0.91 0.66 0.58 0.33 0.20 -0.13 -0.44 -0.71 -1.08 -1.54 -2.15 ±0.00 ±0.11 ±0.15 ±0.16 ±0.13 ±0.13 ±0.14 ±0.14 ±0.12 ±0.12 ±0.14 ±0.15 ±0.20 ±0.13 ±0.16

2

10

0.00 -3.30 -3.68 -4.74 -5.61 -6.36 -7.15 -8.40 -9.81 -12.50 -15.58 -19.18 -25.08 ±0.00 ±0.37 ±0.41 ±0.43 ±0.36 ±0.23 ±0.16 ±0.24 ±0.32 ±0.35 ±0.33 ±0.34 ±0.24

0 0

chemberta-2-5M

0.00 ±0.00

0

0.00 -0.29 -0.12 -0.11 -0.14 -0.16 -0.31 -0.42 -0.61 -0.80 -1.22 -1.46 -1.89 -2.26 -2.73 ±0.00 ±0.15 ±0.17 ±0.18 ±0.12 ±0.28 ±0.26 ±0.19 ±0.09 ±0.15 ±0.21 ±0.22 ±0.14 ±0.15 ±0.15 0.00 0.14 0.40 0.48 0.46 0.36 0.16 0.04 0.02 -0.10 -0.42 -0.61 -0.89 -1.23 -1.70 ±0.00 ±0.16 ±0.12 ±0.17 ±0.14 ±0.19 ±0.12 ±0.11 ±0.15 ±0.12 ±0.11 ±0.16 ±0.19 ±0.18 ±0.13

0

1

2

Relative Layer Depth

3

4

5

6

7

8

0 2

9 10 11 12 13 14

Figure 3: Average relative differences in probing performance (% macro-averaged F1) in probing performance of PT (left) and RI (right) models after fine-tuning on lipophilicity. We group into hydrophilic (top), lipophilic (middle), and other (bottom) groups (see Table 6), with numbers indicating the group size. The RI results for chemberta2 are based on a single model (hence, are the same for all three variants) as all models use the same architecture.

6.5

Discussion

find that models that share the same architecture converge towards the same upper bound shape in terms of probing performance (Figure 14). Similarly, we conduct experiments for the chemberta-3 model which exhibits a strange dome in the upper layers. Further pre-training the model on two different datasets indicates that, again, this dome becomes less pronounced as probing performance also increases in the lower layers (cf. Figure 16). Most notably, we find that the pre-training dataset can make a substantial difference on the resulting probing performance. These experiments showcase how probing may be used to identify signs of undertraining of specific molecular substructures.

Our probing experiments on models before and after fine-tuning on lipophilicity (§6.4) and solubility (Section G) prediction reveal three major findings. First, molecular substructures that are theoretically more relevant for a downstream task undergo larger changes during fine-tuning. Reciprocally, molecular substructures that undergo major changes during fine-tuning for a specific downstream task might indicate a high importance. This might be especially interesting for tasks with a high variability in terms of important molecular substructures such as toxicity prediction. Second, changes are less pronounced compared to pretraining and occur more frequently in the upper layers. Considering the increasing model sizes and consequently, the increasing costs of probing, one way to reduce costs could be to restrict probing to the upper layers as they yield the largest changes. Third, pre-training increases the robustness of molecular substructures in CLMs, making them more likely to be retained during fine-tuning.

7

8

Conclusion

We have employed layer-wise probing to investigate the extent to which chemical language models (CLMs) trained on linearized molecular representations encode important molecular substructures. Our experiments across eight pre-trained and six randomly initialized models show that pre-training generally improves molecular structure awareness, with the most pronounced effects emerging in the upper layers. Although certain molecular substructures are unlearned during pre-training, only a small subset exhibits this behavior consistently across all models. Notably, we find that even randomly initialized models encode ring structures well, suggesting that these are surface-level properties of the input representations and do not require pre-training. Our fine-tuning analysis reveals that, on average, groups of molecular substructures relevant to a downstream task undergo larger representational changes than others. Finally, our further pre-training experiments

Practical Implications

The varying probing performance across different models (especially for the chemberta2 models with different pre-training data sizes, cf. §5) suggests that this might be attributed to a lack of molecules containing a specific molecular substructure during pre-training. To better understand the impact of pre-training data, we further pre-train the models. In particular, we curate pre-training datasets consisting of molecules with molecular substructures on which a model underperforms. Our results show that further pre-training on these molecules does indeed improve a model’s internal representation of them (Section H) . Moreover, we 8

showcase how probing can be used to identify and mitigate gaps using a small and carefully curated dataset.

structure dataset corresponding to a distinct subset of the original 100k molecules, both in terms of size and composition.

9

Limitations of probing Probing remains a debated diagnostic method as it does not indicate whether a feature is used during prediction, but only how extractable the property is from the learned representations. Using linear probes (as in this work), we can therefore only assess the linear separability of the investigated property in these representations. Note, that there is no consensus on what probing classifier to use. While some works argue that utilizing linear classifiers prevents the possibility of memorization (Belinkov, 2022), others propose to instead consider memorization during evaluation (Pimentel et al., 2020).

Limitations

True complexity of lipophilicity and solubility prediction Although Hansch et al. (1977) estimate the lipophilicity of a molecule using a linear combination of its substructures, we note that this approach is an oversimplification of the underlying process. The actual lipophilicity of a molecule also depends on the environment a certain substructure is in (i.e., other neighboring substructures) as well as its depth and position in the molecule and is subject to further research in chemistry. While the logD value (measuring the lipophilicity) is part of the general solubility equation (corrected for ionization at pH 7.4), there are various other factors that influence the logS value (Hill and Young, 2010). Some of these challenges are part of ongoing chemistry research as highlighted by Llompart et al. (2024), who find that accurately predicting the aqueous solubility requires the knowledge of many factors such as the solid-solvated phase transition, solid state, temperature, polymorphism, intermolecular interactions between solutesolvent etc.

More complex probing task Our probing tasks focus on detecting the presence of a molecular substructure, rather than counting occurrences or identifying its location. Consequently, some probing tasks may be easier to learn due to the underlying nature of a molecule. For instance, a molecule may contain multiple occurrences of a single molecular substructure, making it easier to detect its presence (as is the case for aromatic rings that often occur multiple times in many molecules). More challenging tasks could offer additional insights but are subject to future investigation; and moreover, also require respective chemistry research.

Other downstream tasks Our experiments focus on lipophilicity and solubility prediction as the downstream tasks. While there exist other tasks such as predicting the bioactivity and toxicity of molecules, one limiting factors is the number of publicly available molecules that have been studied with the same experimental conditions. Moreover, many tasks are still subject to ongoing chemical research and are not understood well (yet). For instance, a non-trivial challenge in predicting the bioactivity are activity cliffs (Stumpfe et al., 2019), pairs of molecules with highly similar structures— i.e., close proximity in the molecular “landscape”— but different magnitudes in terms of bioactivity, resulting in a steep “cliff”. A better chemical understanding of the underlying process would allow researchers to build models that capture fine-grained structural differences between molecules.

10

Impact Statement

The primary goal of this work is to provide insights on how chemical language models are affected by pre-training and fine-tuning on chemical data. While this work falls under the category of fundamental research without direct implications on downstream applications, the authors acknowledge that some of the findings (e.g., fine-tuning mostly aligns with chemical theory) may lead to research which could be abused to identify molecular substructures that are harmful to the human body. The authors emphasize that the experiments and tasks presented here have no direct connection to harmful applications (including tasks such as predicting the toxicity of molecules).

Effects of random- and downsampling We note that due to the random sampling of the three probing datasets and the downsampling for individual molecular substructures, the molecules across different splits may vary with each molecular sub-

Acknowledgments We thank Shubham Dokania, Max Rausch-Dupont and Afnan Sultan for their helpful discussions and feedback. This work was funded by the Deutsche 9

Forschungsgemeinschaft (DFG, German Research Foundation) – GRK 2853/1 “Neuroexplicit Models of Language, Vision, and Action” - project number 471607914.

Dimitrios Christofidellis, Giorgio Giannone, Jannis Born, Ole Winther, Teodoro Laino, and Matteo Manica. 2023. Unifying molecular and textual representations via multi-task language modelling. In Proceedings of the 40th International Conference on Machine Learning, volume 202 of Proceedings of Machine Learning Research, pages 6140–6157. PMLR.

References

William Curatolo. 1998. Physical chemical properties of oral drug candidates in the discovery and exploratory development settings. Pharmaceutical Science & Technology Today, 1(9):387–393.

Walid Ahmad, Elana Simon, Seyone Chithrananda, Gabriel Grand, and Bharath Ramsundar. 2022. Chemberta-2: Towards chemical foundation models. Preprint, arXiv:2209.01712.

Ana Laura Dias, Latimah Bustillo, and Tiago Rodrigues. 2023. Limitations of representation learning in small molecule property prediction. Nature Communications, 14:6394.

Mohammad Sadegh Akhondzadeh, Vijay Lingam, and Aleksandar Bojchevski. 2023. Probing graph representations. In International Conference on Artificial Intelligence and Statistics, pages 11630–11649. PMLR.

Carl Edwards, Tuan Lai, Kevin Ros, Garrett Honke, Kyunghyun Cho, and Heng Ji. 2022. Translation between molecules and natural language. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pages 375–413, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics.

Nawaf Alampara, Mara Schilling-Wilhelmi, Martiño Ríos-García, Indrajeet Mandal, Pranav Khetarpal, Hargun Singh Grover, NM Anoop Krishnan, and Kevin Maik Jablonka. 2025. Probing the limitations of multimodal language models for chemistry and materials research. Nature computational science, 5(10):952–961.

Carl Edwards, ChengXiang Zhai, and Heng Ji. 2021. Text2Mol: Cross-modal molecule retrieval with natural language queries. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pages 595–607, Online and Punta Cana, Dominican Republic. Association for Computational Linguistics.

Viraj Bagal, Rishal Aggarwal, P. K. Vinod, and U. Deva Priyakumar. 2022. Molgpt: Molecular generation using a transformer-decoder model. Journal of Chemical Information and Modeling, 62(9):2064–2076. PMID: 34694798.

Benedek Fabian, Thomas Edlich, Héléna Gaspar, Marwin Segler, Joshua Meyers, Marco Fiscato, and Mohamed Ahmed. 2020. Molecular representation learning with language models and domain-relevant auxiliary tasks. Preprint, arXiv:2011.13230.

Hartmut Beck, Michael Härter, Bastian Haß, Carsten Schmeck, and Lars Baerfacker. 2022. Small molecules and their impact in drug discovery: A perspective on the occasion of the 125th anniversary of the bayer chemical research laboratory. Drug Discovery Today, 27(6):1560–1574.

Yin Fang, Xiaozhuan Liang, Ningyu Zhang, Kangwei Liu, Rui Huang, Zhuo Chen, Xiaohui Fan, and Huajun Chen. 2024. Mol-instructions: A large-scale biomolecular instruction dataset for large language models. In The Twelfth International Conference on Learning Representations.

Yonatan Belinkov. 2021. Probing classifiers: Promises, shortcomings, and advances. Preprint, arXiv:2102.12452. Yonatan Belinkov. 2022. Probing classifiers: Promises, shortcomings, and advances. Computational Linguistics, 48(1):207–219.

Mohsen Fayyaz, Ehsan Aghazadeh, Ali Modarressi, Hosein Mohebbi, and Mohammad Taher Pilehvar. 2021. Not all models localize linguistic knowledge in the same place: A layer-wise probing on BERToids’ representations. In Proceedings of the Fourth BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 375–388, Punta Cana, Dominican Republic. Association for Computational Linguistics.

Nathan Brown, Marco Fiscato, Marwin H.S. Segler, and Alain C. Vaucher. 2019. Guacamol: Benchmarking models for de novo molecular design. Journal of Chemical Information and Modeling, 59(3):1096– 1108. PMID: 30887799. Raymond E. Carhart, Dennis H. Smith, and R. Venkataraghavan. 1985. Atom pairs as molecular features in structure-activity studies: definition and applications. Journal of Chemical Information and Computer Sciences, 25(2):64–73.

Inken Fender, Jannik Adrian Gut, and Thomas Lemmin. 2025. Beyond performance: how design choices shape chemical language models. Journal of Cheminformatics, 17(1):1–15.

Seyone Chithrananda, Gabriel Grand, and Bharath Ramsundar. 2020. Chemberta: Large-scale selfsupervised pretraining for molecular property prediction. ArXiv, abs/2010.09885.

David Fooshee, Aaron Mood, Eugene Gutman, Mohammadamin Tavakoli, Gregor Urban, Frances Liu, Nancy Huynh, David Van Vranken, and Pierre Baldi. 2018. Deep learning for chemical reaction prediction.

10

Molecular Systems Design & Engineering, 3(3):442– 452.

John Hewitt and Christopher D. Manning. 2019. A structural probe for finding syntax in word representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 4129–4138, Minneapolis, Minnesota. Association for Computational Linguistics.

Veronika Ganeeva, Kuzma Khrabrov, Artur Kadurin, and Elena Tutubalina. 2025. Two steps from hell: Compositionality on chemical LMs. In Findings of the Association for Computational Linguistics: EMNLP 2025, pages 1042–1049, Suzhou, China. Association for Computational Linguistics.

Karl Heyer. 2023. Roberta-zinc-480m. Veronika Ganeeva, Andrey Sakhovskiy, Kuzma Khrabrov, Andrey Savchenko, Artur Kadurin, and Elena Tutubalina. 2024. Lost in translation: Chemical language models and the misunderstanding of molecule structures. In Findings of the Association for Computational Linguistics: EMNLP 2024, pages 12994–13013, Miami, Florida, USA. Association for Computational Linguistics.

Alan P. Hill and Robert J. Young. 2010. Getting physical in drug discovery: a contemporary perspective on solubility and hydrophobicity. Drug Discovery Today, 15(15):648–655. Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2020. The curious case of neural text degeneration. In International Conference on Learning Representations.

Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad AlDahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, Amy Yang, Angela Fan, Anirudh Goyal, Anthony Hartshorn, Aobo Yang, Archi Mitra, Archie Sravankumar, Artem Korenev, Arthur Hinsvark, and 542 others. 2024. The llama 3 herd of models. ArXiv, abs/2407.21783.

Yifan Hou and Mrinmaya Sachan. 2021. Bird’s eye: Probing for linguistic graph structures with a simple information-theoretic approach. In Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), pages 1844–1859, Online. Association for Computational Linguistics.

Kehan Guo, Bozhao Nan, Yujun Zhou, Taicheng Guo, Zhichun Guo, Mihir Surve, Zhenwen Liang, Nitesh V. Chawla, Olaf Wiest, and Xiangliang Zhang. 2024. Can llms solve molecule puzzles? a multimodal benchmark for molecular structure elucidation. In Advances in Neural Information Processing Systems, volume 37, pages 134721–134746. Curran Associates, Inc.

Weihua Hu, Matthias Fey, Hongyu Ren, Maho Nakata, Yuxiao Dong, and Jure Leskovec. 2021. Ogb-lsc: A large-scale challenge for machine learning on graphs. arXiv preprint arXiv:2103.09430. John J. Irwin and Brian K. Shoichet. 2005. Zinc - a free database of commercially available compounds for virtual screening. Journal of Chemical Information and Modeling, 45(1):177–182. PMID: 15667143.

Taicheng Guo, Kehan Guo, Bozhao Nan, Zhenwen Liang, Zhichun Guo, Nitesh V. Chawla, Olaf Wiest, and Xiangliang Zhang. 2023. What can large language models do in chemistry? a comprehensive benchmark on eight tasks. Preprint, arXiv:2305.18365.

John J. Irwin, Khanh G. Tang, Jennifer Young, Chinzorig Dandarchuluun, Benjamin R. Wong, Munkhzul Khurelbaatar, Yurii S. Moroz, John Mayfield, and Roger A. Sayle. 2020. Zinc20—a free ultralarge-scale chemical database for ligand discovery. Journal of Chemical Information and Modeling, 60(12):6065–6073. PMID: 33118813.

Corwin Hansch, Sharon D Rockwell, Priscilla YC Jow, Albert Leo, and Edward E Steller. 1977. Substituent constants for correlation analysis. Journal of medicinal chemistry, 20(2):304–306.

Ross Irwin, Spyridon Dimitriadis, Jiazhen He, and Esben Jannik Bjerrum. 2022. Chemformer: a pretrained transformer for computational chemistry. Machine Learning: Science and Technology, 3(1):1–13.

Marc W. Harrold, issuing body. American Society of Health-System Pharmacists, and Robin M. Zavod. 2023. Basic concepts in medicinal chemistry, 3rd edition. edition. ASHP, Bethesda, MD. Jiazhen He, Eva Nittinger, Christian Tyrchan, Werngard Czechtizky, Atanas Patronov, Esben Jannik Bjerrum, and Ola Engkvist. 2022. Transformer-based molecular optimization beyond matched molecular pairs. Journal of Cheminformatics, 14(1):18.

Yunhui Jang, Jaehyung Kim, and Sungsoo Ahn. 2025. Structural reasoning improves molecular understanding of LLM. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 21016–21036, Vienna, Austria. Association for Computational Linguistics.

Jiazhen He, Huifang You, Emil Sandström, Eva Nittinger, Esben Jannik Bjerrum, Christian Tyrchan, Werngard Czechtizky, and Ola Engkvist. 2021. Molecular optimization by capturing chemist’s intuition using deep neural networks. Journal of Cheminformatics, 13(1):26.

Ganesh Jawahar, Benoît Sagot, and Djamé Seddah. 2019. What does BERT learn about the structure of language? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 3651–3657, Florence, Italy. Association for Computational Linguistics.

11

Chen-Yang Jia, Jing-Yi Li, Ge-Fei Hao, and Guang-Fu Yang. 2020. A drug-likeness toolbox facilitates admet study in drug discovery. Drug Discovery Today, 25(1):248–258.

robust fine-tuning with molecular graph foundation models. NeurIPS. Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2020. Ro{bert}a: A robustly optimized {bert} pretraining approach.

Sunghwan Kim, Jie Chen, Tiejun Cheng, Asta Gindulyte, Jia He, Siqian He, Qingliang Li, Benjamin A Shoemaker, Paul A Thiessen, Bo Yu, Leonid Zaslavsky, Jian Zhang, and Evan E Bolton. 2018. Pubchem 2019 update: improved access to chemical data. Nucleic Acids Research, 47(D1):D1102–D1109.

Yunwu Liu, Ruisheng Zhang, Tongfeng Li, Jing Jiang, Jun Ma, and Ping Wang. 2023a. Molrope-bert: An enhanced molecular representation with rotary position embedding for molecular property prediction. Journal of Molecular Graphics and Modelling, 118:108344.

Gilles Klopman, Shaomeng Wang, and D. M. Balthasar. 1992. Estimation of aqueous solubility of organic molecules by the group contribution approach. application to the study of biodegradation. Journal of Chemical Information and Computer Sciences, 32(5):474–482. PMID: 1400663.

Zequn Liu, Wei Zhang, Yingce Xia, Lijun Wu, Shufang Xie, Tao Qin, Ming Zhang, and Tie-Yan Liu. 2023b. MolXPT: Wrapping molecules with text for generative pre-training. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 2: Short Papers), pages 1606–1616, Toronto, Canada. Association for Computational Linguistics.

Mario Krenn, Florian Häse, AkshatKumar Nigam, Pascal Friederich, and Alan Aspuru-Guzik. 2020. Selfreferencing embedded strings (selfies): A 100% robust molecular string representation. Machine Learning: Science and Technology, 1(4):045024. Greg Landrum, Paolo Tosco, Brian Kelley, Ricardo Rodriguez, David Cosgrove, Riccardo Vianello, sriniker, Peter Gedeck, Gareth Jones, NadineSchneider, Eisuke Kawashima, Dan Nealschneider, Andrew Dalke, Matt Swain, Brian Cole, Samo Turk, Aleksandr Savelev, Alain Vaucher, Maciej Wójcikowski, and 11 others. 2024. rdkit/rdkit: 2024_03_6 (q1 2024) release.

Pol Llompart, Christian Minoletti, Sapark Baybekov, Dragos Horvath, Gilles Marcou, and Alexandre Varnek. 2024. Will we ever be able to accurately predict solubility? Scientific Data, 11(1):303. Jieyu Lu and Yingkai Zhang. 2022. Unified deep learning model for multitask reaction predictions with explanation. Journal of Chemical Information and Modeling, 62(6):1376–1387. PMID: 35266390.

Matthew L. Landry and James J. Crawford. 2020. Logd contributions of substituents commonly used in medicinal chemistry. ACS Medicinal Chemistry Letters, 11(1):72–76.

Gerald M Maggiora. 2006. On outliers and activity cliffs why qsar often disappoints. Journal of chemical information and modeling, 46(4):1535–1535. Favour Danladi Makurvet. 2021. Biologics vs. small molecules: Drug costs and patient access. Medicine in Drug Discovery, 9:100075.

Miguelangel Leon, Yuriy Perezhohin, Fernando Peres, Aleš Popovič, and Mauro Castelli. 2024. Comparing smiles and selfies tokenization for enhanced chemical language modeling. Scientific Reports, 14. Juncai Li and Xiaofei Jiang. 2021. Mol-bert: An effective molecular representation with bert for molecular property prediction. Wireless Communications and Mobile Computing, 2021(1):1–7.

Raimund Mannhold, Gennadiy I. Poda, Claude Ostermann, and Igor V. Tetko. 2009. Calculation of molecular lipophilicity: State-of-the-art and comparison of logp methods on more than 96,000 compounds. Journal of Pharmaceutical Sciences, 98(3):861–893.

Xuan Lin, Long Chen, and Yile Wang. 2026. Attrilensmol: Attribute guided reinforcement learning for molecular property prediction with large language models. Preprint, arXiv:2508.04748.

Eyal Mazuz, Guy Shtar, Bracha Shapira, and Lior Rokach. 2023. Molecule generation using transformers and policy gradient reinforcement learning. Scientific Reports, 13(8799).

Nelson F. Liu, Matt Gardner, Yonatan Belinkov, Matthew E. Peters, and Noah A. Smith. 2019. Linguistic knowledge and transferability of contextual representations. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), pages 1073–1094, Minneapolis, Minnesota. Association for Computational Linguistics.

Amil Merchant, Elahe Rahimtoroghi, Ellie Pavlick, and Ian Tenney. 2020. What happens to BERT embeddings during fine-tuning? In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 33–44, Online. Association for Computational Linguistics. Beata Morak-Młodawska, Małgorzata Jeleń, Emilia Martula, and Rafał Korlacki. 2023. Study of lipophilicity and adme properties of 1,9diazaphenothiazines with anticancer action. International Journal of Molecular Sciences, 24(8).

Shikun Liu, Deyu Zou, Nima Shoghi, Victor Fung, Kai Liu, and Pan Li. 2025. Roft-mol: Benchmarking

12

H. L. Morgan. 1965. The generation of a unique machine description for chemical structures-a technique developed at chemical abstracts service. Journal of Chemical Documentation, 5(2):107–113.

pages 4609–4622, Online. Association for Computational Linguistics. Douglas E. V. Pires, Tom L. Blundell, and David B. Ascher. 2015. pkcsm: Predicting small-molecule pharmacokinetic and toxicity properties using graphbased signatures. Journal of Medicinal Chemistry, 58(9):4066–4072. PMID: 25860834.

Marius Mosbach, Anna Khokhlova, Michael A. Hedderich, and Dietrich Klakow. 2020. On the interplay between fine-tuning and sentence-level probing for linguistic knowledge in pre-trained transformers. In Proceedings of the Third BlackboxNLP Workshop on Analyzing and Interpreting Neural Networks for NLP, pages 68–82, Online. Association for Computational Linguistics.

Jerret Ross, Brian Belgodere, Vijil Chenthamarakshan, Inkit Padhi, Youssef Mroueh, and Payel Das. 2022. Large-scale chemical language representations capture molecular structure and properties. Nature Machine Intelligence, 4(12):1256–1264.

Amin Naghdloo, Dean Tessone, Rajiv M Nagaraju, Brian Zhang, Jeffrey Kang, Shouyi Li, Assad Oberai, James B Hicks, and Peter Kuhn. 2025. Representation learning enables robust single cell phenotyping in whole slide liquid biopsy imaging. Scientific Reports, 15(1):36589.

Shaghayegh Sadeghi, Ali Forooghi, Jianguo Lu, and Alioune Ngom. 2024. Moleval: An evaluation toolkit for molecular embeddings via LLMs. In ICML 2024 Workshop on Efficient and Accessible Foundation Models for Biological Discovery. Franco Scarselli, Marco Gori, Ah Chung Tsoi, Markus Hagenbuchner, and Gabriele Monfardini. 2009. The graph neural network model. IEEE Transactions on Neural Networks, 20(1):61–80.

Ramaswamy Nilakantan, Norman Bauman, J. Scott Dixon, and R. Venkataraghavan. 1987. Topological torsion: a new molecular descriptor for sar applications. comparison with other descriptors. Journal of Chemical Information and Computer Sciences, 27(2):82–85.

Philippe Schwaller, Teodoro Laino, Théophile Gaudin, Peter Bolgar, Christopher A. Hunter, Costas Bekas, and Alpha A. Lee. 2019. Molecular transformer: A model for uncertainty-calibrated chemical reaction prediction. ACS Central Science, 5(9):1572–1583. PMID: 31572784.

OpenAI. 2025. gpt-oss-120b & gpt-oss-20b model card. Preprint, arXiv:2508.10925. Jun-Hyung Park, Yeachan Kim, Mingyu Lee, Hyuntae Park, and SangKeun Lee. 2024a. MolTRES: Improving chemical language representation learning for molecular property prediction. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 14241–14254, Miami, Florida, USA. Association for Computational Linguistics.

Jie Shen and Christos A Nicolaou. 2019. Molecular property prediction: recent trends in the era of artificial intelligence. Drug Discovery Today: Technologies, 32:29–36. Riya Singh, Aryan Amit Barsainyan, Rida Irfan, Connor Joseph Amorin, Stewart He, Tony Davis, Arun Thiagarajan, Shiva Sankaran, Seyone Chithrananda, Walid Ahmad, Derek Jones, Kevin McLoughlin, Hyojin Kim, Anoushka Bhutani, Shreyas Vinaya Sathyanarayana, Venkat Viswanathan, Jonathan E. Allen, and Bharath Ramsundar. 2026. Chemberta-3: an open source training framework for chemical foundation models. Digital Discovery, 5:662–685.

Jun-Hyung Park, Hyuntae Park, Yeachan Kim, Woosang Lim, and SangKeun Lee. 2024b. Moleco: Molecular contrastive learning with chemical language models for molecular property prediction. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing: Industry Track, pages 408–420, Miami, Florida, US. Association for Computational Linguistics.

Dagmar Stumpfe, Huabin Hu, and Jurgen Bajorath. 2019. Evolving concept of activity cliffs. ACS omega, 4(11):14360–14368.

Josh Payne, Mario Srouji, Dian Ang Yap, and Vineet Kosaraju. 2020. Bert learns (and teaches) chemistry. Preprint, arXiv:2007.16012.

Ian Tenney, Dipanjan Das, and Ellie Pavlick. 2019a. BERT rediscovers the classical NLP pipeline. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4593– 4601, Florence, Italy. Association for Computational Linguistics.

Matthew E. Peters, Mark Neumann, Luke Zettlemoyer, and Wen-tau Yih. 2018. Dissecting contextual word embeddings: Architecture and representation. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 1499– 1509, Brussels, Belgium. Association for Computational Linguistics.

Ian Tenney, Patrick Xia, Berlin Chen, Alex Wang, Adam Poliak, R Thomas McCoy, Najoung Kim, Benjamin Van Durme, Sam Bowman, Dipanjan Das, and Ellie Pavlick. 2019b. What do you learn from context? probing for sentence structure in contextualized word representations. In International Conference on Learning Representations.

Tiago Pimentel, Josef Valvoda, Rowan Hall Maudslay, Ran Zmigrod, Adina Williams, and Ryan Cotterell. 2020. Information-theoretic probing for linguistic structure. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics,

13

Prudencio Tossou, Cas Wognum, Michael Craig, Hadrien Mary, and Emmanuel Noutahi. 2024. Realworld molecular out-of-distribution: Specification and investigation. Journal of Chemical Information and Modeling, 64(3):697–711. PMID: 38300258.

like computational chemists. Briefings in Bioinformatics, 23(3):1–13. Jun Xia, Lecheng Zhang, Xiao Zhu, Yue Liu, Zhangyang Gao, Bozhen Hu, Cheng Tan, Jiangbin Zheng, Siyuan Li, and Stan Z. Li. 2023. Understanding the limitations of deep models for molecular property prediction: Insights and solutions. In Thirtyseventh Conference on Neural Information Processing Systems.

Mikhail Volkov, Joseph-André Turk, Nicolas Drizard, Nicolas Martin, Brice Hoffmann, Yann GastonMathé, and Didier Rognan. 2022. On the frustration to predict binding affinities from protein–ligand structures with deep neural networks. Journal of medicinal chemistry, 65(11):7946–7958.

Ziting Xian, Jiawei Gu, Lingbo Li, and Shangsong Liang. 2025. MolRAG: Unlocking the power of large language models for molecular property prediction. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 15513–15531, Vienna, Austria. Association for Computational Linguistics.

Michael A. Walker. 2017. Improvement in aqueous solubility achieved via small molecular changes. Bioorganic & Medicinal Chemistry Letters, 27(23):5100– 5108. Hanchen Wang*, Jean Kaddour*, Shengchao Liu, Jian Tang, Joan Lasenby, and Qi Liu. 2023. Evaluating self-supervised learning for molecular graph embeddings. In NeurIPS 2023, Datasets and Benchmarks Track.

Dongyu Xue, Han Zhang, Dongling Xiao, Yukang Gong, Guohui Chuai, Yu Sun, Hao Tian, Hua Wu, Yukun Li, and Qi Liu. 2021. X-mol: large-scale pre-training for molecular understanding and diverse molecular analysis. bioRxiv.

Sheng Wang, Yuzhi Guo, Yuhong Wang, Hongmao Sun, and Junzhou Huang. 2019. Smiles-bert: Large scale unsupervised pre-training for molecular property prediction. In Proceedings of the 10th ACM International Conference on Bioinformatics, Computational Biology and Health Informatics, BCB ’19, page 429–436, New York, NY, USA. Association for Computing Machinery.

Geyan Ye, Xibao Cai, Houtim Lai, Xing Wang, Junhong Huang, Longyue Wang, Wei Liu, and Xiangxiang Zeng. 2025. Drugassist: a large language model for molecule optimization. Briefings in Bioinformatics, 26(1):1–12. Seongjun Yun, Minbyul Jeong, Raehyun Kim, Jaewoo Kang, and Hyunwoo J Kim. 2019. Graph transformer networks. In Advances in Neural Information Processing Systems, pages 1–11.

Ye Wang, Honggang Zhao, Simone Sciabola, and Wenlu Wang. 2023. cmolgpt: A conditional generative pretrained transformer for target-specific de novo molecular generation. Molecules, 28(11):4430.

Atakan Yüksel, Erva Ulusoy, Atabey Ünlü, and Tunca Doğan. 2023. Selformer: molecular representation learning via selfies language models. Machine Learning: Science and Technology, 4(2):025035.

Michael J. Waring. 2009. Defining optimum lipophilicity and molecular weight ranges for drug candidates—molecular weight dependent lower logd limits based on permeability. Bioorganic & Medicinal Chemistry Letters, 19(10):2844–2851.

Lawrence Zhao, Carl Edwards, and Heng Ji. 2023. What a scientific language model knows and doesn’t know about chemistry. In NeurIPS 2023 AI for Science Workshop.

Michael J. Waring, John Arrowsmith, Andrew R. Leach, Paul D. Leeson, Sam Mandrell, Robert M. Owen, Garry Pairaudeau, William D. Pennie, Stephen D. Pickett, Jibo Wang, Owen Wallace, and Alex Weir. 2015. An analysis of the attrition of drug candidates from four major pharmaceutical companies. Nature Reviews Drug Discovery, 14(7):475–486.

Yichu Zhou and Vivek Srikumar. 2022. A closer look at how fine-tuning changes BERT. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1046–1061, Dublin, Ireland. Association for Computational Linguistics.

David Weininger. 1988. Smiles, a chemical language and information system. 1. introduction to methodology and encoding rules. Journal of Chemical Information and Computer Sciences, 28(1):31–36. Zhenqin Wu, Bharath Ramsundar, Evan N. Feinberg, Joseph Gomes, Caleb Geniesse, Aneesh S. Pappu, Karl Leswing, and Vijay Pande. 2018. Moleculenet: A benchmark for molecular machine learning. Preprint, arXiv:1703.00564.

A

Chemical Background for Molecular Modeling

A.1

Drug Development

Drug development aims at identifying molecules that are both effective against a target disease and can be safely administered to humans (Jia et al., 2020). One crucial bottleneck in drug development is synthesizing potential molecules in the laboratory, which is a time- and cost-intensive process.

Zhenxing Wu, Dejun Jiang, Jike Wang, Xujun Zhang, Hongyan Du, Lurong Pan, Chang-Yu Hsieh, Dongsheng Cao, and Tingjun Hou. 2022. Knowledgebased bert: a method to extract molecular features

14

Machine learning offers the potential to accelerate this process by improving efficiency and reducing costs: for example, by prioritizing candidate compounds with desirable properties, thereby reducing the need to synthesize nonviable molecules. Consequently, molecular property prediction—the estimation of physicochemical, biological and functional properties such as lipophilicity, toxicity, permeability, and reactivity—constitutes a core component of this process (Waring et al., 2015; Pires et al., 2015). In particular, ADMET properties, which characterize the drug-likeness of a compound, are essential for assessing how a drug candidate will interact with the human body. In this work, we focus on small molecules (those with molecular weight ≤ 1000 Da), which possess better ADMET profiles (Beck et al., 2022) and constitute most approved pharmaceuticals (Makurvet, 2021).

canonicalization of SMILES enforces a deterministic graph traversal order, resulting in a one-to-one mapping between molecular graphs and SMILES strings. SMILES encode some graph structural information explicitly. For example, double bonds are represented by =, triple bonds by #. Single bonds, on the other hand, are not explicitly specified. Thus, consecutive atoms are assumed to be connected by either a single or an aromatic bond. Furthermore, as shown in Figure 4, when SMILES are canonicalized, atoms in aromatic rings (i.e., rings with alternating single and double bonds) are represented with lowercase letters.

Molecular property prediction Typically, molecular property prediction encompasses a wide range of tasks such as predicting bioactivity, solubility, permeability, and toxicity which are often addressed by different models. It is important to note that many of these tasks are not classification, but regression tasks. Reliably predicting the magnitude of even a single property can be challenging, as the arrangement and composition of molecular substructures play an important role. In many cases, models must capture fine-grained structural differences between molecules that lead to considerably different behavior. One such example is activity cliffs (Stumpfe et al., 2019), which describe pairs of molecules with highly similar structures—i.e., close proximity in the molecular “landscape”—but different magnitudes in terms of bioactivity, resulting in a steep “cliff”. Activity cliffs have been subject to chemistry research for over a decade, with evolving insights on what molecular substructures constitute them (Maggiora, 2006; Stumpfe et al., 2019; Xia et al., 2023). A.2

COc1ccc(Cl)cc1C(=O)NCCc2ccc(S(=O)(=O)NC(=O)NC3CCCCC3)cc2

Figure 4: An example molecule: Glibenclamide and its SMILES representation. The two aromatic rings are highlighted in green and orange. Compared to the nonaromatic ring (purple), all atoms in the aromatic rings are lowercased.

A.3

Lipophilicity Prediction

Lipophilicity refers to the ability of a chemical compound to dissolve in fat-like solvents (lipids, fats, oils; Morak-Młodawska et al. 2023). It is an important physicochemical property of molecules which correlates with the (oral) absorption, (tissue) distribution, metabolism, excretion, and toxcicity (ADMET) properties of drugs (Mannhold et al., 2009), essential in determining how a candidate drug will interact with the human body (Waring, 2009). Compared to other MPP tasks such as toxicity prediction, the impact of specific molecular substructures is better understood for lipophilicity prediction. For instance, Landry and Crawford

Linearized Representations

In order to process molecules in models or handcrafted algorithms, works have devised various methods. One such method called SMILES proposes to represent molecules as linearized representation of molecular graphs (Weininger, 1988). The standard SMILES encoding is non-unique, resulting in a one-to-many mapping between a molecule and its possible SMILES representations. However, 15

Chemberta Chithrananda et al. (2020) release multiple versions of a six-layer model based on RoBERTa (Liu et al., 2020) trained on different amounts of data sampled. We use the two publicly available models, namely, seyonec/ChemBERTa-zinc-base-v1 (here, referred to as chemberta-base; trained on 100k molecules) and seyonec/chemberta-zinc250k-v1 (chemberta; trained on 250k molecules). Both employ a BPE tokenizer with respective vocabulary sizes of 767 and 52k tokens. While Chemberta has 83,450,880 parameters, chemberta-base has only 44,103,936 parameters.

(2020) study how specific molecular substructures affect the lipophilicity of a molecule. The goal of lipophilicity prediction is to estimate the octanol/water distribution coefficient (i.e., logD at pH 7.4) of a specific molecule. Chemical literature distinguishes between two groups of molecular substructures that are known to affect lipophilicity (Harrold et al., 2023). In particular, hydrophilic substructures such as carboxylic acids substantially decrease a molecule’s logD value while lipophilic substructures such as aromatic rings increase the logD value. We follow this classification of molecular substructures in our analysis and put all other molecular substructures that do not substantially affect lipophilicity into the other group (Section F provides a list of all molecular substructures). A.4

Chemberta-2 In subsequent work, Ahmad et al. (2022) release models pre-trained on a larger set of molecules. We investigate DeepChem/chemberta-5M-MLM, DeepChem/chemberta-10M-MLM and DeepChem/chemberta-77M-MLM, denoted as chemberta-2-5M, chemberta-2-10M and chemberta-2-77M resepectively, as the other models are pre-trained with domain-specific auxiliary objectives5 . chemberta-2-5M/10M/77M models all share the same number of parameters – 3,427,440 parameters. Compared to the earlier version of Chemberta (Chithrananda et al., 2020), these models have only three encoder layers, a hidden layer size of 384, and a BPE tokenizer with a vocabulary size of 600 tokens. The main difference among the three variants is the training data size: 5M, 10M, and 77M molecules, respectively. We note that chemberta2 models exhibit tokenization problems 6 . In particular, some halogens such as chlorine (Cl) and bromine (Br) are tokenized incorrectly.

Aqueous solubility

Aqueous solubility refers to the ability of a molecule to dissolve in water. Similar to lipophilicity, aqueous solubility is a physicochemical property of molecules which affects their biovailability as well as ADME profiles (Llompart et al., 2024). Therefore, predicting a molecule’s solubility is of high importance to the development of orally active drugs/compounds (Klopman et al., 1992). The goal of aqueous solubility prediction is to estimate the logS (log solubility of a molecule in water). A.5

Relation between Lipophilicity and Solubility

Generally, a high lipophilicity is negatively correlated with aqueous solubility, decreasing oral absorption of a drug (Curatolo, 1998). Moreover, both tasks share the same categories of important molecular subgroups—i.e., hydrophilic substructures increase the hydrogen-bonding ability of a molecule, making it more likely to be soluble in water but decreasing lipophilicity while lipophilic substructures increase its hydrophobicity. Although both tasks are closely related, computing the logD and logS values requries the consideration of additional, molecule-specific factors such as the melting point (Hill and Young, 2010).

B

Chemberta-3 Most recently, Singh et al. (2026) released their training framework along with a 12-layer model trained on 100M molecules DeepChem/ChemBERTa-100M-MLM chemberta-3). In contrast to the chemberta-2 models, chemberta-3 is again based on RoBERTa using a selftrained tokenizer with a vocabulary size

Model Selection

In our experiments, we consider the following CLMs trained with the masked language modeling (MLM) objective on datasets of small molecules. Note, that all models except for molformer were trained on molecules from the ZINC dataset (Irwin and Shoichet, 2005):

5

These are the chemberta-5M-MTR, chemberta-10M-MTR and chemberta-77M-MTR where MTR refers to multitask regression of 200 molecular properties. 6 see Issue 1 and Issue 2 and Issue 3

16

of 7,924 tokens. Finally, compared to the chemberta-2 models, both the size of the hidden vectors and the size of the intermediate FFNNs are larger (3072 and 768 dimensions, respectively). chemberta-3 has 92,126,976 parameters.

incorrect conformational information (GetConformer().Is3D() returned False), 3) chemistry problems (rdkit.Chem.DetectChemistryProblems is not empty), 4) other processing errors (e.g., due to unrecoverable rotational or double bond information). Finally, the remaining molecules were annotated with binary labels with 1 denoting the presence of a substructure. To extract functional group information, we used the rdkit.Chem.Fragments module from RDKit. Ring structure labels and bond information were annotated using rdkit.Chem.Lipinski. We then binarized the labels.

Molformer Ross et al. (2022) train a 12 layer RoBERTa-based model using linear attention, rotary positional embeddings and a regexbased tokenizer with a vocabulary size of 2362 tokens. We use the publicly available ibm-research/MoLFormer-XL-both-10pct model7 trained on 100M molecules sampled equally from ZINC and PubChem (Kim et al., 2018) with 44,375,040 parameters.

Runtime estimates The sheer number of instances in the PCQM4Mv2 dataset requires the subsampling of molecules for the probing dataset. For reference, running all 78 probing tasks for a single 14-layer model with 100k instances requires ∼5.4 hours. Scaling this to the whole dataset would require almost 200 hours for only one (out of four) experimental configuration.

roberta-zinc-480m (Heyer, 2023) is a RoBERTabased 14 layer, ∼102M parameter model trained on 480M molecules sampled from the ZINC database (Irwin and Shoichet, 2005) available on Huggingface under entropy/roberta_zinc_480m.

C.2

For the experiments, we select models that are comparable and only differ with respect to their architecture and training data sizes. We thus omit models trained with domain-specific auxiliary objectives such as MolBERT (Fabian et al., 2020), SELFormer (Yüksel et al., 2023), Mol-BERT (Li and Jiang, 2021), K-BERT (Wu et al., 2022) and those that are not publicly available (e.g., SMILESBERT, Wang et al. 2019).

C

Our final probing dataset comprises 78 probing tasks, as summarized in Table 2. The Molecular Substructure column lists the abbreviations used for each probing task, while Description provides a brief explanation of the corresponding task. C.3

Dataset Statistics

Table 3 summarizes the sizes of the downsampled training sets for each probing task. Table 4 reports the class distribution in the corresponding test sets. Note that many molecular substructures are highly infrequent, resulting in significant class imbalance. To address this, we evaluate probing performance using the macro-averaged F1 score, which weighs each class equally. For training the probing classifier, we also downsample instances in the training set (i.e., the probing classifier is always trained on a balanced dataset).

Probing Dataset

We derive our probing dataset from PCQM4Mv2, a publicly available dataset designed for predicting the HOMO-LUMO energy gap. It is part of the open graph benchmark (Hu et al., 2021) and is published under an open source license (CC-BY 4.0). C.1

Probing Tasks

Dataset Preprocessing

In Section 3, we described the preprocessing steps used to obtain the final set of probing datasets. Specifically, all molecules were preprocessed with RDKIT (Landrum et al., 2024)8 , and molecules were discarded if they exhibited any of the following issues: 1) invalid SMILES strings: (MolFromSmiles cannot be created), 2)

D

Probing Setup and Compute Infrastructure

D.1

Probing Setup

For each probing task and each encoder layer of a CLM, we train a logistic regression classifier using scikit learn9 with default parameters and set max_iter=2000. The input features are the hidden representations from the corresponding layer.

7 We note that the best model (Molformer-XL) whose results are reported by (Ross et al., 2022) is unavailable 8 version 2024.3.6

9

17

sklearn.linear_model.LogisticRegression

Molecular Substructure

Description

Molecular Substructure

Description

NHOH NO AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings HAcceptors HDonors Heteroatoms RotatableBonds SaturatedCarbocycles SaturatedHeterocycles SaturatedRings Ring Al_COO Al_OH Al_OH_noTert ArN Ar_COO Ar_N Ar_NH Ar_OH COO COO2 C_O C_O_noCOO C_S Imine NH0 NH1 NH2 N_O Ndealkylation1 Ndealkylation2 Nhpyrrole SH aldehyde alkyl_halide

NHs or OHs nitrogens and oxygens aliphatic (containing at least one non-aromatic bond) carbocycles aliphatic (containing at least one non-aromatic bond) heterocycles aliphatic (containing at least one non-aromatic bond) rings aromatic carbocycles aromatic heterocycles aromatic rings hydrogen bond acceptors hydrogen bond donors heteroatoms rotatable bonds saturated carbocycles saturated heterocycles saturated rings rings aliphatic carboxylic acids aliphatic hydroxyl groups aliphatic hydroxyl groups excluding tert-OH N functional groups attached to aromatics aromatic carboxylic acid aromatic nitrogens aromatic amines aromatic hydroxyl groups carboxylic acids carboxylic acids carbonyl O carbonyl O excluding COOH thiocarbonyl imines tertiary amines secondary amines primary amines hydroxylamine groups XCCNR groups tert-alicyclic amines (no heteroatoms, not quinine-like bridged N) H-pyrrole nitrogens thiol groups aldehydes alkyl halides

allylic_oxid amide amidine aniline aryl_methyl benzene bicyclic diazo ester ether furan guanido halogen hdrzine hdrzone imidazole imide ketone ketone_Topliss lactone methoxy morpholine nitrile nitro nitro_arom nitro_arom_nonortho oxime para_hydroxylation phenol phenol_noOrthoHbond piperdine piperzine priamide pyridine sulfide thiazole thiophene unbrch_alkane urea

allylic oxidation sites excluding steroid dienone amides amidine groups anilines aryl methyl sites for hydroxylation benzene rings bicyclic diazo groups esters ether oxygens (including phenoxy) furan guanidine groups halogens hydrazine groups hydrazone groups imidazole imide groups ketones ketones excluding diaryl, a,b-unsat. dienones, heteroatom on Calpha cyclic esters (lactones) methoxy groups -OCH3 morpholine nitriles nitro groups nitro benzene ring substituents non-ortho nitro benzene ring substituents oxime groups para-hydroxylation sites phenols phenolic OH excluding ortho intramolecular Hbond substituents piperdine piperzine primary amides pyridine thioether thiazole thiophene unbranched alkanes of at least 4 members (excludes halogenated alkanes) urea groups

Table 2: Probing tasks: Molecular substructure abbreviations and descriptions

iments required ∼ 32.5 GPU hours. In total, this amounts to approximately 678.9 CPU-only hours and 867.4 GPU hours.

We apply padding to the longest sequence in each batch. Probing experiments were conducted on CPU-only nodes on the same cluster as described in Section D.2. As an indicative reference, a single probing run for one probing task on a 14-layer model (e.g., roberta-zinc-480m) typically completes in approximately 4.14 minutes on an Intel Xeon E5-2620 CPU. Runtime varies slightly with the number of layers and probing tasks. D.2

E

We provide further evidence and analysis for our findings regarding the effect of pre-training on ring structures (§5.2) and individual molecular substructures (§5.3). Moreover, we perform analysis excluding the strangely behaving chemberta-3 model which we study more extensively in Section I. Excluding chemberta-3 this model, we observe multiple additional patterns which we did not discuss in the main paper (as the analysis in the main paper includes chemberta-3).

Computing Environment

All experiments were performed on a high performance computing cluster with varying CPU and GPU architectures. Fine-tuning experiments were conducted on a single GPU; either NVIDIA Titan X (12GB), NVidia V100 (32GB), or A100 (40GB). Probing experiments were executed on 2 CPUs equipped with AMD EPYC 7662 (64 cores, 512 GB RAM) and did not require a GPU. D.3

Effect of Pre-training (RQ1)

E.1

Ring Structures

For ease of visualization, we omitted the shading that shows the upper and lower quartiles in Figure 1. Figure 6 presents the average probing performance along with its variability for rings, other and all. Notably, rings have the lowest variability across both PT and RI models, while probing performance of other shows greated variability. We provide the 12 ring types used to compute rings in Table 5. We further analyze the representations of rings and other molecular substructures, finding that the representations of both RI ( ) and PT (x) models perform exceptionally well at identifying ring structures compared to all other groups ( and x).

Compute Time

Overall, ∼ 677.9 CPU-only hours were spent for probing. For fine-tuning, the experiments required ∼ 57.9 GPU hours for lipophilicity prediction and 77.3 GPU hours for solubility prediction. The downstream task experiments for the baselines required 692.5 GPU hours for gpt-oss-20b and 7.2 GPU hours for Llama-3.2-3B-Instruct. All baseline experiments using the feature-based models using fingerprints required less than 1 CPU-only hour in total. Finally, the further pre-training exper18

Moreover, both RI and PT models encode ring structures well already at the first encoder layer, suggesting that these surface-level patterns are easy for the models to extract directly from the input, even without pre-training. We also find that the benefit of pre-training diminishes for ring structures compared to that of other molecular substructures (i.e., the gap between x and is much smaller compared to the gap between x and ).

chemberta-base is shown in Figure 7c and Figure 7d, respectively. The three chemberta-2 models trained on 5M, 10M, and 77M molecules are shown in Figure 7e, Figure 7f, and Figure 7g, respectively. Finally, the latest chemberta-3 model’s performance is shown in Figure 7h. We additionally note the following for analysis: PT > RI In addition to the molecular substructures discussed in §5, we further identify consistent improvements across following substructures (when excluding chemberta-3). For hydroxy groups (Al_OH) and aldehyde as well as C_O, although to a lesser degree.

Different rings For different ring types, we find that—in contrast to aliphatic and aromatic rings—pre-training does improve the probing performance for saturated rings for all models except for chemberta-3.

Macro F1

1.0 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 1.0 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2 1.0 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2

Aliphatic Rings (RI)

Aromatic Rings (RI)

F

Aliphatic Rings (PT)

In the following, we provide details about the two downstream tasks, namely, lipophilicity prediction and solubility prediction. For each task, we provide a dataset description, report the hyperparameters considered for tuning, and elaborate the dataset pre-processing along with the final molecular substructures considered for probing. For both tasks, we use the splits provided by Ross et al. (2022).10

Aromatic Rings (PT)

F.1 Saturated Rings (RI)

chemberta chemberta-base chemberta-2-10M

chemberta-2-5M chemberta-2-77M chemberta-3

Lipophilicity

Lipophilicity dataset The lipophilicity dataset originates from MoleculeNet (Wu et al., 2018), an open source dataset (MIT license) and comprises 4,200 molecules and their corresponding logD values. Figure 8a shows the distribution of logD values in the training, validation and test splits. As can be seen, the distributions of the logD value follow a similar shape across the training, validation, and test data with peaks around 0–1. Table 8 shows that the validation and test sets have very similar distributions, with means close to −0.079 and −0.004, while the training set is centered around 0. The validation set values are more dispersed (σ= 1.044) compared to the similar variability of the training (σ= 1.0 and 1.01, respectively). The more negative logD values mean that the validation and test sets are centered slightly towards molecules with lower lipophilicity.

Saturated Rings (PT)

molformer roberta-zinc-480m

0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00

Relative Layer Depth

Figure 5: Probing performance on aliphatic (top), aromatic (mid), and saturated (bottom) rings for randomly initialized (left) and pre-trained (right) models. As can be seen, the differences between randomly initialized and pre-trained models are minimal for aliphatic and aromatic rings (except for chemberta-3, for which we provide an explanation in Section I). This shows that random initializations already encode these ring structures very well.

E.2

Fine-Tuning Setup

Individual Molecular Substructures

Hyperparameters We used the AdamW optimizer with the default settings of β = 0.9 and β = 0.999, ϵ = 1 × 10−8 and weight decay of 0.01, together with a linear decay of the learning rate every 10 epochs with γ = 0.1. We performed

Figure 7 illustrates the layer-wise differences in probing performance after pre-training for individual molecular substructures. Results for molformer are shown in Figure 7a, while roberta-zinc-480m is presented in Figure 7b. The performance of chemberta and

10

19

Available at https://github.com/IBM/molformer

chemberta

1.0

chemberta-base chemberta-2-10M chemberta-2-5M chemberta-2-77M

chemberta-3

molformer

roberta-zinc-480m

0.9

Macro F1

0.8 0.7 0.6 0.5 0.4 0.3 0.2

0.0

0.5

1.0 0.0

0.5

PT (rings) RI (rings) 1.0 0.0 0.5

maj (rings) PT (other) 1.0 0.0 0.5

RI (other) maj (other) 1.0 0.0

PT (all) RI (all) 0.5 1.0 0.0

maj (all) 0.5

1.0 0.0

0.5

1.0 0.0

0.5

1.0

Relative Layer Depth Figure 6: Average probing performance (macro-averaged F1 score ↑) of pre-trained (PT), randomly initialized (RI) models, and the majority class prediction (maj) on molecular substructures with shading between the upper and lower quartiles. We report average performance on all 12 classes of rings (avg rings), all other 66 substructures (avg other) and all 78 substructures (avg all).

grid search across different batch sizes and learning rates (presented in Table 9). We trained the pre-trained for up to 10 epochs. Note, that we did not conduct a separate hyperparameter search for randomly initialized models and instead, used the optimal hyperparameter values obtained for the corresponding pre-trained models. However, we adjusted the maximum number of training epochs for randomly initialized models to 20 epochs. In addition, because the pre-trained chemberta-2 models share the same architecture, we fine-tuned only a single randomly initialized chemberta-2 model, marked with * in Table 10. The final hyperparameters are presented in Table 10.

lipophilicity one (1,128 vs 4,200 molecules). We use the train-validation-test splits (80/10/10) provided by Ross et al. (2022). Figure 8b illustrates the distribution of logS values across the training, validation and test splits of the dataset. As can be seen, the distribution of logS values differs substantially between different splits of the datasets; considerably more than in the lipophilicity prediction dataset. Most notable is the tail difference around a logS value of -8, which occur more frequently in the validation set (∼ 7.5%) compared to the training and the test sets (≤ 2.5%). As shown in Figure 8, while the train and test means are centered around -3 (-3.05 and -2.92, respectively), the validation set has a slightly lower mean of -3.19. Furthermore, the validation set exhibits greater dispersion (standard deviation σ = 2.30) compared to the training (σ = 2.08) and test (σ = 2.06) sets. This might explain lower performance of almost all models on the ESOL dataset, as they were selected on a validation set with a different distribution compared to the training and test sets.

Probing Setup To conduct meaningful analyses of the effect of fine-tuning on molecular substructures, we excluded any molecular substructures that appear fewer than ten times in either the training or test split of the lipophilicity dataset. Table 6 summarizes the final set of 60 molecular substructures, along with their relative frequencies in the training, validation and test splits of the lipophilicity dataset. Because our fine-tuning analyses categorize molecular substructures as either relevant (hydrophilic and lipophilic) or non-relevant (other) (Section A.3), Table 6 also indicates the group assignment (hydrophilic, lipophilic, other) for each molecular substructures. F.2

Hyperparameters We largely followed the finetuning setup described for lipophilicity in the previous section and evaluate the same hyperparameter ranges (Table 9). We only adjusted the number of training epochs to 20, as our initial experiments showed a slower convergence of models on this task. We furthermore used two types of prediction heads: a linear one and a two-layer MLP. Table 11 lists the final hyperparameter values for fine-tuning the models on the ESOL dataset.

Solubility

Solubility dataset We fine-tune all models on the ESOL solubility dataset from the MoleculeNet benchmark (Wu et al., 2018). The dataset comprises 1,127 molecules and their corresponding logS values (log solubility in mols per liter). Notably, the dataset is considerably smaller than the

Probing Setup Our probing setup for solubility follows the general probing setup described in Section 4 and Section D.1. Similar to lipophilicity 20

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

-0.21 ±0.07 0.80 ±0.95 1.80 ±0.24 2.86 ±0.21 0.51 ±0.08 0.20 ±0.08 -2.23 ±0.19 -0.03 ±0.03 -0.82 ±0.70 -0.21 ±0.15 0.30 ±1.23 1.75 ±0.17 1.45 ±0.46 3.46 ±0.57 3.06 ±0.18 0.03 ±0.01 2.40 ±0.04 -0.10 ±0.35 0.29 ±0.26 2.89 ±0.32 2.55 ±0.40 -0.70 ±0.20 -9.65 ±0.67 2.64 ±0.64 2.33 ±0.36 2.12 ±0.22

1.07 ±0.26 -2.37 ±0.45 2.29 ±0.27 3.47 ±0.05 0.37 ±0.16 0.11 ±0.16 -0.35 ±0.07 -0.11 ±0.03 -3.67 ±0.13 1.31 ±0.09 -0.04 ±1.15 4.35 ±0.41 3.05 ±0.46 5.16 ±0.57 3.71 ±0.37 0.01 ±0.01 4.99 ±0.40 2.98 ±0.08 3.15 ±0.26 4.92 ±0.23 3.38 ±0.61 -0.46 ±0.03 -1.32 ±0.61 5.10 ±0.19 4.70 ±0.34 4.82 ±0.27

0.99 ±0.15 -3.86 ±0.65 2.39 ±0.46 3.66 ±0.16 0.37 ±0.14 0.05 ±0.09 -0.45 ±0.08 -0.13 ±0.05 -3.80 ±0.48 0.86 ±0.10 -1.48 ±0.89 5.91 ±0.37 3.65 ±0.44 5.32 ±0.46 4.40 ±0.17 0.03 ±0.03 8.23 ±0.12 3.65 ±0.22 3.75 ±0.05 6.96 ±0.88 5.43 ±0.19 -0.70 ±0.08 -5.40 ±0.73 6.20 ±0.89 7.95 ±0.26 7.64 ±0.49

1.73 ±0.40 -4.97 ±0.98 3.18 ±0.12 3.78 ±0.06 0.16 ±0.07 -0.12 ±0.04 -0.79 ±0.14 -0.15 ±0.04 -4.64 ±0.36 1.85 ±0.19 -1.20 ±1.12 7.20 ±0.25 4.87 ±0.38 5.47 ±0.57 4.28 ±0.37 -0.01 ±0.04 11.14 ±0.25 5.80 ±0.15 5.49 ±0.17 9.64 ±0.47 6.30 ±0.55 -0.95 ±0.10 -3.65 ±0.16 7.87 ±0.64 11.43 ±0.60 11.12 ±0.38

2.70 ±0.07 -6.71 ±1.17 3.29 ±0.15 4.22 ±0.21 0.08 ±0.09 -0.44 ±0.12 -1.10 ±0.08 -0.19 ±0.08 -7.25 ±1.30 2.70 ±0.32 -3.50 ±1.29 7.57 ±0.17 4.51 ±0.38 6.63 ±0.06 4.51 ±0.19 -0.00 ±0.01 14.10 ±0.48 7.28 ±0.33 6.88 ±0.53 15.50 ±0.26 7.85 ±0.70 -1.24 ±0.18 -4.41 ±0.47 11.54 ±0.41 14.70 ±0.44 14.34 ±0.41

2.70 ±0.22 -8.01 ±0.39 3.31 ±0.27 4.45 ±0.09 0.01 ±0.15 -0.46 ±0.18 -1.84 ±0.18 -0.25 ±0.03 -9.05 ±0.20 3.01 ±0.38 -4.16 ±0.79 8.11 ±0.67 4.96 ±0.65 6.62 ±0.29 4.68 ±0.41 -0.10 ±0.06 16.23 ±0.34 7.97 ±0.49 7.40 ±0.67 15.53 ±0.07 9.47 ±0.49 -2.01 ±0.06 -5.29 ±0.44 11.16 ±0.66 17.46 ±0.63 17.84 ±0.66

4.62 ±0.14 -9.77 ±1.18 4.28 ±0.44 4.44 ±0.12 -0.10 ±0.09 -0.34 ±0.16 -2.13 ±0.04 -0.44 ±0.09 -9.49 ±0.69 4.76 ±0.35 -5.24 ±0.29 8.41 ±0.39 6.11 ±0.44 7.31 ±0.07 4.87 ±0.45 -0.17 ±0.01 19.64 ±0.73 11.08 ±0.47 9.99 ±0.13 16.12 ±0.57 10.19 ±0.61 -2.36 ±0.12 -5.27 ±0.23 12.17 ±0.78 20.59 ±0.49 20.44 ±0.87

6.17 ±0.25 -9.66 ±1.19 5.05 ±0.22 4.62 ±0.43 -0.15 ±0.04 -0.14 ±0.24 -2.72 ±0.13 -0.65 ±0.04 -10.22 ±1.53 6.28 ±0.29 -6.47 ±0.76 9.48 ±0.88 6.90 ±0.31 7.57 ±0.62 4.90 ±0.28 -0.25 ±0.05 21.94 ±0.74 12.71 ±0.17 11.75 ±0.23 16.97 ±0.42 12.78 ±0.13 -3.13 ±0.05 -5.91 ±1.05 13.11 ±0.14 24.56 ±0.98 24.47 ±0.97

8.65 ±0.50 -10.73 ±0.95 5.37 ±0.24 4.57 ±0.37 -0.47 ±0.07 -0.34 ±0.29 -3.17 ±0.07 -0.90 ±0.08 -11.56 ±1.49 9.37 ±0.51 -6.46 ±1.15 9.97 ±1.09 7.86 ±0.24 7.54 ±0.47 4.44 ±0.26 -0.40 ±0.02 22.22 ±0.77 14.58 ±0.13 13.80 ±0.07 19.51 ±0.20 13.72 ±0.77 -3.69 ±0.14 -7.15 ±0.09 14.19 ±0.44 25.20 ±0.84 24.75 ±0.33

10.98 ±0.49 -10.37 ±1.21 5.97 ±0.38 4.11 ±0.34 -0.75 ±0.12 -0.54 ±0.15 -3.55 ±0.17 -1.02 ±0.10 -11.25 ±0.40 11.38 ±0.51 -6.44 ±0.49 10.26 ±0.31 8.24 ±0.21 6.85 ±0.22 3.68 ±0.27 -0.66 ±0.06 25.10 ±0.69 17.08 ±0.21 16.58 ±0.40 21.49 ±0.52 16.81 ±0.97 -4.34 ±0.38 -9.15 ±0.09 16.59 ±0.63 28.36 ±0.47 28.27 ±0.28

12.13 ±0.41 -11.52 ±0.75 6.08 ±0.29 4.49 ±0.30 -0.84 ±0.09 -1.02 ±0.33 -3.97 ±0.12 -1.36 ±0.09 -12.20 ±1.33 12.67 ±0.59 -6.55 ±1.00 10.27 ±0.50 8.61 ±0.35 7.26 ±0.68 3.83 ±0.25 -1.11 ±0.16 28.08 ±0.37 19.21 ±0.28 18.73 ±0.20 25.67 ±0.95 18.40 ±1.90 -4.84 ±0.14 -13.08 ±0.91 20.34 ±0.19 31.15 ±0.24 31.08 ±0.43

13.09 ±0.61 -9.88 ±0.83 6.94 ±0.41 5.13 ±0.15 -1.51 ±0.02 -0.29 ±0.17 -3.70 ±0.17 -1.67 ±0.05 -10.30 ±0.73 13.82 ±0.54 -5.03 ±0.50 10.73 ±0.87 9.55 ±0.17 8.60 ±0.56 3.64 ±0.32 -1.70 ±0.27 29.00 ±0.41 20.54 ±0.21 19.45 ±0.39 26.64 ±0.73 20.40 ±1.97 -4.58 ±0.18 -12.03 ±0.82 21.96 ±1.15 32.19 ±0.31 31.49 ±0.30

0

1

2

3

4

5

6

7

8

9

10

11

12

C_O C_O_noCOO C_S Imine NH0 NH1 NH2 N_O Ndealkylation1 Ndealkylation2 Nhpyrrole SH aldehyde alkyl_halide allylic_oxid amide amidine aniline aryl_methyl benzene bicyclic ester ether furan guanido halogen

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

0.91 ±0.06 1.54 ±0.16 -4.52 ±0.96 0.85 ±0.40 -0.75 ±0.14 -0.79 ±0.21 0.73 ±0.34 -2.60 ±0.86 1.91 ±0.66 3.44 ±0.86 -9.65 ±0.67 -2.41 ±0.14 0.62 ±0.40 -4.92 ±0.77 -0.79 ±0.07 3.27 ±0.66 0.05 ±0.48 3.68 ±0.53 4.33 ±1.17 0.23 ±0.09 5.04 ±0.24 1.91 ±0.44 1.38 ±0.38 -25.87 ±0.45 -1.46 ±0.73 -6.55 ±0.23

2.75 ±0.16 2.72 ±0.11 -1.01 ±0.60 4.36 ±0.26 0.79 ±0.16 -0.14 ±0.33 1.30 ±0.65 0.57 ±0.58 2.80 ±0.09 5.42 ±0.69 -1.32 ±0.61 -1.57 ±0.59 1.80 ±0.12 -1.13 ±0.34 0.86 ±0.43 6.09 ±0.50 0.44 ±0.59 2.69 ±0.32 5.32 ±0.37 0.15 ±0.16 8.34 ±0.18 3.64 ±0.57 2.56 ±0.46 -4.58 ±0.90 -0.99 ±0.86 -0.65 ±0.12

3.10 ±0.09 3.14 ±0.19 -1.96 ±0.28 4.92 ±0.31 0.63 ±0.19 -0.45 ±0.25 2.34 ±0.38 1.28 ±0.08 3.96 ±0.36 4.83 ±0.37 -5.40 ±0.73 -2.65 ±0.32 2.67 ±0.28 0.23 ±0.06 1.71 ±0.38 6.95 ±0.24 0.77 ±0.47 3.82 ±0.36 8.07 ±0.40 0.04 ±0.09 9.67 ±0.34 4.96 ±0.36 4.24 ±0.31 -5.46 ±0.73 -0.79 ±0.72 -1.19 ±0.18

4.38 ±0.11 4.80 ±0.32 -1.04 ±0.25 4.44 ±0.15 1.05 ±0.08 0.25 ±0.43 5.04 ±0.48 1.23 ±0.48 7.36 ±0.46 6.06 ±0.14 -3.65 ±0.16 -0.76 ±0.94 4.77 ±0.51 2.77 ±0.11 3.00 ±0.15 9.10 ±0.33 2.07 ±0.60 4.35 ±0.47 10.72 ±0.24 -0.18 ±0.05 9.45 ±0.19 7.43 ±0.60 6.78 ±0.39 -6.31 ±0.34 -0.22 ±0.96 -1.09 ±0.25

6.19 ±0.21 7.79 ±0.43 0.19 ±0.25 5.43 ±0.34 1.28 ±0.30 0.60 ±0.11 8.34 ±0.10 2.04 ±0.36 9.67 ±0.24 7.94 ±0.51 -4.41 ±0.47 0.62 ±0.47 6.29 ±0.28 3.01 ±0.28 3.41 ±0.39 10.05 ±0.23 2.20 ±0.46 5.48 ±0.47 12.21 ±0.28 -0.46 ±0.07 9.43 ±0.31 7.29 ±0.19 7.90 ±0.49 -8.12 ±0.62 -0.42 ±0.82 -2.40 ±0.18

7.69 ±0.20 9.82 ±0.31 -0.59 ±0.22 5.72 ±0.03 0.81 ±0.17 1.21 ±0.25 9.72 ±0.17 3.00 ±0.42 10.82 ±0.36 8.37 ±0.73 -5.29 ±0.44 0.11 ±0.74 7.24 ±0.36 2.59 ±0.39 4.27 ±0.30 11.62 ±0.42 1.81 ±0.52 5.71 ±0.73 13.04 ±0.19 -0.51 ±0.17 9.31 ±0.24 9.40 ±0.12 8.58 ±0.46 -8.55 ±0.64 -0.47 ±0.63 -3.66 ±0.16

8.70 ±0.05 10.99 ±0.34 -1.09 ±0.41 7.58 ±0.19 1.07 ±0.17 2.74 ±0.32 11.74 ±0.41 4.22 ±0.11 12.46 ±0.70 9.07 ±0.75 -5.27 ±0.23 -0.43 ±0.53 9.89 ±0.60 3.56 ±0.36 4.98 ±0.82 12.26 ±0.05 3.15 ±0.73 5.70 ±0.33 13.07 ±0.42 -0.37 ±0.17 9.34 ±0.18 11.52 ±0.73 9.58 ±0.22 -8.77 ±0.81 -0.69 ±0.31 -3.86 ±0.15

9.21 ±0.09 11.79 ±0.36 -0.75 ±0.86 8.98 ±0.33 0.92 ±0.06 3.26 ±0.46 14.05 ±0.23 4.41 ±0.30 14.36 ±0.97 9.57 ±0.92 -5.91 ±1.05 -0.12 ±0.53 13.71 ±0.68 4.78 ±0.23 5.88 ±0.34 13.10 ±0.43 3.40 ±0.75 5.21 ±0.17 13.28 ±0.34 -0.09 ±0.24 9.30 ±0.22 12.34 ±0.36 11.27 ±0.48 -7.41 ±0.83 -1.00 ±0.24 -4.65 ±0.13

9.28 ±0.13 12.57 ±0.10 0.66 ±1.06 9.61 ±0.49 1.59 ±0.24 3.63 ±0.46 16.42 ±0.12 4.85 ±0.74 15.23 ±0.66 9.15 ±0.54 -7.15 ±0.09 0.37 ±1.78 16.43 ±0.71 3.02 ±0.18 6.49 ±0.39 11.87 ±0.34 3.65 ±0.85 4.94 ±0.44 12.79 ±0.49 -0.36 ±0.26 9.02 ±0.32 13.51 ±0.68 11.81 ±0.74 -9.89 ±1.55 -1.58 ±0.67 -5.73 ±0.09

9.45 ±0.33 13.24 ±0.16 2.97 ±0.98 11.53 ±0.16 3.00 ±0.17 6.32 ±0.54 20.23 ±0.25 4.90 ±0.77 13.89 ±0.58 9.32 ±0.41 -9.15 ±0.09 3.92 ±0.75 19.22 ±0.42 2.45 ±0.53 6.46 ±0.38 13.02 ±0.20 3.86 ±0.67 6.14 ±0.12 12.57 ±0.10 -0.49 ±0.11 8.60 ±0.25 15.37 ±0.37 12.23 ±0.60 -10.29 ±1.02 -1.31 ±0.89 -7.09 ±0.27

10.24 ±0.26 14.66 ±0.24 6.54 ±1.41 11.52 ±0.38 2.96 ±0.11 6.56 ±0.73 22.50 ±0.30 4.49 ±0.44 15.65 ±0.44 9.59 ±0.96 -13.08 ±0.91 6.44 ±1.52 21.86 ±0.56 3.03 ±0.23 6.68 ±0.30 14.29 ±0.54 4.32 ±1.25 5.41 ±0.06 13.78 ±0.17 -0.94 ±0.31 8.52 ±0.28 18.07 ±0.47 12.19 ±0.69 -11.48 ±0.34 -0.58 ±1.16 -5.17 ±0.23

11.00 ±0.46 15.31 ±0.19 7.61 ±1.03 15.81 ±0.25 3.88 ±0.11 6.88 ±0.55 22.95 ±0.26 5.23 ±0.34 15.84 ±0.12 10.76 ±0.57 -12.03 ±0.82 6.17 ±1.08 25.17 ±1.39 3.11 ±0.21 7.99 ±0.32 15.42 ±0.40 7.82 ±0.47 9.46 ±0.26 13.85 ±0.63 -0.25 ±0.08 8.29 ±0.16 19.38 ±0.44 12.76 ±0.33 -9.65 ±0.34 0.05 ±1.07 -6.70 ±0.26

0

1

2

3

4

5

6

7

8

9

10

11

12

Layer

hdrzine hdrzone imidazole imide ketone ketone_Topliss lactone methoxy morpholine nitrile nitro nitro_arom nitro_arom_nonortho oxime para_hydroxylation phenol phenol_noOrthoHbond piperdine piperzine priamide pyridine sulfide thiazole thiophene unbrch_alkane urea

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

-1.09 ±0.29 -1.85 ±0.08 -1.11 ±0.42 -0.44 ±0.44 -1.21 ±0.32 -0.60 ±0.28 2.46 ±0.38 4.90 ±0.37 2.69 ±0.52 -11.23 ±1.02 -24.41 ±1.25 -25.39 ±1.08 -21.27 ±0.60 -3.59 ±0.29 1.76 ±0.22 1.65 ±0.19 1.16 ±0.09 2.25 ±0.25 1.05 ±0.18 0.03 ±0.57 1.18 ±0.53 -1.34 ±0.53 -18.88 ±0.43 -28.10 ±0.72 3.58 ±0.46 -1.40 ±0.77

-0.95 ±0.19 -0.26 ±0.39 -0.00 ±0.19 3.54 ±0.09 0.46 ±0.32 0.90 ±0.17 3.35 ±0.50 6.59 ±0.40 4.12 ±0.52 -0.12 ±0.21 -2.46 ±0.44 -3.01 ±1.99 -4.61 ±1.61 0.66 ±0.59 4.49 ±0.27 3.55 ±0.14 3.54 ±0.26 3.50 ±0.59 2.48 ±0.65 2.59 ±0.86 1.80 ±0.51 -0.88 ±0.20 -2.72 ±0.39 -3.79 ±0.76 3.57 ±0.27 -0.14 ±0.97

-1.34 ±0.32 -0.65 ±1.25 -0.73 ±0.35 3.70 ±1.30 0.96 ±0.49 1.41 ±0.38 3.87 ±0.47 9.53 ±0.67 3.89 ±0.93 -0.79 ±0.31 -2.30 ±0.53 -2.57 ±0.32 -2.85 ±0.46 1.00 ±0.39 6.37 ±0.21 4.68 ±0.14 4.42 ±0.09 4.03 ±0.23 2.12 ±0.70 2.96 ±0.20 3.59 ±0.14 -1.40 ±0.44 -3.80 ±0.19 -5.01 ±0.92 4.63 ±0.40 -0.19 ±0.41

-1.00 ±0.63 0.56 ±0.74 -0.65 ±0.29 4.25 ±0.74 3.04 ±0.35 2.82 ±0.35 3.99 ±0.39 17.33 ±0.11 3.43 ±0.96 -0.10 ±0.44 -5.26 ±0.44 -4.37 ±0.33 -4.94 ±2.32 0.95 ±0.25 6.15 ±0.31 5.93 ±0.19 5.51 ±0.54 4.56 ±0.16 3.23 ±0.67 4.90 ±0.18 5.81 ±0.60 -0.93 ±0.28 -4.21 ±0.51 -5.46 ±0.37 6.94 ±0.88 0.34 ±0.40

0.30 ±0.58 0.90 ±0.56 -0.87 ±0.91 4.63 ±0.79 5.27 ±0.24 4.90 ±0.30 3.74 ±0.17 19.13 ±0.12 3.57 ±1.40 -0.65 ±0.35 -5.79 ±1.16 -6.31 ±1.66 -5.17 ±1.74 1.87 ±0.58 7.32 ±0.68 8.91 ±0.22 8.66 ±0.24 4.93 ±0.59 4.09 ±0.72 6.56 ±0.63 5.84 ±0.64 -0.82 ±0.12 -4.61 ±0.73 -5.75 ±0.81 9.87 ±0.36 0.58 ±0.58

2.20 ±0.66 2.06 ±0.98 -1.07 ±0.55 5.16 ±0.39 8.05 ±0.40 7.86 ±0.22 3.72 ±0.21 22.73 ±0.22 4.15 ±0.37 -0.27 ±0.33 -6.26 ±1.68 -5.65 ±0.69 -4.24 ±1.12 2.27 ±0.33 8.61 ±0.78 8.31 ±0.67 8.15 ±0.86 5.17 ±0.36 5.76 ±1.04 7.83 ±0.40 5.40 ±0.42 -0.91 ±0.19 -6.63 ±0.38 -9.44 ±0.96 11.04 ±0.65 0.46 ±0.30

2.79 ±0.82 2.96 ±0.86 0.39 ±0.62 5.90 ±0.31 10.50 ±0.46 10.12 ±0.04 3.84 ±0.29 23.70 ±0.12 4.69 ±1.36 0.81 ±0.13 -5.16 ±1.51 -4.74 ±0.37 -4.07 ±0.95 3.70 ±0.55 9.01 ±0.31 9.02 ±1.17 8.85 ±0.56 5.44 ±0.27 6.23 ±1.60 11.74 ±1.17 7.14 ±0.74 -0.84 ±0.45 -6.04 ±1.38 -10.52 ±1.04 11.58 ±0.69 1.95 ±0.78

2.43 ±0.63 3.52 ±0.90 1.53 ±0.40 8.18 ±0.38 12.19 ±0.41 11.31 ±0.24 4.34 ±0.10 28.52 ±0.35 4.69 ±0.63 2.21 ±0.08 -5.34 ±1.34 -4.18 ±0.23 -3.97 ±1.78 5.37 ±0.61 9.80 ±0.46 10.30 ±0.95 10.04 ±0.93 5.41 ±0.67 5.29 ±1.69 12.79 ±0.83 9.60 ±0.45 -0.81 ±0.38 -6.72 ±0.51 -11.66 ±0.37 13.09 ±0.57 1.43 ±0.59

2.56 ±0.33 3.26 ±0.46 1.92 ±1.06 6.38 ±0.48 14.14 ±0.90 13.53 ±0.50 3.92 ±0.24 28.82 ±0.30 5.06 ±1.27 0.57 ±0.53 -5.91 ±0.29 -6.07 ±0.25 -5.29 ±1.03 6.92 ±1.49 9.89 ±0.27 11.68 ±0.87 12.09 ±0.92 5.33 ±0.54 5.06 ±1.62 15.29 ±0.45 9.47 ±0.51 -1.36 ±0.45 -7.16 ±0.60 -12.64 ±0.90 14.01 ±0.58 0.52 ±1.17

3.19 ±0.42 5.02 ±0.81 2.19 ±0.33 7.60 ±1.15 17.09 ±0.81 16.62 ±0.35 3.76 ±0.61 29.10 ±0.30 6.08 ±0.52 1.26 ±0.43 -5.79 ±0.19 -5.97 ±0.50 -5.70 ±0.98 8.83 ±1.01 9.99 ±0.21 15.05 ±1.02 15.49 ±0.23 5.41 ±0.74 5.75 ±1.41 17.51 ±2.77 10.11 ±0.56 0.47 ±0.05 -6.62 ±0.71 -11.85 ±0.49 13.84 ±1.29 0.94 ±1.12

3.24 ±0.42 5.07 ±0.69 2.73 ±0.30 6.37 ±0.38 19.95 ±0.96 19.03 ±0.34 4.84 ±0.59 30.64 ±0.44 6.44 ±0.15 3.00 ±0.20 -7.44 ±0.93 -6.81 ±0.65 -7.91 ±1.79 9.06 ±1.04 9.20 ±0.49 20.12 ±0.62 20.18 ±0.07 5.41 ±0.61 6.68 ±2.17 19.01 ±0.71 11.80 ±0.49 -0.48 ±0.30 -8.49 ±1.14 -12.23 ±0.61 13.52 ±0.99 0.63 ±0.66

3.38 ±0.95 6.83 ±1.77 3.25 ±0.54 8.98 ±0.53 25.11 ±0.99 23.32 ±1.02 6.32 ±0.54 30.24 ±0.26 4.75 ±0.29 2.16 ±0.51 -5.14 ±1.15 -4.77 ±0.46 -5.11 ±2.26 12.14 ±0.58 10.80 ±0.39 20.96 ±1.07 20.87 ±1.11 5.99 ±0.74 6.21 ±1.86 17.55 ±1.43 15.43 ±0.57 -1.62 ±0.30 -8.91 ±1.37 -13.27 ±1.21 13.65 ±0.97 2.66 ±0.31

0

1

2

3

4

5

6

7

8

9

10

11

12

30

20

10

0

10

Molecular Substructure

Molecular Substructure

NHOH NO AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings HAcceptors HDonors Heteroatoms RotatableBonds SaturatedCarbocycles SaturatedHeterocycles SaturatedRings Ring Al_COO Al_OH Al_OH_noTert ArN Ar_COO Ar_N Ar_NH Ar_OH COO COO2

20

NHOH NO AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings HAcceptors HDonors Heteroatoms RotatableBonds SaturatedCarbocycles SaturatedHeterocycles SaturatedRings Ring Al_COO Al_OH Al_OH_noTert ArN Ar_COO Ar_N Ar_NH Ar_OH COO COO2

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

-0.97 ±0.18 3.93 ±0.17 0.93 ±0.30 0.89 ±0.22 0.45 ±0.18 -0.96 ±0.10 -0.25 ±0.30 0.04 ±0.10 2.42 ±0.19 -0.98 ±0.12 3.35 ±1.23 1.10 ±0.57 0.87 ±0.16 1.01 ±0.22 0.44 ±0.15 -0.02 ±0.01 1.33 ±0.16 1.21 ±0.12 0.91 ±0.16 -1.51 ±0.10 1.48 ±0.16 -1.04 ±0.12 -8.59 ±0.68 -0.49 ±0.13 1.55 ±0.25 1.57 ±0.30

-0.07 ±0.20 -0.90 ±0.39 2.36 ±0.17 2.00 ±0.12 0.66 ±0.17 -1.34 ±0.17 -2.20 ±0.17 -0.02 ±0.11 -0.65 ±0.19 -0.07 ±0.26 0.23 ±0.28 3.11 ±0.45 2.65 ±0.09 1.61 ±0.41 0.62 ±0.26 -0.09 ±0.05 4.27 ±0.55 2.60 ±0.32 2.04 ±0.24 2.90 ±0.90 4.11 ±0.39 -3.25 ±0.36 -7.34 ±0.65 3.43 ±0.44 4.60 ±0.49 4.40 ±0.17

2.16 ±0.16 0.54 ±0.42 3.48 ±0.18 3.21 ±0.12 0.86 ±0.07 -0.90 ±0.09 -0.70 ±0.26 0.04 ±0.12 -0.01 ±0.50 2.33 ±0.19 2.05 ±0.17 3.56 ±0.69 4.30 ±0.12 3.55 ±0.13 1.87 ±0.14 -0.24 ±0.06 8.06 ±0.44 6.43 ±0.13 5.37 ±0.21 8.06 ±0.54 6.38 ±0.69 -1.54 ±0.36 -4.79 ±0.45 8.06 ±0.40 9.21 ±0.48 8.94 ±0.34

1.94 ±0.16 -0.02 ±0.00 5.17 ±0.23 4.41 ±0.27 0.75 ±0.13 -0.18 ±0.07 0.70 ±0.07 0.12 ±0.02 0.49 ±0.54 2.19 ±0.28 3.94 ±0.76 4.96 ±0.60 5.10 ±0.12 5.05 ±0.16 1.96 ±0.06 -0.28 ±0.09 11.84 ±0.22 7.37 ±0.41 6.89 ±0.32 4.96 ±0.21 8.40 ±0.26 -1.11 ±0.28 -7.68 ±0.09 9.83 ±0.75 13.41 ±0.56 13.00 ±0.33

2.16 ±0.17 -1.64 ±0.41 5.94 ±0.05 4.13 ±0.29 0.76 ±0.10 -0.35 ±0.13 0.11 ±0.09 -0.03 ±0.08 -0.01 ±0.76 2.62 ±0.38 1.48 ±0.96 6.54 ±0.47 5.11 ±0.12 4.50 ±0.16 2.18 ±0.30 -0.34 ±0.09 13.99 ±0.31 8.96 ±0.27 7.95 ±0.50 6.80 ±0.37 13.83 ±0.40 -1.34 ±0.13 -6.81 ±0.18 11.19 ±0.26 16.06 ±0.79 15.77 ±0.74

4.92 ±0.30 -1.82 ±0.21 5.52 ±0.25 4.68 ±0.17 0.72 ±0.14 -0.11 ±0.15 -0.05 ±0.08 -0.21 ±0.09 -0.24 ±0.44 5.30 ±0.28 0.59 ±0.18 7.77 ±0.07 5.37 ±0.44 5.50 ±0.22 2.28 ±0.27 -0.34 ±0.05 15.27 ±0.19 8.63 ±0.23 7.56 ±0.23 14.33 ±0.80 16.58 ±0.39 -0.70 ±0.07 -6.14 ±0.16 11.62 ±0.87 16.95 ±0.92 17.04 ±0.51

5.41 ±0.05 -2.12 ±0.25 5.58 ±0.17 4.74 ±0.06 0.83 ±0.10 -0.28 ±0.07 0.22 ±0.21 -0.04 ±0.10 -0.66 ±0.98 5.84 ±0.21 1.53 ±0.22 7.89 ±0.17 6.08 ±0.24 5.58 ±0.17 2.68 ±0.23 -0.30 ±0.05 14.31 ±0.32 9.20 ±0.37 8.62 ±0.29 18.30 ±0.39 12.70 ±0.16 -0.27 ±0.26 -7.15 ±0.04 15.19 ±1.08 15.48 ±0.92 15.30 ±0.67

5.02 ±0.24 -0.25 ±0.56 4.07 ±0.34 3.79 ±0.16 0.67 ±0.08 -0.60 ±0.17 -0.75 ±0.22 -0.20 ±0.02 0.75 ±0.58 5.41 ±0.16 2.76 ±0.50 8.71 ±0.31 4.38 ±0.22 4.26 ±0.10 1.76 ±0.08 -0.38 ±0.08 14.28 ±0.45 9.04 ±0.06 8.35 ±0.30 19.26 ±0.46 11.08 ±0.48 -1.41 ±0.19 -7.86 ±0.44 15.10 ±0.51 14.96 ±0.49 14.95 ±0.76

4.08 ±0.31 -0.47 ±0.92 2.95 ±0.28 3.59 ±0.24 0.35 ±0.09 -0.72 ±0.08 -0.79 ±0.32 -0.39 ±0.04 0.54 ±0.48 4.50 ±0.34 2.72 ±0.81 10.00 ±0.12 3.12 ±0.29 3.98 ±0.37 1.18 ±0.22 -0.62 ±0.04 12.60 ±0.39 8.45 ±0.34 7.59 ±0.42 17.57 ±0.16 9.42 ±0.18 -1.93 ±0.07 -9.63 ±0.73 15.60 ±0.69 13.53 ±0.52 13.47 ±0.43

5.21 4.87 5.17 6.79 6.39 ±0.38 ±0.24 ±0.49 ±0.29 ±0.20 -0.78 1.27 1.90 2.07 1.41 ±0.39 ±0.58 ±0.68 ±0.76 ±1.42 0.84 0.44 0.67 1.59 0.67 ±0.21 ±0.34 ±0.32 ±0.34 ±0.25 2.78 3.04 2.22 2.72 2.02 ±0.16 ±0.16 ±0.14 ±0.18 ±0.09 -0.76 -0.97 -1.05 -0.78 -1.49 ±0.10 ±0.11 ±0.16 ±0.04 ±0.13 -1.37 -2.15 -2.53 -2.98 -3.43 ±0.14 ±0.09 ±0.05 ±0.03 ±0.07 -2.02 -2.94 -3.43 -2.99 -3.24 ±0.21 ±0.22 ±0.17 ±0.03 ±0.25 -1.01 -1.06 -1.09 -1.27 -1.80 ±0.19 ±0.05 ±0.08 ±0.11 ±0.06 0.48 2.36 2.35 2.88 1.71 ±0.69 ±1.02 ±0.36 ±0.41 ±0.77 5.31 5.37 6.02 7.56 6.93 ±0.36 ±0.30 ±0.16 ±0.12 ±0.18 3.36 4.21 4.49 5.40 5.40 ±1.02 ±1.02 ±0.85 ±1.20 ±0.89 9.95 11.53 12.36 12.15 10.70 ±0.32 ±0.35 ±0.62 ±0.44 ±0.16 0.57 0.21 0.58 1.90 0.74 ±0.06 ±0.08 ±0.08 ±0.36 ±0.22 3.05 2.71 2.42 2.77 1.95 ±0.49 ±0.20 ±0.14 ±0.32 ±0.33 -0.06 -0.98 -1.17 -0.97 -1.97 ±0.17 ±0.26 ±0.05 ±0.25 ±0.41 -0.85 -0.95 -1.23 -1.59 -1.96 ±0.16 ±0.06 ±0.10 ±0.29 ±0.21 13.25 16.37 16.93 18.57 18.48 ±0.44 ±0.54 ±0.54 ±0.98 ±0.37 9.27 8.92 9.93 11.29 11.21 ±0.19 ±0.02 ±0.32 ±0.40 ±0.12 8.38 7.91 9.35 10.39 10.40 ±0.23 ±0.16 ±0.74 ±0.30 ±0.40 15.38 15.51 17.08 18.91 19.42 ±0.13 ±0.56 ±0.83 ±1.06 ±0.54 10.50 12.75 13.09 15.77 15.49 ±0.04 ±0.62 ±0.72 ±0.86 ±0.07 -2.74 -3.57 -3.98 -3.52 -3.36 ±0.30 ±0.27 ±0.07 ±0.27 ±0.39 -13.11 -13.07 -13.86 -13.88 -14.46 ±0.42 ±0.87 ±0.03 ±0.60 ±0.76 14.67 12.40 13.28 15.37 14.48 ±0.61 ±0.20 ±1.17 ±1.63 ±0.19 13.89 16.69 17.01 17.99 18.05 ±0.28 ±0.72 ±0.89 ±0.79 ±0.96 13.77 16.41 16.87 18.18 18.23 ±0.38 ±0.76 ±1.10 ±0.88 ±0.50

0

1

2

3

4

5

6

7

8

9

10 11 12 13 14

NHOH NO AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings HAcceptors HDonors Heteroatoms RotatableBonds SaturatedCarbocycles SaturatedHeterocycles SaturatedRings Ring Al_COO Al_OH Al_OH_noTert ArN Ar_COO Ar_N Ar_NH Ar_OH COO COO2

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

1.19 ±0.11 4.34 ±0.47 3.50 ±0.09 2.68 ±0.38 1.49 ±0.26 0.20 ±0.02 1.95 ±0.27 0.50 ±0.07 2.31 ±0.36 1.18 ±0.19 7.60 ±0.82 4.95 ±0.21 2.68 ±0.28 1.83 ±0.08 1.76 ±0.26 -0.04 ±0.07 4.39 ±0.32 2.00 ±0.20 1.80 ±0.19 0.74 ±0.35 4.66 ±0.36 0.71 ±0.27 1.36 ±0.43 2.60 ±0.30 5.17 ±0.23 5.18 ±0.43

1.75 ±0.34 2.67 ±0.42 3.70 ±0.24 3.09 ±0.29 1.11 ±0.02 -0.42 ±0.07 0.91 ±0.21 0.18 ±0.05 2.24 ±0.40 1.95 ±0.23 4.85 ±1.85 5.01 ±0.31 3.31 ±0.18 3.30 ±0.35 2.36 ±0.27 -0.01 ±0.01 4.59 ±0.31 2.22 ±0.10 2.35 ±0.31 5.26 ±0.60 4.97 ±0.31 -0.71 ±0.06 -2.58 ±0.26 4.85 ±0.42 4.57 ±0.18 4.56 ±0.13

0.59 ±0.16 0.04 ±0.19 1.99 ±0.29 1.66 ±0.36 0.80 ±0.19 -0.91 ±0.13 -0.44 ±0.28 -0.06 ±0.10 1.08 ±0.29 0.85 ±0.41 2.20 ±2.28 5.75 ±0.56 2.11 ±0.48 2.24 ±0.32 1.64 ±0.09 -0.01 ±0.03 5.34 ±0.30 3.04 ±0.18 3.04 ±0.53 6.79 ±0.74 5.21 ±0.50 -1.92 ±0.16 -5.79 ±0.22 5.13 ±0.48 5.59 ±0.49 5.83 ±0.50

1.28 ±0.35 -1.14 ±0.82 3.45 ±0.04 2.63 ±0.34 0.93 ±0.20 -1.26 ±0.08 -0.65 ±0.18 -0.00 ±0.10 0.37 ±0.39 1.72 ±0.37 1.73 ±2.27 5.88 ±0.03 3.57 ±0.27 3.39 ±0.16 2.24 ±0.19 -0.04 ±0.04 8.23 ±0.69 2.44 ±0.23 2.35 ±0.10 6.57 ±0.60 8.63 ±0.60 -2.92 ±0.36 -6.67 ±0.78 8.30 ±0.50 8.60 ±0.62 8.96 ±0.52

2.18 ±0.47 -2.62 ±0.55 4.30 ±0.24 4.00 ±0.46 1.12 ±0.10 -1.30 ±0.04 -0.53 ±0.20 0.00 ±0.02 0.13 ±0.27 2.59 ±0.48 2.89 ±2.25 6.29 ±0.14 4.51 ±0.37 4.99 ±0.23 3.08 ±0.07 -0.05 ±0.03 11.75 ±0.48 5.40 ±0.36 4.66 ±0.42 10.41 ±0.71 14.05 ±0.64 -3.01 ±0.13 -5.31 ±0.58 11.44 ±0.32 11.68 ±0.44 11.65 ±0.27

1.14 ±0.25 -2.82 ±0.51 4.50 ±0.12 4.29 ±0.37 1.17 ±0.06 -1.50 ±0.08 -0.87 ±0.05 -0.09 ±0.11 0.01 ±0.79 1.84 ±0.44 2.71 ±2.94 5.46 ±0.53 4.54 ±0.19 5.51 ±0.12 3.01 ±0.12 -0.08 ±0.05 13.15 ±0.70 4.72 ±0.36 4.30 ±0.36 9.49 ±0.22 16.97 ±0.17 -3.84 ±0.02 -5.59 ±0.75 11.39 ±0.76 13.83 ±0.78 14.07 ±0.60

0

1

2

3

4

5

6

C_O C_O_noCOO C_S Imine NH0 NH1 NH2 N_O Ndealkylation1 Ndealkylation2 Nhpyrrole SH aldehyde alkyl_halide allylic_oxid amide amidine aniline aryl_methyl benzene bicyclic ester ether furan guanido halogen

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

1.08 ±0.09 2.06 ±0.18 -2.72 ±0.62 1.94 ±0.29 1.43 ±0.27 1.85 ±0.62 0.24 ±0.42 -1.61 ±0.51 0.12 ±0.35 2.59 ±0.63 1.36 ±0.43 -1.22 ±0.81 2.78 ±0.45 -2.32 ±0.69 0.90 ±0.13 1.48 ±0.13 1.10 ±0.21 1.26 ±0.07 -1.53 ±0.25 0.20 ±0.07 4.19 ±0.21 -0.04 ±0.27 -0.77 ±0.24 -9.78 ±2.28 1.75 ±0.70 -0.29 ±0.23

1.69 ±0.21 2.79 ±0.17 -3.44 ±0.43 2.60 ±0.34 1.02 ±0.20 0.57 ±0.23 4.73 ±0.26 -2.62 ±0.48 1.04 ±0.36 3.32 ±0.14 -2.58 ±0.26 -1.48 ±0.37 2.84 ±0.26 -2.74 ±0.59 1.93 ±0.05 3.83 ±0.38 1.39 ±0.12 3.58 ±0.11 1.08 ±0.38 -0.49 ±0.06 3.84 ±0.17 0.77 ±0.14 -1.59 ±0.23 -16.12 ±0.62 0.39 ±0.91 -1.22 ±0.23

1.84 ±0.15 3.70 ±0.16 -6.38 ±0.68 2.15 ±0.10 0.71 ±0.18 -1.14 ±0.28 5.41 ±0.60 -3.85 ±1.00 1.11 ±0.37 1.59 ±0.78 -5.79 ±0.22 -2.42 ±0.47 3.72 ±0.25 -4.76 ±0.47 1.62 ±0.23 5.21 ±0.66 0.51 ±0.05 2.29 ±0.33 1.55 ±0.08 -0.97 ±0.09 3.10 ±0.24 0.12 ±0.52 -3.90 ±0.04 -21.41 ±0.83 -1.44 ±1.00 -2.69 ±0.15

2.73 ±0.14 5.61 ±0.24 -7.69 ±0.78 3.93 ±0.22 1.23 ±0.19 -2.85 ±0.49 7.48 ±0.39 -4.92 ±1.10 2.01 ±0.21 4.33 ±0.35 -6.67 ±0.78 -3.84 ±0.12 5.62 ±0.49 -5.76 ±0.46 1.76 ±0.39 7.08 ±0.73 0.96 ±0.23 2.22 ±0.34 1.88 ±0.54 -1.28 ±0.05 4.46 ±0.11 -0.14 ±0.74 -5.97 ±0.19 -22.35 ±1.42 -1.86 ±1.26 -4.35 ±0.20

3.76 ±0.14 7.24 ±0.34 -7.00 ±0.58 5.17 ±0.22 1.55 ±0.19 -3.54 ±0.12 9.49 ±0.33 -4.99 ±0.71 1.35 ±0.33 5.05 ±0.44 -5.31 ±0.58 -2.56 ±0.62 8.08 ±1.06 -5.06 ±0.57 2.75 ±0.19 8.23 ±0.51 2.33 ±0.20 2.98 ±0.49 2.00 ±0.07 -1.30 ±0.06 5.06 ±0.16 0.96 ±0.54 -6.45 ±0.15 -21.10 ±1.08 -0.95 ±0.62 -5.03 ±0.35

4.40 ±0.18 8.29 ±0.18 -6.89 ±0.35 6.30 ±0.10 1.43 ±0.10 -4.72 ±0.18 8.18 ±0.49 -4.64 ±1.28 0.27 ±0.26 5.55 ±0.89 -5.59 ±0.75 -3.61 ±0.34 9.59 ±0.75 -5.63 ±0.57 2.82 ±0.13 9.13 ±0.23 1.57 ±0.12 2.52 ±0.82 0.21 ±0.04 -1.50 ±0.01 5.43 ±0.14 1.84 ±0.62 -7.19 ±0.27 -19.68 ±0.96 -1.59 ±0.21 -6.18 ±0.08

0

1

2

3

4

5

6

Layer

hdrzine hdrzone imidazole imide ketone ketone_Topliss lactone methoxy morpholine nitrile nitro nitro_arom nitro_arom_nonortho oxime para_hydroxylation phenol phenol_noOrthoHbond piperdine piperzine priamide pyridine sulfide thiazole thiophene unbrch_alkane urea

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

-2.24 ±0.71 0.75 ±0.32 -3.06 ±0.15 2.39 ±0.46 0.76 ±0.10 0.88 ±0.11 2.37 ±0.46 -6.51 ±0.05 -1.33 ±0.83 -0.26 ±0.37 3.80 ±0.66 3.74 ±1.39 2.24 ±0.96 -0.34 ±0.43 1.10 ±0.22 2.07 ±0.57 2.19 ±0.64 1.64 ±0.13 -1.26 ±1.19 2.27 ±0.59 -3.20 ±0.35 -4.05 ±0.62 -4.81 ±0.53 -10.58 ±1.34 -3.26 ±0.35 0.44 ±0.66

-2.79 ±0.47 0.98 ±0.53 -5.31 ±0.66 2.80 ±0.88 2.65 ±0.26 2.95 ±0.29 3.32 ±0.75 -2.74 ±0.21 -2.63 ±0.32 -0.66 ±1.25 1.31 ±0.87 1.66 ±1.79 2.28 ±0.52 -0.48 ±0.30 1.85 ±0.19 3.59 ±0.45 3.70 ±0.47 1.20 ±0.52 -3.11 ±0.34 2.16 ±0.69 -5.05 ±0.25 -7.25 ±0.42 -8.29 ±0.80 -14.36 ±1.69 -1.47 ±0.73 0.95 ±0.47

-5.08 ±0.40 -1.14 ±1.12 -5.73 ±0.28 2.57 ±0.31 3.03 ±0.31 2.79 ±0.17 1.81 ±0.31 -2.19 ±0.20 -5.03 ±0.65 -1.11 ±0.64 -1.01 ±0.73 -1.29 ±1.56 -3.01 ±0.27 -1.11 ±0.36 1.01 ±0.29 4.03 ±0.55 4.10 ±0.40 -0.90 ±0.41 -5.38 ±1.57 1.59 ±0.78 -8.51 ±0.26 -9.44 ±0.65 -9.56 ±1.20 -18.55 ±1.27 0.09 ±0.49 0.46 ±0.49

-7.36 ±1.40 -0.83 ±0.49 -5.92 ±0.19 4.63 ±0.59 3.99 ±0.13 3.85 ±0.21 2.27 ±0.70 -5.55 ±0.07 -4.60 ±0.42 -1.67 ±0.52 -3.24 ±1.00 -3.39 ±0.69 -5.17 ±0.91 -0.79 ±0.87 0.91 ±0.19 6.16 ±0.44 6.10 ±0.60 -0.03 ±0.25 -4.64 ±1.60 0.21 ±1.00 -9.35 ±0.75 -11.43 ±0.33 -9.67 ±1.14 -19.35 ±0.39 -1.86 ±0.07 1.36 ±0.73

-8.25 ±0.85 -0.36 ±0.60 -4.69 ±0.27 5.67 ±0.94 6.50 ±0.16 6.19 ±0.31 3.85 ±0.38 -8.16 ±0.24 -4.34 ±0.18 -1.02 ±0.42 -3.98 ±1.39 -5.09 ±1.24 -6.95 ±2.41 -0.44 ±0.54 1.23 ±0.35 9.40 ±0.63 9.19 ±0.25 1.19 ±0.13 -4.52 ±1.22 -0.06 ±0.99 -8.74 ±0.70 -12.69 ±0.66 -8.98 ±0.94 -18.73 ±0.65 -4.53 ±0.66 2.06 ±0.90

-8.03 ±0.97 -0.05 ±0.46 -4.83 ±0.67 7.35 ±0.22 8.02 ±0.55 7.81 ±0.55 3.90 ±0.53 -11.58 ±0.36 -4.22 ±0.30 0.62 ±0.47 -2.51 ±1.83 -3.60 ±0.80 -6.22 ±2.37 0.27 ±0.12 1.29 ±0.57 8.95 ±0.48 9.53 ±0.31 0.75 ±0.30 -2.61 ±1.10 -0.36 ±0.81 -7.18 ±0.88 -13.25 ±0.98 -9.01 ±0.63 -18.21 ±0.15 -7.23 ±0.59 2.92 ±0.94

0

1

2

3

4

5

6

15 10 5 0 5 10 15 20

NHOH NO AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings HAcceptors HDonors Heteroatoms RotatableBonds SaturatedCarbocycles SaturatedHeterocycles SaturatedRings Ring Al_COO Al_OH Al_OH_noTert ArN Ar_COO Ar_N Ar_NH Ar_OH COO COO2

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

-2.50 ±0.16 -2.10 ±0.34 -0.17 ±0.09 -0.40 ±0.29 -0.82 ±0.22 -0.83 ±0.17 -2.39 ±0.39 -0.24 ±0.08 -1.77 ±0.62 -2.50 ±0.19 -1.93 ±0.39 0.54 ±0.53 0.70 ±0.41 0.16 ±0.62 0.04 ±0.12 -0.02 ±0.08 2.56 ±0.29 0.39 ±0.12 0.04 ±0.18 -0.66 ±0.12 3.10 ±0.17 -2.53 ±0.13 -2.31 ±0.76 2.14 ±0.46 2.69 ±0.36 2.69 ±0.12

-4.81 ±0.15 -4.18 ±1.08 0.04 ±0.24 -0.86 ±0.14 -0.51 ±0.17 -1.20 ±0.11 -2.71 ±0.15 -0.20 ±0.09 -3.06 ±0.11 -4.65 ±0.31 -1.35 ±1.15 0.89 ±0.21 -0.47 ±0.29 -0.68 ±0.11 -0.17 ±0.32 -0.00 ±0.02 3.58 ±0.20 -1.39 ±0.15 -1.75 ±0.06 -0.59 ±0.63 4.01 ±0.78 -5.51 ±0.35 -4.20 ±0.85 1.22 ±0.69 3.96 ±0.38 3.77 ±0.04

-4.38 ±0.34 -5.46 ±0.41 1.93 ±0.33 0.27 ±0.36 0.51 ±0.14 -1.69 ±0.12 -2.61 ±0.03 -0.17 ±0.12 -2.87 ±0.89 -3.76 ±0.28 -2.01 ±1.04 1.80 ±0.75 1.64 ±0.14 0.76 ±0.11 0.72 ±0.06 -0.04 ±0.13 6.63 ±0.65 0.08 ±0.16 -0.67 ±0.49 -0.22 ±1.19 8.64 ±0.61 -5.78 ±0.16 -6.22 ±0.23 4.30 ±1.01 6.88 ±0.85 7.26 ±0.33

-4.11 ±0.26 -6.71 ±0.35 3.00 ±0.41 1.21 ±0.16 0.70 ±0.09 -2.04 ±0.15 -2.91 ±0.12 -0.29 ±0.02 -3.56 ±1.06 -3.42 ±0.07 -2.12 ±0.94 2.51 ±0.42 3.33 ±0.33 1.93 ±0.07 1.51 ±0.19 -0.03 ±0.05 9.43 ±1.07 0.47 ±0.40 -0.73 ±0.59 0.95 ±1.23 12.94 ±0.97 -6.38 ±0.18 -6.53 ±0.78 5.72 ±0.17 10.08 ±0.94 10.16 ±0.61

-2.85 ±0.35 -7.94 ±0.62 4.21 ±0.23 2.68 ±0.04 0.90 ±0.12 -2.23 ±0.10 -1.58 ±0.16 -0.23 ±0.02 -4.28 ±0.86 -1.96 ±0.39 -3.15 ±1.35 1.37 ±0.60 4.81 ±0.41 3.53 ±0.40 1.92 ±0.29 -0.11 ±0.09 11.12 ±0.70 1.72 ±0.20 0.54 ±0.12 0.65 ±0.74 16.04 ±0.75 -5.56 ±0.30 -5.75 ±0.93 9.00 ±0.89 11.96 ±0.86 12.49 ±0.60

-4.15 ±0.60 -9.74 ±0.72 3.16 ±0.43 2.70 ±0.23 0.66 ±0.12 -1.94 ±0.07 -1.94 ±0.18 -0.32 ±0.06 -5.73 ±0.76 -3.12 ±0.43 -4.25 ±1.08 0.25 ±0.18 4.29 ±0.22 3.79 ±0.36 1.83 ±0.08 -0.20 ±0.09 10.05 ±0.60 1.41 ±0.38 0.40 ±0.55 -2.25 ±0.81 15.47 ±0.37 -5.62 ±0.18 -5.17 ±0.63 6.96 ±1.05 11.34 ±0.67 11.09 ±0.85

0

1

2

3

4

5

6

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

3.30 ±0.33 -0.59 ±0.73 0.45 ±0.45 0.37 ±0.19 0.06 ±0.02 -0.95 ±0.11 -0.83 ±0.15 -0.12 ±0.06 -0.24 ±0.43 3.22 ±0.09 2.69 ±1.79 5.56 ±0.91 0.66 ±0.34 1.84 ±0.53 2.15 ±0.25 0.23 ±0.13 11.15 ±0.51 6.05 ±0.42 5.49 ±0.28 11.45 ±0.67 8.45 ±0.11 -2.87 ±0.19 -4.27 ±0.21 4.63 ±0.11 11.40 ±0.21 11.27 ±0.16

5.78 ±0.13 -0.50 ±0.42 3.75 ±0.42 3.02 ±0.19 -0.06 ±0.12 -2.08 ±0.04 -4.54 ±0.22 -0.19 ±0.02 1.04 ±0.76 6.07 ±0.13 3.30 ±0.77 8.41 ±0.68 5.44 ±0.30 5.80 ±0.34 3.90 ±0.30 0.15 ±0.06 17.34 ±0.50 8.93 ±0.77 8.82 ±0.79 16.10 ±0.29 11.18 ±0.36 -6.64 ±0.06 -4.36 ±0.54 8.17 ±0.12 16.83 ±0.45 16.72 ±0.53

6.31 ±0.26 1.01 ±1.01 3.71 ±0.14 3.72 ±0.15 -0.07 ±0.06 -2.60 ±0.06 -5.48 ±0.09 -0.28 ±0.06 2.68 ±0.60 6.62 ±0.30 5.72 ±1.40 8.99 ±0.56 6.06 ±0.15 8.22 ±0.34 4.84 ±0.18 0.17 ±0.11 21.83 ±0.40 16.08 ±0.14 15.46 ±0.22 14.62 ±0.70 15.99 ±0.38 -7.33 ±0.33 -2.60 ±0.73 9.50 ±0.12 21.19 ±0.42 21.13 ±0.21

0

1

2

3

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

2.90 ±0.27 -2.02 ±0.34 -0.81 ±0.17 0.47 ±0.11 0.31 ±0.00 -1.71 ±0.12 -1.52 ±0.10 -0.14 ±0.02 -0.31 ±0.55 2.66 ±0.16 -0.21 ±0.57 9.45 ±0.56 -0.18 ±0.25 2.58 ±0.17 2.49 ±0.03 0.41 ±0.10 8.59 ±0.46 8.78 ±0.37 8.03 ±0.25 5.15 ±0.20 5.78 ±0.63 -4.72 ±0.11 -3.38 ±0.49 2.88 ±0.20 8.95 ±0.15 8.72 ±0.40

5.87 ±0.32 0.15 ±1.09 1.79 ±0.30 -1.53 ±0.30 -3.20 ±0.15 -5.54 ±0.15 -10.98 ±0.39 -1.81 ±0.07 0.24 ±0.86 6.17 ±0.17 1.75 ±1.50 9.01 ±0.36 2.53 ±0.26 2.91 ±0.53 0.09 ±0.48 -0.17 ±0.05 14.56 ±0.26 12.48 ±0.28 12.33 ±0.14 18.87 ±0.18 10.51 ±0.71 -11.01 ±0.37 -3.28 ±0.69 9.65 ±0.20 14.51 ±0.38 14.21 ±0.14

7.72 ±0.47 -4.42 ±0.56 -0.92 ±0.15 -1.51 ±0.38 -5.55 ±0.08 -6.74 ±0.16 -11.50 ±0.27 -3.41 ±0.13 1.32 ±0.62 8.17 ±0.28 0.97 ±0.32 7.00 ±0.32 1.00 ±0.07 2.67 ±0.24 -1.54 ±0.16 -1.45 ±0.02 29.78 ±0.09 15.93 ±0.22 15.41 ±0.37 16.08 ±0.35 25.01 ±0.70 -12.36 ±0.22 -0.15 ±0.30 16.99 ±0.66 32.36 ±0.58 32.54 ±0.70

0

1

2

3

C_O C_O_noCOO C_S Imine NH0 NH1 NH2 N_O Ndealkylation1 Ndealkylation2 Nhpyrrole SH aldehyde alkyl_halide allylic_oxid amide amidine aniline aryl_methyl benzene bicyclic ester ether furan guanido halogen

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

0

7.28 ±0.24 8.58 ±0.22 -6.84 ±1.13 2.93 ±0.18 4.46 ±0.22 5.52 ±0.09 13.60 ±0.25 4.01 ±0.60 15.17 ±0.54 12.90 ±0.19 -4.36 ±0.54 -8.54 ±1.13 10.44 ±0.44 2.51 ±0.33 3.88 ±0.07 11.57 ±0.32 1.02 ±0.38 8.94 ±0.24 15.02 ±0.09 -2.12 ±0.09 10.62 ±0.38 12.81 ±0.50 11.87 ±0.56 -25.14 ±2.01 -1.35 ±0.28 5.38 ±0.36

9.53 ±0.18 10.94 ±0.19 -2.95 ±1.52 8.93 ±0.60 3.87 ±0.16 5.75 ±0.18 14.13 ±0.05 5.35 ±0.87 15.83 ±0.47 13.88 ±0.84 -2.60 ±0.73 -5.37 ±0.19 15.43 ±0.45 5.04 ±0.06 6.56 ±0.15 14.95 ±0.37 2.29 ±0.29 8.15 ±0.52 12.78 ±0.03 -2.60 ±0.17 11.58 ±0.44 15.02 ±0.65 10.94 ±0.16 -25.86 ±1.60 -1.25 ±0.26 5.06 ±0.32

1

2

3

Layer

hdrzine hdrzone imidazole imide ketone ketone_Topliss lactone methoxy morpholine nitrile nitro nitro_arom nitro_arom_nonortho oxime para_hydroxylation phenol phenol_noOrthoHbond piperdine piperzine priamide pyridine sulfide thiazole thiophene unbrch_alkane urea

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

1.32 ±0.42 -0.09 ±0.56 -4.04 ±0.34 2.10 ±0.06 1.64 ±0.52 2.69 ±0.28 2.92 ±0.09 22.29 ±0.63 0.88 ±0.45 4.88 ±0.85 18.33 ±0.72 16.58 ±1.12 15.74 ±0.54 0.04 ±0.60 1.12 ±0.12 2.78 ±0.28 2.63 ±0.21 1.76 ±0.04 1.36 ±0.43 10.11 ±0.47 -4.99 ±0.27 -2.59 ±0.39 -2.17 ±0.83 -7.41 ±0.25 9.99 ±0.81 2.73 ±0.33

0.75 ±0.69 -0.51 ±0.96 -4.75 ±0.36 7.59 ±1.09 6.42 ±0.61 7.56 ±0.31 3.63 ±0.15 27.62 ±0.64 3.95 ±0.52 2.95 ±0.61 14.24 ±1.41 16.18 ±0.98 15.30 ±1.07 0.24 ±0.10 6.90 ±0.18 5.49 ±0.24 5.27 ±0.28 7.21 ±0.76 5.51 ±1.43 11.66 ±0.51 -6.06 ±0.40 1.74 ±0.51 -11.12 ±0.98 -23.63 ±0.47 10.27 ±0.46 2.77 ±0.78

-0.58 ±0.83 1.45 ±0.95 -3.74 ±0.11 8.23 ±1.22 11.89 ±0.30 12.78 ±0.41 5.27 ±0.57 26.68 ±0.78 4.62 ±0.35 2.68 ±0.62 16.82 ±0.38 16.41 ±1.66 16.09 ±1.49 2.55 ±0.25 7.35 ±0.29 7.15 ±0.50 7.08 ±0.48 7.81 ±0.40 6.80 ±0.98 15.72 ±1.17 -3.51 ±0.41 0.69 ±0.11 -11.60 ±1.27 -23.31 ±0.68 12.19 ±0.64 3.89 ±0.96

0

1

2

3

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

-2.59 ±0.12 -1.42 ±0.51 -4.46 ±0.22 0.80 ±0.52 4.08 ±0.24 4.62 ±0.52 2.16 ±0.19 24.79 ±0.68 -0.89 ±0.59 -12.29 ±0.26 1.41 ±0.59 1.67 ±0.60 3.39 ±0.71 -1.03 ±0.64 0.82 ±0.51 1.05 ±0.08 0.95 ±0.17 2.17 ±0.06 0.50 ±0.61 8.70 ±0.60 -6.09 ±0.39 -1.91 ±0.73 -3.97 ±1.20 -10.15 ±0.44 12.76 ±0.73 1.40 ±0.26

-0.74 ±0.44 -0.49 ±0.20 -5.15 ±0.72 7.55 ±0.88 8.61 ±0.54 9.46 ±0.10 2.07 ±0.51 28.15 ±0.40 1.55 ±0.45 -10.28 ±0.84 10.40 ±1.04 12.32 ±1.23 9.84 ±1.97 2.48 ±0.78 4.69 ±0.35 6.61 ±0.18 6.59 ±0.59 4.34 ±0.66 2.37 ±1.02 13.86 ±0.73 -4.72 ±0.55 -6.68 ±0.48 -11.54 ±1.54 -23.91 ±0.88 10.90 ±0.58 3.19 ±0.63

-2.87 ±0.18 0.90 ±0.36 -6.03 ±0.77 8.45 ±1.13 12.36 ±0.59 12.76 ±0.30 2.43 ±0.47 26.22 ±0.79 0.14 ±0.81 -13.24 ±1.08 -0.67 ±0.84 2.22 ±0.52 2.14 ±1.83 4.32 ±0.61 2.47 ±0.42 12.72 ±0.25 12.17 ±0.63 4.83 ±0.85 3.48 ±0.83 12.16 ±1.58 2.99 ±0.58 -8.92 ±0.27 -11.11 ±1.23 -19.83 ±0.76 9.07 ±1.35 1.79 ±0.26

0

1

2

3

30 20 10 0 10 20

NHOH NO AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings HAcceptors HDonors Heteroatoms RotatableBonds SaturatedCarbocycles SaturatedHeterocycles SaturatedRings Ring Al_COO Al_OH Al_OH_noTert ArN Ar_COO Ar_N Ar_NH Ar_OH COO COO2

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

2.86 ±0.39 9.05 ±0.92 0.98 ±0.44 3.41 ±0.08 0.74 ±0.08 0.46 ±0.08 1.88 ±0.12 -0.03 ±0.04 2.55 ±0.55 2.86 ±0.32 6.91 ±2.26 5.42 ±0.20 1.18 ±0.13 3.74 ±0.30 2.47 ±0.01 0.37 ±0.11 8.71 ±0.08 4.07 ±0.10 3.42 ±0.25 9.64 ±0.48 8.31 ±0.07 0.16 ±0.06 0.47 ±0.35 5.49 ±0.47 9.06 ±0.19 8.56 ±0.10

16.29 ±0.39 11.65 ±1.32 7.49 ±0.32 4.10 ±0.16 0.07 ±0.21 0.83 ±0.18 0.33 ±0.09 -0.17 ±0.01 7.00 ±0.94 16.53 ±0.26 11.87 ±2.12 10.43 ±0.29 8.51 ±0.17 9.09 ±0.47 5.28 ±0.19 0.02 ±0.04 31.47 ±0.49 23.62 ±0.55 23.24 ±0.41 38.01 ±0.77 23.70 ±1.18 0.08 ±0.10 3.98 ±0.20 37.83 ±0.67 32.94 ±0.17 32.57 ±0.33

14.44 ±0.17 5.37 ±0.93 6.21 ±0.55 3.64 ±0.18 -0.44 ±0.11 -0.51 ±0.09 -1.02 ±0.11 -0.77 ±0.14 4.73 ±0.97 14.81 ±0.13 8.55 ±0.97 9.45 ±0.49 8.26 ±0.04 7.32 ±0.32 4.07 ±0.40 -0.03 ±0.09 31.40 ±0.23 19.96 ±0.07 19.73 ±0.15 33.71 ±0.62 23.85 ±0.82 -0.89 ±0.11 7.54 ±0.75 31.45 ±0.59 31.79 ±0.21 31.73 ±0.27

Molecular Substructure

C_O C_O_noCOO C_S Imine NH0 NH1 NH2 N_O Ndealkylation1 Ndealkylation2 Nhpyrrole SH aldehyde alkyl_halide allylic_oxid amide amidine aniline aryl_methyl benzene bicyclic ester ether furan guanido halogen

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

0

5.70 ±0.09 7.76 ±0.34 -5.69 ±1.06 -0.04 ±0.11 1.95 ±0.19 1.76 ±0.21 6.01 ±0.42 1.15 ±0.44 11.77 ±0.18 3.73 ±0.25 -3.38 ±0.49 -4.83 ±0.13 5.78 ±0.27 0.51 ±0.26 2.70 ±0.23 5.37 ±0.25 -1.23 ±0.14 3.87 ±0.34 12.02 ±0.38 -1.78 ±0.06 6.11 ±0.22 10.17 ±0.41 8.46 ±0.47 -21.72 ±2.08 -2.98 ±0.57 -0.33 ±0.37

6.76 ±0.12 9.51 ±0.38 -3.30 ±0.85 5.56 ±0.26 1.64 ±0.19 3.75 ±0.16 13.63 ±0.33 3.17 ±0.50 12.90 ±0.50 9.15 ±1.14 -3.28 ±0.69 -8.77 ±0.44 12.51 ±0.66 -3.35 ±0.29 2.89 ±0.16 13.07 ±0.31 -0.37 ±0.53 7.14 ±0.20 12.69 ±0.16 -5.59 ±0.08 9.17 ±0.40 10.90 ±0.28 7.75 ±0.49 -27.74 ±1.91 -3.12 ±0.40 -5.33 ±0.65

6.81 ±0.37 10.84 ±0.32 -4.74 ±0.81 8.87 ±0.08 2.88 ±0.25 3.16 ±0.12 10.83 ±0.09 4.53 ±1.11 14.13 ±0.46 9.40 ±0.74 -0.15 ±0.30 -8.30 ±0.20 16.42 ±0.08 -2.20 ±0.18 2.39 ±0.59 14.22 ±0.61 -1.31 ±0.14 6.67 ±0.35 9.79 ±0.21 -6.73 ±0.09 8.00 ±0.30 14.61 ±0.57 6.06 ±0.19 -22.55 ±2.17 -4.57 ±0.55 -3.41 ±0.64

1

2

3

Layer

0

1

C_O C_O_noCOO C_S Imine NH0 NH1 NH2 N_O Ndealkylation1 Ndealkylation2 Nhpyrrole SH aldehyde alkyl_halide allylic_oxid amide amidine aniline aryl_methyl benzene bicyclic ester ether furan guanido halogen

2

3

4

5

6

Layer

7

8

9

hdrzine hdrzone imidazole imide ketone ketone_Topliss lactone methoxy morpholine nitrile nitro nitro_arom nitro_arom_nonortho oxime para_hydroxylation phenol phenol_noOrthoHbond piperdine piperzine priamide pyridine sulfide thiazole thiophene unbrch_alkane urea

10 11 12 13 14

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

0.04 ±0.15 1.50 ±0.43 -2.41 ±1.57 0.20 ±0.22 -0.92 ±0.24 -1.99 ±0.24 -2.11 ±0.22 -4.17 ±0.47 -1.37 ±0.20 0.10 ±0.40 -2.31 ±0.76 -2.33 ±0.60 1.48 ±0.25 -4.78 ±0.60 -0.29 ±0.21 0.75 ±0.31 -1.47 ±0.17 -1.60 ±0.05 -0.55 ±0.13 -0.82 ±0.18 1.31 ±0.05 -1.41 ±0.42 -5.54 ±0.41 -16.40 ±1.46 -2.10 ±0.57 -3.36 ±0.12

-0.05 ±0.21 1.83 ±0.30 -6.96 ±0.98 -0.20 ±0.13 -3.12 ±0.07 -7.74 ±0.22 -3.09 ±0.27 -7.63 ±0.66 -2.47 ±0.47 -0.93 ±0.76 -4.20 ±0.85 -5.95 ±0.97 1.99 ±0.77 -10.70 ±0.72 -0.51 ±0.52 0.77 ±0.38 -2.54 ±0.22 -2.93 ±0.23 -1.90 ±0.14 -1.14 ±0.05 1.72 ±0.05 -2.31 ±0.17 -10.92 ±0.12 -20.50 ±1.43 -2.36 ±1.08 -6.38 ±0.14

1.52 ±0.08 4.06 ±0.16 -8.50 ±1.00 0.78 ±0.13 -2.56 ±0.22 -10.21 ±0.17 -3.56 ±0.49 -9.15 ±0.65 -3.02 ±0.08 0.59 ±0.48 -6.22 ±0.23 -10.43 ±1.33 3.30 ±0.32 -13.00 ±0.63 0.15 ±0.12 2.81 ±0.06 -1.91 ±0.56 -4.13 ±0.71 -2.11 ±0.16 -1.74 ±0.09 3.65 ±0.10 -2.91 ±0.22 -12.48 ±0.41 -22.55 ±1.32 -3.37 ±0.26 -9.25 ±0.33

1.84 ±0.10 5.27 ±0.12 -9.80 ±0.73 2.29 ±0.09 -2.07 ±0.25 -10.06 ±0.26 -2.39 ±0.45 -10.03 ±0.43 -3.35 ±0.11 1.57 ±0.57 -6.53 ±0.78 -12.06 ±0.90 5.78 ±0.65 -12.54 ±0.33 0.97 ±0.35 3.83 ±0.60 -2.03 ±0.76 -3.24 ±0.26 -1.21 ±0.11 -2.06 ±0.11 4.56 ±0.17 -3.44 ±0.16 -14.24 ±0.63 -25.18 ±1.29 -3.60 ±0.19 -11.12 ±0.37

2.31 ±0.01 6.16 ±0.16 -8.20 ±1.10 3.84 ±0.18 -1.74 ±0.28 -10.21 ±0.22 -2.85 ±0.01 -10.33 ±0.81 -3.84 ±0.17 1.70 ±0.51 -5.75 ±0.93 -13.11 ±1.29 5.32 ±0.78 -13.18 ±0.80 0.47 ±0.40 4.42 ±0.14 -1.82 ±0.37 -3.28 ±0.18 -2.25 ±0.20 -2.25 ±0.12 4.91 ±0.21 -3.08 ±0.32 -14.43 ±0.27 -25.20 ±0.14 -4.05 ±0.57 -12.06 ±0.21

1.51 ±0.13 5.92 ±0.16 -9.29 ±1.84 3.78 ±0.25 -2.33 ±0.20 -11.37 ±0.14 -4.95 ±0.49 -11.09 ±0.53 -5.17 ±0.49 1.54 ±0.37 -5.17 ±0.63 -14.68 ±1.33 3.12 ±1.05 -15.25 ±0.69 -0.31 ±0.66 4.14 ±0.06 -1.27 ±0.56 -4.57 ±0.50 -4.03 ±0.05 -1.96 ±0.13 4.64 ±0.39 -3.18 ±0.22 -15.76 ±0.11 -27.03 ±0.69 -5.42 ±0.18 -12.78 ±0.26

0

1

2

3

4

5

6

0

1

2

3

C_O C_O_noCOO C_S Imine NH0 NH1 NH2 N_O Ndealkylation1 Ndealkylation2 Nhpyrrole SH aldehyde alkyl_halide allylic_oxid amide amidine aniline aryl_methyl benzene bicyclic ester ether furan guanido halogen

Layer

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

3.48 ±0.19 5.03 ±0.16 -0.37 ±0.23 2.68 ±0.39 3.02 ±0.49 0.63 ±0.41 4.24 ±0.24 1.30 ±0.40 7.92 ±0.33 5.84 ±0.36 0.47 ±0.35 0.20 ±0.56 2.34 ±0.15 2.03 ±0.33 -1.51 ±0.49 7.07 ±0.12 1.70 ±0.44 6.43 ±0.46 10.00 ±0.22 0.37 ±0.05 6.90 ±0.30 7.19 ±0.33 5.76 ±0.43 9.90 ±0.77 3.14 ±0.49 4.67 ±0.67

11.90 ±0.17 15.89 ±0.37 5.29 ±1.27 13.59 ±0.20 8.76 ±0.50 12.32 ±0.13 22.88 ±0.32 9.67 ±0.44 16.65 ±0.53 15.33 ±0.97 3.98 ±0.20 8.39 ±1.30 31.47 ±0.44 2.83 ±0.51 10.01 ±0.59 19.46 ±0.43 4.26 ±0.56 16.50 ±0.28 15.47 ±0.17 0.80 ±0.13 11.70 ±0.24 19.44 ±0.55 14.88 ±0.66 -8.05 ±1.45 -0.06 ±0.90 6.60 ±0.56

10.84 ±0.25 15.14 ±0.24 1.02 ±0.41 13.65 ±0.10 7.22 ±0.59 11.03 ±0.37 19.92 ±0.44 7.94 ±1.15 18.98 ±0.73 14.44 ±0.93 7.54 ±0.75 3.93 ±1.11 25.51 ±0.22 2.99 ±0.46 8.94 ±0.53 20.71 ±0.28 3.59 ±0.86 12.22 ±0.28 12.27 ±0.29 -0.48 ±0.12 11.61 ±0.20 19.35 ±0.61 12.56 ±0.34 -13.30 ±1.75 -1.35 ±0.43 3.66 ±0.49

1

2

3

0

(e) Chemberta-2-5M NHOH NO AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings HAcceptors HDonors Heteroatoms RotatableBonds SaturatedCarbocycles SaturatedHeterocycles SaturatedRings Ring Al_COO Al_OH Al_OH_noTert ArN Ar_COO Ar_N Ar_NH Ar_OH COO COO2

0.37 1.91 3.60 4.40 4.41 5.06 4.70 3.94 3.67 4.34 3.87 4.00 4.56 4.48 ±0.03 ±0.10 ±0.14 ±0.11 ±0.29 ±0.14 ±0.03 ±0.21 ±0.09 ±0.14 ±0.10 ±0.24 ±0.21 ±0.20 1.80 3.36 7.89 9.31 9.19 9.92 9.29 8.75 8.40 8.31 8.19 8.59 9.16 9.34 ±0.10 ±0.23 ±0.17 ±0.31 ±0.32 ±0.52 ±0.36 ±0.24 ±0.19 ±0.15 ±0.55 ±0.46 ±0.47 ±0.33 -4.85 -10.18 -5.35 -7.98 -8.81 -4.96 -1.99 -3.96 -6.40 -8.24 -6.38 -3.06 -2.71 -5.97 ±1.67 ±0.89 ±1.01 ±0.71 ±0.36 ±0.15 ±0.23 ±0.66 ±0.85 ±0.43 ±1.04 ±0.65 ±0.64 ±1.58 -0.65 0.23 2.86 3.16 3.42 4.58 4.90 4.75 4.63 5.08 5.61 6.23 6.98 7.98 ±0.23 ±0.32 ±0.20 ±0.32 ±0.44 ±0.31 ±0.24 ±0.19 ±0.17 ±0.39 ±0.30 ±0.38 ±0.41 ±0.62 0.18 -0.57 0.77 1.26 1.45 1.33 1.58 1.35 0.83 0.55 0.05 0.46 1.33 1.27 ±0.21 ±0.22 ±0.10 ±0.23 ±0.19 ±0.13 ±0.41 ±0.23 ±0.26 ±0.08 ±0.25 ±0.12 ±0.21 ±0.32 -1.59 -4.41 -1.26 -0.86 -0.04 1.26 1.58 1.71 1.39 0.23 -0.33 -0.14 0.75 0.05 ±0.20 ±0.45 ±0.20 ±0.17 ±0.52 ±0.35 ±0.19 ±0.05 ±0.32 ±0.45 ±0.47 ±0.16 ±0.12 ±0.28 0.83 5.06 6.73 7.42 9.62 14.04 16.17 15.71 14.80 12.81 15.07 17.28 17.45 16.67 ±0.22 ±0.50 ±0.37 ±0.15 ±0.66 ±0.47 ±0.48 ±0.17 ±0.37 ±0.37 ±0.05 ±0.15 ±0.37 ±0.50 -4.84 -8.02 -5.31 -4.30 -3.62 -4.09 -2.09 -1.52 -2.00 -1.76 -1.11 -1.04 -0.74 -1.81 ±0.64 ±0.51 ±0.57 ±0.37 ±0.30 ±0.47 ±0.40 ±0.54 ±0.20 ±0.97 ±0.39 ±0.59 ±0.55 ±1.09 -0.34 2.79 2.23 3.32 3.14 3.87 4.29 6.44 6.91 7.22 6.66 8.43 7.52 6.93 ±0.65 ±0.72 ±0.88 ±0.46 ±0.81 ±0.29 ±0.19 ±0.23 ±0.85 ±0.26 ±0.32 ±0.54 ±0.22 ±0.11 -0.05 -0.29 2.94 4.69 5.44 6.47 6.66 6.17 6.33 6.55 6.22 5.31 4.90 3.90 ±0.29 ±0.76 ±0.49 ±0.68 ±0.13 ±0.38 ±1.05 ±0.48 ±0.72 ±0.19 ±0.91 ±0.83 ±0.44 ±0.81 -8.59 -7.34 -4.79 -7.68 -6.81 -6.14 -7.15 -7.86 -9.63 -13.11 -13.07 -13.86 -13.88 -14.46 ±0.68 ±0.65 ±0.45 ±0.09 ±0.18 ±0.16 ±0.04 ±0.44 ±0.73 ±0.42 ±0.87 ±0.03 ±0.60 ±0.76 -1.95 -10.60 -6.66 -3.82 -6.46 -3.59 2.03 2.52 2.02 -1.47 -1.33 0.03 -0.38 -3.29 ±1.49 ±0.45 ±0.75 ±0.81 ±0.46 ±0.63 ±0.49 ±0.60 ±0.90 ±0.16 ±0.25 ±1.44 ±1.11 ±0.65 1.00 2.34 4.90 5.63 6.61 7.01 8.93 9.16 9.69 9.80 11.90 12.86 15.83 19.93 ±0.24 ±0.52 ±0.21 ±0.36 ±0.56 ±0.65 ±0.44 ±0.24 ±0.28 ±0.92 ±0.35 ±0.82 ±1.02 ±1.04 -2.88 -3.18 -5.57 -5.39 -3.61 -0.38 0.39 -1.74 -3.39 -5.34 -4.62 -5.16 -5.36 -5.78 ±0.38 ±0.68 ±0.95 ±0.81 ±1.08 ±0.68 ±0.35 ±0.56 ±0.59 ±0.57 ±0.99 ±1.02 ±0.57 ±0.95 -0.45 0.25 1.45 1.53 2.79 1.97 2.90 3.14 2.55 2.14 2.75 2.67 3.40 3.09 ±0.25 ±0.21 ±0.11 ±0.26 ±0.17 ±0.30 ±0.10 ±0.07 ±0.29 ±0.57 ±0.36 ±0.14 ±0.47 ±0.22 0.30 3.21 8.48 11.16 11.80 12.54 12.12 11.78 11.15 11.66 11.78 12.14 13.72 13.83 ±0.19 ±0.61 ±0.48 ±0.48 ±0.84 ±0.47 ±0.45 ±0.51 ±0.36 ±0.34 ±0.57 ±0.52 ±0.44 ±0.36 -0.05 0.33 0.37 0.75 2.56 1.89 3.75 3.97 4.02 4.26 4.45 4.68 5.83 5.48 ±0.21 ±0.51 ±0.15 ±0.27 ±0.59 ±0.24 ±0.20 ±0.35 ±0.37 ±0.12 ±0.57 ±0.30 ±0.54 ±0.33 -1.90 -0.54 2.87 2.21 4.56 5.33 7.28 6.41 5.86 5.89 5.75 6.74 7.56 7.04 ±0.26 ±0.41 ±0.23 ±0.21 ±0.23 ±0.24 ±0.20 ±0.24 ±0.35 ±0.05 ±0.25 ±0.08 ±0.32 ±0.59 -0.69 1.92 2.03 2.19 2.49 2.64 3.95 3.79 4.04 2.87 2.11 2.09 1.50 2.63 ±0.09 ±0.38 ±0.09 ±0.66 ±0.51 ±0.31 ±0.33 ±0.20 ±0.20 ±0.28 ±0.08 ±0.30 ±0.30 ±0.10 -0.99 -1.40 -0.95 -0.15 -0.36 -0.10 -0.17 -0.58 -0.72 -1.40 -2.13 -2.47 -2.92 -3.45 ±0.23 ±0.03 ±0.11 ±0.01 ±0.11 ±0.06 ±0.11 ±0.11 ±0.10 ±0.10 ±0.13 ±0.00 ±0.08 ±0.10 0.42 2.97 3.57 4.81 4.93 5.09 4.83 4.72 4.34 3.37 2.43 2.65 2.44 2.36 ±0.14 ±0.24 ±0.21 ±0.15 ±0.21 ±0.27 ±0.12 ±0.08 ±0.16 ±0.21 ±0.24 ±0.34 ±0.05 ±0.33 -0.52 -1.33 2.82 3.89 4.33 6.42 7.48 7.84 7.44 8.16 7.62 8.25 9.23 9.35 ±0.17 ±0.22 ±0.28 ±0.35 ±0.34 ±0.15 ±0.24 ±0.25 ±0.28 ±0.47 ±0.25 ±0.19 ±0.26 ±0.46 -2.27 -3.97 -2.79 -1.47 -1.29 -0.50 0.92 1.54 1.29 1.40 1.04 0.88 0.97 0.34 ±0.26 ±0.22 ±0.11 ±0.41 ±0.30 ±0.41 ±0.20 ±0.38 ±0.18 ±0.16 ±0.09 ±0.20 ±0.02 ±0.18 -13.71 -19.33 -15.27 -12.37 -14.70 -11.77 -11.46 -13.51 -11.78 -15.39 -16.83 -17.90 -16.82 -18.55 ±1.14 ±1.64 ±1.13 ±0.73 ±0.57 ±1.00 ±0.81 ±0.34 ±1.07 ±1.07 ±1.19 ±2.29 ±1.95 ±1.77 0.18 -1.24 -2.34 -1.37 -1.21 -0.79 0.28 0.52 -0.12 -1.03 -1.34 -0.38 0.82 -0.62 ±0.17 ±0.47 ±0.27 ±0.28 ±0.22 ±0.54 ±0.58 ±0.34 ±0.46 ±0.10 ±0.37 ±0.70 ±1.04 ±1.05 -3.37 -5.51 -5.63 -4.97 -5.13 -4.13 -3.38 -4.51 -6.56 -7.78 -8.91 -10.08 -10.44 -11.09 ±0.16 ±0.10 ±0.04 ±0.31 ±0.10 ±0.31 ±0.31 ±0.05 ±0.37 ±0.14 ±0.22 ±0.65 ±0.17 ±0.43

0.00 -3.79 -9.28 -7.65 -8.07 -7.91 -5.51 -2.48 -2.09 -2.71 -3.39 -2.66 -2.72 -3.00 -5.12 ±0.00 ±0.63 ±0.24 ±0.70 ±0.07 ±0.77 ±1.24 ±0.98 ±0.41 ±0.42 ±0.52 ±0.34 ±0.20 ±0.36 ±0.46 0.00 -1.36 -2.48 -2.48 -2.01 -2.12 -1.06 -0.02 1.75 0.38 1.19 1.12 1.98 2.74 2.35 ±0.00 ±0.13 ±0.10 ±0.43 ±0.44 ±0.42 ±0.52 ±0.27 ±0.08 ±0.35 ±0.53 ±0.50 ±0.78 ±0.49 ±0.97 0.00 -4.59 -7.04 -6.11 -5.84 -5.28 -5.06 -3.56 -3.51 -3.56 -5.45 -5.78 -6.57 -5.91 -5.23 ±0.00 ±0.15 ±0.28 ±0.25 ±0.38 ±0.40 ±0.17 ±0.58 ±0.42 ±0.25 ±0.14 ±0.63 ±0.12 ±0.67 ±0.31 0.00 0.24 0.86 3.31 5.90 6.02 6.01 5.47 5.82 5.75 5.86 6.56 6.10 7.64 7.63 ±0.00 ±0.08 ±0.39 ±0.48 ±0.36 ±1.16 ±0.77 ±0.67 ±0.17 ±0.33 ±0.67 ±0.67 ±0.80 ±0.89 ±1.32 0.00 0.07 2.68 5.58 9.45 11.13 12.20 11.51 12.37 13.02 12.74 14.13 16.56 20.22 21.25 ±0.00 ±0.11 ±0.38 ±0.24 ±0.25 ±0.36 ±0.26 ±0.45 ±0.06 ±0.29 ±0.40 ±0.48 ±0.51 ±0.65 ±0.60 0.00 0.29 2.66 5.64 8.87 10.51 11.20 10.89 11.62 12.41 11.81 13.11 16.02 19.35 19.81 ±0.00 ±0.46 ±0.16 ±0.12 ±0.61 ±0.47 ±0.44 ±0.86 ±0.38 ±0.59 ±0.52 ±0.42 ±0.32 ±0.07 ±0.46 0.00 1.77 0.67 1.72 2.88 2.79 3.93 4.15 3.42 2.65 4.00 3.65 4.11 6.20 5.96 ±0.00 ±0.12 ±0.66 ±0.42 ±0.12 ±0.46 ±0.44 ±0.41 ±0.31 ±0.81 ±0.39 ±0.59 ±1.10 ±1.32 ±0.76 0.00 -6.24 -1.99 -3.52 -2.62 -0.27 3.29 4.04 7.03 7.88 7.83 7.44 9.81 9.22 9.55 ±0.00 ±0.30 ±0.39 ±0.42 ±0.42 ±0.50 ±0.24 ±0.47 ±0.24 ±0.13 ±0.34 ±0.76 ±0.64 ±0.81 ±0.62 0.00 -2.37 -6.95 -3.94 -2.99 -3.21 -2.44 1.49 -1.01 0.62 -1.59 -2.03 -2.14 -1.04 -2.90 ±0.00 ±0.58 ±0.37 ±0.49 ±0.57 ±1.20 ±1.79 ±2.11 ±1.83 ±0.91 ±0.43 ±0.60 ±0.58 ±0.85 ±1.03 0.00 -6.80 -5.92 1.07 0.09 -0.75 0.08 0.38 2.02 -0.24 0.53 -0.47 -0.06 0.64 -0.58 ±0.00 ±0.43 ±0.50 ±0.40 ±0.39 ±0.39 ±0.24 ±0.28 ±0.30 ±0.42 ±0.11 ±0.82 ±0.28 ±0.56 ±0.31 0.00 -3.09 -3.74 0.53 2.15 0.83 1.08 1.73 -0.72 -1.24 -4.39 -5.37 -5.99 -3.92 -5.15 ±0.00 ±1.26 ±1.49 ±0.85 ±0.52 ±0.53 ±0.79 ±0.52 ±0.98 ±1.22 ±1.12 ±0.90 ±0.76 ±0.97 ±0.56 0.00 -1.39 -5.64 -0.45 1.08 1.32 0.92 1.43 -0.18 -0.89 -4.25 -6.94 -6.52 -4.67 -5.43 ±0.00 ±0.72 ±0.88 ±0.33 ±1.00 ±1.01 ±0.56 ±1.36 ±1.19 ±1.13 ±1.25 ±0.77 ±1.13 ±0.84 ±1.73 0.00 -3.97 -7.26 -2.78 -2.71 -0.90 -1.38 -1.29 -4.69 -5.72 -10.10 -12.09 -11.08 -8.97 -10.31 ±0.00 ±0.09 ±1.58 ±0.89 ±1.23 ±0.61 ±0.66 ±1.01 ±0.42 ±0.18 ±0.91 ±1.26 ±1.89 ±2.62 ±2.31 0.00 -0.65 -1.54 -1.56 -1.07 -1.55 -1.73 -0.18 0.08 -0.00 1.24 2.28 3.23 4.32 3.31 ±0.00 ±0.31 ±0.17 ±0.46 ±0.27 ±0.63 ±0.04 ±0.67 ±0.52 ±0.24 ±0.46 ±0.89 ±0.96 ±1.00 ±0.58 0.00 -0.74 -0.91 -1.17 1.84 1.72 2.21 2.25 1.08 1.49 0.31 -1.48 -0.50 -1.60 -2.38 ±0.00 ±0.29 ±0.13 ±0.15 ±0.04 ±0.23 ±0.33 ±0.06 ±0.26 ±0.30 ±0.32 ±0.20 ±0.49 ±0.05 ±0.15 0.00 -0.09 1.53 6.03 9.11 11.41 13.31 17.18 15.72 14.95 13.22 11.80 14.70 16.50 14.89 ±0.00 ±0.55 ±0.61 ±0.45 ±0.39 ±0.27 ±0.23 ±0.87 ±0.39 ±0.68 ±1.25 ±0.24 ±1.25 ±0.64 ±1.13 0.00 -0.27 1.62 5.79 8.79 10.86 12.48 16.74 16.00 15.26 13.59 12.71 15.21 16.55 15.31 ±0.00 ±0.20 ±0.47 ±0.51 ±0.65 ±0.38 ±0.45 ±0.64 ±0.63 ±1.23 ±1.33 ±0.86 ±1.65 ±2.21 ±2.38 0.00 -1.06 -2.40 0.32 2.01 1.80 3.36 4.51 3.05 3.20 2.50 2.21 0.98 0.77 -0.59 ±0.00 ±0.42 ±0.79 ±0.90 ±0.10 ±0.63 ±1.15 ±0.28 ±0.48 ±1.04 ±0.68 ±0.71 ±0.85 ±0.57 ±0.88 0.00 -1.61 -4.87 -1.34 -1.15 -0.77 -0.44 1.51 -0.81 -0.99 -1.91 -1.38 -0.67 1.25 -0.84 ±0.00 ±0.99 ±1.52 ±0.89 ±1.08 ±0.92 ±1.22 ±0.97 ±0.96 ±1.74 ±1.02 ±1.15 ±1.62 ±0.68 ±0.48 0.00 -1.25 2.68 3.45 4.44 3.97 3.81 4.64 7.96 4.32 5.74 5.70 6.57 13.80 10.71 ±0.00 ±1.02 ±0.85 ±0.52 ±1.24 ±2.12 ±2.60 ±1.50 ±2.00 ±1.90 ±0.43 ±0.61 ±1.68 ±2.28 ±0.69 0.00 -7.85 -7.67 -3.63 -2.25 -0.73 0.41 1.50 0.27 0.13 -1.63 -4.22 -4.79 -4.14 -2.83 ±0.00 ±0.41 ±0.64 ±0.67 ±0.99 ±0.62 ±0.99 ±0.62 ±0.48 ±0.67 ±1.14 ±1.13 ±0.58 ±0.37 ±0.95 0.00 -5.16 -14.47 -10.13 -5.54 -6.56 -3.62 0.41 3.09 4.67 2.96 3.47 4.69 3.89 1.41 ±0.00 ±0.25 ±0.36 ±0.45 ±0.21 ±0.16 ±0.34 ±0.42 ±0.55 ±0.24 ±0.53 ±0.86 ±0.64 ±0.59 ±0.49 0.00 -6.64 -11.15 -10.05 -7.79 -8.52 -6.74 -4.98 -5.77 -5.60 -9.44 -10.59 -10.69 -8.44 -9.49 ±0.00 ±1.26 ±0.76 ±0.71 ±1.06 ±0.48 ±0.71 ±0.47 ±0.97 ±0.71 ±0.63 ±1.64 ±0.99 ±0.36 ±1.31 0.00 -14.98 -20.55 -19.82 -12.45 -15.53 -13.41 -9.02 -12.63 -9.56 -14.17 -15.43 -16.13 -15.79 -17.12 ±0.00 ±0.60 ±0.31 ±0.49 ±0.79 ±0.23 ±0.74 ±0.72 ±0.48 ±1.01 ±0.40 ±0.29 ±0.55 ±0.94 ±1.27 0.00 2.54 1.33 2.00 6.74 6.24 6.83 7.42 9.73 11.50 11.41 12.70 12.64 11.91 11.41 ±0.00 ±0.31 ±0.42 ±0.63 ±0.60 ±0.62 ±0.64 ±0.24 ±0.45 ±0.09 ±0.25 ±0.26 ±0.30 ±0.64 ±0.90 0.00 0.60 1.19 2.57 3.01 3.49 3.12 3.20 3.00 1.83 3.31 3.29 3.13 4.34 4.05 ±0.00 ±0.26 ±0.36 ±0.22 ±0.62 ±0.22 ±1.22 ±0.34 ±0.83 ±0.47 ±0.67 ±0.57 ±0.55 ±1.41 ±1.01

0

1

2

3

4

5

6

7

8

30

20

10

0

10

20

9 10 11 12 13 14

hdrzine hdrzone imidazole imide ketone ketone_Topliss lactone methoxy morpholine nitrile nitro nitro_arom nitro_arom_nonortho oxime para_hydroxylation phenol phenol_noOrthoHbond piperdine piperzine priamide pyridine sulfide thiazole thiophene unbrch_alkane urea

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

-5.58 ±0.33 -1.03 ±0.18 -3.29 ±0.58 2.37 ±0.80 0.64 ±0.12 0.50 ±0.09 1.14 ±0.25 -7.29 ±0.32 -3.61 ±0.20 -1.14 ±0.34 -3.31 ±1.08 -2.40 ±1.28 -0.95 ±1.17 -0.91 ±0.33 0.33 ±0.24 0.57 ±0.41 0.87 ±0.38 -1.49 ±0.16 -2.05 ±1.02 1.65 ±0.64 -4.42 ±0.38 -10.59 ±0.71 -6.71 ±0.64 -13.14 ±0.39 -7.61 ±0.10 0.33 ±1.03

-9.04 ±0.69 -2.11 ±0.44 -5.40 ±0.67 1.01 ±0.75 0.54 ±0.15 0.94 ±0.20 0.93 ±0.30 -12.20 ±0.40 -4.13 ±0.63 -5.26 ±0.72 -2.25 ±0.99 -4.12 ±1.86 -6.08 ±2.54 -2.92 ±0.49 0.01 ±0.21 0.36 ±0.59 0.43 ±0.37 -3.20 ±0.98 -7.47 ±0.99 -0.23 ±0.75 -9.30 ±0.85 -15.89 ±0.49 -9.40 ±0.48 -19.93 ±0.80 -10.56 ±0.74 -0.10 ±0.84

-11.67 ±1.10 -2.78 ±0.64 -6.13 ±0.27 3.60 ±1.14 2.40 ±0.09 2.34 ±0.09 1.47 ±0.41 -15.51 ±0.35 -4.52 ±0.83 -4.01 ±0.68 -1.72 ±2.51 -3.68 ±0.44 -4.22 ±0.86 -2.94 ±0.62 -0.41 ±0.47 3.29 ±0.06 3.29 ±0.34 -3.41 ±0.66 -8.08 ±1.29 -1.74 ±0.26 -10.61 ±0.45 -18.79 ±0.47 -11.03 ±0.54 -22.19 ±0.54 -13.89 ±1.09 0.02 ±1.35

-12.20 ±1.44 -3.03 ±1.03 -6.39 ±0.11 5.42 ±1.25 3.71 ±0.45 3.90 ±0.17 1.29 ±0.34 -18.44 ±0.16 -5.09 ±1.20 -6.12 ±0.69 -2.45 ±1.35 -2.72 ±0.38 -4.25 ±1.76 -1.75 ±0.94 -1.20 ±0.44 4.61 ±0.49 4.11 ±0.63 -2.65 ±0.51 -8.46 ±1.32 -1.54 ±0.20 -13.11 ±0.63 -20.00 ±0.66 -12.84 ±0.95 -24.02 ±0.28 -16.36 ±1.06 0.67 ±1.15

-13.34 ±0.43 -3.04 ±0.18 -6.78 ±0.08 6.07 ±1.12 4.61 ±0.17 4.49 ±0.33 2.49 ±0.31 -20.36 ±0.07 -5.03 ±1.06 -6.86 ±0.48 -1.80 ±2.04 -2.49 ±0.51 -5.03 ±1.46 -1.08 ±0.91 -0.35 ±0.28 9.14 ±0.40 9.04 ±0.73 -2.17 ±0.59 -7.51 ±0.75 -0.99 ±0.21 -11.59 ±0.84 -21.97 ±0.48 -11.76 ±1.03 -23.72 ±0.50 -18.32 ±0.61 1.11 ±1.24

-14.70 ±0.84 -3.46 ±0.44 -7.85 ±0.91 6.74 ±1.23 4.84 ±0.56 4.61 ±0.71 2.73 ±0.49 -22.40 ±0.01 -5.47 ±1.23 -7.77 ±0.56 -1.82 ±2.11 -1.94 ±0.43 -5.03 ±1.58 -1.62 ±0.51 -0.88 ±0.36 7.66 ±0.33 7.46 ±0.64 -2.90 ±0.61 -8.97 ±0.81 -0.93 ±0.20 -10.75 ±1.08 -23.08 ±0.67 -12.27 ±0.58 -25.75 ±0.73 -20.32 ±0.65 0.49 ±1.34

0

1

2

3

4

5

6

15 10 5 0 5 10 15 20 25

(d) Chemberta-base

4.87 ±0.21 5.62 ±0.41 -4.78 ±0.69 -0.14 ±0.45 0.91 ±0.24 1.18 ±0.25 7.06 ±0.16 1.83 ±0.47 9.87 ±0.08 2.05 ±0.35 -4.27 ±0.21 -4.79 ±1.16 5.15 ±0.27 0.92 ±0.20 0.51 ±0.19 3.31 ±0.23 1.43 ±0.37 6.30 ±0.41 12.66 ±0.55 -1.04 ±0.05 4.52 ±0.05 7.95 ±0.25 6.31 ±0.36 -9.80 ±1.12 2.66 ±0.42 4.35 ±0.76

Molecular Substructure

Molecular Substructure

(c) Chemberta NHOH NO AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings HAcceptors HDonors Heteroatoms RotatableBonds SaturatedCarbocycles SaturatedHeterocycles SaturatedRings Ring Al_COO Al_OH Al_OH_noTert ArN Ar_COO Ar_N Ar_NH Ar_OH COO COO2

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

(b) Roberta-zinc-480m

Molecular Substructure

Molecular Substructure

(a) Molformer

C_O C_O_noCOO C_S Imine NH0 NH1 NH2 N_O Ndealkylation1 Ndealkylation2 Nhpyrrole SH aldehyde alkyl_halide allylic_oxid amide amidine aniline aryl_methyl benzene bicyclic ester ether furan guanido halogen

Layer

hdrzine hdrzone imidazole imide ketone ketone_Topliss lactone methoxy morpholine nitrile nitro nitro_arom nitro_arom_nonortho oxime para_hydroxylation phenol phenol_noOrthoHbond piperdine piperzine priamide pyridine sulfide thiazole thiophene unbrch_alkane urea

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

0.54 ±0.46 0.58 ±0.45 -4.23 ±0.21 4.01 ±0.47 1.80 ±0.42 2.81 ±0.29 4.43 ±0.36 18.68 ±0.50 3.12 ±0.19 0.23 ±0.41 14.28 ±0.79 11.91 ±0.49 8.50 ±0.43 2.49 ±1.05 3.73 ±0.15 3.62 ±0.29 3.28 ±0.18 5.43 ±0.31 1.86 ±0.21 9.81 ±0.53 14.91 ±0.61 -1.44 ±0.27 1.68 ±1.10 4.82 ±0.09 6.63 ±0.55 3.31 ±0.31

2.56 ±0.78 5.95 ±0.43 -1.14 ±0.61 13.27 ±2.09 15.69 ±0.03 16.27 ±0.13 7.18 ±0.26 29.06 ±0.32 9.80 ±0.35 2.88 ±0.29 26.97 ±0.54 27.49 ±0.38 23.44 ±0.71 8.16 ±0.36 12.13 ±0.13 32.49 ±0.43 32.65 ±0.59 10.42 ±0.40 8.61 ±1.53 28.46 ±2.39 13.06 ±0.66 2.71 ±0.81 -6.39 ±1.27 -7.41 ±0.23 11.69 ±0.91 6.59 ±0.32

0.25 ±0.26 3.89 ±0.51 0.88 ±0.49 12.56 ±1.01 19.93 ±0.35 19.91 ±0.14 7.03 ±0.73 27.13 ±0.44 6.02 ±0.44 -0.81 ±0.46 19.90 ±0.91 21.21 ±0.38 17.93 ±1.35 7.00 ±0.66 10.82 ±0.38 27.58 ±0.77 27.18 ±1.08 8.80 ±1.10 7.87 ±0.44 19.55 ±2.92 15.26 ±1.05 -1.35 ±0.62 -5.60 ±1.69 -13.04 ±1.13 10.86 ±0.96 6.21 ±0.97

0

1

2

3

30

20

10

0

10

(f) Chemberta-2-10M hdrzine hdrzone imidazole imide ketone ketone_Topliss lactone methoxy morpholine nitrile nitro nitro_arom nitro_arom_nonortho oxime para_hydroxylation phenol phenol_noOrthoHbond piperdine piperzine priamide pyridine sulfide thiazole thiophene unbrch_alkane urea

40 30 20 10 0 10 20

NHOH NO AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings HAcceptors HDonors Heteroatoms RotatableBonds SaturatedCarbocycles SaturatedHeterocycles SaturatedRings Ring Al_COO Al_OH Al_OH_noTert ArN Ar_COO Ar_N Ar_NH Ar_OH COO COO2

(g) Chemberta-2-77M

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

-15.20 ±0.37 -16.77 ±0.39 -11.80 ±0.25 -9.69 ±0.35 -11.95 ±0.24 -16.27 ±0.14 -16.16 ±0.45 -9.85 ±0.08 -15.73 ±0.22 -14.89 ±0.57 -14.07 ±0.12 -5.21 ±0.23 -12.76 ±0.16 -9.51 ±0.32 -12.13 ±0.36 -8.26 ±0.02 -8.12 ±0.50 -8.10 ±0.26 -7.87 ±0.63 -10.15 ±0.35 -2.95 ±0.55 -17.40 ±0.63 -33.07 ±0.22 -10.12 ±0.73 -10.06 ±0.39 -10.12 ±0.35

-14.35 ±0.06 -14.09 ±0.12 -10.45 ±0.14 -10.35 ±0.24 -10.87 ±0.23 -12.69 ±0.11 -16.54 ±0.40 -7.68 ±0.19 -13.31 ±0.23 -14.22 ±0.25 -13.08 ±0.51 -8.32 ±0.49 -10.30 ±0.54 -8.92 ±0.33 -10.69 ±0.34 -5.39 ±0.11 -5.77 ±0.68 -8.24 ±0.33 -8.58 ±0.39 -9.12 ±0.19 -1.02 ±0.46 -17.94 ±0.50 -30.07 ±0.52 -8.86 ±0.58 -6.93 ±0.19 -6.60 ±0.25

-13.81 ±0.27 -13.66 ±0.59 -10.71 ±0.43 -10.55 ±0.69 -11.10 ±0.38 -13.90 ±0.37 -17.88 ±0.10 -7.99 ±0.08 -12.69 ±0.40 -13.43 ±0.13 -12.70 ±0.49 -7.12 ±0.32 -11.32 ±0.46 -9.59 ±0.55 -11.06 ±0.33 -7.99 ±0.05 -6.93 ±0.31 -7.90 ±0.12 -8.28 ±0.22 -10.14 ±0.37 -2.53 ±0.27 -19.26 ±0.29 -29.29 ±0.56 -7.62 ±0.51 -7.97 ±0.27 -7.86 ±0.13

-11.40 ±0.19 -16.79 ±0.45 -8.89 ±0.37 -7.40 ±0.18 -8.05 ±0.14 -10.86 ±0.14 -17.09 ±0.41 -5.23 ±0.17 -15.03 ±1.28 -11.10 ±0.27 -14.84 ±0.18 -6.83 ±0.24 -9.21 ±0.64 -7.65 ±0.22 -9.15 ±0.05 -4.55 ±0.19 -8.24 ±0.54 -5.94 ±0.46 -6.07 ±0.35 -9.52 ±0.12 -2.01 ±0.41 -19.15 ±0.38 -24.48 ±0.63 -8.30 ±0.37 -8.45 ±0.25 -8.43 ±0.35

-11.08 ±0.04 -15.24 ±0.37 -5.24 ±0.17 -2.75 ±0.20 -1.97 ±0.12 -9.95 ±0.06 -14.56 ±0.16 -2.57 ±0.11 -12.80 ±0.38 -10.46 ±0.28 -12.23 ±0.15 -1.05 ±0.20 -5.84 ±0.21 -3.43 ±0.04 -4.22 ±0.17 -0.74 ±0.04 -3.17 ±0.31 -6.35 ±0.73 -6.08 ±0.59 -2.62 ±0.58 1.25 ±0.58 -16.29 ±0.16 -22.35 ±0.49 -1.74 ±0.30 -2.75 ±0.41 -2.63 ±0.36

-2.93 ±0.28 -4.15 ±0.29 1.13 ±0.47 1.43 ±0.24 0.62 ±0.11 -3.02 ±0.09 -3.99 ±0.29 -0.39 ±0.05 -2.92 ±0.43 -2.29 ±0.52 -3.18 ±0.62 2.47 ±0.39 0.74 ±0.53 1.25 ±0.26 0.32 ±0.14 -0.20 ±0.08 4.62 ±0.18 0.91 ±0.36 0.47 ±0.48 2.42 ±0.40 6.54 ±0.54 -6.33 ±0.03 -10.31 ±0.81 1.67 ±0.62 5.56 ±0.60 5.43 ±0.34

-1.89 ±0.37 -1.00 ±0.77 3.29 ±0.56 2.78 ±0.24 0.76 ±0.13 -2.16 ±0.02 -2.06 ±0.15 -0.21 ±0.08 0.59 ±1.26 -1.36 ±0.25 0.11 ±0.38 3.93 ±0.22 3.22 ±0.38 3.27 ±0.44 1.52 ±0.10 -0.23 ±0.13 7.44 ±0.41 2.26 ±0.56 2.15 ±0.48 2.06 ±0.96 8.69 ±0.90 -4.40 ±0.36 -8.02 ±1.01 5.49 ±0.29 7.57 ±0.18 7.62 ±0.40

-1.80 ±0.38 -2.83 ±1.56 4.63 ±0.05 3.67 ±0.21 0.93 ±0.10 -1.23 ±0.03 -0.43 ±0.03 -0.29 ±0.10 0.36 ±0.65 -1.37 ±0.13 -1.08 ±1.36 4.07 ±0.35 4.83 ±0.66 4.66 ±0.20 2.18 ±0.18 -0.52 ±0.15 8.00 ±0.56 1.59 ±0.72 2.19 ±0.70 4.16 ±0.47 8.42 ±0.87 -2.43 ±0.09 -8.89 ±0.95 5.75 ±0.21 7.25 ±0.24 7.88 ±0.73

-0.32 ±0.36 -0.62 ±0.72 4.82 ±0.40 4.94 ±0.14 0.71 ±0.16 -0.88 ±0.11 0.02 ±0.06 0.09 ±0.05 1.02 ±0.96 0.16 ±0.26 0.19 ±1.28 4.90 ±0.45 4.12 ±0.13 5.43 ±0.21 1.71 ±0.10 -1.24 ±0.20 9.61 ±1.87 3.80 ±0.71 3.62 ±0.65 4.82 ±0.74 12.14 ±0.33 -1.95 ±0.34 -9.38 ±0.86 8.11 ±0.81 11.26 ±0.62 11.26 ±0.83

-0.63 ±0.36 -1.86 ±1.05 4.18 ±0.64 4.89 ±0.39 0.66 ±0.14 -1.12 ±0.05 -0.37 ±0.18 -0.11 ±0.04 0.34 ±0.99 -0.29 ±0.13 -0.24 ±1.78 3.59 ±0.86 5.26 ±0.28 5.90 ±0.39 1.90 ±0.19 -0.49 ±0.14 11.46 ±0.32 4.19 ±0.62 3.52 ±0.56 5.79 ±0.33 13.95 ±1.25 -1.93 ±0.26 -10.20 ±0.33 5.64 ±0.27 11.37 ±0.17 11.19 ±0.33

-2.03 ±0.18 -0.15 ±0.61 1.97 ±0.29 3.88 ±0.21 0.50 ±0.09 -0.89 ±0.20 -1.31 ±0.37 -0.32 ±0.16 0.71 ±0.74 -1.63 ±0.34 0.50 ±0.93 2.48 ±0.70 3.28 ±0.42 4.58 ±0.61 1.05 ±0.27 -0.19 ±0.08 8.67 ±0.25 1.89 ±0.86 1.32 ±0.98 6.41 ±0.52 11.11 ±0.96 -2.10 ±0.27 -14.39 ±0.69 4.46 ±0.14 9.08 ±0.37 8.93 ±0.05

-2.30 ±0.31 -2.63 ±0.85 1.07 ±0.38 2.59 ±0.21 -0.43 ±0.11 -2.12 ±0.15 -4.28 ±0.40 -0.87 ±0.14 -1.05 ±1.20 -2.11 ±0.30 0.16 ±1.04 -0.06 ±0.04 1.01 ±0.21 2.60 ±0.68 -1.06 ±0.18 -0.10 ±0.02 5.14 ±0.17 -0.27 ±0.63 -0.50 ±0.58 6.38 ±0.77 7.69 ±1.46 -4.93 ±0.28 -13.93 ±0.86 2.59 ±0.25 5.79 ±0.98 4.55 ±0.51

0

1

2

3

4

5

6

7

8

9

10

11

12

C_O C_O_noCOO C_S Imine NH0 NH1 NH2 N_O Ndealkylation1 Ndealkylation2 Nhpyrrole SH aldehyde alkyl_halide allylic_oxid amide amidine aniline aryl_methyl benzene bicyclic ester ether furan guanido halogen

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

-8.71 ±0.29 -4.64 ±0.17 -19.06 ±1.21 -7.43 ±0.30 -11.54 ±0.54 -18.23 ±0.20 -15.74 ±0.50 -10.32 ±1.15

-6.19 ±0.27 -4.36 ±0.18 -18.36 ±1.23 -6.91 ±0.08 -10.77 ±0.10 -20.69 ±0.06 -12.81 ±0.78 -11.15 ±0.52

-6.76 ±0.40 -5.20 ±0.13 -17.95 ±0.63 -6.73 ±0.41 -10.45 ±0.27 -19.76 ±0.13 -13.87 ±0.48 -10.45 ±1.31

-7.93 ±0.25 -5.14 ±0.14 -20.66 ±0.55 -6.30 ±0.09 -10.81 ±0.34 -19.26 ±0.04 -10.44 ±0.06 -9.96 ±0.69

-5.90 ±0.09 -4.20 ±0.44 -18.81 ±0.82 -7.44 ±0.25 -9.51 ±0.02 -16.72 ±0.15 -6.95 ±0.23 -8.94 ±1.33

0.65 ±0.19 2.60 ±0.23 -11.03 ±0.22 -0.58 ±0.34 -2.71 ±0.16 -8.37 ±0.52 -0.87 ±0.17 -6.07 ±0.49

0.49 ±0.30 2.49 ±0.09 -9.59 ±0.88 2.35 ±0.74 -0.53 ±0.17 -6.27 ±0.46 0.51 ±0.15 -5.84 ±0.50

0.55 ±0.35 2.59 ±0.20 -8.58 ±0.75 2.28 ±0.33 -0.47 ±0.27 -5.93 ±0.16 0.58 ±0.72 -5.65 ±

0.97 ±0.18 3.70 ±0.20 -8.77 ±0.66 3.21 ±0.21 1.28 ±0.26 -4.86 ±0.50 3.86 ±0.33

0.71 ±0.21 3.65 ±0.24 -10.29 ±0.52 3.80 ±0.24 0.60 ±0.19 -5.50 ±0.23 2.06 ±0.45

-0.36 ±0.02 1.90 ±0.17 -10.69 ±1.05 4.38 ±0.41 -0.43 ±0.27 -6.74 ±0.35 1.97 ±0.36

-2.35 ±0.19 0.18 ±0.15 -12.24 ±0.92 3.52 ±0.24 -0.66 ±0.28 -7.91 ±0.08 1.04 ±0.20

0

1

2

3

4

5

6

7

8

9

10

11

12

h Chembe a 3

F gure 7 Effect of pre-tra n ng on mo ecu ar substructure encod ng n CLMs We p o he ayer-w se d fference n performance (% macro-averaged F1) on each prob ng ask be ween each pre- ra ned mode and s random y n a zed coun erpar Red deno es mprovemen n prob ng performance af er pre- ra n ng B ue des gna es degrada on n prob ng performance af er pre- ra n ng

prediction (cf Section F 1) we exclude probing tasks for substructures appearing fewer than ten mes n e her he ra n ng or es sp of he ESOL da ase As a resu our ESOL fine- un ng ana yses are based on 39 molecular substructures listed in Table 7 F3

a common set of substructures difficult For example ∼ 90% molecules in the lipophilicity dataset contain AromaticHeterocycles whereas AromaticHeterocycles are found in only 10%15% of molecules in ESOL Some substructures appearing frequently in the lipophilicity dataset (e g pyridine COO and methoxy) are absent from the ESOL dataset altogether and vice versa (unbrch-alkane imide and allylic-oxid) We note that this also results in slight differences for groups of hydroph c poph c and o her mo ecular substructures (cf Table 6 and Table 7)

Lipophilicity and Solubility: Differences

After preprocessing we analyze both datasets in erms of he d s r bu on of mo ecu ar subs ruc ures as well as SMILES strings (i e molecules) Overlap in molecular substructures We observe that the distributions of substructures between the lipophilicity and ESOL datasets are quite different making a direct comparison on

Overlap in SMILES We further inspected the poph c y and so ub y da ase s for over app ng SMILES strings which would have enabled us 21

train set 3360 molecules

valid set 420 molecules

test set 420 molecules

train set 901 molecules

valid set 113 molecules

12.0%

8.0%

7.0%

14.0%

14.0%

12.0%

12.0%

17.5%

10.0%

20.0%

15.0%

6.0% 10.0%

10.0%

8.0%

8.0%

6.0%

6.0%

8.0% 15.0%

12.5%

% molecules

% molecules

5.0%

4.0%

6.0%

10.0% 10.0%

3.0%

7.5%

4.0%

2.0%

4.0%

4.0%

2.0%

2.0%

5.0% 5.0%

1.0%

0.0%

test set 113 molecules

25.0%

20.0%

3

2

1

0

1

2

0.0%

3

2

1

0

1

2

0.0%

2.0% 2.5%

3

2

1

0

1

2

0.0%

12

10

8

6

4

2

0

2

0.0%

8

6

4

logD

logS

(a) logD

(b) logS

2

0

0.0%

8

6

4

2

0

Figure 8: The distribution of predicted values across the training, validation and test splits of the lipophilicity (8a) and solubility (ESOL) (8b) datasets.

to investigate how the two tasks influence the encoding of substructures present in a molecule. In particular, if the effects are reversed. However, we identified only 37 overlapping molecules in the training sets and a single one in the test sets across the two datasets. In light of this, we decide against a direct comparison between the effect of fine-tuning on lipophilicity and ESOL on substructure encoding. Instead, we conduct separate analyses for the two datasets, based on 60 and 39 substructures, respectively.

We can furthermore observe that the effect of fine-tuning on the probing performance is noticeably smaller compared to pre-training (similar as we observed for lipophilicity prediction). Interestingly, we see several common patterns with respect to how fine-tuning on solubility affects RI and PT models. Namely, we find that all RI models improve on hydrophilic substructures during finetuning. However, in contrast to PT models, these changes are more prominent in lower and middle layers.

G

G.2

Effect of Fine-Tuning (RQ2)

The heatmaps in Figure 11 present the layerwise differences in probing performance after finetuning on solubility prediction for each substructure (Figure 10 presents the respective heatmaps for lipophilicity prediction). First, our results show that fine-tuning on solubility improves encoding of NO (nitrogens and oxygens), Heteroatoms and HAcceptors across all PT models. In addition, all models except for chemberta-2-10M exhibit improved probing performance on phenols (phenol, phenol_noOrthoHbond), various hydroxyl groups (i.e., aliphatic hydroxyl groups (Al_OH), aromatic hydroxyl groups (Ar_OH), aliphatic hydroxyl groups excluding tert-OH (Ar_OH_noTert)), as well as aromatic nitrogens (Ar_N). These results are consistent with chemical theory. More specifically, Walker (2017) state that replacing a carbon atom with polar heteroatoms N or O is one of the most common approaches to increasing solubility. Furthermore, Harrold et al. (2023) note that the ability of polar hydroxyl groups (OH) to form hydrogen bonds leads to an increase in aqueous solubility.

Similar to lipophilicity prediction (§6), we analyse the probing performance on solubility prediction for different groups and individual molecular substructures. G.1

Solubility: Individual Analysis

Solubility: Group Analysis

Figure 9 shows the changes (before and after finetuning on solubility) in terms of probing performance. We observe similar trends: fine-tuning leads to larger changes in task-relevant groups (lipophilic and hydrophilic), concentrated in the upper layers. In contrast to lipophilicity, finetuning on solubility has a more pronounced effect on hydrophilic substructures, as highlighted by the darker red shading. This corroborates our lipophilicity results that molecular substructure learning is consistent with chemical theory. This still holds when we observe chemberta-2-10M, which exhibits a consistent unlearning effect on both downstream tasks: For solubility prediction, the unlearning effect is stronger for the lipophilic groups while for lipophilicity prediction it is stronger for the hydrophilic groups. 22

chemberta-base

0.00 ±0.00

0.98 ±0.24

2.10 ±0.16

1.13 ±0.24

1.04 ±0.33

2.28 ±0.25

2.85 ±0.21

lipophilic (11)

0.00 ±0.00

-0.02 ±0.08

-0.10 ±0.09

-0.60 ±0.15

-0.87 ±0.11

-0.80 ±0.18

-0.42 ±0.17

other (11)

0.00 ±0.00

1.12 ±0.13

1.00 ±0.15

-0.08 ±0.17

-0.72 ±0.21

-0.59 ±0.12

-0.88 ±0.40

0

1

2

3

4

5

6

chemberta-2-10M

hydrophilic (17)

0.00 ±0.00

-2.69 ±0.37

-2.43 ±0.83

-1.73 ±0.44

lipophilic (11)

0.00 ±0.00

-1.63 ±0.13

-3.98 ±0.08

-1.25 ±0.11

other (11)

0.00 ±0.00

-1.88 ±0.11

-3.00 ±0.62

-2.71 ±0.37

0

1

2

3

chemberta-2-77M

hydrophilic (17)

0.00 ±0.00

0.74 ±0.12

4.03 ±0.24

-2.20 ±0.29

lipophilic (11)

0.00 ±0.00

0.16 ±0.09

0.31 ±0.17

1.35 ±0.27

other (11)

0.00 ±0.00

-0.27 ±0.12

0.12 ±0.38

-5.27 ±0.43

0

1

2

3

2 1 0

0.00 ±0.00

-0.36 ±0.17

-0.99 ±0.25

1.14 ±0.45

2.61 ±0.54

4.22 ±0.75

4.60 ±0.50

0.00 ±0.00

-1.90 ±0.11

-3.20 ±0.19

-1.69 ±0.10

-0.65 ±0.18

-0.66 ±0.13

-0.20 ±0.12

0.00 ±0.00

0.35 ±0.20

-0.85 ±0.15

-0.62 ±0.24

-0.03 ±0.30

0.44 ±0.29

0.86 ±0.55

0

1

2

3

4

5

6

0 1 2 3

chemberta-2-5M

0.00 ±0.00

-5.55 ±0.31

6.09 ±0.49

3.53 ±0.36

0.00 ±0.00

-6.36 ±0.14

-2.13 ±0.06

-1.04 ±0.16

0.00 ±0.00

-6.27 ±0.22

-2.02 ±0.22

-2.73 ±0.33

0

1

2

3

chemberta-3

2.5 0.0 2.5 5.0

0.00 0.05 0.02 -1.00 -0.75 -1.09 0.37 0.42 0.73 0.90 1.90 2.51 3.62 ±0.00 ±0.17 ±0.12 ±0.18 ±0.14 ±0.14 ±0.20 ±0.35 ±0.46 ±0.41 ±0.49 ±0.37 ±0.51

0.00 -0.19 0.34 0.52 0.33 0.71 1.11 1.03 1.36 1.37 2.24 2.39 2.07 hydrophilic (17) ±0.00 ±0.31 ±0.24 ±0.28 ±0.25 ±0.18 ±0.33 ±0.26 ±0.32 ±0.42 ±0.35 ±0.24 ±0.45

2

0.00 0.05 -0.13 -0.07 0.59 0.75 1.13 1.41 1.48 2.90 3.04 2.75 3.15 2.29 2.20 ±0.00 ±0.06 ±0.09 ±0.08 ±0.20 ±0.10 ±0.15 ±0.28 ±0.22 ±0.29 ±0.26 ±0.34 ±0.27 ±0.24 ±0.47

0.00 0.05 0.01 0.10 0.11 0.16 0.33 0.50 0.49 0.55 0.84 1.05 0.41 lipophilic (11) ±0.00 ±0.11 ±0.11 ±0.13 ±0.13 ±0.08 ±0.11 ±0.20 ±0.15 ±0.11 ±0.09 ±0.08 ±0.14

1

0.00 -0.06 0.02 -0.03 0.08 -0.01 -0.06 0.03 0.22 0.71 0.64 0.59 1.27 1.17 1.79 ±0.00 ±0.07 ±0.09 ±0.11 ±0.05 ±0.06 ±0.09 ±0.05 ±0.07 ±0.11 ±0.16 ±0.11 ±0.11 ±0.09 ±0.19

0.00 0.01 0.25 0.40 0.27 0.37 0.38 0.52 0.49 0.45 0.71 0.83 0.20 other (11) ±0.00 ±0.14 ±0.16 ±0.28 ±0.23 ±0.18 ±0.19 ±0.33 ±0.36 ±0.16 ±0.16 ±0.40 ±0.35 0 1 2 3 4 5 6 7 8 9 10 11 12

0

molformer

0.00 -0.09 -0.13 -0.08 -1.57 -0.43 -0.33 -0.44 -0.16 0.20 0.47 0.57 1.83 ±0.00 ±0.08 ±0.10 ±0.13 ±0.19 ±0.12 ±0.11 ±0.10 ±0.13 ±0.06 ±0.13 ±0.13 ±0.12 0.00 0.12 0.25 -0.21 -0.44 -0.95 0.13 0.12 0.35 0.48 0.57 -0.08 1.45 ±0.00 ±0.15 ±0.05 ±0.11 ±0.19 ±0.20 ±0.18 ±0.19 ±0.47 ±0.34 ±0.33 ±0.47 ±0.44

0

1

2

3

4

5

6

7

8

roberta-zinc-480m

9

0

1

2

3

4

5

6

7

8

2 5 0 5

2 0

10 11 12

0.00 -0.08 -0.18 -0.12 0.06 0.01 0.15 0.05 0.31 0.57 0.23 0.01 0.43 0.14 0.05 ±0.00 ±0.07 ±0.09 ±0.06 ±0.12 ±0.09 ±0.12 ±0.23 ±0.15 ±0.20 ±0.09 ±0.21 ±0.22 ±0.29 ±0.42

Relative Layer Depth

chemberta

4 2 0

Category

Category

chemberta hydrophilic (17)

9 10 11 12 13 14

3

chemberta-base

hydrophilic (17)

0.00 ±0.00

1.69 ±0.26

1.12 ±0.36

0.78 ±0.26

-0.02 ±0.25

-0.84 ±0.25

-1.43 ±0.21

lipophilic (11)

0.00 ±0.00

-1.46 ±0.13

-1.83 ±0.14

-2.71 ±0.18

-3.42 ±0.15

-4.01 ±0.11

-3.75 ±0.15

other (11)

0.00 ±0.00

0.19 ±0.15

-0.08 ±0.14

-0.62 ±0.18

-1.52 ±0.27

-2.37 ±0.21

-3.27 ±0.30

0

1

2

3

4

5

6

chemberta-2-10M

hydrophilic (17)

0.00 ±0.00

2.47 ±0.24

2.31 ±0.28

1.07 ±0.19

lipophilic (11)

0.00 ±0.00

-2.85 ±0.26

-3.33 ±0.13

-4.03 ±0.13

other (11)

0.00 ±0.00

-2.05 ±0.26

-1.54 ±0.15

-1.91 ±0.11

0

1

2

3

chemberta-2-77M

hydrophilic (17)

0.00 ±0.00

2.47 ±0.24

2.31 ±0.28

1.07 ±0.19

lipophilic (11)

0.00 ±0.00

-2.85 ±0.26

-3.33 ±0.13

-4.03 ±0.13

other (11)

0.00 ±0.00

-2.05 ±0.26

-1.54 ±0.15

-1.91 ±0.11

0

1

2

3

molformer

0.00 0.41 0.25 0.50 0.56 0.49 0.41 0.52 0.54 0.49 0.48 0.50 0.47 hydrophilic (17) ±0.00 ±0.14 ±0.25 ±0.12 ±0.10 ±0.15 ±0.17 ±0.14 ±0.13 ±0.20 ±0.19 ±0.22 ±0.12

2

0.00 -0.15 -0.08 -0.02 -0.00 0.03 0.03 0.03 0.07 0.05 0.05 0.09 0.05 lipophilic (11) ±0.00 ±0.11 ±0.04 ±0.10 ±0.12 ±0.07 ±0.07 ±0.09 ±0.10 ±0.08 ±0.07 ±0.19 ±0.16

1

0.00 -0.22 -0.10 0.06 0.11 0.19 0.29 0.29 0.45 0.49 0.52 0.50 0.50 other (11) ±0.00 ±0.32 ±0.33 ±0.14 ±0.18 ±0.20 ±0.20 ±0.17 ±0.20 ±0.22 ±0.24 ±0.25 ±0.17 0 1 2 3 4 5 6 7 8 9 10 11 12

0

0 2 4

0.00 ±0.00

-5.37 ±0.18

-7.36 ±0.24

-8.78 ±0.27

-9.65 ±0.24

-11.78 ±0.31

-16.62 ±0.25

0.00 ±0.00

-7.57 ±0.19

-9.44 ±0.12

-10.23 ±0.10

-10.73 ±0.14

-12.23 ±0.16

-16.57 ±0.17

0.00 ±0.00

-5.39 ±0.15

-7.30 ±0.26

-8.44 ±0.29

-9.04 ±0.20

-10.91 ±0.14

-15.59 ±0.33

0

1

2

3

4

5

6

2 0 2 4 2

chemberta-2-5M

0.00 ±0.00

2.47 ±0.24

2.31 ±0.28

1.07 ±0.19

0.00 ±0.00

-2.85 ±0.26

-3.33 ±0.13

-4.03 ±0.13

0.00 ±0.00

-2.05 ±0.26

-1.54 ±0.15

-1.91 ±0.11

0

1

2

3

chemberta-3

0.00 1.54 0.77 0.01 -1.46 -2.62 -4.07 -5.55 -7.03 -8.83 -11.18 -13.72 -17.61 ±0.00 ±0.32 ±0.20 ±0.18 ±0.23 ±0.25 ±0.29 ±0.29 ±0.32 ±0.44 ±0.39 ±0.39 ±0.47

0 2 4 0.4 0.2 0.0 0.2

0 5 10 15 2 0 2 4 0 5 10 15

0.00 -1.31 -1.78 -2.10 -2.93 -2.94 -3.69 -4.65 -5.96 -7.56 -9.33 -11.48 -14.97 ±0.00 ±0.11 ±0.09 ±0.10 ±0.17 ±0.04 ±0.10 ±0.17 ±0.13 ±0.15 ±0.14 ±0.19 ±0.21 0.00 0.47 -0.40 -0.98 -1.95 -2.80 -4.07 -5.56 -7.10 -8.77 -10.71 -13.25 -16.92 ±0.00 ±0.13 ±0.09 ±0.10 ±0.21 ±0.17 ±0.18 ±0.29 ±0.31 ±0.34 ±0.39 ±0.56 ±0.68

0

1

2

3

4

5

6

7

8

roberta-zinc-480m

9

10 11 12

0.00 0.61 0.92 0.84 0.84 0.61 0.53 0.21 0.12 -0.20 -0.45 -0.83 -1.19 -1.59 -2.29 ±0.00 ±0.12 ±0.16 ±0.16 ±0.10 ±0.10 ±0.15 ±0.13 ±0.08 ±0.11 ±0.13 ±0.16 ±0.16 ±0.13 ±0.11 0.00 -0.13 0.07 0.03 -0.02 -0.09 -0.14 -0.16 -0.27 -0.31 -0.56 -0.71 -0.97 -1.27 -1.57 ±0.00 ±0.11 ±0.08 ±0.07 ±0.08 ±0.10 ±0.10 ±0.10 ±0.05 ±0.09 ±0.05 ±0.08 ±0.05 ±0.10 ±0.10 0.00 0.26 0.43 0.52 0.59 0.54 0.39 0.24 0.16 -0.03 -0.27 -0.42 -0.58 -0.89 -1.21 ±0.00 ±0.12 ±0.09 ±0.14 ±0.21 ±0.15 ±0.10 ±0.13 ±0.14 ±0.12 ±0.19 ±0.13 ±0.18 ±0.18 ±0.14

0

Relative Layer Depth

1

2

3

4

5

6

7

8

9 10 11 12 13 14

0 1 2

Figure 9: Average differences in probing performance of PT (left) and RI (right) models after fine-tuning on solubility (ESOL). We group into hydrophilic (top), lipophilic (middle), and other (bottom) groups (see Table 7). The number of substructures in the corresponding group is indicated next to it. The RI results for chemberta-2 are based on a single model (hence, are the same for all three variants) as all models use the same architecture.

PT<RI In Section 5.3, we saw that pretraining degraded probing performance on substructures such as thiophene, thiazole and furan across all models. We hypothesize that the models were undertrained on these substructures. While we cannot conduct a frequency analysis of these substructures due to the unavailability of the pre-training data, further pre-training models on molecules that contain these molecular substructures should allow models to somewhat recover from the low performance. For this experiment, we select three models—chemberta-2-5M (3 layers), chemberta (6 layers) and molformer (12 layers)—each representing a small, mid-sized, and large model.

Phenols on the other hand have a moderate effect on solubility due to the presence of a lipophilic aromatic ring and a hydrophilic hydroxyl group. In summary, after fine-tuning on solubility, we observe improvements on hydrophilic substructures known to positively contribute to the solubility of a molecule. For decreased performance, we do not observe such systematic patterns. While in general there are no consistent trends for RI models, fine-tuning on solubility prediction leads to an improvement of RI models on NO, Heteroatoms, HAcceptors in lower layers. While this might suggest that RI models might be also capable of capturing task-relevant substructures, we do not observe such a trend for lipophilicity prediction.

H

Chemberta-2 and Halogens Our analysis in Section 5.3 showed that chemberta-2-5M and chemberta-2-10M improved on halogens after pretraining, whereas performance degraded for all other models (which generally showed a very high probing performance). Moreover, in Section 6.4 we found that, unlike the other models, fine-tuning on lipophilicity led to further improvements on halogens for chemberta-2-5M and chemberta-2-10M. We thus hypothesize the following:

Practical Implications: Further Pre-Training

Our probing experiments have allowed us to identify three interesting patterns. First, we saw that pre-training leads to worse encodings of thiophene, thiazole and furan in all pre-trained models compared to the randomly initialized ones. Second, all chemberta-2 models use the same architecture and exhibit consistent improvements in probing performance on halogens while we observe the reverse for all other models. Finally, we observed diverging patterns in substructure encoding for chemberta-2-5M and chemberta-2-10M which differ only in terms of pre-training data. We conjecture that these differences might stem from different compositions of the pre-training data and conduct experiments to see if further pre-training models on specific datasets can mitigate the low probing performance.

1. Further pre-training on halogen-containing molecules will improve chemberta-2-5M and chemberta-2-10M’s performance on halogens. 2. As further pre-training improves probing performance for halogens, we expect smaller probing performance gains on halogens after fine-tuning on lipophilicity. Same architecture, different performance We conjecture that probing can furthermore be used to identify molecular substructures that a model has 23

seen less during pre-training and that further pretraining can be used to mitigate this gap. We study this hypothesis on two models with the same architecture (chemberta-2-5M and chemberta-2-10M) which exhibit large differences in probing performance on phenol (Figure 15, left). We conjecture that this is a result of the difference in the pretraining data (and its phenol distribution). H.1

ing performance on the corresponding substructure for chemberta-2-5M, chemberta and molformer. First, we observe an improvement in probing performance across all three substructures and models. While further pre-training benefits the chemberta model most (originally with a pretraining dataset of only 250k molecules), large improvements for molformer are generally observed in the upper layers. Compared to the other two models, chemberta-2-5M exhibits more variation in the magnitude of improvement. Second, further pre-training has a mixed effect on the downstream performance on lipophilicity prediction. Whereas molformer’s downstream performance is negatively affected in all three cases, chemberta-2-5M mostly benefits from further pretraining. In contrast, chemberta’s performance improves when further pre-trained on furan and is negligibly affected by pre-training on thiazole and thiophene. This may change under extensive hyperparameter tuning.

Pre-training Dataset

For further pre-training, we subsample from the Guacamol dataset (Brown et al., 2019), a dataset that is designed for benchmarking de novo molecular design. It comes with pre-defined train, validation and test splits with 1,273,104, 79,568 and 238,706 molecules, respectively. Pre-processing We first canonicalize all SMILES strings and discard 3,217 molecules which appear in the downstream lipophilicity and solubility datasets. We then annotate the canonicalized and cleaned training and validation splits of Guacamol with binary labels (where 1 denotes the presence of a substructure) using RDKIT. For each substructure of interest, we select SMILES strings containing this substructure from the training and validation sets. Having obtained the substructure-containing subsets, we randomly sample 50,000 and 3,000 molecules from them. H.2

Chemberta-2 and Halogens Figure 14 presents the results of further pre-training of models from the chemberta-2 family on halogen data. First, we observe slight performance improvements for chemberta-2-5/10M and substantial performance gains for chemberta-2-77M when probing for halogens. Interestingly, we find that the probing performance on halogens converges towards a joint upper bound which is also shared in the randomly initialized chemberta-2 model. This is in stark contrast to all other models shown in Figure 13 which have a substantially higher performance in the upper layers. One reason for this might be the tokenization issues of the chemberta-2 models which result in turning halogens such as Cl and Br into C and B. 11 Second, improved probing performance on halogens after further pre-training makes their encoding more robust, resulting in smaller changes after finetuning, detailed in Table 13. We conjecture that these changes might be more prominent without the tokenization issues. Finally, similar to further pre-training on furan, thiazole and thiophene on downstream performance, the results in Table 14 demonstrate that pre-training on more halogen data has a mixed effect on downstream performance. Whereas chemberta-2-5M improves on the downstream task and chemberta-2-10M’s performance is sub-

Experimental Setup

Pre-training We conducted pre-training for all models on the subset of data sampled from Guacamol for 20 epochs (15,625 steps) with a batch size of 64. We used the AdamW optimizer with the default settings of β1 = 0.9 and β2 = 0.999, ϵ = 1 × 10−8 and weight decay of 0.01. We use a scheduler with a linear learning rate decay of 0.1 and 100 warm-up steps. Probing Setup Our probing setup is identical to that summarized in Section D.1. Finetuning Setup For fine-tuning on lipophilicity prediction, we follow the setup described in Section F.1 and use the best hyperparameters for pre-trained models listed in Table 10. H.3

Results

PT<RI Figure 12 shows the impact of further pre-training on more data with specific substructures (furan, thiazole, thiophene) on the prob-

11

24

see Issue 1 and Issue 2 and Issue 3

I.1

stantially negatively affected, chemberta-2-77M remains largely unchanged. One reason for this disparity despite all models sharing the same architecture might be the different hyperparameters which were not tuned individually for further pre-trained models. Tables 10 and 11 show different optimal hyperparameters for different chemberta-2 models across both tasks. Hence, extensively tuning the hyperparameters for the further pre-trained models might mitigate drops in the performance.

To obtain the training data, we subsample the Guacamol (Brown et al., 2019) and ZINC-100M (Singh et al., 2026) datasets. Guacamol comes with pre-defined train, validation and test splits with 1,273,104, 79,568 and 238,706 molecules, respectively. Singh et al. (2026) curated ZINC-100M by sampling 100M molecules from the ZINC20 database (Irwin et al., 2020). I.2

Same architecture, different performance Figure 15 shows the effect of pre-training chemberta-2-5M and chemberta-2-10M on more data containing phenol substructures. Although further pre-training on data containing phenol improved the respective probing performance, the performance gap between the two models persisted. We conjecture that the difference in the amount of original pre-training data (5 million vs 10 million) might be responsible for this and that longer pre-training might further reduce this gap. Again, Table 15 shows that further pretraining can both negatively and positively affect downstream performance. H.4

Data Preprocessing and Sampling

Guacamol We first canonicalize all SMILES strings. We then discard any molecules which appear in the downstream lipophilicity and solubility datasets. In the case of Guacamol, we discard 3,217 molecules. We then randomly sample 100,000 and 10,000 molecules from the training and validation sets, respectively. ZINC-100M Due to the large size of ZINC100M (∼7.7GB), we first performed random sampling on the byte positions, obtaining 100,000 and 10,000 molecules for the training and validation sets, respectively. We then canonicalize the SMILES strings and check for SMILES strings which overlap with the those in the fine-tuning datasets.

Discussion

I.3

Our results suggest that substructure probing can serve as a diagnostic tool to identify whether a model was undertrained on specific substructures. With the additional pre-training experiments, we have demonstrated that further pre-training on infrequent substructures can mitigate the negative effects of undertraining to some extent and increase robustness during fine-tuning. Based on our findings, systematically studying the pre-training dynamics of CLMs via probing could be a promising endeavor for future work.

I

Datasets

Pre-Training Setup

We follow the same pre-training setup described in Section H.2. However, we train the model for 32 epochs (50k steps). I.4

Effect of Further Pre-Training

Figure 16 shows the effect of further pre-training chemberta-3 on subsets of Guacamol and ZINC100M. Interestingly, the model further pre-trained on Guacamol exhibits substantially better probing performance compared to the model trained on ZINC across all molecular substructures. Further pre-training on ZINC benefits other (non-ring) substructures more than rings, with improved performance mostly in upper layers. Importantly, the lower layers which are characterized by substantially lower performance remain mostly unaffected by further pre-training on ZINC while showing a slight improvement for Guacamol. To better understand these effects, we analyze the frequency distributions of molecular substructures in Figure 17 across both datasets, finding that they are mostly similar. Surprisingly, the actual overlap

The Peculiar Case of Chemberta-3

Chemberta-3 exhibits a peculiar "dome" in the average probing performance at lower layers (Figure 1). We hypothesize that this might stem from the irregularities during the pre-training process. To investigate this as well as the impact of different training data, we further pre-train chemberta-3 on subsets of the Guacamol (Brown et al., 2019) (CC BY-SA 3.0) and ZINC-100M (Singh et al., 2026) datasets. 25

in terms of molecules (i.e., SMILES string) is almost zero with only four overlapping strings in the training split. This suggest that small differences in the training frequency of molecular substructures may already substantially affect the encodings during pre-training. Furthermore, there might be other effects beyond molecular substructures that need to be studied in future work. Nonetheless, we conclude that further pretraining a model on a small, carefully curated dataset can already mitigate some of the negative probing performances.

J

from which we use the lipophilicity and solubility datasets—and evaluate four different vector sizes dim∈{300, 512, 1024, 2048} For each of the extracted fingerprints, we then train a logistic regression model (LR), a support vector machine (SVM), and a gradient boosted tree (XGB). For the SVM, we additionally tune different values for c={0.00001, 0.0001, 0.001, 0.01, 0.1, 1, 4, 16, 64, 256, 1024} using the validation set. Results. Table 16 and Table 17 show the performance (RMSE) of different fingerprinting algorithms across different models for lipophilicity and solubility, respectively. Most notably, the SVM consistently performs best across all dimensions, fingerprints, and tasks, showing the best performance for atom pairs fingerprinting (ATFP) with a dimension of 2048 and a c of 16. Interestingly, we find that Morgan fingerprints (MFP) perform most stable across different dimensions and best on average. Finally, we find that for lipophilicity prediction the SVM and XGB both benefit from higher dimensions, while this is the opposite for the LR model. Similarly, the SVM and XGB both perform robustly across all dimensions and fingerprints, while LR performance varies a lot; especially for higher dimensions.

Lipophilicity and Solubility Prediction with Other Models

We provide additional experiments on the downstream tasks using models other than CLMs. More specifically, we investigate models using traditional chemical features, referred to as “molecular fingerprints” and more recent, decoder-only foundation models. J.1

Models Trained on Fingerprints

Fingerprints are a common way of representing molecules in chemistry to be used as features in machine learning models. They are often generated using hand-crafted algorithms that output a binarized feature vector representing the presence of specific molecular substructures in a molecule. One of the most famous methods is Morgan’s algorithm (Morgan, 1965), with various other methods that have been developed over time. In this work, we evaluate four different types of fingerprints:

J.2

Foundation Models

Recently, an increasing number of works have investigated the potential of foundation models for chemistry tasks (Guo et al., 2023, 2024). While they find that some do perform decently on MPP tasks, they primarily focus on classification tasks. In this work, we provide complementary results on two regression tasks (i.e., lipophilicity and solubility prediction).

MFP Morgan fingerprints as introduced by Morgan (1965). RDFP The fingerprinting method provided in RDKit (Landrum et al., 2024).

Experimental setup. We conduct experiments for two large language models, namely Llama-3.2-3B-Instruct (Grattafiori et al., 2024) and gpt-oss-20B (OpenAI, 2025). We prompt each model in a zero-shot setting, asking it to predict the logD (or logS) value of the given molecule (basic). We further evaluate three additional setups to accommodate for the model’s lack of chemical knowledge. First, we provide explanations on which functional groups increase and decrease the logD (or logS) value (expl). Second, we provide a list of extracted molecular substructures present in the molecule (hint). Finally, we pass both piece of information to the model (both). For generation, we use nucleus

ATFP Atom pair fingerprints introduced by (Carhart et al., 1985). In the RDKit implementation, an atom is represented by a tuple of the atomic number, number of pi electrons, and the degree of the atom, with the option of adding chirality information. TTFP Topoligical torsion fingerprints are similar to atom pair fingerprints but use 4-atom sequences to capture more local graph structure. (Nilakantan et al., 1987). We extract fingerprints using a radius of 2, which is equivalent to the same value used in MoleculeNet (Wu et al., 2018)—the benchmark 26

sampling (Holtzman et al., 2020) with p = 0.95, a temperature of 0.8 and a maximum token budget of 4,096 tokens. We evaluate all three reasoning levels for the gpt-oss-20B model, i.e., low, medium, and high. For the high reasoning level, we find that the model tends to generate very lengthy responses that exceed the token budget. We thus further evaluate token budgets of 8,192 and 16,384. All experiments were conducted on a high performance computing cluster with 4 NVIDIA A100 (40GB). Each GPU ran for ≈112 hours (∼4.7 days in total).

(“This is the molecule”) shifted accordingly. The setting both combines all three templates (one of the basic templates, chemical explanations, and preprocessed functional groups). Basic Prompt (lipophilicity) System: You are now working as an excellent expert in chemistry and drug discovery. User: Predict the lipophilicity the following molecule. The lipohilicity of a molecule is defined by the logD value, the logarithmic form of the distribution coefficient D. The molecule is represented using the canonical form of its SMILES string.

Results. Table 18 and Table 19 provide the results of our experiments for lipophilicity prediction and solubility prediction, respectively. Overall, we can see that all LLMs perform worse compared to the models that use CLMs or fingerprints. We further find that a high reasoning level does not gpt-oss-20b automatically lead to an improve performance but instead, can produce responses that exceed the token budget (as it consistently happens for the high reasoning level). Interestingly, providing the models with either explanations or information about the present functional groups does improve their performance, however providing both leads to a worse performance for lipophilicity prediction. This follows the findings by Ganeeva et al. (2024, 2025) who find that LLMs might still be lacking in terms of compositionality (as they seem incapable of putting together the provided explanation and the functional groups). In contrast, we do not consistently observe this behavior for solubility prediction, as providing both even leads to the best result (for gpt-oss-20b4,096 , low). However, we also see that the performance of smaller models or lower and medium reasoning levels is substantially worse for solubility prediction compared to lipophilicity prediction. Figure 18 showcases an example response for the best performing LLM (gpt-oss-20b4 , 096, with hints and medium reasoning level). As can be seen, the model seemingly “reasons” about the task, but assigns the wrong sign to the predicted logD value, indicating that it does not have actual knowledge about the task.

This is the molecule: {canonicalized SMILES string}

Basic Prompt (solubility) System: You are now working as an excellent expert in chemistry and drug discovery. User: Predict the solubility of the following molecule. The solubility of a molecule is measured by the log solubility in mols per liter. The molecule is represented using the canonical form of its SMILES string. This is the molecule: {canonicalized SMILES string}

+ Chemical Explanations: For prediction it is also important to consider the following functional groups that affect {lipophilicity, solubility}. These are functional groups that increase {lipophilicity, solubility}: {List of groups that increase logD or logS value} These are functional groups that decrease {lipophilicity, solubility}: {List of groups that decrease logD or logS value}

Prompt templates. In the following, we provide all prompt templates that we used in our experiments. Note, that the basic prompt (we provide individual templates for lipophilicity and solubility prediction) is always present, and that the respective sub-prompts are appended accordingly. We always provide the molecule last with the prefix

This is the molecule: {canonicalized SMILES string}

27

+ Preprocessed Functional Groups: As a hint, you are also given the functional groups that are present in the molecule. Following functional groups are present in the molecule: {all present groups} This is the molecule: {canonicalized SMILES string}

28

Probing task

NHOH NO AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticRings AromaticCarbocycles AromaticHeterocycles HAcceptors HDonors Heteroatoms RotatableBonds SaturatedRings SaturatedCarbocycles SaturatedHeterocycles Ring Al_COO Al_OH Al_OH_noTert ArN Ar_COO Ar_N Ar_NH Ar_OH COO COO2 C_O C_O_noCOO C_S Imine ketone NH0 NH1 NH2 N_O Ndealkylation1 Ndealkylation2 Nhpyrrole SH

Training Split Size Seed 42

Seed 77

Seed 4

54574 5308 44174 66372 99440 76818 75432 65176 4662 53418 2076 15078 70664 34032 42976 23322 14092 52366 48914 10872 3992 56262 10330 8548 18034 19182 68424 52970 870 28506 15014 75328 70862 39752 3462 14322 13014 10330 3550

54936 5374 43876 66602 99306 77146 75536 64620 4778 53890 2060 15180 70306 33544 43060 23484 14152 52306 48874 10808 3956 55878 10088 8220 18050 19154 68080 52798 812 28680 15276 75390 70678 39668 3568 14354 13062 10088 3444

54628 5454 44168 66996 99924 77070 76012 64144 4826 53548 2156 15118 71046 33866 43562 23270 13996 52222 48840 11102 3878 55370 9820 8362 17816 18944 68792 53526 792 28360 15558 76000 71058 40030 3484 14098 13032 9820 3436

Probing task

aldehyde alkyl_halide allylic_oxid amide amidine aniline aryl_methyl benzene bicyclic ester ether furan guanido halogen hdrzine hdrzone imidazole imide ketone_Topliss lactone methoxy morpholine nitrile nitro nitro_arom nitro_arom_nonortho oxime para_hydroxylation phenol phenol_noOrthoHbond piperdine piperzine pyridine priamide sulfide thiazole thiophene unbrch_alkane urea

Training Split Size Seed 42

Seed 77

Seed 4

4128 12224 19950 19672 9086 27334 42070 75432 35830 15866 63050 5358 2322 35422 3384 1754 7152 1254 13072 1930 22138 2316 11304 4444 2410 1182 1570 12858 5432 5354 10696 3064 18140 576 9854 3874 5610 12118 758

4054 12168 20470 19528 9024 27340 41994 75530 35808 15856 63536 5272 2364 35528 3348 1792 7336 1166 13170 1940 22602 2170 11414 4384 2418 1200 1476 12690 5382 5284 10720 3168 18244 556 9744 3852 5394 12130 752

4228 12022 20198 19746 9054 27794 41350 76012 35426 15890 63826 5252 2346 35384 3280 1944 7074 1206 13448 2090 22642 2256 11520 4396 2368 1116 1556 12736 5422 5306 10628 3176 18118 572 9946 3782 5370 11878 768

Table 3: Probing training dataset statistics. Each row shows the size of the undersampled training split for each probing task across three random samples. Some substructures occur infrequently in the original PCQM4Mv2 dataset, resulting in smaller training sets for certain tasks (e.g., priamide, urea).

29

Probing task

NHOH NO AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticRings AromaticCarbocycles AromaticHeterocycles HAcceptors HDonors Heteroatoms RotatableBonds SaturatedRings SaturatedCarbocycles SaturatedHeterocycles Ring Al_COO Al_OH Al_OH_noTert ArN Ar_COO Ar_N Ar_NH Ar_OH COO COO2 C_O C_O_noCOO C_S Imine ketone NH0 NH1 NH2 N_O Ndealkylation1 Ndealkylation2 Nhpyrrole SH

% of molecules w/substructure Seed 42

Seed 77

Seed 4

61.20 95.38 17.38 29.50 42.71 64.66 49.26 27.84 95.94 62.17 97.82 88.03 26.55 11.27 17.39 85.53 7.17 21.76 20.09 4.68 2.06 23.54 5.59 5.36 9.20 9.24 42.93 35.73 1.06 12.85 10.04 57.19 28.09 12.85 1.97 3.51 3.62 5.59 3.23

61.51 95.48 17.34 29.54 42.77 64.69 48.84 28.32 96.14 62.59 97.92 88.13 26.68 11.21 17.46 85.84 7.14 21.98 20.36 4.50 2.09 23.74 5.83 5.11 9.19 9.24 43.13 36.13 1.01 12.77 9.84 57.01 28.98 12.63 1.95 3.48 3.61 5.83 3.26

60.86 95.42 17.29 30.05 43.22 64.67 49.48 27.61 96.03 61.91 97.84 87.86 26.68 11.04 17.59 86.00 7.42 21.84 20.27 4.27 2.11 22.96 5.48 4.88 9.48 9.54 43.42 36.20 1.01 13.32 9.61 57.49 28.75 12.24 2.06 3.54 3.87 5.48 3.12

Probing task

aldehyde alkyl_halide allylic_oxid amide amidine aniline aryl_methyl benzene bicyclic ester ether furan guanido halogen hdrzine hdrzone imidazole imide ketone_Topliss lactone methoxy morpholine nitrile nitro nitro_arom nitro_arom_nonortho oxime para_hydroxylation phenol phenol_noOrthoHbond piperdine piperzine pyridine priamide sulfide thiazole thiophene unbrch_alkane urea

% of molecules w/substructure Seed 42

Seed 77

Seed 4

2.04 6.27 12.45 14.00 5.04 13.87 17.68 49.20 24.58 10.78 30.86 2.37 1.56 18.18 2.75 3.01 3.17 1.52 7.95 1.38 10.91 1.12 5.19 4.65 2.81 1.77 1.71 10.64 3.87 3.75 3.48 1.17 7.32 1.42 6.33 1.45 2.36 7.11 1.76

2.02 6.12 12.53 14.27 5.41 13.58 17.71 48.78 24.25 11.16 31.51 2.66 1.64 17.68 2.65 3.13 3.29 1.57 7.81 1.44 11.30 1.18 5.31 4.46 2.69 1.71 1.76 10.49 3.75 3.65 3.38 1.23 7.14 1.49 6.22 1.55 2.44 7.14 1.70

2.00 6.12 12.56 14.69 5.44 13.42 18.09 49.43 24.30 11.09 31.12 2.63 1.59 17.57 2.61 3.33 3.16 1.62 7.54 1.39 11.19 1.14 5.36 4.46 2.75 1.71 1.74 10.81 3.58 3.50 3.58 1.21 6.98 1.47 6.25 1.46 2.44 7.00 1.71

Table 4: Class distributions in the probing test set. Each row shows the percentage of molecules containing a molecular substructure probed for in each of the three randomly drawn test samples.

Molecular Substructure AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticRings AromaticCarbocycles AromaticHeterocycles SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene bicyclic Ring Table 5: List of the 12 ring types used for computing the average probing performance for rings shown in Figure 1.

30

lipophilic hydrophilic other

Molecular Substructure

% of molecules w/substructure train val test

AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen pyridine thiazole thiophene

14.79 48.87 56.58 89.70 73.42 99.02 12.59 40.27 47.74 89.70 43.51 29.64 6.31 6.04

16.67 44.29 55.24 90.00 76.43 98.33 14.05 35.48 45.48 90.00 40.00 28.57 6.43 4.76

19.29 48.33 61.43 90.48 74.05 98.33 17.14 40.00 52.86 90.48 40.95 29.29 7.14 6.90

NHOH NO Heteroatoms HAcceptors HDonors Al_COO Al_OH Al_OH_noTert Ar_COO Ar_NH Ar_OH COO COO2 NH0 NH1 NH2 amide aniline ether morpholine para_hydroxylation phenol phenol_noOrthoHbond piperdine piperzine priamide

86.37 99.97 99.97 99.94 86.37 7.71 13.30 11.19 3.36 16.85 8.66 11.07 11.07 87.68 65.48 19.55 51.25 53.01 45.92 6.22 16.93 6.70 6.58 18.21 9.64 4.76

88.57 100.00 100.00 100.00 88.81 5.48 16.43 14.05 5.71 19.05 11.19 11.19 11.19 83.81 68.57 20.71 49.29 58.33 44.05 7.38 18.81 8.81 8.81 15.48 8.81 4.29

85.00 100.00 100.00 100.00 85.00 6.90 10.00 8.81 4.29 19.52 7.86 10.95 10.95 84.76 66.90 20.71 52.38 53.33 42.38 6.90 21.19 6.67 6.19 17.62 10.00 6.90

RotatableBonds Ring ArN Ar_N C_O C_O_noCOO Imine Ndealkylation1 Ndealkylation2 Nhpyrrole alkyl_halide aryl_methyl bicyclic imidazole ketone ketone_Topliss methoxy nitrile sulfide urea

97.65 99.94 9.97 69.43 63.15 56.43 1.93 9.43 16.37 16.85 8.39 25.77 53.57 9.26 4.85 4.08 16.88 7.83 4.97 4.35

96.19 99.76 12.38 71.90 61.43 55.71 1.67 7.14 11.19 19.05 6.19 27.14 54.29 8.10 6.90 5.71 15.48 7.62 6.90 4.05

96.90 99.52 8.10 69.05 63.57 57.38 2.62 11.67 14.29 19.52 7.14 25.00 53.33 7.86 7.38 6.43 17.62 6.67 4.76 3.81

Table 6: Molecular substructure frequencies in the lipophilicity dataset: Percentage of molecules in the training, validation and test splits of the lipophilicity dataset containing each of the 60 specific molecular substructures.

31

lipophilic hydrophilic other

Molecular Substructure

% of molecules w/substructure train val test

AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen

12.99 14.65 25.08 49.72 15.32 58.60 9.66 9.32 17.20 49.72 29.41

13.27 9.73 22.12 48.67 10.62 55.75 8.85 7.96 16.81 48.67 30.97

15.04 16.81 26.55 53.98 19.47 61.95 12.39 12.39 21.24 53.98 27.43

NHOH NO Heteroatoms HAcceptors HDonors Al_OH Al_OH_noTert Ar_OH NH0 NH1 amide aniline ester ether para_hydroxylation phenol phenol_noOrthoHbond

43.17 71.37 86.13 72.25 43.62 13.98 11.54 7.88 26.75 18.87 18.87 15.76 8.88 21.31 10.43 6.99 6.88

34.51 67.26 84.07 68.14 34.51 15.93 12.39 4.42 27.43 14.16 15.93 10.62 7.08 14.16 7.96 4.42 4.42

53.98 83.19 89.38 83.19 53.98 11.50 9.73 13.27 33.63 30.97 28.32 23.89 8.85 22.12 15.04 11.50 11.50

RotatableBonds Ring Ar_N C_O C_O_noCOO allylic_oxid aryl_methyl bicyclic imide unbrch_alkane urea

68.04 71.81 14.21 34.63 34.63 11.43 12.32 22.09 4.99 10.65 6.44

67.26 69.03 8.85 31.86 31.86 12.39 14.16 18.58 7.08 12.39 6.19

68.14 75.22 15.93 43.36 43.36 12.39 15.04 28.32 11.50 9.73 12.39

Table 7: Molecular substructure frequencies in the ESOL dataset: Percentage of molecules in the training, validation and test splits of the ESOL dataset containing each of the 39 specific molecular substructure.

32

mean std min 25% 50% 75% max

train

logD valid

test

0.000000 1.000149 -3.091151 -0.632465 0.145283 0.755773 1.926575

-0.078882 1.043810 -3.024248 -0.718184 0.061654 0.707686 1.926575

-0.004254 1.014424 -3.024248 -0.665917 0.157827 0.755773 1.926575

mean std min 25% 50% 75% max

train

logS valid

test

-3.047764 2.076592 -11.600000 -4.300000 -2.900000 -1.614000 1.580000

-3.186230 2.294987 -8.710000 -4.570000 -3.000000 -1.456000 1.110000

-2.919593 2.061854 -9.332000 -3.955000 -2.630000 -1.600000 1.100000

Table 8: Summary statistics for the distribution of predicted values – logD (lipophilicity) and logS (solubility)– in the training, validation and test splits of the lipophilicity and solubility datasets, respectively.

Task

Lipo ESOL

Regression Head

LR LR, 2L-MLP

Batch size

Learning rate

# Epochs

PT

RI

[8, 16, 32, 64] [8, 16, 32, 64]

[0.01, 1 × 10−3 , 1 × 10−4 , 5 × 10−4 , 1 × 10−5 , 5 × 10−5 ] [0.01, 1 × 10−3 , 1 × 10−4 , 5 × 10−4 , 1 × 10−5 , 5 × 10−5 ]

10 20

20 20

Table 9: Hyperparameter ranges explored. For randomly initialized models, we did not perform extensive hyperparameter tuning, instead adopting the optimal batch size and learning rates from their pre-trained counterparts. However, we increased the number of epochs. For ESOL, we evaluate both, linear regression (LR) and 2-layer MLP (2L-MLP) heads.

Model

Batch size

Learning rate

Best Epoch

PT

RI

PT

RI

PT

RI

chemberta-base chemberta

16 16

16 16

1 × 10−4 1 × 10−5

1 × 10−4 1 × 10−5

8 9

18 10

chemberta-2-5M chemberta-2-10M chemberta-2-77M

8 32 32

32* 32* 32*

5 × 10−5 5 × 10−4 5 × 10−4

5 × 10−4 * 5 × 10−4 * 5 × 10−4 *

8 10 7

20* 20* 20*

molformer roberta-zinc-480m chemberta-3

16 32 64

16 32 64

1 × 10−4 1 × 10−5 1 × 10−4

1 × 10−4 1 × 10−5 1 × 10−4

10 7 10

20 17 10

Table 10: Final hyperparameter values for pre-trained and randomly initialized models fine-tuned for lipophilicity. Both pre-trained and ranodmly initalized models share the hyperparameters except for the best epoch. Note that chemberta-2-5M, chemberta-2-10M, chemberta-2-77M models share the same architecture. Thus, we fine-tune a single randomly initialized model for all of them (chemberta-2), selecting the average optimal hyperparameters across all three pre-trained models (*).

33

Model

Head type

Batch size

Learning rate

Best Epoch

PT

RI

PT

RI

PT

RI

chemberta-base chemberta

linear linear

16 8

16 8

1 × 10−4 5 × 10−5

1 × 10−4 5 × 10−5

19 12

18 17

chemberta-2-5M chemberta-2-10M chemberta-2-77M

linear MLP linear

64 8 8

8* 8* 8*

0.001 5 × 10−4 5 × 10−4

5 × 10−4 * 5 × 10−4 * 5 × 10−4 *

19 20 6

19* 19* 19*

molformer roberta-zinc-480m chemberta-3

linear linear linear

16 32 32

16 32 32

5 × 10−5 5 × 10−5 1 × 10−4

5 × 10−5 5 × 10−5 1 × 10−4

18 16 19

19 20 20

Table 11: Final hyperparameter values for pre-trained and randomly initialized models fine-tuned on ESOL. Both pre-trained and ranodmly initalized models share the hyperparameters except for the best epoch. Note that chemberta-2-5M, chemberta-2-10M, chemberta-2-77M models share the same architecture. Thus, we fine-tune a single randomly initialized model for all of them (chemberta-2), selecting the average optimal hyperparameters across all three pre-trained models (*).

Model

molformer chemberta chemberta-2-5M

RMSE furan

thiazole

thiophene

0.596 (+0.031) 0.641 (-0.034) 0.661 (-0.003)

0.597 (+0.032) 0.669 (-0.006) 0.631 (-0.033)

0.594 (+0.029) 0.667 (-0.008) 0.646 (-0.018)

Table 12: Fine-tuning performance (RMSE, lower is better) of models further pre-trained on data containing furan, thiazole, thiophene on the lipophilicity dataset. green indicates improvement in RMSE (lower is better) while red denotes increased RMSE. Gray indicates negligible changes (<0.01) in RMSE.

Model

Layer

chemberta-2-5M

0 1 2 3 0 1 2 3 0 1 2 3

chemberta-2-10M

chemberta-2-77M

∆ macro F1 ∆ macro F1 (halogens) (PT) 0.000 -0.209 0.015 2.955 0.000 -1.252 -1.395 2.407 0.000 -1.062 -0.419 4.381

0.000 -0.700 0.499 0.573 0.000 -2.341 -0.212 -1.257 0.000 -2.606 -1.297 -2.869

Table 13: Further pretraining on halogens: Layer-wise difference in probing performance on halogens for chemberta-2 models (left) and their counterparts further pre-trained on halogens (right) after fine-tuning.

34

0.00 ±0.00

-0.24 ±0.53

-0.23 ±0.33

0.07 ±0.20

-0.36 ±0.39

0.34 ±0.24

0.53 ±0.19

-0.70 ±0.18

-0.58 ±0.19

0.67 ±0.16

0.58 ±0.43

0.95 ±0.48

1.36 ±0.20

0.00 ±0.00

0.13 ±0.27

-0.09 ±0.08

0.16 ±0.17

0.28 ±0.19

0.26 ±0.10

0.64 ±0.16

0.05 ±0.16

-0.61 ±0.27

-0.58 ±0.17

-0.00 ±0.36

0.40 ±0.12

-0.09 ±0.13

0.00 ±0.00

-0.06 ±0.16

0.00 ±0.02

-0.12 ±0.06

-0.14 ±0.02

-0.14 ±0.08

-0.14 ±0.12

-0.30 ±0.06

-0.39 ±0.05

-0.61 ±0.04

-0.46 ±0.10

-0.25 ±0.12

-0.34 ±0.01

0.00 ±0.00

0.00 ±0.16

-0.00 ±0.07

-0.09 ±0.07

0.05 ±0.07

0.22 ±0.15

0.23 ±0.19

0.17 ±0.31

0.27 ±0.21

0.51 ±0.22

0.83 ±0.10

1.24 ±0.15

0.44 ±0.06

0.00 ±0.00

0.16 ±0.08

-0.00 ±0.10

0.06 ±0.04

-0.06 ±0.04

-0.32 ±0.14

-1.03 ±0.34

-1.04 ±0.07

-0.38 ±0.15

-0.65 ±0.05

-0.19 ±0.20

-0.25 ±0.10

-0.25 ±0.26

0.00 ±0.00

0.00 ±0.01

0.03 ±0.04

0.03 ±0.06

-0.15 ±0.05

-0.09 ±0.03

-0.07 ±0.12

-0.07 ±0.11

0.14 ±0.07

0.29 ±0.16

0.31 ±0.10

0.61 ±0.12

0.49 ±0.04

0.00 ±0.00

-0.18 ±0.27

-0.29 ±0.49

-0.38 ±0.11

-0.71 ±0.14

0.56 ±0.10

0.33 ±0.25

-0.30 ±0.36

-0.10 ±0.53

0.73 ±0.43

0.44 ±0.07

1.07 ±0.44

0.87 ±0.31

0.00 ±0.00

0.01 ±0.09

-0.45 ±0.08

0.32 ±0.31

0.68 ±0.55

-0.06 ±0.31

-0.00 ±0.21

-0.44 ±0.34

-0.45 ±0.13

-0.71 ±0.12

0.63 ±0.07

1.26 ±0.10

0.13 ±0.39

0.00 ±0.00

-0.28 ±0.36

-0.17 ±0.21

-0.32 ±0.19

0.18 ±0.20

0.18 ±0.17

-0.24 ±0.24

-0.30 ±0.18

-0.30 ±0.09

-0.70 ±0.31

0.21 ±0.11

0.16 ±0.12

-0.19 ±0.11

0.00 ±0.00

0.02 ±0.10

-0.04 ±0.03

-0.12 ±0.04

0.08 ±0.09

0.22 ±0.16

0.20 ±0.09

0.22 ±0.25

0.27 ±0.19

0.49 ±0.20

0.82 ±0.06

1.17 ±0.13

0.42 ±0.17

0.00 ±0.00

0.09 ±0.45

0.28 ±0.05

0.10 ±0.09

-0.65 ±0.12

-0.53 ±0.14

-1.56 ±0.22

-1.41 ±0.19

-0.04 ±0.16

0.56 ±0.43

2.74 ±0.09

0.14 ±0.17

0.08 ±0.06

0.00 ±0.00

-0.45 ±0.11

-0.16 ±0.63

0.05 ±0.26

-1.72 ±0.45

-0.26 ±0.44

-0.10 ±0.36

0.33 ±0.77

0.67 ±0.44

1.15 ±0.52

1.20 ±0.46

1.39 ±0.41

0.09 ±0.55

0.00 ±0.00

0.35 ±0.25

1.59 ±0.63

-1.10 ±0.99

-1.13 ±1.67

-2.03 ±0.75

-2.74 ±0.79

-2.36 ±1.67

-0.90 ±1.17

-1.78 ±0.63

-1.31 ±0.48

-0.86 ±1.56

-0.83 ±0.90

0.00 ±0.00

0.59 ±1.06

0.26 ±1.40

-2.54 ±0.54

-4.77 ±0.21

-6.13 ±0.81

-5.92 ±0.46

-3.95 ±0.63

-1.65 ±0.85

-1.28 ±0.91

-1.56 ±1.66

-2.92 ±1.80

-3.74 ±0.87

0.00 ±0.00

0.26 ±0.13

0.15 ±0.25

0.52 ±0.30

-0.38 ±0.12

0.33 ±0.27

0.45 ±0.26

-0.85 ±0.45

-0.54 ±0.14

-0.53 ±0.45

-0.17 ±0.25

-0.31 ±0.11

-0.92 ±0.17

0.00 ±0.00

0.71 ±1.01

2.03 ±1.16

1.23 ±0.78

0.15 ±0.52

0.53 ±0.67

-0.43 ±1.38

-2.05 ±0.53

-2.51 ±0.14

-1.43 ±0.81

-1.56 ±0.66

-0.95 ±0.13

-3.10 ±1.38

0.00 ±0.00

0.21 ±0.15

0.02 ±0.15

0.59 ±0.47

-0.50 ±0.28

0.31 ±0.34

0.34 ±0.38

-0.83 ±0.46

-0.48 ±0.31

-0.87 ±0.52

0.04 ±0.08

-0.75 ±0.47

-1.20 ±0.14

0.00 ±0.00

0.19 ±0.10

-0.21 ±0.55

-0.72 ±0.62

-0.51 ±0.55

-0.60 ±0.08

-0.60 ±1.08

0.54 ±1.37

1.67 ±0.44

3.27 ±0.24

4.43 ±0.17

3.46 ±0.50

6.11 ±0.40

0.00 ±0.00

-0.55 ±0.41

-0.46 ±0.34

-0.26 ±0.34

-0.58 ±0.21

-0.25 ±0.30

0.40 ±0.32

-0.86 ±0.56

-1.41 ±0.47

-0.47 ±0.27

0.30 ±0.57

0.38 ±0.09

-0.21 ±0.52

0.00 ±0.00

-0.34 ±0.09

-0.43 ±0.76

-0.22 ±0.14

-0.21 ±0.28

-0.13 ±0.10

0.49 ±0.17

-0.63 ±0.44

-1.05 ±0.20

0.04 ±0.23

0.40 ±0.65

0.22 ±0.55

0.02 ±0.38

0

1

2

3

4

5

6

7

8

9

10

11

12

Ar_COO Ar_NH Ar_OH COO COO2 NH0 NH1 NH2 amide aniline ether morpholine para_hydroxylation phenol phenol_noOrthoHbond piperdine piperzine priamide NO Heteroatoms

0.00 ±0.00

-0.18 ±0.28

-0.42 ±0.16

-1.34 ±0.48

-0.42 ±0.51

-0.28 ±0.23

-0.75 ±0.37

1.09 ±0.23

1.13 ±0.38

1.96 ±0.84

3.29 ±0.79

1.89 ±1.36

6.77 ±1.23

0.00 ±0.00

-0.08 ±0.38

0.22 ±0.47

-0.16 ±0.17

-5.97 ±0.28

-4.76 ±1.08

-5.40 ±1.09

-4.39 ±0.43

-5.24 ±1.46

-5.12 ±0.29

-4.23 ±0.34

-5.10 ±0.22

-5.22 ±0.29

0.00 ±0.00

-0.29 ±0.44

-0.65 ±0.42

-0.62 ±0.73

-0.99 ±0.33

-1.72 ±0.31

-1.33 ±0.60

-0.92 ±0.25

-0.75 ±0.30

-0.63 ±0.58

-0.24 ±1.02

-3.42 ±0.78

-2.95 ±1.38

0.00 ±0.00

0.14 ±0.17

-0.16 ±0.64

-0.67 ±0.15

-0.64 ±0.16

-0.31 ±0.16

-0.35 ±0.20

0.67 ±0.72

0.76 ±0.80

3.25 ±1.03

4.53 ±0.39

3.12 ±0.36

3.92 ±0.60

0.00 ±0.00

0.37 ±0.08

-0.08 ±0.14

-0.38 ±0.13

-0.51 ±0.24

-0.23 ±0.36

-0.57 ±0.69

0.78 ±1.02

0.38 ±0.73

3.69 ±0.36

4.49 ±0.58

2.77 ±0.31

4.39 ±0.59

0.00 ±0.00

0.04 ±0.10

0.18 ±0.36

-0.19 ±0.26

0.08 ±0.16

0.14 ±0.43

-0.09 ±0.37

-0.46 ±0.24

-0.15 ±0.06

-0.27 ±0.24

-0.19 ±0.37

-0.16 ±0.18

-0.94 ±0.28

0.00 ±0.00

0.54 ±0.21

0.08 ±0.02

0.38 ±0.23

-0.80 ±0.41

-0.56 ±0.35

-1.30 ±0.15

-1.87 ±0.10

-1.47 ±0.38

-0.88 ±0.36

-0.77 ±0.20

-0.23 ±0.59

-0.50 ±0.29

0.00 ±0.00

0.34 ±0.13

0.27 ±0.53

0.29 ±0.14

0.06 ±0.45

-0.36 ±0.36

-0.85 ±0.72

-1.95 ±0.47

-2.26 ±0.05

-1.87 ±0.40

-1.52 ±0.31

-1.30 ±0.52

-1.69 ±0.90

0.00 ±0.00

-0.57 ±0.54

-0.08 ±0.37

-0.50 ±0.29

-0.92 ±0.15

0.60 ±0.30

0.34 ±0.32

-0.98 ±0.46

-1.53 ±0.11

-0.98 ±0.24

-1.12 ±0.24

-0.79 ±0.33

-2.24 ±0.11

0.00 ±0.00

0.25 ±0.33

-0.58 ±0.32

-1.04 ±0.45

-0.66 ±0.18

-0.84 ±0.07

-0.60 ±0.45

-0.78 ±0.17

-1.26 ±0.55

-1.47 ±0.16

-0.45 ±0.17

-0.51 ±0.27

-3.26 ±0.32

0.00 ±0.00

0.13 ±0.17

0.23 ±0.15

0.54 ±0.31

0.02 ±0.06

0.82 ±0.19

0.56 ±0.29

0.22 ±0.12

-0.22 ±0.18

-0.12 ±0.33

0.58 ±0.17

0.10 ±0.16

-1.27 ±0.40

0.00 ±0.00

0.20 ±0.20

-0.54 ±0.85

-0.51 ±0.42

0.01 ±0.38

0.34 ±0.75

-0.64 ±0.34

-0.39 ±0.76

0.71 ±0.13

1.07 ±0.61

1.54 ±0.93

1.59 ±1.09

0.61 ±0.78

0.00 ±0.00

-0.51 ±0.21

-0.71 ±0.31

-0.27 ±0.29

0.25 ±0.46

-0.17 ±0.76

0.14 ±0.48

-0.02 ±0.16

0.41 ±0.46

0.23 ±0.59

1.34 ±0.30

2.05 ±0.42

0.28 ±0.40

0.00 ±0.00

-0.32 ±0.46

-0.32 ±0.33

-0.77 ±0.24

-0.45 ±0.09

-1.15 ±0.41

-0.99 ±0.31

-0.05 ±0.73

0.61 ±0.57

-0.15 ±0.29

0.24 ±1.10

-3.96 ±0.22

-4.69 ±1.44

0.00 ±0.00

0.09 ±0.35

-0.13 ±0.48

0.01 ±0.52

-0.30 ±0.17

-0.94 ±0.39

-0.66 ±0.13

-0.31 ±0.54

-0.01 ±0.42

-0.93 ±0.26

-0.02 ±0.74

-4.34 ±0.32

-4.02 ±1.32

0.00 ±0.00

-0.42 ±0.55

-0.11 ±0.13

-0.79 ±0.44

0.18 ±0.44

0.56 ±0.14

0.07 ±0.54

0.20 ±0.58

0.41 ±0.72

-0.04 ±0.69

0.90 ±0.43

1.12 ±0.24

0.93 ±0.02

0.00 ±0.00

0.11 ±0.43

0.09 ±0.54

-0.19 ±0.39

0.37 ±0.31

1.04 ±1.91

-0.74 ±0.83

-1.18 ±0.64

-0.49 ±0.70

1.41 ±1.00

2.70 ±0.73

1.88 ±0.50

2.58 ±0.58

0.00 ±0.00

0.03 ±0.07

-0.87 ±0.77

-0.12 ±0.36

-0.11 ±0.73

-0.47 ±1.08

-0.11 ±0.37

-1.76 ±1.16

-1.36 ±2.08

-1.71 ±0.90

1.51 ±3.55

1.24 ±1.06

-0.52 ±2.44

0.00 ±0.00

0.33 ±0.87

2.52 ±0.66

1.94 ±0.44

1.87 ±0.91

1.46 ±0.91

1.30 ±1.07

1.19 ±1.20

0.07 ±0.36

1.00 ±0.36

0.21 ±1.06

0.22 ±0.63

-1.94 ±1.44

0.00 ±0.00

-0.19 ±1.46

1.51 ±0.60

1.20 ±0.79

-1.15 ±0.22

-1.08 ±0.42

-2.01 ±0.98

-2.66 ±0.29

-2.26 ±0.82

-2.39 ±0.18

-2.19 ±0.50

-2.40 ±0.28

-3.77 ±0.49

0

1

2

3

4

5

6

7

8

9

10

11

12

Layer

RotatableBonds Ring ArN Ar_N C_O C_O_noCOO Imine Ndealkylation1 Ndealkylation2 Nhpyrrole alkyl_halide aryl_methyl bicyclic imidazole ketone ketone_Topliss methoxy nitrile sulfide urea

0.00 ±0.00

0.38 ±0.47

0.23 ±0.48

0.40 ±0.22

-0.51 ±0.41

-0.60 ±0.20

-0.72 ±0.26

-0.38 ±0.39

-0.00 ±0.72

-0.56 ±0.71

0.51 ±0.12

0.74 ±0.36

0.17 ±0.60

0.00 ±0.00

0.00 ±0.01

0.01 ±0.00

0.01 ±0.00

-0.02 ±0.06

-0.07 ±0.03

0.05 ±0.03

-0.06 ±0.04

-0.21 ±0.15

-0.17 ±0.09

-0.40 ±0.17

-0.17 ±0.17

0.41 ±0.35

0.00 ±0.00

-0.19 ±0.22

-0.08 ±0.33

-0.57 ±0.39

-0.20 ±0.67

-2.95 ±0.43

-2.30 ±0.58

-2.63 ±0.44

-2.30 ±0.66

-4.35 ±0.70

-4.38 ±0.57

-7.27 ±1.09

-8.28 ±0.35

0.00 ±0.00

-0.10 ±0.11

0.00 ±0.07

0.11 ±0.08

-0.45 ±0.13

-0.66 ±0.12

-0.86 ±0.23

-0.99 ±0.15

-0.27 ±0.04

-0.66 ±0.07

-0.28 ±0.25

-0.51 ±0.16

-0.53 ±0.14

0.00 ±0.00

-0.13 ±0.23

-0.01 ±0.26

0.06 ±0.05

-0.31 ±0.14

0.34 ±0.40

-0.05 ±0.07

-0.06 ±0.06

0.21 ±0.30

0.63 ±0.27

0.37 ±0.08

-0.04 ±0.14

-0.86 ±0.30

0.00 ±0.00

-0.35 ±0.26

-0.09 ±0.33

0.06 ±0.07

0.12 ±0.01

0.81 ±0.46

-0.07 ±0.19

-0.07 ±0.09

-0.01 ±0.39

0.62 ±0.16

1.02 ±0.23

0.84 ±0.13

-0.22 ±0.18

0.00 ±0.00

-0.31 ±0.54

-0.59 ±0.23

-0.95 ±0.33

0.12 ±0.20

0.46 ±0.42

-0.18 ±0.41

-0.24 ±0.27

-0.32 ±0.36

0.22 ±0.31

-0.21 ±0.42

0.81 ±0.13

-1.24 ±0.23

0.00 ±0.00

-0.33 ±0.34

0.09 ±0.27

0.09 ±0.45

0.27 ±0.41

0.14 ±0.38

-0.28 ±0.13

-1.10 ±0.68

-1.38 ±0.59

-0.75 ±0.24

0.66 ±0.64

0.99 ±0.06

0.35 ±0.64

0.00 ±0.00

-0.14 ±0.70

-0.62 ±0.69

-0.33 ±0.41

0.03 ±0.40

-0.03 ±0.18

-0.34 ±0.76

-0.65 ±0.42

-0.44 ±0.16

0.01 ±0.59

0.99 ±0.45

0.86 ±0.60

0.85 ±0.16

0.00 ±0.00

-0.08 ±0.38

0.22 ±0.47

-0.16 ±0.17

-5.97 ±0.28

-4.76 ±1.08

-5.40 ±1.09

-4.39 ±0.43

-5.24 ±1.46

-5.12 ±0.29

-4.23 ±0.34

-5.10 ±0.22

-5.22 ±0.29

0.00 ±0.00

-0.10 ±0.48

0.37 ±0.25

0.40 ±0.06

-1.82 ±0.25

-1.32 ±0.19

-1.85 ±0.20

-1.49 ±0.61

-1.48 ±0.63

-1.23 ±0.81

0.91 ±0.69

0.21 ±0.34

0.27 ±0.93

0.00 ±0.00

-0.54 ±0.31

-0.58 ±0.49

0.18 ±0.55

-0.19 ±0.35

-0.45 ±0.27

-1.34 ±0.55

-0.76 ±0.10

-1.18 ±0.31

-1.81 ±0.12

-1.36 ±0.56

-0.39 ±0.43

-1.28 ±0.31

0.00 ±0.00

-0.36 ±0.17

-0.59 ±0.05

-0.41 ±0.07

-0.02 ±0.12

-0.14 ±0.13

-0.17 ±0.04

-0.21 ±0.10

-0.32 ±0.16

-0.56 ±0.27

-0.09 ±0.07

-0.88 ±0.21

-0.52 ±0.19

0.00 ±0.00

-0.28 ±0.10

-0.02 ±0.24

0.27 ±0.41

0.34 ±0.45

0.52 ±0.72

0.75 ±0.28

0.12 ±0.13

-0.01 ±0.48

-0.96 ±0.61

0.40 ±0.45

-0.59 ±0.93

-1.06 ±0.19

0.00 ±0.00

-0.41 ±0.26

-0.45 ±0.49

0.13 ±0.05

0.17 ±0.21

0.32 ±0.48

-0.24 ±0.58

-0.98 ±0.36

-0.43 ±0.22

-0.49 ±0.63

-0.19 ±0.49

-1.47 ±0.32

-4.40 ±0.48

0.00 ±0.00

-0.66 ±0.15

-0.40 ±0.25

-0.10 ±0.30

-0.35 ±0.42

0.41 ±0.33

-0.34 ±0.36

-1.40 ±0.34

-0.71 ±0.34

-0.31 ±0.31

-0.05 ±0.15

-1.36 ±0.58

-4.11 ±0.98

0.00 ±0.00

-0.73 ±0.04

0.11 ±0.47

0.11 ±0.90

-0.73 ±0.15

-0.39 ±0.17

-0.25 ±0.26

-0.04 ±0.72

0.24 ±0.28

0.03 ±0.28

0.55 ±0.13

1.19 ±0.29

0.73 ±0.41

0.00 ±0.00

-0.40 ±1.01

0.78 ±0.48

-0.33 ±0.31

-0.37 ±1.05

-0.20 ±0.57

1.73 ±0.64

1.87 ±0.21

0.93 ±0.09

1.81 ±0.77

1.27 ±0.49

0.69 ±0.62

0.72 ±0.37

0.00 ±0.00

0.22 ±0.30

-0.06 ±0.13

0.01 ±0.34

-0.37 ±0.49

-0.70 ±0.30

-0.68 ±0.61

-0.12 ±0.22

-0.60 ±0.87

-1.13 ±0.72

-0.16 ±0.55

-1.56 ±0.38

-4.65 ±0.48

0.00 ±0.00

-0.14 ±0.60

-0.25 ±0.96

0.57 ±0.27

0.85 ±1.02

-0.40 ±0.76

-0.30 ±0.69

-2.07 ±1.03

-1.33 ±1.41

0.06 ±1.19

-0.59 ±0.76

0.18 ±0.91

-1.35 ±1.02

0

1

2

3

4

5

6

7

8

9

10 11 12

10 5 0 5 10

Molecular Substructure

Molecular Substructure

AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen pyridine thiazole thiophene NHOH HAcceptors HDonors Al_COO Al_OH Al_OH_noTert

15

AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen pyridine thiazole thiophene NHOH HAcceptors HDonors Al_COO Al_OH Al_OH_noTert

0.00 ±0.00

-0.01 ±0.38

0.12 ±0.35

0.37 ±0.31

0.21 ±0.15

0.29 ±0.29

0.12 ±0.46

0.01 ±0.19

0.20 ±0.12

0.18 ±0.14

-0.17 ±0.20

-0.49 ±0.12

-1.10 ±0.10

0.00 ±0.00

-0.29 ±0.47

0.17 ±0.18

0.38 ±0.25

0.54 ±0.20

0.74 ±0.18

0.71 ±0.08

0.55 ±0.15

0.65 ±0.23

0.45 ±0.25

0.16 ±0.18

-0.03 ±0.12

-0.60 ±0.10

0.00 ±0.00

-0.66 ±0.17

-0.43 ±0.09

-0.21 ±0.12

-0.41 ±0.10

-0.44 ±0.08

-0.50 ±0.05

-0.58 ±0.18

-0.56 ±0.12

-0.73 ±0.09

-0.90 ±0.10

-1.12 ±0.07

-1.42 ±0.13

0.00 ±0.00

-0.39 ±0.09

-0.30 ±0.16

-0.40 ±0.14

-0.33 ±0.06

-0.33 ±0.04

-0.24 ±0.14

-0.29 ±0.02

-0.34 ±0.14

-0.38 ±0.16

-0.56 ±0.13

-0.81 ±0.15

-1.20 ±0.19

0.00 ±0.00

-0.22 ±0.14

-0.23 ±0.06

-0.33 ±0.07

-0.31 ±0.09

-0.42 ±0.04

-0.78 ±0.17

-1.09 ±0.12

-1.57 ±0.17

-2.44 ±0.22

-3.62 ±0.02

-5.07 ±0.16

-6.77 ±0.16

0.00 ±0.00

0.03 ±0.04

0.01 ±0.06

-0.01 ±0.01

-0.03 ±0.03

-0.01 ±0.04

0.00 ±0.03

-0.02 ±0.04

0.01 ±0.01

0.01 ±0.01

-0.02 ±0.06

-0.05 ±0.00

-0.00 ±0.03

0.00 ±0.00

0.27 ±0.17

0.26 ±0.33

0.16 ±0.09

0.37 ±0.48

0.27 ±0.28

-0.01 ±0.04

0.03 ±0.10

0.25 ±0.10

0.08 ±0.15

-0.30 ±0.19

-0.30 ±0.22

-0.69 ±0.44

0.00 ±0.00

0.18 ±0.12

0.52 ±0.30

0.82 ±0.28

0.70 ±0.66

0.81 ±0.29

0.31 ±0.18

0.49 ±0.22

0.50 ±0.23

0.32 ±0.07

-0.02 ±0.13

-0.23 ±0.21

-0.43 ±0.20

0.00 ±0.00

-0.24 ±0.07

0.20 ±0.32

0.19 ±0.20

0.20 ±0.37

0.09 ±0.31

0.10 ±0.14

0.23 ±0.25

0.17 ±0.52

0.02 ±0.47

-0.22 ±0.56

-0.40 ±0.44

-0.67 ±0.08

0.00 ±0.00

-0.38 ±0.09

-0.30 ±0.14

-0.41 ±0.12

-0.36 ±0.06

-0.36 ±0.04

-0.29 ±0.05

-0.33 ±0.04

-0.39 ±0.06

-0.37 ±0.13

-0.51 ±0.13

-0.78 ±0.20

-1.23 ±0.13

0.00 ±0.00

0.18 ±0.03

0.11 ±0.08

-0.05 ±0.09

-0.05 ±0.08

-0.21 ±0.04

-0.38 ±0.04

-0.57 ±0.05

-0.85 ±0.15

-1.82 ±0.14

-2.81 ±0.31

-3.61 ±0.08

-5.92 ±0.32

0.00 ±0.00

0.61 ±0.22

0.15 ±0.62

0.30 ±0.18

-0.22 ±0.35

-0.02 ±0.42

-0.27 ±0.15

0.09 ±0.21

-0.63 ±0.58

-1.01 ±0.54

-1.47 ±0.49

-1.07 ±0.23

-2.05 ±0.21

0.00 ±0.00

0.24 ±0.26

-0.45 ±0.18

-0.33 ±0.30

-0.44 ±0.50

-0.69 ±0.65

-1.32 ±0.71

-2.20 ±0.83

-3.92 ±0.89

-5.78 ±0.90

-9.05 ±1.00

-12.43 ±0.61

-16.77 ±0.77

0.00 ±0.00

-0.60 ±0.27

-1.09 ±0.54

-1.41 ±0.41

-1.79 ±0.48

-2.21 ±0.75

-3.27 ±0.99

-4.98 ±1.22

-7.01 ±1.40

-9.65 ±1.28

-14.17 ±1.35

-18.03 ±0.75

-23.68 ±0.47

0.00 ±0.00

0.21 ±0.16

0.32 ±0.05

0.74 ±0.24

0.39 ±0.35

0.47 ±0.42

0.12 ±0.43

0.08 ±0.52

-0.17 ±0.13

-0.56 ±0.07

-1.07 ±0.19

-1.46 ±0.29

-2.33 ±0.45

0.00 ±0.00

-1.94 ±0.35

-1.67 ±0.52

-2.62 ±0.81

-4.12 ±1.01

-6.89 ±1.03

-8.73 ±0.11

-9.46 ±0.37

-11.47 ±0.23

-13.68 ±0.43

-16.19 ±0.49

-18.48 ±0.55

-21.47 ±0.41

0.00 ±0.00

0.12 ±0.06

0.51 ±0.06

0.56 ±0.23

0.40 ±0.41

0.47 ±0.36

0.26 ±0.27

0.00 ±0.47

-0.27 ±0.20

-0.55 ±0.17

-0.98 ±0.22

-1.54 ±0.47

-2.31 ±0.47

0.00 ±0.00

0.41 ±0.54

0.28 ±0.66

0.10 ±0.74

-0.28 ±0.69

-0.48 ±0.42

-0.52 ±0.77

-0.84 ±0.49

-1.12 ±0.62

-1.07 ±0.30

-1.09 ±0.44

-1.27 ±0.28

-1.35 ±0.40

0.00 ±0.00

-0.13 ±0.41

0.74 ±0.17

0.61 ±0.24

0.61 ±0.25

0.72 ±0.29

0.79 ±0.30

0.50 ±0.36

0.28 ±0.12

0.17 ±0.22

-0.32 ±0.38

-0.84 ±0.24

-1.20 ±0.08

0.00 ±0.00

0.15 ±0.38

0.59 ±0.45

0.58 ±0.20

0.54 ±0.40

0.43 ±0.38

0.71 ±0.53

0.33 ±0.12

0.18 ±0.19

-0.01 ±0.06

-0.08 ±0.22

-0.55 ±0.27

-1.26 ±0.34

0

1

2

3

4

5

6

7

8

9

10

11

12

0.00 -0.06 -0.23 -0.28 -0.13 -0.31 -0.33 -0.00 -0.05 1.18 0.98 1.49 1.20 0.30 0.74 ±0.00 ±0.08 ±0.05 ±0.05 ±0.38 ±0.04 ±0.19 ±0.25 ±0.09 ±0.07 ±0.31 ±0.19 ±0.17 ±0.14 ±0.19 0.00 -0.07 0.07 0.01 0.13 0.01 -0.27 -0.26 0.28 0.65 0.38 0.73 0.81 0.38 1.22 ±0.00 ±0.11 ±0.08 ±0.08 ±0.20 ±0.08 ±0.04 ±0.05 ±0.12 ±0.18 ±0.05 ±0.17 ±0.29 ±0.15 ±0.22 0.00 -0.03 0.04 -0.01 0.12 0.13 0.03 0.03 -0.03 0.26 0.40 0.26 -0.43 -0.56 -0.19 ±0.00 ±0.04 ±0.13 ±0.06 ±0.13 ±0.08 ±0.01 ±0.05 ±0.11 ±0.05 ±0.05 ±0.08 ±0.08 ±0.09 ±0.28 0.00 0.02 -0.12 -0.16 -0.03 -0.05 -0.20 -0.02 0.26 0.69 0.45 0.53 0.82 0.97 1.33 ±0.00 ±0.04 ±0.10 ±0.04 ±0.08 ±0.11 ±0.05 ±0.05 ±0.07 ±0.07 ±0.10 ±0.02 ±0.06 ±0.18 ±0.19 0.00 -0.10 0.02 -0.12 -0.21 -0.05 -0.24 -0.28 0.35 0.90 0.83 0.95 -0.24 -0.29 -0.12 ±0.00 ±0.06 ±0.06 ±0.10 ±0.20 ±0.04 ±0.25 ±0.09 ±0.05 ±0.15 ±0.10 ±0.15 ±0.28 ±0.11 ±0.12 0.00 -0.03 0.05 -0.00 0.03 -0.08 0.06 0.01 -0.03 0.21 0.46 0.40 0.13 0.23 0.22 ±0.00 ±0.05 ±0.05 ±0.05 ±0.01 ±0.09 ±0.03 ±0.06 ±0.06 ±0.05 ±0.19 ±0.09 ±0.15 ±0.04 ±0.05 0.00 -0.15 -0.08 -0.02 -0.01 -0.40 -0.27 -0.11 -0.22 1.28 0.82 1.61 1.59 0.16 0.84 ±0.00 ±0.08 ±0.04 ±0.13 ±0.29 ±0.09 ±0.21 ±0.09 ±0.06 ±0.36 ±0.27 ±0.15 ±0.08 ±0.34 ±0.10 0.00 -0.12 0.24 0.03 -0.09 -0.06 -0.13 -0.20 0.47 0.89 0.80 0.74 0.22 -0.12 0.17 ±0.00 ±0.13 ±0.15 ±0.17 ±0.10 ±0.11 ±0.21 ±0.18 ±0.32 ±0.24 ±0.20 ±0.26 ±0.28 ±0.10 ±0.42 0.00 -0.04 0.15 0.03 0.15 0.09 -0.04 -0.07 0.23 0.65 0.46 0.38 -0.34 -0.67 -0.70 ±0.00 ±0.02 ±0.03 ±0.12 ±0.18 ±0.11 ±0.08 ±0.22 ±0.25 ±0.09 ±0.04 ±0.20 ±0.23 ±0.25 ±0.35 0.00 -0.00 -0.09 -0.16 -0.10 -0.06 -0.21 -0.06 0.24 0.66 0.47 0.46 0.73 0.99 1.38 ±0.00 ±0.07 ±0.07 ±0.10 ±0.09 ±0.07 ±0.05 ±0.08 ±0.10 ±0.07 ±0.12 ±0.04 ±0.10 ±0.21 ±0.24 0.00 0.08 -0.08 -0.22 -0.33 -0.12 0.14 0.41 0.69 2.14 2.54 2.76 2.20 2.47 1.47 ±0.00 ±0.21 ±0.18 ±0.20 ±0.19 ±0.20 ±0.17 ±0.20 ±0.25 ±0.18 ±0.37 ±0.09 ±0.63 ±0.14 ±0.29 0.00 -0.22 0.06 -0.40 -0.54 -0.35 -0.16 -0.21 1.41 1.96 1.27 1.85 1.54 1.37 0.90 ±0.00 ±0.13 ±0.25 ±0.15 ±0.25 ±0.30 ±0.25 ±0.21 ±0.24 ±0.35 ±0.33 ±0.55 ±0.34 ±0.16 ±0.31 0.00 0.12 -0.05 -0.19 -0.15 -0.59 -0.85 0.67 1.45 1.11 1.75 1.57 1.65 -0.59 -0.40 ±0.00 ±0.09 ±0.02 ±0.41 ±0.24 ±0.40 ±0.27 ±0.33 ±0.90 ±0.98 ±0.32 ±0.57 ±0.18 ±0.82 ±0.08 0.00 0.36 -0.10 -0.51 -0.55 -1.45 -1.48 0.24 2.30 2.50 2.65 2.88 3.67 2.91 3.49 ±0.00 ±0.25 ±0.19 ±0.27 ±0.31 ±0.43 ±0.51 ±0.92 ±0.11 ±0.60 ±0.68 ±1.13 ±1.25 ±0.54 ±0.61 0.00 -0.23 0.09 0.03 0.19 0.09 0.03 0.05 0.25 1.38 2.09 2.41 2.43 1.37 1.51 ±0.00 ±0.05 ±0.08 ±0.09 ±0.01 ±0.03 ±0.16 ±0.05 ±0.17 ±0.31 ±0.15 ±0.08 ±0.14 ±0.18 ±0.05 0.00 -0.48 -0.31 0.32 0.38 0.65 1.57 2.25 1.74 2.93 2.14 2.09 1.46 0.23 0.53 ±0.00 ±0.10 ±0.24 ±0.19 ±0.04 ±0.33 ±0.23 ±0.41 ±0.35 ±0.36 ±0.91 ±0.25 ±0.49 ±0.26 ±0.41 0.00 -0.07 0.04 0.03 0.07 0.00 0.24 -0.15 0.37 1.58 2.25 2.26 2.34 1.34 1.40 ±0.00 ±0.09 ±0.11 ±0.11 ±0.09 ±0.04 ±0.11 ±0.22 ±0.21 ±0.19 ±0.12 ±0.06 ±0.02 ±0.15 ±0.24 0.00 -0.04 -0.24 -0.17 0.55 -0.45 -0.59 0.06 0.56 2.26 5.16 5.01 5.73 4.49 4.49 ±0.00 ±0.09 ±0.18 ±0.22 ±0.19 ±0.15 ±0.25 ±0.22 ±0.36 ±0.35 ±0.33 ±0.11 ±0.51 ±0.94 ±0.41 0.00 -0.18 0.06 0.01 0.30 -0.09 0.08 -0.34 0.59 2.26 3.50 3.59 3.40 1.90 1.42 ±0.00 ±0.07 ±0.19 ±0.18 ±0.18 ±0.11 ±0.08 ±0.03 ±0.44 ±0.51 ±0.29 ±0.16 ±0.08 ±0.53 ±0.29 0.00 -0.16 0.12 0.00 0.11 -0.22 0.04 -0.49 0.71 2.32 3.17 3.59 2.80 1.96 1.68 ±0.00 ±0.15 ±0.20 ±0.06 ±0.03 ±0.28 ±0.19 ±0.18 ±0.31 ±0.26 ±0.20 ±0.21 ±0.07 ±0.20 ±0.33

0

1

2

3

4

5

6

7

8

9

Ar_COO Ar_NH Ar_OH COO COO2 NH0 NH1 NH2 amide aniline ether morpholine para_hydroxylation phenol phenol_noOrthoHbond piperdine piperzine priamide NO Heteroatoms

10 11 12 13 14

0.00 -0.38 0.20 -0.03 -0.11 -0.68 -0.87 -0.88 -0.46 0.50 0.99 -0.38 -1.80 -2.80 -2.09 ±0.00 ±0.12 ±0.32 ±0.14 ±0.39 ±0.14 ±0.26 ±0.38 ±0.40 ±0.68 ±0.61 ±1.03 ±0.19 ±0.66 ±0.81 0.00 -0.16 0.17 -0.26 0.40 0.51 0.58 1.60 0.87 3.89 4.60 4.05 3.76 1.86 2.17 ±0.00 ±0.12 ±0.24 ±0.11 ±0.34 ±0.15 ±0.24 ±0.22 ±0.45 ±0.78 ±0.63 ±0.23 ±0.06 ±0.79 ±0.27 0.00 -0.11 -0.16 -0.20 0.62 -0.67 -0.16 0.11 0.79 2.94 5.81 4.39 4.28 3.29 3.32 ±0.00 ±0.00 ±0.15 ±0.19 ±0.18 ±0.08 ±0.17 ±0.08 ±0.22 ±0.28 ±0.51 ±0.12 ±0.09 ±0.04 ±0.18 0.00 -0.16 -0.32 -0.44 0.91 -0.79 -0.34 0.15 0.84 3.05 5.85 4.49 4.54 3.43 3.37 ±0.00 ±0.07 ±0.13 ±0.21 ±0.22 ±0.31 ±0.18 ±0.33 ±0.23 ±0.25 ±0.49 ±0.14 ±0.26 ±0.20 ±0.24 0.00 -0.12 -0.16 -0.17 0.02 -0.09 -0.04 -0.26 0.56 0.96 0.72 0.81 0.29 -0.18 -0.31 ±0.00 ±0.09 ±0.24 ±0.15 ±0.17 ±0.17 ±0.13 ±0.21 ±0.13 ±0.04 ±0.17 ±0.21 ±0.19 ±0.22 ±0.08 0.00 -0.29 0.19 0.10 0.45 0.16 -0.16 -0.06 -0.61 0.51 1.13 1.59 0.97 0.92 0.91 ±0.00 ±0.08 ±0.16 ±0.18 ±0.10 ±0.12 ±0.21 ±0.08 ±0.26 ±0.11 ±0.32 ±0.08 ±0.14 ±0.14 ±0.14 0.00 -0.16 -0.47 0.24 0.32 -0.58 -1.27 -1.32 -0.22 0.56 2.13 2.74 1.34 0.39 0.46 ±0.00 ±0.14 ±0.20 ±0.24 ±0.17 ±0.09 ±0.28 ±0.32 ±0.16 ±0.24 ±0.45 ±0.20 ±0.04 ±0.20 ±0.52 0.00 -0.30 -0.11 0.18 0.47 0.40 0.05 -0.20 0.30 1.35 0.60 -0.18 -0.28 -1.01 -1.74 ±0.00 ±0.21 ±0.19 ±0.12 ±0.26 ±0.08 ±0.22 ±0.14 ±0.14 ±0.22 ±0.41 ±0.37 ±0.23 ±0.36 ±0.27 0.00 -0.27 -0.27 -0.31 0.47 -0.23 -0.15 -0.47 0.49 0.88 1.43 1.95 1.58 1.05 1.33 ±0.00 ±0.16 ±0.25 ±0.22 ±0.14 ±0.16 ±0.06 ±0.01 ±0.05 ±0.21 ±0.11 ±0.47 ±0.16 ±0.37 ±0.18 0.00 0.15 0.03 -0.05 0.49 -0.18 -0.28 -0.06 -0.42 0.21 0.55 0.64 0.31 0.28 0.59 ±0.00 ±0.01 ±0.07 ±0.14 ±0.04 ±0.19 ±0.09 ±0.04 ±0.15 ±0.17 ±0.18 ±0.07 ±0.03 ±0.12 ±0.18 0.00 0.18 0.14 0.00 0.04 -0.00 -0.07 -0.57 0.50 3.45 1.93 1.84 0.94 -0.71 -0.16 ±0.00 ±0.07 ±0.03 ±0.06 ±0.13 ±0.11 ±0.28 ±0.92 ±0.30 ±0.41 ±0.48 ±0.84 ±1.05 ±0.34 ±1.05 0.00 -0.13 -0.02 -0.34 -0.37 -0.32 -0.14 0.01 0.94 1.73 2.03 2.12 1.84 1.31 0.68 ±0.00 ±0.08 ±0.07 ±0.03 ±0.07 ±0.13 ±0.10 ±0.28 ±0.05 ±0.12 ±0.17 ±0.41 ±0.46 ±0.51 ±0.60 0.00 -0.11 0.01 -0.61 0.25 0.84 1.18 2.98 2.39 6.26 7.56 5.59 5.03 2.14 3.33 ±0.00 ±0.10 ±0.12 ±0.22 ±0.34 ±0.07 ±0.11 ±0.52 ±0.73 ±1.05 ±1.03 ±0.69 ±0.82 ±0.90 ±0.63 0.00 0.02 0.06 -0.39 0.04 0.72 1.49 2.41 1.77 6.45 7.20 5.75 5.57 2.85 3.51 ±0.00 ±0.04 ±0.06 ±0.01 ±0.23 ±0.15 ±0.41 ±0.28 ±0.41 ±0.96 ±0.75 ±0.34 ±0.57 ±0.49 ±0.50 0.00 -0.12 0.07 -0.10 -0.02 -0.11 0.31 -0.20 0.52 1.72 1.01 1.77 2.25 1.43 2.43 ±0.00 ±0.09 ±0.14 ±0.26 ±0.04 ±0.22 ±0.28 ±0.37 ±0.52 ±0.41 ±0.48 ±0.13 ±0.50 ±0.46 ±0.15 0.00 -0.11 0.01 0.81 0.20 0.10 1.42 -0.18 1.64 2.52 1.56 1.97 1.55 2.28 2.73 ±0.00 ±0.11 ±0.08 ±0.23 ±0.14 ±0.37 ±0.38 ±0.63 ±0.29 ±0.49 ±0.16 ±0.91 ±1.23 ±0.39 ±0.47 0.00 -0.26 -0.62 -0.11 0.48 -1.05 0.37 1.78 2.49 4.16 1.83 2.64 2.58 0.97 1.33 ±0.00 ±0.11 ±0.24 ±0.38 ±0.24 ±0.26 ±0.36 ±1.15 ±0.92 ±1.59 ±0.35 ±0.55 ±1.00 ±0.15 ±0.79 0.00 -0.42 -0.04 0.05 0.65 0.79 1.72 2.98 2.05 2.35 2.34 1.23 0.53 -0.09 -0.59 ±0.00 ±0.10 ±0.18 ±0.40 ±0.25 ±0.07 ±0.13 ±0.43 ±0.09 ±0.51 ±0.40 ±0.26 ±0.31 ±0.41 ±1.33 0.00 -0.44 -0.18 -0.63 -0.45 0.37 1.58 1.52 1.77 2.18 0.77 1.43 1.10 0.53 -0.70 ±0.00 ±0.20 ±0.04 ±0.18 ±0.27 ±0.43 ±0.44 ±0.17 ±0.59 ±0.48 ±0.31 ±0.62 ±0.07 ±0.27 ±1.20

0

1

2

3

4

5

Layer

6

7

8

9

RotatableBonds Ring ArN Ar_N C_O C_O_noCOO Imine Ndealkylation1 Ndealkylation2 Nhpyrrole alkyl_halide aryl_methyl bicyclic imidazole ketone ketone_Topliss methoxy nitrile sulfide urea

10 11 12 13 14

0.00 -0.04 -0.21 -0.06 0.16 -0.07 -0.14 -0.35 -0.64 0.44 0.15 0.26 -0.36 -0.43 -0.42 ±0.00 ±0.16 ±0.08 ±0.03 ±0.32 ±0.21 ±0.23 ±0.10 ±0.15 ±0.08 ±0.04 ±0.53 ±0.69 ±0.49 ±0.41 0.00 -0.00 -0.01 0.07 0.02 -0.01 0.03 -0.02 0.13 0.13 0.11 0.01 -0.27 -0.31 -0.42 ±0.00 ±0.01 ±0.07 ±0.10 ±0.04 ±0.07 ±0.01 ±0.09 ±0.06 ±0.09 ±0.19 ±0.07 ±0.15 ±0.19 ±0.25 0.00 -0.48 -0.51 -0.37 0.55 -0.11 -3.18 -3.69 -2.65 -0.64 0.82 2.61 1.81 0.10 0.92 ±0.00 ±0.10 ±0.38 ±0.10 ±0.20 ±0.14 ±0.21 ±0.34 ±0.78 ±0.79 ±0.27 ±1.25 ±0.47 ±1.20 ±1.32

15

0.00 0.06 -0.09 -0.15 -0.20 -0.17 -0.09 -0.01 0.70 1.17 0.89 0.68 0.09 -0.28 -0.24 ±0.00 ±0.02 ±0.10 ±0.09 ±0.12 ±0.13 ±0.13 ±0.17 ±0.27 ±0.23 ±0.16 ±0.11 ±0.33 ±0.13 ±0.39 0.00 -0.01 0.02 0.01 0.12 0.45 0.07 -0.14 0.33 0.81 0.55 0.60 0.39 -0.03 -0.25 ±0.00 ±0.02 ±0.09 ±0.06 ±0.06 ±0.04 ±0.04 ±0.04 ±0.09 ±0.17 ±0.06 ±0.09 ±0.15 ±0.15 ±0.15 0.00 -0.03 -0.01 -0.02 0.14 0.22 0.23 0.11 0.53 1.50 1.46 1.09 0.50 0.16 -0.42 ±0.00 ±0.12 ±0.16 ±0.14 ±0.05 ±0.03 ±0.07 ±0.05 ±0.23 ±0.23 ±0.36 ±0.13 ±0.09 ±0.16 ±0.10

10

0.00 -0.14 0.03 -0.25 -0.12 -0.10 -0.45 -0.24 0.19 0.75 0.62 1.22 1.28 0.95 0.32 ±0.00 ±0.12 ±0.11 ±0.12 ±0.23 ±0.20 ±0.09 ±0.15 ±0.10 ±0.16 ±0.14 ±0.44 ±0.16 ±0.44 ±0.42 0.00 -0.02 0.20 -0.11 0.20 -0.14 -0.11 -0.01 -0.06 0.75 -0.42 -0.43 -0.78 -0.09 -0.44 ±0.00 ±0.19 ±0.16 ±0.08 ±0.27 ±0.20 ±0.25 ±0.21 ±0.28 ±0.85 ±0.21 ±0.30 ±0.31 ±0.19 ±0.33 0.00 -0.28 0.14 0.01 -0.22 -0.30 -0.33 -0.51 -0.26 0.75 0.12 0.56 0.27 0.10 0.28 ±0.00 ±0.12 ±0.16 ±0.13 ±0.30 ±0.25 ±0.03 ±0.25 ±0.20 ±0.27 ±0.40 ±0.62 ±0.20 ±0.45 ±0.81

5

0.00 -0.38 0.20 -0.03 -0.11 -0.68 -0.87 -0.88 -0.46 0.50 0.99 -0.38 -1.80 -2.80 -2.09 ±0.00 ±0.12 ±0.32 ±0.14 ±0.39 ±0.14 ±0.26 ±0.38 ±0.40 ±0.68 ±0.61 ±1.03 ±0.19 ±0.66 ±0.81 0.00 0.04 0.01 -0.13 -0.02 -0.07 0.45 0.26 1.15 2.90 3.53 3.27 3.76 2.93 1.71 ±0.00 ±0.12 ±0.22 ±0.24 ±0.18 ±0.31 ±0.14 ±0.26 ±0.16 ±0.38 ±0.13 ±0.75 ±0.49 ±0.58 ±0.16 0.00 -0.08 -0.11 -0.24 0.22 -0.66 -0.58 -0.22 0.63 0.98 0.62 -0.14 -0.46 0.17 -0.80 ±0.00 ±0.14 ±0.16 ±0.17 ±0.04 ±0.26 ±0.10 ±0.26 ±0.20 ±0.12 ±0.34 ±0.15 ±0.26 ±0.28 ±0.19

0

0.00 0.07 0.20 -0.05 0.11 0.08 -0.07 -0.22 0.20 0.33 0.29 0.65 0.12 -0.22 -0.65 ±0.00 ±0.04 ±0.03 ±0.04 ±0.00 ±0.04 ±0.02 ±0.13 ±0.05 ±0.03 ±0.05 ±0.15 ±0.21 ±0.19 ±0.19 0.00 -0.36 -0.12 -0.13 -0.37 -0.19 -0.06 -0.49 0.37 0.97 0.56 1.01 0.85 0.25 -0.79 ±0.00 ±0.05 ±0.20 ±0.16 ±0.21 ±0.20 ±0.38 ±0.32 ±0.56 ±0.05 ±0.15 ±0.23 ±0.13 ±0.26 ±0.09

5

0.00 -0.01 -0.01 -0.21 -0.03 -0.55 -0.11 -0.26 0.15 0.98 1.11 1.00 0.48 -0.35 -0.31 ±0.00 ±0.10 ±0.08 ±0.11 ±0.27 ±0.28 ±0.19 ±0.12 ±0.21 ±0.46 ±0.63 ±0.07 ±0.16 ±0.18 ±0.52 0.00 -0.04 -0.07 -0.19 -0.11 -0.51 -0.35 -0.04 0.05 1.05 1.10 0.83 0.60 -0.75 -0.26 ±0.00 ±0.09 ±0.08 ±0.16 ±0.17 ±0.22 ±0.15 ±0.20 ±0.29 ±0.19 ±0.13 ±0.10 ±0.42 ±0.52 ±0.29 0.00 -0.07 -0.46 -0.36 0.16 -0.11 -0.53 -0.50 -1.17 -0.73 0.42 0.67 0.27 0.53 0.21 ±0.00 ±0.09 ±0.08 ±0.32 ±0.06 ±0.26 ±0.04 ±0.24 ±0.38 ±0.39 ±0.13 ±0.29 ±0.32 ±0.26 ±0.11

10

0.00 -0.86 0.12 -0.23 0.18 0.26 0.43 -0.39 0.55 2.33 1.48 2.19 0.16 0.37 0.17 ±0.00 ±0.06 ±0.27 ±0.46 ±0.08 ±0.07 ±0.43 ±0.02 ±0.38 ±0.21 ±0.34 ±0.19 ±0.52 ±0.20 ±1.21 0.00 -0.04 0.14 -0.34 0.46 -0.34 0.21 1.26 0.87 1.92 3.27 3.44 3.52 2.75 2.20 ±0.00 ±0.07 ±0.07 ±0.27 ±0.23 ±0.12 ±0.31 ±0.24 ±0.52 ±0.31 ±1.05 ±0.81 ±1.03 ±0.81 ±0.22 0.00 -0.22 -0.01 0.03 0.42 -0.14 -0.30 0.00 -0.15 1.17 0.58 0.59 -0.14 -0.43 -0.31 ±0.00 ±0.09 ±0.08 ±0.15 ±0.06 ±0.43 ±0.21 ±0.37 ±0.74 ±0.33 ±0.54 ±0.68 ±0.08 ±0.91 ±0.38

0

1

2

3

4

5

6

7

8

AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen pyridine thiazole thiophene NHOH HAcceptors HDonors Al_COO Al_OH Al_OH_noTert

9 10 11 12 13 14

0.00 0.20 0.63 0.31 0.13 0.42 0.27 0.22 0.19 -0.32 -0.48 -0.86 -1.14 -1.53 -2.09 ±0.00 ±0.07 ±0.31 ±0.24 ±0.01 ±0.17 ±0.28 ±0.02 ±0.23 ±0.07 ±0.07 ±0.16 ±0.15 ±0.02 ±0.17 0.00 -0.10 0.32 0.24 0.15 -0.33 -0.38 -0.39 -0.73 -0.22 -0.60 -0.78 -1.29 -1.81 -1.79 ±0.00 ±0.38 ±0.07 ±0.13 ±0.18 ±0.27 ±0.15 ±0.10 ±0.03 ±0.04 ±0.11 ±0.11 ±0.20 ±0.27 ±0.10 0.00 -0.28 -0.06 -0.11 -0.07 -0.18 -0.14 -0.30 -0.48 -0.49 -0.57 -0.91 -1.13 -1.15 -1.40 ±0.00 ±0.22 ±0.09 ±0.09 ±0.10 ±0.04 ±0.08 ±0.03 ±0.09 ±0.08 ±0.10 ±0.04 ±0.03 ±0.09 ±0.11 0.00 0.00 0.09 0.14 0.12 0.05 0.04 0.14 0.19 0.08 -0.04 -0.19 -0.21 -0.27 -0.38 ±0.00 ±0.19 ±0.14 ±0.09 ±0.04 ±0.08 ±0.15 ±0.03 ±0.10 ±0.07 ±0.17 ±0.12 ±0.09 ±0.05 ±0.08 0.00 -0.68 -0.75 -0.51 -0.69 -0.74 -0.90 -0.87 -1.09 -1.18 -1.43 -1.49 -2.28 -2.78 -3.15 ±0.00 ±0.05 ±0.06 ±0.19 ±0.13 ±0.10 ±0.23 ±0.09 ±0.17 ±0.18 ±0.18 ±0.28 ±0.09 ±0.05 ±0.12 0.00 -0.03 0.05 -0.01 -0.02 -0.04 0.01 -0.05 -0.06 0.02 -0.16 -0.01 -0.07 -0.08 -0.13 ±0.00 ±0.03 ±0.02 ±0.05 ±0.04 ±0.06 ±0.09 ±0.08 ±0.12 ±0.07 ±0.03 ±0.04 ±0.04 ±0.02 ±0.03 0.00 0.10 0.62 0.66 0.26 0.46 0.51 0.38 0.19 0.03 -0.53 -0.65 -0.67 -1.08 -1.73 ±0.00 ±0.09 ±0.15 ±0.17 ±0.23 ±0.10 ±0.15 ±0.30 ±0.10 ±0.29 ±0.06 ±0.25 ±0.07 ±0.31 ±0.34 0.00 -0.30 0.08 -0.01 0.12 -0.23 -0.20 -0.58 -0.53 -0.34 -0.45 -0.65 -1.41 -1.99 -2.14 ±0.00 ±0.07 ±0.13 ±0.24 ±0.21 ±0.33 ±0.39 ±0.28 ±0.12 ±0.28 ±0.14 ±0.13 ±0.04 ±0.13 ±0.05 0.00 -0.15 -0.19 -0.23 -0.23 -0.28 -0.38 -0.31 -0.60 -0.57 -0.73 -0.94 -1.18 -1.48 -1.83 ±0.00 ±0.23 ±0.17 ±0.13 ±0.05 ±0.17 ±0.17 ±0.16 ±0.12 ±0.08 ±0.14 ±0.03 ±0.09 ±0.06 ±0.17 0.00 0.06 0.10 0.10 0.17 0.03 0.11 0.19 0.25 0.08 -0.07 -0.16 -0.18 -0.26 -0.31 ±0.00 ±0.06 ±0.10 ±0.12 ±0.08 ±0.08 ±0.09 ±0.16 ±0.05 ±0.11 ±0.13 ±0.11 ±0.04 ±0.06 ±0.02 0.00 -0.29 -0.18 -0.25 -0.18 -0.20 -0.47 -0.17 -0.27 -0.47 -1.12 -1.18 -1.12 -1.51 -2.34 ±0.00 ±0.19 ±0.13 ±0.03 ±0.20 ±0.02 ±0.06 ±0.07 ±0.07 ±0.07 ±0.13 ±0.16 ±0.12 ±0.11 ±0.29 0.00 -1.04 -1.25 -1.17 -0.87 -0.41 -0.71 -1.29 -1.77 -1.86 -2.19 -2.82 -3.41 -3.56 -3.93 ±0.00 ±0.47 ±0.28 ±0.18 ±0.25 ±0.29 ±0.20 ±0.47 ±0.08 ±0.12 ±0.40 ±0.25 ±0.33 ±0.24 ±0.08 0.00 -0.52 -0.43 0.18 -0.07 -0.03 -0.65 -1.02 -1.51 -3.03 -4.21 -4.25 -5.05 -5.87 -7.24 ±0.00 ±0.32 ±0.36 ±0.50 ±0.15 ±0.90 ±0.94 ±0.54 ±0.36 ±0.35 ±0.13 ±0.46 ±0.48 ±0.24 ±0.17 0.00 -1.00 -0.75 -0.83 -0.82 -0.74 -1.39 -1.80 -2.27 -2.91 -4.52 -5.58 -7.34 -8.31 -9.70 ±0.00 ±0.38 ±0.66 ±0.68 ±0.48 ±0.81 ±0.74 ±0.51 ±0.30 ±0.57 ±0.85 ±0.85 ±0.40 ±0.55 ±0.57 0.00 0.11 0.07 0.43 -0.03 -0.57 -0.50 -1.12 -1.10 -1.11 -1.39 -1.56 -1.84 -1.87 -2.75 ±0.00 ±0.07 ±0.08 ±0.21 ±0.13 ±0.11 ±0.07 ±0.08 ±0.07 ±0.08 ±0.04 ±0.08 ±0.22 ±0.16 ±0.05 0.00 1.26 1.69 1.48 1.68 1.21 1.24 0.91 0.75 0.23 -0.35 -0.39 -1.12 -1.56 -2.74 ±0.00 ±0.32 ±0.20 ±0.71 ±0.14 ±0.31 ±0.15 ±0.41 ±0.28 ±0.25 ±0.42 ±0.43 ±0.47 ±0.35 ±0.26 0.00 0.16 0.04 0.43 0.04 -0.63 -0.40 -1.18 -1.07 -1.28 -1.33 -1.64 -1.98 -1.82 -2.74 ±0.00 ±0.33 ±0.20 ±0.21 ±0.20 ±0.07 ±0.19 ±0.10 ±0.07 ±0.08 ±0.37 ±0.12 ±0.19 ±0.12 ±0.17 0.00 1.34 1.32 0.98 0.87 0.83 0.71 0.40 0.27 -0.05 -0.25 -0.48 -0.81 -1.22 -1.72 ±0.00 ±0.28 ±0.24 ±0.24 ±0.25 ±0.10 ±0.21 ±0.14 ±0.15 ±0.47 ±0.41 ±0.37 ±0.30 ±0.34 ±0.20 0.00 0.61 0.57 0.58 0.11 -0.25 -0.19 -0.21 -0.28 -0.62 -1.39 -1.52 -1.46 -1.59 -2.44 ±0.00 ±0.21 ±0.11 ±0.18 ±0.06 ±0.16 ±0.15 ±0.13 ±0.17 ±0.33 ±0.16 ±0.25 ±0.46 ±0.33 ±0.19 0.00 0.68 0.86 0.75 0.67 0.07 -0.01 0.00 -0.24 -0.74 -1.06 -1.28 -1.14 -1.13 -1.91 ±0.00 ±0.28 ±0.10 ±0.11 ±0.09 ±0.13 ±0.11 ±0.16 ±0.10 ±0.24 ±0.37 ±0.34 ±0.50 ±0.18 ±0.10

0

1

2

3

4

5

6

7

8

9

-0.21 ±0.22

0.18 ±0.08

0.06 ±0.32

-0.07 ±0.18

1.14 ±0.29

0.64 ±0.63

0.00 ±0.00

-0.27 ±0.16

0.22 ±0.23

-0.11 ±0.35

-0.69 ±0.09

0.65 ±0.15

0.73 ±0.20

0.00 ±0.00

-0.04 ±0.01

0.01 ±0.06

-0.21 ±0.05

0.00 ±0.04

-0.14 ±0.08

-0.13 ±0.02

0.00 ±0.00

0.08 ±0.08

0.18 ±0.04

0.05 ±0.07

0.13 ±0.06

0.33 ±0.06

0.02 ±0.10

0.00 ±0.00

0.02 ±0.06

0.44 ±0.17

-0.32 ±0.25

0.14 ±0.20

0.67 ±0.21

0.43 ±0.22

0.00 ±0.00

-0.03 ±0.02

0.04 ±0.05

-0.00 ±0.06

0.00 ±0.10

-0.04 ±0.02

0.05 ±0.13

0.00 ±0.00

-0.05 ±0.24

0.01 ±0.28

0.26 ±0.15

0.07 ±0.59

1.23 ±0.25

1.20 ±0.25

0.00 ±0.00

-0.18 ±0.14

-0.07 ±0.06

-0.19 ±0.36

-0.64 ±0.07

1.09 ±0.12

0.87 ±0.15

0.00 ±0.00

-0.19 ±0.28

-0.50 ±0.07

-0.45 ±0.08

-0.38 ±0.29

-0.15 ±0.08

-0.27 ±0.00

0.00 ±0.00

0.02 ±0.04

0.25 ±0.07

0.10 ±0.06

0.11 ±0.10

0.37 ±0.06

0.02 ±0.15

0.00 ±0.00

0.07 ±0.23

0.47 ±0.10

0.12 ±0.51

0.40 ±0.27

-0.91 ±0.49

-0.12 ±0.33

0.00 ±0.00

-0.15 ±0.16

0.57 ±0.15

0.72 ±0.16

0.14 ±0.43

2.40 ±0.66

0.66 ±0.20

0.00 ±0.00

0.37 ±0.50

1.20 ±0.24

-0.29 ±0.16

-1.10 ±0.52

0.34 ±0.60

-0.31 ±0.97

0.00 ±0.00

0.41 ±0.32

2.59 ±0.24

0.40 ±0.24

0.28 ±0.62

1.47 ±1.12

-0.07 ±0.94

0.00 ±0.00

0.00 ±0.00

0.00 ±0.00

0.00 ±0.00

0.00 ±0.00

0.00 ±0.00

0.00 ±0.00

0.00 ±0.00

0.90 ±0.14

1.78 ±0.21

1.29 ±0.16

1.33 ±0.26

1.66 ±0.63

0.54 ±0.35

0.00 ±0.00

0.11 ±0.24

-0.28 ±0.05

-0.38 ±0.14

-0.53 ±0.24

1.09 ±0.20

1.68 ±0.30

0.00 ±0.00

0.34 ±0.18

0.68 ±0.26

0.12 ±0.15

0.02 ±0.75

2.77 ±0.41

6.12 ±1.03

0.00 ±0.00

0.70 ±0.04

0.65 ±0.44

-0.94 ±0.19

-0.21 ±0.18

0.90 ±0.22

2.63 ±0.31

0.00 ±0.00

0.72 ±0.19

0.83 ±0.28

-0.74 ±0.48

-0.40 ±0.27

0.48 ±0.21

2.19 ±0.11

0

1

2

3

4

5

6

0.00 ±0.00

-1.67 ±0.13

-4.60 ±0.16

-4.57 ±0.22

-2.39 ±0.43

-3.35 ±0.58

-2.50 ±0.51

0.00 ±0.00

-2.05 ±0.19

-3.89 ±0.30

-4.69 ±0.10

-3.13 ±0.45

-3.96 ±0.24

-3.90 ±0.13

0.00 ±0.00

-1.46 ±0.22

-4.08 ±0.30

-4.54 ±0.16

-3.20 ±0.10

-3.63 ±0.18

-3.48 ±0.13

Ar_COO Ar_NH Ar_OH COO COO2 NH0 NH1 NH2 amide aniline ether morpholine para_hydroxylation phenol phenol_noOrthoHbond piperdine piperzine priamide NO Heteroatoms

0.00 ±0.00

0.73 ±0.03

1.46 ±0.06

0.66 ±0.15

2.68 ±1.02

3.39 ±0.76

6.37 ±0.80

0.00 ±0.00

0.15 ±0.15

0.58 ±0.13

0.48 ±0.43

2.18 ±0.69

1.62 ±0.63

-0.37 ±0.42

0.00 ±0.00

1.23 ±0.33

1.70 ±0.23

1.04 ±0.37

-0.87 ±0.94

0.07 ±0.53

1.02 ±1.51

0.00 ±0.00

0.52 ±0.12

0.41 ±0.07

0.41 ±0.37

-0.00 ±0.14

3.33 ±0.52

5.22 ±0.35

0.00 ±0.00

0.62 ±0.04

0.54 ±0.23

0.33 ±0.27

0.04 ±0.31

3.87 ±0.97

5.46 ±0.33

0.00 ±0.00

0.25 ±0.25

0.59 ±0.16

0.46 ±0.06

0.42 ±0.25

0.13 ±0.18

0.45 ±0.09

0.00 ±0.00

-0.41 ±0.29

-0.41 ±0.12

-0.52 ±0.29

-0.98 ±0.19

-0.37 ±0.28

0.68 ±0.39

0.00 ±0.00

1.78 ±0.20

2.84 ±0.35

0.15 ±0.23

-1.08 ±0.12

-2.28 ±0.17

0.20 ±0.49

0.00 ±0.00

0.81 ±0.26

1.66 ±0.14

1.45 ±0.26

1.49 ±0.09

0.18 ±0.45

-1.25 ±0.41

0.00 ±0.00

0.86 ±0.11

2.64 ±0.24

1.86 ±0.41

0.85 ±0.81

0.21 ±0.69

1.34 ±0.62

0.00 ±0.00

0.28 ±0.08

0.11 ±0.22

0.23 ±0.24

0.11 ±0.11

2.32 ±0.29

2.43 ±0.17

0.00 ±0.00

0.53 ±0.18

0.49 ±0.50

0.17 ±0.10

0.11 ±0.67

0.89 ±0.45

0.55 ±1.17

0.00 ±0.00

0.09 ±0.19

0.32 ±0.08

-0.02 ±0.02

-0.02 ±0.58

-0.35 ±0.52

-1.41 ±0.94

0.00 ±0.00

0.74 ±0.07

1.38 ±0.08

0.98 ±0.61

-0.13 ±0.36

0.59 ±0.68

2.40 ±0.59

0.00 ±0.00

0.74 ±0.36

1.30 ±0.28

1.07 ±0.52

0.10 ±0.08

1.14 ±0.43

1.66 ±1.35

0.00 ±0.00

0.13 ±0.29

0.98 ±0.19

1.10 ±0.37

-0.17 ±0.28

0.79 ±0.08

1.26 ±0.73

0.00 ±0.00

0.23 ±0.32

0.50 ±0.39

0.09 ±0.41

-1.13 ±0.31

0.72 ±0.75

-0.42 ±0.66

0.00 ±0.00

0.44 ±0.18

2.30 ±0.72

-0.75 ±0.67

1.28 ±0.62

0.99 ±0.72

2.30 ±0.20

0.00 ±0.00

0.50 ±0.86

1.01 ±1.74

0.77 ±1.33

0.46 ±0.80

0.66 ±1.14

0.09 ±0.15

0.00 ±0.00

1.37 ±0.38

2.70 ±0.54

1.90 ±0.93

0.90 ±0.88

-0.41 ±0.38

-0.47 ±1.39

0

1

2

3

4

5

6

Layer

RotatableBonds Ring ArN Ar_N C_O C_O_noCOO Imine Ndealkylation1 Ndealkylation2 Nhpyrrole alkyl_halide aryl_methyl bicyclic imidazole ketone ketone_Topliss methoxy nitrile sulfide urea

0.13 ±0.38

0.82 ±0.17

-0.40 ±0.44

-1.00 ±0.02

-1.72 ±0.09

-2.23 ±0.12

0.00 ±0.00

-0.03 ±0.03

-0.03 ±0.02

-0.01 ±0.01

0.01 ±0.04

-0.06 ±0.04

-0.11 ±0.15

0.00 ±0.00

2.12 ±0.13

5.46 ±0.37

2.14 ±0.92

1.98 ±0.52

-0.38 ±0.47

1.82 ±1.11

0.00 ±0.00

0.09 ±0.06

0.59 ±0.13

-0.16 ±0.27

0.03 ±0.09

0.87 ±0.28

1.12 ±0.13

0.00 ±0.00

0.53 ±0.13

0.05 ±0.12

-0.10 ±0.14

-0.67 ±0.10

-1.13 ±0.06

-1.59 ±0.10

0.00 ±0.00

0.81 ±0.29

0.75 ±0.07

0.17 ±0.11

-0.41 ±0.21

-0.37 ±0.19

-0.98 ±0.17

0.00 ±0.00

0.27 ±0.13

-0.11 ±0.16

0.46 ±0.18

0.43 ±0.12

-0.52 ±0.11

-0.68 ±0.14

0.00 ±0.00

1.38 ±0.33

1.39 ±0.15

0.17 ±0.12

0.28 ±0.28

0.61 ±0.65

0.95 ±0.65

0.00 ±0.00

-0.02 ±0.24

0.49 ±0.12

1.17 ±0.40

0.22 ±0.42

1.23 ±0.45

0.99 ±1.18

0.00 ±0.00

0.15 ±0.15

0.58 ±0.13

0.48 ±0.43

2.18 ±0.69

1.62 ±0.63

-0.37 ±0.42

0.00 ±0.00

-0.22 ±0.38

0.21 ±0.31

-1.21 ±0.66

-0.10 ±1.24

-1.13 ±1.15

-0.48 ±0.66

0.00 ±0.00

1.94 ±0.06

2.52 ±0.20

0.46 ±0.12

-0.86 ±0.17

-0.89 ±0.31

0.91 ±0.34

0.00 ±0.00

-0.01 ±0.11

0.02 ±0.11

-0.88 ±0.08

-0.64 ±0.26

-0.69 ±0.31

-1.17 ±0.20

0.00 ±0.00

0.37 ±0.27

1.52 ±0.35

0.26 ±0.23

0.07 ±0.30

1.38 ±0.35

1.38 ±1.03

0.00 ±0.00

0.08 ±0.18

0.33 ±0.12

0.21 ±0.27

0.47 ±0.31

-0.18 ±0.17

-0.44 ±0.66

0.00 ±0.00

0.10 ±0.26

0.45 ±0.12

0.34 ±0.13

0.47 ±0.15

0.37 ±0.49

0.12 ±0.51

0.00 ±0.00

2.60 ±0.18

2.57 ±0.27

1.35 ±0.51

0.70 ±0.15

-0.17 ±0.14

3.87 ±0.44

0.00 ±0.00

-0.87 ±0.39

-0.53 ±0.42

-1.65 ±0.61

-1.99 ±1.09

-0.88 ±0.56

-3.50 ±1.43

0.00 ±0.00

0.08 ±0.26

1.86 ±0.26

0.72 ±0.75

0.44 ±0.72

0.85 ±0.33

0.42 ±0.46

0.00 ±0.00

0.12 ±0.10

0.42 ±0.08

0.67 ±0.33

0.18 ±0.77

0.48 ±1.34

0.13 ±1.10

0

1

2

3

4

5

6

0.00 ±0.00

0.21 ±0.11

-2.48 ±0.19

-4.88 ±0.56

-4.64 ±0.32

-3.84 ±0.44

-3.57 ±0.13

0.00 ±0.00

-0.19 ±0.07

-0.38 ±0.08

-0.79 ±0.03

-1.23 ±0.12

-1.47 ±0.19

-1.40 ±0.22

0.00 ±0.00

0.44 ±0.45

1.61 ±0.94

1.64 ±0.71

0.58 ±0.95

-0.64 ±0.19

-0.42 ±0.63

7.5 5.0 2.5 0.0 2.5 5.0 7.5 10.0

AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen pyridine thiazole thiophene NHOH HAcceptors HDonors Al_COO Al_OH Al_OH_noTert

0.00 ±0.00

0.26 ±0.33

0.55 ±0.15

0.13 ±0.02

-0.24 ±0.20

-0.61 ±0.12

-1.08 ±0.07

0.00 ±0.00

-0.96 ±0.29

-1.09 ±0.14

-1.10 ±0.21

-1.79 ±0.14

-2.13 ±0.42

-2.47 ±0.45

0.00 ±0.00

-0.68 ±0.13

-0.85 ±0.08

-0.91 ±0.23

-1.10 ±0.10

-1.59 ±0.18

-1.68 ±0.07

0.00 ±0.00

-0.06 ±0.08

-0.17 ±0.01

-0.09 ±0.10

-0.28 ±0.07

-0.36 ±0.02

-0.61 ±0.14

0.00 ±0.00

-2.05 ±0.23

-2.34 ±0.12

-2.58 ±0.14

-2.97 ±0.12

-3.81 ±0.20

-4.38 ±0.05

0.00 ±0.00

-0.27 ±0.06

-0.21 ±0.02

-0.27 ±0.11

-0.25 ±0.04

-0.30 ±0.07

-0.40 ±0.06

0.00 ±0.00

-0.09 ±0.36

0.24 ±0.34

0.26 ±0.25

-0.19 ±0.27

-0.82 ±0.12

-1.24 ±0.11

0.00 ±0.00

-1.18 ±0.20

-1.16 ±0.38

-1.04 ±0.61

-1.71 ±0.22

-2.16 ±0.24

-2.34 ±0.41

0.00 ±0.00

-1.15 ±0.16

-0.95 ±0.20

-1.18 ±0.20

-1.51 ±0.09

-1.96 ±0.22

-2.25 ±0.08

0.00 ±0.00

-0.07 ±0.10

-0.16 ±0.04

-0.12 ±0.13

-0.32 ±0.01

-0.38 ±0.09

-0.56 ±0.09

0.00 ±0.00

-1.50 ±0.42

-1.04 ±0.35

-1.18 ±0.26

-0.88 ±0.27

-1.74 ±0.22

-2.12 ±0.16

0.00 ±0.00

-3.10 ±0.19

-2.89 ±0.50

-3.49 ±0.23

-4.59 ±0.22

-5.50 ±0.58

-6.42 ±0.51

0.00 ±0.00

-1.38 ±0.58

-3.01 ±0.28

-4.01 ±0.64

-4.16 ±0.91

-5.77 ±1.62

-5.82 ±1.61

0.00 ±0.00

-3.10 ±0.74

-4.17 ±1.21

-5.43 ±0.61

-7.26 ±0.93

-9.18 ±1.02

-9.79 ±0.39

0.00 ±0.00

-0.63 ±0.17

-0.22 ±0.12

-0.85 ±0.27

-1.10 ±0.18

-1.96 ±0.31

-2.64 ±0.20

0.00 ±0.00

0.97 ±0.56

0.85 ±0.66

-0.05 ±0.21

-0.23 ±0.32

-1.22 ±0.86

-1.59 ±0.36

0.00 ±0.00

-0.45 ±0.22

-0.10 ±0.03

-0.78 ±0.28

-1.16 ±0.07

-1.92 ±0.24

-2.67 ±0.12

0.00 ±0.00

2.25 ±0.08

2.30 ±0.06

1.61 ±0.12

1.58 ±0.19

1.16 ±0.24

0.22 ±0.28

0.00 ±0.00

-0.04 ±0.41

0.68 ±0.30

0.46 ±0.28

-0.14 ±0.11

-0.53 ±0.04

-2.08 ±0.22

0.00 ±0.00

0.18 ±0.07

1.07 ±0.11

0.67 ±0.17

-0.05 ±0.30

-0.66 ±0.08

-1.69 ±0.15

0

1

2

3

4

5

6

0.00 ±0.00

-6.43 ±0.37

-8.82 ±0.04

-8.70 ±0.17

-9.37 ±0.31

-10.14 ±0.27

-13.61 ±0.26

-7.04 ±0.34

-9.38 ±0.42

-10.35 ±0.32

-10.51 ±0.54

-11.20 ±0.46

-14.42 ±0.37

-8.13 ±0.10

-10.58 ±0.05

-10.90 ±0.17

-11.38 ±0.14

-12.48 ±0.25

-16.62 ±0.27

Molecular Substructure

(e) Chemberta (PT) AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen pyridine thiazole thiophene NHOH HAcceptors HDonors Al_COO Al_OH Al_OH_noTert

0.00 ±0.00

-0.36 ±0.07

-1.55 ±0.14

-1.89 ±0.16

-1.89 ±0.07

-1.71 ±0.08

-1.66 ±0.18

0.00 ±0.00

-3.13 ±0.19

-5.74 ±0.16

-8.09 ±0.25

-5.88 ±0.14

-6.17 ±0.34

-5.14 ±0.13

0.00 ±0.00

-0.41 ±0.08

-0.94 ±0.18

-1.35 ±0.05

-1.24 ±0.04

-1.37 ±0.14

-0.83 ±0.14

0.00 ±0.00

-2.23 ±0.38

-4.36 ±0.50

-4.22 ±0.28

-2.61 ±0.49

-3.56 ±0.41

-2.72 ±0.21

0.00 ±0.00

-2.58 ±0.23

-4.64 ±0.32

-4.73 ±0.26

-3.02 ±0.29

-3.99 ±0.10

-3.68 ±0.37

0.00 ±0.00

-2.53 ±0.09

-5.34 ±0.39

-5.24 ±0.02

-3.90 ±0.19

-4.16 ±0.31

-4.05 ±0.21

0.00 ±0.00

-0.29 ±0.04

-1.64 ±0.13

-1.95 ±0.22

-1.91 ±0.13

-1.66 ±0.12

-1.64 ±0.04

0.00 ±0.00

-3.74 ±0.31

-10.26 ±0.06

-8.51 ±0.48

-6.93 ±0.51

-5.99 ±0.51

-5.40 ±0.20

0.00 ±0.00

-3.42 ±0.36

-4.07 ±0.55

-4.70 ±0.64

-1.32 ±0.17

-1.86 ±0.13

-2.27 ±0.06

0.00 ±0.00

-0.63 ±0.51

-2.00 ±0.47

-3.73 ±0.35

-3.00 ±1.11

-3.27 ±1.29

-1.84 ±1.23

0.00 ±0.00

-3.31 ±0.66

-6.60 ±0.82

-4.32 ±0.54

-2.96 ±0.11

-1.40 ±0.42

0.96 ±0.59

0.00 ±0.00

-2.00 ±0.20

-1.17 ±0.36

-1.84 ±0.41

-0.71 ±0.33

0.45 ±0.23

1.19 ±0.47

0.00 ±0.00

-1.53 ±0.39

-0.39 ±0.50

-3.71 ±0.22

-3.68 ±1.35

-3.70 ±0.90

-2.57 ±0.92

0.00 ±0.00

-1.96 ±0.33

-1.07 ±0.41

-1.78 ±0.26

-0.63 ±0.11

0.34 ±0.25

0.86 ±0.35

0.00 ±0.00

-1.52 ±0.05

-2.14 ±0.40

-6.24 ±0.62

-0.95 ±1.23

4.41 ±0.98

5.30 ±0.90

0.00 ±0.00

-2.10 ±0.21

-1.61 ±0.14

-3.22 ±0.40

-2.30 ±0.45

1.63 ±0.21

2.27 ±0.36

0.00 ±0.00

-1.83 ±0.25

-1.73 ±0.27

-2.96 ±0.36

-1.95 ±0.20

1.32 ±0.34

2.03 ±0.23

0

1

2

3

4

5

6

Ar_COO Ar_NH Ar_OH COO COO2 NH0 NH1 NH2 amide aniline ether morpholine para_hydroxylation phenol phenol_noOrthoHbond piperdine piperzine priamide NO Heteroatoms

-0.04 ±0.15

0.23 ±0.10

0.21 ±0.10

0.16 ±0.16

-0.05 ±0.21

-0.11 ±0.30

-0.33 ±0.20

-0.51 ±0.24

-0.59 ±0.41

-0.89 ±0.23

-1.18 ±0.13

-0.75 ±0.27

-0.10 ±0.06

-0.40 ±0.10

-0.22 ±0.05

-0.73 ±0.09

-1.65 ±0.26

-2.37 ±0.34

-5.04 ±0.56

-8.50 ±0.52

-13.55 ±0.69

-18.34 ±0.13

-25.85 ±0.46

0.00 ±0.00

0.92 ±0.28

2.13 ±0.24

2.03 ±0.15

1.82 ±0.52

1.41 ±0.31

1.06 ±0.12

1.12 ±0.12

0.97 ±0.13

0.81 ±0.26

0.53 ±0.20

-0.43 ±0.23

-0.87 ±0.20

0.00 ±0.00

0.18 ±0.29

0.29 ±0.61

0.23 ±0.55

0.19 ±0.41

0.19 ±0.37

-0.26 ±0.41

-0.42 ±0.63

-0.90 ±0.46

-0.77 ±0.41

-1.01 ±0.19

-1.45 ±0.26

-1.59 ±0.17

0.00 ±0.00

0.02 ±0.30

0.48 ±0.39

0.24 ±0.54

0.01 ±0.50

-0.19 ±0.34

-0.21 ±0.31

-0.58 ±0.54

-0.92 ±0.35

-1.17 ±0.38

-1.24 ±0.29

-1.85 ±0.34

-2.10 ±0.19

0.00 ±0.00

-0.12 ±0.10

0.19 ±0.29

0.47 ±0.29

0.81 ±0.18

0.96 ±0.21

0.84 ±0.05

0.69 ±0.16

0.45 ±0.09

0.12 ±0.11

-0.34 ±0.08

-0.96 ±0.16

-1.99 ±0.25

0.00 ±0.00

-0.10 ±0.25

-0.45 ±0.18

-0.27 ±0.49

-0.40 ±0.27

-0.25 ±0.05

-0.40 ±0.42

-0.61 ±0.32

-1.06 ±0.28

-1.61 ±0.20

-2.36 ±0.30

-3.07 ±0.42

-4.12 ±0.38

0.00 ±0.00

-0.25 ±0.39

0.66 ±0.15

0.78 ±0.27

0.52 ±0.47

0.61 ±0.27

0.56 ±0.26

0.64 ±0.11

0.24 ±0.05

0.47 ±0.28

0.22 ±0.09

-0.46 ±0.23

-1.36 ±0.32

0.00 ±0.00

0.11 ±0.27

2.12 ±0.56

2.65 ±0.19

3.49 ±0.40

3.52 ±0.58

3.57 ±0.50

3.60 ±0.29

3.46 ±0.26

2.66 ±0.09

2.33 ±0.31

1.99 ±0.16

1.35 ±0.40

0.00 ±0.00

-0.74 ±0.34

-0.98 ±0.31

0.18 ±0.25

0.35 ±0.33

0.66 ±0.56

0.72 ±0.54

0.14 ±0.34

-0.21 ±0.19

-0.66 ±0.54

-1.12 ±0.59

-1.78 ±0.55

-2.56 ±0.52

0.00 ±0.00

-0.29 ±0.41

1.27 ±0.09

1.35 ±0.16

1.42 ±0.20

1.37 ±0.04

1.30 ±0.11

1.13 ±0.18

1.20 ±0.09

1.06 ±0.27

0.48 ±0.03

0.21 ±0.14

-0.26 ±0.22

0.00 ±0.00

1.49 ±0.30

1.07 ±0.29

1.20 ±0.49

0.77 ±0.74

0.49 ±0.51

0.26 ±0.45

0.02 ±0.46

-0.07 ±0.34

-0.24 ±0.30

-0.15 ±0.29

-0.34 ±0.28

-0.67 ±0.45

0.00 ±0.00

-0.02 ±0.13

-0.08 ±0.14

-0.06 ±0.11

-0.27 ±0.16

-0.17 ±0.22

-0.12 ±0.31

-0.21 ±0.17

-0.41 ±0.18

-0.49 ±0.27

-0.78 ±0.31

-1.16 ±0.15

-1.82 ±0.11

0.00 ±0.00

0.50 ±0.07

1.80 ±0.15

1.82 ±0.13

1.84 ±0.08

1.41 ±0.28

1.06 ±0.14

0.99 ±0.11

0.90 ±0.25

0.73 ±0.30

0.35 ±0.31

-0.11 ±0.44

-0.95 ±0.18

0.00 ±0.00

0.43 ±0.31

1.64 ±0.19

1.66 ±0.29

1.19 ±0.17

1.17 ±0.35

0.92 ±0.41

0.70 ±0.26

0.47 ±0.48

0.30 ±0.54

0.20 ±0.42

-0.32 ±0.40

-0.95 ±0.21

0.00 ±0.00

0.12 ±0.13

-0.17 ±0.66

0.23 ±0.18

0.13 ±0.24

0.21 ±0.25

0.38 ±0.16

0.27 ±0.46

0.33 ±0.31

0.05 ±0.31

-0.14 ±0.53

-0.30 ±0.62

-0.55 ±0.93

0.00 ±0.00

-0.02 ±0.29

0.35 ±0.47

-0.59 ±0.19

-0.45 ±0.44

-0.35 ±0.38

-0.19 ±0.68

-0.38 ±0.61

-0.64 ±0.70

-0.64 ±0.52

-1.39 ±0.44

-1.22 ±0.99

-1.42 ±0.82

0.00 ±0.00

-0.68 ±0.56

0.02 ±0.51

0.51 ±0.50

0.50 ±0.32

0.04 ±0.73

0.21 ±0.45

-0.04 ±0.61

-0.12 ±0.44

-0.39 ±0.43

-0.54 ±0.38

-0.76 ±0.55

-1.15 ±0.51

0.00 ±0.00

-1.96 ±0.59

-1.08 ±0.13

-1.28 ±0.12

-3.01 ±0.52

-5.53 ±0.68

-7.33 ±0.97

-9.02 ±0.84

-11.08 ±0.69

-13.12 ±0.96

-15.30 ±0.68

-18.00 ±0.49

-20.50 ±0.53

0.00 ±0.00

-1.28 ±0.64

-1.26 ±0.31

-2.54 ±0.64

-3.84 ±0.22

-5.89 ±0.31

-7.30 ±0.79

-8.23 ±0.70

-9.77 ±0.25

-11.19 ±1.61

-13.58 ±1.51

-14.77 ±1.55

-16.13 ±1.21

0

1

2

3

4

5

6

7

8

9

10

11

12

Layer

RotatableBonds Ring ArN Ar_N C_O C_O_noCOO Imine Ndealkylation1 Ndealkylation2 Nhpyrrole alkyl_halide aryl_methyl bicyclic imidazole ketone ketone_Topliss methoxy nitrile sulfide urea

0.00 ±0.00

-1.95 ±0.09

-0.94 ±0.31

0.08 ±0.32

-0.37 ±0.16

0.01 ±0.40

-0.38 ±0.57

-0.35 ±0.44

-0.48 ±0.40

-0.51 ±0.74

-0.79 ±0.84

-1.48 ±0.65

-2.46 ±0.78

0.00 ±0.00

-0.05 ±0.06

-0.04 ±0.02

-0.00 ±0.02

-0.07 ±0.04

-0.05 ±0.03

-0.02 ±0.01

-0.07 ±0.02

-0.11 ±0.03

-0.15 ±0.04

-0.19 ±0.04

-0.16 ±0.01

-0.22 ±0.01

0.00 ±0.00

-0.16 ±0.08

0.67 ±0.16

0.73 ±0.64

0.47 ±0.49

0.44 ±0.52

0.30 ±0.59

0.26 ±0.23

0.04 ±0.36

0.04 ±0.61

-0.65 ±0.54

-1.10 ±0.72

-2.08 ±0.60

0.00 ±0.00

-0.04 ±0.06

-0.10 ±0.03

-0.06 ±0.07

-0.07 ±0.05

-0.26 ±0.05

-0.27 ±0.02

-0.46 ±0.08

-0.67 ±0.15

-0.89 ±0.15

-1.44 ±0.05

-2.20 ±0.05

-3.11 ±0.14

0.00 ±0.00

-0.76 ±0.23

-0.11 ±0.17

0.01 ±0.05

-0.22 ±0.09

-0.13 ±0.29

-0.19 ±0.40

-0.30 ±0.15

-0.53 ±0.17

-0.76 ±0.17

-1.13 ±0.14

-1.39 ±0.14

-1.83 ±0.17

0.00 ±0.00

-0.10 ±0.15

0.85 ±0.08

0.66 ±0.02

0.83 ±0.22

1.16 ±0.23

0.96 ±0.22

0.83 ±0.08

0.79 ±0.16

0.50 ±0.05

0.25 ±0.14

0.17 ±0.06

-0.23 ±0.15

0.00 ±0.00

0.22 ±0.28

0.43 ±0.09

0.95 ±0.09

1.08 ±0.24

0.94 ±0.45

1.09 ±0.38

0.97 ±0.17

1.05 ±0.06

0.96 ±0.19

0.81 ±0.25

0.54 ±0.23

0.03 ±0.31

0.00 ±0.00

-0.30 ±0.41

0.88 ±0.15

0.68 ±0.42

0.37 ±0.40

0.61 ±0.55

0.33 ±0.32

0.25 ±0.42

0.05 ±0.58

0.00 ±0.30

0.00 ±0.00

-0.22 ±0.21

-0.09 ±0.68

-0.48 ±0.44

-0.41 ±0.52

0.04 ±0.67

0.31 ±0.56

0.42 ±0.54

0.33 ±0.21

0.01 ±0.40

-0.58 ±0.32

-0.53 ±0.57

-0.68 ±0.54

0.00 ±0.00

-0.75 ±0.27

-0.10 ±0.06

-0.40 ±0.10

-0.22 ±0.05

-0.73 ±0.09

-1.65 ±0.26

-2.37 ±0.34

-5.04 ±0.56

-8.50 ±0.52

-13.55 ±0.69

-18.34 ±0.13

-25.85 ±0.46

0.00 ±0.00

-0.22 ±0.50

-0.06 ±0.12

0.20 ±0.26

0.47 ±0.14

0.36 ±0.05

0.43 ±0.17

0.37 ±0.26

0.34 ±0.48

-0.09 ±0.59

-0.19 ±0.40

-0.94 ±0.54

-1.53 ±0.44

0.00 ±0.00

-1.61 ±1.24

0.17 ±0.26

0.22 ±0.18

-0.11 ±0.30

-0.12 ±0.43

-0.09 ±0.22

0.16 ±0.33

-0.09 ±0.43

-0.79 ±0.31

-0.96 ±0.37

-1.98 ±0.04

-2.50 ±0.29

-0.53 ±0.24

-0.88 ±0.48

-1.12 ±0.27

0.00 ±0.00

0.13 ±0.11

0.38 ±0.10

0.50 ±0.14

0.60 ±0.07

0.72 ±0.15

0.46 ±0.22

0.67 ±0.24

0.49 ±0.14

0.21 ±0.10

-0.17 ±0.10

-0.56 ±0.03

-1.40 ±0.14

0.00 ±0.00

-0.09 ±0.04

0.15 ±0.16

-0.01 ±0.21

-0.19 ±0.12

-0.24 ±0.29

-0.49 ±0.21

-0.54 ±0.31

-0.71 ±0.22

-0.97 ±0.38

-1.65 ±0.13

-2.08 ±0.12

-2.91 ±0.06

0.00 ±0.00

0.04 ±0.24

0.11 ±0.10

0.10 ±0.16

0.04 ±0.27

-0.14 ±0.29

-0.32 ±0.38

-0.01 ±0.47

-0.17 ±0.34

-0.27 ±0.52

-0.45 ±0.43

-0.71 ±0.33

-0.63 ±0.33

0.00 ±0.00

-0.13 ±0.28

0.13 ±0.23

0.10 ±0.37

-0.36 ±0.26

-0.62 ±0.34

-0.56 ±0.26

-0.58 ±0.41

-0.74 ±0.50

-0.67 ±0.31

-0.61 ±0.19

-0.70 ±0.15

-0.55 ±0.13

0.00 ±0.00

-0.81 ±0.24

1.78 ±0.13

1.37 ±0.22

1.31 ±0.02

1.35 ±0.63

1.39 ±0.45

0.90 ±0.62

1.00 ±0.61

0.61 ±0.51

-0.05 ±0.30

-0.51 ±0.60

-1.13 ±0.41

0.00 ±0.00

-0.31 ±0.09

-0.39 ±0.09

-0.40 ±0.24

-0.21 ±0.17

-0.06 ±0.37

0.01 ±0.21

-0.07 ±0.13

-0.87 ±0.19

-1.50 ±0.43

-2.92 ±0.28

-4.40 ±0.06

-7.29 ±0.07

0.00 ±0.00

-0.08 ±0.12

-0.20 ±0.07

-0.21 ±0.10

-0.22 ±0.09

0.08 ±0.11

0.00 ±0.07

0.02 ±0.04

-0.02 ±0.15

-0.21 ±0.20

-0.54 ±0.22

-1.13 ±0.21

-2.13 ±0.21

0.00 ±0.00

-0.26 ±0.33

0.44 ±0.45

0.89 ±0.31

0.51 ±0.73

0.64 ±0.45

0.71 ±0.55

0.19 ±0.74

0.18 ±0.82

-0.19 ±0.55

-0.24 ±0.60

-0.11 ±1.00

-0.19 ±0.78

0

1

2

3

4

5

6

7

8

9

10 11 12

10 5 0 5 10 15 20 25

0.00 1.30 1.47 1.42 1.84 2.75 2.49 2.06 1.86 1.15 0.97 0.77 0.52 -0.19 -0.58 ±0.00 ±0.23 ±0.35 ±0.33 ±0.21 ±0.16 ±0.27 ±0.11 ±0.08 ±0.18 ±0.11 ±0.03 ±0.03 ±0.14 ±0.20 0.00 0.27 0.28 0.32 0.34 -0.27 0.01 -0.14 -0.05 -0.43 -0.82 -1.30 -1.71 -1.98 -2.36 ±0.00 ±0.18 ±0.24 ±0.17 ±0.29 ±0.19 ±0.20 ±0.07 ±0.19 ±0.08 ±0.27 ±0.28 ±0.17 ±0.49 ±0.47 0.00 1.69 2.69 2.04 2.43 2.54 2.24 1.92 1.94 1.52 1.47 1.08 0.50 -0.64 -1.40 ±0.00 ±0.25 ±0.19 ±0.20 ±0.23 ±0.23 ±0.42 ±0.32 ±0.24 ±0.10 ±0.41 ±0.52 ±0.31 ±0.40 ±0.31 0.00 1.92 1.54 1.22 1.15 1.13 0.93 0.46 0.39 -0.10 -0.25 -0.42 -0.81 -1.12 -1.62 ±0.00 ±0.15 ±0.12 ±0.01 ±0.19 ±0.23 ±0.18 ±0.10 ±0.17 ±0.23 ±0.50 ±0.41 ±0.49 ±0.36 ±0.33 0.00 2.10 1.41 1.18 1.07 1.08 0.85 0.62 0.52 0.09 -0.12 -0.32 -0.72 -0.89 -1.49 ±0.00 ±0.33 ±0.26 ±0.17 ±0.20 ±0.24 ±0.20 ±0.08 ±0.09 ±0.20 ±0.31 ±0.42 ±0.48 ±0.30 ±0.18 0.00 -0.19 -0.17 -0.22 -0.28 -0.37 -0.61 -0.90 -0.76 -0.98 -1.23 -1.84 -2.21 -2.73 -2.68 ±0.00 ±0.11 ±0.15 ±0.10 ±0.18 ±0.17 ±0.16 ±0.05 ±0.11 ±0.18 ±0.16 ±0.17 ±0.08 ±0.09 ±0.27 0.00 -0.51 -0.50 -0.08 -0.39 -0.61 -0.76 -1.13 -1.19 -1.43 -1.98 -3.11 -3.34 -3.67 -4.83 ±0.00 ±0.24 ±0.20 ±0.11 ±0.05 ±0.13 ±0.36 ±0.11 ±0.24 ±0.18 ±0.38 ±0.13 ±0.10 ±0.17 ±0.13 0.00 0.54 1.34 2.09 1.55 1.20 1.07 1.17 0.92 0.34 0.31 0.38 0.17 -0.58 -1.54 ±0.00 ±0.07 ±0.17 ±0.11 ±0.08 ±0.31 ±0.20 ±0.12 ±0.12 ±0.30 ±0.06 ±0.27 ±0.03 ±0.25 ±0.45 0.00 0.20 0.70 0.83 1.02 0.72 0.67 0.50 0.20 -0.32 -0.35 -1.00 -1.44 -1.46 -2.12 ±0.00 ±0.19 ±0.10 ±0.17 ±0.06 ±0.16 ±0.09 ±0.18 ±0.05 ±0.11 ±0.22 ±0.34 ±0.19 ±0.18 ±0.15 0.00 0.74 1.02 0.98 1.15 1.16 0.75 0.74 0.80 0.45 0.35 0.40 -0.07 -0.45 -1.28 ±0.00 ±0.38 ±0.06 ±0.16 ±0.29 ±0.28 ±0.15 ±0.08 ±0.27 ±0.11 ±0.16 ±0.17 ±0.39 ±0.38 ±0.32 0.00 -0.82 -0.37 -0.47 -0.23 -0.34 -0.37 -0.48 -0.89 -1.03 -1.31 -1.62 -1.57 -1.93 -2.37 ±0.00 ±0.08 ±0.26 ±0.28 ±0.27 ±0.13 ±0.18 ±0.21 ±0.11 ±0.43 ±0.39 ±0.08 ±0.14 ±0.03 ±0.39 0.00 -0.13 0.58 0.45 0.28 -0.54 -0.57 -0.18 -0.45 -0.46 -0.76 -0.33 -0.60 -1.25 -1.23 ±0.00 ±0.25 ±0.49 ±0.38 ±0.44 ±0.22 ±0.17 ±0.31 ±0.35 ±0.16 ±0.23 ±0.23 ±0.31 ±0.21 ±0.42 0.00 0.40 0.36 0.26 0.29 0.15 0.02 -0.01 -0.01 -0.17 -0.18 -0.49 -0.55 -0.83 -0.74 ±0.00 ±0.21 ±0.09 ±0.05 ±0.04 ±0.07 ±0.04 ±0.10 ±0.04 ±0.12 ±0.09 ±0.25 ±0.09 ±0.12 ±0.09 0.00 1.46 2.25 1.70 1.80 1.98 1.93 1.46 1.67 1.23 1.34 1.07 0.38 -0.11 -1.01 ±0.00 ±0.15 ±0.09 ±0.39 ±0.25 ±0.24 ±0.27 ±0.48 ±0.22 ±0.30 ±0.30 ±0.16 ±0.08 ±0.07 ±0.18 0.00 1.50 2.29 1.67 1.66 1.87 1.81 1.36 1.31 1.27 1.19 0.99 0.54 -0.14 -0.98 ±0.00 ±0.23 ±0.04 ±0.12 ±0.23 ±0.18 ±0.19 ±0.40 ±0.16 ±0.29 ±0.38 ±0.29 ±0.30 ±0.21 ±0.47 0.00 -0.27 0.02 0.17 0.20 -0.31 -0.12 -0.42 -0.39 -0.20 -0.94 -1.00 -1.56 -1.76 -2.30 ±0.00 ±0.33 ±0.15 ±0.40 ±0.18 ±0.48 ±0.35 ±0.33 ±0.14 ±0.16 ±0.27 ±0.28 ±0.86 ±0.15 ±0.16 0.00 1.05 1.64 0.98 0.57 -0.07 0.07 0.06 -0.60 -0.58 -1.63 -2.04 -2.50 -3.95 -4.36 ±0.00 ±0.14 ±0.21 ±0.55 ±0.11 ±0.64 ±0.58 ±0.51 ±0.59 ±0.54 ±0.39 ±0.55 ±0.40 ±0.15 ±0.27 0.00 2.44 2.45 2.59 2.23 1.72 1.40 1.22 0.73 0.22 -0.11 0.05 -0.44 -0.97 -1.05 ±0.00 ±0.11 ±0.41 ±0.30 ±0.59 ±0.32 ±0.28 ±0.18 ±0.36 ±0.20 ±0.41 ±0.23 ±0.11 ±0.41 ±0.76 0.00 1.36 1.66 1.44 1.50 1.31 1.16 0.87 0.56 -0.33 -0.90 -1.39 -1.99 -2.71 -3.35 ±0.00 ±0.54 ±0.70 ±0.44 ±0.37 ±0.29 ±0.30 ±0.29 ±0.16 ±0.32 ±0.21 ±0.12 ±0.08 ±0.20 ±0.13 0.00 1.21 1.86 1.77 2.09 1.33 1.28 0.59 0.19 0.04 -0.76 -1.57 -2.37 -3.59 -4.36 ±0.00 ±0.04 ±0.39 ±0.33 ±0.30 ±0.43 ±0.63 ±0.30 ±0.09 ±0.33 ±0.52 ±0.62 ±0.48 ±0.43 ±0.24

0

1

2

3

4

5

Layer

6

7

8

9

RotatableBonds Ring ArN Ar_N C_O C_O_noCOO Imine Ndealkylation1 Ndealkylation2 Nhpyrrole alkyl_halide aryl_methyl bicyclic imidazole ketone ketone_Topliss methoxy nitrile sulfide urea

10 11 12 13 14

0.00 0.13 0.35 0.34 0.16 0.59 0.23 0.39 0.53 0.26 -0.51 -0.65 -1.11 -1.27 -1.63 ±0.00 ±0.45 ±0.27 ±0.09 ±0.23 ±0.09 ±0.19 ±0.26 ±0.38 ±0.23 ±0.19 ±0.25 ±0.06 ±0.27 ±0.23 0.00 -0.01 -0.02 -0.01 0.00 -0.01 0.00 0.02 -0.01 -0.00 -0.05 -0.03 -0.02 -0.02 -0.02 ±0.00 ±0.01 ±0.04 ±0.01 ±0.03 ±0.02 ±0.01 ±0.00 ±0.01 ±0.01 ±0.05 ±0.02 ±0.05 ±0.01 ±0.02 0.00 1.26 2.45 2.72 1.86 1.98 1.66 2.07 2.00 1.43 1.39 1.14 0.59 -0.10 -1.31 ±0.00 ±0.21 ±0.23 ±0.27 ±0.18 ±0.52 ±0.08 ±0.31 ±0.34 ±0.25 ±0.25 ±0.66 ±0.70 ±0.59 ±0.47 0.00 -1.00 -0.72 -0.44 -0.50 -0.61 -0.83 -1.00 -1.21 -1.12 -1.67 -2.09 -2.42 -2.79 -3.01 ±0.00 ±0.09 ±0.21 ±0.05 ±0.08 ±0.22 ±0.10 ±0.13 ±0.07 ±0.07 ±0.05 ±0.13 ±0.16 ±0.05 ±0.10

20

0.00 0.03 0.19 0.20 0.21 0.15 0.12 -0.23 -0.22 -0.35 -0.31 -0.56 -0.99 -1.30 -1.53 ±0.00 ±0.17 ±0.16 ±0.20 ±0.04 ±0.08 ±0.10 ±0.09 ±0.07 ±0.25 ±0.05 ±0.14 ±0.10 ±0.17 ±0.12 0.00 0.20 0.61 0.71 0.99 0.87 0.80 0.23 0.54 0.62 0.73 0.38 0.26 -0.12 -0.62 ±0.00 ±0.25 ±0.11 ±0.11 ±0.09 ±0.05 ±0.17 ±0.28 ±0.30 ±0.18 ±0.09 ±0.20 ±0.14 ±0.03 ±0.14 0.00 0.19 0.62 0.54 0.35 0.25 0.12 -0.44 -0.24 -0.56 -0.56 -1.09 -1.15 -1.58 -2.15 ±0.00 ±0.39 ±0.10 ±0.35 ±0.32 ±0.07 ±0.20 ±0.07 ±0.23 ±0.21 ±0.12 ±0.09 ±0.43 ±0.46 ±0.18 0.00 0.25 0.44 0.48 0.59 0.06 -0.06 0.51 0.23 0.62 0.34 -0.19 -0.60 -1.01 -0.98 ±0.00 ±0.48 ±0.59 ±0.34 ±0.07 ±0.41 ±0.23 ±0.27 ±0.25 ±0.31 ±0.30 ±0.45 ±0.14 ±0.30 ±0.22 0.00 -0.44 -0.14 -0.23 -0.31 -0.04 -0.14 -0.68 -0.74 -0.43 -1.00 -1.00 -0.95 -1.57 -2.30 ±0.00 ±0.16 ±0.09 ±0.25 ±0.10 ±0.31 ±0.31 ±0.32 ±0.17 ±0.51 ±0.20 ±0.37 ±0.29 ±0.44 ±0.23

15 10

0.00 0.27 0.28 0.32 0.34 -0.27 0.01 -0.14 -0.05 -0.43 -0.82 -1.30 -1.71 -1.98 -2.36 ±0.00 ±0.18 ±0.24 ±0.17 ±0.29 ±0.19 ±0.20 ±0.07 ±0.19 ±0.08 ±0.27 ±0.28 ±0.17 ±0.49 ±0.47 0.00 -1.32 -0.78 -0.32 -0.31 -0.06 -0.63 -0.34 -0.47 -0.42 -1.49 -1.54 -1.39 -1.89 -3.02 ±0.00 ±0.65 ±0.26 ±0.64 ±0.45 ±0.73 ±0.43 ±0.25 ±0.54 ±0.40 ±0.35 ±0.12 ±0.10 ±0.19 ±0.35 0.00 1.76 1.16 0.97 1.13 0.89 0.80 1.13 0.88 0.83 0.53 0.34 0.19 0.06 -0.41 ±0.00 ±0.11 ±0.21 ±0.09 ±0.06 ±0.16 ±0.34 ±0.09 ±0.16 ±0.28 ±0.14 ±0.16 ±0.39 ±0.41 ±0.34 0.00 0.73 0.94 0.40 0.69 0.52 0.36 0.30 0.16 0.06 -0.11 -0.21 -0.42 -0.72 -0.81 ±0.00 ±0.14 ±0.10 ±0.11 ±0.19 ±0.10 ±0.13 ±0.09 ±0.02 ±0.12 ±0.24 ±0.22 ±0.27 ±0.05 ±0.23 0.00 -0.55 -0.20 -0.16 -0.16 -0.22 -0.84 -1.48 -1.14 -1.91 -3.05 -2.98 -3.52 -4.07 -4.95 ±0.00 ±0.26 ±0.33 ±0.47 ±0.18 ±0.50 ±0.24 ±0.43 ±0.48 ±0.42 ±0.32 ±0.35 ±0.63 ±0.47 ±0.35

5 0

0.00 0.50 0.57 0.81 0.71 0.77 0.67 0.50 0.78 0.55 0.50 0.28 -0.08 -0.38 -0.69 ±0.00 ±0.17 ±0.22 ±0.15 ±0.04 ±0.17 ±0.45 ±0.23 ±0.28 ±0.25 ±0.24 ±0.26 ±0.45 ±0.33 ±0.28

5

0.00 0.66 0.59 0.87 0.81 0.75 0.61 0.37 0.64 0.66 0.53 0.22 -0.10 -0.22 -0.73 ±0.00 ±0.21 ±0.19 ±0.21 ±0.23 ±0.07 ±0.33 ±0.32 ±0.22 ±0.27 ±0.43 ±0.52 ±0.50 ±0.32 ±0.38 0.00 -0.27 1.09 1.71 1.65 1.06 1.08 1.33 1.50 1.67 1.22 1.50 1.43 1.12 0.84 ±0.00 ±0.44 ±0.33 ±0.51 ±0.44 ±0.23 ±0.29 ±0.22 ±0.05 ±0.25 ±0.13 ±0.35 ±0.22 ±0.26 ±0.35 0.00 0.85 0.87 1.55 1.64 1.61 1.29 0.95 0.60 0.22 -0.23 -0.25 -0.69 -1.39 -1.83 ±0.00 ±0.46 ±0.33 ±0.29 ±0.25 ±0.37 ±0.19 ±0.08 ±0.17 ±0.12 ±0.29 ±0.47 ±0.43 ±0.54 ±0.41

10

0.00 -0.71 -0.74 -1.48 -1.25 -1.68 -2.58 -2.74 -3.12 -3.24 -3.15 -3.88 -4.59 -4.41 -5.39 ±0.00 ±0.31 ±0.29 ±0.32 ±0.42 ±0.14 ±0.15 ±0.17 ±0.35 ±0.29 ±0.26 ±0.36 ±0.22 ±0.32 ±0.45 0.00 0.31 0.53 0.56 0.61 0.59 0.45 0.09 -0.19 -0.45 -0.60 -0.37 -0.51 -0.90 -1.06 ±0.00 ±0.20 ±0.29 ±0.43 ±0.31 ±0.20 ±0.17 ±0.16 ±0.08 ±0.19 ±0.34 ±0.40 ±0.52 ±0.38 ±0.41

0

1

2

3

4

5

6

7

8

9 10 11 12 13 14

Ar_COO Ar_NH Ar_OH COO COO2 NH0 NH1 NH2 amide aniline ether morpholine para_hydroxylation phenol phenol_noOrthoHbond piperdine piperzine priamide NO Heteroatoms

0.00 ±0.00

1.67 ±0.09

1.80 ±0.27

2.42 ±0.33

2.44 ±0.34

1.49 ±0.41

1.38 ±0.29

0.00 ±0.00

-1.91 ±0.87

-2.05 ±0.25

-2.08 ±0.14

-2.73 ±0.24

-2.77 ±0.12

-3.18 ±0.19

0.00 ±0.00

1.78 ±0.15

2.12 ±0.33

2.14 ±0.39

2.16 ±0.23

2.05 ±0.19

1.11 ±0.20

0.00 ±0.00

2.66 ±0.40

2.43 ±0.22

1.47 ±0.06

1.29 ±0.31

0.86 ±0.27

0.17 ±0.06

0.00 ±0.00

2.94 ±0.56

2.30 ±0.18

1.28 ±0.19

1.37 ±0.12

0.95 ±0.25

0.32 ±0.29

0.00 ±0.00

-0.69 ±0.17

-0.96 ±0.30

-1.24 ±0.20

-2.07 ±0.11

-2.53 ±0.16

-3.01 ±0.20

0.00 ±0.00

-1.16 ±0.51

-1.03 ±0.49

-1.99 ±0.13

-3.02 ±0.49

-4.08 ±0.23

-4.19 ±0.26

0.00 ±0.00

3.43 ±0.04

3.60 ±0.24

2.69 ±0.21

2.23 ±0.29

1.70 ±0.23

1.27 ±0.39

0.00 ±0.00

1.68 ±0.24

1.74 ±0.16

1.52 ±0.17

1.15 ±0.25

0.38 ±0.55

0.07 ±0.16

0.00 ±0.00

0.70 ±0.35

1.65 ±0.28

0.92 ±0.14

0.75 ±0.11

-0.02 ±0.13

-0.36 ±0.24

0.00 ±0.00

-2.35 ±0.31

-1.61 ±0.28

-1.74 ±0.11

-2.20 ±0.41

-3.62 ±0.22

-3.41 ±0.38

0.00 ±0.00

-0.40 ±0.54

-0.90 ±0.79

-1.75 ±0.62

-2.22 ±0.58

-3.41 ±0.50

-3.37 ±0.47

0.00 ±0.00

-0.14 ±0.11

-0.11 ±0.05

0.14 ±0.32

0.01 ±0.06

-0.21 ±0.27

-0.41 ±0.22

0.00 ±0.00

1.30 ±0.45

1.70 ±0.12

1.55 ±0.44

1.47 ±0.27

1.34 ±0.21

0.51 ±0.29

0.00 ±0.00

1.13 ±0.04

1.44 ±0.15

1.53 ±0.42

1.64 ±0.27

1.14 ±0.31

0.34 ±0.43

0.00 ±0.00

-0.58 ±0.37

-0.17 ±0.17

-0.43 ±0.29

-1.08 ±0.25

-0.92 ±0.16

-1.23 ±0.45

0.00 ±0.00

-1.07 ±0.43

-0.68 ±0.25

-1.60 ±0.87

-2.74 ±1.25

-2.77 ±0.78

-3.01 ±1.63

0.00 ±0.00

5.80 ±0.24

6.42 ±0.49

4.92 ±0.65

4.78 ±0.96

3.83 ±0.50

3.29 ±0.31

0.00 ±0.00

0.86 ±0.20

0.97 ±0.24

-0.16 ±0.24

-0.36 ±0.40

-1.79 ±0.80

-2.71 ±0.71

0.00 ±0.00

1.98 ±0.41

1.47 ±0.57

0.08 ±0.52

-0.97 ±0.21

-1.33 ±0.26

-1.70 ±0.27

0

1

2

3

4

5

6

Layer

RotatableBonds Ring ArN Ar_N C_O C_O_noCOO Imine Ndealkylation1 Ndealkylation2 Nhpyrrole alkyl_halide aryl_methyl bicyclic imidazole ketone ketone_Topliss methoxy nitrile sulfide urea

0.00 ±0.00

0.54 ±0.22

0.83 ±0.15

-0.17 ±0.18

-0.59 ±0.11

-1.51 ±0.37

-2.44 ±0.47

0.00 ±0.00

-0.09 ±0.07

-0.09 ±0.05

-0.02 ±0.02

-0.05 ±0.01

-0.09 ±0.06

-0.15 ±0.05

0.00 ±0.00

3.30 ±0.20

6.48 ±0.23

5.27 ±0.16

4.70 ±0.68

3.97 ±0.60

3.70 ±0.37

0.00 ±0.00

-2.63 ±0.14

-2.41 ±0.37

-2.99 ±0.38

-3.84 ±0.14

-4.70 ±0.25

-5.10 ±0.02

0.00 ±0.00

-0.21 ±0.01

-0.33 ±0.11

-0.29 ±0.02

-0.84 ±0.17

-1.29 ±0.10

-1.46 ±0.13

0.00 ±0.00

0.51 ±0.22

0.65 ±0.26

0.75 ±0.09

0.22 ±0.20

-0.49 ±0.17

-0.63 ±0.20

0.00 ±0.00

0.05 ±0.12

-0.11 ±0.14

-0.36 ±0.16

-0.82 ±0.30

-1.12 ±0.16

-1.38 ±0.19

0.00 ±0.00

2.04 ±0.37

2.04 ±0.26

0.95 ±0.53

1.05 ±0.15

0.58 ±0.15

0.18 ±0.50

0.00 ±0.00

-0.70 ±0.24

-1.10 ±0.16

-1.20 ±0.28

-1.69 ±0.59

-1.85 ±0.53

-1.71 ±0.68

0.00 ±0.00

-1.91 ±0.87

-2.05 ±0.25

-2.08 ±0.14

-2.73 ±0.24

-2.77 ±0.12

-3.18 ±0.19

0.00 ±0.00

-2.59 ±0.76

-1.96 ±0.37

-2.99 ±0.56

-2.29 ±0.62

-3.31 ±0.53

-4.51 ±0.28

0.00 ±0.00

0.99 ±0.21

1.64 ±0.43

1.09 ±0.26

0.89 ±0.10

0.46 ±0.12

-0.53 ±0.05

0.00 ±0.00

-0.30 ±0.31

0.02 ±0.04

-0.04 ±0.20

-0.18 ±0.23

-0.33 ±0.14

-0.77 ±0.06

0.00 ±0.00

-2.10 ±0.32

-2.48 ±0.01

-2.55 ±0.48

-3.38 ±0.33

-3.82 ±0.54

-5.07 ±0.20

0.00 ±0.00

0.50 ±0.26

0.70 ±0.23

0.60 ±0.25

0.54 ±0.12

0.26 ±0.11

-0.05 ±0.07

0.00 ±0.00

0.36 ±0.08

0.48 ±0.21

0.30 ±0.29

0.36 ±0.17

0.09 ±0.22

0.19 ±0.18

0.00 ±0.00

1.79 ±0.12

3.89 ±0.13

3.29 ±0.15

3.58 ±0.21

2.38 ±0.16

1.83 ±0.41

0.00 ±0.00

-0.98 ±0.17

0.42 ±0.19

0.08 ±0.58

-0.85 ±0.56

-0.81 ±0.49

-1.09 ±0.59

0.00 ±0.00

-2.58 ±0.12

-4.00 ±0.09

-4.46 ±0.21

-4.72 ±0.33

-5.76 ±0.74

-7.78 ±0.24

0.00 ±0.00

0.36 ±0.24

0.31 ±0.37

0.00 ±0.91

-0.23 ±0.76

-0.17 ±0.70

-0.29 ±0.84

0

1

2

3

4

5

6

20 15 10 5 0 5 10

(f) Chemberta (RI)

0.00 ±0.00

-0.55 ±0.37

-1.62 ±1.07

-5.22 ±0.15

-1.53 ±0.23

0.94 ±0.73

2.68 ±0.95

0.00 ±0.00

-1.37 ±0.60

-5.36 ±0.97

-13.46 ±0.38

-13.02 ±0.85

-6.03 ±0.13

-7.85 ±0.60

0.00 ±0.00

-0.60 ±0.75

-1.70 ±0.35

-5.19 ±0.21

-5.88 ±0.43

-3.03 ±0.57

-1.45 ±0.29

0.00 ±0.00

-0.96 ±0.27

-2.24 ±0.53

-6.87 ±0.63

-0.19 ±0.41

4.33 ±0.52

5.13 ±0.60

0.00 ±0.00

-1.06 ±0.20

-2.21 ±0.17

-6.98 ±0.51

-0.04 ±0.75

4.35 ±0.29

5.64 ±0.97

0.00 ±0.00

-1.54 ±0.23

-1.99 ±0.03

-4.72 ±0.11

-3.14 ±0.17

-3.62 ±0.19

-2.79 ±0.10

0.00 ±0.00

-3.90 ±0.28

-1.12 ±0.06

-1.77 ±0.24

-0.66 ±0.18

-0.41 ±0.35

0.16 ±0.28

0.00 ±0.00

0.83 ±0.07

2.02 ±0.19

2.30 ±0.55

3.32 ±0.38

3.39 ±0.36

4.11 ±0.45

0.00 ±0.00

-1.54 ±0.25

-1.52 ±0.32

-4.50 ±0.68

-2.51 ±0.98

-2.78 ±0.50

-1.33 ±0.37

0.00 ±0.00

-0.88 ±0.17

0.04 ±0.21

-1.22 ±0.28

0.34 ±0.20

-1.09 ±0.15

-1.49 ±0.13

0.00 ±0.00

-2.21 ±0.30

-1.76 ±0.36

-3.83 ±0.33

-1.46 ±0.49

-2.43 ±0.50

-1.91 ±0.46

0.00 ±0.00

-1.78 ±0.23

-4.54 ±0.48

-4.22 ±0.70

-0.81 ±0.31

-0.95 ±0.75

-0.93 ±0.42

0.00 ±0.00

-0.54 ±0.12

-1.13 ±0.09

-2.19 ±0.38

-1.70 ±0.26

-1.22 ±0.17

-0.56 ±0.34

0.00 ±0.00

-0.54 ±0.47

-1.26 ±0.32

-4.90 ±1.17

-5.25 ±0.79

-3.69 ±1.32

-2.43 ±0.18

0.00 ±0.00

-0.83 ±0.68

-1.29 ±0.17

-4.00 ±0.75

-4.41 ±0.91

-2.90 ±1.21

-2.32 ±0.59

0.00 ±0.00

-1.69 ±0.36

-1.19 ±0.32

-1.70 ±0.36

-0.38 ±0.50

-0.89 ±0.10

-0.92 ±0.49

0.00 ±0.00

-2.67 ±0.64

-0.81 ±0.42

-1.59 ±0.32

-1.49 ±0.37

-0.95 ±1.08

-0.33 ±0.84

0.00 ±0.00

-0.37 ±1.08

-0.20 ±0.61

-0.14 ±1.13

3.06 ±0.61

2.05 ±1.26

0.68 ±0.39

0.00 ±0.00

-0.24 ±0.12

0.85 ±0.88

-2.21 ±1.36

-2.50 ±0.36

-2.26 ±1.45

-1.37 ±0.77

0.00 ±0.00

-0.69 ±0.79

-0.68 ±0.80

-2.38 ±0.99

-3.19 ±0.53

-3.03 ±0.24

-2.17 ±0.30

0

1

2

3

4

5

6

Layer

-0.45 ±0.30

0.00 ±0.00

(d) Roberta-zinc-480m (RI) 0.00 ±0.00

Molecular Substructure

Molecular Substructure

0.00 ±0.00

Ar_COO Ar_NH Ar_OH COO COO2 NH0 NH1 NH2 amide aniline ether morpholine para_hydroxylation phenol phenol_noOrthoHbond piperdine piperzine priamide NO Heteroatoms

10 11 12 13 14

(c) Roberta-zinc-480m (PT) AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen pyridine thiazole thiophene NHOH HAcceptors HDonors Al_COO Al_OH Al_OH_noTert

0.00 ±0.00

(b) Molformer(RI)

0.00 -0.12 0.18 -0.09 1.19 0.85 0.26 1.46 3.04 4.47 10.16 12.91 14.73 12.63 11.44 ±0.00 ±0.01 ±0.23 ±0.21 ±0.43 ±0.12 ±0.44 ±0.40 ±0.35 ±0.81 ±1.20 ±0.18 ±1.47 ±1.37 ±1.41

Molecular Substructure

Molecular Substructure

(a) Molformer (PT) AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen pyridine thiazole thiophene NHOH HAcceptors HDonors Al_COO Al_OH Al_OH_noTert

Ar_COO Ar_NH Ar_OH COO COO2 NH0 NH1 NH2 amide aniline ether morpholine para_hydroxylation phenol phenol_noOrthoHbond piperdine piperzine priamide NO Heteroatoms

RotatableBonds Ring ArN Ar_N C_O C_O_noCOO Imine Ndealkylation1 Ndealkylation2 Nhpyrrole alkyl_halide aryl_methyl bicyclic imidazole ketone ketone_Topliss methoxy nitrile sulfide urea

0.00 ±0.00

-2.41 ±0.18

-4.18 ±0.14

-7.12 ±0.13

-4.17 ±0.41

-4.51 ±0.13

-3.59 ±0.05

0.00 ±0.00

0.41 ±0.09

-1.42 ±0.13

-4.42 ±0.22

-4.28 ±0.33

-4.92 ±0.10

-2.12 ±0.15

0.00 ±0.00

-0.36 ±0.21

-2.00 ±0.18

-4.62 ±0.06

-3.26 ±0.05

-3.88 ±0.30

-1.75 ±0.09

0.00 ±0.00

-1.51 ±0.30

-2.75 ±0.14

-3.57 ±0.20

-3.50 ±0.22

-3.61 ±0.27

-2.98 ±0.32

0.00 ±0.00

0.77 ±0.46

1.12 ±0.78

1.08 ±0.18

2.05 ±0.79

0.36 ±0.54

0.77 ±0.48

0.00 ±0.00

-2.81 ±0.17

-3.08 ±0.33

-4.26 ±0.17

-0.87 ±0.56

-1.41 ±0.57

-2.42 ±0.55

0.00 ±0.00

-1.37 ±0.60

-5.36 ±0.97

-13.46 ±0.38

-13.02 ±0.85

-6.03 ±0.13

-7.85 ±0.60

0.00 ±0.00

-1.90 ±0.91

-5.72 ±0.96

-5.09 ±0.26

-4.70 ±0.11

-3.25 ±0.48

-1.49 ±0.36

0.00 ±0.00

0.68 ±0.17

1.24 ±0.19

0.90 ±0.42

-1.33 ±0.39

-1.93 ±0.33

-2.52 ±0.27

0.00 ±0.00

-2.11 ±0.15

-5.22 ±0.11

-6.48 ±0.20

-4.89 ±0.36

-3.43 ±0.19

-2.82 ±0.29

0.00 ±0.00

-1.32 ±0.80

-2.46 ±0.63

-5.49 ±0.70

-4.21 ±0.88

-3.34 ±0.34

-2.45 ±0.67

0.00 ±0.00

-1.69 ±0.28

-1.76 ±0.44

-3.68 ±0.15

-1.96 ±0.48

-2.93 ±0.37

-2.60 ±0.89

0.00 ±0.00

-1.36 ±0.23

-1.66 ±0.57

-3.19 ±0.26

-1.71 ±0.29

-2.21 ±0.29

-2.53 ±0.36

0.00 ±0.00

4.00 ±0.24

9.15 ±0.58

5.44 ±0.89

4.71 ±0.43

2.76 ±0.60

1.94 ±0.32

0.00 ±0.00

-10.10 ±0.29

-10.22 ±0.31

-10.56 ±0.71

-11.07 ±1.19

-9.05 ±1.36

-6.49 ±0.83

0.00 ±0.00

-2.55 ±0.31

-2.73 ±0.63

-3.81 ±0.65

-3.39 ±0.79

-2.42 ±1.05

-1.98 ±0.11

0.00 ±0.00

-0.36 ±0.85

-1.45 ±1.43

-2.40 ±1.13

-0.57 ±1.43

-0.97 ±0.39

-0.19 ±0.77

0

1

2

3

4

5

6

15 10 5 0 5 10 15 20

AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen pyridine thiazole thiophene NHOH HAcceptors HDonors Al_COO Al_OH Al_OH_noTert

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

-7.01 ±0.17

-6.62 ±0.17

-7.19 ±0.25

-7.69 ±0.14

-9.28 ±0.01

-13.25 ±0.08

0.00 ±0.00

-12.44 ±0.43

-14.15 ±0.30

-16.27 ±0.33

-17.14 ±0.22

-19.11 ±0.37

-24.58 ±0.51

0.00 ±0.00

-3.28 ±0.12

-4.77 ±0.10

-5.32 ±0.16

-5.16 ±0.16

-6.70 ±0.07

-11.12 ±0.08

0.00 ±0.00

-6.45 ±0.51

-9.80 ±0.12

-9.32 ±0.48

-9.74 ±0.39

-10.76 ±0.25

-14.30 ±0.60

0.00 ±0.00

-7.82 ±0.73

-11.05 ±0.30

-11.13 ±0.32

-11.01 ±0.36

-12.36 ±0.49

-15.89 ±0.22

0.00 ±0.00

-8.89 ±0.24

-12.11 ±0.12

-12.26 ±0.27

-12.47 ±0.33

-13.39 ±0.19

-17.02 ±0.15

0.00 ±0.00

-6.96 ±0.12

-6.66 ±0.22

-7.26 ±0.30

-7.70 ±0.13

-9.38 ±0.04

-13.36 ±0.18

0.00 ±0.00

-8.84 ±0.29

-9.87 ±0.25

-13.78 ±0.36

-15.90 ±0.42

-19.76 ±0.29

-28.08 ±0.35

0.00 ±0.00

-9.52 ±0.41

-12.22 ±0.74

-14.44 ±0.79

-16.07 ±0.42

-18.27 ±0.65

-24.50 ±0.83

0.00 ±0.00

-8.02 ±0.89

-10.55 ±0.57

-13.73 ±0.36

-15.62 ±0.28

-17.87 ±0.28

-24.78 ±0.62

0.00 ±0.00

-12.83 ±0.78

-21.37 ±0.47

-24.70 ±0.38

-26.90 ±0.29

-30.71 ±0.39

-39.52 ±0.39

0.00 ±0.00

-8.18 ±0.24

-9.83 ±0.29

-10.64 ±0.44

-11.41 ±0.21

-12.51 ±0.45

-16.09 ±0.56

0.00 ±0.00

-6.39 ±0.39

-9.20 ±1.02

-11.80 ±0.82

-12.50 ±0.96

-16.09 ±0.44

-22.71 ±0.47

0.00 ±0.00

-8.16 ±0.49

-9.84 ±0.25

-10.38 ±0.50

-11.27 ±0.28

-12.39 ±0.28

-15.84 ±0.24

0.00 ±0.00

-3.91 ±0.52

-4.44 ±0.11

-5.98 ±0.28

-7.41 ±0.35

-9.74 ±0.35

-13.52 ±0.46

0.00 ±0.00

-4.99 ±0.30

-6.

0

1

2

3

4

5

6

(g) Chemberta-base (PT)

h Chembe a ba e R

Chembe a 2 5M PT

Chembe a 2 R

k Chembe a 2 10M PT

Chembe a 2 R

m Chembe a 2 77M PT

n Chembe a 2 R

o Chembe a 3 PT

p Chembe a 3 R

F gure 10 Effect of fine-tun ng on poph c ty on mo ecu ar substructure encod ng n CLMs We repor he ayer-w se d fference n prob ng performance (macro-averaged F1 n % ) af er fine- un ng across 60 asks (cf Tab e 6) F gures 10 10 and 10n correspond o he same mode arch ec ure PT deno es o he pre- ra ned mode s ( ef ) wh e RI refers he same mode s bu w h random y n a zed we gh s (r gh ) Improvemen s n prob ng 35 performance af er fine- un ng are shown n red wh e degrada on s nd ca ed n b ue

0.00 ±0.00

-0.11 ±0.14

-0.07 ±0.18

0.50 ±0.48

0.43 ±0.25

0.44 ±0.28

0.81 ±0.19

0.58 ±0.40

0.45 ±0.20

1.10 ±0.29

1.54 ±0.34

1.59 ±0.12

1.12 ±0.40

0.00 ±0.00

0.31 ±0.26

-0.23 ±0.42

0.06 ±0.28

0.15 ±0.22

0.24 ±0.19

0.61 ±0.07

0.80 ±0.11

0.38 ±0.14

0.40 ±0.13

1.12 ±0.16

1.34 ±0.09

0.39 ±0.09

0.00 ±0.00

0.05 ±0.18

0.06 ±0.05

-0.01 ±0.08

0.03 ±0.05

-0.09 ±0.13

0.02 ±0.09

0.03 ±0.11

-0.20 ±0.04

-0.33 ±0.04

-0.13 ±0.18

0.01 ±0.07

-0.41 ±0.01

0.00 ±0.00

-0.03 ±0.19

0.08 ±0.22

-0.06 ±0.10

0.18 ±0.08

0.21 ±0.09

0.22 ±0.29

0.47 ±0.31

0.73 ±0.12

0.80 ±0.21

0.87 ±0.16

1.65 ±0.25

0.99 ±0.11

0.00 ±0.00

-0.23 ±0.15

0.14 ±0.05

0.06 ±0.04

0.05 ±0.15

0.16 ±0.11

0.35 ±0.23

0.58 ±0.09

0.96 ±0.20

0.45 ±0.11

0.46 ±0.10

0.31 ±0.13

0.61 ±0.18

0.00 ±0.00

-0.02 ±0.01

0.03 ±0.05

0.02 ±0.06

-0.05 ±0.01

0.01 ±0.05

0.02 ±0.08

-0.01 ±0.13

0.17 ±0.13

0.26 ±0.12

0.37 ±0.09

0.56 ±0.18

0.64 ±0.14

0.00 ±0.00

0.29 ±0.41

-0.00 ±0.20

0.09 ±0.09

-0.20 ±0.33

0.55 ±0.28

0.61 ±0.47

1.15 ±0.76

0.92 ±0.35

1.32 ±0.42

1.96 ±0.26

2.66 ±0.30

1.45 ±0.35

0.00 ±0.00

-0.02 ±0.21

-0.37 ±0.13

0.12 ±0.19

0.38 ±0.17

0.06 ±0.10

0.54 ±0.16

0.25 ±0.31

0.01 ±0.61

0.10 ±0.20

0.82 ±0.27

1.37 ±0.31

0.67 ±0.30

0.00 ±0.00

0.09 ±0.08

0.14 ±0.05

0.03 ±0.10

0.42 ±0.06

0.36 ±0.21

0.52 ±0.20

0.52 ±0.13

0.11 ±0.16

0.05 ±0.06

0.47 ±0.20

0.90 ±0.15

0.14 ±0.40

0.00 ±0.00

-0.01 ±0.05

0.05 ±0.16

-0.02 ±0.14

0.26 ±0.09

0.22 ±0.17

0.31 ±0.15

0.55 ±0.25

0.72 ±0.18

0.77 ±0.14

0.88 ±0.07

1.57 ±0.18

0.94 ±0.19

0.00 ±0.00

0.23 ±0.19

0.33 ±0.09

0.33 ±0.19

-0.45 ±0.43

-0.42 ±0.11

-0.44 ±0.16

0.56 ±0.18

1.15 ±0.23

1.10 ±0.23

0.85 ±0.28

-0.39 ±0.10

-2.05 ±0.35

0.00 ±0.00

-0.09 ±0.24

0.31 ±0.52

0.74 ±0.29

0.08 ±0.13

0.57 ±0.14

1.05 ±0.57

0.43 ±0.67

0.86 ±0.21

1.33 ±0.21

1.79 ±0.34

1.75 ±0.35

1.22 ±0.23

0.00 ±0.00

-0.27 ±1.13

1.49 ±0.54

1.38 ±0.73

1.20 ±0.16

1.18 ±0.16

1.74 ±0.91

2.06 ±0.74

2.23 ±0.94

2.27 ±0.39

1.89 ±0.60

3.50 ±0.73

2.87 ±1.06

0.00 ±0.00

-0.28 ±0.33

0.28 ±0.27

0.74 ±0.32

0.15 ±0.21

0.72 ±0.02

0.95 ±0.43

0.60 ±0.53

1.06 ±0.11

1.20 ±0.05

1.91 ±0.47

1.70 ±0.18

1.20 ±0.19

0.00 ±0.00

-0.19 ±0.40

0.01 ±0.20

0.22 ±0.44

0.17 ±0.15

0.81 ±0.14

1.03 ±0.19

0.55 ±0.31

1.33 ±0.26

1.35 ±0.50

2.93 ±0.47

3.70 ±0.20

2.82 ±0.43

0.00 ±0.00

0.01 ±0.26

-0.22 ±0.33

0.38 ±0.05

0.48 ±0.25

0.72 ±0.44

0.94 ±0.66

0.74 ±0.25

1.56 ±0.16

1.24 ±0.20

2.61 ±0.31

3.28 ±0.17

2.44 ±0.48

0.00 ±0.00

0.20 ±0.27

-0.39 ±0.15

-0.16 ±0.55

0.11 ±0.28

0.19 ±0.31

1.02 ±0.91

0.53 ±0.88

1.24 ±0.12

1.22 ±0.58

3.08 ±0.69

2.56 ±0.62

4.52 ±0.73

0.00 ±0.00

-0.31 ±0.16

0.29 ±0.26

0.13 ±0.17

-0.15 ±0.12

-0.30 ±0.30

0.36 ±0.13

0.16 ±0.15

0.08 ±0.14

0.39 ±0.13

0.48 ±0.14

0.89 ±0.36

0.83 ±0.06

0.00 ±0.00

0.25 ±0.37

0.20 ±0.06

0.74 ±0.35

0.07 ±0.42

0.74 ±0.08

0.62 ±0.23

0.47 ±0.38

0.91 ±0.45

0.80 ±0.33

0.82 ±0.10

1.36 ±0.39

0.44 ±0.10

0.00 ±0.00

-0.87 ±0.09

0.40 ±0.38

0.25 ±0.15

-0.02 ±0.34

1.26 ±0.18

1.09 ±0.26

1.02 ±0.23

0.33 ±0.60

0.68 ±0.14

0.81 ±0.17

0.61 ±0.24

-0.79 ±0.42

0

1

2

3

4

5

6

7

8

9

10

11

12

0.00 ±0.00

-0.22 ±0.19

-0.27 ±0.06

0.02 ±0.20

0.44 ±0.27

aniline ester ether para_hydroxylation phenol phenol_noOrthoHbond NO Heteroatoms RotatableBonds Ring Ar_N C_O C_O_noCOO allylic_oxid aryl_methyl bicyclic imide unbrch_alkane urea

Layer

0.00 ±0.00

-0.04 ±0.25

-0.20 ±0.32

-0.25 ±0.39

-0.29 ±0.22

-0.03 ±0.34

0.16 ±0.34

0.67 ±0.36

0.52 ±0.40

0.77 ±0.32

2.16 ±0.22

1.92 ±0.17

2.35 ±0.05

0.00 ±0.00

0.04 ±0.20

0.48 ±0.08

0.44 ±0.29

-0.05 ±0.75

0.63 ±0.45

1.03 ±0.13

1.53 ±0.25

1.28 ±0.65

1.00 ±0.77

2.24 ±0.80

1.51 ±0.49

0.55 ±0.58

0.00 ±0.00

0.13 ±0.14

0.35 ±0.28

0.55 ±0.26

0.59 ±0.29

1.05 ±0.07

1.00 ±0.08

0.83 ±0.30

1.18 ±0.21

1.05 ±0.27

2.10 ±0.15

2.38 ±0.15

1.12 ±0.29

0.00 ±0.00

-0.33 ±0.12

-0.59 ±0.30

-0.52 ±0.62

0.72 ±0.02

0.98 ±0.29

1.37 ±0.37

1.69 ±0.32

1.43 ±0.11

1.29 ±0.14

1.97 ±0.27

2.15 ±0.44

2.10 ±0.41

0.00 ±0.00

0.00 ±0.06

-0.21 ±0.19

-0.05 ±0.22

0.20 ±0.42

0.24 ±0.64

1.06 ±0.80

0.18 ±0.52

1.65 ±0.66

1.62 ±1.54

3.98 ±0.88

3.93 ±0.47

4.83 ±0.77

0.00 ±0.00

0.21 ±0.01

-0.03 ±0.09

0.26 ±0.22

0.15 ±0.55

0.15 ±0.27

1.37 ±0.82

0.38 ±0.51

1.20 ±1.14

0.98 ±0.94

3.82 ±1.33

3.67 ±1.07

5.10 ±1.44

0.00 ±0.00

-1.09 ±0.64

1.67 ±0.99

1.95 ±0.55

1.44 ±0.34

1.43 ±0.62

2.38 ±1.14

3.25 ±1.09

3.20 ±0.31

3.66 ±0.41

2.66 ±0.95

3.89 ±0.23

2.04 ±1.20

0.00 ±0.00

-0.53 ±1.00

1.86 ±0.69

2.09 ±1.23

0.77 ±1.05

1.81 ±0.32

1.67 ±0.72

2.38 ±0.69

3.05 ±0.74

2.49 ±1.29

2.87 ±0.72

1.89 ±0.30

1.57 ±1.34

0.00 ±0.00

0.62 ±0.29

1.06 ±0.49

1.00 ±0.30

0.10 ±0.52

0.05 ±0.37

0.33 ±0.28

0.77 ±0.49

0.68 ±0.66

-0.10 ±0.51

0.47 ±0.26

0.68 ±0.28

0.01 ±0.43

0.00 ±0.00

0.00 ±0.01

0.00 ±0.03

0.00 ±0.01

-0.04 ±0.04

-0.07 ±0.08

-0.03 ±0.01

-0.13 ±0.02

-0.16 ±0.09

-0.00 ±0.06

0.01 ±0.16

0.15 ±0.06

0.49 ±0.32

0.00 ±0.00

-0.16 ±0.07

0.05 ±0.09

-0.02 ±0.06

-0.34 ±0.03

0.02 ±0.17

0.51 ±0.04

0.69 ±0.17

0.92 ±0.12

0.48 ±0.15

0.58 ±0.37

0.34 ±0.19

0.83 ±0.26

0.00 ±0.00

-0.28 ±0.22

0.18 ±0.14

0.23 ±0.09

-0.11 ±0.21

0.87 ±0.21

0.55 ±0.11

0.75 ±0.06

0.56 ±0.34

0.42 ±0.19

0.40 ±0.16

0.12 ±0.21

-0.69 ±0.41

0.00 ±0.00

-0.34 ±0.31

0.13 ±0.27

0.40 ±0.12

0.26 ±0.28

0.98 ±0.26

0.95 ±0.15

1.19 ±0.02

0.82 ±0.43

0.78 ±0.43

1.09 ±0.14

0.70 ±0.09

-0.45 ±0.19

0.00 ±0.00

0.21 ±0.11

0.56 ±0.16

0.18 ±0.48

0.23 ±0.37

0.44 ±0.51

0.76 ±0.35

1.08 ±0.74

1.66 ±0.19

1.20 ±0.38

1.51 ±0.02

1.55 ±0.14

0.10 ±0.30

0.00 ±0.00

0.31 ±0.08

-0.10 ±0.32

1.16 ±0.46

0.55 ±0.55

0.18 ±0.30

-0.04 ±0.24

0.34 ±0.07

0.28 ±0.60

0.45 ±0.47

0.38 ±0.33

0.60 ±0.53

0.41 ±0.05

0.00 ±0.00

-0.19 ±0.20

-0.02 ±0.38

-0.13 ±0.09

0.06 ±0.13

0.25 ±0.05

0.31 ±0.03

0.21 ±0.05

0.22 ±0.12

0.20 ±0.33

0.68 ±0.19

0.38 ±0.15

0.70 ±0.05

0.00 ±0.00

0.30 ±0.13

1.06 ±0.40

0.19 ±0.96

0.23 ±0.64

-0.25 ±0.33

0.30 ±0.68

0.18 ±0.68

-0.71 ±0.55

0.06 ±0.24

0.11 ±0.53

1.25 ±1.41

-0.16 ±1.11

0.00 ±0.00

-0.33 ±0.26

0.21 ±0.18

1.13 ±0.27

1.36 ±0.61

1.29 ±0.06

0.60 ±0.18

0.80 ±0.30

0.98 ±0.42

0.45 ±0.19

2.20 ±0.42

1.87 ±0.61

0.67 ±0.34

0.00 ±0.00

-0.02 ±0.50

-0.36 ±0.45

0.22 ±0.50

0.68 ±0.56

0.30 ±0.58

-0.03 ±0.28

-0.20 ±0.92

0.11 ±1.34

1.03 ±0.50

0.40 ±0.46

1.54 ±0.68

0.27 ±1.03

0

1

2

3

4

5

6

7

8

9

10

11

12

20 15 10 5 0

Molecular Substructure

Molecular Substructure

AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen NHOH HAcceptors HDonors Al_OH Al_OH_noTert Ar_OH NH0 NH1 amide

5 10

AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen NHOH HAcceptors HDonors Al_OH Al_OH_noTert Ar_OH NH0 NH1 amide

0.00 ±0.00

-0.22 ±0.04

-0.02 ±0.11

0.05 ±0.15

0.29 ±0.35

0.32 ±0.06

0.28 ±0.28

0.33 ±0.12

0.22 ±0.05

0.43 ±0.24

0.40 ±0.20

0.36 ±0.16

0.21 ±0.17

0.00 ±0.00

-0.14 ±0.07

0.02 ±0.13

-0.18 ±0.31

-0.13 ±0.13

0.03 ±0.22

0.25 ±0.07

0.04 ±0.16

0.05 ±0.13

-0.13 ±0.15

0.04 ±0.14

0.34 ±0.14

0.16 ±0.23

0.00 ±0.00

-0.30 ±0.12

-0.18 ±0.09

0.04 ±0.00

-0.09 ±0.03

-0.05 ±0.02

-0.08 ±0.11

-0.09 ±0.07

0.02 ±0.08

0.01 ±0.07

0.04 ±0.03

0.00 ±0.04

0.03 ±0.03

0.00 ±0.00

-0.29 ±0.07

-0.17 ±0.09

-0.09 ±0.05

-0.13 ±0.04

-0.14 ±0.14

-0.06 ±0.09

-0.11 ±0.16

-0.10 ±0.07

-0.05 ±0.05

0.02 ±0.03

-0.05 ±0.16

-0.01 ±0.10

0.00 ±0.00

0.07 ±0.09

0.03 ±0.03

0.01 ±0.02

-0.02 ±0.05

0.00 ±0.03

-0.03 ±0.05

-0.07 ±0.04

-0.09 ±0.09

-0.06 ±0.08

-0.01 ±0.08

-0.09 ±0.04

-0.06 ±0.02

0.00 ±0.00

0.02 ±0.05

0.02 ±0.04

0.01 ±0.02

0.03 ±0.01

0.00 ±0.03

0.01 ±0.04

-0.01 ±0.03

0.01 ±0.00

0.03 ±0.03

0.02 ±0.01

-0.00 ±0.03

0.05 ±0.02

0.00 ±0.00

-0.43 ±0.16

-0.22 ±0.17

-0.13 ±0.24

-0.11 ±0.22

0.08 ±0.02

-0.09 ±0.08

0.06 ±0.32

0.14 ±0.16

0.05 ±0.20

-0.27 ±0.19

-0.05 ±0.15

-0.21 ±0.32

0.00 ±0.00

0.05 ±0.43

-0.19 ±0.15

0.02 ±0.15

0.13 ±0.33

0.19 ±0.15

0.09 ±0.17

0.35 ±0.03

0.42 ±0.38

0.25 ±0.25

0.29 ±0.16

0.55 ±0.72

0.60 ±0.54

0.00 ±0.00

-0.24 ±0.08

-0.09 ±0.16

0.12 ±0.11

0.14 ±0.25

0.03 ±0.14

0.03 ±0.02

0.07 ±0.05

0.27 ±0.15

0.20 ±0.22

0.23 ±0.19

0.30 ±0.11

0.27 ±0.10

0.00 ±0.00

-0.31 ±0.03

-0.15 ±0.10

-0.10 ±0.04

-0.12 ±0.08

-0.09 ±0.08

-0.08 ±0.09

-0.07 ±0.13

-0.03 ±0.12

-0.05 ±0.09

-0.01 ±0.08

-0.07 ±0.15

-0.04 ±0.08

0.00 ±0.00

0.14 ±0.08

0.10 ±0.11

0.05 ±0.06

-0.01 ±0.10

-0.05 ±0.03

-0.02 ±0.03

-0.15 ±0.08

-0.13 ±0.08

-0.17 ±0.08

-0.22 ±0.09

-0.34 ±0.10

-0.49 ±0.08

0.00 ±0.00

-0.08 ±0.06

-0.14 ±0.05

0.05 ±0.13

0.01 ±0.09

0.14 ±0.13

-0.19 ±0.49

0.08 ±0.19

0.18 ±0.18

0.19 ±0.08

0.04 ±0.23

0.27 ±0.22

-0.06 ±0.11

0.00 ±0.00

2.41 ±0.31

1.92 ±0.64

2.48 ±0.17

2.45 ±0.18

2.12 ±0.45

1.97 ±0.73

2.13 ±0.46

1.49 ±0.43

0.96 ±0.49

0.65 ±0.20

0.73 ±0.51

0.27 ±0.20

0.00 ±0.00

-0.12 ±0.03

0.08 ±0.30

-0.02 ±0.09

0.03 ±0.11

0.07 ±0.15

-0.10 ±0.30

-0.08 ±0.40

0.03 ±0.35

0.08 ±0.18

0.02 ±0.13

0.08 ±0.43

-0.01 ±0.17

0.00 ±0.00

-0.20 ±0.02

0.13 ±0.25

-0.01 ±0.25

0.27 ±0.15

0.18 ±0.12

0.32 ±0.39

0.77 ±0.48

0.60 ±0.37

0.75 ±0.21

0.83 ±0.22

0.79 ±0.31

1.03 ±0.28

0.00 ±0.00

-0.09 ±0.25

0.05 ±0.36

0.03 ±0.30

0.19 ±0.40

-0.07 ±0.33

0.33 ±0.51

0.49 ±0.20

0.51 ±0.23

0.45 ±0.22

0.67 ±0.20

0.58 ±0.17

0.61 ±0.19

0.00 ±0.00

0.23 ±0.18

-0.05 ±0.06

-0.05 ±0.24

0.01 ±0.30

-0.14 ±0.22

-0.35 ±0.15

-0.13 ±0.14

0.21 ±0.08

0.41 ±0.10

0.40 ±0.19

0.20 ±0.21

0.33 ±0.16

0.00 ±0.00

0.13 ±0.17

0.20 ±0.08

0.01 ±0.18

0.19 ±0.20

0.28 ±0.08

0.07 ±0.11

0.19 ±0.10

-0.04 ±0.23

0.25 ±0.28

0.34 ±0.10

0.19 ±0.11

0.32 ±0.10

0.00 ±0.00

-0.09 ±0.20

-0.04 ±0.01

0.49 ±0.31

0.13 ±0.24

0.26 ±0.12

0.29 ±0.09

0.41 ±0.25

0.50 ±0.18

0.22 ±0.06

0.31 ±0.21

0.43 ±0.25

0.28 ±0.29

0.00 ±0.00

0.20 ±0.20

0.82 ±1.00

1.08 ±0.18

1.43 ±0.25

1.27 ±0.19

1.56 ±0.24

1.49 ±0.35

2.01 ±0.31

1.38 ±0.54

1.44 ±0.09

1.53 ±0.24

1.80 ±0.27

0

1

2

3

4

5

6

7

8

9

10

11

12

0.00 ±0.00

0.20 ±0.07

0.27 ±0.28

0.22 ±0.02

0.19 ±0.23

-0.03 ±0.04

-0.11 ±0.35

0.05 ±0.18

-0.42 ±0.13

1.06 ±0.20

0.54 ±0.53

0.28 ±0.22

1.63 ±0.04

1.20 ±0.23

1.87 ±0.25

0.00 ±0.00

0.02 ±0.16

-0.03 ±0.05

-0.13 ±0.11

-0.02 ±0.16

-0.08 ±0.05

-0.36 ±0.16

-0.17 ±0.12

0.07 ±0.18

0.29 ±0.02

0.13 ±0.09

-0.17 ±0.30

1.45 ±0.11

1.47 ±0.14

2.06 ±0.15

0.00 ±0.00

-0.01 ±0.05

-0.00 ±0.17

0.03 ±0.08

0.17 ±0.08

0.07 ±0.07

-0.09 ±0.11

-0.00 ±0.07

-0.00 ±0.06

0.26 ±0.09

0.50 ±0.04

0.05 ±0.07

0.21 ±0.17

0.26 ±0.08

0.90 ±0.08

0.00 ±0.00

-0.09 ±0.07

0.16 ±0.06

-0.04 ±0.05

0.14 ±0.05

0.08 ±0.10

0.05 ±0.01

0.19 ±0.08

0.32 ±0.05

0.82 ±0.16

0.88 ±0.13

1.34 ±0.04

1.57 ±0.05

1.65 ±0.13

2.32 ±0.10

0.00 ±0.00

-0.19 ±0.04

0.14 ±0.04

0.38 ±0.31

0.19 ±0.05

0.37 ±0.11

0.16 ±0.03

0.06 ±0.11

0.44 ±0.09

0.19 ±0.22

0.19 ±0.21

0.53 ±0.12

1.43 ±0.25

1.53 ±0.19

1.84 ±0.09

0.00 ±0.00

0.01 ±0.02

-0.05 ±0.02

-0.04 ±0.04

0.01 ±0.08

-0.06 ±0.10

0.06 ±0.03

-0.02 ±0.05

0.01 ±0.06

0.23 ±0.03

0.53 ±0.10

0.36 ±0.07

0.23 ±0.08

0.47 ±0.09

0.80 ±0.12

0.00 ±0.00

-0.27 ±0.22

0.18 ±0.17

-0.07 ±0.10

-0.08 ±0.13

-0.27 ±0.16

0.01 ±0.26

-0.23 ±0.18

0.20 ±0.13

1.49 ±0.18

0.62 ±0.27

0.74 ±0.39

2.00 ±0.19

0.99 ±0.40

2.60 ±0.73

0.00 ±0.00

-0.06 ±0.14

-0.10 ±0.14

-0.03 ±0.18

0.46 ±0.15

0.01 ±0.15

-0.23 ±0.12

-0.13 ±0.08

0.06 ±0.11

0.50 ±0.38

0.36 ±0.43

0.04 ±0.26

0.89 ±0.11

1.10 ±0.21

1.79 ±0.18

0.00 ±0.00

0.11 ±0.02

-0.07 ±0.12

-0.05 ±0.05

0.03 ±0.11

-0.14 ±0.24

-0.28 ±0.17

-0.21 ±0.05

-0.27 ±0.14

0.43 ±0.06

0.18 ±0.12

-0.26 ±0.16

0.86 ±0.17

0.60 ±0.25

1.38 ±0.29

0.00 ±0.00

0.00 ±0.10

0.15 ±0.02

-0.08 ±0.08

0.14 ±0.08

0.03 ±0.08

0.06 ±0.03

0.12 ±0.05

0.32 ±0.07

0.80 ±0.13

0.85 ±0.15

1.25 ±0.06

1.54 ±0.05

1.61 ±0.17

2.33 ±0.11

0.00 ±0.00

0.04 ±0.08

-0.13 ±0.32

-0.23 ±0.09

-0.24 ±0.16

0.32 ±0.25

0.26 ±0.24

0.65 ±0.15

0.84 ±0.06

1.74 ±0.31

2.27 ±0.01

2.28 ±0.18

2.15 ±0.40

2.05 ±0.22

1.82 ±0.16

0.00 ±0.00

-0.04 ±0.13

-0.16 ±0.24

-0.30 ±0.27

0.14 ±0.09

0.04 ±0.08

0.03 ±0.14

-0.20 ±0.19

-0.40 ±0.27

1.31 ±0.14

1.52 ±0.20

1.23 ±0.10

2.00 ±0.06

1.10 ±0.16

0.90 ±0.26

0.00 ±0.00

-0.13 ±0.07

-0.17 ±0.24

0.61 ±0.23

1.84 ±0.17

2.56 ±0.35

3.79 ±0.50

4.37 ±0.87

4.86 ±0.81

5.50 ±0.70

4.61 ±0.45

3.10 ±0.87

3.28 ±0.90

2.69 ±0.63

2.81 ±0.76

0.00 ±0.00

0.12 ±0.06

-0.05 ±0.34

-0.17 ±0.14

0.23 ±0.15

0.06 ±0.08

0.11 ±0.03

-0.23 ±0.04

-0.22 ±0.32

1.32 ±0.19

1.59 ±0.30

1.23 ±0.13

1.62 ±0.16

1.02 ±0.21

0.74 ±0.53

0.00 ±0.00

0.29 ±0.18

-0.18 ±0.24

-0.56 ±0.09

0.25 ±0.14

0.08 ±0.13

0.18 ±0.18

-0.41 ±0.14

-0.02 ±0.37

3.18 ±0.42

3.29 ±0.05

3.11 ±0.12

3.38 ±0.19

1.46 ±0.28

0.39 ±0.14

0.00 ±0.00

0.39 ±0.12

-0.11 ±0.12

-0.67 ±0.27

0.00 ±0.11

-0.31 ±0.15

0.06 ±0.26

-1.03 ±0.29

-0.31 ±0.44

3.00 ±0.15

3.03 ±0.26

2.96 ±0.21

2.85 ±0.06

1.39 ±0.25

0.67 ±0.14

0.00 ±0.00

0.22 ±0.01

0.10 ±0.06

-0.69 ±0.33

-0.20 ±0.72

0.19 ±0.30

0.55 ±0.15

1.44 ±0.17

1.33 ±0.44

3.19 ±0.71

5.13 ±0.42

5.14 ±0.90

5.37 ±0.19

3.07 ±0.96

3.29 ±1.17

0.00 ±0.00

-0.12 ±0.10

0.02 ±0.25

0.06 ±0.10

0.40 ±0.27

0.36 ±0.18

0.32 ±0.09

0.43 ±0.20

0.19 ±0.03

0.53 ±0.19

0.27 ±0.25

-0.14 ±0.04

0.41 ±0.20

0.03 ±0.08

-0.03 ±0.19

0.00 ±0.00

0.12 ±0.05

-0.16 ±0.13

-0.23 ±0.30

-0.08 ±0.24

-0.16 ±0.10

-0.46 ±0.05

0.21 ±0.30

-0.22 ±0.35

0.11 ±0.09

0.69 ±0.22

0.77 ±0.14

1.38 ±0.30

1.13 ±0.18

0.84 ±0.17

0.00 ±0.00

-0.10 ±0.16

-0.75 ±0.30

-0.44 ±0.20

-0.02 ±0.25

0.30 ±0.11

0.31 ±0.09

-0.02 ±0.11

0.20 ±0.27

1.53 ±0.10

0.75 ±0.26

0.63 ±0.14

1.81 ±0.28

0.84 ±0.41

0.24 ±0.05

0

1

2

3

4

5

6

7

8

9

10

11

12

13

14

aniline ester ether para_hydroxylation phenol phenol_noOrthoHbond NO Heteroatoms RotatableBonds Ring Ar_N C_O C_O_noCOO allylic_oxid aryl_methyl bicyclic imide unbrch_alkane urea

Layer

-0.11 ±0.03

-0.64 ±0.15

-0.32 ±0.20

0.59 ±0.30

-0.35 ±0.14

-0.14 ±0.19

0.57 ±0.32

1.73 ±0.32

2.56 ±0.28

2.14 ±0.51

2.23 ±0.19

3.30 ±0.29

2.34 ±0.31

2.71 ±0.14

0.00 ±0.00

-0.15 ±0.04

-0.14 ±0.20

0.21 ±0.22

0.35 ±0.09

-0.21 ±0.35

0.13 ±0.09

0.00 ±0.24

-0.48 ±0.25

0.78 ±0.36

1.29 ±0.35

1.31 ±0.38

1.75 ±0.45

1.75 ±0.35

0.71 ±0.83

0.00 ±0.00

-0.31 ±0.09

-0.38 ±0.07

0.19 ±0.06

0.53 ±0.17

-0.09 ±0.18

0.01 ±0.04

0.32 ±0.21

-0.05 ±0.17

0.61 ±0.24

0.74 ±0.30

1.25 ±0.16

0.88 ±0.22

1.23 ±0.12

0.72 ±0.14

0.00 ±0.00

-0.26 ±0.06

-0.15 ±0.25

-0.48 ±0.18

0.04 ±0.09

0.03 ±0.18

-0.12 ±0.07

-0.14 ±0.14

0.25 ±0.16

1.74 ±0.23

1.76 ±0.11

2.87 ±0.36

2.71 ±0.32

3.13 ±0.27

3.89 ±0.48

0.00 ±0.00

0.34 ±0.20

-0.01 ±0.10

-0.52 ±0.29

0.37 ±0.53

1.46 ±0.32

1.43 ±0.30

3.83 ±0.16

4.22 ±0.50

6.34 ±0.14

6.95 ±0.94

6.20 ±0.76

7.51 ±0.69

5.28 ±0.49

6.73 ±1.42

0.00 ±0.00

0.51 ±0.15

-0.19 ±0.03

-0.47 ±0.14

0.59 ±0.65

1.09 ±0.21

1.56 ±0.37

3.91 ±0.40

3.10 ±0.56

5.20 ±0.77

6.44 ±1.03

6.35 ±0.14

6.93 ±0.40

5.45 ±0.69

6.25 ±1.25

0.00 ±0.00

-0.06 ±0.09

0.29 ±0.18

0.87 ±0.26

2.09 ±0.33

3.27 ±0.32

5.05 ±0.49

5.43 ±0.32

5.04 ±0.31

5.42 ±0.47

5.89 ±0.18

3.68 ±0.51

3.08 ±0.49

2.99 ±0.50

3.10 ±1.10

0.00 ±0.00

0.15 ±0.15

0.51 ±0.20

1.63 ±0.27

2.88 ±0.45

4.42 ±0.30

6.42 ±0.35

5.58 ±1.11

5.91 ±0.90

6.95 ±1.08

5.61 ±0.46

4.93 ±1.11

5.36 ±0.99

4.04 ±0.61

3.35 ±1.04

0.00 ±0.00

-0.19 ±0.14

0.07 ±0.26

0.01 ±0.20

-0.01 ±0.29

0.18 ±0.03

0.09 ±0.13

-0.65 ±0.18

0.20 ±0.06

-0.30 ±0.20

-0.13 ±0.04

-0.96 ±0.35

-1.47 ±0.71

-1.73 ±0.19

-1.92 ±0.11

0.00 ±0.00

-0.01 ±0.03

-0.03 ±0.00

-0.01 ±0.12

-0.09 ±0.02

-0.13 ±0.09

-0.09 ±0.01

-0.01 ±0.02

0.05 ±0.09

0.16 ±0.08

0.02 ±0.14

-0.02 ±0.01

0.17 ±0.12

0.26 ±0.36

0.43 ±0.14

0.00 ±0.00

-0.23 ±0.06

-0.09 ±0.11

0.61 ±0.14

0.52 ±0.15

0.44 ±0.18

0.21 ±0.04

0.25 ±0.16

0.49 ±0.31

0.63 ±0.16

0.06 ±0.17

0.61 ±0.16

1.31 ±0.35

0.69 ±0.22

1.19 ±0.24

0.00 ±0.00

0.07 ±0.08

-0.30 ±0.05

-0.41 ±0.10

-0.23 ±0.01

0.38 ±0.10

0.12 ±0.13

0.03 ±0.06

0.39 ±0.13

0.71 ±0.10

0.44 ±0.03

0.51 ±0.31

0.84 ±0.11

0.27 ±0.18

0.17 ±0.20

0.00 ±0.00

0.24 ±0.03

-0.21 ±0.12

-0.65 ±0.16

-0.29 ±0.06

0.24 ±0.10

0.00 ±0.22

0.02 ±0.02

0.16 ±0.27

1.15 ±0.14

1.00 ±0.21

1.10 ±0.29

1.17 ±0.29

0.44 ±0.15

0.07 ±0.24

0.00 ±0.00

-0.15 ±0.15

-0.21 ±0.12

-0.35 ±0.12

-0.15 ±0.03

-0.38 ±0.11

-0.52 ±0.14

-0.10 ±0.05

-0.15 ±0.14

0.22 ±0.43

-0.09 ±0.26

-0.64 ±0.12

0.19 ±0.18

-0.01 ±0.43

-0.01 ±0.32

0.00 ±0.00

-0.04 ±0.24

-0.51 ±0.25

-0.33 ±0.24

-0.18 ±0.17

-0.58 ±0.14

0.03 ±0.45

-0.20 ±0.38

0.14 ±0.21

0.12 ±0.11

-0.40 ±0.27

-0.12 ±0.08

-0.13 ±0.21

0.25 ±0.29

0.26 ±0.09

0.00 ±0.00

-0.04 ±0.12

-0.24 ±0.03

-0.13 ±0.09

0.08 ±0.11

-0.03 ±0.04

-0.07 ±0.12

-0.05 ±0.06

0.16 ±0.05

0.12 ±0.09

0.24 ±0.05

0.20 ±0.07

-0.08 ±0.19

0.77 ±0.18

1.00 ±0.20

0.00 ±0.00

-0.38 ±0.10

-0.16 ±0.05

-0.91 ±0.06

-0.38 ±0.26

-0.57 ±0.37

1.11 ±0.30

0.37 ±0.40

1.00 ±0.49

1.30 ±0.63

0.07 ±0.15

-1.09 ±0.12

1.48 ±0.68

0.07 ±0.49

-1.32 ±1.25

0.00 ±0.00

-0.01 ±0.19

0.13 ±0.26

1.13 ±0.04

1.49 ±0.09

0.93 ±0.15

0.66 ±0.09

0.56 ±0.76

0.85 ±0.40

0.44 ±0.13

0.26 ±0.20

0.25 ±0.13

-0.22 ±0.03

-0.23 ±0.19

-0.64 ±0.85

0.00 ±0.00

-0.18 ±0.15

-0.44 ±0.15

-0.33 ±0.14

-0.04 ±0.38

-0.42 ±0.07

0.14 ±0.17

0.32 ±0.28

0.10 ±0.04

1.74 ±0.56

1.08 ±0.26

0.31 ±0.75

1.50 ±0.38

0.81 ±1.16

1.26 ±1.06

0

1

2

3

4

5

6

7

8

9

10

11

12

13

14

10 5 0 5 10 15

AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen NHOH HAcceptors HDonors Al_OH Al_OH_noTert Ar_OH NH0 NH1 amide

0.63 ±0.31

0.31 ±0.24

0.13 ±0.01

0.42 ±0.17

-0.14 ±0.10

-0.11 ±0.21

-0.73 ±0.05

-1.70 ±0.27

-1.95 ±0.25

-0.81 ±0.63

0.00 ±0.00

-0.32 ±0.17

-0.53 ±0.22

-1.38 ±0.22

-2.28 ±0.37

-2.03 ±0.03

-1.53 ±0.17

0.00 ±0.00

0.05 ±0.05

0.09 ±0.06

-0.31 ±0.10

-0.78 ±0.10

-0.89 ±0.03

-1.12 ±0.04

0.00 ±0.00

0.00 ±0.06

0.34 ±0.13

-0.05 ±0.08

0.11 ±0.11

0.27 ±0.12

0.35 ±0.15

0.00 ±0.00

0.57 ±0.08

0.51 ±0.14

-0.39 ±0.26

0.08 ±0.05

1.43 ±0.22

1.59 ±0.04

0.00 ±0.00

0.00 ±0.01

-0.01 ±0.06

0.06 ±0.09

-0.10 ±0.06

-0.17 ±0.01

-0.05 ±0.06

0.00 ±0.00

-0.27 ±0.14

-0.42 ±0.17

-1.18 ±0.53

-1.21 ±0.31

-1.96 ±0.61

-0.47 ±0.39

0.00 ±0.00

-0.17 ±0.19

-0.86 ±0.38

-1.19 ±0.23

-2.14 ±0.26

-1.57 ±0.32

-1.19 ±0.21

0.00 ±0.00

-0.57 ±0.17

-0.85 ±0.18

-1.20 ±0.14

-1.90 ±0.14

-2.56 ±0.24

-2.03 ±0.15

0.00 ±0.00

-0.08 ±0.08

0.39 ±0.09

0.01 ±0.07

0.12 ±0.04

0.29 ±0.05

0.39 ±0.21

0.00 ±0.00

0.64 ±0.30

0.40 ±0.07

-0.30 ±0.36

0.23 ±0.12

0.37 ±0.27

0.28 ±0.27

0.00 ±0.00

0.12 ±0.03

-0.05 ±0.21

-0.71 ±0.40

-2.40 ±0.37

-1.15 ±0.22

-0.28 ±0.15

0.00 ±0.00

3.18 ±0.04

6.13 ±0.16

4.68 ±0.68

5.18 ±0.43

7.59 ±0.45

7.60 ±0.69

0.00 ±0.00

0.04 ±0.32

0.09 ±0.30

-0.51 ±0.38

-2.23 ±0.15

-1.51 ±0.22

-0.63 ±0.13

0.00 ±0.00

0.91 ±0.09

2.18 ±0.34

0.95 ±0.15

2.43 ±0.14

1.72 ±0.13

3.54 ±0.49

0.00 ±0.00

0.90 ±0.07

1.89 ±0.43

0.57 ±0.15

1.50 ±0.25

1.48 ±0.17

2.80 ±0.46

0.00 ±0.00

2.49 ±0.26

4.04 ±0.24

2.72 ±0.43

2.11 ±1.08

4.80 ±0.94

5.50 ±0.60

0.00 ±0.00

-0.33 ±0.22

-0.46 ±0.21

-1.27 ±0.08

-1.73 ±0.18

-1.44 ±0.25

-0.65 ±0.09

0.00 ±0.00

-0.87 ±0.32

-0.42 ±0.20

-2.26 ±0.52

-4.04 ±0.08

-2.95 ±0.35

-1.41 ±0.16

0.00 ±0.00

-0.47 ±0.30

0.30 ±0.14

-0.14 ±0.12

-2.87 ±0.50

-2.16 ±0.06

-3.41 ±0.53

0

1

2

3

4

5

6

0.00 ±0.00

-2.33 ±0.22

-2.52 ±0.29

-1.92 ±0.20

0.04 ±0.20

-0.45 ±0.30

0.49 ±0.31

0.00 ±0.00

-1.73 ±0.01

-2.26 ±0.25

-1.43 ±0.32

0.71 ±0.36

0.21 ±0.18

-0.32 ±0.13

0.00 ±0.00

-1.28 ±0.11

-2.03 ±0.17

-1.38 ±0.10

-0.54 ±0.11

-0.62 ±0.10

-0.68 ±0.09

aniline ester ether para_hydroxylation phenol phenol_noOrthoHbond NO Heteroatoms RotatableBonds Ring Ar_N C_O C_O_noCOO allylic_oxid aryl_methyl bicyclic imide unbrch_alkane urea

Layer

0.77 ±0.11

1.55 ±0.05

-0.26 ±0.15

-2.12 ±0.23

-0.53 ±0.11

0.08 ±0.54

0.00 ±0.00

0.69 ±0.19

0.05 ±0.25

0.41 ±0.48

2.30 ±0.55

2.52 ±0.54

0.78 ±0.57

0.00 ±0.00

0.56 ±0.26

1.53 ±0.30

0.43 ±0.20

2.37 ±0.35

3.49 ±0.28

3.38 ±0.15

0.00 ±0.00

-0.03 ±0.16

-0.18 ±0.13

-0.60 ±0.37

-0.31 ±0.34

-0.59 ±0.17

0.35 ±0.26

0.00 ±0.00

1.78 ±0.39

3.92 ±0.32

2.57 ±0.42

3.29 ±0.41

6.23 ±0.82

8.06 ±0.52

0.00 ±0.00

1.94 ±0.46

3.80 ±0.37

2.40 ±0.23

3.03 ±1.14

5.91 ±0.25

7.46 ±0.31

0.00 ±0.00

2.60 ±0.64

5.17 ±0.07

4.23 ±0.29

4.72 ±0.18

7.57 ±0.11

7.06 ±0.35

0.00 ±0.00

2.35 ±0.95

6.15 ±0.76

5.91 ±1.04

6.43 ±0.99

7.73 ±0.46

8.22 ±0.75

0.00 ±0.00

0.59 ±0.26

0.83 ±0.18

-0.59 ±0.61

-1.49 ±0.30

-3.01 ±0.36

-3.67 ±0.50

0.00 ±0.00

-0.02 ±0.02

-0.04 ±0.03

-0.09 ±0.03

-0.11 ±0.05

-0.17 ±0.07

-0.25 ±0.08

0.00 ±0.00

0.27 ±0.19

0.48 ±0.03

-0.91 ±0.15

0.06 ±0.06

1.93 ±0.39

2.54 ±0.24

0.00 ±0.00

0.41 ±0.11

0.78 ±0.21

0.14 ±0.11

-1.54 ±0.11

-1.45 ±0.13

-2.57 ±0.27

0.00 ±0.00

1.01 ±0.13

1.77 ±0.30

0.82 ±0.22

-1.50 ±0.22

-1.73 ±0.18

-3.62 ±0.20

0.00 ±0.00

-0.70 ±0.17

-0.20 ±0.02

-0.39 ±0.13

-1.03 ±0.29

-2.14 ±0.22

-1.43 ±0.33

0.00 ±0.00

3.24 ±0.26

3.00 ±0.28

0.72 ±0.28

0.10 ±0.32

0.48 ±0.27

0.07 ±0.11

0.00 ±0.00

-0.57 ±0.16

-0.65 ±0.03

-1.69 ±0.27

-1.96 ±0.32

-1.50 ±0.31

-0.68 ±0.08

0.00 ±0.00

-1.23 ±0.45

-1.52 ±0.50

-1.89 ±0.17

-2.50 ±0.28

-1.83 ±0.24

-3.58 ±1.46

0.00 ±0.00

9.80 ±0.37

7.02 ±0.13

3.44 ±0.35

3.51 ±0.64

4.03 ±0.37

4.22 ±0.68

0.00 ±0.00

-0.46 ±0.33

-0.43 ±0.26

-0.43 ±0.43

-1.49 ±0.68

-1.12 ±0.48

-0.72 ±0.29

0

1

2

3

4

5

6

0.00 ±0.00

-1.94 ±0.52

-3.59 ±0.38

-1.65 ±0.51

0.76 ±0.31

0.92 ±0.44

1.56 ±0.36

0.00 ±0.00

-2.26 ±0.18

-1.35 ±0.52

-0.16 ±0.13

0.21 ±0.13

1.01 ±0.56

0.51 ±0.46

0.00 ±0.00

-1.00 ±0.34

0.09 ±0.33

1.14 ±0.26

3.49 ±0.57

2.67 ±0.50

3.34 ±0.36

0.00 ±0.00

-1.56 ±0.04

-3.11 ±0.35

-1.42 ±0.44

-0.50 ±0.63

-0.52 ±0.41

0.09 ±0.59

0.00 ±0.00

1.76 ±0.41

1.26 ±0.16

4.62 ±0.12

5.03 ±0.37

8.54 ±0.86

8.97 ±0.10

0.00 ±0.00

1.25 ±0.49

1.01 ±0.23

4.35 ±0.43

4.94 ±0.97

8.83 ±1.97

8.52 ±0.70

0.00 ±0.00

1.68 ±0.11

3.10 ±0.85

4.39 ±0.52

5.90 ±0.62

8.25 ±0.37

9.64 ±1.26

0.00 ±0.00

2.37 ±0.06

1.79 ±0.76

4.47 ±1.62

6.05 ±1.81

9.29 ±2.94

9.09 ±1.95

0.00 ±0.00

1.18 ±0.14

-1.58 ±0.11

-2.56 ±0.22

-1.98 ±0.23

-1.20 ±0.35

-0.44 ±0.41

0.00 ±0.00

-0.00 ±0.08

-0.07 ±0.08

-0.14 ±0.02

-0.20 ±0.01

-0.22 ±0.10

-0.44 ±0.08

0.00 ±0.00

-4.20 ±0.41

-3.86 ±0.20

-1.12 ±0.19

0.42 ±0.21

1.25 ±0.13

1.23 ±0.04

0.00 ±0.00

0.20 ±0.16

-1.27 ±0.27

-0.09 ±0.05

-0.38 ±0.15

0.99 ±0.10

0.96 ±0.24

10 5 0 5 10 15 20

-1.18 ±0.05

-2.75 ±0.07

-0.73 ±0.20

-0.49 ±0.21

-0.29 ±0.12

-0.20 ±0.18

-4.16 ±0.29

-4.37 ±0.46

-1.16 ±0.16

-0.02 ±0.11

0.69 ±0.07

0.96 ±0.10

0.00 ±0.00

-0.53 ±0.06

-1.00 ±0.16

-0.59 ±0.20

-0.68 ±0.20

-0.48 ±0.17

-0.23 ±0.12

0.00 ±0.00

-2.31 ±0.18

-2.19 ±0.68

-2.05 ±0.39

-0.34 ±0.25

-1.10 ±0.07

-0.07 ±0.36

0.00 ±0.00

-1.76 ±0.34

-2.44 ±0.39

-1.07 ±0.20

0.82 ±0.25

0.55 ±0.51

0.00 ±0.10

0.00 ±0.00

-2.03 ±0.21

-2.87 ±0.39

-1.93 ±0.08

-1.16 ±0.25

-1.18 ±0.25

-1.14 ±0.10

0.00 ±0.00

-1.17 ±0.01

-2.84 ±0.04

-0.78 ±0.15

-0.55 ±0.18

-0.37 ±0.16

-0.21 ±0.15

0.00 ±0.00

-2.42 ±0.24

-9.96 ±0.40

-5.58 ±0.33

-4.98 ±0.78

-4.26 ±0.22

-0.82 ±0.40

0.00 ±0.00

-2.58 ±0.09

-4.11 ±0.24

-0.87 ±0.58

1.17 ±0.43

1.80 ±0.20

2.55 ±0.24

0.00 ±0.00

1.71 ±0.41

2.42 ±1.09

4.09 ±0.83

6.30 ±1.95

8.48 ±1.55

9.15 ±0.97

0.00 ±0.00

-2.58 ±0.17

-4.07 ±0.28

-0.82 ±0.70

1.12 ±0.42

1.50 ±0.19

2.24 ±0.14

0.00 ±0.00

1.01 ±0.07

-0.37 ±0.40

1.89 ±0.21

3.48 ±0.15

5.95 ±0.11

6.30 ±0.33

0.00 ±0.00

1.43 ±0.29

-0.78 ±0.55

1.31 ±0.12

3.65 ±0.22

5.77 ±0.50

6.32 ±0.41

0.00 ±0.00

1.91 ±0.49

1.15 ±0.16

5.39 ±1.52

5.59 ±0.14

9.97 ±0.58

9.40 ±1.23

0.00 ±0.00

-1.94 ±0.17

-2.61 ±0.44

-2.72 ±0.19

-0.47 ±0.29

0.04 ±0.17

0.69 ±0.12

0.00 ±0.00

-4.94 ±0.32

-5.92 ±0.34

-3.02 ±0.10

-1.74 ±0.31

-0.61 ±0.24

0.13 ±0.28

0.00 ±0.00

-0.50 ±0.09

-1.77 ±0.31

-1.64 ±0.37

-0.68 ±0.57

-0.13 ±0.63

-0.37 ±0.53

0

1

2

3

4

5

6

aniline ester ether para_hydroxylation phenol phenol_noOrthoHbond NO Heteroatoms RotatableBonds Ring Ar_N C_O C_O_noCOO allylic_oxid aryl_methyl bicyclic imide unbrch_alkane urea

Layer

0.00 ±0.00

0.67 ±0.37

-0.89 ±0.30

-0.24 ±0.20

-1.01 ±0.08

0.74 ±0.16

0.20 ±0.22

0.00 ±0.00

-1.95 ±0.28

-3.44 ±0.29

-4.33 ±0.40

-2.73 ±0.33

-2.13 ±0.11

-1.00 ±0.03

0.00 ±0.00

2.53 ±0.23

0.30 ±0.19

0.77 ±0.15

1.41 ±0.13

1.36 ±0.37

2.40 ±0.12

0.00 ±0.00

-1.17 ±0.26

-0.98 ±0.15

-1.92 ±0.31

-1.07 ±0.18

-1.09 ±0.06

0.36 ±0.15

0.00 ±0.00

-1.11 ±0.82

-1.20 ±0.44

-2.40 ±0.62

-1.25 ±1.12

-1.15 ±0.12

-1.63 ±1.77

0.00 ±0.00

7.75 ±0.30

4.26 ±0.34

5.24 ±0.70

5.83 ±0.16

6.00 ±0.08

7.21 ±0.19

0.00 ±0.00

-0.03 ±0.47

-0.57 ±0.61

-0.06 ±0.63

0.67 ±0.20

0.25 ±1.05

0.65 ±1.16

0

1

2

3

4

5

6

10

5

0

5

10

-5.93 ±0.44

-3.00 ±0.17

-1.72 ±0.57

0.00 ±0.00

-5.92 ±0.28

-1.62 ±0.26

-1.36 ±0.17

0.00 ±0.00

-8.69 ±0.33

-2.73 ±0.16

-1.86 ±0.06

0.00 ±0.00

-5.72 ±0.16

-2.10 ±0.23

0.90 ±0.21

0.00 ±0.00

-6.46 ±0.48

0.70 ±0.26

2.99 ±0.25

0.00 ±0.00

-2.29 ±0.07

-0.70 ±0.07

-0.38 ±0.08

0.00 ±0.00

-6.64 ±0.16

-3.09 ±0.24

-2.73 ±0.39

0.00 ±0.00

-6.20 ±0.08

-0.48 ±0.18

-1.63 ±0.48

0.00 ±0.00

-8.81 ±0.22

-3.16 ±0.23

-3.12 ±0.16

0.00 ±0.00

-5.69 ±0.18

-2.11 ±0.21

0.93 ±0.16

0.00 ±0.00

-7.65 ±0.13

-5.15 ±0.25

-3.40 ±0.24

0.00 ±0.00

-6.31 ±0.35

5.16 ±0.29

3.68 ±0.10

0.00 ±0.00

-5.95 ±0.36

8.17 ±1.03

5.70 ±0.45

0.00 ±0.00

-6.13 ±0.11

5.15 ±0.49

3.44 ±0.21

0.00 ±0.00

-7.14 ±0.60

12.19 ±1.32

2.69 ±0.04

0.00 ±0.00

-7.05 ±0.40

10.93 ±1.17

2.29 ±0.18

0.00 ±0.00

-3.46 ±0.25

15.51 ±0.99

12.58 ±0.90

0.00 ±0.00

-5.94 ±0.51

-2.28 ±0.09

-1.57 ±0.19

0.00 ±0.00

-6.22 ±0.41

-1.47 ±0.27

-1.35 ±0.16

0.00 ±0.00

-6.47 ±0.17

-0.78 ±0.21

-1.76 ±0.54

0

1

2

3

0.00 ±0.00

-1.75 ±0.35

-6.01 ±0.20

-2.12 ±0.37

0.00 ±0.00

-3.57 ±0.18

-3.00 ±0.08

-2.17 ±0.20

aniline ester ether para_hydroxylation phenol phenol_noOrthoHbond NO Heteroatoms RotatableBonds Ring Ar_N C_O C_O_noCOO allylic_oxid aryl_methyl bicyclic imide unbrch_alkane urea

-8.57 ±0.21

0.00 ±0.00

Layer

-4.06 ±0.33

-7.06 ±0.23

-3.05 ±0.52

-1.59 ±0.13

-3.72 ±0.36

0.00 ±0.00

-9.28 ±0.20

-1.50 ±0.06

-0.72 ±0.24

0.00 ±0.00

1.25 ±0.29

-0.37 ±0.29

1.36 ±0.49

0.00 ±0.00

-2.21 ±0.13

15.05 ±0.54

12.57 ±0.72

0.00 ±0.00

-2.08 ±0.47

14.98 ±0.39

12.30 ±0.26

0.00 ±0.00

-5.88 ±0.09

13.50 ±0.68

6.68 ±0.64

0.00 ±0.00

-5.82 ±1.42

14.84 ±1.72

8.87 ±1.51

0.00 ±0.00

-6.06 ±0.42

-7.19 ±0.46

-8.75 ±0.28

0.00 ±0.00

-0.03 ±0.08

-0.08 ±0.04

-0.31 ±0.09

0.00 ±0.00

-6.56 ±0.06

2.54 ±0.23

5.10 ±0.18

0.00 ±0.00

-9.48 ±0.27

0.79 ±0.04

-1.20 ±0.03

0.00 ±0.00

-8.87 ±0.29

2.72 ±0.32

0.96 ±0.27

0.00 ±0.00

-9.89 ±0.23

-6.00 ±0.09

-6.92 ±0.15

0.00 ±0.00

-6.00 ±0.77

-6.23 ±0.24

-6.03 ±0.32

0.00 ±0.00

-5.20 ±0.07

-3.73 ±0.52

-2.78 ±0.41

0.00 ±0.00

-1.41 ±0.39

1.84 ±0.04

-0.17 ±1.01

0.00 ±0.00

-9.98 ±0.53

-7.47 ±0.41

-9.69 ±0.95

0.00 ±0.00

-5.54 ±0.46

0.57 ±0.66

-0.22 ±0.21

0

1

2

3

0.00 ±0.00

-4.83 ±0.31

-6.99 ±0.35

-2.50 ±0.38

0.00 ±0.00

-0.62 ±0.20

-3.29 ±0.38

-4.19 ±0.65

0.00 ±0.00

-1.75 ±0.27

-0.61 ±0.21

0.07 ±0.02

10

0

10

20

30

-1.49 ±0.12

-3.39 ±0.25

-2.46 ±0.13

-0.97 ±0.02

-4.12 ±0.13

-0.61 ±0.13

0.00 ±0.00

-0.40 ±0.05

-1.78 ±0.11

-0.48 ±0.17

0.00 ±0.00

0.04 ±0.06

-1.16 ±0.10

0.21 ±0.15

0.00 ±0.00

-1.20 ±0.35

-6.49 ±0.15

-1.35 ±0.11

0.00 ±0.00

-3.93 ±0.34

-4.44 ±0.11

-1.27 ±0.11

0.00 ±0.00

-2.52 ±0.22

-5.79 ±0.29

-2.17 ±0.42

0.00 ±0.00

-0.92 ±0.06

-4.19 ±0.10

-0.54 ±0.11

0.00 ±0.00

-1.20 ±0.11

-3.39 ±0.30

-0.81 ±0.18

0.00 ±0.00

-2.39 ±0.19

-4.44 ±0.26

-4.30 ±0.04

0.00 ±0.00

-2.09 ±0.43

3.92 ±0.84

3.65 ±0.24

0.00 ±0.00

-2.27 ±0.14

-4.40 ±0.33

-4.69 ±0.30

0.00 ±0.00

-2.37 ±0.09

-1.65 ±0.18

-0.83 ±0.12

0.00 ±0.00

-2.22 ±0.24

-2.41 ±0.13

-1.61 ±0.41

0.00 ±0.00

-3.20 ±0.48

-7.24 ±1.20

-6.60 ±0.21

0.00 ±0.00

-2.16 ±0.26

-3.88 ±0.35

-2.54 ±0.33

0.00 ±0.00

-0.88 ±0.26

-4.92 ±0.10

-3.99 ±0.37

0.00 ±0.00

-5.08 ±0.16

-1.99 ±0.60

-5.07 ±0.19

0

1

2

3

0.00 ±0.00

0.63 ±0.18

-2.62 ±0.20

0.71 ±0.55

0.00 ±0.00

-0.04 ±0.29

1.59 ±0.39

-1.85 ±0.53

0.00 ±0.00

-0.31 ±0.03

-1.27 ±0.06

-1.33 ±0.25

0.00 ±0.00

0.02 ±0.03

0.19 ±0.18

3.27 ±0.23

0.00 ±0.00

0.95 ±0.13

5.72 ±0.53

8.54 ±0.25

0.00 ±0.00

0.02 ±0.04

-0.01 ±0.09

1.08 ±0.14

0.00 ±0.00

0.38 ±0.12

-1.99 ±0.11

0.28 ±0.12

0.00 ±0.00

0.04 ±0.26

0.98 ±0.25

-0.70 ±1.03

0.00 ±0.00

-0.27 ±0.06

-0.80 ±0.11

-0.56 ±0.21

Layer

0.00 ±0.00

-2.73 ±0.19

-5.51 ±0.27

0.31 ±0.55

0.00 ±0.00

-2.40 ±0.17

-2.88 ±0.93

-2.22 ±1.24

0.00 ±0.00

-2.13 ±0.17

-2.64 ±1.36

-2.26 ±1.74

0.00 ±0.00

-4.78 ±0.82

2.77 ±1.02

2.97 ±0.81

0.00 ±0.00

-3.86 ±1.62

4.82 ±3.57

4.38 ±0.56

0.00 ±0.00

-1.61 ±0.27

-5.15 ±0.06

-7.36 ±0.14

0.00 ±0.00

-0.08 ±0.04

-0.71 ±0.12

-2.66 ±0.32

0.00 ±0.00

-0.19 ±0.07

-1.26 ±0.22

-0.94 ±0.19

0.00 ±0.00

-3.45 ±0.18

1.70 ±0.09

0.19 ±0.30

0.00 ±0.00

-3.57 ±0.21

0.22 ±0.33

-1.06 ±0.03

0.00 ±0.00

-3.19 ±0.33

-7.93 ±0.24

-4.80 ±0.78

0.00 ±0.00

-2.25 ±0.18

-4.87 ±0.16

-3.16 ±0.64

0.00 ±0.00

-2.21 ±0.26

-4.42 ±0.15

-0.69 ±0.05

0.00 ±0.00

-2.24 ±0.08

-3.23 ±2.23

-1.60 ±0.26

0.00 ±0.00

0.28 ±0.41

-5.53 ±0.23

-5.60 ±0.90

0.00 ±0.00

-2.13 ±0.17

-1.79 ±0.33

-2.10 ±1.10

0

1

2

3

0.00 ±0.00

-0.33 ±0.08

1.24 ±0.56

-3.29 ±0.60

0.00 ±0.00

0.13 ±0.29

0.24 ±0.48

-12.74 ±0.24

0.00 ±0.00

0.24 ±0.27

4.33 ±0.44

0.46 ±0.11

0.00 ±0.00

-0.00 ±0.38

-1.75 ±0.37

2.55 ±0.40

0.00 ±0.00

0.61 ±0.28

5.87 ±0.56

-0.41 ±0.19

0.00 ±0.00

0.62 ±0.29

6.23 ±0.52

0.53 ±1.08

0.00 ±0.00

4.90 ±0.40

11.65 ±0.42

7.17 ±0.80

0.00 ±0.00

3.22 ±0.48

8.36 ±1.15

5.14 ±0.60

0.00 ±0.00

-0.84 ±0.22

-3.09 ±0.17

-11.22 ±0.29

0.00 ±0.00

-0.01 ±0.03

-2.50 ±0.20

-8.98 ±0.39

10

5

0

5

10

15

0.01 ±0.07

0.19 ±0.24

3.28 ±0.15

0.29 ±0.07

1.44 ±0.56

2.17 ±0.40

0.00 ±0.00

-0.33 ±0.18

2.41 ±0.31

-2.51 ±0.30

0.00 ±0.00

3.76 ±0.38

8.56 ±0.68

6.15 ±0.76

0.00 ±0.00

-0.01 ±0.07

2.22 ±0.06

-2.86 ±0.39

0.00 ±0.00

-0.69 ±0.18

4.78 ±0.37

-7.27 ±0.29

0.00 ±0.00

-0.96 ±0.08

3.31 ±0.28

-8.28 ±0.46

0.00 ±0.00

0.96 ±0.33

7.14 ±0.35

-6.01 ±1.01

0.00 ±0.00

-0.09 ±0.18

2.53 ±0.10

0.00 ±0.00

0.15 ±0.19

1.60 ±0.37

-2.37 ±0.47

0.00 ±0.00

0.46 ±0.28

-0.26 ±0.27

-12.29 ±0.74

0

1

2

3

-1.36 ±0.18

Layer

0.00 ±0.00

1.12 ±0.14

7.45 ±0.19

9.97 ±0.31

0.00 ±0.00

-0.50 ±0.14

2.98 ±0.12

-4.09 ±0.16

0.00 ±0.00

-0.55 ±0.26

3.04 ±0.11

-7.04 ±0.33

0.00 ±0.00

-0.62 ±0.24

-0.29 ±0.03

-2.15 ±0.58

0.00 ±0.00

0.36 ±0.27

-0.55 ±0.27

-4.14 ±0.26

0.00 ±0.00

-1.36 ±0.09

-4.33 ±0.52

-4.30 ±0.40

0.00 ±0.00

-0.16 ±0.11

-0.34 ±0.14

-9.25 ±1.17

0.00 ±0.00

-0.70 ±0.24

-0.70 ±0.39

-10.11 ±1.51

0.00 ±0.00

0.31 ±0.48

-0.38 ±1.38

-6.61 ±0.84

0

1

2

3

0

10

20

30

(m) Chemberta-2-77M (PT) AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen NHOH HAcceptors HDonors Al_OH Al_OH_noTert Ar_OH NH0 NH1 amide

0.00 ±0.00

0.15 ±0.10

-0.11 ±0.16

-0.17 ±0.29

-2.29 ±0.57

-0.69 ±0.31

0.05 ±0.11

-0.47 ±0.27

-0.62 ±0.17

-0.44 ±0.20

0.74 ±0.38

1.98 ±0.45

2.36 ±0.10

0.00 ±0.00

-0.24 ±0.21

0.22 ±0.11

-0.56 ±0.45

-1.66 ±0.22

-1.63 ±0.10

-0.68 ±0.05

-0.40 ±0.14

0.17 ±0.08

-0.06 ±0.19

0.39 ±0.34

0.63 ±0.10

1.15 ±0.22

0.00 ±0.00

0.10 ±0.13

-0.14 ±0.13

-0.29 ±0.37

-1.83 ±0.21

-1.14 ±0.23

-0.08 ±0.08

-0.11 ±0.04

-0.07 ±0.03

0.37 ±0.10

0.49 ±0.05

0.12 ±0.13

0.73 ±0.08

0.00 ±0.00

0.02 ±0.10

-0.44 ±0.20

-0.01 ±0.09

-1.35 ±0.13

-0.36 ±0.38

-1.44 ±0.07

-0.79 ±0.05

-0.56 ±0.08

-0.13 ±0.07

0.22 ±0.10

-0.17 ±0.22

0.91 ±0.07

0.00 ±0.00

0.11 ±0.09

0.09 ±0.26

-0.10 ±0.35

-0.13 ±0.26

0.03 ±0.11

-2.13 ±0.19

-0.90 ±0.00

-1.08 ±0.17

0.11 ±0.10

0.65 ±0.03

0.94 ±0.21

3.35 ±0.35

0.00 ±0.00

-0.26 ±0.22

0.01 ±0.18

0.31 ±0.19

-1.36 ±0.21

-0.25 ±0.07

0.05 ±0.08

-0.11 ±0.02

0.16 ±0.06

-0.04 ±0.05

0.28 ±0.07

0.18 ±0.10

0.37 ±0.12

0.00 ±0.00

-0.24 ±0.03

-0.50 ±0.44

-0.42 ±0.09

-2.50 ±0.69

-0.43 ±0.26

0.70 ±0.43

0.47 ±0.10

0.44 ±0.41

0.99 ±0.13

0.45 ±0.13

0.92 ±0.30

2.91 ±0.44

0.00 ±0.00

-0.20 ±0.26

0.23 ±0.23

-0.13 ±0.11

-1.14 ±0.10

-1.06 ±0.38

-0.42 ±0.14

-0.87 ±0.30

-0.32 ±0.07

0.36 ±0.19

0.47 ±0.13

0.29 ±0.21

1.17 ±0.22

0.00 ±0.00

-0.30 ±0.25

-0.20 ±0.36

-0.21 ±0.29

-1.44 ±0.40

-0.84 ±0.30

0.10 ±0.04

-0.50 ±0.11

-0.07 ±0.11

0.71 ±0.08

0.69 ±0.29

-0.10 ±0.28

1.01 ±0.22

0.00 ±0.00

-0.02 ±0.08

-0.50 ±0.27

-0.00 ±0.16

-1.51 ±0.15

-0.29 ±0.35

-1.39 ±0.13

-0.80 ±0.05

-0.64 ±0.08

-0.18 ±0.08

0.22 ±0.02

-0.23 ±0.22

0.94 ±0.17

0.00 ±0.00

-0.10 ±0.21

-0.04 ±0.25

0.76 ±0.14

-2.07 ±0.46

1.96 ±0.12

1.61 ±0.20

-0.34 ±0.21

0.79 ±0.38

0.54 ±0.24

0.60 ±0.19

1.72 ±0.52

5.19 ±0.30

0.00 ±0.00

0.30 ±0.46

0.27 ±0.29

-1.22 ±0.25

-1.28 ±0.11

0.44 ±0.33

0.24 ±0.34

0.38 ±0.17

0.52 ±0.16

1.58 ±0.35

2.45 ±0.23

3.01 ±0.10

2.93 ±0.24

0.00 ±0.00

0.14 ±0.10

0.03 ±0.29

-1.57 ±0.37

-0.17 ±0.52

-0.62 ±0.29

1.81 ±0.62

0.78 ±0.75

1.05 ±0.31

1.17 ±0.66

2.32 ±1.22

1.61 ±0.66

3.03 ±0.53

0.00 ±0.00

0.32 ±0.72

0.37 ±0.39

-1.15 ±0.09

-1.16 ±0.47

0.12 ±0.36

0.16 ±0.50

0.40 ±0.21

0.65 ±0.02

1.54 ±0.14

2.57 ±0.11

2.83 ±0.25

2.76 ±0.22

0.00 ±0.00

0.35 ±0.08

0.19 ±0.38

-0.64 ±0.27

-2.27 ±0.46

-0.18 ±0.21

2.29 ±0.23

1.86 ±0.44

2.34 ±0.18

2.66 ±0.42

1.40 ±0.19

3.59 ±0.60

4.30 ±0.63

0.00 ±0.00

0.17 ±0.06

0.43 ±0.16

-0.37 ±0.59

-2.37 ±0.25

-0.44 ±0.19

1.96 ±0.44

1.65 ±0.37

2.04 ±0.38

2.30 ±0.39

0.91 ±0.21

2.74 ±0.80

3.22 ±0.39

0.00 ±0.00

-0.25 ±0.25

0.35 ±0.33

-1.37 ±0.14

-0.88 ±0.31

-1.15 ±0.22

2.57 ±0.34

0.58 ±0.61

0.85 ±0.68

2.73 ±0.32

6.80 ±1.03

8.21 ±0.63

9.36 ±0.51

0.00 ±0.00

-0.40 ±0.36

-0.50 ±0.17

-0.89 ±0.32

-0.77 ±0.23

-1.31 ±0.26

-0.41 ±0.10

-0.30 ±0.04

0.17 ±0.13

-0.80 ±0.36

0.07 ±0.08

-0.39 ±0.40

0.30 ±0.20

0.00 ±0.00

0.46 ±0.16

0.06 ±0.05

-1.24 ±0.31

-0.52 ±0.13

-2.74 ±0.28

-1.50 ±0.30

-0.07 ±0.05

-0.38 ±0.29

-0.95 ±0.30

-0.78 ±0.20

-0.61 ±0.36

0.91 ±0.35

0.00 ±0.00

-0.19 ±0.35

-0.72 ±0.26

-0.57 ±0.58

-0.54 ±0.14

-0.46 ±0.25

0.25 ±0.37

-0.32 ±0.36

1.02 ±0.60

-0.77 ±0.38

-0.02 ±0.47

-0.19 ±0.22

0.89 ±0.15

0

1

2

3

4

5

6

7

8

9

10

11

12

aniline ester ether para_hydroxylation phenol phenol_noOrthoHbond NO Heteroatoms RotatableBonds Ring Ar_N C_O C_O_noCOO allylic_oxid aryl_methyl bicyclic imide unbrch_alkane urea

-0.44 ±0.31

-0.56 ±0.35

-0.40 ±0.15

-0.17 ±0.12

0.31 ±0.17

0.50 ±0.35

0.42 ±0.27

0.69 ±0.35

0.73 ±0.39

0.63 ±0.54

0.63 ±0.31

0.55 ±0.27

0.61 ±0.17

0.71 ±0.33

0.47 ±0.54

0.49 ±0.43

0.58 ±0.11

0.62 ±0.18

0.55 ±0.10

0.50 ±0.20

0.45 ±0.26

0.32 ±0.34

0.34 ±0.44

0.50 ±0.29

0.35 ±0.37

0.36 ±0.12

0.30 ±0.30

-0.08 ±0.11

-0.10 ±0.24

0.12 ±0.36

0.09 ±0.38

0.14 ±0.43

0.09 ±0.37

0.08 ±0.32

0.20 ±0.24

0.34 ±0.25

0.33 ±0.20

0.14 ±0.23

0.25 ±0.29

0.40 ±0.22

0.20 ±0.09

0.15 ±0.25

0.39 ±0.12

0.59 ±0.23

0.78 ±0.41

0.72 ±0.18

0.79 ±0.14

0.66 ±0.18

0.00 ±0.00

0.01 ±0.06

0.12 ±0.22

0.16 ±0.24

0.19 ±0.31

0.34 ±0.51

0.11 ±0.43

0.21 ±0.43

0.41 ±0.26

0.62 ±0.34

0.82 ±0.21

1.02 ±0.09

0.99 ±0.27

0.00 ±0.00

1.96 ±0.35

1.61 ±0.35

1.28 ±0.22

1.25 ±0.04

1.07 ±0.33

1.05 ±0.60

1.16 ±0.47

0.97 ±0.41

0.69 ±0.30

0.88 ±0.25

0.70 ±0.12

0.43 ±0.39

0.00 ±0.00

2.30 ±0.61

2.08 ±0.41

2.39 ±0.57

2.62 ±0.16

1.74 ±0.29

1.17 ±0.39

1.09 ±0.54

0.79 ±0.59

-0.07 ±0.87

-0.54 ±0.86

-0.48 ±1.01

-0.51 ±0.57

0.00 ±0.00

-1.24 ±0.18

-0.66 ±0.46

0.00 ±0.38

0.14 ±0.18

0.49 ±0.21

0.43 ±0.36

0.53 ±0.20

0.66 ±0.21

1.02 ±0.47

1.21 ±0.49

1.19 ±0.41

0.90 ±0.40

0.00 ±0.00

0.00 ±0.03

0.02 ±0.01

0.05 ±0.01

0.02 ±0.02

0.02 ±0.02

0.03 ±0.01

0.00 ±0.01

-0.00 ±0.01

-0.04 ±0.04

-0.05 ±0.03

-0.05 ±0.05

-0.02 ±0.05

0.00 ±0.00

0.06 ±0.07

-0.02 ±0.08

0.05 ±0.09

-0.03 ±0.09

0.01 ±0.08

0.01 ±0.05

-0.07 ±0.08

-0.02 ±0.06

-0.00 ±0.03

-0.04 ±0.05

-0.06 ±0.03

-0.06 ±0.05

0.00 ±0.00

0.29 ±0.11

0.22 ±0.14

0.24 ±0.17

0.35 ±0.21

0.57 ±0.14

0.65 ±0.10

0.86 ±0.02

1.16 ±0.11

1.27 ±0.15

1.36 ±0.21

1.47 ±0.06

1.85 ±0.14

0.00 ±0.00

0.26 ±0.15

0.14 ±0.16

0.18 ±0.06

0.34 ±0.03

0.62 ±0.13

0.63 ±0.27

0.60 ±0.22

0.89 ±0.33

1.11 ±0.12

1.19 ±0.16

1.06 ±0.10

1.50 ±0.25

0.00 ±0.00

-0.02 ±0.04

-0.25 ±0.12

-0.09 ±0.07

0.06 ±0.03

0.03 ±0.34

0.23 ±0.48

0.26 ±0.31

0.39 ±0.37

0.26 ±0.37

0.23 ±0.51

0.19 ±0.34

0.15 ±0.53

0.00 ±0.00

-1.17 ±1.08

-0.40 ±1.18

-0.26 ±0.24

-0.34 ±0.21

-0.70 ±0.41

-0.63 ±0.50

-0.62 ±0.44

-0.25 ±0.40

-0.38 ±0.28

-0.13 ±0.29

-0.30 ±0.31

-0.43 ±0.25

0.00 ±0.00

0.07 ±0.16

0.25 ±0.08

0.19 ±0.12

0.14 ±0.09

0.15 ±0.16

0.06 ±0.11

0.04 ±0.15

0.15 ±0.34

0.32 ±0.20

0.21 ±0.29

0.24 ±0.12

0.16 ±0.17

0.00 ±0.00

-0.28 ±0.34

-0.30 ±0.16

0.19 ±0.49

0.53 ±0.44

0.37 ±0.64

0.50 ±0.56

0.69 ±0.39

0.44 ±0.58

0.38 ±0.74

0.42 ±0.72

0.23 ±0.25

0.17 ±0.34

0.00 ±0.00

-0.17 ±0.10

-0.15 ±0.13

-0.00 ±0.11

-0.29 ±0.14

0.09 ±0.28

0.35 ±0.11

-0.02 ±0.27

0.27 ±0.14

0.32 ±0.19

0.42 ±0.59

0.75 ±0.34

0.81 ±0.48

0.00 ±0.00

-0.19 ±0.59

0.06 ±0.39

0.13 ±0.16

0.26 ±0.58

0.49 ±0.54

0.98 ±0.13

0.94 ±0.54

1.26 ±0.59

1.11 ±0.51

0.90 ±0.67

0.74 ±0.88

0.47 ±0.45

0

1

2

3

4

5

6

7

8

9

10

11

12

5

0

5

10

15

-0.38 ±0.15

-0.39 ±0.10

-0.73 ±0.03

-0.22 ±0.04

-0.60 ±0.11

-0.78 ±0.11

-1.29 ±0.20

-1.81 ±0.27

-1.79 ±0.10

-0.14 ±0.08

-0.30 ±0.03

-0.48 ±0.09

-0.49 ±0.08

-0.57 ±0.10

-0.91 ±0.04

-1.13 ±0.03

-1.15 ±0.09

-1.40 ±0.11

0.00 ±0.00

0.00 ±0.19

0.09 ±0.14

0.14 ±0.09

0.12 ±0.04

0.05 ±0.08

0.04 ±0.15

0.14 ±0.03

0.19 ±0.10

0.08 ±0.07

-0.04 ±0.17

-0.19 ±0.12

-0.21 ±0.09

-0.27 ±0.05

-0.38 ±0.08

0.00 ±0.00

-0.68 ±0.05

-0.75 ±0.06

-0.51 ±0.19

-0.69 ±0.13

-0.74 ±0.10

-0.90 ±0.23

-0.87 ±0.09

-1.09 ±0.17

-1.18 ±0.18

-1.43 ±0.18

-1.49 ±0.28

-2.28 ±0.09

-2.78 ±0.05

-3.15 ±0.12

0.00 ±0.00

-0.03 ±0.03

0.05 ±0.02

-0.01 ±0.05

-0.02 ±0.04

-0.04 ±0.06

0.01 ±0.09

-0.05 ±0.08

-0.06 ±0.12

0.02 ±0.07

-0.16 ±0.03

-0.01 ±0.04

-0.07 ±0.04

-0.08 ±0.02

-0.13 ±0.03

0.00 ±0.00

0.10 ±0.09

0.62 ±0.15

0.66 ±0.17

0.26 ±0.23

0.46 ±0.10

0.51 ±0.15

0.38 ±0.30

0.19 ±0.10

0.03 ±0.29

-0.53 ±0.06

-0.65 ±0.25

-0.67 ±0.07

-1.08 ±0.31

-1.73 ±0.34

0.00 ±0.00

-0.30 ±0.07

0.08 ±0.13

-0.01 ±0.24

0.12 ±0.21

-0.23 ±0.33

-0.20 ±0.39

-0.58 ±0.28

-0.53 ±0.12

-0.34 ±0.28

-0.45 ±0.14

-0.65 ±0.13

-1.41 ±0.04

-1.99 ±0.13

-2.14 ±0.05

0.00 ±0.00

-0.15 ±0.23

-0.19 ±0.17

-0.23 ±0.13

-0.23 ±0.05

-0.28 ±0.17

-0.38 ±0.17

-0.31 ±0.16

-0.60 ±0.12

-0.57 ±0.08

-0.73 ±0.14

-0.94 ±0.03

-1.18 ±0.09

-1.48 ±0.06

-1.83 ±0.17

0.00 ±0.00

0.06 ±0.06

0.10 ±0.10

0.10 ±0.12

0.17 ±0.08

0.03 ±0.08

0.11 ±0.09

0.19 ±0.16

0.25 ±0.05

0.08 ±0.11

-0.07 ±0.13

-0.16 ±0.11

-0.18 ±0.04

-0.26 ±0.06

-0.31 ±0.02

0.00 ±0.00

-0.29 ±0.19

-0.18 ±0.13

-0.25 ±0.03

-0.18 ±0.20

-0.20 ±0.02

-0.47 ±0.06

-0.17 ±0.07

-0.27 ±0.07

-0.47 ±0.07

-1.12 ±0.13

-1.18 ±0.16

-1.12 ±0.12

-1.51 ±0.11

-2.34 ±0.29

0.00 ±0.00

0.11 ±0.07

0.07 ±0.08

0.43 ±0.21

-0.03 ±0.13

-0.57 ±0.11

-0.50 ±0.07

-1.12 ±0.08

-1.10 ±0.07

-1.11 ±0.08

-1.39 ±0.04

-1.56 ±0.08

-1.84 ±0.22

-1.87 ±0.16

-2.75 ±0.05

0.00 ±0.00

1.26 ±0.32

1.69 ±0.20

1.48 ±0.71

1.68 ±0.14

1.21 ±0.31

1.24 ±0.15

0.91 ±0.41

0.75 ±0.28

0.23 ±0.25

-0.35 ±0.42

-0.39 ±0.43

-1.12 ±0.47

-1.56 ±0.35

-2.74 ±0.26

0.00 ±0.00

0.16 ±0.33

0.04 ±0.20

0.43 ±0.21

0.04 ±0.20

-0.63 ±0.07

-0.40 ±0.19

-1.18 ±0.10

-1.07 ±0.07

-1.28 ±0.08

-1.33 ±0.37

-1.64 ±0.12

-1.98 ±0.19

-1.82 ±0.12

-2.74 ±0.17

0.00 ±0.00

0.61 ±0.21

0.57 ±0.11

0.58 ±0.18

0.11 ±0.06

-0.25 ±0.16

-0.19 ±0.15

-0.21 ±0.13

-0.28 ±0.17

-0.62 ±0.33

-1.39 ±0.16

-1.52 ±0.25

-1.46 ±0.46

-1.59 ±0.33

-2.44 ±0.19

0.00 ±0.00

0.68 ±0.28

0.86 ±0.10

0.75 ±0.11

0.67 ±0.09

0.07 ±0.13

-0.01 ±0.11

0.00 ±0.16

-0.24 ±0.10

-0.74 ±0.24

-1.06 ±0.37

-1.28 ±0.34

-1.14 ±0.50

-1.13 ±0.18

-1.91 ±0.10

0.00 ±0.00

1.69 ±0.25

2.69 ±0.19

2.04 ±0.20

2.43 ±0.23

2.54 ±0.23

2.24 ±0.42

1.92 ±0.32

1.94 ±0.24

1.52 ±0.10

1.47 ±0.41

1.08 ±0.52

0.50 ±0.31

-0.64 ±0.40

-1.40 ±0.31

0.00 ±0.00

-0.19 ±0.11

-0.17 ±0.15

-0.22 ±0.10

-0.28 ±0.18

-0.37 ±0.17

-0.61 ±0.16

-0.90 ±0.05

-0.76 ±0.11

-0.98 ±0.18

-1.23 ±0.16

-1.84 ±0.17

-2.21 ±0.08

-2.73 ±0.09

-2.68 ±0.27

0.00 ±0.00

-0.51 ±0.24

-0.50 ±0.20

-0.08 ±0.11

-0.39 ±0.05

-0.61 ±0.13

-0.76 ±0.36

-1.13 ±0.11

-1.19 ±0.24

-1.43 ±0.18

-1.98 ±0.38

-3.11 ±0.13

-3.34 ±0.10

-3.67 ±0.17

-4.83 ±0.13

0.00 ±0.00

0.20 ±0.19

0.70 ±0.10

0.83 ±0.17

1.02 ±0.06

0.72 ±0.16

0.67 ±0.09

0.50 ±0.18

0.20 ±0.05

-0.32 ±0.11

-0.35 ±0.22

-1.00 ±0.34

-1.44 ±0.19

-1.46 ±0.18

-2.12 ±0.15

0

1

2

3

4

5

6

7

8

9

10

11

12

13

14

aniline ester ether para_hydroxylation phenol phenol_noOrthoHbond NO Heteroatoms RotatableBonds Ring Ar_N C_O C_O_noCOO allylic_oxid aryl_methyl bicyclic imide unbrch_alkane urea

Layer

0.00 ±0.00

0.74 ±0.38

1.02 ±0.06

0.98 ±0.16

1.15 ±0.29

1.16 ±0.28

0.75 ±0.15

0.74 ±0.08

0.80 ±0.27

0.45 ±0.11

0.35 ±0.16

0.40 ±0.17

-0.07 ±0.39

-0.45 ±0.38

-1.28 ±0.32

0.00 ±0.00

0.57 ±0.22

0.62 ±0.23

0.69 ±0.29

0.81 ±0.28

0.70 ±0.24

0.72 ±0.17

0.30 ±0.14

0.22 ±0.24

-0.17 ±0.22

0.17 ±0.21

-0.18 ±0.22

-0.51 ±0.24

-0.74 ±0.07

-1.22 ±0.21

0.00 ±0.00

-0.82 ±0.08

-0.37 ±0.26

-0.47 ±0.28

-0.23 ±0.27

-0.34 ±0.13

-0.37 ±0.18

-0.48 ±0.21

-0.89 ±0.11

-1.03 ±0.43

-1.31 ±0.39

-1.62 ±0.08

-1.57 ±0.14

-1.93 ±0.03

-2.37 ±0.39

0.00 ±0.00

0.40 ±0.21

0.36 ±0.09

0.26 ±0.05

0.29 ±0.04

0.15 ±0.07

0.02 ±0.04

-0.01 ±0.10

-0.01 ±0.04

-0.17 ±0.12

-0.18 ±0.09

-0.49 ±0.25

-0.55 ±0.09

-0.83 ±0.12

-0.74 ±0.09

0.00 ±0.00

1.46 ±0.15

2.25 ±0.09

1.70 ±0.39

1.80 ±0.25

1.98 ±0.24

1.93 ±0.27

1.46 ±0.48

1.67 ±0.22

1.23 ±0.30

1.34 ±0.30

1.07 ±0.16

0.38 ±0.08

-0.11 ±0.07

-1.01 ±0.18

0.00 ±0.00

1.50 ±0.23

2.29 ±0.04

1.67 ±0.12

1.66 ±0.23

1.87 ±0.18

1.81 ±0.19

1.36 ±0.40

1.31 ±0.16

1.27 ±0.29

1.19 ±0.38

0.99 ±0.29

0.54 ±0.30

-0.14 ±0.21

-0.98 ±0.47

0.00 ±0.00

1.36 ±0.54

1.66 ±0.70

1.44 ±0.44

1.50 ±0.37

1.31 ±0.29

1.16 ±0.30

0.87 ±0.29

0.56 ±0.16

-0.33 ±0.32

-0.90 ±0.21

-1.39 ±0.12

-1.99 ±0.08

-2.71 ±0.20

-3.35 ±0.13

0.00 ±0.00

1.21 ±0.04

1.86 ±0.39

1.77 ±0.33

2.09 ±0.30

1.33 ±0.43

1.28 ±0.63

0.59 ±0.30

0.19 ±0.09

0.04 ±0.33

-0.76 ±0.52

-1.57 ±0.62

-2.37 ±0.48

-3.59 ±0.43

-4.36 ±0.24

0.00 ±0.00

0.13 ±0.45

0.35 ±0.27

0.34 ±0.09

0.16 ±0.23

0.59 ±0.09

0.23 ±0.19

0.39 ±0.26

0.53 ±0.38

0.26 ±0.23

-0.51 ±0.19

-0.65 ±0.25

-1.11 ±0.06

-1.27 ±0.27

-1.63 ±0.23

0.00 ±0.00

-0.01 ±0.01

-0.02 ±0.04

-0.01 ±0.01

0.00 ±0.03

-0.01 ±0.02

0.00 ±0.01

0.02 ±0.00

-0.01 ±0.01

-0.00 ±0.01

-0.05 ±0.05

-0.03 ±0.02

-0.02 ±0.05

-0.02 ±0.01

-0.02 ±0.02

0.00 ±0.00

-1.00 ±0.09

-0.72 ±0.21

-0.44 ±0.05

-0.50 ±0.08

-0.61 ±0.22

-0.83 ±0.10

-1.00 ±0.13

-1.21 ±0.07

-1.12 ±0.07

-1.67 ±0.05

-2.09 ±0.13

-2.42 ±0.16

-2.79 ±0.05

-3.01 ±0.10

0.00 ±0.00

0.03 ±0.17

0.19 ±0.16

0.20 ±0.20

0.21 ±0.04

0.15 ±0.08

0.12 ±0.10

-0.23 ±0.09

-0.22 ±0.07

-0.35 ±0.25

-0.31 ±0.05

-0.56 ±0.14

-0.99 ±0.10

-1.30 ±0.17

-1.53 ±0.12

0.00 ±0.00

0.20 ±0.25

0.61 ±0.11

0.71 ±0.11

0.99 ±0.09

0.87 ±0.05

0.80 ±0.17

0.23 ±0.28

0.54 ±0.30

0.62 ±0.18

0.73 ±0.09

0.38 ±0.20

0.26 ±0.14

-0.12 ±0.03

-0.62 ±0.14

0.00 ±0.00

0.05 ±0.21

-0.13 ±0.32

0.00 ±0.10

-0.01 ±0.25

0.07 ±0.13

-0.05 ±0.19

-0.26 ±0.17

-0.26 ±0.35

-0.80 ±0.15

-0.97 ±0.10

-1.33 ±0.04

-1.28 ±0.24

-1.56 ±0.27

-1.83 ±0.14

0.00 ±0.00

1.76 ±0.11

1.16 ±0.21

0.97 ±0.09

1.13 ±0.06

0.89 ±0.16

0.80 ±0.34

1.13 ±0.09

0.88 ±0.16

0.83 ±0.28

0.53 ±0.14

0.34 ±0.16

0.19 ±0.39

0.06 ±0.41

-0.41 ±0.34

0.00 ±0.00

0.73 ±0.14

0.94 ±0.10

0.40 ±0.11

0.69 ±0.19

0.52 ±0.10

0.36 ±0.13

0.30 ±0.09

0.16 ±0.02

0.06 ±0.12

-0.11 ±0.24

-0.21 ±0.22

-0.42 ±0.27

-0.72 ±0.05

-0.81 ±0.23

0.00 ±0.00

0.70 ±0.12

1.13 ±0.12

1.16 ±0.29

1.22 ±0.56

1.19 ±0.47

1.34 ±0.10

1.00 ±0.03

0.90 ±0.12

0.50 ±0.29

0.24 ±0.56

0.04 ±0.45

-0.24 ±0.48

-0.60 ±0.53

-1.20 ±0.46

0.00 ±0.00

-0.04 ±0.30

0.73 ±0.17

1.81 ±0.37

2.01 ±0.64

1.69 ±0.44

1.02 ±0.31

0.97 ±0.46

0.65 ±0.33

0.12 ±0.47

-0.21 ±0.52

-0.09 ±0.21

0.11 ±0.50

-0.54 ±0.15

-1.18 ±0.24

0.00 ±0.00

0.31 ±0.20

0.53 ±0.29

0.56 ±0.43

0.61 ±0.31

0.59 ±0.20

0.45 ±0.17

0.09 ±0.16

-0.19 ±0.08

-0.45 ±0.19

-0.60 ±0.34

-0.37 ±0.40

-0.51 ±0.52

-0.90 ±0.38

-1.06 ±0.41

0

1

2

3

4

5

6

7

8

9

10

11

12

13

14

20 15 10 5 0 5 10

AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen NHOH HAcceptors HDonors Al_OH Al_OH_noTert Ar_OH NH0 NH1 amide

0.00 ±0.00

0.29 ±0.09

-0.75 ±0.39

-1.44 ±0.15

-2.40 ±0.41

-3.31 ±0.37

-3.01 ±0.43

0.00 ±0.00

-2.11 ±0.36

-2.51 ±0.34

-3.58 ±0.45

-4.60 ±0.29

-4.76 ±0.28

-4.83 ±0.43

0.00 ±0.00

-1.85 ±0.30

-2.35 ±0.10

-3.19 ±0.17

-3.96 ±0.17

-4.80 ±0.11

-4.77 ±0.17

0.00 ±0.00

-0.29 ±0.10

-0.73 ±0.04

-1.08 ±0.08

-1.41 ±0.10

-1.91 ±0.15

-1.80 ±0.11

0.00 ±0.00

-2.64 ±0.31

-3.16 ±0.20

-4.82 ±0.05

-5.55 ±0.05

-6.48 ±0.26

-5.89 ±0.23

0.00 ±0.00

-0.66 ±0.02

-0.46 ±0.06

-0.48 ±0.07

-0.77 ±0.08

-1.11 ±0.06

-1.24 ±0.08

0.00 ±0.00

0.13 ±0.09

-0.64 ±0.33

-1.47 ±0.39

-2.35 ±0.25

-3.28 ±0.17

-3.25 ±0.04

0.00 ±0.00

-2.17 ±0.27

-2.34 ±0.40

-3.04 ±0.57

-4.09 ±0.25

-4.13 ±0.01

-4.28 ±0.35

0.00 ±0.00

-2.01 ±0.31

-2.54 ±0.20

-3.40 ±0.27

-4.32 ±0.53

-5.07 ±0.08

-4.65 ±0.38

0.00 ±0.00

-0.32 ±0.02

-0.71 ±0.09

-1.07 ±0.10

-1.44 ±0.07

-1.94 ±0.14

-1.85 ±0.09

0.00 ±0.00

-4.48 ±0.17

-3.97 ±0.05

-6.29 ±0.11

-6.67 ±0.23

-7.34 ±0.05

-5.72 ±0.24

0.00 ±0.00

-1.22 ±0.08

-1.83 ±0.23

-2.69 ±0.18

-3.35 ±0.24

-3.76 ±0.57

-3.84 ±0.35

0.00 ±0.00

4.64 ±0.96

4.90 ±1.11

3.67 ±0.64

3.05 ±0.55

1.92 ±0.85

0.82 ±0.92

0.00 ±0.00

-0.88 ±0.13

-1.59 ±0.18

-2.54 ±0.40

-3.01 ±0.41

-3.66 ±0.46

-3.67 ±0.33

0.00 ±0.00

3.11 ±0.38

2.35 ±0.39

1.92 ±0.21

0.50 ±0.21

-0.47 ±0.27

-0.98 ±0.16

0.00 ±0.00

3.42 ±0.35

2.81 ±0.29

2.44 ±0.30

1.18 ±0.07

0.27 ±0.18

-0.32 ±0.14

0.00 ±0.00

3.51 ±0.40

3.32 ±0.41

2.41 ±0.75

1.62 ±0.34

0.57 ±0.35

0.46 ±0.23

0.00 ±0.00

-1.09 ±0.02

-1.93 ±0.12

-2.07 ±0.33

-2.89 ±0.15

-3.28 ±0.04

-3.43 ±0.26

0.00 ±0.00

-2.06 ±0.59

-3.70 ±0.40

-4.44 ±0.33

-5.12 ±0.38

-6.26 ±0.15

-5.54 ±0.09

0.00 ±0.00

2.22 ±0.26

2.13 ±0.18

3.26 ±0.23

2.12 ±0.50

1.14 ±0.37

-0.05 ±0.41

0

1

2

3

4

5

6

0.00 ±0.00

-6.43 ±0.37

-8.82 ±0.04

-8.70 ±0.17

-9.37 ±0.31

-10.14 ±0.27

-13.61 ±0.26

-7.04 ±0.34

-9.38 ±0.42

-10.35 ±0.32

-10.51 ±0.54

-11.20 ±0.46

-14.42 ±0.37

-8.13 ±0.10

-10.58 ±0.05

-10.90 ±0.17

-11.38 ±0.14

-12.48 ±0.25

-16.62 ±0.27

aniline ester ether para_hydroxylation phenol phenol_noOrthoHbond NO Heteroatoms RotatableBonds Ring Ar_N C_O C_O_noCOO allylic_oxid aryl_methyl bicyclic imide unbrch_alkane urea

Layer

0.00 ±0.00

1.37 ±0.26

0.73 ±0.27

0.22 ±0.20

-0.63 ±0.18

-1.32 ±0.26

-2.48 ±0.23

0.00 ±0.00

0.01 ±0.62

-0.62 ±0.54

-0.16 ±0.30

-0.64 ±0.61

-2.14 ±0.54

-3.60 ±0.29

0.00 ±0.00

-2.71 ±0.09

-3.61 ±0.22

-3.95 ±0.04

-4.34 ±0.17

-5.84 ±0.06

-6.09 ±0.38

0.00 ±0.00

0.26 ±0.14

0.09 ±0.13

-0.13 ±0.33

-0.44 ±0.24

-0.73 ±0.37

-1.23 ±0.28

0.00 ±0.00

3.04 ±0.32

3.25 ±0.54

3.15 ±0.34

2.24 ±0.40

1.60 ±0.42

1.09 ±0.52

0.00 ±0.00

3.10 ±0.23

3.14 ±0.33

2.74 ±0.37

2.38 ±0.28

1.64 ±0.59

1.08 ±0.38

0.00 ±0.00

5.19 ±0.55

4.16 ±0.50

3.55 ±0.27

2.57 ±0.54

1.60 ±0.04

0.17 ±0.42

0.00 ±0.00

6.80 ±0.72

5.46 ±1.49

5.88 ±1.14

4.46 ±1.12

4.37 ±0.87

3.31 ±0.74

0.00 ±0.00

-0.54 ±0.55

-0.72 ±0.47

-2.04 ±0.42

-3.44 ±0.47

-4.75 ±0.48

-6.19 ±0.59

0.00 ±0.00

-0.69 ±0.06

-0.89 ±0.09

-1.17 ±0.25

-1.61 ±0.22

-1.61 ±0.05

-1.46 ±0.13

0.00 ±0.00

-3.02 ±0.22

-3.77 ±0.36

-4.73 ±0.27

-5.41 ±0.14

-6.70 ±0.17

-6.39 ±0.17

0.00 ±0.00

0.75 ±0.12

0.83 ±0.16

0.38 ±0.05

-0.37 ±0.06

-1.01 ±0.09

-1.46 ±0.12

0.00 ±0.00

2.89 ±0.22

3.28 ±0.24

2.68 ±0.11

1.59 ±0.04

0.60 ±0.11

-0.24 ±0.27

0.00 ±0.00

-0.86 ±0.13

-0.47 ±0.24

-1.27 ±0.19

-1.74 ±0.33

-2.11 ±0.23

-2.52 ±0.25

0.00 ±0.00

1.34 ±0.30

0.45 ±0.28

-0.50 ±0.33

-1.30 ±0.25

-2.11 ±0.06

-3.51 ±0.22

0.00 ±0.00

-2.03 ±0.21

-2.32 ±0.08

-3.51 ±0.08

-4.46 ±0.26

-4.95 ±0.23

-5.30 ±0.18

0.00 ±0.00

-1.05 ±0.32

-0.54 ±0.41

-0.17 ±0.55

-1.13 ±0.18

-2.14 ±0.19

-4.14 ±0.41

0.00 ±0.00

5.86 ±0.52

3.87 ±0.47

3.17 ±0.44

1.45 ±0.75

-0.33 ±0.56

-2.51 ±1.04

0.00 ±0.00

-0.57 ±0.27

-0.65 ±0.30

0.38 ±0.52

-0.24 ±0.88

-0.98 ±0.67

-2.24 ±0.75

0

1

2

3

4

5

6

30

20

10

0

10

AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen NHOH HAcceptors HDonors Al_OH Al_OH_noTert Ar_OH NH0 NH1 amide

0.00 ±0.00 0.00 ±0.00 0.00 ±0.00

-7.01 ±0.17

-6.62 ±0.17

-7.19 ±0.25

-7.69 ±0.14

-9.28 ±0.01

-13.25 ±0.08

0.00 ±0.00

-12.44 ±0.43

-14.15 ±0.30

-16.27 ±0.33

-17.14 ±0.22

-19.11 ±0.37

-24.58 ±0.51

0.00 ±0.00

-3.28 ±0.12

-4.77 ±0.10

-5.32 ±0.16

-5.16 ±0.16

-6.70 ±0.07

-11.12 ±0.08

0.00 ±0.00

-6.45 ±0.51

-9.80 ±0.12

-9.32 ±0.48

-9.74 ±0.39

-10.76 ±0.25

-14.30 ±0.60

0.00 ±0.00

-7.82 ±0.73

-11.05 ±0.30

-11.13 ±0.32

-11.01 ±0.36

-12.36 ±0.49

-15.89 ±0.22

0.00 ±0.00

-8.89 ±0.24

-12.11 ±0.12

-12.26 ±0.27

-12.47 ±0.33

-13.39 ±0.19

-17.02 ±0.15

0.00 ±0.00

-6.96 ±0.12

-6.66 ±0.22

-9.38 ±0.04

-13.36 ±0.18

-7.26 ±0.30

-7.70 ±0.13

0.00 ±0.00

-8.84 ±0.29

-9.87 ±0.25

-13.78 ±0.36

-15.90 ±0.42

-19.76 ±0.29

-28.08 ±0.35

0.00 ±0.00

-8.18 ±0.24

-9.83 ±0.29

-10.64 ±0.44

-11.41 ±0.21

-12.51 ±0.45

-16.09 ±0.56

0.00 ±0.00

-6.39 ±0.39

-9.20 ±1.02

-11.80 ±0.82

-12.50 ±0.96

-16.09 ±0.44

-22.71 ±0.47

0.00 ±0.00

-8.16 ±0.49

-9.84 ±0.25

-10.38 ±0.50

-11.27 ±0.28

-12.39 ±0.28

-15.84 ±0.24

0.00 ±0.00

-4.99 ±0.30

-6.54 ±0.12

-8.28 ±0.50

-9.29 ±0.45

-11.45 ±0.31

-15.05 ±0.42

0.00 ±0.00

-4.23 ±0.14

-6.18 ±0.50

-7.86 ±0.15

-9.49 ±0.35

-11.18 ±0.13

-14.45 ±0.62

0.00 ±0.00

-0.68 ±0.56

-3.33 ±0.76

-4.56 ±0.96

-4.97 ±0.59

-7.48 ±0.82

-11.79 ±0.38

0.00 ±0.00

-5.93 ±0.31

-7.94 ±0.08

-9.47 ±0.32

-10.14 ±0.21

-11.35 ±0.14

-14.64 ±0.09

0.00 ±0.00

-12.03 ±0.43

-13.30 ±0.33

-14.94 ±0.27

-16.20 ±0.40

-18.11 ±0.18

-23.16 ±0.07

0.00 ±0.00

-5.51 ±0.26

-7.49 ±0.19

-8.19 ±0.13

-8.44 ±0.50

-10.58 ±0.49

-14.92 ±0.06

0

1

2

3

4

5

6

aniline ester ether para_hydroxylation phenol phenol_noOrthoHbond NO Heteroatoms RotatableBonds Ring Ar_N C_O C_O_noCOO allylic_oxid aryl_methyl bicyclic imide unbrch_alkane urea

Layer

0.00 ±0.00

-5.26 ±0.23

-9.46 ±0.38

-16.26 ±0.85

0.00 ±0.00

-3.64 ±0.51

-4.57 ±0.27

-5.59 ±0.32

-6.60 ±0.48

-7.68 ±0.42

-13.65 ±0.66

0.00 ±0.00

-8.28 ±0.12

-11.49 ±0.25

-14.24 ±0.33

-15.34 ±0.36

-17.03 ±0.14

-22.04 ±0.14

0.00 ±0.00

-2.79 ±0.10

-4.09 ±0.36

-4.68 ±0.09

-5.99 ±0.05

-8.04 ±0.05

-14.25 ±0.24

0.00 ±0.00

-1.54 ±0.75

-5.23 ±0.26

-3.00 ±0.42

-6.29 ±0.36

-3.70 ±1.04

-7.09 ±0.54

-4.20 ±0.88

-6.35 ±1.04

-10.59 ±0.63

0.00 ±0.00

-0.99 ±0.34

-3.31 ±0.49

-3.82 ±0.50

-3.89 ±0.56

-6.39 ±0.61

-10.18 ±0.48

0.00 ±0.00

-6.85 ±0.56

-10.45 ±0.51

-13.31 ±0.38

-14.33 ±0.79

-17.82 ±1.13

-24.72 ±0.44

0.00 ±0.00

-5.77 ±0.44

-9.40 ±0.24

-11.57 ±0.50

-12.93 ±0.51

-16.30 ±0.51

-22.10 ±0.75

0.00 ±0.00

-5.11 ±0.65

-8.45 ±0.05

-10.32 ±0.25

-10.91 ±0.16

-13.14 ±0.35

-18.59 ±0.12

0.00 ±0.00

-5.39 ±0.23

-7.31 ±0.19

-7.31 ±0.18

-5.57 ±0.14

-7.97 ±0.27

-14.47 ±0.25

0.00 ±0.00

-13.87 ±0.36

-14.53 ±0.29

-16.74 ±0.39

-18.09 ±0.24

-20.21 ±0.28

-25.97 ±0.31

0.00 ±0.00

-9.48 ±0.16

-13.41 ±0.32

0.00 ±0.00

-6.50 ±0.42

-7.62 ±0.32

-7.94 ±0.03

-8.19 ±0.48

-9.13 ±0.32

-13.21 ±0.23

0.00 ±0.00

-10.53 ±0.19

-7.42 ±0.23

-10.75 ±0.52

-8.63 ±0.27

-10.82 ±0.35

-8.58 ±0.17

-10.83 ±0.27

-8.50 ±0.10

-11.25 ±0.55

-14.84 ±0.31

0.00 ±0.00

-2.79 ±0.25

-4.43 ±0.50

-6.78 ±0.54

-8.46 ±0.21

-10.61 ±0.28

-15.77 ±0.24

0.00 ±0.00

-1.77 ±0.12

-4.41 ±0.07

-6.05 ±0.28

-6.63 ±0.10

-8.69 ±0.27

-13.36 ±0.09

0.00 ±0.00

-3.00 ±0.39

-4.76 ±0.15

-5.06 ±0.47

-5.70 ±0.26

-8.34 ±0.36

-12.41 ±0.08

0.00 ±0.00

-0.91 ±0.26

-5.79 ±0.88

-8.54 ±1.14

-11.03 ±0.76

-13.88 ±0.57

-17.34 ±0.95

0.00 ±0.00

-2.04 ±0.38

-3.60 ±0.66

-4.73 ±0.34

-5.56 ±0.48

-7.32 ±0.57

-12.16 ±1.06

0

1

2

3

4

5

6

20 10 0 10 20 30 40

AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen NHOH HAcceptors HDonors Al_OH Al_OH_noTert Ar_OH NH0 NH1 amide

0.00 ±0.00

-1.21 ±0.24

-2.75 ±0.13

0.00 ±0.00

-6.01 ±0.46

-5.23 ±0.21

-5.73 ±0.12

0.00 ±0.00

-8.73 ±0.09

-7.42 ±0.29

-8.50 ±0.21

0.00 ±0.00

-1.12 ±0.15

-1.47 ±0.10

-2.17 ±0.14

0.00 ±0.00

-0.20 ±0.11

-1.42 ±0.04

-2.08 ±0.05

0.00 ±0.00

-0.05 ±0.02

-0.08 ±0.02

-3.63 ±0.22

-0.11 ±0.04

0.00 ±0.00

-1.17 ±0.53

-2.74 ±0.37

-3.19 ±0.39

0.00 ±0.00

-2.35 ±0.12

-2.16 ±0.39

-2.61 ±0.42

0.00 ±0.00

-3.62 ±0.13

-4.62 ±0.03

-4.78 ±0.20

0.00 ±0.00

-1.21 ±0.08

-1.59 ±0.16

-2.17 ±0.08

0.00 ±0.00

-5.65 ±0.87

-7.15 ±0.24

-9.37 ±0.30

0.00 ±0.00

-1.16 ±0.41

-1.44 ±0.46

-2.09 ±0.28

0.00 ±0.00

10.86 ±0.36

9.38 ±1.09

7.66 ±0.79

0.00 ±0.00

-1.16 ±0.33

-1.41 ±0.12

-2.21 ±0.11

0.00 ±0.00

-3.50 ±0.04

-3.66 ±0.08

-3.33 ±0.19

0.00 ±0.00

-4.19 ±0.12

-4.09 ±0.06

-4.08 ±0.15

0.00 ±0.00

3.95 ±0.37

4.19 ±0.55

2.99 ±0.34

0.00 ±0.00

-1.27 ±0.23

-1.17 ±0.46

-1.63 ±0.23

0.00 ±0.00

-2.96 ±0.72

-2.35 ±0.36

-2.47 ±0.54

0.00 ±0.00

-1.54 ±0.27

2.41 ±0.53

1.91 ±0.49

0

1

2

3

0.00 ±0.00

-1.21 ±0.24

-2.75 ±0.13

-3.63 ±0.22

0.00 ±0.00

-6.01 ±0.46

-5.23 ±0.21

-5.73 ±0.12

aniline ester ether para_hydroxylation phenol phenol_noOrthoHbond NO Heteroatoms RotatableBonds Ring Ar_N C_O C_O_noCOO allylic_oxid aryl_methyl bicyclic imide unbrch_alkane urea

0.00 ±0.00

-0.59 ±0.40

1.46 ±0.34

1.37 ±0.48

0.00 ±0.00

3.06 ±0.45

2.71 ±0.49

1.84 ±0.35

0.00 ±0.00

-0.48 ±0.46

-2.89 ±0.36

-4.42 ±0.15

0.00 ±0.00

0.88 ±0.10

0.71 ±0.27

-0.83 ±0.23

0.00 ±0.00

3.52 ±0.59

3.57 ±0.22

2.74 ±0.33

0.00 ±0.00

3.09 ±0.56

3.32 ±0.36

2.30 ±0.38

0.00 ±0.00

17.51 ±0.87

14.47 ±1.05

8.42 ±0.59

0.00 ±0.00

15.99 ±0.82

14.00 ±0.49

10.02 ±0.57

0.00 ±0.00

-4.99 ±1.01

-4.92 ±0.39

-6.00 ±0.32

0.00 ±0.00

-1.08 ±0.16

-0.37 ±0.09

-0.82 ±0.19

0.00 ±0.00

-0.51 ±0.07

-0.63 ±0.11

-0.85 ±0.08

0.00 ±0.00

-3.63 ±0.10

-4.72 ±0.40

-3.76 ±0.22

0.00 ±0.00

-2.48 ±0.20

-3.53 ±0.52

-2.65 ±0.21

0.00 ±0.00

-6.06 ±0.25

-5.30 ±0.25

-5.23 ±0.27

0.00 ±0.00

6.07 ±0.12

3.83 ±0.16

1.62 ±0.07

0.00 ±0.00

-4.62 ±0.17

-0.47 ±0.12

-1.56 ±0.18

0.00 ±0.00

-0.10 ±0.30

1.64 ±0.20

0.68 ±0.32

0.00 ±0.00

-4.46 ±0.23

-1.30 ±0.26

-1.03 ±0.41

0.00 ±0.00

-0.67 ±0.21

-1.12 ±0.49

-1.41 ±0.39

0

1

2

3

0.00 ±0.00

-0.59 ±0.40

1.46 ±0.34

1.37 ±0.48

0.00 ±0.00

3.06 ±0.45

2.71 ±0.49

1.84 ±0.35

0.00 ±0.00

-0.48 ±0.46

-2.89 ±0.36

-4.42 ±0.15

Layer

10

0

10

20

30

AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen NHOH HAcceptors HDonors Al_OH Al_OH_noTert Ar_OH NH0 NH1 amide

0.00 ±0.00

-8.73 ±0.09

-7.42 ±0.29

-8.50 ±0.21

0.00 ±0.00

-1.12 ±0.15

-1.47 ±0.10

-2.17 ±0.14

0.00 ±0.00

-0.20 ±0.11

-1.42 ±0.04

-2.08 ±0.05

0.00 ±0.00

-0.05 ±0.02

-0.08 ±0.02

-0.11 ±0.04

0.00 ±0.00

-1.17 ±0.53

-2.74 ±0.37

-3.19 ±0.39

0.00 ±0.00

-2.35 ±0.12

-2.16 ±0.39

-2.61 ±0.42

0.00 ±0.00

-3.62 ±0.13

-4.62 ±0.03

-4.78 ±0.20

0.00 ±0.00

-1.21 ±0.08

-1.59 ±0.16

-2.17 ±0.08

0.00 ±0.00

-5.65 ±0.87

-7.15 ±0.24

-9.37 ±0.30

0.00 ±0.00

-1.16 ±0.41

-1.44 ±0.46

-2.09 ±0.28

0.00 ±0.00

10.86 ±0.36

9.38 ±1.09

7.66 ±0.79

0.00 ±0.00

-1.16 ±0.33

-1.41 ±0.12

-2.21 ±0.11

0.00 ±0.00

-3.50 ±0.04

-3.66 ±0.08

-3.33 ±0.19

0.00 ±0.00

-4.19 ±0.12

-4.09 ±0.06

-4.08 ±0.15

0.00 ±0.00

3.95 ±0.37

4.19 ±0.55

2.99 ±0.34

0.00 ±0.00

-1.27 ±0.23

-1.17 ±0.46

-1.63 ±0.23

0.00 ±0.00

-2.96 ±0.72

-2.35 ±0.36

-2.47 ±0.54

0.00 ±0.00

-1.54 ±0.27

2.41 ±0.53

1.91 ±0.49

0

1

2

3

0.00 ±0.00

-1.21 ±0.24

-2.75 ±0.13

-3.63 ±0.22

0.00 ±0.00

-6.01 ±0.46

-5.23 ±0.21

-5.73 ±0.12

0.00 ±0.00

-8.73 ±0.09

-7.42 ±0.29

-8.50 ±0.21

0.00 ±0.00

-1.12 ±0.15

-1.47 ±0.10

-2.17 ±0.14

0.00 ±0.00

-0.20 ±0.11

-1.42 ±0.04

-2.08 ±0.05

0.00 ±0.00

-0.05 ±0.02

-0.08 ±0.02

-0.11 ±0.04

0.00 ±0.00

-1.17 ±0.53

-2.74 ±0.37

-3.19 ±0.39

0.00 ±0.00

-2.35 ±0.12

-2.16 ±0.39

-2.61 ±0.42

0.00 ±0.00

-3.62 ±0.13

-4.62 ±0.03

-4.78 ±0.20

aniline ester ether para_hydroxylation phenol phenol_noOrthoHbond NO Heteroatoms RotatableBonds Ring Ar_N C_O C_O_noCOO allylic_oxid aryl_methyl bicyclic imide unbrch_alkane urea

0.00 ±0.00

0.88 ±0.10

0.71 ±0.27

-0.83 ±0.23

0.00 ±0.00

3.52 ±0.59

3.57 ±0.22

2.74 ±0.33

0.00 ±0.00

3.09 ±0.56

3.32 ±0.36

2.30 ±0.38

0.00 ±0.00

17.51 ±0.87

14.47 ±1.05

8.42 ±0.59

0.00 ±0.00

15.99 ±0.82

14.00 ±0.49

10.02 ±0.57

0.00 ±0.00

-4.99 ±1.01

-4.92 ±0.39

-6.00 ±0.32

0.00 ±0.00

-1.08 ±0.16

-0.37 ±0.09

-0.82 ±0.19

0.00 ±0.00

-0.51 ±0.07

-0.63 ±0.11

-0.85 ±0.08

0.00 ±0.00

-3.63 ±0.10

-4.72 ±0.40

-3.76 ±0.22

0.00 ±0.00

-2.48 ±0.20

-3.53 ±0.52

-2.65 ±0.21

0.00 ±0.00

-6.06 ±0.25

-5.30 ±0.25

-5.23 ±0.27

0.00 ±0.00

6.07 ±0.12

3.83 ±0.16

1.62 ±0.07

0.00 ±0.00

-4.62 ±0.17

-0.47 ±0.12

-1.56 ±0.18

0.00 ±0.00

-0.10 ±0.30

1.64 ±0.20

0.68 ±0.32

0.00 ±0.00

-4.46 ±0.23

-1.30 ±0.26

-1.03 ±0.41

0.00 ±0.00

-0.67 ±0.21

-1.12 ±0.49

-1.41 ±0.39

0

1

2

3

0.00 ±0.00

-0.59 ±0.40

1.46 ±0.34

1.37 ±0.48

0.00 ±0.00

3.06 ±0.45

2.71 ±0.49

1.84 ±0.35

Layer

10

0

10

20

30

AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen NHOH HAcceptors HDonors Al_OH Al_OH_noTert Ar_OH NH0 NH1 amide

0.00 ±0.00

-1.21 ±0.08

-1.59 ±0.16

-2.17 ±0.08

0.00 ±0.00

-5.65 ±0.87

-7.15 ±0.24

-9.37 ±0.30

0.00 ±0.00

-1.16 ±0.41

-1.44 ±0.46

-2.09 ±0.28

0.00 ±0.00

10.86 ±0.36

9.38 ±1.09

7.66 ±0.79

0.00 ±0.00

-1.16 ±0.33

-1.41 ±0.12

-2.21 ±0.11

0.00 ±0.00

-3.50 ±0.04

-3.66 ±0.08

-3.33 ±0.19

0.00 ±0.00

-4.19 ±0.12

-4.09 ±0.06

-4.08 ±0.15

0.00 ±0.00

3.95 ±0.37

4.19 ±0.55

2.99 ±0.34

0.00 ±0.00

-1.27 ±0.23

-1.17 ±0.46

-1.63 ±0.23

0.00 ±0.00

-2.96 ±0.72

-2.35 ±0.36

-2.47 ±0.54

0.00 ±0.00

-1.54 ±0.27

2.41 ±0.53

1.91 ±0.49

0

1

2

3

aniline ester ether para_hydroxylation phenol phenol_noOrthoHbond NO Heteroatoms RotatableBonds Ring Ar_N C_O C_O_noCOO allylic_oxid aryl_methyl bicyclic imide unbrch_alkane urea

0.00 ±0.00

-0.48 ±0.46

-2.89 ±0.36

-4.42 ±0.15

0.00 ±0.00

0.88 ±0.10

0.71 ±0.27

-0.83 ±0.23

0.00 ±0.00

3.52 ±0.59

3.57 ±0.22

2.74 ±0.33

0.00 ±0.00

3.09 ±0.56

3.32 ±0.36

2.30 ±0.38

0.00 ±0.00

17.51 ±0.87

14.47 ±1.05

8.42 ±0.59

0.00 ±0.00

15.99 ±0.82

14.00 ±0.49

10.02 ±0.57

0.00 ±0.00

-4.99 ±1.01

-4.92 ±0.39

-6.00 ±0.32

0.00 ±0.00

-1.08 ±0.16

-0.37 ±0.09

-0.82 ±0.19

0.00 ±0.00

-0.51 ±0.07

-0.63 ±0.11

-0.85 ±0.08

0.00 ±0.00

-3.63 ±0.10

-4.72 ±0.40

-3.76 ±0.22

0.00 ±0.00

-2.48 ±0.20

-3.53 ±0.52

-2.65 ±0.21

0.00 ±0.00

-6.06 ±0.25

-5.30 ±0.25

-5.23 ±0.27

0.00 ±0.00

6.07 ±0.12

3.83 ±0.16

1.62 ±0.07

0.00 ±0.00

-4.62 ±0.17

-0.47 ±0.12

-1.56 ±0.18

0.00 ±0.00

-0.10 ±0.30

1.64 ±0.20

0.68 ±0.32

0.00 ±0.00

-4.46 ±0.23

-1.30 ±0.26

-1.03 ±0.41

0.00 ±0.00

-0.67 ±0.21

-1.12 ±0.49

-1.41 ±0.39

0

1

2

3

Layer

10

0

10

20

30

(n) Chemberta-2 (RI)

0.00 ±0.00

0.02 ±0.13

-0.70 ±0.21

-0.91 ±0.06

0.00 ±0.32

-0.87 ±0.57

-2.68 ±0.41

-0.67 ±0.13

-0.87 ±0.37

-0.38 ±0.11

0.13 ±0.38

0.43 ±0.26

2.63 ±0.46

0.00 ±0.00

0.07 ±0.38

0.29 ±0.10

-0.39 ±0.06

-0.48 ±0.22

-1.98 ±0.05

-0.63 ±0.23

0.86 ±0.10

1.26 ±0.20

1.24 ±0.89

1.45 ±0.42

0.36 ±0.21

0.81 ±0.30

0.00 ±0.00

0.12 ±0.27

0.49 ±0.25

-0.20 ±0.34

0.75 ±0.29

-3.65 ±0.04

-1.11 ±0.20

-0.52 ±0.05

-0.77 ±0.38

-1.54 ±0.16

-0.19 ±0.22

0.56 ±0.31

1.21 ±0.33

0.00 ±0.00

0.11 ±0.22

0.17 ±0.00

-0.73 ±0.39

-1.56 ±0.16

-1.68 ±0.39

-1.57 ±0.06

0.25 ±0.14

-0.60 ±0.46

-1.06 ±0.61

-1.71 ±0.77

0.11 ±0.34

1.85 ±1.08

0.00 ±0.00

-0.21 ±0.25

0.47 ±0.30

-1.19 ±0.19

-0.39 ±0.34

-1.70 ±0.26

0.90 ±0.46

0.48 ±1.04

0.82 ±1.71

2.01 ±1.27

5.62 ±1.43

8.92 ±0.62

10.35 ±1.08

0.00 ±0.00

-0.28 ±0.16

0.42 ±0.49

-1.02 ±0.20

-0.40 ±0.30

-1.44 ±0.36

0.98 ±0.17

0.90 ±0.46

0.76 ±1.14

2.85 ±1.40

6.26 ±1.16

7.77 ±1.20

9.93 ±0.52

0.00 ±0.00

-0.04 ±0.15

-0.65 ±0.21

-1.75 ±0.18

-0.63 ±0.58

-0.54 ±0.48

1.49 ±0.77

0.28 ±0.23

2.08 ±1.26

0.81 ±0.77

2.07 ±1.30

0.09 ±0.37

3.05 ±1.12

0.00 ±0.00

0.18 ±0.18

-0.68 ±0.31

-1.

0

1

2

3

4

5

6

7

8

9

10

11

12

o Chembe a 3 PT

-2.09 ±0.17

-0.33 ±0.27 -0.18 ±0.04

(l) Chemberta-2 (RI) 10

Molecular Substructure

Molecular Substructure

0.00 ±0.00 0.00 ±0.00

aniline ester ether para_hydroxylation phenol phenol_noOrthoHbond NO Heteroatoms RotatableBonds Ring Ar_N C_O C_O_noCOO allylic_oxid aryl_methyl bicyclic imide unbrch_alkane urea

-1.53 ±0.02

0.15 ±0.18 -0.07 ±0.10

(k) Chemberta-2-10M (PT) AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen NHOH HAcceptors HDonors Al_OH Al_OH_noTert Ar_OH NH0 NH1 amide

-0.84 ±0.10

0.79 ±0.14

0.24 ±0.24 0.11 ±0.07

0.09 ±0.20

(j) Chemberta-2 (RI)

Molecular Substructure

Molecular Substructure

0.00 ±0.00 0.00 ±0.00

aniline ester ether para_hydroxylation phenol phenol_noOrthoHbond NO Heteroatoms RotatableBonds Ring Ar_N C_O C_O_noCOO allylic_oxid aryl_methyl bicyclic imide unbrch_alkane urea

-1.14 ±0.15

0.24 ±0.13 -0.11 ±0.09

(i) Chemberta-2-5M (PT) AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen NHOH HAcceptors HDonors Al_OH Al_OH_noTert Ar_OH NH0 NH1 amide

-0.99 ±0.06

0.33 ±0.22

0.20 ±0.21 0.12 ±0.22

0.00 ±0.00

(h) Chemberta-base (RI)

0.00 ±0.00

Molecular Substructure

Molecular Substructure

0.00 ±0.00

-0.86 ±0.16

0.32 ±0.07 -0.06 ±0.09

(g) Chemberta-base (PT) AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen NHOH HAcceptors HDonors Al_OH Al_OH_noTert Ar_OH NH0 NH1 amide

-3.37 ±0.05

0.31 ±0.13

0.00 ±0.00 0.00 ±0.00

(f) Chemberta (RI)

Molecular Substructure

Molecular Substructure

0.00 ±0.00 0.00 ±0.00

-0.48 ±0.07

-0.10 ±0.38 -0.28 ±0.22

(e) Chemberta (PT) AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen NHOH HAcceptors HDonors Al_OH Al_OH_noTert Ar_OH NH0 NH1 amide

-0.45 ±0.14

0.00 ±0.00

(d) Roberta-zinc-480m (RI)

0.00 ±0.00

Molecular Substructure

Molecular Substructure

0.00 ±0.00

-0.32 ±0.07

0.00 ±0.00 0.00 ±0.00

(c) Roberta-zinc-480m (PT) AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen NHOH HAcceptors HDonors Al_OH Al_OH_noTert Ar_OH NH0 NH1 amide

Layer

0.00 ±0.00

(b) Molformer(RI)

0.00 ±0.00

Molecular Substructure

Molecular Substructure

(a) Molformer (PT) AliphaticCarbocycles AliphaticHeterocycles AliphaticRings AromaticCarbocycles AromaticHeterocycles AromaticRings SaturatedCarbocycles SaturatedHeterocycles SaturatedRings benzene halogen NHOH HAcceptors HDonors Al_OH Al_OH_noTert Ar_OH NH0 NH1 amide

aniline ester ether para_hydroxylation phenol phenol_noOrthoHbond NO Heteroatoms RotatableBonds Ring Ar_N C_O C_O_noCOO allylic_oxid aryl_methyl bicyclic imide unbrch_alkane urea

p Chembe a 3 R

F gure 11 Effect of fine-tun ng on so ub ty (ESOL) on mo ecu ar substructure encod ng n CLMs We repor he ayer-rw se d fference n prob ng performance (macro-averaged F1 n % ) af er fine- un ng across 39 asks (cf Tab e 7) F gures 10 10 and 10n correspond o he same mode arch ec ure PT deno es o he pre- ra ned mode s ( ef ) wh e RI refers he same mode s bu w h random y n a zed we gh s (r gh ) Improvemen s n prob ng 36 performance af er fine- un ng are shown n red wh e degrada ons are nd ca ed n b ue

Model type

1.0

PT

chemberta-2-5M furan

further PT

chemberta furan

molformer furan

0.8 0.6 0.4 0.2

Macro F1

1.0

0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00

chemberta-2-5M thiazole

chemberta thiazole

molformer thiazole

0.8 0.6 0.4 0.2 1.0

0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00

chemberta-2-5M thiophene

chemberta thiophene

molformer thiophene

0.8 0.6 0.4 0.2

0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00

Relative Layer Depth Figure 12: Probing performance of pre-trained models (PT) and those further pre-trained (further PT) on data containing furan (row 1), thiazole (row 2) and thiophene (row 3) on the respective substructure. Across all models, we observe an improvement in probing performance on the specific substructure after additional pre-training with data containing the corresponding substructure.

Model

RMSE

chemberta-2-5M chemberta-2-10M chemberta-2-77M

0.646 (-0.018) 0.699 (+0.108) 0.635 (+0.003)

Table 14: Fine-tuning performance (RMSE, lower is better) of chemberta-2-5M, chemberta-2-10M and chemberta-2-77M further pre-trained on data containing halogens. green indicates improvement in RMSE (lower is better) while red denotes increased RMSE. Gray indicates negligible changes (<0.01) in RMSE.

37

halogen (RI)

1.0

halogen (PT)

0.9 0.8

Macro F1

0.7 0.6 0.5 0.4 chemberta chemberta-base chemberta-2-10M

0.3 0.2

0.0

0.2

0.4

0.6

chemberta-2-5M chemberta-2-77M chemberta-3

0.8

1.0

0.0

molformer roberta-zinc-480m

0.2

0.4

0.6

0.8

1.0

Relative Layer Depth Figure 13: Layerwise probing performance (macro F1 score) on halogens for randomly initialized (RI) and pre-trained (PT) models. We observe that all models except for chemberta-2-5M and chemberta-2-10M unlearn halogens in the middle and upper layers.

Macro F1

Model type

1.0 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2

PT

chemberta-2-5M halogen

further PT

chemberta-2-10M halogen

chemberta-2-77M halogen

0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00 0.00 0.25 0.50 0.75 1.00

Relative Layer Depth

Macro F1

Figure 14: Probing performance of pre-trained models (PT) chemberta-2-5M, chemberta-2-10M and chemberta-2-77M and their counterparts further pre-trained (further PT) on data containing halogens on the halogen substructures.

1.0 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2

phenol

phenol

chemberta2_5M chemberta2_10M

0.0

0.2

0.4

0.6

0.8

1.0

chemberta2_5M_phenol chemberta2_10M_phenol

0.0

0.2

0.4

0.6

0.8

1.0

Relative Layer Depth Figure 15: Probing performance of pre-trained models chemberta-2-5M and chemberta-2-10M (left) and the same models further pre-trained on data containing phenols (right) on phenols. We observe that while further pre-training on data containing phenol leads to better probing performance on the corresponding substructure, the performance gap between the two models does not close. We hypothesize that longer pre-training might narrow this gap.

38

Model

RMSE

chemberta-2-5M chemberta-2-10M

0.622 (-0.042) 0.618 (+0.027)

Macro F1

Table 15: Fine-tuning performance (RMSE, lower is better) of chemberta-2-5M and chemberta-2-10M further pre-trained on data containing phenol. green indicates improvement in RMSE (lower is better) while red denotes increased RMSE.

1.0 0.9 0.8 0.7 0.6 0.5 0.4 0.3

rings substructures

0.00

0.25

0.50

0.75

other substructures

1.00 0.00

0.25

0.50

0.75

1.00

Relative Layer Depth chemberta-3 (RI) chemberta-3 (PT)

chemberta3-guacamol (PT) chemberta3-zinc (PT)

Figure 16: Effect of further pre-training chemberta-3 on different data. chemberta-3 (PT) denotes the pretrained model, chemberta-3-zinc (PT) is the model further pre-trained on a 100k subset of Guacamol, while chemberta-3-guacamol (PT) has been further pre-trained on a subset of ZINC-100M. chemberta-3 (RI) is the randomly initialized chemberta-3.

Training set Guacamol Zinc

100

Relative Frequency

80

60

40

20

0 0

20

40

60

80

Substructure id (alphabetically sorted)

100

Figure 17: Comparison of relative frequency distributions of molecular substructures in the training splits of the subsampled ZINC and Guacamol datasets used for further pre-training chemberta-3.

39

Model

dim

MFP

RDFP

ATFP

TTFP

Avg

LR

300 512 1024 2048 Avg

0.84 0.84 0.85 1.09 0.91

1.02 1.00 0.98 1.20 1.05

0.88 0.91 0.95 1.27 1.00

0.90 0.87 0.95 1.05 0.94

0.91 0.91 0.93 1.15 -

SVM

300 512 1024 2048 Avg

0.72 0.71 0.68 0.67 0.70

0.94 0.89∗ 0.75 0.72 0.83

0.75∗ 0.73∗ 0.68 0.64∗ 0.70

0.77 0.73 0.70 0.69 0.72

0.80 0.77 0.70 0.68 -

XGB

300 512 1024 2048 Avg

0.76 0.73 0.72 0.72 0.73

0.97 0.91 0.83 0.79 0.88

0.80 0.76 0.74 0.69 0.75

0.80 0.80 0.74 0.73 0.77

0.83 0.80 0.76 0.73 -

Avg

-

0.78

0.92

0.82

0.81

-

Table 16: Test set performance [RMSE (↓)] on lipophilicity. We evaluate four different methods for generating fingerprints (MFP, RDFP, ATFP, TTFP) and evaluate dimension sizes of 300, 512, 1024, and 2048. For the SVM, c=4 worked best except for the ones marked with ∗ (c=16). Interestingly, we find that both XGB and SVM perform better for larger dimensions, while LR performs better for lower dimensions.

Model

dim

MFP

RDFP

ATFP

TTFP

Avg

LR

300 512 1024 2048 Avg

1.81 2.06 4.74 2.53 2.79

1.63 1.93 6.87 2.61 3.26

1.44 2.11 367.63 5.84 94.26

1.39 1.87 11.18 4.03 4.62

1.57 1.99 97.61 3.75 -

SVM

300 512 1024 2048 Avg

1.19c=256 1.06c=256 1.06c=64 1.00c=256 1.08

1.02c=64 0.91c=256 0.90c=1024 0.88c=1024 0.93

0.93c=16 1.00c=4 0.92c=4 0.83c=16 0.92

0.99c=16 1.14c=4 1.20c=16 1.11c=16 1.11

1.03 1.03 1.02 0.96 -

XGB

300 512 1024 2048 Avg

1.23 1.12 1.13 1.17 1.16

1.16 1.06 1.01 0.86 1.02

1.01 1.01 1.04 0.93 1.00

1.08 1.14 1.16 1.17 1.13

1.12 1.08 1.09 1.03 -

Avg

-

1.66

1.74

32.06

2.29

-

Table 17: Test set performance [RMSE (↓)] on ESOL (solubility prediction). We evaluate four different methods for generating fingerprints (MFP, RDFP, ATFP, TTFP) and evaluate dimension sizes of 300, 512, 1024, and 2048. For the SVM, we report the best-performing c for each configuration. We find that both XGB and SVM maintain a robust performance around 1.00 across all dimensions, while LR performance varies substantially, especially for dim = 1, 024.

40

Model

reasoning

basic

SVM(c=16) + ATFP molformer

expl

hint

both

61.13 2.66 2.80 246.42 43.04 376.55

42.06 2.82 2.53 19.95 27.67 337.18

25.93 2.84 2.94 98.97 24.60 15.45

0.64 0.57

Llama-3.2-3B-Instruct gpt-oss-20b4,096 gpt-oss-20b4,096 gpt-oss-20b4,096 gpt-oss-20b8,192 gpt-oss-20b16,384

low medium high high high

188.24 2.89 3.11 237,981.69 139.36 61.84

Table 18: Test set performance [RMSE (↓)] on lipophilicity. We find that the gpt-oss-20b4,096 model using the medium reasoning level together with hints performs best out of all LLMs and that providing explanations as well as hints can improve the performance. Nonetheless, all LLMs are substantially outperformed by the SVM using fingerprints and the molformer model.

Model

reasoning

basic

SVM(c=16) + ATFP molformer Llama-3.2-3B-Instruct gpt-oss-20b4,096 gpt-oss-20b4,096 gpt-oss-20b4,096 gpt-oss-20b8,192 gpt-oss-20b16,384

expl

hint

both

477.07 9.51 14.11 95.61 42.34 19.39

940,720,868.84 8.96 23.31 36.32 193.03 60.72

0.83 0.59 low medium high high high

233.11 9.27 48.95 193.50 23.53 32.94

213.79 31.97 51.68 41.91 44.42 83.57

Table 19: Test set performance [RMSE (↓)] on esol (solubility prediction). We find that the gpt-oss-20b4,096 model using the low reasoning level and both (hints and explanations) performs best. Interestingly, we find that especially for larger contexts, providing explanations or hints can deteriorate the performance. Again, all LLMs are substantially outperformed by the SVM using fingerprints and the molformer model.

41

We need to predict logD value (probably logD at physiological pH, e.g., logD7.4). This is a computational estimate. We are given a SMILES: C#Cc1cccc(Nc2nc3cc(C(=O)O)ccc3c3cncnc23)c1 Let’s parse: It’s a heteroaromatic system fused? Actually: C#C-? Let’s analyze. The base skeleton: phenyl ring (c1cccc...) with a -C#C substituent at position 1? The SMILES: C#C c1 c c c c (N c2 n c3 c c (C(=O)O) c c c3 c3 c n c n c2 3) c1. So there is a phenyl ring with a -C#C group (propyne). At the para position (maybe 4) there is an anilide linking to a bicyclic heteroaromatic (imidazo[...])? Actually c2nc3cc(C(=O)O)ccc3c3cncnc23. That seems to be a fused bicyclic heteroaromatic containing two rings: a benzodiazine? It appears to be a quinazoline ring (a benzene fused with a diazine). Quinazoline: fused benzene and pyrimidine. So we have a 4-? It attaches via an aniline NH to the quinazoline ring. Also there is a carboxylic acid group (C(=O)O) attached to the quinazoline ring. The entire molecule contains one carboxylic acid (pKa 4.5). At physiological pH, it will be deprotonated (anionic). This will significantly reduce logD. Also there is one NH (aniline). There are hetero nitrogens: 4? In quinazoline ring: two ring nitrogens (positions 1 and 3). Additional ring may have two more? Wait: c3cncnc23: That is a fused pyrimidine? Actually c3 c n c n c 2 3. So ring 3 has two nitrogens. So total hetero atoms: 4 ring nitrogens, 1 amine N, 1 carboxylate O’s (2). So 5 nitrogens, 3 oxygens. So high hetero content. Also there’s a nitrile (C#C) group: actually it’s an alkyne (C#C) group attached to phenyl; no hetero. So logP (neutral) might be moderate-high due to aromatic rings (∼3?). But logD7.4 will be lower due to deprotonated carboxylate (negative). Many hetero nitrogens may also be basic? In quinazoline, the ring nitrogens are not protonated at physiological? They are aromatic; they can act as weak bases (pKa 2-3). So mostly neutral. The aniline NH is weakly basic; likely neutral. So overall, the logD7.4 may be around 1.5? Actually carboxylate reduces lipophilicity drastically. For example, carboxylate-containing molecules like ibuprofen logD7.4 ∼3.5? Wait ibuprofen is acid (pKa 4.5). Ibuprofen has logP 3.5, logD7.4 ∼3.5? Actually at neutral, logD decreases a bit; but still high due to hydrophobic. But our molecule has 3 aromatic rings (two fused hetero aromatic + phenyl). That gives high aromatic surface. Carboxylate reduces logP by ∼2 units maybe. For a 3-ring acid, logD7.4 maybe ∼2.5-3.0. Example: diclofenac (two phenyl rings + acid) has logP 4.5, logD7.4 4.0? Actually diclofenac logP 4.5, logD7.4 4.0? Not sure. Another acid with hetero: warfarin logP 3.1, logD7.4 ∼2.5. So our compound may be similar to warfarin. Thus I’d guess logD7.4 ∼2.0-2.5. Let’s pick 2.3. But maybe due to hetero nitrogens, more hydrophilic, so logD lower ∼1.8. Let’s compute approximate cLogP using fragments: Each phenyl ring ∼2.3; hetero ring reduces; carboxylate reduces ∼1.5. Hard. I’ll answer: logD ≈ 1.9. Given the question: P̈redict the lipophilicity the following molecule.S̈o I should provide a single numeric value. I’ll answer 2.0 (±0.5). Provide explanation. Let’s produce final.assistantfinal**Predicted logD (pH ≈ 7.4)** ≈ **2.0** The molecule contains three aromatic rings (one phenyl and one fused benzodiazine) that give it a relatively high baseline lipophilicity, but the presence of a carboxylic acid (which will be largely anionic at physiological pH) and several hetero nitrogen atoms dramatically reduces its distribution into the octanol phase. A logD value around 2.0 is consistent with similar tricyclic heteroaromatic acids (e.g., warfarin, diclofenac-like scaffolds) when measured at neutral to slightly basic conditions.

Figure 18: Response of gpt-oss-20b4,096 with medium reasoning level and preprocessed functional groups as hints. While the model seemingly “reasons” about possible implications of different molecular subgroups, it predicts a logD value of 2.0 while the true value lies at -1.067.

42

Record · ID 332549 · SHA-256 3a106b02c0443b7c
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.