Conceptio › Archive › NCBI PubMed Central
NCBI PubMed Centralopen access

Leveraging Transfer Learning for Predicting Protein-Small-Molecule Interaction Predictions.

Wang J et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
machine learning systems

Leveraging Transfer Learning for Predicting Protein–Small-Molecule Interaction Predictions - PMC Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice J Chem Inf Model . Author manuscript; available in PMC: 2026 Apr 19. Published in final edited form as: J Chem Inf Model. 2025 Mar 24;65(7):3262–3269. doi: 10.1021/acs.jcim.4c02256 Search in PMC Search in PubMed View in NLM Catalog Add to search Leveraging Transfer Learning for Predicting Protein–Small-Molecule Interaction Predictions Jian Wang Jian Wang 1 Department of Neuroscience and Experimental Therapeutics, Penn State College of Medicine, Hershey, Pennsylvania 17033, United States Find articles by Jian Wang 1 , Nikolay V Dokholyan Nikolay V Dokholyan 2 Department of Neuroscience and Experimental Therapeutics, Penn State College of Medicine, Hershey, Pennsylvania 17033, United States; Department of Chemistry and Department of Biomedical Engineering, Pennsylvania State University, University Park, Pennsylvania 16802, United States Find articles by Nikolay V Dokholyan 2 Author information Article notes Copyright and License information 1 Department of Neuroscience and Experimental Therapeutics, Penn State College of Medicine, Hershey, Pennsylvania 17033, United States 2 Department of Neuroscience and Experimental Therapeutics, Penn State College of Medicine, Hershey, Pennsylvania 17033, United States; Department of Chemistry and Department of Biomedical Engineering, Pennsylvania State University, University Park, Pennsylvania 16802, United States Author Contributions Jian Wang contributed to the conceptualization, methodology, model development, data analysis, writing of the original draft, and visualization of the study. Nikolay V. Dokholyan provided supervision, resources, writing review and editing, project administration, and funding acquisition. ✉ Corresponding Author: Nikolay V. Dokholyan – Department of Neuroscience and Experimental Therapeutics, Penn State College of Medicine, Hershey, Pennsylvania 17033, United States; Department of Chemistry and Department of Biomedical Engineering, Pennsylvania State University, University Park, Pennsylvania 16802, United States; [email protected] Issue date 2025 Apr 14. PMC Copyright notice PMCID: PMC13091560  NIHMSID: NIHMS2157797  PMID: 40127309 The publisher's version of this article is available at J Chem Inf Model Previous version available: This article is based on a previously available preprint posted on bioRxiv on October 14, 2024: " Leveraging Transfer Learning for Predicting Protein-Small Molecule Interactions ". Abstract A complex web of intermolecular interactions defines and regulates biological processes. Understanding this web has been particularly challenging because of the sheer number of actors in biological systems: ~10 4 proteins in a typical human cell offer plausible 10 8 interactions. This number grows rapidly if we consider metabolites, drugs, nutrients, and other biological molecules. The relative strength of interactions also critically affects these biological processes. However, the small and often incomplete data sets (10 3 –10 4 protein–ligand interactions) traditionally used for binding affinity predictions limit the ability to capture the full complexity of these interactions. To overcome this challenge, we developed Yuel 2, a novel neural network-based approach that leverages transfer learning to address the limitations of small data sets. Yuel 2 is pretrained on a large-scale data set to learn intricate structural features and then fine-tuned on specialized data sets like PDBbind to enhance the predictive accuracy and robustness. We show that Yuel 2 predicts multiple binding affinity metrics, K d , K i , and IC 50 , between proteins and small molecules, offering a comprehensive representation of molecular interactions crucial for drug design and development. Graphical Abstract INTRODUCTION The life of cells and organisms is defined by interactions between molecules that reside inside cells and those that are introduced, such as drugs, metabolites, nutrients, and toxins. This complex web of interactions results in spatiotemporal organization of the life processes. Understanding the basic principles underlying these interactions and being able to predict them have profound implications for biology and medicine. Yet, the fundamental physical uncertainties and the large space of interacting agents limit our abilities to fully comprehend this complex web of interactions that define biological life. A typical human cell contains approximately 10 4 different proteins, each of which can engage in numerous interactions with other proteins, as well as with various small molecules. These interactions are essential for the proper functioning of biological systems, influencing everything from cellular signaling to metabolic pathways. However, the potential number of interactions within a single cell, potentially reaching 10 8 , poses a significant challenge to our understanding and predictive capabilities. This complexity is further compounded by the fact that these interactions vary in strength and specificity, critically affecting the stability and dynamics of biological networks. Predicting these interactions and their relative affinities remains one of the most daunting tasks in computational biology and drug discovery. Traditional approaches to predicting protein–small-molecule interactions have relied heavily on data sets with experimental protein–small-molecule binding affinities, such as PDBbind 1 and BindingDB. 2 Despite the advances in deep learning methods and their demonstrated higher accuracy across various data sets, traditional docking software such as AutoDock Vina, 3 Glide SP, 4 and MedusaDock 5 , 6 , 7 remain prevalent in virtual screening and target identification. 8 The persistent use of conventional docking methods rather than machine learning methods can be attributed to the limited size of existing data sets used to train machine learning models. These data sets typically comprise 10 3 to 10 4 protein–ligand interactions. For example, the PDBbind 1 -refined data set comprises 5316 protein–ligand pairs, and the PDBbind 1 general data set comprises 14,127 protein–ligand pairs. The small size and incomplete nature of these data sets restrict the ability of computational models to generalize across the vast interaction space. As a result, predictions made by these models may lack the accuracy and robustness needed for practical applications in drug design and other areas of biomedical research. In contrast, traditional docking software relies on physical force fields, offering a robust theoretical approach for predicting binding interactions. However, the accuracy of these traditional scoring functions is often constrained by their reliance on fundamental physical approximations of interatomic interactions. In addition to the challenge of limited data set size in training machine learning models, both traditional docking methods and neural networks also face difficulties in accurately predicting complex phenomena like activity cliffs, 9 , 10 where small molecular changes can drastically impact binding affinities. Encountered in complex biochemical systems, these cliffs result in the inaccurate prediction of binding affinities and limit the applicability of the existing methods. Overcoming these challenges is essential for ensuring that a prediction method navigates the intricate landscape of activity cliffs to provide comprehensive and reliable insights into the dynamic interplay between proteins and bioactive compounds. Finally, while many existing methods focus on predicting a single metric, such as the dissociation constant ( K d ), a comprehensive approach that also predicts the inhibition constant ( K i ) and the half-maximal inhibitory concentration (IC 50 ) is significantly more practical. Each of these metrics provides unique insights into different aspects of interaction dynamics. K d measures the affinity of the ligand for the protein at equilibrium, K i indicates the potency of an inhibitor, and IC 50 quantifies how much of a substance is needed to inhibit a biological process by one-half. Predicting all kinds of metrics using only one method would offer a more detailed and accurate representation of the protein–ligand interaction, which is essential for drug design and development. Existing methods often fall short by addressing only one of these critical metrics or even mixing them, thus limiting their applicability and precision. Here, we propose Yuel 2, which utilizes transfer learning to address the limitations posed by small and incomplete data sets. We create a large-scale data set by systematically docking multiple proteins with multiple ligands, resulting in a comprehensive data set containing all possible protein–ligand pairs. This approach generates an expansive data set that significantly surpasses the size of current data sets like PDBbind, which may only include a limited subset of possible interactions. By training Yuel 2 on this large-scale data set, we enable the model to learn the intricate structural features of both proteins and ligands at a high level. Once the neural network has been pretrained on this vast data set to predict docking scores, we fine-tune the model on smaller, specialized data sets, such as PDBbind, to refine its predictive capabilities for specific binding affinities. This transfer learning strategy leverages the broader structural knowledge acquired during the initial training phase and adapts it to more focused tasks, enhancing the model’s generalizability and accuracy. In addition, the original version, Yuel, 11 relied solely on protein sequences and 2D ligand structures, addressing the overfitting issues in protein–small-molecule interaction predictions. In contrast, Yuel 2 integrates protein structural information, with the aim of enhancing predictive accuracy. Finally, Yuel 2 extends beyond the conventional aim of predicting K d to encompass K i and IC 50 values for protein–small-molecule interactions and also aims to address the challenge of accurately predicting activity cliffs. The challenge of addressing activity cliffs arises primarily from the limited number of samples available in current experimental protein–ligand data sets, where compounds with only slight structural variations often exhibit significant differences in binding affinities to the same protein. Yuel 2 utilizes a pretraining data test that provides a more diverse set of samples, including those with subtle structural differences. RESULTS Pretraining with Large-Scale Docking Data Set. Yuel 2 is composed of Yuel-SE (Structural Encoder) for pretraining and Yuel-AP (Affinity Predictor) for fine-tuning ( Figure 1 ). In the pretraining phase, Yuel-SE processes protein pocket structures and small-molecule 2D structures and then subjects the processed features to Yuel-AP, which predicts the docking scores. Yuel 2 constructs a fully connected graph between the protein pocket and the small molecule. This approach allows Yuel-SE to learn comprehensive structural features from a large-scale data set. During the fine-tuning phase, Yuel-AP is trained with the fixed parameters of Yuel-SE on a smaller experimental data set to predict binding affinities. As part of the pretraining step in our transfer learning strategy, we evaluate the ability of Yuel 2 to predict docking scores generated by two docking software, MedusaDock and AutoDock Vina. To compile a comprehensive data set for the pretraining, we begin by expanding the PDBbind data set through systematic docking, pairing each protein with multiple ligands to create a large and diverse set of protein–ligand interactions. For each of these pairs, MedusaDock and AutoDock are employed to compute binding scores. Medusa-Score, the scoring function of MedusaDock, includes various energy components such as VDW_A (van der Waals attraction), VDW_R (van der Waals repulsion), SOLV (solvation), SB (side chain–backbone), HB_BB (hydrogen bond between backbone and backbone), HB_SB (hydrogen bond between side chain and backbone), and HB_SS (hydrogen bond between side chain and side chain). Similarly, AutoDock Vina scores are composed of energy terms like Gauss 1, Gauss 2, repulsion, hydrophobic, and hydrogen-bonding energies. Yuel 2 is pretrained on this extensive data set to predict these energy terms, demonstrating high accuracy. Yuel 2 predicts the energy terms using a single output layer, which produces a vector in which each element corresponds to a specific energy term. Specifically, Yuel 2 achieves a correlation coefficient of 0.97 for predicting VDW_A ( Figure 2 ), a major component of MedusaScore ( Figure S1 ), and a correlation coefficient of 0.94 for predicting the total MedusaScore ( Figure S2 ). Yuel 2 also showed strong predictive performance for other energy components ( Figure 2 ) and achieved an accuracy of 0.8 in predicting AutoDock Vina docking scores ( Figure 2 ). These findings emphasize the capability of Yuel 2 to replicate and even enhance the predictive performance of traditional docking scores, underscoring its potential as a powerful tool for virtual screening. Figure 1. Open in a new tab Architecture of Yuel 2. (a) We adopt transfer learning to train Yuel 2. Yuel is composed of a structure encoder (Yuel-SE) and an affinity predictor (Yuel-AP). We compile the original PDBbind data set to a pretraining data set and a fine-tuning data set (Methods). Yuel 2 is first trained on the pretraining data set, and then Yuel-SE parameters will be kept fixed, and the parameters of Yuel-AP will be retrained on the fine-tuning data set. (b) Both proteins and small molecules are converted to graphs. Each residue in the protein and each atom in a small molecule are represented by nodes. Residue contacts in the protein and atom bonds in the small molecule are represented by edges in graphs. (c) Yuel-SE first encodes the protein pocket structure and small-molecule structure to graphs. The two graphs will be merged into a large graph, which will then be subject to several graph attention layers. The latent vector generated by Yuel-SE will be subject to Yuel-AP, which is essentially several fully connected layers. Figure 2. Open in a new tab Yuel 2 recapitulates the energy items in MedusaScore and AutoDock energy. MedusaScore, the scoring function of MedusaDock, includes various energy components such as VDW_A (van der Waals attraction), VDW_R (van der Waals repulsion), SOLV (solvation), SB (side chain–backbone), HB_BB (hydrogen bond between backbone and backbone), HB_SB (hydrogen bond between side chain and backbone), and HB_SS (hydrogen bond between side chain and side chain). AutoDock Vina scores are composed of energy terms like Gauss 1, Gauss 2, repulsion, hydrophobic, and hydrogen-bonding energies. Yuel 2 is pretrained on this extensive data set to predict these energy terms, demonstrating high accuracy. Fine-Tuning for Protein–Small-Molecule Binding Affinity Prediction. Using the CASF-2016 12 test set, we evaluate the scoring power of Yuel 2 by computing the binding affinity of a protein–ligand complex based on its structure. The correlation between the experimental binding affinity data and the values predicted by Yuel 2 is depicted in Figure 3a , with a Pearson correlation coefficient of 0.85. A comparison of several deep/machine learning models 13 – 16 published in recent years, along with their scoring power on the same test set, is provided in Figure 3b . Most models achieve a Pearson correlation coefficient from 0.82 to 0.84, indicating that Yuel 2 performs at least comparably to the best among them. Figure 3. Open in a new tab Evaluation of the scoring power and ranking power of Yuel 2. (a) Correlation between the experimental binding affinity data and the predicted binding affinity in the CASF-2016 data set. The Pearson correlation coefficient is 0.85. (b) Comparison of the average accuracy of Yuel 2 on the CASF-2016 data set compared to the other methods. (c) Comparison of the Spearman correlation coefficient of Yuel 2 on the CASF-2016 data set compared to the other methods. (d) Pearson’s correlation coefficient of Yuel 2 tested on the K d , K i , and IC 50 data sets. Unlike most other models, which directly take 3D protein–ligand complex structures as input, Yuel 2 uses only 3D binding pocket structures and 2D ligand structures. Predicting protein–ligand binding affinity without precise binding mode information is inherently more challenging. The scoring power of deep learning models relying on crystal protein–ligand complex structures would decrease significantly if complex structures derived from molecular docking were used instead. Although the scoring power test provided by the CASF-2016 benchmark is a fundamental assessment of a protein–ligand interaction scoring function’s quality, it is not very meaningful to “predict” binding affinity based on crystal complex structures because obtaining these structures is more difficult than measuring binding affinity experimentally. Therefore, our Yuel 2 model is more robust and efficient for predicting protein–ligand binding affinity without requiring prior acquisition or generation of the corresponding protein–ligand complex structures. Another assessment provided by CASF-2016 is the ranking power test. Ranking power focuses on ranking ligands that are binding to a particular target protein by their binding affinity. This test better reflects the practical application of a scoring model. In the ranking power test, Yuel 2 outperforms other top-ranked nondocking methods in the CASF-2016 test set ( Figure 3c ). Yuel 2 achieved an average Spearman correlation coefficient of 0.682 on 57 target proteins and an AUROC of 0.86 in the CASF-2016 test set ( Figure S4 ). Finally, we evaluated the ability of Yuel 2 to predict multiple binding metrics. The PDBbind general data set provides protein–ligand pairs characterized by three distinct metrics: K d , K i , and IC 50 . To evaluate Yuel 2, we train and test the model using the PDBbind general data set, with an 80:20 split between the training and test sets. The results ( Figure 3d ) show that Yuel 2 achieves average Pearson correlation coefficients of 0.81 for K d , 0.85 for K i , and 0.75 for IC 50 . These results are obtained by using multitask learning, and we have also tested Yuel 2 by using single-task learning ( Figure S6 ), which are comparable to multitask learning. Overall, Yuel 2 demonstrates Pearson correlation coefficients ranging from 0.75 to 0.85, indicating a robust performance in predicting binding affinities across different metrics and data sets. Evaluation of the Performance on the Activity Cliff Data Set. To rigorously evaluate Yuel 2’s ability to address activity cliffs, we compile a specialized data set from BindingDB, which is known for its extensive collection of protein–ligand binding affinities. Our goal is to assess Yuel 2’s ability to accurately predict binding affinities for ligands that are structurally similar but exhibit significantly different affinities for the same protein. We select proteins from the BindingDB data set that have multiple associated ligands with experimentally determined binding affinities, which ensures a diverse set of ligand interactions for each protein. For each selected protein, we identified pairs of ligands with high structural similarity. Structural similarity is quantified using FP2 fingerprints and Tanimoto coefficients. For each ligand pair, we calculate the difference in their binding affinities and multiply it by their structural similarity, resulting in a metric we term the Similarity-Weighted Affinity Difference (SWAD). This metric highlights pairs where structural similarity is high but binding affinity differences are large, making them challenging cases for prediction methods. To evaluate Yuel 2’s predictive performance, we calculate the SWAD for all ligand pairs within the same protein. We used Yuel 2 to predict the binding affinities for these ligand pairs. We compute the predicted SWAD values by taking the difference in predicted binding affinities and multiplying by the structural similarity of the ligand pairs. The predicted SWAD values are compared with the true SWAD values. The Pearson correlation coefficient between the predicted and true SWAD values is calculated to quantify the predictive accuracy of Yuel 2. Our analysis yields a Pearson correlation coefficient of 0.65 ( Figure 4 ), indicating Yuel 2’s ability to differentiate between ligands with similar structures that have different binding affinities. As activity cliffs represent a prominent challenge in predicting protein–ligand interactions, the ability of Yuel 2 to differentiate between small molecules on these activity cliffs represents a major milestone in our understanding of molecular interactions. Figure 4. Open in a new tab Evaluation of Yuel 2 on the activity cliff data set. (a) Predicted SWAD vs the true SWAD for all pairs of ligands in the activity cliff data set. (b) Distribution of the Pearson correlation coefficient of the predicted SWAD in the activity cliff data set. We have also evaluated Yuel 2 on an activity cliff benchmark data set, 17 comparing its performance to ACGCN, 17 a state-of-the-art model specifically designed for this task. The benchmark data set includes three target proteins—thrombin, Mu opioid receptor, and melanocortin receptor 4—each with a large number of annotated small-molecule ligands enriched with matched molecular pairs. For these targets, Yuel 2 achieves a performance comparable to that of ACGCN ( Tables S1 – S3 ). The activity cliff data set focuses on well-defined targets, and ACGCN is explicitly trained for each data set, optimizing its performance for these specific scenarios. In our comparison, we adopt the same training strategy as ACGCN, which primarily highlights the capabilities of model architecture rather than the advantages of Yuel 2’s transfer learning approach. Yuel 2’s transfer learning framework is designed to leverage pretrained representations from diverse data sets, enabling it to generalize across a wide range of proteins and small molecules. However, this strength is less pronounced in the activity cliff benchmark, where the focus is on specific, well-defined targets. Despite this, the comparable performance of Yuel 2 to ACGCN demonstrates its robustness, even in scenarios where the transfer learning advantage is not fully utilized. DISCUSSION Almost all the existing protein–ligand affinity prediction methods rely solely on the PDBbind data set for training, which leads to a significant limitation due to the scarcity of available data. This constraint hinders the generalizability and robustness of these models. While transfer learning has been widely applied in various fields, Yuel 2 uniquely leverages it by employing a two-step training process, where Yuel 2 is first pretrained on a large-scale data set to predict docking scores and subsequently fine-tuned with a smaller, specialized data set. This methodology enables Yuel 2 to capture a more comprehensive representation of protein–ligand interactions, distinguishing it from conventional approaches. The pretraining phase allows the model to learn from a diverse set of protein–ligand pairs, capturing a broad range of structural and interaction features. Fixing the parameters of Yuel-SE during the fine-tuning phase ensures that the model retains the learned structural knowledge while adapting to specific binding affinity metrics in a smaller data set. This strategy not only improves the model’s ability to predict binding affinities accurately but also enhances its robustness across different types of data. Transfer learning thus provides a mechanism to overcome the limitations imposed by small data sets and ensures that Yuel 2 benefits from the richness of both large-scale and targeted data. Predicting multiple metrics, including K d , K i , and IC 50 , provides a distinct advantage over models focused on a single metric. The comprehensive approach captures a broader spectrum of interactions between proteins and small molecules, offering a more nuanced understanding of their binding dynamics. While a model predicting only K d may excel in specific scenarios, incorporating IC 50 and K i predictions enhances versatility in selecting effective drugs and a direct view of the impact of a small molecule on cell functioning. For example, IC 50 represents the inhibitor concentration at which the response is reduced by half, which provides insights into inhibitory potency. In addition, we acknowledge that each metric carries distinct pathway-specific information— K d represents intrinsic binding affinity, K i reflects enzymatic inhibition, and IC 50 is influenced by experimental conditions. To effectively integrate these related but distinct measurements, Yuel 2 employs a two-phase training strategy: during pretraining, the model learns fundamental protein–ligand interaction patterns, while fine-tuning enables it to capture the unique characteristics of each metric. This approach ensures that the model is generalized across different affinity types while preserving the specificity of individual binding metrics. However, the IC 50 predictions made by Yuel 2 are also implicitly tied to the particular experimental settings as in the training set, which is a limitation of Yuel 2. Multitask fine-tuning typically excels when tasks share underlying patterns or dependencies, as it can leverage shared representations to improve generalization. For example, predicting related biochemical properties like K d and K i could benefit from multitask learning if their underlying mechanisms are strongly correlated. However, in our study, the lack of significant differences suggests that while K d , K i , and IC 50 are related, they measure distinct biochemical properties with unique mechanisms. This limits the potential benefits of shared representations in a multitask framework. Yuel 2’s adept handling of the activity cliff problem holds significant implications for advancing research in various domains. Activity cliffs, where small chemical modifications lead to disproportionately large changes in bioactivity, pose a formidable challenge in the design of bioactive compounds. By effectively navigating and mitigating activity cliffs, Yuel 2 contributes to the identification and optimization of compounds with subtle structural variations, ensuring a more comprehensive understanding of their bioactivity. This capability is crucial not only for drug discovery but also for other fields, such as nutrition research, where fine-tuning the molecular structures of bioactive compounds, such as nutrients or nutraceuticals, can enhance their efficacy while minimizing unintended effects. This breakthrough accelerates the development of nutritionally beneficial compounds and promotes a more targeted and efficient approach to designing functional foods and supplements with optimal health benefits. While Yuel 2 currently focuses on protein–ligand affinity prediction, its architecture and training strategy can be extended to RNA–ligand interactions in the future. RNA–ligand binding plays an increasingly important role in areas such as RNA-based therapeutics and gene regulation. By incorporating RNA-specific features into future versions of Yuel 2, we could provide valuable insights into RNA–ligand affinity prediction. Compared to protein–ligand affinity prediction, data on RNA–ligand complexes are more limited. Therefore, applying our proposed two-phase strategy to expand the training data set would be particularly well suited to address this challenge. Finally, Yuel 2 can predict binding affinities using only the separate protein and ligand structures as inputs without requiring the docking software to obtain the complex structure. This feature significantly enhances the speed and efficiency of our approach for it eliminates the need for time-consuming docking processes. Many existing neural network models necessitate a complex structure as an input, thereby requiring the integration of docking procedures. In contrast, our models bypass this requirement, offering a streamlined and rapid prediction workflow. METHODS Architecture of Yuel 2. Yuel 2 comprises two main components: Yuel Structural Encoder (Yuel-SE) and Yuel Affinity Predictor (Yuel-AP). Yuel-SE is designed to extract meaningful structural features from both the protein pocket and the small molecule. During the pretraining phase, Yuel-SE is combined with Yuel-AP and trained on a large-scale data set to predict docking scores generated by traditional docking software like MedusaDock and AutoDock Vina. This extensive pretraining allows Yuel-SE to learn the intricate structural patterns and interactions between proteins and ligands. Yuel-SE requires two inputs to estimate interactions between a specific compound and the protein target: (i) the compound’s 2D structure (SMILES 18 or INCHI 19 code) and (ii) the pocket structures. We employ an in-house C++ program to extract the pocket structure of a given protein. For proteins with unknown binding pockets, we use the FPocket 20 program to predict the position of the binding pocket. We used a graph structure to represent the pocket structure. Each node in the graph represents the “C α ” atom in a protein residue, and each edge in the graph represents the contact between two residues with features such as whether the two residues are neighboring and the distance between the two residue C α atoms. Each residue corresponds to a column in the BLOSUM62 matrix. The features of nonstandard amino acids are zero-initialized. Given the SMILES of the compound, Yuel-SE first employs rdkit 21 to represent the structure by a graph ( N , V , E ), where N is the number of nodes, V is the feature vector of each atom, and E is the feature vector of each bond. The feature of each atom is the concatenation of the one-hot encoding of the atom type, number of bonds, bond type, mass, and charge vectors. The feature of each bond is its bond order. The graphs of the protein and the small molecule are merged into an intact graph by adding the edges between each node in the protein graph and each node in the small-molecule graph ( Figure 1 ). The merged graph is then subject to 8 graph attention layers. The node features after the graph attention layers are then subject to linear layers to predict a latent vector, which is used as the input for Yuel-AP. After pretraining, Yuel-SE’s parameters are fixed, and Yuel-AP is fine-tuned on a smaller data set, the PDBbind general set, to predict binding affinities more accurately. Yuel-AP consists of six linear layers, taking as input the latent ligand predicted by Yuel-SE and directly predicting the overall binding affinity. Preparation of the Fine-Tuning Data Set. The PDBbind “general set” is a comprehensive collection of protein–ligand complexes with the available 3D structures from the Protein Data Bank (PDB) and the corresponding experimental binding affinity data (i.e., K d , K i , and IC 50 ) curated from the literature. This data set has been extensively used for developing deep learning and machine learning models aimed at predicting protein–ligand binding affinity. For our study, we used the PDBbind general set (v.2020), composed of a total of 19,443 protein–ligand complexes, as the initial pool for training Yuel 2. The PDBbind general set provides the protein PDB structure file and small-molecule SDF and Mol2 files as well as the PDB structure file of the pocket. We used an in-house C ++ program to process the pocket PDB structure to convert it to a graph file. We used RDKit to process the small-molecule Mol2 file and remove those that could not pass the sanitize operation, resulting in 19,386 successfully processed protein–ligand complexes from the PDBbind general set (v.2020). Pretraining and Fine-Tuning Data Set. To compile a large-scale data set for pretraining, we randomly selected 500 PDB IDs from the PDBbind general set. For each selected PDB ID, we performed MedusaDock docking attempts and AutoDock Vina docking attempts to obtain the MedusaScore energies and AutoDock scores. The docked file and the docking scores are stored in a Sqlite 3 database file. This approach yielded a total of 500 × 500 = 250,000 protein–ligand pairs, creating a diverse and extensive data set for pretraining. The PDBbind general set is used as the fine-tuning data set. We split it into the training and test sets with a ratio of 8:2. Finally, we use the CASF-2016 data set for the comparison of Yuel 2 to other methods. To ensure the reliability of our model, we excluded 285 protein–ligand complexes in the CASF-2016 test set from our training set. Activity Cliff Data Set. To evaluate Yuel 2’s performance in addressing activity cliffs, where structurally similar compounds exhibit markedly different binding affinities for a single protein target, we propose the following approach to compile a data set. We begin with the BindingDB data set, which provides a comprehensive repository of protein–ligand binding affinities. The process involves several steps to isolate relevant examples of activity cliffs. First, we identify proteins with a significant number of binding affinities reported for multiple ligands in the BindingDB data set. This ensures that the data set will contain a sufficient variety of compounds for each protein target. For each selected protein, we extracted a subset of ligands that have similar chemical structures. This is achieved by clustering compounds based on their molecular fingerprints or using chemical similarity metrics. Within each cluster of structurally similar compounds, we identify cases where there is substantial variation in binding affinities. These instances are characterized by a high discrepancy in affinity values despite the close structural resemblance among the compounds. We then construct a data set that includes these pairs of similar compounds with divergent binding affinities. Each entry should contain the protein target, structurally similar compounds, and their corresponding binding affinities. We format the binding affinity information alongside the molecular representations of the ligands and protein structure. Ensuring that the data set is split into training and validation subsets, we may properly assess Yuel 2’s ability to accurately predict binding affinities and handle activity cliffs. We propose a new metric, SWAD, to evaluate the performance of Yuel 2 on the activity cliff data set. For each ligand pair for a specific protein, we calculate the difference in their binding affinities to the protein and multiply it by their structural similarity SWAD = K d 2 - K d 1 ⋅ S 12 where K d 1 and K d 2 represent the binding affinity of ligand 1 and ligand 2 to the protein, and S 12 represents the structural similarity of the two ligands. Here, we use the Tanimoto score of the FP2 fingerprints of the two ligands as a structural similarity. Model Evaluation. In the context of computational chemistry and drug discovery, “scoring power” refers to the ability of a scoring model to predict a binding score that exhibits a linear correlation with experimental protein–ligand binding data, based on valid inputs. In our study, we evaluated the scoring power of Yuel 2 using the CASF-2016 benchmark, which is widely recognized for assessing scoring functions. The test set in CASF-2016 comprises 285 diverse protein–ligand complexes with high-quality crystal structures and experimental binding data. In alignment with the CASF-2016 protocol, our evaluation metrics include the Pearson correlation coefficient between the experimental binding data and predicted values. Another related evaluation is the ranking power test. “Ranking power” assesses the ability of a scoring model to correctly rank known ligands of a common target protein. Unlike scoring power, ranking power measures how well the model ranks ligands based on their binding data, rather than expecting a linear correlation between predicted binding data and actual values. For each target protein, we computed the Spearman correlation coefficient between the experimental binding data and predicted values. The overall ranking power was determined by averaging the Spearman correlation coefficient values across all 57 target proteins. Supplementary Material SI NIHMS2157797-supplement-SI.pdf (559.4KB, pdf) ASSOCIATED CONTENT Supporting Information The Supporting Information is available free of charge at https://pubs.acs.org/doi/10.1021/acs.jcim.4c02256 . Extra testing results of Yuel 2 ( PDF ) ACKNOWLEDGMENTS We acknowledge support from the National Institutes for Health (R35 GM134864), the National Science Foundation (2040667), and the Passan Foundation. This project was also supported by the Penn State College of Medicine’s Artificial Intelligence and Biomedical Informatics Program. We appreciate Congzhou Sha and Alicia Xie for their enthusiastic proofreading of the manuscript. Footnotes The authors declare no competing financial interest. Complete contact information is available at: https://pubs.acs.org/10.1021/acs.jcim.4c02256 Contributor Information Jian Wang, Department of Neuroscience and Experimental Therapeutics, Penn State College of Medicine, Hershey, Pennsylvania 17033, United States. Nikolay V. Dokholyan, Department of Neuroscience and Experimental Therapeutics, Penn State College of Medicine, Hershey, Pennsylvania 17033, United States; Department of Chemistry and Department of Biomedical Engineering, Pennsylvania State University, University Park, Pennsylvania 16802, United States Data Availability Statement Source codes and test data are deposited at: https://bitbucket.org/dokhlab/yuel2 . REFERENCES (1). Liu Z; et al. PDB-wide collection of binding data: current status of the PDBbind database. Bioinformatics 2015, 31, 405. [ DOI ] [ PubMed ] [ Google Scholar ] (2). Liu T; Lin Y; Wen X; Jorissen RN; Gilson MK BindingDB: a web-accessible database of experimentally determined protein–ligand binding affinities. Nucleic Acids Res. 2007, 35, D198–D201. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] (3). Trott O; Olson AJ AutoDock Vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading. J. Comput. Chem. 2010, 31, 455–461. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] (4). Friesner RA; et al. Glide: a new approach for rapid, accurate docking and scoring. 1. Method and assessment of docking accuracy. J. Med. Chem. 2004, 47, 1739–1749. [ DOI ] [ PubMed ] [ Google Scholar ] (5). Wang J; Dokholyan NV MedusaDock 2.0: Efficient and Accurate Protein–Ligand Docking With Constraints. J. Chem. Inf. Model. 2019, 59, 2509–2515. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] (6). Ding F; Yin S; Dokholyan NV Rapid flexible docking using a stochastic rotamer library of ligands. J. Chem. Inf. Model. 2010, 50, 1623–1632. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] (7). Fan M; Wang J; Jiang H; Feng Y; Mahdavi M; Madduri K; Kandemir MT; Dokholyan NV GPU-accelerated flexible molecular docking. J. Phys. Chem. B 2021, 125, 1049–1060. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] (8). Chirasani VR; Wang J; Sha C; Raup-Konsavage W; Vrana K; Dokholyan NV Whole proteome mapping of compound-protein interactions. Current Research in Chemical Biology 2022, 2, 100035. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] (9). Cruz-Monteagudo M; et al. Activity cliffs in drug discovery: Dr Jekyll or Mr Hyde? Drug Discovery Today 2014, 19, 1069–1080. [ DOI ] [ PubMed ] [ Google Scholar ] (10). Van Tilborg D; Alenicheva A; Grisoni F Exposing the limitations of molecular machine learning with activity cliffs. J. Chem. Inf. Model. 2022, 62, 5938–5951. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] (11). Wang J; Dokholyan NV Yuel: Improving the Generalizability of Structure-Free Compound-Protein Interaction Prediction. J. Chem. Inf. Model. 2022, 62, 463–471. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] (12). Su M; et al. Comparative assessment of scoring functions: the CASF-2016 update. J. Chem. Inf. Model. 2019, 59, 895–913. [ DOI ] [ PubMed ] [ Google Scholar ] (13). Zhang X; et al. Planet: a multi-objective graph neural network model for protein–ligand binding affinity prediction. J. Chem. Inf. Model. 2024, 64, 2205–2220. [ DOI ] [ PubMed ] [ Google Scholar ] (14). Liu X; Feng H; Wu J; Xia K Dowker complex based machine learning (DCML) models for protein-ligand binding affinity prediction. PLoS Comput. Biol. 2022, 18, No. e1009943. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] (15). Meng Z; Xia K Persistent spectral–based machine learning (PerSpect ML) for protein-ligand binding affinity prediction. Sci. Adv. 2021, 7, No. eabc5329. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] (16). Wee J; Xia K Ollivier Persistent Ricci Curvature-Based Machine Learning for the Protein-Ligand Binding Affinity Prediction. J. Chem. Inf. Model. 2021, 61, 1617–1626. [ DOI ] [ PubMed ] [ Google Scholar ] (17). Park J; Sung G; Lee S; Kang S; Park C ACGCN: Graph Convolutional Networks for Activity Cliff Prediction between Matched Molecular Pairs. J. Chem. Inf. Model. 2022, 62, 2341–2351. [ DOI ] [ PubMed ] [ Google Scholar ] (18). Weininger D SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules. J. Chem. Inf. Comput. Sci. 1988, 28, 31–36. [ Google Scholar ] (19). Stein SE; Heller SR; Tchekhovskoi DV An open standard for chemical structure representation: The IUPAC chemical identifier. In Proceedings of the International Chemical Information Conference |15th|; NIST: Nimes, FR, 2003. [ Google Scholar ] (20). Le Guilloux V; Schmidtke P; Tuffery P Fpocket: an open source platform for ligand pocket detection. BMC Bioinf. 2009, 10, 168. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] (21). Landrum G RDKit: A software suite for cheminformatics, computational chemistry, and predictive modeling, 2013. https://www.rdkit.org/RDKit_Overview.pdf . Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials SI NIHMS2157797-supplement-SI.pdf (559.4KB, pdf) Data Availability Statement Source codes and test data are deposited at: https://bitbucket.org/dokhlab/yuel2 . ACTIONS View on publisher site PDF (1007.8 KB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Record · ID 30534 · SHA-256 6e8719a120486c6c
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.