Conceptio › Archive › NCBI PubMed Central
NCBI PubMed Centralopen access

Protocol update to: Protocol to generate dual-target compounds using a transformer-based chemical language model.

Srinivasan S et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
machine learning systems

Protocol update to: Protocol to generate dual-target compounds using a transformer-based chemical language model - PMC Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice STAR Protoc . 2026 Apr 16;7(2):104492. doi: 10.1016/j.xpro.2026.104492 Search in PMC Search in PubMed View in NLM Catalog Add to search Protocol update to: Protocol to generate dual-target compounds using a transformer-based chemical language model Sanjana Srinivasan Sanjana Srinivasan 1 Department of Life Science Informatics and Data Science, B-IT, LIMES Program Unit Chemical Biology and Medicinal Chemistry, Rheinische Friedrich-Wilhelms-Universität, Friedrich-Hirzebruch-Allee 5/6, 53115 Bonn, Germany 2 Lamarr Institute for Machine Learning and Artificial Intelligence, Friedrich-Hirzebruch-Allee 5/6, 53115 Bonn, Germany Find articles by Sanjana Srinivasan 1, 2, 3, ∗ , Jürgen Bajorath Jürgen Bajorath 1 Department of Life Science Informatics and Data Science, B-IT, LIMES Program Unit Chemical Biology and Medicinal Chemistry, Rheinische Friedrich-Wilhelms-Universität, Friedrich-Hirzebruch-Allee 5/6, 53115 Bonn, Germany 2 Lamarr Institute for Machine Learning and Artificial Intelligence, Friedrich-Hirzebruch-Allee 5/6, 53115 Bonn, Germany Find articles by Jürgen Bajorath 1, 2, 4, ∗∗ Author information Article notes Copyright and License information 1 Department of Life Science Informatics and Data Science, B-IT, LIMES Program Unit Chemical Biology and Medicinal Chemistry, Rheinische Friedrich-Wilhelms-Universität, Friedrich-Hirzebruch-Allee 5/6, 53115 Bonn, Germany 2 Lamarr Institute for Machine Learning and Artificial Intelligence, Friedrich-Hirzebruch-Allee 5/6, 53115 Bonn, Germany ∗ Corresponding author [email protected] ∗∗ Corresponding author [email protected] 3 Technical contact 4 Lead contact Collection date 2026 Jun 19. © 2026 The Author(s) This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). PMC Copyright notice PMCID: PMC13098578  PMID: 41999606 Summary We present a protocol to generate dual-target compounds (DT-CPDs) and triple-target compounds (TT-CPDs) interacting with two and three target proteins, respectively, using transformer-based chemical language models. We describe steps for installing software, preparing data, and pre-training the model on pairs of single-target compounds (ST-CPDs), which bind to an individual protein, and DT-CPDs. We then detail procedures for assembling ST-CPD and corresponding DT- or TT-CPD data for fine-tuning on specific protein pairs or triplets and for evaluating model performance on hold-out test sets. For complete details on the use and execution of this protocol, please refer to Srinivasan et al. 1 for DT-CPDs and to Srinivasan et al. 2 for TT-CPDs. This protocol is an update to Srinivasan et al. 3 Subject areas: Bioinformatics, Chemistry, Computer sciences Graphical abstract Open in a new tab Highlights • Guidance on the generative design of dual- and triple-target compounds • Steps for applying transformer-based chemical language models • Instructions for model pre-training and fine-tuning Publisher’s note: Undertaking any experimental protocol requires adherence to local institutional guidelines for laboratory safety and ethics. We present a protocol to generate dual-target compounds (DT-CPDs) and triple-target compounds (TT-CPDs) interacting with two and three target proteins, respectively, using transformer-based chemical language models. We describe steps for installing software, preparing data, and pre-training the model on pairs of single-target compounds (ST-CPDs), which bind to an individual protein, and DT-CPDs. We then detail procedures for assembling ST-CPD and corresponding DT- or TT-CPD data for fine-tuning on specific protein pairs or triplets and for evaluating model performance on hold-out test sets. Before you begin This protocol adapts a transformer-based encoder-decoder architecture from a previous study on the prediction of activity cliffs 4 for the generation of DT- and TT-CPDs for target protein pairs and triplets, respectively. Multi-target compounds with well-defined activity are highly desirable for drug discovery targeting multi-factorial diseases via polypharmacology. 1 , 2 The key feature of the transformer model is the attention mechanism in both the encoder and decoder layers. The multi-head self-attention sublayers in the encoder and decoder enable the model to focus on different segments of the input sequences, learn relationships between tokens in sequence embeddings including the output of the first decoder module, thus gaining long-term memory. 5 Chemical language models (CLMs), like natural language models in general, require input in the form of sequential data, thus making SMILES 6 strings a suitable molecular representation for CLMs. Initially, the SMILES strings of the input sequences are tokenized to generate a vocabulary. During training, the tokenized SMILES strings of the source and the target sequences and their positional embeddings are submitted to the encoder and decoder, respectively. This allows the model to learn the chemical “language” and relationships between the sequences. In the generation phase, only the source sequence is provided as input to generate the corresponding output sequence. In this way, target sequences with implicit semantic features (desired properties) can be designed. Herein, we derive transformer models to learn mappings of ST-CPDs to DT- or TT-CPDs. The transformer is trained in two phases. Firstly, the model undergoes pre-training on a large dataset of ST- and corresponding DT-CPDs covering an extensive target space, enabling the model to learn the syntax of the SMILES strings and the chemical space of DT-CPDs. For generative design of new DT-CPDs, this is followed by fine-tuning when the model is specifically trained to generate DT-CPDs for a specific protein pair. 1 Different from DT-CPDs, only limited amounts of known TT-CPDs with a severely restricted target space are available for learning, essentially precluding pre-training of a generative model based on ST- and corresponding TT-CPDs. Therefore, to generate new candidate TT-CPDs, it was attempted to use the pre-trained ST-/DT-CPD model introduced above for fine-tuning on ST-/TT-CPD pairs using a transfer learning strategy. 2 The key criterion for the evaluation of a fine-tuned model was its ability to regenerate known TT-CPDs or structural analogues of these TT-CPDs that were excluded from training, as detailed in the original publication. 2 Fine-tuned model variants were shown to learn specific TT chemical space and successfully generate TT-CPDs. 2 Figure 1 shows representative examples of designed DT- and TT-CPDs. Figure 1. Open in a new tab Exemplary ST-CPDs and a correctly predicted (exactly reproduced) corresponding DT-CPD (top) and TT-CPD (bottom) Target proteins are specified. In the following, the pre-training and the fine-tuning steps required to generate desired target sequences are detailed using exemplary data from the original publications. 1 , 2 All required code is developed using Python and its libraries. The reported execution times are specific to the machine’s configurations specified in materials and equipment section. Depending on the machine, datasets, and specific software configurations, the execution times are likely to vary. Installation Timing: 10 min 1. Install Python 3 (python = 3.7.4) and the libraries required to run the code. Use a conda virtual environment for easier dependency management. a. Download and install Anaconda following the installation guidelines from https://docs.anaconda.com/anaconda/install/ if conda is unavailable. Note: For a smaller version of Anaconda, install miniconda from https://docs.anaconda.com/miniconda/ Troubleshooting 1 . b. Download the scripts from the link: https://github.com/sana1312/dt_trans . c. Set up the conda environment using the environment.yml file in the repository. i. Navigate to the root folder containing the environment.yml file. ii. Open the environment.yml file in a text editor and change the name parameter to a preferred name (default environment name: dt_trans). d. Create the conda environment. >conda env create -f environment.yml Troubleshooting 2 . Alternatives: Install the packages provided in the key resources table if preferred instead of using a conda environment. e. Activate the conda environment. >conda activate dt_trans (or tt_trans) Troubleshooting 3 . Key resources table REAGENT or RESOURCE SOURCE IDENTIFIER Deposited data Compound activity data ChEMBL (release 33) https://doi.org/10.6019/CHEMBL.database.33 Confirmed or likely colloidal aggregators Aggregator advisor Publication: https://doi.org/10.1021/acs.jcim.0c00675 Database: https://zinc20.docking.org/catalogs/aggregators/ Software and algorithms dt_trans Original DT-CPD paper by Srinivasan and Bajorath 1 Github: https://github.com/sana1312/dt_trans Zenodo: https://doi.org/10.5281/zenodo.14310465 tt_trans Original TT-CPD paper by Srinivasan and Bajorath 2 Github: https://github.com/sana1312/dt_trans/tree/tt_trans Zenodo: https://doi.org/10.5281/zenodo.18495653 Pandas 1.0.0 GitHub https://github.com/pandas-dev/pandas Lilly Medchem Rules GitHub Publication: https://doi.org/10.1021/jm301008n Github: https://github.com/IanAWatson/Lilly-Medchem-Rules NumPy 1.17.3 GitHub https://github.com/numpy/numpy PyTorch 2.10.0 GitHub https://github.com/pytorch/pytorch RDKit 2020.03.2.0 GitHub https://github.com/rdkit/rdkit Scikit-learn 0.21.3 GitHub https://github.com/scikit-learn/scikit-learn TensorboardX 2.0 GitHub https://github.com/lanpa/tensorboardX Other AMD EPYC MILAN 7513 @ 2.60 GHz N/A N/A NVIDIA A40 N/A N/A 48 GB RAM N/A N/A Ubuntu 22.04.3 LTS Operating System N/A N/A Open in a new tab Materials and equipment Computational resources Component Brand Model/Capabilities/Version CPU AMD EPYC MILAN 7513 @ 2.60 GHz (32-Core) GPU NVIDIA A40 (48 GB GDDR6, 10752 CUDA Cores, 300 W Max Power, PCIe Gen 4 x16), CUDA Toolkit version 12.6 Update 3 RAM Any 48 GB Operating System Linux Ubuntu 22.04.3 LTS Open in a new tab Step-by-step method details This section provides detailed methods for data preparation, data processing, training, and evaluation of the transformer models. We describe how to properly format the data to be used by the transformer model and how to build the vocabulary. Subsequently, the steps required for pre-training and subsequent fine-tuning of the model are explained. Data preparation and preprocessing Timing: 10 min The steps define the organization of the data by the user. 1. Prepare the dataset as a comma-separated file (.csv) file with two columns named ‘Source_Mol’ (ST-CPDs) and ‘Target_Mol’ (DT-CPDs). The columns should contain SMILES representations of the source and the target compounds, respectively. An exemplary data file is shown in Figure 2 . Note: Compound activity data curation 1 resulted in 120,195 unique qualifying compounds with activity against 1747 unique target proteins. For all possible target pairs, a search for DT-CPDs was carried out. DT-CPDs were detected for 75,280 target pairs involving 1457 unique targets. A subset of 7747 target pairs included targets from different families. 1 Note: For DT-CPD design, 1 three versions of the datasets are created, differing in the similarity between the source and target molecules. Therefore, we calculate Tanimoto similarity 7 based on extended connectivity fingerprints 8 (diameter=4, nbits=2048) of compounds. The calculations were implemented using RDKit. 9 The models derived from these three datasets correspond to the “0-”, “25-”, and “50-models”, where the compound pairings have at least 0%, 25%, and 50% similarity, respectively. The generation of alternative datasets with different similarity levels is optional. Note: For TT-CPD design, 2 a Tanimoto similarity 6 threshold of 50%, calculated using the folded extended connectivity fingerprint s (diameter=4, nbits=2048), was applied to the compounds during both the pre-training and fine-tuning phases. 2 Only input (ST) and output (DT or TT) compound pairs meeting this threshold were selected. All calculations were carried out using RDKit. The resulting dataset for pre-training consisted of 188,677 pairs of ST- and DT-CPDs (including 24,109 unique DT-CPDs). 2. Build the vocabulary from the data. a. Navigate to the Chemical_Language_Model folder. >cd Chemical_Language_Model b. Generate the vocabulary by running the preprocess script. >python preprocess.py --input-data-folder <path_to_folder> --data-file-name <file_name.csv> Note: The following parameters are required. i. INPUT_DATA_FOLDER: Path to the folder containing the data file. ii. DATA_FILE_NAME: The name of the input data file with .csv extension. Troubleshooting 4 . Note: The output file, vocab.pkl can be accessed in the <INPUT_DATA_FOLDER>. 3. After the generation of vocab.pkl, split the dataset from Step 1 for pre-training and fine-tuning. Note: For DT-CPD design, the provided data folder contains exemplary datasets for pre-training and fine-tuning. The compound pair data corresponding to specific target pairs are separated for the fine-tuning phase, while the remaining compound pairs are utilized for pre-training. For details on the target pairs selected for fine-tuning, please, see the original publication. 1 Note: For TT-CPD design, the provided data folder contains exemplary datasets for pre-training and fine-tuning. The pre-training data consists of ST-/DT-CPD pairs while the fine-tuning data comprises ST-/ TT-CPD pairs for specific target triples. For details concerning the target triples selected for fine-tuning and the number of corresponding ST- and TT-CPDs, please, see the original publication. 2 Note: The SMILES data are extracted from ChEMBL database 10 (release 33) and curated using ChEMBL’s internal filter criteria and additional publicly available filter methods, as reported in the original publication. 1 Figure 2. Open in a new tab Exemplary .csv file containing the SMILES sequences of source (input) and target (output) compounds Files with this format are provided in the data folder. Data splitting Timing: 5 min This section details the process of splitting the data into training, validation, and test sets. 4. Run the script for splitting the data. >python split_data.py -–input-data-folder <path_to_folder> --data-file-name <file_name.csv> Troubleshooting 4 . Note: The parameter options are provided below, with each parameter labeled as required or optional. a. INPUT_DATA_FOLDER: Path to the folder containing the input file. The output files are saved in the same folder (required). b. DATA_FILE_NAME: The name of the input file with .csv extension (required). c. SPLIT_TEST: Boolean, default value: false. Then, the data is split into training and validation sets. Use this option if a test set based on a particular criterion is already available (for example, if test compounds are excluded from the training phase, according to the original publication 1 ). Set this parameter to true if the data should be randomly divided into train, validation, and test sets (optional). 5. Repeat the process to generate the datasets for pre-training and fine-tuning. Pre-training Timing: 5–7 days Note: Given the computational requirements for transformer pre-training, the use of a GPU is highly recommended (see key resources table ). Note: Given these computational requirements, as an alternative, we also provide the pre-trained model and the vocabulary file in the code and data depositions (dt_trans/data/model.zip or tt-trans/pretrained_model.zip; see key resources table ) The model can immediately be fine-tuned (beginning at epoch 35) following the fine-tuning section below. In this section, the main steps for pre-training the model and further evaluating the results are explained. 6. Execute the training script to start the training process. >python train.py -–data-path <path_to_folder> Troubleshooting 4 . Troubleshooting 5 . Note: The parameters are defined in ‘configurations/opts.py’ in the training parameters section. Default values can be modified directly in the opts.py file or provided as arguments when running the command given above. Each parameter is annotated as required or optional. Default values will be used for the optional parameters, if not provided as arguments. a. DATA_PATH: Path to the folder containing the input files. Make sure that the prepared input (.csv) files and the vocab.pkl are available in this folder (required). b. SAVE_DIRECTORY: Path to the folder saving the output files, default ‘train’ (optional). c. BATCH_SIZE: Batch size to be used for training, default 64 (optional). d. NUM_EPOCH: Number of training epochs, default 200 (optional). e. STARTING_EPOCH: Training from the given starting epoch, default 1 (optional). f. CUDA_DEVICE: Cuda device to use for training, default 0 (optional). g. N: Number of encoder and decoder layers, default 6 (optional). h. H: Number of attention heads, default 8 (optional). i. D-MODEL: Model and embedding dimension, default 256 (optional). j. D-FF: Dimension for the feed-forward network, default 2048 (optional). k. DROPOUT: Dropout rate, default 0.1, (optional). l. LABEL_SMOOTHING: Epsilon value for label smoothing, default 0.0 (optional). m. FACTOR: Factor multiplied to the learning rate scheduler for the NoamOpt, default 1.0 (optional). 3 n. WARMUP_STEPS: Number of warmup steps for custom decay, default 4000 (optional). o. ADAM_BETA1: Beta1 parameter for Adam optimizer, default 0.9 (optional). p. ADAM_BETA2: Beta2 parameter for Adam optimizer, default 0.98 (optional). q. ADAM_EPS: Eps parameter for Adam optimizer, default 1e -9 (optional). Note: This step saves a checkpoint of the model at each epoch in the <SAVE_DIRECTORY>/checkpoint folder. Note: In our publication reporting TT-CPD design, 2 the following inference step for pre-training was omitted (and only performed after fine-tuning). However, this step can be carried out following pre-training if a user would like to determine the proportions of chemically valid and unique compounds among all generated SMILES strings. 7. Locate the train_model.log in the folder containing the output files from Step 1. Here, the training and validation loss for each epoch are reported. Select the epoch with the smallest validation loss for the next steps. 8. Execute generate.py to generate target sequences. >python generate.py -–data-path <path_to_folder> --test-file-name <file_name> --model-path <path_to_model> --epoch <epoch_number> Troubleshooting 4 . Troubleshooting 5 . Troubleshooting 6 . Note: The parameters are defined in ‘configurations/opts.py’ in the evaluation parameters section. Similar to the training parameters, the default values can be changed directly in the opts.py or given as an argument when calling the above command. The required and optional parameters are annotated as such. a. DATA_PATH: Path to the folder containing the test and vocabulary files (required). b. TEST_FILE_NAME: Name of the test file without .csv extension (required). c. SAVE_DIRECTORY: Directory to save the result files, default ‘evaluation’ (optional). d. MODEL_PATH: Path to the folder containing the model to be used for evaluation (required). See Step 2. e. EPOCH: Epoch number to use for evaluation (required). See Step 2. f. BATCH_SIZE: Batch size to use for evaluation, default 64 (optional). g. NUM_SAMPLES: Number of sequences or compounds to generate, default 50 (optional). Note: Over a maximum of 100 trials, the model samples a maximum of 50 unique compounds. If the NUM_SAMPLES parameter is set to values greater than 100, the MAX_TRIALS parameter in the generate.py script must be set accordingly. h. DECODE_TYPE: Decode strategy ‘greedy’ or ‘multinomial’, default ‘multinomial’ (optional). Note: The result ‘generated_molecules.csv’ is found in the <SAVE_DIRECTORY>. 9. Execute the script for analyzing the results. >python result_analysis.py –-result-path <path_to_result_file> -- train-path <path_to_train_file> Troubleshooting 4 . Note: The parameters listed below are contained in results_analysis.py and labeled as optional or required. a. RESULT_PATH: Path to the output file from the generation step with .csv extension, see Step 3 (required). b. TRAIN_PATH: Path to the train file with .csv extension (required). c. SAVE_DIRECTORY: Path to the folder to save the results, default ‘analysis’ (optional). d. CALCULATE_SIMILARITY: Boolean, default value: false. If set to true, this option saves a file containing the average Tanimoto similarity, nearest-neighbor similarity, and the SMILES of the nearest neighbor training compound for each generated compound, compared to the training set compounds. For a large number of generated compounds, this step might take several minutes, as it involves comparing each generated compound with all training compounds. Note: Columns are added to the original results file from the analysis and saved. Anaysis.txt contains a summary of the results. Similarity.tsv contains the results of the similarity calculations. The files are found in the <SAVE_DIRECTORY>. Fine-tuning Timing: 1–2 days (for fine-tuning on a single dataset) Here, the main steps for the fine-tuning phase of a pre-trained model are detailed. For fine-tuning , we re-use the scripts from the pre-training phase for the fine-tuning data sets. Please, refer to the corresponding steps specified above for details on the parameter options and settings. 10. Execute train.py for training on the fine-tuning data. >python train.py -–data-path <path_to_folder> --starting-epoch <epoch_number> Note: Fine-tuning starts from the optimized (best performing) version of the pre-trained model (see Steps 2 and 3 of the pre-training phase). Thus, the STARTING_EPOCH parameter setting is crucial for resuming training during the fine-tuning phase. Please, ensure that the vocab.pkl file from the earlier steps resides in the input folder. Troubleshooting 4 . Troubleshooting 5 . 11. Following the training phase, navigate to train_model.log and select the model version with the smallest validation loss. 12. Execute the evaluation script. >python generate.py -–data-path <path_to_folder> --test-file-name <file_name> --model-path <path_to_model> --epoch <epoch_number> Troubleshooting 4 . Troubleshooting 5 . Troubleshooting 6 . 13. Execute the analysis script. >python result_analysis.py –-result-path <path_to_result_file> -- train-path <path_to_train_file> Troubleshooting 4 . 14. Repeat steps 1 and 2 for each fine-tuning dataset. Expected outcomes The results of model evaluation and analysis are saved in the designated folders. During the evaluation, the fine-tuned models generate output (target) DT- or TT-CPDs for the corresponding input (source) ST-CPDs. SMILES strings of the output compounds are recorded together with two measures: number of trials and valid output strings (including duplicates), as illustrated for DT-CPDs in Figure 3 . Figure 3. Open in a new tab Results file with model predictions (Predicted_smi_1, Predicted_smi_2, etc.) The final two columns on the right contain the total and the valid counts of the generated SMILES, respectively. During the analysis, we calculate the count and percentage of unique (non-repetitive) and novel candidate DT-or TT-CPDs (distinct from training compounds) for each source compound, which are added to the results. To determine reproducibility, we compare the complete sets of known test compounds and newly generated candidates for each activity class used for fine-tuning and determine the number of known target compounds that are exactly reproduced by the model. The statistics are displayed on the console and saved to a .txt file. Figure 4 shows a console with exemplary results statistics. Figure 4. Open in a new tab Console display of results statistics for an exemplary dataset The overall reproducibility, average validity, uniqueness, and novelty (highlighted in red) are reported. Limitations Transformer models typically require extensive computational resources and long training times, especially for large datasets, as reported herein. These models also require sequential input data such as SMILES strings that are generated from molecular graphs. Accordingly, it is difficult to encode three-dimensional structures or different conformations of compounds. Currently, for multi-target compounds, transformer pre-training is essentially limited to DT-CPDs because for known compounds with activity against three or more unrelated targets, insufficient data is available for model derivation. This may change in the future, providing an opportunity, for example, to train a transformer for the generation of compounds with defined triple-target activity. However, fine-tuning of the pre-trained DT-CPD transformer on confined numbers of available TT-CPDs was sufficient to reproduce known TT-CPDs from hold-out test sets and generate many new candidates, indicating that transfer learning enabled the model to chart TT chemical space. However, benchmark calculations are required to determine reproducibility rates relative to the total number of sampled compounds. Although the model generated candidate compounds identical or highly similar to the known TT-CPDs, the desired multi-target activity of novel candidate compounds must ultimately be evaluated and confirmed experimentally. Troubleshooting Problem 1 Related to installation : the user might encounter “conda not recognized as an internal or external command”, while working with the WindowsPowerShell. Potential solution Please, ensure sure that the Anaconda (or Miniconda) installation paths are added to the system’s environment variables. Initialize conda in the PowerShell by running the following command: >conda init and reopen the terminal. Alternatively, you can use Anaconda Prompt. Problem 2 Related to installation : the environment.yml file is derived for use across different platforms. However, if the generation of the conda environment fails and reportds “PackagesNotFoundError”, the listed packages might need to be installed (or reinstalled). Potential solution If the problematic packages are not listed in the key resources table , remove them from the environment.yml file and recreate the environment. If the problem still persists, manually install the packages listed in the key resources table using pip or conda: pip install <package_name> or conda install <package_name>. Problem 3 Related to installation : the user might encounter “EnvironmentNameNotFound” when attempting to activate the environment. Potential solution Please, verify that there are no typographical errors in the name of the environment when running conda activate <environment_name>. Alternatively, list the environment using conda env list. Problem 4 Related to data preparation and preprocessing , data splitting , pre-training and Fine-tuning: if the command fails with a “FileNotFoundError”, indicating a required file is missing, consider the following solution. Potential solution Please, ensure that there are no typographical errors in the folder name or input file name and verify that the correct extensions are used and that all necessary input files reside in the given folder. Problem 5 Related to pre-training , fine-tuning : command fails with “FileNotFoundError”: missing vocab.pkl. Potential solution To correct this error, ensure that the vocab.pkl file is contained in the input folder. Problem 6 Related to pre-training , fine-tuning : model loading fails with “FileNotFoundError”: cannot find <model_epoch_number.pt> Potential solution Please, ensure that the correct model.pt file is present in the input folder. For instance, if selecting epoch 75, the model_75.pt and model_74.pt files should be contained in the input folder. Resource availability Lead contact Requests for further information, resources, and software should be directed to and will be answered by the lead contact, Jürgen Bajorath ( [email protected] ). Technical contact Questions about technical specifics of the protocol and its execution should be directed to and will be answered by the technical contact, Sanjana Srinivasan ( [email protected] ). Materials availability Not applicable. Data and code availability • The dt_trans code and datasets are available on GitHub at https://github.com/sana1312/dt_trans . • The tt_trans code and datasets are available on GitHub at https://github.com/sana1312/dt_trans/tree/tt_trans . Acknowledgments The authors would like to thank Hengwei Chen and Martin Vogt for their helpful discussions. Author contributions Conceptualization, S.S. and J.B.; software, S.S.; investigation, S.S.; formal analysis, S.S. and J.B.; writing – original draft, S.S.; writing – editing and finalizing the draft, S.S. and J.B. Declaration of interests The authors declare no competing interests. Declaration of generative AI and AI-assisted technologies in the writing process During the preparation of the original manuscript, AI tools were used for spelling and grammar checks and refinement of texts. The authors reviewed and edited the content as required and take full responsibility for the final content of the published article. Contributor Information Sanjana Srinivasan, Email: [email protected]. Jürgen Bajorath, Email: [email protected]. References 1. Srinivasan S., Bajorath J. Generation of dual-target compounds using a transformer chemical language model. Cell Rep. Phys. Sci. 2024;5 doi: 10.1016/j.xcrp.2024.102255. [ DOI ] [ Google Scholar ] 2. Srinivasan S., Bajorath J. Chemical language models for generating compounds with triple-target activity. Cell Rep. Phys. Sci. 2026;7 doi: 10.1016/j.xcrp.2025.103054. [ DOI ] [ Google Scholar ] 3. Srinivasan S., Bajorath J. Protocol to generate dual-target compounds using a transformer chemical language model. STAR Protoc. 2025;6 doi: 10.1016/j.xpro.2024.103584. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 4. Chen H., Bajorath J. Designing highly potent compounds using a chemical language model. Sci. Rep. 2023;13:7412. doi: 10.1038/s41598-023-34683-x. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 5. Vaswani A., Shazeer N., Parmar N., Uszkoreit J., Jones L., Gomez A.N., Kaiser L., Polosukhin I. Attention Is All You Need. 2017. [ DOI ] 6. Weininger D. SMILES, a chemical language and information system. 1. Introduction to methodology and encoding rules. J. Chem. Inf. Comput. Sci. 1988;28:31–36. doi: 10.1021/ci00057a005. [ DOI ] [ Google Scholar ] 7. Bajusz D., Rácz A., Héberger K. Why is Tanimoto index an appropriate choice for fingerprint-based similarity calculations? J. Cheminf. 2015;7:20. doi: 10.1186/s13321-015-0069-3. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 8. Rogers D., Hahn M. Extended-Connectivity Fingerprints. J. Chem. Inf. Model. 2010;50:742–754. doi: 10.1021/ci100050t. [ DOI ] [ PubMed ] [ Google Scholar ] 9. RDKit. https://www.rdkit.org/ . 10. Zdrazil B., Felix E., Hunter F., Manners E.J., Blackshaw J., Corbett S., de Veij M., Ioannidis H., Lopez D.M., Mosquera J.F., et al. The ChEMBL Database in 2023: a drug discovery platform spanning multiple bioactivity data types and time periods. Nucleic Acids Res. 2024;52:D1180–D1192. doi: 10.1093/nar/gkad1004. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Data Availability Statement • The dt_trans code and datasets are available on GitHub at https://github.com/sana1312/dt_trans . • The tt_trans code and datasets are available on GitHub at https://github.com/sana1312/dt_trans/tree/tt_trans . Articles from STAR Protocols are provided here courtesy of Elsevier ACTIONS View on publisher site PDF (3.1 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Record · ID 122797 · SHA-256 9a90196d737efe42
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.