In silico discovery of umami peptides: From sequence-based intelligent models to structure–function mechanistic insight - PMC Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Curr Res Food Sci . 2026 Apr 9;12:101404. doi: 10.1016/j.crfs.2026.101404 Search in PMC Search in PubMed View in NLM Catalog Add to search In silico discovery of umami peptides: From sequence-based intelligent models to structure–function mechanistic insight Xiaolong Li Xiaolong Li a The Clinical Hospital of Chengdu Brain Science Institute, School of Life Science and Technology, University of Electronic Science and Technology of China, Chengdu, 610054, China Find articles by Xiaolong Li a , Xinwei Luo Xinwei Luo a The Clinical Hospital of Chengdu Brain Science Institute, School of Life Science and Technology, University of Electronic Science and Technology of China, Chengdu, 610054, China Find articles by Xinwei Luo a , Sijia Xie Sijia Xie a The Clinical Hospital of Chengdu Brain Science Institute, School of Life Science and Technology, University of Electronic Science and Technology of China, Chengdu, 610054, China Find articles by Sijia Xie a , Yijie Wei Yijie Wei a The Clinical Hospital of Chengdu Brain Science Institute, School of Life Science and Technology, University of Electronic Science and Technology of China, Chengdu, 610054, China Find articles by Yijie Wei a , Feitong Hong Feitong Hong a The Clinical Hospital of Chengdu Brain Science Institute, School of Life Science and Technology, University of Electronic Science and Technology of China, Chengdu, 610054, China Find articles by Feitong Hong a , Yuqing Jiang Yuqing Jiang a The Clinical Hospital of Chengdu Brain Science Institute, School of Life Science and Technology, University of Electronic Science and Technology of China, Chengdu, 610054, China Find articles by Yuqing Jiang a , Xinrui Zhong Xinrui Zhong a The Clinical Hospital of Chengdu Brain Science Institute, School of Life Science and Technology, University of Electronic Science and Technology of China, Chengdu, 610054, China Find articles by Xinrui Zhong a , Xueqin Xie Xueqin Xie a The Clinical Hospital of Chengdu Brain Science Institute, School of Life Science and Technology, University of Electronic Science and Technology of China, Chengdu, 610054, China Find articles by Xueqin Xie a , Caiyi Ma Caiyi Ma b School of Computer Science and Technology, Aba Teachers College, Aba, 623002, China Find articles by Caiyi Ma b , Yuduo Hao Yuduo Hao a The Clinical Hospital of Chengdu Brain Science Institute, School of Life Science and Technology, University of Electronic Science and Technology of China, Chengdu, 610054, China Find articles by Yuduo Hao a , Fuying Dao Fuying Dao c School of Biological Sciences, Nanyang Technological University, Singapore, 639798, Singapore Find articles by Fuying Dao c , Kun Yang Kun Yang a The Clinical Hospital of Chengdu Brain Science Institute, School of Life Science and Technology, University of Electronic Science and Technology of China, Chengdu, 610054, China Find articles by Kun Yang a, ⁎ , Hao Lin Hao Lin a The Clinical Hospital of Chengdu Brain Science Institute, School of Life Science and Technology, University of Electronic Science and Technology of China, Chengdu, 610054, China Find articles by Hao Lin a, ⁎⁎ , Hongyan Lai Hongyan Lai d Chongqing Key Laboratory of Big Data for Bio Intelligence, Chongqing University of Posts and Telecommunications, Chongqing, 400065, China Find articles by Hongyan Lai d, ⁎⁎⁎ , Hao Lyu Hao Lyu a The Clinical Hospital of Chengdu Brain Science Institute, School of Life Science and Technology, University of Electronic Science and Technology of China, Chengdu, 610054, China Find articles by Hao Lyu a, ⁎⁎⁎⁎ Author information Article notes Copyright and License information a The Clinical Hospital of Chengdu Brain Science Institute, School of Life Science and Technology, University of Electronic Science and Technology of China, Chengdu, 610054, China b School of Computer Science and Technology, Aba Teachers College, Aba, 623002, China c School of Biological Sciences, Nanyang Technological University, Singapore, 639798, Singapore d Chongqing Key Laboratory of Big Data for Bio Intelligence, Chongqing University of Posts and Telecommunications, Chongqing, 400065, China ⁎ Corresponding author. [email protected] ⁎⁎ Corresponding author. [email protected] ⁎⁎⁎ Corresponding author. [email protected] ⁎⁎⁎⁎ Corresponding author. [email protected] Received 2026 Jan 29; Revised 2026 Mar 21; Accepted 2026 Apr 8; Collection date 2026. © 2026 The Authors. Published by Elsevier B.V. This is an open access article under the CC BY-NC-ND license (http://creativecommons.org/licenses/by-nc-nd/4.0/). PMC Copyright notice PMCID: PMC13092066 PMID: 42011233 Abstract Umami peptides are naturally occurring bioactive peptides with distinctive umami taste and promising applications in the food industry. However, conventional identification methods based on sensory evaluation, chromatography, and mass spectrometry are time-consuming, costly, and low-throughput. Recent advances in artificial intelligence (AI) have enabled efficient computational frameworks for large-scale umami peptide screening. This review summarizes recent progress in AI-driven umami peptide prediction and mechanistic insights from molecular docking analyses. We overview available umami peptide databases, feature representation strategies, and state-of-the-art AI models, including machine learning, deep learning, and multi-model fusion approaches. In addition, molecular docking studies elucidating the interactions between umami peptides and taste receptors, particularly T1R1/T1R3, are discussed to support rational peptide design. Finally, current challenges and future perspectives in AI-assisted umami peptide research are highlighted, providing guidance for the intelligent development and application of umami peptides in the food industry. Keywords: Umami peptides, Machine learning, Deep learning, Molecular docking, Homology modeling Graphical abstract Open in a new tab Highlights • Comprehensive review of AI-driven umami peptide discovery approaches. • Summarizes key databases enabling large-scale umami peptide screening. • Compares machine learning and deep learning frameworks for accurate umami peptide prediction. • Integrates molecular docking to reveal peptide–receptor interaction mechanisms. • Proposes future AI strategies for rational umami peptide design and validation. 1. Introduction With the deepening concept of healthy food and the growing market demand for natural flavors, umami peptides, short-chain linear peptides with a molecular weight of 150-3000 Da, have become a research focus due to their natural origin and multiple bioactivities, showing significant application potential ( An et al., 2024 ; Zhang et al., 2019 ). Studies have shown that these peptides not only enhance the inherent umami intensity and saltiness of food ( Song et al., 2023 ) but also exhibit various physiological benefits, including antihypertensive, antioxidant, hypoglycemic, and hypolipidemic properties ( Hao et al., 2020 ; Wu et al., 2025 ; Jingcheng Zhang et al., 2024 ; Jincheng Zhang et al., 2023a ). As such, food-derived umami peptides (e.g., glutamine peptides) serve not only as flavoring substrates but also as key ingredients for developing low-sodium foods and health-promoting products, underscoring their substantial research value and wide-ranging application potential in food industry. Traditional screening of umami peptides has primarily relied on structure-activity relationship and frequency analysis of umami motifs ( Wu et al., 2025 ), supported by chromatographic separation, mass spectrometry identification, and sensory evaluation. However, these wet-laboratory approaches are often limited by low throughput, high cost, lengthy procedures, and subjective variability in results. For instance, while Gel Filtration Chromatography (GFC) and Reversed-Phase High-Performance Liquid Chromatography (RP-HPLC) enable peptides purification, they require extensive optimization and are not suited for high-throughput applications ( Gu et al., 2024a ). Sensory evaluation, though critical, is prone to variability due to evaluator experience, preference, and environmental conditions, compromising reproducibility and efficiency ( Bu et al., 2021 ). Furthermore, the structural complexity of umami peptides—where activity depends on amino acid composition, chain length, spatial conformation, and T1R1/T1R3 receptor interactions—poses additional challenges. Traditional methods often lack the capacity to systematically deconvolute these structure-function relationships, hindering in-depth mechanistic insights ( M. Chen et al., 2021a ). Thus, there is a pressing need to shift from these inefficient conventional approaches to precise and high-throughput screening technologies to advance umami peptide discovery. Driven by the necessity to overcome the inherent bottlenecks of conventional wet-lab techniques, artificial intelligence (AI) and computational chemistry have emerged as transformative tools for high-throughput screening ( Mahapatra et al., 2024 ). Unlike traditional empirical methods that heavily rely on physical sample availability and physical trial-and-error, machine learning (ML) and deep learning (DL) algorithms can autonomously extract complex, non-linear structure-activity relationships from massive sequence datasets. By encoding amino acid sequences into high-dimensional numerical features, these AI-driven approaches enable the in silico virtual screening of millions of potential candidates. This fundamentally bypasses the need for laborious chromatographic separations and eliminates subjective human error during initial candidate identification ( Tan et al., 2019 ; Xia et al., 2025 ). Consequently, this paradigm shift from empirical observation to data-driven rational design has led to the establishment of robust computational prediction models, facilitating significant breakthroughs in umami peptide screening efficiency ( Ashikhmina et al., 2024 ; Charoenkwan et al., 2021 ; Charoenkwan et al., 2020 ; Z. Cui et al., 2023a ; Z. Cui et al., 2023b ; Indiran et al., 2024 ; Ji et al., 2025 ; J. Jiang et al., 2023 ; L. Jiang et al., 2022 ; M. Liu et al., 2024 ; Qi et al., 2023 ; Jingcheng Zhang et al., 2023b ). In parallel, while AI excels at large-scale sequence classification, advanced molecular docking and structural simulation technologies complement these models by elucidating the precise binding mechanisms and spatial conformations at the atomic level ( Dang et al., 2019a ; Gu et al., 2024c ; Spaggiari et al., 2020 ). Furthermore, the convergence of predictive AI modeling and mechanistic simulations has shifted the research paradigm from mere sequence-based screening to systematic functional analysis. A consolidated high-throughput workflow combining ML/DL modeling, molecular docking, and sensory evaluation is now widely applied in umami peptide screening (e.g., glutamate peptides ( Li et al., 2023 )), injecting innovative momentum into food science ( Fig. 1 ). However, systematic reviews of these AI-driven works remain scarce, particularly concerning the comparison of computational identification models, the analysis of their strengths and limitations, as well as the application of molecular docking techniques. Fig. 1. Open in a new tab A streamlined in-silico pipeline for the identification and development of umami peptides, key stages including: (i) data curation from dedicated databases; (ii) feature extraction and numerical representation of peptide sequences; (iii) building AI-based prediction models and screening promising candidates; (iv) structure modeling and molecular docking for validating and elucidating interaction mechanism. Created in BioRender. To offer a comprehensive understanding of the latest advancements, along with methodological and practical insights for umami peptide research, we conducted a systematic survey of in-silico approaches for screening and predicting umami peptides, and organized findings into the following four key aspects: (i) overview of current umami peptide databases and their applications; (ii) strategies for umami peptide sequence feature representation; (iii) typical applications of traditional ML and DL algorithms in umami peptide screening, with a focus on evaluating model performance and applicability in terms of feature extraction, prediction accuracy, and generalization capability; and (iv) the complementary role of receptor modeling and molecular docking in deciphering the structure and action mechanism of umami peptides. Moreover, we outlined the remaining challenges, such as inconsistent data quality, limited model generalization, and high experimental validation costs, and a lack of systematic research in key aspects. Future efforts should focus on constructing unified computational frameworks and exploring large-scale language models for discriminative and generative tasks, thereby promoting deeper integration of intelligent algorithms and experimental methods to accelerate umami peptide development and translational application. 2. Current landscape of umami peptide databases High-quality umami peptide databases are essential for enabling efficient AI-based screening and prediction, as they provide the foundational data required for model training, testing, and validation. Currently, dedicated databases specifically for umami peptides remain limited in number and scope. Nevertheless, substantial information on umami peptide is available within broader bioactive peptide databases and taste-related compound repositories ( Table 1 ). Table 1. Overview of key umami peptide databases. Database Scope Number of umami peptide Key features Key limitations URL Reference BIOPEP-UWM Comprehensive bioactive peptides, including sensory peptides 102 Sequence, Source, Activity category (e.g., umami), Reference; Receiving new submission Lack of quantitative sensory data; Updates rely on newly submitted data http://www.uwm.edu.pl/biochemia/index.php/pl/biopep Iwaniak et al. (2016) ChemTastesDB Broad-spectrum taste compounds Including “umami” class, no specified number Nine taste categories, Multi-label annotation, SMILES, InChI, Chemical info; Suitable for multi-taste peptide research Relatively limited peptide data; Lack of quantitative data on umami peptide https://doi.org/10.5281/zenodo.5747393 Rojas et al. (2022) TastePeptidesDB (TastePeptides-Meta) Taste peptide, focus on umami/bitter 2414 taste entries (umami/bitter ∼66.2%), Expanded to 2926 taste entries Sequence, Taste type, Validation status, SMILES, Reference, Update time, Synergistic effect data; Including prediction tool (Umami YYDS) Data distribution is heavily skewed towards umami and bitter; High-precision intensity data remains scarce http://www.tastepeptides-meta.com/ Z. Cui et al. (2025b) UmamiMeta (ummpep2024) Umami peptide 957 umami and 589 non-umami peptides (Filtered) The largest umami/non-umami peptide dataset; Integrated from multi-source data; Providing analysis and prediction tools Negative samples are mainly bitter peptides; Lacking in quantitative intensity data https://hwwlab.com/Webserver/umamimeta He et al. (2025) Open in a new tab 2.1. Comprehensive databases with umami peptide data BIOPEP-UWM is a widely used resource in the bioactive peptide field ( Iwaniak et al., 2016 ), accessible at https://biochemia.uwm.edu.pl/biopep/start_biopep.ph p. This database covers four categories of molecules, including proteins, bioactive peptides, allergenic proteins with their epitopes, as well as sensory peptides and amino acids. It contains a total of 5549 bioactive peptides, 97 of which are umami peptides. BIOPEP-UWM provides detailed annotation information for peptides, such as amino acid sequence, biological source (e.g., food origin, enzymatic hydrolysates), activity type (including umami), and relevant literature references. To support cheminformatics analysis, it also offers essential cheminformatics data, including chemical and monoisotopic masses, SMILES representations, and InChIKeys, along with quantitative activity metrics (e.g., I C 50 or E C 50 values) for quantitative structure-activity relationship (QSAR) modeling studies. ChemTastesDB is a taste compound database with broader coverage that includes both peptides and other small molecule chemicals ( Rojas et al., 2022 ). A distinguishing feature of this database is its granular taste classification system, categorizing compounds into nine taste categories, including umami. The database contains 2944 curated organic and inorganic tastants, 98 of which are classified as umami. Users can access ChemTastesDB data via Zenodo platform ( https://zenodo.org/records/5747393 ). Each entry is annotated with standardized identifiers such as PubChem CID, CAS registry numbers, and canonical SMILES strings, and also includes 3D molecular structures in HyperChem (.hin) format to facilitate structural analysis. Given that peptides often exhibit multiple taste properties, the multi-label taste annotations in ChemTastesDB are especially valuable for analyzing the complexity of peptide flavors and constructing more accurate predictive models. 2.2. Dedicated umami peptide databases As research in taste peptide, particularly umami peptides, progresses, several specialized databases have been developed to systematically collect, manage, and query umami peptide information. These databases provide targeted resources for advancing the understanding of umami peptides. The TastePeptides-Meta platform ( http://www.tastepeptides-meta.com/ ), developed by Cui et al., represents the first open-source multidimensional repository in food sensory science that integrates taste peptides, structural derivatives, and synergistic effects. Its core component, TastePeptidesDB (TPDB), systematically documents peptides with various taste activities, including umami, bitter, sweet, salty, sour, and kokumi ( Z. Cui et al., 2025a ). As of 2025, the TPDB submodule comprises 2414 taste entries, with umami (728 entries) and bitter (871 entries) peptides together accounting for approximately 66.2% of the total. This database provides extensive metadata for each entry, covering amino acid sequences, taste profiles, experimental validation status (e.g., in vitro validation), chemical structure information (e.g., Canonical SMILES), original references, and update records. Additionally, the TastePeptides-Meta platform has expanded its total collection to 2926 peptides, and additionally incorporated 975 peptide structural derivatives and 954 entries on peptide synergistic effects ( Z. Cui et al., 2025b ). The ummpep2024 dataset, a pivotal component of UmamiMeta platform ( https://hwwlab.com/Webserver/umamimeta ), represents the largest specialized umami peptide dataset to date ( He et al., 2025 ). It integrates public data from various sources, including TastePeptidesDB, TastePepMap, BIOPEP-UWM, and specialized projects such as UMPred-FRL and iUmami-SCM, as well as extensive literature mining. Ummpep2024 dataset initially contains 972 umami and 608 non-umami peptides. After excluding sequences exceeding 15 amino acid residues and applying rigorous screening, 957 umami and 589 non-umami peptides are retained for model training. This curated dataset provides a comprehensive and reliable data foundation for machine learning-based umami peptide research. 2.3. Persisting challenges and limitations Despite the expansion of umami peptide-related databases, the availability of high-quality, experimentally validated data remains limited, particularly for quantitative measures such as precise umami intensity or recognition thresholds. This scarcity directly constrains the training effectiveness and generalization ability of complex machine learning models, such as QSAR models aimed at quantitatively predicting umami intensity. The utility of a database hinges on data accuracy and timeliness, yet sustained maintenance requires consistent human and financial resources. As a result, some databases face the risk of update stagnation or even discontinuation ( Minkiewicz et al., 2019 ). Furthermore, discrepancies in data formats, annotation standards (e.g., varying descriptions of umami intensity), and intensity units across different databases hinder effective data integration, cross-dataset/model comparison, and reproducible research. Existing databases also exhibit biases in coverage, often emphasizing umami peptides from specific sources or structural classes, while underrepresenting emerging sources or peptides with atypical structural features. This gap presents a challenge for comprehensive, unbiased peptide discovery and analysis. To circumvent the inherent constraints in data scale and quality discussed above, researchers have been compelled to seek methodological breakthroughs, directly catalyzing a generational evolution in feature representation technologies (detailed in Section 2 ). The scarcity of large-scale annotated data renders models relying on high-dimensional handcrafted features susceptible to overfitting. To circumvent this bottleneck, recent studies have pivoted toward transfer learning and pre-trained dynamic sequence embeddings. By leveraging massive unannotated protein corpora, these approaches extract context-aware semantics, significantly reducing the reliance on labeled umami data. Furthermore, to mitigate the poor generalization caused by structural biases in existing databases, researchers have increasingly integrated 3D structural encodings and docking-derived features. By capturing intrinsic biophysical interactions rather than superficial sequence distributions, these advanced descriptors enhance model robustness when evaluating atypical peptides. Therefore, the evolution of feature representation in umami peptide research is fundamentally driven by the imperative to compensate for existing data deficiencies. 3. Feature representation methods for umami peptides Building upon the challenges of data scarcity and quantitative imbalance exposed in current databases, transforming biological information into highly informative numerical features has become a pivotal strategy for enhancing model generalization. In computational prediction of umami peptides, this fundamental process, referred to as feature representation, is no longer merely a preliminary step but a strategic necessity to compensate for limited experimental data. With the advancement of this field, particularly through the adoption of deep learning techniques, feature encoding strategies have evolved from traditional handcrafted descriptors into a diverse, multi-level system designed to capture latent structural and functional signals ( Fig. 2 ). Fig. 2. Open in a new tab Evolution of feature representation methods for umami peptides. This diagram illustrates the transition from traditional handcrafted feature engineering to modern representation learning. Traditional methods generate peptide-level features, subdivided into composition-based, physicochemical property-based, and structure-based. Modern approaches produce amino acid-level features, categorized into static embeddings and dynamic sequence embeddings. Created in BioRender. Drawing on feature classification schemes in antimicrobial peptide research, current umami peptide features can be broadly grouped into two major categories: peptide-level features and amino acid-level features. The former encode global characteristics of the whole peptide chain, whereas the latter focus on local residues-level information ( Singh et al., 2022 ). Together, they capture complementary aspects of flavor-relevant signals ( Table 2 ). Table 2. Feature representation methods applied in umami peptide prediction. Feature Category Feature Sub-category Specific Feature Method Description Reference Peptide-Level Features Composition-based AAC (amino acid composition) Calculates the occurrence frequency of the 20 standard amino acids. ( Charoenkwan et al., 2021 ; Ji et al., 2025 ; Qi et al., 2023 ) DPC (dipeptide composition) Calculates the occurrence frequency of the 400 possible dipeptides. ( Charoenkwan et al., 2021 ; Qi et al., 2023 ) DDE (dipeptide deviation from expected mean) Quantifies the deviation of dipeptide frequency from the expected random distribution. Qi et al. (2023) BFP (binary profile feature) One-hot encoding for amino acids. Qi et al. (2023) Physicochemical Property-based CTD (composition-transition-distribution) Describes the Composition (C), Transition (T), and Distribution (D) of physicochemical properties (e.g., hydrophobicity, polarity) in the sequence. ( Charoenkwan et al., 2021 ; Ji et al., 2025 ; Qi et al., 2023 ) Pse-AAC (pseudo-amino acid composition) Combines AAC with sequence-order information and physicochemical properties. ( Charoenkwan et al., 2021 ; Ji et al., 2025 ) APAAC (amphiphilic pseudo-amino acid composition) A variant of Pse-AAC focusing on hydrophobicity and hydrophilicity to calculate sequence-order correlations. Charoenkwan et al. (2021) OPF (overlapping property feature) Encodes AAs based on 10 overlapping physicochemical groups. Qi et al. (2023) MDs (molecular descriptors) Calculates global molecular attributes (e.g., MolLogP, VSA, charge) using cheminformatics tools (e.g., RDKit). ( Z. Cui et al., 2023a ; Z. Cui et al., 2023b ) Structure-based FPs (molecular fingerprints) Encodes chemical structure into a vector. Types include: Morgan/ECFP, FCFP, MACCS, Topological, Atom Pairs, Avalon. ( Z. Cui et al., 2023a ; M. Liu et al., 2024 ) DA (docking analysis) Uses docking scores or interacting residues from peptide-receptor (T1R1/T1R3) simulation as features. ( Ashikhmina et al., 2024 ; Z. Cui et al., 2023a ; M. Liu et al., 2024 ) Amino Acid-Level Features Static Embeddings Dynamic Sequence Embeddings Word2vec/FastText Assigns a fixed, context-independent vector to each amino acid (used for comparison). ( Jingcheng Zhang et al., 2023b ) UniRep (mLSTM-based) Encodes sequence into a fixed-length global vector using an mLSTM network pre-trained on large-scale sequences. J. Jiang et al. (2023) BERT Bidirectional transformer encoder generating context-dependent dynamic embeddings for each amino acid. ( L. Jiang et al., 2022 ; Jingcheng Zhang et al., 2023b ) ProtBert A BERT model pre-trained specifically on massive protein sequences (e.g., UniRef100). ( Ashikhmina et al., 2024 ; Indiran et al., 2024 ; Ji et al., 2025 ) ESM-2 A transformer model pre-trained on large-scale protein sequences (e.g., UniRef50). Indiran et al. (2024) Open in a new tab 3.1. Peptide-level features based on traditional feature engineering Peptide-level features are calculated from the entire peptide sequence, integrating information on amino acid composition, physicochemical properties, and structural characteristics. They typically yield fixed-length vectors that provide a global description of each umami peptide. 3.1.1. Representation of amino acid composition Based on the 20 standard amino acids (including A, C, D, E, F, G, H, I, K, L, M, N, P, Q, R, S, T, V, W, Y) and their permutations, composition-based features summarize the frequencies of individual amino acids or amino acid combinations across a peptide, thereby providing a global statistical representation. The most widely used composition-based encodings for umami peptides include Amino acid composition (AAC), dipeptide composition (DPC) ( Charoenkwan et al., 2021 ), dipeptide deviation from expected mean (DDE) and binary profile feature (BFP) ( Wei et al., 2018 ). (1) Amino Acid Composition (AAC) AAC generates a 20-dimensional feature vector by quantifying the frequency of 20 standard amino acids to reflect key physicochemical properties, such as overall charge and hydrophobicity, which dictate the intrinsic interactions between peptides and taste receptors: A A C = N ( i ) L , i ∈ { A , C , D , E , F , G , H , … , Y } (1) where N ( i ) is the count of amino acid i , and L is the sequence length. (2) Dipeptide Composition (DPC) DPC generates a 400-dimensional feature vector by quantifying the frequency of all 400 possible dipeptides to capture local sequence order and spatial proximity. With umami peptides primarily comprising short-chain dipeptides and tripeptides, this characterization of micro-structural motifs is essential for the structural fitting of peptides into the specific recognition pockets of umami receptors. D P C = N ( r , s ) L − 1 , r , s ∈ { A , C , D , E , F , G , H , … , Y } (2) where N ( r , s ) denotes the count of dipeptide ( r , s ) . (3) Dipeptide Deviation from Expected Mean (DDE) DDE measures the deviation between the observed frequency of each dipeptide and its expected frequency under a random codon distribution. It was initially proposed for discriminating linear B-cell epitopes from non-epitopes ( Saravanan and Gautham, 2015 ). The DDE feature vector is derived from the actual dipeptide composition (DC), theoretical mean (TM), and theoretical variance (TV): D C ( r , s ) = N ( r , s ) L − 1 ( r , s ∈ { A , C , D , … , Y } ) (3) T M ( r , s ) = C r C n × C s C n (4) T V ( r , s ) = T M ( r , s ) ( 1 − T M ( r , s ) ) L − 1 (5) here, C r and C s are the numbers of codons encoding amino acids r and s , respectively, and C n = 61 is the total number of sense codons. Finally, the 400-dimensional DDE feature is defined as: D D E ( r , s ) = D C ( r , s ) − T M ( r , s ) T V ( r , s ) (6) This statistical derivation highlights specific dipeptide combinations that manifest at anomalous frequencies relative to random expectations, thereby effectively filtering out background evolutionary noise. Consequently, this approach successfully isolates highly conserved functional motifs that are specifically enriched for triggering umami taste perception. (4) Binary Profile Feature (BFP) BFP encodes each amino acid by a 20-dimensional one-hot vector. Only one element is 1, indicating the identity of the residue, and all others are 0. For instance, amino acid A in the alphabet is encoded as (1, 0, …, 0), C as (0, 1, 0, …, 0), etc. Each residue contributes a 20-dimensional binary vector, thus, an umami peptide of length L will be represented as a L × 20 feature vector. 3.1.2. Descriptors of physicochemical properties This type of features focuses on physicochemical properties of umami peptides. Representative methods include overlapping property feature (OPF), composition-transition-distribution (CTD), pseudo-amino acid composition (Pse-AAC), amphiphilic pseudo-amino acid composition (APAAC) and molecular descriptors (MDs). (1) Overlapping Property Feature (OPF) OPF first divides the 20 standard amino acids into 10 property-based groups, such as Aromatic, Negative, Positive, Polar, Hydrophobic, Aliphatic, Tiny, Charged, Small, and Proline. Because some amino acids possess two or more physicochemical properties, they may be belong to more than one group ( Dou et al., 2014 ; Wei et al., 2017 ). Each amino acid in the peptide sequence is then encoded as a 10-dimensional 0/1 vector, where each element indicates membership in a given group (1 = in group; 0 = not in group) ( Wei et al., 2018 ). (2) Composition-Transition-Distribution (CTD) CTD captures both the composition and spatial distribution of residues with specific physicochemical properties. Originally proposed by Dubchak et al. for protein folding class prediction ( Dubchak et al., 1999 ), CTD consists of three components: Composition (CTDC) calculating the proportion of residues in each property group; Transition (CTDT) measuring the frequency of transitions between different groups along the sequence; Distribution (CTDD) describing positions of residues from each group at specified percentiles of the sequence. The 20 amino acids are first partitioned into groups according to a chosen property (e.g., hydrophobicity, polarity, charge, or secondary structure propensity). For example, for hydrophobicity they may be grouped as hydrophobic (HY), neutral (NE), and hydrophilic (PO). The corresponding formulas are: C ( r ) = N ( r ) L , r ∈ { N E , P O , H Y } (7) where N ( r ) is the total number of amino acids in group r . This equation quantifies the specific fraction of a peptide sequence exhibiting a defined physicochemical property, thus providing a macroscopic overview of its global chemical nature. T ( r , s ) = N ( r , s ) + N ( s , r ) L − 1 , r , s ∈ { ( N E , H Y ) , ( P O , N E ) , ( H Y , P O ) } (8) where N ( r , s ) refers to the number of adjacent residue pairs with group types r and s . This transition metric rigorously characterizes the boundary regions between diverse chemical environments along the peptide chain, effectively modeling the intrinsic folding patterns of the molecule. D ( r ) = ( L ( r , 1 ) L , L ( r , 2 ) L , L ( r , 3 ) L , L ( r , 4 ) L , L ( r , 5 ) L ) , r ∈ { N E , P O , H Y } (9) where L ( r , k ) denotes the sequence position of the k -th percentile (1st, 25th, 50th, 75th, and last occurrence) of residues belonging to group r . Together, CTDC, CTDT, and CTDD provide a global description of the composition, transition frequencies, and spatial distribution patterns of amino acid groups. This distribution function elucidates how specific chemical properties are spatially spread across the peptide sequence, revealing the concentration patterns of crucial functional groups. (3) Pse-AAC and APAAC AAC and DPC ignore sequence-order information. To address this limitation, Pse-AAC and APAAC embed both composition and sequence-order correlations derived from physicochemical properties, generate more informative feature vectors ( Chou, 2001 ). For Pse-AAC, the feature vector has dimension of 20 + λ , where the first 20 components describe normalized amino acid frequencies and the remaining λ components encode sequence-order correlation factors: { X c ( i ) = N i 1 + ω ∑ j = 1 λ θ j X c ( 20 + j ) = ω θ j 1 + ω ∑ j = 1 λ θ j θ j = 1 L − j ∑ i = 1 N − j ( P i − P i + j ) 2 ( i = 1 , 2 , … , 20 ; j = 1 , 2 , … , λ ; λ < L ) (10) here, N i represents the count of the i -th amino acid, ω is a weight factor balancing amino acid composition and order information. θ j represents the j -tier sequence correlation factor, where P i and P i + j denote the normalized physicochemical properties of residues at positions i and i + j , respectively. By uniquely integrating fundamental amino acid composition with long-range sequence order correlations, this formulation generates a robust mathematical proxy for the three-dimensional structural features of the peptide. This analytical capacity is absolutely essential, as it effectively captures the distant chemical interactions indispensable for maintaining the biologically active conformation. APAAC is a variant of Pse-AAC that specifically exploits hydrophobicity and hydrophilicity to capture amphiphilic patterns. The feature vector also has dimension of 20 + λ , with: { X c ( i ) = N i 1 + ω ∑ j = 1 λ τ j X c ( 20 + j ) = ω τ j 1 + ω ∑ j = 1 λ τ j τ j = 1 L − j ∑ i = 1 N − j P a m p h i ( i ) · P a m p h i ( i + j ) ( i = 1 , 2 , … , 20 ; j = 1 , 2 , … , λ ; λ < L ) (11) where P a m p h i ( i ) is an amphiphilicity-related property (e.g., hydrophobicity/hydrophilicity) of the residue at position i . As in Pse-AAC, the first 20 components capture normalized amino acid frequencies, while the remaining λ components encode amphiphilicity-based sequence-order correlations, with ω controlling the relative contribution of composition versus sequence order. Designed to explicitly quantify the amphiphilic character of the peptide, this variant systematically measures the periodic distribution of hydrophobic and hydrophilic residues. Furthermore, this structural periodicity directly dictates the behavior of the peptide at the solvent interface and simultaneously governs its anchoring mechanism into the complex amphipathic environment of the receptor binding site. (4) Molecular Descriptors (MDs) MDs quantify various physicochemical properties of peptides (e.g., molecular weight, LogP, charge distribution) and summarize overall molecular attributes. In umami peptide studies, the open-source cheminformatics toolkit RDKit ( https://www.rdkit.org/ ) is commonly used to compute such descriptors. RDKit supports 2D and 3D molecular operations, descriptor and fingerprint generation, similarity assessment, and molecular visualization. For example, Cui et al. ( Z. Cui et al., 2023b ) calculated 208 RDKit descriptors for a curated umami peptide dataset and applied recursive feature elimination with cross-validation to select 8 informative descriptors as input to their YYDS model. These 8 descriptors were divided into three categories: (i) solubility-related features (including MolLogP and SMR_VSA1), (ii) charge and van der Waals radius-related features (including VSA_EState5, VSA_EState6, VSA_EState7, MinEStateIndex, and PEOE_VSA14), and (iii) molecular weight-related features (e.g., BCUT2D_MWLOW). Their results indicated that solubility-related descriptors such as MR_VSA1 and LogP played a dominant role in distinguishing umami peptides. 3.1.3. Encoding of structural information The umami properties of peptides depend not only on sequence and physicochemical features but are also strongly influenced by 3D structural characteristics. Therefore, transforming complex structural information into computational model-processable numerical features is crucial for umami peptide screening and for elucidating structure-function relationships. Two commonly used structural encodings are molecular fingerprints (FPs) and docking analysis (DA) features. (1) Molecular Fingerprints (FPs) FPs encode molecular structures into fixed-length binary vector or count vectors, enabling efficient comparison of molecular similarity. Five types of FPs commonly used in umami peptide research, including atom pairs fingerprint (APF) ( Carhart et al., 1985 ), morgan fingerprints (MFs) ( Morgan, 1965 ), RDKit topological fingerprint (TP), avalon fingerprint (AF) ( Gedeck et al., 2006 ), and molecular access system (MACCS) fingerprint ( Durant et al., 2002 ), which can all be calculated using RDKit package. APF computes binary fingerprints by analyzing distances and chemical properties between atom pairs, effectively capturing both spatial and chemical characteristics. MFs are circular (extended-connectivity) fingerprints based on molecular topology and a specified radius, they map the SMILES representation of a peptide to a fixed-length binary vector. Common variants include ECFP (atom-invariant-based) and FCFP (feature-invariant-based). TP is a binary fingerprint derived from topological relationships and torsional information, particularly of ring systems, capturing stereochemical and interaction-relevant features. AF enumerates paths and feature classes in the molecular graph, comprehensively encoding atom, bond, and ring information to support accurate predictions. MACCS keys are widely used structural keys based on predefined substructures and functional group fragments, frequently employed in drug similarity analysis and virtual screening. For instance, in the TPDM model, Cui et al. ( Z. Cui et al., 2023b ) systematically compared different fingerprints and found that a specific Morgan fingerprint variant, FCFP4 (with radius of 2), outperformed topological and MACCS fingerprints, achieving superior precision, AUC, accuracy, and recall. (2) Docking Analysis (DA) Molecular docking simulates the interaction between a ligand (e.g., an umami peptide) and a receptor (e.g., T1R1/T1R3), yielding quantitative features (such as binding energies and specific number of bonds) and qualitative information (interaction types and interacting residues). These simulation-derived features can be directly used as inputs for machine learning models and provide biophysical interpretable signals for basic taste molecule recognition. A core feature from docking analysis is the binding affinity or docking score, typically reported as a negative binding energy (e.g., kcal/mol), where larger absolute values indicate stronger binding between the peptide and the receptor. Docking analysis also identifies key receptor residues engaged in hydrogen bonding, hydrophobic contacts and other non-covalent interactions. For example, in the TPDM model developed by Cui et al. ( Z. Cui et al., 2023a ), docking scores were obtained using Smina, a fork of AutoDock Vina optimized for scoring function development and high-performance energy minimization. A single docking score cutoff (e.g., −7.1 kcal/mol for umami) was found to yield limited predictive accuracy (only about 0.548). In contrast, integrating docking scores (e.g., the average of the top N rounds) with interacting-residue features substantially improved model performance. Specifically, using averaged scores and contact statistics from the top 14 rounds produced an optimal model. Moreover, TPDM provided mechanistic insights through two complementary analyses. First, statistical analysis of co-contact residues across the 14 docking rounds identified key hydrogen bond recognition pockets in T1R1, including 107S-109S, 148S-154T, and 247F-249A residues. Second, shapley additive explanation (SHAP) analysis of the trained model highlighted important interacting residues in T1R1/T1R3 and T2R14, linking specific contact patterns to predictive importance. A further development in docking-based feature representation is to abstract the 3D peptide-protein complex into a graph, where atoms or residues are nodes and covalent/non-covalent interactions between them are edges. Graph neural networks (GNNs) can directly process such graph-structured data ( Yin et al., 2024 ). For example, Wang et al., (2021) proposed GNN-DOVE, which extracts the protein interface region, represents it as a graph, and uses atomic chemical properties and inter-atomic distances as node and edge features, respectively, to predict docking quality. This graph-based representation retains rich geometric and topological information and forms a foundation for deeply integrating computational chemistry with deep learning in umami peptide research. 3.2. Amino acid-level features based on representation learning Peptide-level features have long served as the foundation for computational analysis of umami peptides. However, traditional handcrafted descriptors, which focus on global properties, often overlook crucial sequence-order information. To overcome this limitation, current advances in representation learning have shifted toward amino acid-level feature extraction. This approach enhances the granularity of feature representation from entire peptides to individual amino acid residues, enabling more informative and accurate sequence encoding ( Bengio et al., 2013 ; F. Cui et al., 2022 ). This section reviews various representation strategies inspired by natural language processing (NLP) techniques ( Khurana et al., 2023 ), where peptide sequences are analogized to “sentences” and amino acids to “words” for vectorization. We begin with static embeddings, which assign fixed, context-independent vectors to each residue, and proceed to dynamic sequence embeddings that generate context-aware representations. Among dynamic approaches, we compare two core architectures: (i) early recurrent neural network (RNN)-based methods, which capture sequential dependencies and provide a global peptide representation; (ii) more recent Transformer-based models, which produce high-dimensional, context-sensitive representations for each residue. Transformer-based methods currently represent the state-of-the-art in peptide-related tasks, significantly enhancing predictive accuracy and mechanistic interpretability. 3.2.1. Static word embeddings Static embeddings are a classical representation technique derived from NLP, where a fixed, context-independent vector is assigned to each element (each amino acid type in this case) in the vocabulary. Under this paradigm, the representation for a given residue (e.g., Alanine “A”) remains constant, irrespective of its position within the peptide or its adjacent amino acids. Notable methods in this category include Word2vec ( Mikolov et al., 2013 ) and FastText ( Joulin et al., 2016 ). 3.2.2. Dynamic sequence embedding The functional role of an amino acid in umami peptides is highly influenced by its local sequence context. Static embeddings, due to their context-agnostic nature, fail to capture these local dependencies. In contrast, dynamic sequence embeddings generate residue-specific representations that are conditioned on the surrounding sequence environment. This capability significantly improves the modeling of peptide structure-function relationships of peptides and forms the basis for many modern umami peptide prediction methods. Dynamic embedding techniques can be categorized into two main types based on their architecture and output characteristics: RNN-based global sequence embeddings, which capture long-range dependencies across the entire sequence, and Transformer-based contextual residue embeddings, which provide context-sensitive representations for each individual residue. (1) RNN-based Global Sequence Embedding RNNs are specialized architectures for processing sequential data, leveraging internal recurrence to capture temporal (or sequential) dependencies ( X. Chen et al., 2021b ). Long short-term memory (LSTM) networks, an advanced variant of RNNs, address the vanishing gradient problem by incorporating gating mechanisms, such as forget and input gates, allowing them to learn and retain long-range dependencies within sequences. This capability is essential for accurately analyzing peptide functions. The iUmami-DRL model, for instance, employed the Unified Representation (UniRep) model for feature extraction ( J. Jiang et al., 2023 ). UniRep is architecturally based on a multiplicative LSTM (mLSTM) network ( Alley et al., 2019 ) and was pre-trained on 24 million protein sequences from the UniRef50 database. It encodes peptide sequences of arbitrary length into a fixed-length, 1900-dimensional feature vector, effectively summarizing global sequence information into an abstract representation suitable for downstream machine learning classifiers. (2) Transformer-based Contextual Residue Embedding The advent of Transformer architecture has markedly advanced contextual embedding techniques in biological sequence representation learning. Unlike RNNs, Transformer-based models generate dynamic, context-sensitive vector representations for each amino acid by integrating information across the entire sequence context. This residue-level encoding captures not only intrinsic amino acid properties but also local and global inter-residue correlations. Models such as BERT (bidirectional encoder representations from transformers) leverage self-attention mechanisms to capture long-range dependencies and global contextual correlations among residues ( Devlin, 2018 ). Building on this, Jiang et al. and Zhang et al. developed IUP-BERT ( L. Jiang et al., 2022 ) and Umami-BERT ( Jingcheng Zhang et al., 2023b ), respectively. Both models transform umami peptide sequences into 768-dimensional feature vectors but adopt different distinct downstream strategies: IUP-BERT applies feature selection (e.g., LGBM), to reduce dimensionality before employing an SVM classifier, whereas Umami-BERT utilizes an end-to-end approach, feeding the full feature set into an Inception network for classification. In comparative evaluations, BERT-based dynamic representations substantially outperformed static embeddings (e.g., Word2vec, FastText), underscoring their superior capacity to model contextual sequence determinants. Recent research has increasingly shifted toward pre-trained protein-specific language models, leveraging transfer learning from large-scale sequence corpora to acquire biologically meaningful feature representations ( Weiss et al., 2016 ). For example, UmamiPreDL ( Indiran et al., 2024 ) employs ProtBERT (pre-trained on BFD and UniProt) and ESM-2 (pre-trained on UniRef50) to extract residue-level features, which are subsequently pooled into fixed-length representations. Similarly, Umami-gcForest ( Ji et al., 2025 ) and CatBoost-BERT model ( Ashikhmina et al., 2024 ) utilize ProtBERT-derived embeddings, often in combination with handcrafted features (e.g., AAC, CTD, PAAC), to enhance predictive performance. These models benefit from bidirectional pre-training on extensive datasets (e.g., UniRef100, BFD), enabling the embeddings to encapsulate deep structural and functional priors ( Elnaggar et al., 2021 ; Rives et al., 2021 ). Moreover, a key advantage of Transformer models lies in the interpretability afforded by their attention mechanism. Attention weights can quantify residue interdependence and visualize interaction patterns via attention maps. For instance, Umami-BERT ( Jingcheng Zhang et al., 2023b ) analysis revealed that alanine (A), aspartic acid (D), glutamic acid (E), and glycine (G) contribute most significantly to umami taste. Furthermore, the model identified context-dependent motifs, including self-attention among D, E, and G residues, specific D–E interactions decaying with distance, and asymmetric local preferences (such as A's directional coupling with G and acidic residues). These insights not only enhance the understanding of umami peptide function but also demonstrate the interpretability advantages of Transformer-based models. In summary, the evolution from static to dynamic embeddings, the adoption of protein-specific pre-trained models, and the increased interpretability of Transformer-based approaches represent significant advancements in umami peptide feature representation. These NLP-inspired methodologies have not only improved predictive performance but also provided powerful computational tools for decoding the complex mechanisms underlying sequence-function relationships in peptides. 4. AI paradigms for intelligent identification and design of umami peptides Following numerical representation of umami peptide sequences, the construction of high-precision predictive models has become essential for enabling high-throughput screening. The evolution of machine learning models for umami peptides exhibits distinct developmental phases ( Fig. 3 ). Early models were based on shallow statistical features, characterized by simple structures and good interpretability. Subsequently, research shifted toward integrating sophisticated feature engineering with ensemble learning to enhance predictive performance. More recently, deep learning has emerged as a dominant paradigm, facilitating end-to-end learning from raw sequences to prediction outcomes. This section systematically reviews representative models and methodological innovations along this developmental trajectory. Fig. 3. Open in a new tab Evolution of paradigms in ML/DL-based models for umami peptide prediction. This flowchart outlines three developmental stages: Phase I encompasses classical machine learning methods, beginning with early statistical models such as iUmami-SCM and progressing to complex feature engineering and ensemble learning models such as UMPred-FRL and TPDM. Phase II marks the shift to deep learning, including hybrid architectures such as Umami-MRNN, decoupled paradigms such as iUP-BERT, and end-to-end models such as Umami-BERT. Phase III highlights the latest advancements in generative models, such as CatBoost-BERT, which focus on inverse design rather than simple prediction. Created in BioRender. 4.1. Classical machine learning driven by feature engineering Classical machine learning, with its established algorithms, high efficiency, and interpretability, has become a mainstream approach in data-driven umami peptide research. This trajectory began with simple, single-classifier models and progressively evolved towards more sophisticated ensemble learning frameworks. 4.1.1. Linear models based on shallow statistical features Early computational research on umami peptide prediction primarily adopted single-classifier frameworks, focusing on optimizing feature engineering and algorithm selection to extract umami-related patterns from shallow sequence features such as amino acid composition. A representative model, iUmami-SCM ( Charoenkwan et al., 2020 ), pioneered umami peptide prediction using only sequence information. This model employed the scoring card method (SCM), which assigned propensity scores to each dipeptide by statistically comparing their frequencies in umami versus non-umami peptides. A genetic algorithm (GA) was then applied to optimize these scores. During prediction, the weighted scores of all dipeptides in a query sequence are summed and compared against a threshold for classification. A key advantage of this approach is its high interpretability: the scores reflect the contribution of specific dipeptides (e.g., glutamate-rich DE and EE) to umami taste, offering intuitive guidance for peptide screening and design. Despite demonstrating the feasibility of sequence-based prediction, the performance of iUmami-SCM is constrained by limited data and simplistic features, increasing the risk of overfitting and limiting generalizability. Furthermore, such linear frequency-based models struggle to capture complex non-linear relationships between residues, particularly in longer peptides. 4.1.2. Performance leap via multi-dimensional features and ensemble learning To overcome the limitations of early statistical models, the field progressed to a second stage emphasizing multi-dimensional feature engineering and ensemble learning. This phase marked the transition from relying on singular sequence features to employing more sophisticated methods that integrate multiple descriptors and combine various learning models for improved robustness. The Umami_YYDS model exemplifies refined feature engineering ( Z. Cui et al., 2023b ). It utilizes gradient boosting decision tree (GBDT) classifier and addresses class imbalance using SMOTE. The model achieved 89.6% accuracy and an AUC of 98% AUC on its calibration set. Initially, 278 molecular descriptors were extra computed using RDKit toolkit, then rigorously filtered via variance thresholding, statistical testing, and recursive feature elimination with cross-validation (RFECV) based on random forest, resulting in eight key descriptors. SHAP analysis revealed solubility (represented by MolLogP and SMR_VSA1), as the most critical discriminator, highlighting the association between high hydrophilicity and umami taste. Secondary factors included charge properties and van der Waals surface area (e.g., VSA_EState5/6/7, MinEStateIndex, SMR_VSA1, PEOE_VSA14), followed by molecular weight (BCUT2D_MWLOW) consistent with conventional screening criteria for umami peptides ( Gao et al., 2021 ; Happel et al., 2025 ; Yang et al., 2024 ). In ensemble learning, a notable strategy uses predictions from multiple base learners as meta-features to train a meta-classifier. The UMPred-FRL model implements this through feature representation learning (FRL) strategy ( Charoenkwan et al., 2021 ). It constructed 42 base learners by combining seven common feature encoding methods (including AAC, DPC, PAAC, APAAC, CTDC, CTDT, CTDD) with six classical ML classifiers (Extra Trees (ET), K-Nearest Neighbors (KNN), Logistic Regression (LR), Partial Least Squares (PLS), Random Forest (RF), and Support Vector Machine (SVM)). Output probability scores from these learners were integrated into probabilistic features (PF), optimized via a genetic algorithm, and used to train a final SVM meta-predictor. t-SNE visualization confirmed that UMPred-FRL's feature space better separated umami and non-umami peptides, and the model significantly outperformed baseline predictors in accuracy, sensitivity, and Matthews Correlation Coefficient (MCC). The TPDM model further advances ensemble learning by constructing an integrated framework based on multi-modal feature fusion. Specifically, it builds an ensemble model from sub-models trained on three complementary data modalities: DA, MD, and FP ( Z. Cui et al., 2023a ). These modalities capture different aspects of peptide–receptor interactions, including docking-derived binding characteristics, physicochemical properties, and structural topological patterns. Across these feature spaces, nineteen binary classifiers (e.g., k-NN, SVM, RF, XGB, GTB, LR, GNB, SGD) were trained, and the top seven sub-models were selected to construct the final ensemble model using SVM. By integrating physicochemical, structural, and interaction-level information, TPDM provides a more comprehensive representation of peptide–receptor recognition mechanisms. While both UMPred-FRL and TPDM employ SVM as a meta-classifier, they differ in ensemble philosophy: UMPred-FRL integrates model-level predictions from diverse feature–algorithm combinations, whereas TPDM emphasizes the fusion of heterogeneous feature modalities. In this sense, TPDM represents an intermediate methodological stage between traditional feature-engineering-based machine learning and emerging representation-learning approaches. By explicitly integrating multiple complementary feature modalities derived from biological knowledge and computational modeling, TPDM begins to approximate the multi-source information integration commonly observed in modern deep learning frameworks. Although multi-source feature integration is not exclusive to deep learning, the strategy adopted in TPDM reflects an increasing emphasis on integrating heterogeneous biological information to construct richer feature representations. This concept is further advanced in modern deep learning frameworks through automatic representation learning. Thus, TPDM can be viewed as a transitional strategy that extends traditional feature engineering toward more integrated representation paradigms in computational peptide prediction. 4.2. Paradigm shift facilitated by deep learning Although ensemble methods have significantly improved prediction accuracy, they remain highly reliant on manual feature extraction, requiring domain expertise. The advent of deep learning has instigated a fundamental paradigm shift, enabling end-to-end learning directly from raw sequences to prediction results. Current research reflects diverse and progressive modeling strategies, which we categorize into four main paradigms. 4.2.1. Hybrid architecture of manual features and deep learning classifiers In this paradigm, manually engineered peptide-level features are used as inputs to deep learning classifiers, which capture complex, non-linear relationships among the features. This approach combines the interpretability of traditional features with the powerful fitting capability of deep learning. The Umami-MRNN model is a typical example ( Qi et al., 2023 ), combining multi-layer perceptron (MLP) and RNN. It extracts statistical features (AAC, DPC, DDE, CTD), for the MLP branch, and BFP and OPF features for the RNN branch features to capture sequential information. The outputs of both branches are fused via weighted averaging for final classification. This hybrid design leverages both statistical and sequential attributes, achieving an average accuracy of 90.5% on an independent test set. Similarly, the Mlp4Umami module within the VmmScore framework uses an MLP classifier with Morgan fingerprints (computed via RDKit) as input ( M. Liu et al., 2024 ). By directly feeding structural fingerprints into the MLP, this approach avoids complex sequence feature learning and achieved 94% accuracy and precision, outperforming other algorithms. Compared to linear models such as iUmami-SCM, the performance gain in this paradigm stems from the non-linear modeling capacity of deep learning, enabling more comprehensive pattern recognition within handcrafted feature spaces. 4.2.2. Leveraging dynamic embeddings with traditional machine learning classifiers With advances in NLP, the focus of deep learning has expanded from classification to feature extraction. This strategy employs large-scale pre-trained models to encode umami peptide sequences into high-dimensional dynamic embeddings, which are then classified using traditional machine learning algorithms. This decoupling enhances feature generalizability while retaining the interpretability and stability of classical classifiers. For example, the iUmami-DRLF model uses the pre-trained UniRep model to encode umami peptides into 1900-dimensional feature vectors ( J. Jiang et al., 2023 ). After SMOTE balancing and feature selection (ANOVA, LGBM, mutual information), a 177-dimensional feature subset is classified using logistic regression (LR). This approach enriches feature semantics while maintaining model interpretability. Similarly, iUP-BERT ( L. Jiang et al., 2022 ) employs BERT to produce 768-dimensional sequence embeddings, reduces dimensionality to 139 via SMOTE and LGBM, and uses an SVM for umami peptide classification. SVM's robustness on high-dimensional, small-sample data makes it well-suited for processing refined BERT features. Both iUmami-DRLF and iUP-BERT significantly outperformed manual feature-based methods, achieving accuracies of 92.1% and 89.9%, and MCC values of 0.815 and 0.774, respectively. The Umami-gcForest ( Ji et al., 2025 ) model combining traditional (AAC, CTD, PAAC) and dynamic embedding (from ProtBERT) features into a 1338-dimensional vector. After mutual information-based feature selection, the optimized subset is classified using gcForest, a deep forest model with a cascade structure of multi-grained decision trees. In independent testing for umami peptide prediction, this integration achieved 91.07% accuracy and MCC of 0.822 on independent testing data, benefiting from multi-type feature fusion and the cascade structure's ability to learn deep discriminative patterns. 4.2.3. End-to-end learning paradigm The End-to-End learning paradigm achieves synergistic optimization by seamlessly marrying dynamic feature embedding with final classification task within a unified deep learning networks, thus fully harnessing its innate self-learning and hierarchical feature representation capabilities. The Umami-BERT model, proposed by Zhang et al. ( Jingcheng Zhang et al., 2023b ), implements this cutting-edge paradigm using BERT as a feature encoder, pre-trained on 26,000 multi-source peptide sequences to generate 768-dimensional embeddings. Its key innovation is an Inception-based classifier that leverages parallel multi-scale convolutions to capture diverse feature patterns, enhancing the decoding of abstract BERT representations. This model achieved 93.23% accuracy (MCC 0.78) on balanced umami/non-umami data and 95.00% (MCC 0.85) on unbalanced data, outperforming static embedding methods such as Word2Vec and FastText. Similarly, UmamiPreDL compared different combinations of pre-trained models (ProtBert, ESM-2) and classifiers (CNN, DNN), finding that the ProtBert-CNN architecture performed best, achieving 95% and 94% accuracy in 5-fold cross-validation and independent testing, respectively ( Indiran et al., 2024 ). The CNN's local receptive fields effectively identify conserved motifs, while ProtBert's bidirectional encoding supplies rich biological context, together enabling deep integration of feature learning and classification. 4.2.4. Cascaded deep learning models for inverse design While discriminative models excel in predictive classification, they cannot perform inverse design, namely generating sequences from desired functional properties. Cascaded deep learning models address this challenge through multi-stage pipelines that map target taste attributes to peptide sequences. The CatBoost-BERT model exemplifies this emerging direction ( Ashikhmina et al., 2024 ). Its two-stage cascade first uses CatBoost to predict amino acid composition and peptide length distribution of peptides from a target umami intensity (0–100%), generating candidate sequences. The second stage utilizes a fine-tuned ProtBERT model (based on BERT-Large) to regress the umami intensity of these candidates, verifying whether they meet design specifications. With mean squared error (MSE) values of 0.31 (validation) and 0.39 (test), the model demonstrates feasible inverse inference capability. This decoupled strategy simplifies the inverse design problem, though it carries a risk of error propagation between stages. Nevertheless, CatBoost-BERT provides a computationally tractable pathway for the rational umami peptide design. 4.3. Integrated screening strategies via multi-model synergy Recent practical development of umami peptides has moved beyond single-model reliance toward multi-tool, multi-modal integrated screening. This strategy significantly improves prediction efficiency and accuracy, accelerating the discovery of novel umami peptides and deepening insights into their mechanisms and potential functions. For instance, Chen et al. ( H. Chen et al., 2024b ) integrated six prediction tools (e.g., iUmami-SCM, UMPred-FRL, Umami_YYDS, BIOPEP-UWM) to identify candidate peptides from wheat gluten hydrolysates. Combined with bioinformatics tools (ToxinPred, AllerTOP, Innovagen) for safety, allergenicity, and solubility profiling, they screened six novel umami peptides, with QQLPQFEE showing the lowest sensory threshold and EEDQ exhibiting the strongest umami intensity. Sensory experiments further confirmed that these peptides also enhanced saltiness. Similarly, Gu et al., (2024b) applied an ensemble framework (including UMPred-FRL, Umami_YYDS, Umami-MRNN) to screen 627 peptides from Cheddar cheese. After toxicity and allergenicity assessment, six safe umami candidates were identified, among which EKNRLNFLK had the strongest umami taste (threshold: 0.08 mmol/L), while LEEL and DERF exhibited significant umami-enhancing effects. In porcine bone broth studies, Liu et al. ( Q. Liu et al., 2023 ) jointly used UMPred-FRL and Umami-MRNN, achieving 86.7% accuracy. Molecular docking identified nine novel umami peptides (e.g., KKMFETES, threshold: 0.125 mmol/L) that strongly interacted with T1R1/T1R3 receptor residues (e.g., Ser146, His121, Glu277). Furthermore, Fu et al., (2025) combined multiple predictors (iUmami-SCM, Umami_YYDS, TastePeptide-DM) with ToxinPred to screen three novel umami peptides from oyster hydrolysates, including PQFAPEED (threshold: 0.24 mg/mL). Collectively, these studies demonstrate that integrating multiple ML/DL-based predictors with bioinformatics tools establishes an efficient virtual screening workflow for umami peptides. This paradigm enables rapid identification of high-potential candidates from large peptide libraries and significantly enhances experimental validation rates, laying a solid foundation for exploring structure-function relationships. 5. Integrating molecular docking in umami peptide identification and mechanism analysis The growing adoption of machine learning and deep learning for high-throughput screening of umami peptides has intensified the need for robust validating of prediction results and in-depth mechanistic interpretation. As ML/DL-based computational prediction models cannot fully recapitulate complex biological recognition processes, integrating molecular docking to evaluate peptide-receptor interactions from three-dimensional structural perspective has become essential step in the umami peptide research pipeline. 5.1. Theoretical foundations and key components of molecular docking Molecular docking is a computational simulation technique that predicts the optimal binding mode between a ligand and a receptor based on their three-dimensional structures, through conformational sampling and scoring of binding affinity ( Salmaso and Moro, 2018 ; Tao et al., 2020 ). By simulating geometric matching, electrostatic complementarity, and energy matching between molecules, this method provides atomic-level structural perspective on the interaction mechanisms of bioactive peptides with target receptors ( Majid et al., 2022 ; Wen et al., 2020 ; Zhao et al., 2023 ). The increasing application of molecular docking in umami peptide research hinges on a precise understanding of three critical elements: receptor modeling, conformational sampling algorithms, and scoring functions ( Fig. 4 ). Fig. 4. Open in a new tab Conceptual Framework and Core Components of Molecular Docking. This schematic illustrates the three foundational elements of molecular docking: (i) Receptor modeling, often established via homology modeling (e.g., PDB ID: 1EWK ) and subsequently validated through structural quality metrics including Ramachandran plots, ERRAT, and ProSA-web Z-scores; (ii) Conformational sampling to comprehensively explore the ligand (umami peptide) binding landscape by employing systematic, stochastic, or deterministic search algorithms; and (iii) Scoring functions (categorized into force-field-based, empirical, and knowledge-based approaches) used to evaluate binding energies for ranking poses and identifying the optimal binding mode. Created in BioRender. 5.1.1. Umami receptor structure and homology modeling Umami perception is initiated by ligands (such as L-glutamate or umami peptides) bind to specific receptors, inducing conformational changes that activate downstream signaling pathways. Currently, eight potential umami taste receptors have been identified ( Jiang et al., 2022 ), including the T1R1/T1R3 heterodimer, metabotropic glutamate receptors (brain-mGluR1, brain-mGluR4, taste-mGluR1, taste-mGluR4), the extracellular-calcium-sensing receptor (CaSR), and G protein-coupled receptors (GPCR class C: GPRC6A; class A: GPR92/GPR93/LPAR5). Among these, the T1R1/T1R3 heterodimer is considered the primary umami receptor due to their broad responsiveness to amino acids. The Venus Flytrap Domain (VFTD) of the T1R1 subunit has been experimentally confirmed as the principal ligand-binding domain ( Dang et al. 2019a , 2019b ). As the crystal structure of the T1R1/T1R3 complex remains unresolved, homology modeling has become the predominant approach for constructing its 3D structure ( Spaggiari et al., 2020 ). Based on the evolutionary concept that protein structure is more conserved than sequence, this method uses homologous proteins with known structures to model the target receptor. Taking human T1R1 ( NP_619642.2 ) as an example, the modeling process typically involves retrieving its sequence from databases such as NCBI, identifying homologous templates via BLAST searches against structural databases (e.g., the Protein Data Bank, PDB), and model building. The crystal structure of mGluR1 (PDB ID: 1EWK ) is frequently selected as a template due to its coverage of “closed-open/active” conformational transitions. Template selection involves a trade-off: functionally analogous templates (e.g., fish taste receptors) may better capture taste-relevant conformations, while evolutionarily closer templates (e.g., human mGluRs) often yield higher precision in conserved regions. This inherent variability leads to subtle structural differences among T1R1 models across studies, impacting the comparability of docking results and suggesting future re-evaluation may be necessary once experimental structures become available. Following sequence alignment and 3D coordinate construction utilizing specialized software such as MODELLER, rigorous model quality assessment is imperative. Key tools and metrics include Ramachandran plots (assessing residue dihedral angle plausibility), ERRAT program (evaluating non-bonded atomic interactions), and the ProSA-web Z-score (determining overall structural quality). These steps are essential for ensuring the stereochemical rationality and structural reliability of the homology model. Notably, Liu et al. ( M. Liu et al., 2024 ) adopted a multi-receptor strategy, systematically evaluating the binding of 96 peptides to T1R1/T1R3 and five other candidate receptors (GRM1, GRM4, GPRC6A, GPR92, CaSR). Subsequent analysis, integrating conformational optimization (MedusaGraph) and scoring (ITScoreAff), revealed that GPR92 and GRM1 exhibited superior binding affinities compared the canonical T1R1/T1R3 receptor. This finding expands the potential receptor repertoire for umami peptides and offers new directions for screening and mechanistic research. 5.1.2. Classifications and selection strategies of conformational sampling algorithms After a reliable receptor structure is obtained, the main challenge is to efficiently sample the conformational space of the ligand within the binding site. Conformational sampling systematically or stochastically explores ligand translations, rotations, and torsions across the receptor surface or within defined pockets ( Ciemny et al., 2018 ; Ferreira et al., 2015 ; Maia et al., 2020 ). For flexible peptides, the number of possible binding modes grows exponentially, hence, developing efficient sampling algorithms is critical to balance accuracy and computational cost. Conformational sampling algorithms are generally categorized into three types ( Du et al., 2023 ): (i) Systematic search samples all degrees of freedom at fixed intervals. Though theoretically thorough, its cost scales exponentially with flexibility, limiting its use for long peptides (ii) Stochastic search (e.g., Monte Carlo). randomly perturbs ligand conformation and accepts/rejects states based on energy criteria (e.g., Metropolis), offering a good efficiency–accuracy balance. (iii) Deterministic search relies on molecular dynamics simulations. This approach is computationally intensive and sensitive to initial conditions, often trapping in local minima; it is thus frequently combined with other methods for local refinement (e.g., CDOCKER, GOLD, and Glide). In practice, the choice of conformational sampling method should be aligned with research objectives: high-throughput virtual screening typically favors fast, coarse sampling, while mechanistic studies require thorough, fine-grained sampling. For umami peptides, hybrid strategies are often employed to enhance both coverage and efficiency. 5.1.3. Principles and classifications of scoring functions After generating candidate binding poses via conformational sampling, scoring functions are required to evaluate and rank their energetic and geometric compatibility, identifying the pose most likely to represent the physiological binding mode. A scoring function is essentially a mathematical model estimating the binding free energy or a related proxy for the ligand-receptor complex ( Abdelkader and Kim, 2024 ; Gagnon et al., 2016 ; Maia et al., 2020 ; Salmaso and Moro, 2018 ). Lower scores generally indicate more stable conformations. Based on construction principles, scoring functions are classified into three major categories ( Guedes et al., 2018 ): (i) Force-field-based methods use classical molecular mechanics (e.g., CHARMM, AMBER) to compute van der Waals and electrostatic contributions. They have clear physical interpretation but often neglect solvation and entropy, and are computationally expensive. (ii) Empirical scoring functions decompose binding free energy into weighted terms (such as hydrogen bonds, hydrophobics, electrostatics, and rotational entropy penalties) fitted to experimental binding data. Their speed makes them popular in virtual screening. (iii) Knowledge-based potentials derive from statistical analyses of atom–atom contacts in known protein–ligand complexes (from PDB), under the assumption that frequent interactions are energetically favorable. They can capture non-trivial interaction patterns missed by physical models. To mitigate the limitations of single scoring functions, consensus scoring strategies that combine multiple scoring methods are often employed. For example, AutoDock Vina integrates empirical and knowledge-based scores to improve result consistency. Notably, scoring functions are primarily utilized for ranking ligands rather than predicting precise binding affinities, making consensus strategies a robust approach for improving virtual screening reliability. 5.1.4. Recommended strategies for conformational sampling and scoring functions of umami peptides For umami peptides characterized by high conformational flexibility, selecting an appropriate combination of sampling algorithms and scoring functions is crucial. In practical applications, AutoDock Vina is currently the most widely used molecular docking tool ( H. Chen et al., 2024b ; Z. Cui et al., 2023a ; Gu et al., 2024b ). This software integrates an iterated local search algorithm combined with gradient-based optimization and a hybrid scoring function based on empirical and knowledge-based terms ( Trott and Olson, 2010 ). Consequently, this specific combination not only efficiently navigates the massive conformational space of peptides but also provides a reliable ranking of binding affinities, making it highly suitable for the large-scale virtual screening of umami peptides. However, to further improve the accuracy of binding poses and deeply elucidate interaction mechanisms, it is generally recommended to employ additional docking strategies following the initial screening. For instance, peptide-optimized global docking tools, such as HPEPDOCK or CABS-dock ( Kurcinski et al., 2015 ), or high-precision refinement tools utilizing deterministic search algorithms and force-field-based scoring functions, such as FlexPepDock or CDOCKER ( London et al., 2011 ), can better address the local dynamics of peptide backbones and side chains. Furthermore, adopting a consensus docking strategy, which integrates results from distinct docking protocols, is an effective method to eliminate the inherent biases of individual scoring functions and enhance the overall reliability of predictions. Table 3 summarizes the widely used molecular docking tools for flexible short peptides, highlighting their core algorithms and scoring functions. Table 3. Recommended combinations of conformational sampling algorithms and scoring functions for flexible short peptides. Software/Web Server Sampling Algorithm Scoring Function Rationale for Selection in Umami Peptide Research AutoDock Vina Iterated local search with gradient-based optimization (BFGS) Hybrid empirical scoring function Provides an optimal balance between computational speed and accuracy, making it the preferred choice for high-throughput virtual screening of large peptide libraries. CDOCKER Stochastic (Simulated annealing) + Deterministic (MDS) Force-field-based (Soft-core potential) Delivers high-precision structural optimization for flexible ligands, which is highly suitable for refining initial poses and exploring detailed structure-activity mechanisms. HPEPDOCK Systematic search (Hierarchical searching) Consensus (Knowledge-based + Empirical) Specifically optimized for peptide-protein global docking without requiring a predefined active site, capturing the global binding landscape of unknown umami peptides. FlexPepDock Stochastic search (Monte Carlo with energy minimization) Force-field-based (Rosetta score) Enables fully flexible refinement of both the peptide backbone and side chains, yielding near-native, high-resolution complex conformations for mechanistic elucidation. CABS-dock Replica-exchange Monte Carlo with coarse-grained CABS model Knowledge-based statistical potentials Accommodates fully flexible receptors and ligands during global docking, compensating for the high conformational flexibility typically observed in short peptides. Open in a new tab 5.2. Application of molecular docking in mechanism research of umami peptides Serving as a critical bridge between sequence prediction and biophysical reality, molecular docking plays multifaceted roles in validating and mechanistically elucidating umami peptides. By analyzing the lowest-energy complex models, researchers can visualize and quantify the network of non-covalent interactions between umami peptides and key residues within the receptor's binding pocket, revealing the atomic-level structural basis of taste recognition. 5.2.1. Quantitative analysis of interaction patterns Consistent evidence from molecular docking studies shows that umami peptide-receptor complexes are stabilized by multiple non-covalent interaction forces, with hydrogen bonds playing a dominant role. For instance, Fu et al., (2025) reported that hydrogen bonds accounted for 57.89%–68.42% of interactions in oyster-derived umami peptides. Similarly, Gu et al., (2024b) observed hydrogen bond occurrence frequencies of 66.67%–77.78% in peptides from Cheddar cheese, far exceeding other forces. Work on porcine collagen peptides ( Gu et al., 2024c ) further demonstrated that umami-enhancing peptides (e.g., GESMTDGF, DGC, CRD) formed significantly more hydrogen bonds (19–41) than inactive ones (13–14), structurally corroborating their sensory function. Beyond hydrogen bonds, hydrophobic interactions, electrostatic forces (e.g., salt bridges, electrostatic attraction), and van der Waals forces also contribute to the binding stability. Chen et al. ( H. Chen et al., 2024a ) and Fu et al., (2025) both described umami peptide-receptor complex interaction networks involving hydrogen bonds, hydrophobic effects, and electrostatic forces. Yu et al., (2021) and Zhao et al., (2023) specifically highlighted salt bridges between peptide acidic/basic groups and receptor residues (e.g., Arg151, Asp147), underscoring the role of electrostatic interactions. Through surface force analysis, Tang et al., (2024) identified substantial solvent-accessible surface (SAS) matching and multiple short-range non-covalent contacts between Antarctic krill-derived umami peptides and T1R1/T1R3 receptor, confirming the contribution of van der Waals forces. 5.2.2. Systematic identification of key binding sites Molecular docking also enables the precise identification of key amino acid residues mediating umami peptide-receptor binding, revealing conserved binding hotspots. In the T1R1 subunit, residues such as Asp147, Arg151, Ser107, His71, and Gln52 are frequently identified as core binding sites, with Asp147 and Arg151 acting as universal “molecular anchors” ( Yu et al., 2021 ; Zhao et al., 2023 ). In the T1R3 subunit, Ser146, His121/145, Glu277, and Ala302 are repeatedly validated as critical interacting residues ( Chen et al., 2024a ; Q. Liu et al., 2023 ; T. Zhang et al., 2022 ). These highly conserved regions collectively form the structural basis for umami peptide recognition. 5.2.3. Structure-activity relationships (SAR) analysis Molecular docking not only identifies receptor hotspots but also clarifies how peptide structural features dictate binding mode and activity. Studies show that the position and distribution of charged residues within the peptide are critical determinants of binding affinity. For instance, N-terminal acidic and C-terminal basic groups enhance umami taste in short peptides ( Yu et al., 2021 ); a C-terminal arginine (Arg) boosts umami intensity ( Tang et al., 2024 ), and specific acidic residues (e.g., C-terminal Asp) form key electrostatic interactions with the receptor ( Fu et al., 2025 ). These findings consistently demonstrate that peptide amino acid composition and physicochemical properties directly shape interactions with receptor hotspots, influencing taste perception and offering clear guidance for rational umami peptide design. In summary, molecular docking serves not only to validate machine/deep learning predictions but also as a core tool for revealing molecular mechanisms, elucidating SAR, and exploring peptide multifunctionality. It provides a structural biology foundation for the targeted design and functional optimization of umami peptides. 6. Discussion and future perspectives Umami peptides, as natural bioactive compounds, offer significant value for sensory enhancement and physiological regulation. In addition to delivering the essential savory taste, these peptides provide notable health benefits, such as sodium reduction and antihypertensive effects. With their safety and multifunctionality, aligned with the clean-label trends, accelerating the discovery and development of umami peptides has become a priority for the modern food industry. This review systematically summarizes the current applications and recent advances of AI technology in umami peptide research. We begin by evaluating existing umami peptide databases, highlighting both their foundational value and inherent limitations. We then explore diverse feature representation strategies, from conventional descriptors such as amino acid composition and physicochemical properties to deep learning-based dynamic embeddings. Subsequently, we analyze the application and evolution of classical machine learning, advanced deep learning architectures, and their hybrid frameworks in the identification, classification, and functional development of umami peptides. Finally, we discuss the critical role of molecular docking in elucidating atomic-level interaction mechanisms between umami peptides and taste receptors. The computational framework integrating AI prediction models with molecular docking validation has increasingly become the core paradigm for high-throughput screening and rational design of umami peptides. Although deep learning models demonstrate considerable advantages in high-throughput tasks (e.g., processing tens of thousands of peptide sequences) and high precision (with AUC typically >0.9) in umami peptide prediction, several challenges persist in practice. First, the predictive accuracy of machine learning and deep learning models is limited by the scarcity of quantitative sensory data and the insufficient representativeness of negative samples in existing databases. For example, UmamiPreDL is hindered by limited training data, while negative samples for Umami-gcForest are largely derived from bitter peptides, potentially introducing classification bias. Additionally, the inherent “black box” nature of deep learning models complicates their application in elucidating complex structure-activity relationships. While models such as CatBoost-BERT improve accuracy by leveraging high-dimensional features, they face challenges in clarifying the binding mechanisms between umami peptides and T1R1/T1R3 receptors, creating a trade-off between interpretability and predictive power. Furthermore, the reliability of molecular docking is impacted by issues such as the insufficient accuracy of homology-modeled structures and biases in scoring functions when estimating the binding free energy of flexible peptide chains. Future advances in umami peptide research will rely on a multi-layered integration of data, algorithms, and validation to establish a robust, self-improving research closed loop. At the data level, there is an urgent need to build standardized, quantitatively annotated databases. These repositories should include rigorously labeled positive and negative samples, as well as normalized sensory intensity and threshold measures. Furthermore, they must integrate detailed source and processing metadata with experimental outputs aligned with mass spectrometry, chromatography, and receptor activity assays. Consequently, multimodal datasets equipped with confidence metrics are essential to improve model generalizability and reduce class bias. At the algorithmic level, priority should be given to predictive frameworks that fuse heterogeneous information. Such frameworks must combine sequence-derived features, physicochemical properties, and structural representations through multi-modal learning and advanced architectures. Specifically, representation learning based on protein language models, graph neural networks, and structure-aware models is especially promising for capturing the complex determinants of umami activity, transcending the limitations of manual feature engineering. Within this framework, emerging AI tools act as pivotal enablers. For instance, ProtGPT2 and ProGen support the de novo design and high-throughput generation of candidate peptides ( Ferruz et al., 2022 ; Madani et al., 2023 ). Meanwhile, advanced structural predictors such as AlphaFold2 overcome the long-standing limitations of traditional homology modeling ( Jumper et al., 2021 ). They achieve this by directly predicting atomic-resolution receptor models from sequences (for example, taste receptors such as the T1R1/T1R3 heterodimer), while concurrently providing residue-level confidence estimates to guide downstream selection and refinement. Additionally, geometric deep-learning docking methods such as DiffDock and GNINA significantly improve interaction modeling and candidate prioritization ( Corso et al., 2022 ; McNutt et al., 2021 ). In the validation phase, algorithmic predictions should first be screened using structure-aware docking and rapid scoring. Subsequently, these candidates should be refined through physics-based simulations and free-energy estimates. Finally, the top candidates must be validated using receptor activity assays, biophysical binding experiments, and standardized sensory tests. Ultimately, all resulting biophysical and perceptual data should be systematically fed back into the models via active learning or iterative refinement. This comprehensive feedback mechanism completes the design–structure–interaction–validation loop, thereby accelerating rational discovery and industrial translation. CRediT authorship contribution statement Xiaolong Li: Visualization, Writing - original draft, Conceptualization, Investigation, Writing - review & editing. Xinwei Luo, Sijia Xie, Yijie Wei, Feitong Hong, Yuqing Jiang, Xinrui Zhong: Conceptualization, Investigation, Writing - review & editing. Xueqin Xie, Caiyi Ma, Yuduo Hao, Fuying Dao: Visualization, Writing - review & editing. Kun Yang, Hao Lin: Conceptualization, Resources, Supervision, Funding acquisition, Writing - review & editing. Hongyan Lai, Hao Lyu: Conceptualization, Supervision, Writing - review & editing. Funding This work was supported by the National Natural Science Foundation of China (82130112, 62402089), Science and Technology Department of Sichuan Province (2025ZNSFSC1465), China Postdoctoral Science Foundation (2023TQ0047, GZC20230380), and Natural Science Foundation of Chongqing (CSTB2024NSCQ-MSX1241). Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. Handling Editor: Professor Georgios Leontidis Contributor Information Kun Yang, Email: [email protected]. Hao Lin, Email: [email protected]. Hongyan Lai, Email: [email protected]. Hao Lyu, Email: [email protected]. Data availability No data was used for the research described in the article. References Abdelkader G.A., Kim J.-D. Advances in protein-ligand binding affinity prediction via deep learning: a comprehensive study of datasets, data preprocessing techniques, and model architectures. Curr. Drug Targets. 2024;25(15):1041–1065. doi: 10.2174/0113894501330963240905083020. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Alley E.C., Khimulya G., Biswas S., AlQuraishi M., Church G.M. Unified rational protein engineering with sequence-based deep representation learning. Nat. Methods. 2019;16(12):1315–1322. doi: 10.1038/s41592-019-0598-1. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] An J., Wicaksana F., Woo M.W., Liu C., Tian J., Yao Y. Current food processing methods for obtaining umami peptides from protein-rich foods: a review. Trends Food Sci. Technol. 2024;153 [ Google Scholar ] Ashikhmina M.S., Zenkin A.M., Pantiukhin I.S., Litvak I.G., Nesterov P.V., Dutta K.…Orlova O.Y. Uncovering the taste features: applying machine learning and molecular docking approaches to predict umami taste intensity of peptides. Food Biosci. 2024;62 [ Google Scholar ] Bengio Y., Courville A., Vincent P. Representation learning: a review and new perspectives. IEEE Trans. Pattern Anal. Mach. Intell. 2013;35(8):1798–1828. doi: 10.1109/TPAMI.2013.50. [ DOI ] [ PubMed ] [ Google Scholar ] Bu Y., Liu Y., Luan H., Zhu W., Li X., Li J. Characterization and structure–activity relationship of novel umami peptides isolated from Thai fish sauce. Food Funct. 2021;12(11):5027–5037. doi: 10.1039/d0fo03326j. [ DOI ] [ PubMed ] [ Google Scholar ] Carhart R.E., Smith D.H., Venkataraghavan R. Atom pairs as molecular features in structure-activity studies: definition and applications. J. Chem. Inf. Comput. Sci. 1985;25(2):64–73. [ Google Scholar ] Charoenkwan P., Nantasenamat C., Hasan M.M., Moni M.A., Manavalan B., Shoombuatong W. UMPred-FRL: a new approach for accurate prediction of umami peptides using feature representation learning. Int. J. Mol. Sci. 2021;22(23) doi: 10.3390/ijms222313124. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Charoenkwan P., Yana J., Nantasenamat C., Hasan M.M., Shoombuatong W. iUmami-SCM: a novel sequence-based predictor for prediction and analysis of umami peptides using a scoring card method with propensity scores of dipeptides. J. Chem. Inf. Model. 2020;60(12):6666–6678. doi: 10.1021/acs.jcim.0c00707. [ DOI ] [ PubMed ] [ Google Scholar ] Chen H., Zhao H., Li C., Zhou C., Chen J., Xu W.…Luo D. Exploration of bioactive umami peptides from wheat gluten: umami mechanism, antioxidant activity, and potential disease target sites. Foods. 2024;13(23):3805. doi: 10.3390/foods13233805. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Chen H., Zhou C., Jiang G., Yi J., Guan J., Xu M.…Luo D. Screening and characterization of umami peptides from enzymatic and fermented products of wheat gluten using machine learning. LWT. 2024;208 [ Google Scholar ] Chen M., Gao X., Pan D., Xu S., Zhang H., Sun Y.…Dang Y. Taste characteristics and umami mechanism of novel umami peptides and umami-enhancing peptides isolated from the hydrolysates of Sanhuang Chicken. Eur. Food Res. Technol. 2021;247(7):1633–1644. [ Google Scholar ] Chen X., Li C., Bernards M.T., Shi Y., Shao Q., He Y. Sequence-based peptide identification, generation, and property prediction with deep learning: a review. Mol. Syst. Design Eng. 2021;6(6):406–428. [ Google Scholar ] Chou K.C. Prediction of protein cellular attributes using pseudo‐amino acid composition. Proteins: Struct., Funct., Bioinf. 2001;43(3):246–255. doi: 10.1002/prot.1035. [ DOI ] [ PubMed ] [ Google Scholar ] Ciemny M., Kurcinski M., Kamel K., Kolinski A., Alam N., Schueler-Furman O., Kmiecik S. Protein–peptide docking: opportunities and challenges. Drug Discov. Today. 2018;23(8):1530–1537. doi: 10.1016/j.drudis.2018.05.006. [ DOI ] [ PubMed ] [ Google Scholar ] Corso G., Stärk H., Jing B., Barzilay R., Jaakkola T. Diffdock: diffusion steps, twists, and turns for molecular docking. arXiv preprint arXiv:2210.01776. 2022 [ Google Scholar ] Cui F., Li S., Zhang Z., Sui M., Cao C., Hesham A.E.-L., Zou Q. DeepMC-iNABP: deep learning for multiclass identification and classification of nucleic acid-binding proteins. Comput. Struct. Biotechnol. J. 2022;20:2020–2028. doi: 10.1016/j.csbj.2022.04.029. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Cui Z., Qi C., Zhou T., Yu Y., Wang Y., Zhang Z.…Liu Y. Artificial intelligence and food flavor: how AI models are shaping the future and revolutionary technologies for flavor food development. Compr. Rev. Food Sci. Food Saf. 2025;24(1) doi: 10.1111/1541-4337.70068. [ DOI ] [ PubMed ] [ Google Scholar ] Cui Z., Zhang N., Zhou T., Zhou X., Meng H., Yu Y.…Liu Y. Conserved sites and recognition mechanisms of T1R1 and T2R14 receptors revealed by ensemble docking and molecular descriptors and fingerprints combined with machine learning. J. Agric. Food Chem. 2023;71(14):5630–5645. doi: 10.1021/acs.jafc.3c00591. [ DOI ] [ PubMed ] [ Google Scholar ] Cui Z., Zhang Z., Zhou T., Zhou X., Zhang Y., Meng H.…Liu Y. A TastePeptides-Meta system including an umami/bitter classification model Umami_YYDS, a TastePeptidesDB database and an open-source package Auto_Taste_ML. Food Chem. 2023;405 doi: 10.1016/j.foodchem.2022.134812. [ DOI ] [ PubMed ] [ Google Scholar ] Cui Z., Zhou T., Wang S., Blank I., Gu J., Zhang D.…Liu Y. TastePeptides-Meta: a one-stop platform for taste peptides and their structural derivatives, including taste properties, interactions, and prediction models. J. Agric. Food Chem. 2025;73(16):9817–9826. doi: 10.1021/acs.jafc.4c12922. [ DOI ] [ PubMed ] [ Google Scholar ] Dang Y., Hao L., Cao J., Sun Y., Zeng X., Wu Z., Pan D. Molecular docking and simulation of the synergistic effect between umami peptides, monosodium glutamate and taste receptor T1R1/T1R3. Food Chem. 2019;271:697–706. doi: 10.1016/j.foodchem.2018.08.001. [ DOI ] [ PubMed ] [ Google Scholar ] Dang Y., Hao L., Zhou T., Cao J., Sun Y., Pan D. Establishment of new assessment method for the synergistic effect between umami peptides and monosodium glutamate using electronic tongue. Food Res. Int. 2019;121:20–27. doi: 10.1016/j.foodres.2019.03.001. [ DOI ] [ PubMed ] [ Google Scholar ] Devlin J. Bert: pre-training of deep bidirectional transformers for language understanding/arXiv preprint. arXiv preprint arXiv:1810.04805. 2018 [ Google Scholar ] Dou Y., Yao B., Zhang C. PhosphoSVM: prediction of phosphorylation sites by integrating various protein sequence attributes with a support vector machine. Amino Acids. 2014;46(6):1459–1469. doi: 10.1007/s00726-014-1711-5. [ DOI ] [ PubMed ] [ Google Scholar ] Du Z., Comer J., Li Y. Bioinformatics approaches to discovering food-derived bioactive peptides: reviews and perspectives. TrAC, Trends Anal. Chem. 2023;162 [ Google Scholar ] Dubchak I., Muchnik I., Mayor C., Dralyuk I., Kim S.H. Recognition of a protein fold in the context of the SCOP classification. Proteins: Struct., Funct., Bioinf. 1999;35(4):401–407. [ PubMed ] [ Google Scholar ] Durant J.L., Leland B.A., Henry D.R., Nourse J.G. Reoptimization of MDL keys for use in drug discovery. J. Chem. Inf. Comput. Sci. 2002;42(6):1273–1280. doi: 10.1021/ci010132r. [ DOI ] [ PubMed ] [ Google Scholar ] Elnaggar A., Heinzinger M., Dallago C., Rehawi G., Wang Y., Jones L.…Steinegger M. Prottrans: toward understanding the language of life through self-supervised learning. IEEE Trans. Pattern Anal. Mach. Intell. 2021;44(10):7112–7127. doi: 10.1109/TPAMI.2021.3095381. [ DOI ] [ PubMed ] [ Google Scholar ] Ferreira L.G., Dos Santos R.N., Oliva G., Andricopulo A.D. Molecular docking and structure-based drug design strategies. Molecules. 2015;20(7):13384–13421. doi: 10.3390/molecules200713384. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Ferruz N., Schmidt S., Höcker B. ProtGPT2 is a deep unsupervised language model for protein design. Nat. Commun. 2022;13(1):4348. doi: 10.1038/s41467-022-32007-7. Epub 2022/07/28. PubMed PMID: 35896542. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Fu B., Li M., Chang Z., Yi J., Cheng S., Du M. Identification of novel umami peptides from oyster hydrolysate and the mechanisms underlying their taste characteristics using machine learning. Food Chem. 2025;473 doi: 10.1016/j.foodchem.2025.142970. [ DOI ] [ PubMed ] [ Google Scholar ] Gagnon J.K., Law S.M., Brooks III C.L. Flexible CDOCKER: development and application of a pseudo‐explicit structure‐based docking method within CHARMM. J. Comput. Chem. 2016;37(8):753–762. doi: 10.1002/jcc.24259. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Gao J., Fang D., Kimatu B.M., Chen X., Wu X., Du J.…An X. Analysis of umami taste substances of morel mushroom (Morchella sextelata) hydrolysates derived from different enzymatic systems. Food Chem. 2021;362 doi: 10.1016/j.foodchem.2021.130192. [ DOI ] [ PubMed ] [ Google Scholar ] Gedeck P., Rohde B., Bartels C. QSAR− how good is it in practice? Comparison of descriptor sets on an unbiased cross section of corporate data sets. J. Chem. Inf. Model. 2006;46(5):1924–1936. doi: 10.1021/ci050413p. [ DOI ] [ PubMed ] [ Google Scholar ] Gu Y., Niu Y., Zhang J., Sun B., Mao X., Liu Z., Zhang Y. Identification of novel umami peptides from yeast protein through enzymatic, sensory, and in silico approaches. J. Agric. Food Chem. 2024;72(36):20014–20027. doi: 10.1021/acs.jafc.3c08346. [ DOI ] [ PubMed ] [ Google Scholar ] Gu Y., Zhang J., Niu Y., Sun B., Liu Z., Mao X., Zhang Y. Screening and characterization of novel umami peptides in Cheddar cheese using peptidomics and bioinformatics approaches. LWT. 2024;194 [ Google Scholar ] Gu Y., Zhang J., Niu Y., Sun B., Liu Z., Mao X., Zhang Y. Virtual screening and characteristics of novel umami peptides from porcine type I collagen. Food Chem. 2024;434 doi: 10.1016/j.foodchem.2023.137386. [ DOI ] [ PubMed ] [ Google Scholar ] Guedes I.A., Pereira F.S., Dardenne L.E. Empirical scoring functions for structure-based virtual screening: applications, critical aspects, and challenges. Front. Pharmacol. 2018;9:1089. doi: 10.3389/fphar.2018.01089. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Hao L., Gao X., Zhou T., Cao J., Sun Y., Dang Y., Pan D. Angiotensin I-converting enzyme (ACE) inhibitory and antioxidant activity of umami peptides after in vitro gastrointestinal digestion. J. Agric. Food Chem. 2020;68(31):8232–8241. doi: 10.1021/acs.jafc.0c02797. [ DOI ] [ PubMed ] [ Google Scholar ] Happel K., Zeller L., Hammer A.K., Zorn H. Umami enhancing properties of enzymatically hydrolyzed mycelium of Flammulina velutipes cultured on potato pulp. Food Sci. Nutr. 2025;13(4) doi: 10.1002/fsn3.70128. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] He Y., Tian Z., Zheng J., Wang H., Han L., Han W. Prediction of umami peptides based on a large language model of proteins. J. Chem. Inf. Model. 2025;65(8):3955–3962. doi: 10.1021/acs.jcim.4c02394. [ DOI ] [ PubMed ] [ Google Scholar ] Indiran A.P., Fatima H., Chattopadhyay S., Ramadoss S., Radhakrishnan Y. UmamiPreDL: deep learning model for umami taste prediction of peptides using BERT and CNN. Comput. Biol. Chem. 2024;111 doi: 10.1016/j.compbiolchem.2024.108116. [ DOI ] [ PubMed ] [ Google Scholar ] Iwaniak A., Minkiewicz P., Darewicz M., Sieniawski K., Starowicz P. BIOPEP database of sensory peptides and amino acids. Food Res. Int. 2016;85:155–161. doi: 10.1016/j.foodres.2016.04.031. [ DOI ] [ PubMed ] [ Google Scholar ] Ji S., Wu J., An F., Lou M., Zhang T., Guo J.…Wu R. Umami-gcForest: construction of a predictive model for umami peptides based on deep forest. Food Chem. 2025;464 doi: 10.1016/j.foodchem.2024.141826. [ DOI ] [ PubMed ] [ Google Scholar ] Jiang J., Li J., Li J., Pei H., Li M., Zou Q., Lv Z. A machine learning method to identify umami peptide sequences by using multiplicative LSTM embedded features. Foods. 2023;12(7):1498. doi: 10.3390/foods12071498. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Jiang L., Jiang J., Wang X., Zhang Y., Zheng B., Liu S.…Xiang D. IUP-BERT: identification of umami peptides based on BERT features. Foods. 2022;11(22):3742. doi: 10.3390/foods11223742. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Joulin A., Grave E., Bojanowski P., Mikolov T. Bag of tricks for efficient text classification. CoRR abs/1607. 2016 (2016). arXiv preprint arXiv:1607.01759. [ Google Scholar ] Jumper J., Evans R., Pritzel A., Green T., Figurnov M., Ronneberger O.…Potapenko A. Highly accurate protein structure prediction with AlphaFold. Nature. 2021;596(7873):583–589. doi: 10.1038/s41586-021-03819-2. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Khurana D., Koli A., Khatter K., Singh S. Natural language processing: state of the art, current trends and challenges. Multimed. Tool. Appl. 2023;82(3):3713–3744. doi: 10.1007/s11042-022-13428-4. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Kurcinski M., Jamroz M., Blaszczyk M., Kolinski A., Kmiecik S. CABS-dock web server for the flexible docking of peptides to proteins without prior knowledge of the binding site. Nucleic Acids Res. 2015;43(W1):W419–W424. doi: 10.1093/nar/gkv456. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Li C., Yang D., Li L., Wang Y., Chen S., Zhao Y., Lin W. Comparison of the taste mechanisms of umami and bitter peptides from fermented mandarin fish (Chouguiyu) based on molecular docking and electronic tongue technology. Food Funct. 2023;14(21):9671–9680. doi: 10.1039/d3fo02697c. [ DOI ] [ PubMed ] [ Google Scholar ] Liu M., Yang J., He Y., Cao F., Li W., Han W. VmmScore: an umami peptide prediction and receptor matching program based on a deep learning approach. Comput. Biol. Med. 2024;179 doi: 10.1016/j.compbiomed.2024.108814. [ DOI ] [ PubMed ] [ Google Scholar ] Liu Q., Gao X., Pan D., Liu Z., Xiao C., Du L.…Zou Y. Rapid screening based on machine learning and molecular docking of umami peptides from porcine bone. J. Sci. Food Agric. 2023;103(8):3915–3925. doi: 10.1002/jsfa.12319. [ DOI ] [ PubMed ] [ Google Scholar ] London N., Raveh B., Cohen E., Fathi G., Schueler-Furman O. Rosetta FlexPepDock web server—high resolution modeling of peptide–protein interactions. Nucleic Acids Res. 2011;39(Suppl. l_2):W249–W253. doi: 10.1093/nar/gkr431. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Madani A., Krause B., Greene E.R., Subramanian S., Mohr B.P., Holton J.M.…Socher R. Large language models generate functional protein sequences across diverse families. Nat. Biotechnol. 2023;41(8):1099–1106. doi: 10.1038/s41587-022-01618-2. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Mahapatra M., Sahu C., Mohapatra S. Trends of artificial intelligence (AI) use in drug targets, discovery and development: current status and future perspectives. Curr. Drug Targets. 2024 doi: 10.2174/0113894501322734241008163304. [ DOI ] [ PubMed ] [ Google Scholar ] Maia E.H.B., Assis L.C., De Oliveira T.A., Da Silva A.M., Taranto A.G. Structure-based virtual screening: from classical to artificial intelligence. Front. Chem. 2020;8:343. doi: 10.3389/fchem.2020.00343. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Majid A., Lakshmikanth M., Lokanath N., Priyadarshini C.P. Generation, characterization and molecular binding mechanism of novel dipeptidyl peptidase-4 inhibitory peptides from sorghum bicolor seed protein. Food Chem. 2022;369 doi: 10.1016/j.foodchem.2021.130888. [ DOI ] [ PubMed ] [ Google Scholar ] McNutt A.T., Francoeur P., Aggarwal R., Masuda T., Meli R., Ragoza M.…Koes D.R. Gnina 1.0: molecular docking with deep learning. J. Cheminf. 2021;13(1):43. doi: 10.1186/s13321-021-00522-2. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Mikolov T., Sutskever I., Chen K., Corrado G.S., Dean J. Distributed representations of words and phrases and their compositionality. Adv. Neural Inf. Process. Syst. 2013;26 [ Google Scholar ] Minkiewicz P., Iwaniak A., Darewicz M. BIOPEP-UWM database of bioactive peptides: current opportunities. Int. J. Mol. Sci. 2019;20(23):5978. doi: 10.3390/ijms20235978. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Morgan H.L. The generation of a unique machine description for chemical structures-a technique developed at chemical abstracts service. J. Chem. Doc. 1965;5(2):107–113. [ Google Scholar ] Qi L., Du J., Sun Y., Xiong Y., Zhao X., Pan D.…Gao X. Umami-MRNN: deep learning-based prediction of umami peptide using RNN and MLP. Food Chem. 2023;405 [ Google Scholar ] Rives A., Meier J., Sercu T., Goyal S., Lin Z., Liu J.…Ma J. Biological structure and function emerge from scaling unsupervised learning to 250 million protein sequences. Proc. Natl. Acad. Sci. 2021;118(15) doi: 10.1073/pnas.2016239118. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Rojas C., Ballabio D., Sarmiento K.P., Jaramillo E.P., Mendoza M., García F. ChemTastesDB: a curated database of molecular tastants. Food Chem.: Mol. Sci. 2022;4 doi: 10.1016/j.fochms.2022.100090. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Salmaso V., Moro S. Bridging molecular docking to molecular dynamics in exploring ligand-protein recognition process: an overview. Front. Pharmacol. 2018;9:923. doi: 10.3389/fphar.2018.00923. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Saravanan V., Gautham N. Harnessing computational biology for exact linear B-cell epitope prediction: a novel amino acid composition-based feature descriptor. OMICS A J. Integr. Biol. 2015;19(10):648–658. doi: 10.1089/omi.2015.0095. [ DOI ] [ PubMed ] [ Google Scholar ] Singh V., Shrivastava S., Kumar Singh S., Kumar A., Saxena S. StaBle-ABPpred: a stacked ensemble predictor based on biLSTM and attention mechanism for accelerated discovery of antibacterial peptides. Briefings Bioinf. 2022;23(1) doi: 10.1093/bib/bbab439. bbab439. [ DOI ] [ PubMed ] [ Google Scholar ] Song S., Zhuang J., Ma C., Feng T., Yao L., Ho C.-T., Sun M. Identification of novel umami peptides from Boletus edulis and its mechanism via sensory analysis and molecular simulation approaches. Food Chem. 2023;398 doi: 10.1016/j.foodchem.2022.133835. [ DOI ] [ PubMed ] [ Google Scholar ] Spaggiari G., Di Pizio A., Cozzini P. Sweet, umami and bitter taste receptors: state of the art of in silico molecular modeling approaches. Trends Food Sci. Technol. 2020;96:21–29. [ Google Scholar ] Tan J.-X., Lv H., Wang F., Dao F.-Y., Chen W., Ding H. A survey for predicting enzyme family classes using machine learning methods. Curr. Drug Targets. 2019;20(5):540–550. doi: 10.2174/1389450119666181002143355. [ DOI ] [ PubMed ] [ Google Scholar ] Tang W., Feng Y., Tian J., Wang Z., He J., Liu J. Extraction, characterization, and molecular docking study of umami peptides from Antarctic krill. Food Biosci. 2024;61 [ Google Scholar ] Tao X., Huang Y., Wang C., Chen F., Yang L., Ling L.…Chen X. Recent developments in molecular docking technology applied in food science: a review. Int. J. Food Sci. Technol. 2020;55(1):33–45. [ Google Scholar ] Trott O., Olson A.J. AutoDock Vina: improving the speed and accuracy of docking with a new scoring function, efficient optimization, and multithreading. J. Comput. Chem. 2010;31(2):455–461. doi: 10.1002/jcc.21334. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Wang X., Flannery S.T., Kihara D. Protein docking model evaluation by graph neural networks. Front. Mol. Biosci. 2021;8 doi: 10.3389/fmolb.2021.647915. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Wei L., Xing P., Shi G., Ji Z., Zou Q. Fast prediction of protein methylation sites using a sequence-based feature selection technique. IEEE ACM Trans. Comput. Biol. Bioinf. 2017;16(4):1264–1273. doi: 10.1109/TCBB.2017.2670558. [ DOI ] [ PubMed ] [ Google Scholar ] Wei L., Zhou C., Chen H., Song J., Su R. ACPred-FL: a sequence-based predictor using effective feature representation to improve the prediction of anti-cancer peptides. Bioinformatics. 2018;34(23):4007–4016. doi: 10.1093/bioinformatics/bty451. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Weiss K., Khoshgoftaar T.M., Wang D. A survey of transfer learning. J. Big Data. 2016;3(1):9. [ Google Scholar ] Wen C., Zhang J., Zhang H., Duan Y., Ma H. Plant protein-derived antioxidant peptides: isolation, identification, mechanism of action and application in food systems: a review. Trends Food Sci. Technol. 2020;105:308–322. [ Google Scholar ] Wu Y., Shi Y., Qiu Z., Zhang J., Liu T., Shi W., Wang X. Food-derived umami peptides: bioactive ingredients for enhancing flavor. Crit. Rev. Food Sci. Nutr. 2025;65(28):6012–6028. doi: 10.1080/10408398.2024.2435592. [ DOI ] [ PubMed ] [ Google Scholar ] Xia R., Qiao Y., Xu H., Hou Z., Qian G., Wang Y.…Xin G. Unlocking the potential of the umami taste-presenting compounds: a review of the health benefits, metabolic mechanisms and intelligent detection strategies. Food Rev. Int. 2025;41(1):323–343. [ Google Scholar ] Yang Z., Li W., Yang R., Qu L., Piao C., Mu B.…Zhao C. Exploring novel umami peptides from bovine bone soups using Nano-HPLC-MS/MS and molecular docking. Foods. 2024;13(18):2870. doi: 10.3390/foods13182870. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Yin S., Mi X., Shukla D. Leveraging machine learning models for peptide–protein interaction prediction. RSC Chem. Biol. 2024;5(5):401–417. doi: 10.1039/d3cb00208j. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Yu Z., Kang L., Zhao W., Wu S., Ding L., Zheng F.…Li J. Identification of novel umami peptides from myosin via homology modeling and molecular docking. Food Chem. 2021;344 doi: 10.1016/j.foodchem.2020.128728. [ DOI ] [ PubMed ] [ Google Scholar ] Zhang J., He W., Liang L., Sun B., Zhang Y. Study on the saltiness-enhancing mechanism of chicken-derived umami peptides by sensory evaluation and molecular docking to transmembrane channel-like protein 4 (TMC4) Food Res. Int. 2024;182 doi: 10.1016/j.foodres.2024.114139. [ DOI ] [ PubMed ] [ Google Scholar ] Zhang J., Liang L., Zhang L., Zhou X., Sun B., Zhang Y. ACE inhibitory activity and salt-reduction properties of umami peptides from chicken soup. Food Chem. 2023;425 doi: 10.1016/j.foodchem.2023.136480. [ DOI ] [ PubMed ] [ Google Scholar ] Zhang J., Sun-Waterhouse D., Su G., Zhao M. New insight into umami receptor, umami/umami-enhancing peptides and their derivatives: a review. Trends Food Sci. Technol. 2019;88:429–438. [ Google Scholar ] Zhang J., Yan W., Zhang Q., Li Z., Liang L., Zuo M., Zhang Y. Umami-BERT: an interpretable BERT-based model for umami peptides prediction. Food Res. Int. 2023;172 doi: 10.1016/j.foodres.2023.113142. [ DOI ] [ PubMed ] [ Google Scholar ] Zhang T., Hua Y., Zhou C., Xiong Y., Pan D., Liu Z., Dang Y. Umami peptides screened based on peptidomics and virtual screening from Ruditapes philippinarum and Mactra veneriformis clams. Food Chem. 2022;394 doi: 10.1016/j.foodchem.2022.133504. [ DOI ] [ PubMed ] [ Google Scholar ] Zhao W., Su L., Huo S., Yu Z., Li J., Liu J. Virtual screening, molecular docking and identification of umami peptides derived from Oncorhynchus mykiss. Food Sci. Hum. Wellness. 2023;12(1):89–93. [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Data Availability Statement No data was used for the research described in the article. Articles from Current Research in Food Science are provided here courtesy of Elsevier ACTIONS View on publisher site PDF (5.1 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top