ABSTRACT
Abstract
A computer-implemented method for generating a set of candidate variant amino acid sequences of an antibody, a nanobody, or a fragment thereof, having binding ability to a target protein, may comprise: (a) obtaining a set of seed amino acid sequences; and (b) processing the set of seed amino acid sequences using a first trained machine learning algorithm to generate the set of candidate amino acid sequences, wherein the first trained machine learning algorithm is trained with first training data comprising a set of training amino acid sequences for the target protein, wherein the first trained machine learning algorithm is further trained through a transfer learning method using a second trained machine learning algorithm, wherein the second trained machine learning algorithm is trained with second training data comprising a set of training amino acid sequences for a second target protein, wherein the second target protein is different from the target protein.
Description
CROSS-REFERENCE
This application is a continuation of International Application No. PCT/US2022/044754, filed Sep. 26, 2022, which claims the benefit of U.S. Provisional Application No. 63/248,761, filed Sep. 27, 2021, U.S. Provisional Application No. 63/318,037, filed Mar. 9, 2022, U.S. Provisional Application No. 63/332,418, filed Apr. 19, 2022, U.S. Provisional Application No. 63/395,487, filed Aug. 5, 2022, and U.S. Provisional Application No. 63/397,603, filed Aug. 12, 2022, each of which is incorporated by reference herein in its entirety.
BACKGROUND
Antibodies play an important role in therapeutics discovery and vaccine manufacturing due to their ability to bind to one and/or more target antigens. The discovery and generation of antibodies and nanobodies may be performed via in-silico methods.
SUMMARY
The binding site of an antibody/nanobody may include a region referred to as the complementarity-determining region (CDR). Amongst CDRs, the third complementarity-determining region (CDR) of the heavy chain (CDR-H3) of an antibody/nanobody and/or fragment may be the region of highest sequence diversity, which may play a significant role in antigen recognition and dictates binding specificity and affinity. A deep neural network capable of learning from complex and highly variant sequences may be developed to model CDR-H3 sequences and/or to design CDR-H3 sequences mutations. However, increasing the depth of the neural network by adding more layers may lead to challenges such as overfitting and vanishing gradient problems. Moreover, existing deep learning approaches may need to be trained with large training datasets, and may be incapable of discovering and generating new antibodies/nanobodies and/or fragments when facing data scarcity for novel targets. In addition, there are a number of key therapeutics and developability properties to screen for identifying the lead antibody/nanobody candidates. The key therapeutics and developability properties may include but are not limited to antibody specificity to one or multiple targets, cross-reactivity to one or multiple targets, viscosity, clearance, stability, solubility, affinity, affinity maturation, epitope mapping, yield, aggregation, function immunogenicity, humanization, and humaneness. Therefore, there remains a need for tools that can rapidly analyze antibodies/nanobodies and/or fragments for their therapeutic and developability properties, in order to identify the lead candidates, perform analysis of the antibodies/nanobodies for their key biophysical and therapeutic properties, such as specificity and cross-reactivity, and/or repurpose the antibodies/nanobodies for different targets.
In addition, accurate prediction of antibody-antigen binding affinity requires time-consuming and expensive wet-lab experimentations. Therefore, a deep machine-learning method that can rapidly assess the antibody-antigen binding landscape is of significant importance.
Furthermore, besides the CDR-H3, the variable regions (Fv) in antibodies/nanobodies play a significant role in their specificity and binding affinity. However, existing artificial intelligence (AI) methods for antibody discovery mainly focus on antibodies' CDR-H3 region and not on the antibodies' variable fragment (Fv) regions. Thus, a deep machine-learning method that can predict antibody variable regions (Fvs), humanized Fvs, and antibody variable region mutations is of substantial importance.
Moreover, an essential task in antibody therapeutics discovery is to rapidly identify whether an antibody can recognize a specific epitope on an antigen surface and bind to it. Characterization of antibody-paratope:antigen-epitope interaction provides valuable mechanistic insights that can enhance antibody therapeutics' effectiveness. However, other methods may be overly costly and labor-intensive as repetitive wet-lab experimentations are required to sufficiently explore the antibody sequence space. Moreover, due to the extended time and the significant cost associated with those methods, such analysis may be used at the late stage of the discovery and development process and on a limited number of antibody candidates or antigenic variants. This reduces the number of goal shots among the antibody sequences galaxy and increases the risk of failure. Therefore, a deep machine-learning method that can enable an early stage, high-throughput, and efficient method for rapid and accurate antibody analysis, which relies on only amino acid sequences is of critical importance.
Moreover, the design of new therapeutic antibodies requires accurate paratope prediction for optimized antibody paratope-antigen epitope interactions that can be achieved through 3D structure modeling. In-silico antibody discovery for novel targets begins with antibody amino acid sequence data without antibody 3D structure data. The accurate modeling of antigen epitope-antibody paratope requires the model to incorporate the antigen epitope 3D structure to predict the antibody paratope 3D structure. Moreover, besides the backbone residue region in the antibody paratope, the surface side-chain conformation of the antibody paratope plays an important role in the binding interaction of the antibody paratope-antigen epitope. Therefore, the method may need to model both surface side chains and backbone residues in predicting antibody paratope 3D structure. However, the existing artificial intelligence (AI) methods developed for designing new antibody therapeutics do not incorporate the antigen 3D structure and do not model the surface side-chain conformation. Thus, there is a need for rapid and accurate antibody paratope design by the autoregressive generation of antibody paratope sequences while recursively improving antibody paratopes' backbones and side chain structures.
Described herein are systems and techniques for precisely designing and developing antibodies/nanobodies for existing and/or new targets. In some embodiments, the techniques may be used to design and develop antibodies/nanobodies and/or fragments against one target. In other embodiments, the techniques may be used to design and develop antibodies/nanobodies and/or fragments against multiple targets.
In some aspects, the present disclosure provides a deep learning-based antibody design tool. The antibody design tool may include but is not limited to a variational autoencoder (VAE) and residual neural network (Resnet), VAE-Resnet-based algorithm equipped with a transfer learning technique that generates antibody CDR3 sequences (e.g., particularly for novel and difficult targets for which the training data is often scarce), a generative adversarial network (GAN) and reinforcement learning (RL) GAN-RL-based algorithm that generates humanized antibody variable regions, and a conditioned deep learning graph neural network (GNN)-based 3D structure predictor that accurately designs antibodies while predicting antibody paratopes' 3D structures.
In some aspects, the present disclosure provides a deep learning-based antibody analysis tool. The antibody analysis tool may include but is not limited to a VAE-Resnet-based algorithm that reduces the dimensionality of antibodies by extracting the key features, analyzes antibodies for their key biological and key therapeutics (e.g., specificity and cross-reactivity), repurposes existing antibodies for other targets, and improves the performance of clustering models; a deep learning-based epitope mapping algorithm that predicts the epitopes using only amino acid sequences for early stage antibody assessment and for accurately identifying the antibody's mechanism of action; a deep learning-based classifier that accurately classifies the generated antibodies based on their binding ability to a given target; BioOne-Hot preprocessing for improving the antibody classification; and an AI-based automated and multimodal screening method to identify a short list of the best-in-class antibodies for their key therapeutic and biophysical properties.
A deep machine-learning engine is developed and trained using a variety of antibodies/nanobodies known to their targets. In some embodiments, the targets may include but are not limited to antigens. In other embodiments, the deep machine-learning engine may predict antibodies/nanobodies sequences for new target antigens that were and/or were not among the antibodies/nanobodies sequences used for training. In other embodiments, a deep learning model may be used to perform epitope mapping/epitope prediction at the early stage of discovery using only amino acid sequences. In another embodiment, the deep machine-learning engine utilizes a given antigen epitope data to predict the 3D structure of the antibody/nanobody paratopes that bind to the given antigen epitope. The deep machine-learning engine predicts the antibody/ nanobody paratope 3D structure by modeling the antigen epitope-antibody paratope interactions. The predicted antibody/ nanobody paratope 3D structure comprises the backbone residue region and the surface side chains. In other embodiments, the deep machine-learning engine may predict the antibody paratope binding affinity by modeling the antigen epitope-antibody paratope interactions. In other embodiments, the deep machine-learning method may employ transfer-learning techniques to design antibodies/nanobodies for a novel target lacking sufficient training data. In some embodiments, the techniques may be used to design antibodies/nanobodies and/or antibodies/nanobodies mutations and/or to screen their biophysical properties. In another embodiment, the biophysical properties include but are not limited to net-charge, hydrophobicity, viscosity, clearance, solubility, stability, binding affinity, affinity maturation, isoelectric points, specificity, cross-reactivity, immunogenicity, humaneness, humanization, developability, manufacturability, half-life, pharmacokinetic, yield, aggregation, function, and antibody epitope mapping to identify the lead candidates for their therapeutics developability. In other embodiments, the techniques may be used for automated and multimodal screening to identify a short list of the best-in-class antibodies for their key therapeutic and biophysical properties. The deep-learning engine can reduce dimensions in the antibodies/nanobodies sequences data, and/or extracts the key biophysical features and/or group antibodies/nanobodies sequence data into clusters and/or improve the clustering of antibodies/nanobodies sequences and/or recognize hidden patterns across the clusters in their latent spaces and/or repurpose the antibodies/nanobodies sequences recognized from the hidden patterns from one target to be used against multiple targets. In some embodiments, the deep machine-learning engine can recognize antibodies/nanobodies having specificity to one and/or multiple targets from the hidden patterns. In another embodiment, the deep machine-learning engine can identify antibodies/nanobodies having cross-reactivity to one and/or multiple targets from the hidden patterns. In other embodiments, the deep machine-learning engine can map the epitopes of antibodies/nanobodies and/or find out where antibodies/nanobodies bind to their antigens. The deep-learning engine can expand the library of antibody/nanobody lead candidates for one and/or multiple targets. The deep-learning engine can repurpose the antibody/nanobody sequences of one target for another target by finding the hidden patterns across antibody/nanobody sequence libraries.
In an aspect, the present disclosure provides a computer-implemented method for generating a set of candidate variant amino acid sequences of an antibody, a nanobody, or a fragment thereof, having binding ability to a target protein, comprising: (a) obtaining a set of seed amino acid sequences; and (b) processing the set of seed amino acid sequences using a first trained machine learning algorithm to generate the set of candidate variant amino acid sequences, wherein the first trained machine learning algorithm is trained with first training data comprising a set of training amino acid sequences for the target protein, wherein the first trained machine learning algorithm is further trained through a transfer learning method using a second trained machine learning algorithm, wherein the second trained machine learning algorithm is trained with second training data comprising a set of training amino acid sequences for a second target protein, wherein the second target protein is different from the target protein.
In some embodiments, the set of seed amino acid sequences comprises antibody variable regions (Fvs). In some embodiments, the Fvs comprise complementarity determining regions (CDRs) or frameworks (FWRs). In some embodiments, the Fvs comprise CDRs. In some embodiments, the CDRs comprise CDR1, CDR2, or CDR3. In some embodiments, the CDRs comprise CDR3. In some embodiments, the Fvs comprise heavy chains (VH) or light chains (VL). In some embodiments, the Fvs comprise frameworks (FWRs). In some embodiments, the FWRs comprise FWR1, FWR2, FWR3, or FWR4. The method of claim 1 , wherein the target protein comprises at least a portion of a target antigen. In some embodiments, the target antigen comprises at least a portion of SARS-CoV-2 virus.
In some embodiments, the second target protein is selected from the group consisting of Ranibizumab, Trastuzumab, chicken ovalbumin, and yeast display scFV. In some embodiments, the second target protein comprises at least two of Ranibizumab, Trastuzumab, chicken ovalbumin, and yeast display scFV. In some embodiments, the second target protein comprises at least three of Ranibizumab, Trastuzumab, chicken ovalbumin, and yeast display scFV. In some embodiments, the second target protein comprises Ranibizumab, Trastuzumab, chicken ovalbumin, and yeast display scFV.
In some embodiments, the set of training amino acid sequences for the second target protein has a larger number of amino acid sequences than the set of training amino acid sequences for the target protein.
In some embodiments, the method further comprises, prior to (b), pre-processing the set of seed amino acid sequences to have a same sequence length. In some embodiments, the method further comprises, prior to (b), pre-processing the set of seed amino acid sequences at least in part by generating a numerical representation of the set of seed amino acid sequences. In some embodiments, the numerical representation comprises a matrix. In some embodiments, the matrix comprises a one-hot encoding of the set of amino acid sequences. In some embodiments, the matrix comprises a one-hot encoding of amino acid bio-physicochemical values. In some embodiments, the amino acid bio-physicochemical values are selected from isoelectric point, volume, hydrophobicity, solubility, charge, and solvent-accessible surface area (SASA). In some embodiments, the method further comprises, prior to (b), performing an embedding pre-processing on the set of seed amino acid sequences. In some embodiments, the matrix comprises a one-hot encoding of amino acid bio-physicochemical values that improves a classification models (e.g., Decision Tree (DT), Support Vector Machines (SVM), Random Forests (RF), and k-nearest neighbors (KNN)).
In some embodiments, the first trained machine learning algorithm or the second trained machine learning algorithm comprises a deep learning model. In some embodiments, th
CROSS-REFERENCE
This application is a continuation of International Application No. PCT/US2022/044754, filed Sep. 26, 2022, which claims the benefit of U.S. Provisional Application No. 63/248,761, filed Sep. 27, 2021, U.S. Provisional Application No. 63/318,037, filed Mar. 9, 2022, U.S. Provisional Application No. 63/332,418, filed Apr. 19, 2022, U.S. Provisional Application No. 63/395,487, filed Aug. 5, 2022, and U.S. Provisional Application No. 63/397,603, filed Aug. 12, 2022, each of which is incorporated by reference herein in its entirety.
BACKGROUND
Antibodies play an important role in therapeutics discovery and vaccine manufacturing due to their ability to bind to one and/or more target antigens. The discovery and generation of antibodies and nanobodies may be performed via in-silico methods.
SUMMARY
The binding site of an antibody/nanobody may include a region referred to as the complementarity-determining region (CDR). Amongst CDRs, the third complementarity-determining region (CDR) of the heavy chain (CDR-H3) of an antibody/nanobody and/or fragment may be the region of highest sequence diversity, which may play a significant role in antigen recognition and dictates binding specificity and affinity. A deep neural network capable of learning from complex and highly variant sequences may be developed to model CDR-H3 sequences and/or to design CDR-H3 sequences mutations. However, increasing the depth of the neural network by adding more layers may lead to challenges such as overfitting and vanishing gradient problems. Moreover, existing deep learning approaches may need to be trained with large training datasets, and may be incapable of discovering and generating new antibodies/nanobodies and/or fragments when facing data scarcity for novel targets. In addition, there are a number of key therapeutics and developability properties to screen for identifying the lead antibody/nanobody candidates. The key therapeutics and developability properties may include but are not limited to antibody specificity to one or multiple targets, cross-reactivity to one or multiple targets, viscosity, clearance, stability, solubility, affinity, affinity maturation, epitope mapping, yield, aggregation, function immunogenicity, humanization, and humaneness. Therefore, there remains a need for tools that can rapidly analyze antibodies/nanobodies and/or fragments for their therapeutic and developability properties, in order to identify the lead candidates, perform analysis of the antibodies/nanobodies for their key biophysical and therapeutic properties, such as specificity and cross-reactivity, and/or repurpose the antibodies/nanobodies for different targets.
In addition, accurate prediction of antibody-antigen binding affinity requires time-consuming and expensive wet-lab experimentations. Therefore, a deep machine-learning method that can rapidly assess the antibody-antigen binding landscape is of significant importance.
Furthermore, besides the CDR-H3, the variable regions (Fv) in antibodies/nanobodies play a significant role in their specificity and binding affinity. However, existing artificial intelligence (AI) methods for antibody discovery mainly focus on antibodies' CDR-H3 region and not on the antibodies' variable fragment (Fv) regions. Thus, a deep machine-learning method that can predict antibody variable regions (Fvs), humanized Fvs, and antibody variable region mutations is of substantial importance.
Moreover, an essential task in antibody therapeutics discovery is to rapidly identify whether an antibody can recognize a specific epitope on an antigen surface and bind to it. Characterization of antibody-paratope:antigen-epitope interaction provides valuable mechanistic insights that can enhance antibody therapeutics' effectiveness. However, other methods may be overly costly and labor-intensive as repetitive wet-lab experimentations are required to sufficiently explore the antibody sequence space. Moreover, due to the extended time and the significant cost associated with those methods, such analysis may be used at the late stage of the discovery and development process and on a limited number of antibody candidates or antigenic variants. This reduces the number of goal shots among the antibody sequences galaxy and increases the risk of failure. Therefore, a deep machine-learning method that can enable an early stage, high-throughput, and efficient method for rapid and accurate antibody analysis, which relies on only amino acid sequences is of critical importance.
Moreover, the design of new therapeutic antibodies requires accurate paratope prediction for optimized antibody paratope-antigen epitope interactions that can be achieved through 3D structure modeling. In-silico antibody discovery for novel targets begins with antibody amino acid sequence data without antibody 3D structure data. The accurate modeling of antigen epitope-antibody paratope requires the model to incorporate the antigen epitope 3D structure to predict the antibody paratope 3D structure. Moreover, besides the backbone residue region in the antibody paratope, the surface side-chain conformation of the antibody paratope plays an important role in the binding interaction of the antibody paratope-antigen epitope. Therefore, the method may need to model both surface side chains and backbone residues in predicting antibody paratope 3D structure. However, the existing artificial intelligence (AI) methods developed for designing new antibody therapeutics do not incorporate the antigen 3D structure and do not model the surface side-chain conformation. Thus, there is a need for rapid and accurate antibody paratope design by the autoregressive generation of antibody paratope sequences while recursively improving antibody paratopes' backbones and side chain structures.
Described herein are systems and techniques for precisely designing and developing antibodies/nanobodies for existing and/or new targets. In some embodiments, the techniques may be used to design and develop antibodies/nanobodies and/or fragments against one target. In other embodiments, the techniques may be used to design and develop antibodies/nanobodies and/or fragments against multiple targets.
In some aspects, the present disclosure provides a deep learning-based antibody design tool. The antibody design tool may include but is not limited to a variational autoencoder (VAE) and residual neural network (Resnet), VAE-Resnet-based algorithm equipped with a transfer learning technique that generates antibody CDR3 sequences (e.g., particularly for novel and difficult targets for which the training data is often scarce), a generative adversarial network (GAN) and reinforcement learning (RL) GAN-RL-based algorithm that generates humanized antibody variable regions, and a conditioned deep learning graph neural network (GNN)-based 3D structure predictor that accurately designs antibodies while predicting antibody paratopes' 3D structures.
In some aspects, the present disclosure provides a deep learning-based antibody analysis tool. The antibody analysis tool may include but is not limited to a VAE-Resnet-based algorithm that reduces the dimensionality of antibodies by extracting the key features, analyzes antibodies for their key biological and key therapeutics (e.g., specificity and cross-reactivity), repurposes existing antibodies for other targets, and improves the performance of clustering models; a deep learning-based epitope mapping algorithm that predicts the epitopes using only amino acid sequences for early stage antibody assessment and for accurately identifying the antibody's mechanism of action; a deep learning-based classifier that accurately classifies the generated antibodies based on their binding ability to a given target; BioOne-Hot preprocessing for improving the antibody classification; and an AI-based automated and multimodal screening method to identify a short list of the best-in-class antibodies for their key therapeutic and biophysical properties.
A deep machine-learning engine is developed and trained using a variety of antibodies/nanobodies known to their targets. In some embodiments, the targets may include but are not limited to antigens. In other embodiments, the deep machine-learning engine may predict antibodies/nanobodies sequences for new target antigens that were and/or were not among the antibodies/nanobodies sequences used for training. In other embodiments, a deep learning model may be used to perform epitope mapping/epitope prediction at the early stage of discovery using only amino acid sequences. In another embodiment, the deep machine-learning engine utilizes a given antigen epitope data to predict the 3D structure of the antibody/nanobody paratopes that bind to the given antigen epitope. The deep machine-learning engine predicts the antibody/ nanobody paratope 3D structure by modeling the antigen epitope-antibody paratope interactions. The predicted antibody/ nanobody paratope 3D structure comprises the backbone residue region and the surface side chains. In other embodiments, the deep machine-learning engine may predict the antibody paratope binding affinity by modeling the antigen epitope-antibody paratope interactions. In other embodiments, the deep machine-learning method may employ transfer-learning techniques to design antibodies/nanobodies for a novel target lacking sufficient training data. In some embodiments, the techniques may be used to design antibodies/nanobodies and/or antibodies/nanobodies mutations and/or to screen their biophysical properties. In another embodiment, the biophysical properties include but are not limited to net-charge, hydrophobicity, viscosity, clearance, solubility, stability, binding affinity, affinity maturation, isoelectric points, specificity, cross-reactivity, immunogenicity, humaneness, humanization, developability, manufacturability, half-life, pharmacokinetic, yield, aggregation, function, and antibody epitope mapping to identify the lead candidates for their therapeutics developability. In other embodiments, the techniques may be used for automated and multimodal screening to identify a short list of the best-in-class antibodies for their key therapeutic and biophysical properties. The deep-learning engine can reduce dimensions in the antibodies/nanobodies sequences data, and/or extracts the key biophysical features and/or group antibodies/nanobodies sequence data into clusters and/or improve the clustering of antibodies/nanobodies sequences and/or recognize hidden patterns across the clusters in their latent spaces and/or repurpose the antibodies/nanobodies sequences recognized from the hidden patterns from one target to be used against multiple targets. In some embodiments, the deep machine-learning engine can recognize antibodies/nanobodies having specificity to one and/or multiple targets from the hidden patterns. In another embodiment, the deep machine-learning engine can identify antibodies/nanobodies having cross-reactivity to one and/or multiple targets from the hidden patterns. In other embodiments, the deep machine-learning engine can map the epitopes of antibodies/nanobodies and/or find out where antibodies/nanobodies bind to their antigens. The deep-learning engine can expand the library of antibody/nanobody lead candidates for one and/or multiple targets. The deep-learning engine can repurpose the antibody/nanobody sequences of one target for another target by finding the hidden patterns across antibody/nanobody sequence libraries.
In an aspect, the present disclosure provides a computer-implemented method for generating a set of candidate variant amino acid sequences of an antibody, a nanobody, or a fragment thereof, having binding ability to a target protein, comprising: (a) obtaining a set of seed amino acid sequences; and (b) processing the set of seed amino acid sequences using a first trained machine learning algorithm to generate the set of candidate variant amino acid sequences, wherein the first trained machine learning algorithm is trained with first training data comprising a set of training amino acid sequences for the target protein, wherein the first trained machine learning algorithm is further trained through a transfer learning method using a second trained machine learning algorithm, wherein the second trained machine learning algorithm is trained with second training data comprising a set of training amino acid sequences for a second target protein, wherein the second target protein is different from the target protein.
In some embodiments, the set of seed amino acid sequences comprises antibody variable regions (Fvs). In some embodiments, the Fvs comprise complementarity determining regions (CDRs) or frameworks (FWRs). In some embodiments, the Fvs comprise CDRs. In some embodiments, the CDRs comprise CDR1, CDR2, or CDR3. In some embodiments, the CDRs comprise CDR3. In some embodiments, the Fvs comprise heavy chains (VH) or light chains (VL). In some embodiments, the Fvs comprise frameworks (FWRs). In some embodiments, the FWRs comprise FWR1, FWR2, FWR3, or FWR4. The method of claim 1 , wherein the target protein comprises at least a portion of a target antigen. In some embodiments, the target antigen comprises at least a portion of SARS-CoV-2 virus.
In some embodiments, the second target protein is selected from the group consisting of Ranibizumab, Trastuzumab, chicken ovalbumin, and yeast display scFV. In some embodiments, the second target protein comprises at least two of Ranibizumab, Trastuzumab, chicken ovalbumin, and yeast display scFV. In some embodiments, the second target protein comprises at least three of Ranibizumab, Trastuzumab, chicken ovalbumin, and yeast display scFV. In some embodiments, the second target protein comprises Ranibizumab, Trastuzumab, chicken ovalbumin, and yeast display scFV.
In some embodiments, the set of training amino acid sequences for the second target protein has a larger number of amino acid sequences than the set of training amino acid sequences for the target protein.
In some embodiments, the method further comprises, prior to (b), pre-processing the set of seed amino acid sequences to have a same sequence length. In some embodiments, the method further comprises, prior to (b), pre-processing the set of seed amino acid sequences at least in part by generating a numerical representation of the set of seed amino acid sequences. In some embodiments, the numerical representation comprises a matrix. In some embodiments, the matrix comprises a one-hot encoding of the set of amino acid sequences. In some embodiments, the matrix comprises a one-hot encoding of amino acid bio-physicochemical values. In some embodiments, the amino acid bio-physicochemical values are selected from isoelectric point, volume, hydrophobicity, solubility, charge, and solvent-accessible surface area (SASA). In some embodiments, the method further comprises, prior to (b), performing an embedding pre-processing on the set of seed amino acid sequences. In some embodiments, the matrix comprises a one-hot encoding of amino acid bio-physicochemical values that improves a classification models (e.g., Decision Tree (DT), Support Vector Machines (SVM), Random Forests (RF), and k-nearest neighbors (KNN)).
In some embodiments, the first trained machine learning algorithm or the second trained machine learning algorithm comprises a deep learning model. In some embodiments, the deep learning model comprises a deep generative model. In some embodiments, the deep generative model comprises at least one of an Autoencoder (AE), a Variational Autoencoder (VAE) and a Residual Neural Network (Resnet), a Generative Adversarial Network (GAN) and a Reinforcement Learning (RL), and a Graph Neural Network (GNN). In some embodiments, the deep generative model comprises at least two of the GAN, the AE, the VAE, the Resnet, the RL and the GNN. In some embodiments, the deep generative model comprises at least three of the GAN, the AE, the VAE, the Resnet, the RL, and the GNN. In some embodiments, the deep generative model comprises the GAN, the VAE, the Resnet, the RL, and the GNN.
In some embodiments, the first trained machine learning algorithm or the second trained machine learning algorithm comprises an encoder and a decoder. In some embodiments, the encoder comprises at least one of residual layers, convolutional layers, and dense layers. In some embodiments, the encoder comprises at least two of residual layers, convolutional layers, and dense layers. In some embodiments, the encoder comprises residual layers, convolutional layers, and dense layers.
In some embodiments, the method further comprises performing, using the encoder, a dimensionality reduction of the set of seed amino acid sequences. In some embodiments, the dimensionality reduction comprises generating a two-dimensional feature space. In some embodiments, the method further comprises improving the clustering features of the set of seed amino acid sequences based on sequence similarity or sequence diversity. In some embodiments, the clustering comprises using K-means or Gaussian mixture model (GMM).
In some embodiments, the method further comprises repurposing of the set of seed amino acid sequences. In some embodiments, the repurposing comprises varying the set of seed amino acid sequences that has binding to one target to have binding to an additional target.
In some embodiments, the dimensionality reduction comprises principal component analysis (PCA), t-distributed Stochastic Neighbor Embedding (t-SNE), uniform manifold approximation and projection (UMAP), variational autoencoder (VAE), or single-cell Consensus Clusters of Encoded Subspaces (scCCESS). In some embodiments, the method further comprises reconstructing, using the decoder, an input sequence from the two-dimensional feature space. In some embodiments, the decoder comprises at least one of convolutional layers and convolutional transpose layers. In some embodiments, the decoder comprises convolutional layers and convolutional transpose layers.
In some embodiments, the first trained machine learning algorithm and the second trained machine learning algorithm are trained without use of structural data.
In some embodiments, the method further comprises processing the set of candidate variant amino acid sequences to determine a therapeutic property selected from single specificity, multiple specificity, single cross-reactivity, and multiple cross-reactivity.
In some embodiments, the method further comprises processing the set of candidate variant amino acid sequences to determine molecular characteristics of the set of candidate variant amino acid sequences. In some embodiments, the molecular characteristics comprise at least one of amino acid distribution, amino acid frequency, amino acid length distribution, sequence similarity, sequence diversity, total charge distribution, hydrophobicity, net charge, stability, solubility, isoelectric point, solvent accessible surface area (SASA), binding affinity, epitope-paratrope interaction, epitope prediction, humanization, humaneness, developability, manufacturability, half-life, pharmacokinetic, yield, aggregation, function, clearance, viscosity, immunogenicity, single-target specificity, multi-target specificity, single-target cross-reactivity, and multi-target cross-reactivity. In some embodiments, the method further comprises determining sequence similarity between the set of candidate variant amino acid sequences and the set of seed amino acid sequences. In some embodiments, the method further comprises determining the sequence similarity using at least one of a Bilingual Evaluation Understudy (BLEU) method, a Jensen-Shannon divergence (JSD) method, and a Needleman-Wunch (NW) method. In some embodiments, the method further comprises determining sequence diversity between the set of candidate variant amino acid sequences and the set of seed amino acid sequences. In some embodiments, the method further comprises determining the sequence diversity at least in part by measuring a number of shared n-grams for one or more values of n. In some embodiments, the shared n-grams comprise 2-grams, 3-grams, and 4-grams.
In some embodiments, the method further comprises determining specificity or binding affinity of the set of candidate variant amino acid sequences to the target protein.
In some embodiments, the method further comprises filtering out at least one of the set of candidate variant amino acid sequences having undesired molecular characteristics. In some embodiments, the method further comprises selecting or ranking at least one of the set of candidate variant amino acid sequences based at least in part on having desired molecular characteristics. In some embodiments, the method further comprises validating at least one of the selected or ranked candidate variant amino acid sequences.
In some embodiments, the set of candidate variant amino acid sequences correspond to an antibody, a nanobody, or a fragment thereof. In some embodiments, the antibody is a monoclonal antibody, a bispecific antibody, or a trispecific antibody.
In some embodiments, the method further comprises determining a three-dimensional structure of the antibody, the nanobody, or the fragment thereof, based at least in part on the set of candidate variant amino acid sequences.
In some embodiments, the method further comprises predicting a humanized antibody, nanobody, or fragment variable region based at least in part on the set of candidate variant amino acid sequences.
In some embodiments, the method further comprises performing a deep learning-based classification to group the set of candidate variant amino acid sequences having binding or specificity to a same target. In some embodiments, the deep learning-based classification comprises a deep neural network (DNN). In some embodiments, the method further comprises performing automated and multimodal screening of the set of candidate variant amino acid sequences. In some embodiments, the automated and multimodal screening comprises using an ensemble model to identify a shortlist of antibodies, nanobodies, or fragments thereof with desirable therapeutics properties.
In another aspect, the present disclosure provides a computer-implemented method for generating a set of candidate variant amino acid sequences mutations of an antibody, a nanobody, or a fragment thereof, having binding ability to a target protein, comprising: (a) obtaining a set of seed amino acid sequences; and (b) processing the set of seed amino acid sequences using a trained machine learning algorithm to generate the set of candidate variant amino acid sequences, wherein the trained machine learning algorithm comprises at least one of a generative adversarial network (GAN) and Reinforcement Learning (RL).
In some embodiments, the trained machine learning algorithm is trained with training data comprising a set of training amino acid sequences for the target protein. In some embodiments, the set of training amino acid sequences comprises antibody variable regions (Fvs). In some embodiments, the Fvs comprise complementarity determining regions (CDRs) or frameworks (FWRs). In some embodiments, the Fvs comprise CDRs. In some embodiments, the CDRs comprise CDR1, CDR2, or CDR3. In some embodiments, the CDRs comprise CDR3. In some embodiments, the Fvs comprise heavy chains (VH) or light chains (VL). In some embodiments, the Fvs comprise frameworks (FWRs). In some embodiments, the FRs comprise FWR1, FWR2, FWR3, or FWR4. In some embodiments, the target protein comprises at least a portion of a target antigen. In some embodiments, the target antigen comprises at least a portion of SARS-CoV-2 virus.
In some embodiments, the second target protein is selected from the group consisting of Ranibizumab, Trastuzumab, chicken ovalbumin, and yeast display scFV. In some embodiments, the second target protein comprises at least two of Ranibizumab, Trastuzumab, chicken ovalbumin, and yeast display scFV. In some embodiments, the second target protein comprises at least three of Ranibizumab, Trastuzumab, chicken ovalbumin, and yeast display scFV. In some embodiments, the second target protein comprises Ranibizumab, Trastuzumab, chicken ovalbumin, and yeast display scFV.
In some embodiments, the method further comprises, prior to (b), pre-processing the set of seed amino acid sequences to have a same sequence length. In some embodiments, the method further comprises, prior to (b), pre-processing the set of seed amino acid sequences at least in part by generating a numerical representation of the set of seed amino acid sequences. In some embodiments, the numerical representation comprises a matrix. In some embodiments, the matrix comprises a one-hot encoding of the set of amino acid sequences. In some embodiments, the matrix comprises a one-hot encoding of amino acid bio-physicochemical values. In some embodiments, the amino acid bio-physicochemical values are selected from isoelectric point, volume, hydrophobicity, solubility, charge, and solvent-accessible surface area (SASA). In some embodiments, the method further comprises, prior to (b), performing an embedding pre-processing on the set of seed amino acid sequences.
In some embodiments, the trained machine learning algorithm comprises the GAN and the RL. In some embodiments, the trained machine learning algorithm is trained without use of structural data.
In some embodiments, the method further comprises processing the set of candidate variant amino acid sequences to determine a therapeutic property selected from single specificity, multiple specificity, single cross-reactivity, and multiple cross-reactivity.
In some embodiments, the method further comprises processing the set of candidate variant amino acid sequences to determine molecular characteristics of the set of candidate variant amino acid sequences. In some embodiments, the molecular characteristics comprise at least one of amino acid distribution, amino acid frequency, amino acid length distribution, sequence similarity, sequence diversity, total charge distribution, hydrophobicity, net charge, stability, solubility, isoelectric point, solvent accessible surface area (SASA), binding affinity, epitope-paratrope interaction, epitope prediction, humanization, humaneness, developability, manufacturability, half-life, pharmacokinetic, yield, aggregation, function, clearance, viscosity, immunogenicity, single-target specificity, multi-target specificity, single-target cross-reactivity, and multi-target cross-reactivity. In some embodiments, the method further comprises determining sequence similarity between the set of candidate variant amino acid sequences and the set of seed amino acid sequences. In some embodiments, the method further comprises determining the sequence similarity using at least one of a Bilingual Evaluation Understudy (BLEU) method, a Jensen-Shannon divergence (JSD) method, and a Needleman-Wunch (NW) method. In some embodiments, the method further comprises determining sequence diversity between the set of candidate variant amino acid sequences and the set of seed amino acid sequences. In some embodiments, the method further comprises determining the sequence diversity at least in part by measuring a number of shared n-grams for one or more values of n. In some embodiments, the shared n-grams comprise 2-grams, 3-grams, and 4-grams.
In some embodiments, the method further comprises determining specificity or binding affinity of the set of candidate variant amino acid sequences to the target protein.
In some embodiments, the method further comprises filtering out at least one of the set of candidate variant amino acid sequences having undesired molecular characteristics. In some embodiments, the method further comprises selecting or ranking at least one of the set of candidate variant amino acid sequences based at least in part on having desired molecular characteristics. In some embodiments, the method further comprises validating at least one of the selected or ranked candidate variant amino acid sequences.
In some embodiments, the set of candidate variant amino acid sequences correspond to an antibody, a nanobody, or a fragment thereof. In some embodiments, the antibody is a monoclonal antibody, a bispecific antibody, or a trispecific antibody.
In some embodiments, the method further comprises predicting an epitope based at least in part on the set of candidate variant amino acid sequences.
In some embodiments, the method further comprises determining a three-dimensional structure of the antibody, the nanobody, or the fragment thereof, based at least in part on the set of candidate variant amino acid sequences.
In some embodiments, the method further comprises predicting a humanized antibody variable region based at least in part on the set of candidate variant amino acid sequences.
In some embodiments, the method further comprises performing a deep learning-based classification to group the set of candidate variant amino acid sequences having binding or specificity to a same target. In some embodiments, the deep learning-based classification comprises a deep neural network (DNN). In some embodiments, the method further comprises performing automated and multimodal screening of the set of candidate variant amino acid sequences. In some embodiments, the automated and multimodal screening comprises using an ensemble model to identify a shortlist of antibodies, nanobodies, or fragments thereof with desirable therapeutics properties.
In another aspect, the present disclosure provides computer-implemented method for predicting an amino acid sequence or structure of an antibody paratrope or fragment thereof, having binding ability to a target protein, comprising: (a) obtaining a structure and amino acid sequences of an epitope of the target protein; and (b) processing the structure and amino acid sequences of the epitope of the target protein using a trained machine learning algorithm to predict the amino acid sequence or structure of the antibody paratrope or fragment thereof.
In some embodiments, the target protein comprises at least a portion of a target antigen. In some embodiments, the target antigen comprises at least a portion of SARS-CoV-2 virus.
In some embodiments, the trained machine learning algorithm is trained with training data comprising a set of training amino acid sequences or structures for the target protein. In some embodiments, the trained machine learning algorithm comprises a deep learning model. In some embodiments, the deep learning model comprises a deep generative model. In some embodiments, the deep generative model comprises at least one of a generative adversarial network (GAN), an autoencoder (AE), a variational autoencoder (VAE), a residual neural network (Resnet), Reinforcement Learning (RL), and a Graph Neural Network (GNN). In some embodiments, the deep generative model comprises at least two of the GAN, the AE, the VAE, the Resnet, the RL, and the GNN. In some embodiments, the deep generative model comprises at least three of the GAN, the VAE, the Resnet, the RL, and the GNN. In some embodiments, the deep generative model comprises at least four of the GAN, the VAE, the Resnet, the RL, and the GNN. In some embodiments, the deep generative model comprises the GAN, the VAE, the Resnet, the RL, and the GNN.
In some embodiments, (a) comprises obtaining a 3-dimensional structure and sequences of the target antigen, and wherein (b) comprises predicting the 3-dimensional structure of the antibody paratrope of fragment thereof.
In some embodiments, the method further comprises processing the predicted sequence or structure of the antibody paratrope of fragment thereof to determine molecular characteristics of the antibody paratrope of fragment thereof. In some embodiments, the molecular characteristics comprise at least one of amino acid distribution, amino acid frequency, amino acid length distribution, sequence similarity, sequence diversity, total charge distribution, hydrophobicity, net charge, stability, solubility, isoelectric point, solvent accessible surface area (SASA), binding affinity, epitope-paratrope interaction, epitope prediction, humanization, humaneness, developability, manufacturability, half-life, pharmacokinetic, yield, aggregation, function, clearance, viscosity, immunogenicity, single-target specificity, multi-target specificity, single-target cross-reactivity, and multi-target cross-reactivity.
In some embodiments, the method further comprises processing determining specificity or binding affinity of the antibody paratrope of fragment thereof to the target antigen.
In some embodiments, the method further comprises processing validating the predicted sequence or structure of the antibody paratrope or fragment thereof.
In some embodiments, (b) comprises predicting the sequence and the structure of the antibody paratrope or fragment thereof.
In some embodiments, the method further comprises using a deep learning-based model to analyze antigen epitope-antibody paratope interactions. In some embodiments, the deep learning-based model uses only amino acid sequences to predict epitopes, and wherein the method further comprises identifying antibodies, nanobodies, and fragments thereof with a desired mechanism of action.
In some embodiments, the method comprises quantitative metrics or qualitative metrics. In some embodiments, the quantitative metrics or qualitative metrics comprise at least one of accuracy, precision, perplexity (PPL), Area under the ROC Curve (AUC), amino acid recovery (AAR), and contact amino acid recovery (CAAR). In some embodiments, the method further comprises measuring at least one of contact accuracy, interface RMSD, and ligand RMSD.
In some embodiments, the method further comprises analyzing physicochemical attributes of the predicted antibody paratope or fragment thereof. In some embodiments, the physicochemical attributes comprise at least one of hydrophobicity, polarity, exposed surface, and accessibility.
In some embodiments, the method further comprises using a deep learning-based automated and multimodal screening method to identify the desirable paratope or fragment thereof faster, more accurately, and more efficiently. In some embodiments, the method further comprises classifying antibodies, nanobodies, or fragments thereof, based on their binding capabilities. In some embodiments, the classifying comprises use of a deep neural network (DNN).
Another aspect of the present disclosure provides a non-transitory computer readable medium comprising machine executable code that, upon execution by one or more computer processors, implements any of the methods above or elsewhere herein.
Another aspect of the present disclosure provides a system comprising one or more computer processors and computer memory coupled thereto. The computer memory comprises machine executable code that, upon execution by the one or more computer processors, implements any of the methods above or elsewhere herein.
Additional aspects and advantages of the present disclosure will become readily apparent to those skilled in this art from the following detailed description, wherein only illustrative embodiments of the present disclosure are shown and described. As will be realized, the present disclosure is capable of other and different embodiments, and its several details are capable of modifications in various obvious respects, all without departing from the disclosure. Accordingly, the drawings and description are to be regarded as illustrative in nature, and not as restrictive.
INCORPORATION BY REFERENCE
All publications, patents, and patent applications mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent, or patent application was specifically and individually indicated to be incorporated by reference. To the extent publications and patents or patent applications incorporated by reference contradict the disclosure contained in the specification, the specification is intended to supersede and/or take precedence over any such contradictory material.
BRIEF DESCRIPTION OF THE DRAWINGS
The novel features of the invention are set forth with particularity in the appended claims. A better understanding of the features and advantages of the present invention will be obtained by reference to the following detailed description that sets forth illustrative embodiments, in which the principles of the invention are utilized, and the accompanying drawings (also âfigureâ and âFIG.â herein), of which:
FIG. 1 illustrates an overview representation of an example system that identifies candidate amino acid sequences using a variational autoencoder (VAE) Resnet transfer learning (VAEResTL) deep generative model trained on initial amino acid sequence(s).
FIG. 2 A is a graph demonstrating the total loss during training for the VAEResTL model.
FIG. 2 B is a graph demonstrating the total loss during training for the VAERes model.
FIG. 3 A provides a Visualize Comparison of molecular characteristics, amino acid distribution, between Seed (SARS-CoV-2 seed sequences, Yellow), VAERes (generated sequences by VAERes for SARS-CoV-2, Violet) and VAEResTL (generated sequences by VAEResTL for SARS-CoV-2, Turquoise).
FIG. 3 B provides a Visualize Comparison of molecular characteristics, amino acid length distribution, between Seed (SARS-CoV-2 seed sequences, Yellow), VAERes (generated sequences by VAERes for SARS-CoV-2, Violet) and VAEResTL (generated sequences by VAEResTL for SARS-CoV-2, Turquoise).
FIG. 3 C provides a Visualize Comparison of molecular characteristics, total charge distribution, between Seed (SARS-CoV-2 seed sequences, Yellow),VAERes (generated sequences by VAERes for SARS-CoV-2, Violet) and VAEResTL (generated sequences by VAEResTL for SARS-CoV-2, Turquoise).
FIG. 3 D provides a Visualize Comparison of molecular characteristics, Eisenberg hydrophobicity, between Seed (SARS-CoV-2 seed sequences, Yellow),VAERes (generated sequences by VAERes for SARS-CoV-2, Violet) and VAEResTL (generated sequences by VAEResTL for SARS-CoV-2, Turquoise).
FIG. 3 E provides a Visualize Comparison of molecular characteristics, Eisenberg Hydrophobic moment, between Seed (SARS-CoV-2 seed sequences, Yellow),VAERes (generated sequences by VAERes for SARS-CoV-2, Violet) and VAEResTL (generated sequences by VAEResTL for SARS-CoV-2, Turquoise).
FIG. 4 A 1 is a Machine Learning (ML) Visualization that demonstrates a position proposed map. Heatmap visualization showing the count of observed seed sequences for SARS-CoV-2.
FIG. 4 B 1 is a Machine Learning (ML) Visualization that demonstrates a position proposed map. Heatmap visualization showing the count of VAERes-generated sequences or SARS-CoV-2 from (left) and to (right) each amino acid at each sequence position.
FIG. 4 C 1 is a Machine Learning (ML) Visualization that demonstrates a position proposed map. Heatmap visualization showing the count of VAEResTL-generated sequences for SARS-CoV-2 from (left) and to (right) each amino acid at each sequence position (seed and VAERes-generated and VAEResTL-generated sequences having length 20).
FIG. 4 A 2 is a Machine Learning (ML) Visualization that demonstrates sequence logos for Seeds SARS-CoV-2.
FIG. 4 B 2 is a Machine Learning (ML) Visualization that demonstrates the VAERes-generated sequences for SARS-CoV-2.
FIG. 4 C 2 is a Machine Learning (ML) Visualization that demonstrates the VAEResTL-generated sequences for SARS-CoV-2 are based on residue frequency.
FIG. 5 A illustrates In-Silico Screening of predicted CDR-H3 sequences for further assessment, Net charge of CDR-H3 sequence.
FIG. 5 B illustrates In-Silico Screening of predicted CDR-H3 sequences for further assessment, CDR-H3 hydrophobicity.
FIG. 5 C illustrates In-Silico Screening of predicted CDR-H3 sequences for further assessment, CamSol solubility score.
FIG. 5 D illustrates In-Silico Screening of predicted CDR-H3 sequences for further assessment, CDR-H3 stability score.
FIG. 5 E illustrates In-Silico Screening of predicted CDR-H3 sequences for further assessment, The minimum NetMHCIIpan % Rank (<2 strong binder; <10 weak binder) across all possible 15-mers for a given CDR-H3 sequences and across all 26 HLA alleles.
FIG. 5 F illustrates In-Silico Screening of predicted CDR-H3 sequences for further assessment, The average NetMHCIIpan % Rank across all possible 15-mers for a given CDR-H3 sequences and across all 26 HLA alleles.
FIG. 6 illustrates an overview representation of an example system that identifies amino acid sequences based on their single or multiple target specificity and/or antibody amino acid sequences' cross-reactivity to one or multiple targets using a variational autoencoder (VAE) Resnet dimensionality reduction (VAEResDR) model, which is a deep learning-based feature extractor model trained on amino acid sequence(s).
FIG. 7 A is a visualization plot of dimensionality reduction (DR) for the baseline PCA. CDR-H3 sequences for three target antigens in 5 clusters were processed.
FIG. 7 B is a visualization plot of dimensionality reduction for the baseline t-SNE. CDR-H3 sequences for three target antigens in 5 clusters were processed.
FIG. 7 C is a visualization plot of dimensionality reduction for the baseline UMAP. CDR-H3 sequences for three target antigens in 5 clusters were processed.
FIG. 7 D is a visualization plot of dimensionality reduction for the baseline VAE. CDR-H3 sequences for three target antigens in 5 clusters were processed.
FIG. 7 E is a visualization plot of dimensionality reduction for the baseline scCCESS (AE). CDR-H3 sequences for three target antigens in 5 clusters were processed.
FIG. 7 F is a visualization plot of dimensionality reduction for the baseline VAEDR. CDR-H3 sequences for three target antigens in 5 clusters were processed.
FIG. 7 G is a visualization plot of dimensionality reduction for the VAEResDR. CDR-H3 sequences for three target antigens in 5 clusters were processed.
FIG. 7 H is a visualization plot of dimensionality reduction for baseline PCA. CDR-H3 sequences for four target antigens in 6 clusters were processed.
FIG. 7 I is a visualization plot of dimensionality reduction for baseline t-SNE. CDR-H3 sequences for four target antigens in 6 clusters were processed.
FIG. 7 J is a visualization plot of dimensionality reduction for baseline uniform manifold approximation and projection (UMAP). CDR-H3 sequences for four target antigens in 6 clusters were processed.
FIG. 7 K is a visualization plot of dimensionality reduction for baseline VAE. CDR-H3 sequences for four target antigens in 6 clusters were processed.
FIG. 7 L is a visualization plot of dimensionality reduction for baseline scCCESS (AE). CDR-H3 sequences for four target antigens in 6 clusters were processed.
FIG. 7 M is a visualization plot of dimensionality reduction for the baseline VAEDR. CDR-H3 sequences for four target antigens in 6 clusters were processed.
FIG. 7 N is a visualization plot of dimensionality reduction for the VAEResDR. CDR-H3 sequences for four target antigens in 6 clusters were processed.
FIG. 8 A is a bar chart showing Clustering Comparison to demonstrate the dimensionality reduction (DR) methods' performance on CDR-H3 sequences in 5 clusters.
FIG. 8 B is a bar chart showing Clustering Comparison to demonstrate the DR methods' performance on CDR-H3 sequences in 6 clusters.
FIG. 9 A illustrates a VAEResDR two-dimensional Hidden Patterns Analysis. 2-dimensional visualization by VAEResDR method for four targets, 6 clusters. CDR-H3 Sequences sharing the same latent space are marked as group (a) and CDR-H3 sequences in different latent spaces are marked as group (b) and (c).
FIG. 9 B illustrates a VAEResDR two-dimensional Hidden Patterns Analysis, including the biophysical metric of hydrophobicity (H).
FIG. 9 C illustrates a VAEResDR two-dimensional Hidden Patterns Analysis, including the biophysical metric of net-charge.
FIG. 9 D illustrates a VAEResDR two-dimensional Hidden Patterns Analysis, including the biophysical metric of Iso-electric Point (ISO).
FIG. 9 E illustrates a VAEResDR two-dimensional Hidden Patterns Analysis, including the machine learning (ML) visualization Heatmap plot for group (a).
FIG. 9 F illustrates a VAEResDR two-dimensional Hidden Patterns Analysis, including the ML visualization Heatmap plot for group (b).
FIG. 9 G illustrates a VAEResDR two-dimensional Hidden Patterns Analysis, including the ML visualization Heatmap plot for group (c).
FIG. 10 A is an amino acid distribution plot comparing Seed sequences (SARS-CoV-2 seed sequences, orange), (OVA seed sequences, Turquoise), (Yeast display scFV seed sequences, Violet), (Rani seed sequences, Yellow), and includes a schematic of the comparison of the training seed amino acid sequences.
<div id="p-0108" num="0107" class="de
CLAIMS
Claims ( 29 )
1 .- 117 . (canceled)
118 . A computer-implemented method for generating a set of candidate variant amino acid sequences of an antibody, a nanobody, or a fragment thereof, having binding ability to protein targets, comprising:
(a) obtaining a set of seed amino acid sequences; and (b) processing the set of seed amino acid sequences using a first trained machine learning algorithm to generate the set of novel candidate variant amino acid sequences, wherein the first trained machine learning algorithm is trained with the first training data comprising a set of training amino acid sequences for the protein targets, wherein the first trained machine learning algorithm is further trained through a transfer learning method using a second trained machine learning algorithm, wherein the second trained machine learning algorithm is trained with second training data comprising a set of training amino acid sequences for a second protein target, wherein the second protein target is different from the first protein targets.
119 . The method of claim 118 , wherein the set of seed amino acid sequences comprises antibody variable regions (Fvs).
120 . The method of claim 119 , wherein the Fvs comprise complementarity determining regions (CDRs).
121 . The method of claim 119 , wherein the Fvs comprise heavy chains (VH) or light chains (VL).
122 . The method of claim 119 , wherein the Fvs comprise frameworks (FWRs).
123 . The method of claim 118 , wherein the protein targets comprise at least a portion of a target antigen.
124 . The method of claim 118 , wherein the set of training amino acid sequences for the second protein target has a smaller number of training data than the set of training amino acid sequences for the first protein targets.
125 . The method of claim 118 , further comprises, prior to (b), pre-processing the set of seed amino acid sequences (i) to have the same sequence length or (ii) at least in part by generating a numerical representation of the set of seed amino acid sequences.
126 . The method of claim 125 , wherein the numerical representation comprises a matrix, wherein the matrix comprises a one-hot encoding and/or an embedding pre-processing of (i) the set of amino acid sequences or (ii) amino acid bio-physicochemical values.
127 . The method of claim 126 , wherein the amino acid bio-physicochemical values are selected from isoelectric point, volume, hydrophobicity, solubility, stability, charge, solvent-accessible surface area (SASA), Immunogenicity, and humanness score.
128 . The method of claim 118 , wherein the first trained machine learning algorithm or the second trained machine learning algorithm comprises a deep learning model.
129 . The method of claim 128 , wherein the deep learning model comprises a deep generative AI model, wherein the deep generative AI model comprises at least one of a generative adversarial network (GAN), an autoencoder (AE), a variational autoencoder (VAE), and a residual neural network (Resnet), and a Reinforcement Learning (RL).
130 . The method of claim 118 , wherein the first trained machine learning algorithm or the second trained machine learning algorithm comprises an encoder and a decoder.
131 . The method of claim 130 , further comprising performing, using the encoder, a dimensionality reduction of the set of seed amino acid sequences.
132 . The method of claim 131 , further comprising extracting a set of features associated with the set of seed amino acid sequences generated by the encoder, and reconstructing the set of features using a decoder thereby generating new candidate variant amino acid sequences.
133 . The method of claim 131 , further comprising clustering features of the set of seed amino acid sequences based on sequence similarity or sequence diversity.
134 . The method of claim 133 , wherein the clustering comprises generating a two-dimensional plot indicative of single specificity or multi-specificity of the set of seed amino acid sequences.
135 . The method of claim 118 , wherein the first trained machine learning algorithm and the second trained machine learning algorithm are trained without the use of structural data.
136 . The method of claim 118 , further comprising processing the set of generated candidate variant amino acid sequences to determine a therapeutic property selected from single specificity, multiple specificity, single cross-reactivity, and multiple cross-reactivity.
137 . The method of claim 118 , further comprising processing the set of generated candidate variant amino acid sequences to determine molecular characteristics of the set of generated candidate variant amino acid sequences.
138 . The method of claim 118 , wherein the molecular characteristics comprise at least one of amino acid distribution, amino acid frequency, amino acid length distribution, sequence similarity, sequence diversity, total charge distribution, hydrophobicity, net charge, stability, solubility, isoelectric point, solvent accessible surface area (SASA), binding affinity, affinity maturation, epitope-paratrope interaction, epitope prediction, humaneness, humanization, developability, manufacturability, half-life, pharmacokinetic, yield, aggregation, function, clearance, viscosity, immunogenicity, single-target specificity, multi-target specificity, single-target cross-reactivity, and multi-target cross-reactivity.
139 . The method of claim 138 , further comprising determining sequence similarity between the set of generated candidate variant amino acid sequences and the set of seed amino acid sequences.
140 . The method of claim 138 , further comprising determining sequence diversity between the set of generated candidate variant amino acid sequences and the set of seed amino acid sequences.
141 . The method of claim 138 , further comprising determining specificity or binding affinity of the set of generated candidate variant amino acid sequences to the protein target.
142 . The method of claim 138 , further comprising (i) selecting or ranking at least one of the set of generated candidate variant amino acid sequences based at least in part on having desired molecular characteristics, or (ii) filtering out at least one of the set of generated candidate variant amino acid sequences having undesired molecular characteristics.
143 . The method of claim 118 , wherein the set of generated candidate variant amino acid sequences correspond to an antibody, a nanobody, or a fragment thereof.
144 . The method of claim 118 , further comprising predicting a humanized antibody, nanobody, or fragment variable region based at least in part on the set of generated candidate variant amino acid sequences.
145 . A computer-implemented method for generating a set of candidate variant amino acid sequences of an antibody, a nanobody, or a fragment thereof, having binding ability to a protein target, comprising:
(a) obtaining a set of seed amino acid sequences; and (b) processing the set of seed amino acid sequences using a trained machine learning algorithm to generate the set of candidate variant amino acid sequences, wherein the trained machine learning algorithm comprises at least one of a generative adversarial network (GAN) and Reinforcement Learning (RL).
US18/615,028
2021-09-27
2024-03-25
Machine learning for designing antibodies and nanobodies in-silico
Pending
US20240379248A1
( en )
Priority Applications (1)
Application Number
Priority Date
Filing Date
Title
US18/615,028
US20240379248A1
( en )
2021-09-27
2024-03-25
Machine learning for designing antibodies and nanobodies in-silico
Applications Claiming Priority (7)
Application Number
Priority Date
Filing Date
Title
US202163248761P
2021-09-27
2021-09-27
US202263318037P
2022-03-09
2022-03-09
US202263332418P
2022-04-19
2022-04-19
US202263395487P
2022-08-05
2022-08-05
US202263397603P
2022-08-12
2022-08-12
PCT/US2022/044754
WO2023049466A2
( en )
2021-09-27
2022-09-26
Machine learning for designing antibodies and nanobodies in-silico
US18/615,028
US20240379248A1
( en )
2021-09-27
2024-03-25
Machine learning for designing antibodies and nanobodies in-silico
Related Parent Applications (1)
Application Number
Title
Priority Date
Filing Date
PCT/US2022/044754
Continuation
WO2023049466A2
( en )
2021-09-27
2022-09-26
Machine learning for designing antibodies and nanobodies in-silico
Publications (1)
Publication Number
Publication Date
US20240379248A1
true
US20240379248A1 ( en )
2024-11-14
Family
ID=85719642
Family Applications (1)
Application Number
Title
Priority Date
Filing Date
US18/615,028
Pending
US20240379248A1
( en )
2021-09-27
2024-03-25
Machine learning for designing antibodies and nanobodies in-silico
Country Status (2)
Country
Link
US
( 1 )
US20240379248A1
( en )
WO
( 1 )
WO2023049466A2
( en )
Cited By (2)
* Cited by examiner, â Cited by third party
Publication number
Priority date
Publication date
Assignee
Title
US12367329B1
( en )
*
2024-06-06
2025-07-22
EvolutionaryScale, PBC
Protein binder search
CN120722757A
( en )
*
2025-08-27
2025-09-30
ä¸å大å¦
A turboshaft engine identification and predictive control method based on MRR-KELM
Families Citing this family (13)
* Cited by examiner, â Cited by third party
Publication number
Priority date
Publication date
Assignee
Title
US11049590B1
( en )
2020-02-12
2021-06-29
Peptilogics, Inc.
Artificial intelligence engine architecture for generating candidate drugs
US11403316B2
( en )
2020-11-23
2022-08-02
Peptilogics, Inc.
Generating enhanced graphical user interfaces for presentation of anti-infective design spaces for selecting drug candidates
US11512345B1
( en )
2021-05-07
2022-11-29
Peptilogics, Inc.
Methods and apparatuses for generating peptides by synthesizing a portion of a design space to identify peptides having non-canonical amino acids
WO2024238129A1
( en )
*
2023-05-16
2024-11-21
Genentech, Inc.
Clearance prediction according to antibody property analysis
CN116959576A
( en )
*
2023-05-18
2023-10-27
è ¾è®¯ç§æï¼æ·±å³ï¼æéå ¬å¸
Antibody sequence generation method, device, computer equipment and storage medium
US20260018249A1
( en )
*
2023-05-31
2026-01-15
Amazon Technologies, Inc.
Peptide manufacturability determination
EP4721077A1
( en )
*
2023-06-05
2026-04-08
Sanofi
Predicting properties of single variable domains using machine-learning models
WO2025022002A1
( en )
*
2023-07-26
2025-01-30
Alchemab Therapeutics Ltd
Analysis of antigen-binding proteins
CN117291138B
( en )
*
2023-11-22
2024-02-13
å ¨è¯æºé ææ¯æéå ¬å¸
Method, apparatus and medium for generating layout elements
US20250232830A1
( en )
*
2024-01-11
2025-07-17
Amgen Inc.
Methods and systems for viscosity prediction and protein engineering
WO2025193716A1
( en )
*
2024-03-11
2025-09-18
Livemed Health Inc.
Systems, methods, and devices for message control
CN117952961B
( en )
*
2024-03-25
2024-06-07
æ·±å³å¤§å¦
Training and application method and device of image prediction model and readable storage medium
WO2025255259A1
( en )
*
2024-06-06
2025-12-11
Generate Biomedicines, Inc.
Machine learning-guided generation of cross-reactive neutralizing antigen binding molecules against viral proteins
Family Cites Families (4)
* Cited by examiner, â Cited by third party
Publication number
Priority date
Publication date
Assignee
Title
US10431325B2
( en )
*
2012-08-03
2019-10-01
Novartis Ag
Methods to identify amino acid residues involved in macromolecular binding and uses therefor
WO2020208555A1
( en )
*
2019-04-09
2020-10-15
Eth Zurich
Systems and methods to classify antibodies
US20220270711A1
( en )
*
2019-08-02
2022-08-25
Flagship Pioneering Innovations Vi, Llc
Machine learning guided polypeptide design
WO2021138548A1
( en )
*
2020-01-02
2021-07-08
Spring Discovery, Inc.
Methods, systems, and tools for longevity-related applications
2022
2022-09-26
WO
PCT/US2022/044754
patent/WO2023049466A2/en
not_active
Ceased
2024
2024-03-25
US
US18/615,028
patent/US20240379248A1/en
active
Pending
Cited By (2)
* Cited by examiner, â Cited by third party
Publication number
Priority date
Publication date
Assignee
Title
US12367329B1
( en )
*
2024-06-06
2025-07-22
EvolutionaryScale, PBC
Protein binder search
CN120722757A
( en )
*
2025-08-27
2025-09-30
ä¸å大å¦
A turboshaft engine identification and predictive control method based on MRR-KELM
Also Published As
Publication number
Publication date
WO2023049466A2
( en )
2023-03-30
WO2023049466A3
( en )
2023-09-14
Similar Documents
Publication
Publication Date
Title
WO2023049466A2
( en )
2023-03-30
Machine learning for designing antibodies and nanobodies in-silico
Pittala et al.
2020
Learning context-aware structural representations to predict antigen and antibody binding interfaces
Lv et al.
2021
Anticancer peptides prediction with deep representation learning features
Sahu et al.
2022
Artificial intelligence (AI) in drugs and pharmaceuticals
Wilman et al.
2022
Machine-designed biotherapeutics: opportunities, feasibility and advantages of deep learning in computational antibody discovery
Andrews et al.
2020
Structural aspects and prediction of calmodulin-binding proteins
Xu et al.
2020
Deep dive into machine learning models for protein engineering
Joubbi et al.
2024
Antibody design using deep learning: from sequence and structure design to affinity maturation
Chen et al.
2021
Sequence-based peptide identification, generation, and property prediction with deep learning: a review
Brizuela et al.
2025
AI methods for antimicrobial peptides: progress and challenges
Pushkaran et al.
2024
From understanding diseases to drug design: can artificial intelligence bridge the gap?
Zeng et al.
2023
Recent progress in antibody epitope prediction
Singh et al.
2025
Learning the language of antibody hypervariability
Zhang et al.
2022
DeepStack-DTIs: predicting drugâtarget interactions using LightGBM feature selection and deep-stacked ensemble classifier
Kusuma et al.
2019
Prediction of ATP-binding sites in membrane proteins using a two-dimensional convolutional neural network
Nedyalkova et al.
2023
Sequence-based prediction of plant allergenic proteins: machine learning classification approach
EP4427224A1
( en )
2024-09-11
Systems and methods for polymer sequence prediction
WO2025022002A1
( en )
2025-01-30
Analysis of antigen-binding proteins
Dai et al.
2025
Predicting metal-binding proteins and structures through integration of evolutionary-scale and physics-based modeling
Ali et al.
2023
IGPredâHDnet: Prediction of Immunoglobulin Proteins Using Graphical Features and the Hierarchal Deep LearningâBased Approach
Sriwastava et al.
2015
Proteinâprotein interaction site prediction in Homo sapiens and E. coli using an interaction-affinity based membership function in fuzzy SVM
Peng et al.
2023
Generative diffusion models for antibody design, docking, and optimization
Han et al.
2021
Quality assessment of protein docking models based on graph neural network
Li et al.
2024
Prediction of paratopeâepitope pairs using convolutional neural networks
Han et al.
2024
GeoNet enables the accurate prediction of protein-ligand binding sites through interpretable geometric deep learning
Legal Events
Date
Code
Title
Description
2024-08-05
AS
Assignment
Owner name : MARWELL BIO INC., CALIFORNIA
Free format text : ASSIGNMENT OF ASSIGNORS INTEREST;ASSIGNORS:HEMMATIAN, ZARA;CHAVOSH, ADRIAN ALIREZA;REEL/FRAME:068180/0909
Effective date : 20240418
2024-08-31
STPP
Information on status: patent application and granting procedure in general
Free format text : DOCKETED NEW CASE - READY FOR EXAMINATION
2025-12-12
STPP
Information on status: patent application and granting procedure in general
Free format text : NON FINAL ACTION COUNTED, NOT YET MAILED
2025-12-23
STPP
Information on status: patent application and granting procedure in general
Free format text : NON FINAL ACTION MAILED
2025-12-29
STPP
Information on status: patent application and granting procedure in general
Free format text : NON FINAL ACTION MAILED