ConceptioArchiveNCBI PubMed Central
NCBI PubMed Centralopen access

XA-Novo: high-throughput mass spectrometry-based de novo sequencing technology for monoclonal antibodies and antibody mixtures.

Xiong Y et al. · ncbi_pmc
NCBI PubMed Central · Papers · License: Open Access
Open Source ↗Direct PDF ↓
machine learning systems

Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Nat Commun . 2026 Mar 12;17:3391. doi: 10.1038/s41467-026-70496-y Search in PMC Search in PubMed View in NLM Catalog Add to search XA-Novo: high-throughput mass spectrometry-based de novo sequencing technology for monoclonal antibodies and antibody mixtures Yueting Xiong Yueting Xiong 1 State Key Laboratory of Vaccines for Infectious Diseases, Xiang An Biomedicine Laboratory, National Institute for Data Science in Health and Medicine, School of Public Health, Xiamen University, Fujian, China Find articles by Yueting Xiong 1, ✉, # , Wenbin Jiang Wenbin Jiang 1 State Key Laboratory of Vaccines for Infectious Diseases, Xiang An Biomedicine Laboratory, National Institute for Data Science in Health and Medicine, School of Public Health, Xiamen University, Fujian, China 2 School of Informatics, Xiamen University, Fujian, China Find articles by Wenbin Jiang 1, 2, # , Jin Xiao Jin Xiao 1 State Key Laboratory of Vaccines for Infectious Diseases, Xiang An Biomedicine Laboratory, National Institute for Data Science in Health and Medicine, School of Public Health, Xiamen University, Fujian, China Find articles by Jin Xiao 1, # , Qingfang Bu Qingfang Bu 1 State Key Laboratory of Vaccines for Infectious Diseases, Xiang An Biomedicine Laboratory, National Institute for Data Science in Health and Medicine, School of Public Health, Xiamen University, Fujian, China Find articles by Qingfang Bu 1 , Jingyi Wang Jingyi Wang 1 State Key Laboratory of Vaccines for Infectious Diseases, Xiang An Biomedicine Laboratory, National Institute for Data Science in Health and Medicine, School of Public Health, Xiamen University, Fujian, China Find articles by Jingyi Wang 1 , Zhenjian Jiang Zhenjian Jiang 3 Department of Pathology, Zhongshan Hospital, Fudan University (Xiamen Branch), Fujian, China Find articles by Zhenjian Jiang 3 , Ling Luo Ling Luo 1 State Key Laboratory of Vaccines for Infectious Diseases, Xiang An Biomedicine Laboratory, National Institute for Data Science in Health and Medicine, School of Public Health, Xiamen University, Fujian, China Find articles by Ling Luo 1 , Xiaoqing Chen Xiaoqing Chen 1 State Key Laboratory of Vaccines for Infectious Diseases, Xiang An Biomedicine Laboratory, National Institute for Data Science in Health and Medicine, School of Public Health, Xiamen University, Fujian, China Find articles by Xiaoqing Chen 1 , Yijie Qiu Yijie Qiu 1 State Key Laboratory of Vaccines for Infectious Diseases, Xiang An Biomedicine Laboratory, National Institute for Data Science in Health and Medicine, School of Public Health, Xiamen University, Fujian, China Find articles by Yijie Qiu 1 , Yangtao Wu Yangtao Wu 1 State Key Laboratory of Vaccines for Infectious Diseases, Xiang An Biomedicine Laboratory, National Institute for Data Science in Health and Medicine, School of Public Health, Xiamen University, Fujian, China Find articles by Yangtao Wu 1 , Fan Liu Fan Liu 1 State Key Laboratory of Vaccines for Infectious Diseases, Xiang An Biomedicine Laboratory, National Institute for Data Science in Health and Medicine, School of Public Health, Xiamen University, Fujian, China Find articles by Fan Liu 1 , Rongshan Yu Rongshan Yu 1 State Key Laboratory of Vaccines for Infectious Diseases, Xiang An Biomedicine Laboratory, National Institute for Data Science in Health and Medicine, School of Public Health, Xiamen University, Fujian, China 2 School of Informatics, Xiamen University, Fujian, China 4 Aginome Scientific, Xiamen, Fujian, China Find articles by Rongshan Yu 1, 2, 4, ✉ , Ningshao Xia Ningshao Xia 1 State Key Laboratory of Vaccines for Infectious Diseases, Xiang An Biomedicine Laboratory, National Institute for Data Science in Health and Medicine, School of Public Health, Xiamen University, Fujian, China Find articles by Ningshao Xia 1, ✉ , Quan Yuan Quan Yuan 1 State Key Laboratory of Vaccines for Infectious Diseases, Xiang An Biomedicine Laboratory, National Institute for Data Science in Health and Medicine, School of Public Health, Xiamen University, Fujian, China Find articles by Quan Yuan 1, ✉ Author information Article notes Copyright and License information 1 State Key Laboratory of Vaccines for Infectious Diseases, Xiang An Biomedicine Laboratory, National Institute for Data Science in Health and Medicine, School of Public Health, Xiamen University, Fujian, China 2 School of Informatics, Xiamen University, Fujian, China 3 Department of Pathology, Zhongshan Hospital, Fudan University (Xiamen Branch), Fujian, China 4 Aginome Scientific, Xiamen, Fujian, China ✉ Corresponding author. # Contributed equally. Received 2025 Jan 16; Accepted 2026 Feb 26; Collection date 2026. © The Author(s) 2026 Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article's Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article's Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/ . PMC Copyright notice PMCID: PMC13066000  PMID: 41820386 Abstract Elucidating antibody sequences by mass spectrometry-based de novo sequencing is essential but remains technically challenging. Here we present XA-Novo, an accurate and high-throughput de novo sequencing solution that integrates a single-pot multi-enzymatic gradient digestion method with a beam search-based assembler (Fusion) to reconstruct full-length antibody sequences directly from bottom-up mass spectrometry data. Benchmarking across well-characterized antibodies from multiple species demonstrates that XA-Novo outperforms commercial solutions in identification sensitivity, sequence completeness, and reconstruction accuracy. Furthermore, XA-Novo successfully reconstructs six immunotherapeutic antibodies with unknown sequences, and in vitro/vivo assays validate that these generated antibodies exhibit functionality equivalent to their commercial counterparts. Moreover, XA-Novo achieves over 99.54% accurate sequence coverage in distinguishing mixed COVID-19 neutralizing antibodies, exceeding the performance of current assemblers reported for single-antibody sequencing. Overall, XA-Novo establishes a reliable, scalable, and broadly applicable workflow for routine antibody sequencing, thereby accelerating both fundamental antibody research and therapeutic antibody development. Subject terms: Proteomics, Protein sequencing, Mass spectrometry XA-Novo is a mass spectrometry–only de novo antibody sequencing workflow that reconstructs full-length antibodies from single samples or mixtures, enabling rapid sequence recovery and functional re-expression for antibody research and drug discovery. Introduction Antibodies play a crucial role in host defense against pathogens by recognizing and neutralizing foreign entities, such as pathogenic bacteria and viruses 1 , 2 . In addition to directly neutralizing these pathogens, antibodies coordinate a broader immune response by recruiting effector functions for mechanisms including antibody-dependent cellular cytotoxicity and complement-dependent cytotoxicity 1 . Moreover, antibodies generated in response to an infection can persist in circulation for extended periods, facilitating rapid and robust responses upon re-encountering the same pathogen due to the recall ability of memory B cells 3 . The multifaceted functionality of antibodies underscores their critical significance across a broad spectrum of disciplines, encompassing immunology, clinical chemistry, biochemistry, therapeutics, and medicine 4 – 6 . Consequently, they have emerged as invaluable tools in numerous research endeavors and practical applications. The antibody sequence information is crucial for comprehending the structural foundation of antigen-antibody binding, recognition, and interaction 7 , 8 . Additionally, sequence-based recombination, expression, or engineering modifications can be used to achieve constant region substitution of antibodies from different species, thereby altering antibody specificity and enabling the design of diverse antibodies. However, obtaining complete sequences for monoclonal antibodies (mAbs) remains a significant challenge. Traditional hybridoma technology 9 , 10 , which involves antibody screening and sequencing, typically requires three to six months to complete and incurs substantial time and cost. Researchers have made progress in isolating B cells directly from blood or bone marrow, extracting DNA or RNA, and creating sequencing libraries using second-generation sequencing technologies independent of the hybridoma method. This advancement has reduced the timeline for obtaining sequences to as short as one month 11 – 13 . Despite these advancements, the time-consuming nature of these methods and the need for complementary information from both genetic and protein levels remain challenges that must be addressed. Additionally, issues such as bacterial contamination or decreased viability during the antibody screening process can render it impossible to obtain sequence information, severely limiting flexibility in antibody screening efforts. Furthermore, since antibodies are secreted into bodily fluids and mucus without a direct connection to their producing B cells, questions arise regarding the quantitative relationship between the secreted antibody pool and its underlying B-cell population—along with potential sampling biases inherent in current antibody sequencing strategies. Mass spectrometry (MS)-based de novo sequencing of secreted antibodies serves as a valuable complementary approach that can effectively address several challenges faced by conventional strategies, which rely on cloning and sequencing of coding mRNAs 14 – 16 . For example, the MS-based de novo sequencing approach can directly target polypeptide products in biological fluids to acquire antibody sequences. This method has successfully sequenced mAbs from lost hybridoma cell lines, thereby establishing the foundation for the next generation of serological multi-antibody de novo sequencing. Nevertheless, despite these advancements, MS-based de novo sequencing still encounters significant impediments such as large sample requirements, low throughput, unanticipated missing ion information, and difficulties in sequencing and assembly accuracy. During antibody sequencing, the existence of isomers with similar mass (e.g., leucine [Leu] and isoleucine [Ile]) or amino acid combinations that yield identical results (e.g., AG = Q, GG = N) can easily lead to erroneous sequence assembly 17 – 19 . These issues are already encountered when sequencing a single antibody and become even more pronounced during the de novo sequencing of multiple antibodies. Therefore, selecting appropriate peptides from ambiguous sequences and implementing accurate amino acid error correction are essential for developing more precise, efficient high-throughput methods for antibody sequencing. To address these challenges, we propose XA-Novo, an integrated pipeline combining single-pot multi-enzymatic gradient digestion (SP-MEGD) with a beam-search-based Fusion assembler for accurate end-to-end reconstruction of complete antibody sequences. We benchmarked XA-Novo using data from multiple known antibodies across various species, which included samples containing only a single antibody as well as mixtures of COVID-19 neutralizing antibodies (NAbs) comprising two or three antibodies per sample. The results demonstrated that XA-Novo outperforms commercial solutions in terms of identification capacity, data completeness, and accuracy of antibody reconstruction. Furthermore, XA-Novo successfully reconstructed six immunotherapeutic antibodies with unknown sequences; in vitro assays confirmed that the generated antibodies exhibit functionality equivalent to their commercial counterparts, thereby further validating the effectiveness of our solution. In conclusion, we have established a precise, robust, high-throughput de novo sequencing solution that advances antibody research by enhancing both the accuracy and efficiency of complex antibody sequencing. Results XA-Novo delivers high-throughput, mixture-compatible de novo antibody sequencing by MS-only proteomics The comprehensive de novo sequencing solution provided by the XA-Novo platform is illustrated in Fig. 1 . Specifically, XA-Novo was developed based on SP-MEGD sample preparation for obtaining sufficient peptide-level information, followed by liquid chromatography-mass spectrometry (LC-MS/MS) analysis for generating mass spectra, then employing deep learning-based de novo peptide sequencing to generate peptide reads, and finally using Fusion assembler to assemble these peptides into complete heavy and light chains. The sequencing results obtained through experimental methods were initially validated via intact mass analysis and subsequently confirmed through bioactivity or functional assays. Our mAb sequencing method is capable of functioning with as little as 50 µg of purified mAb. This method is not only applicable to traditional mAb sequencing but also facilitates high-throughput mixture sequencing involving multiple antibodies. Our study provides three key contributions that leverage these advancements, thereby making antibody sequencing at the protein level via MS more attainable: Extensive experimental results demonstrate the accuracy and robustness of our solution for traditional mAb sequencing, thereby facilitating the acquisition of precise antibody sequence information for routine applications. XA-Novo enables higher-throughput full-length de novo mAb sequencing within a bottom-up proteomics workflow, while maintaining high accuracy, thereby improving the efficiency of mAb research. XA-Novo supports the simultaneous analysis of multiple mAbs under identical experimental conditions, reducing sequencing costs and improving resource utilization in large-scale antibody research and development by optimizing enzyme and reagent usage and MS instrument time. Fig. 1. XA-Novo: a mass spectrometry-based de novo sequencing workflow for monoclonal antibodies. Open in a new tab a Overview of the XA-Novo pipeline integrating experimental sample preparation and computational sequence assembly. Monoclonal antibody samples (50–200 μg) were processed using a single pot multi-enzymatic gradient digestion (SP-MEGD) method, involving denaturation and reduction (30 mins), alkylation (30 mins), and ultrafiltration (10 mins). Five proteases were used in a time-staggered gradient digestion, with aliquots collected at 2nd, 4th, and 6th hour and pooled for analysis. Peptides were analyzed by LC-MS/MS using a dual-fragmentation strategy (HCD and EThcD). De novo sequencing analysis was performed via a deep learning model and the Fusion assembler, which applies a beam search strategy to optimize sequence reconstruction. (Created in BioRender. XIONG, Y. (2026) https://BioRender.com/04z86q3 ). b In monoclonal antibody analysis, XA-Novo enables full-sequence reconstruction and validation through intact mass profiling and functional assays. (Created in BioRender. XIONG, Y. (2026) https://BioRender.com/04z86q3 ). c XA-Novo enables accurate reconstruction of paired heavy- and light-chain antibody sequences from complex antibody mixtures. (Created in BioRender. XIONG, Y. (2026) https://BioRender.com/04z86q3 ). SP-MEGD sample preparation enables sufficient peptide-level information To tackle the challenge of unexpectedly missing ion information and incomplete sequence coverage, we developed a method termed SP-MEGD. This approach employs a five-protease gradient digestion, with each enzyme utilized in a single pot (Fig. 1a ). Samples were collected at two-hour intervals over a total duration of six hours within an integrated reactor. This approach is designed to minimize sample loss and enhance the availability of antibody ion information. During SP-MEGD method development, we optimized and evaluated crucial aspects of antibody preparation using the S2P6 mAb with published sequences 20 . We prioritized sequence coverage and accuracy as the primary performance metric (refer to Fig. 2a for a detailed explanation). We aimed to identify the best combination of reagents for protein denaturation and reduction, minimizing sample loss compared to the conventional five-step process 21 . We compared the lysis efficiencies of three compatible lysis buffers: a. 2% DOC + 20 mM TCEP + 200 mM Tris-HCl, pH 8.0; b. 8 M Urea + 20 mM TCEP + 100 mM Tris-HCl, pH 8.0; c. 6 M Gu·HCl + 20 mM DTT. The results showed that the Gu·HCl buffer provided high coverage rates for heavy chains in both high-energy collision dissociation (HCD) and electron-transfer high-energy collision dissociation (EThcD) fragmentation modes when combined with various enzymes - trypsin, chymotrypsin, pepsin, elastase, and Asp-N (Fig. 2b ). However, coverage rates for the light chain from the combination of Gu·HCl buffer and Asp-N were not high due to fewer cleavage sites in the tested antibody light chain sequence, resulting in limited peptide information. Therefore, Gu·HCl buffer was selected for lysing antibody samples, considering its overall effectiveness in achieving high coverage. Fig. 2. Development and optimization of the SP-MEGD method for antibody sequencing. Open in a new tab a Antibody sequence assembly was evaluated based on three key metrics: coverage, accuracy, and accurate coverage. b Sequence coverage of S2P6 heavy and light chains using different digestion buffers (DOC, urea, Gu·HCl) combined with various enzymes and MS fragmentation modes (HCD and EThcD). c Numbers of peptide-spectrum matches (PSMs) and identified peptides for anti-mouse Ly-6G antibody using five proteases at enzyme-to-protein ratios of 1:20 and 1:50. Statistical analysis was performed utilizing a paired t-test ( n =  3). d Distribution of PSMs for anti-mouse Ly-6G antibody across different confidence thresholds; each bar represents the average of three replicates ( n =  3). e Average sequence coverage of S2P6 heavy and light chains at different digestion times (2, 4, 6, 8, and 16 hours), with triplicate measurements ( n =  3). f Comparison of missed cleavage rates at different digestion times using trypsin. The left panel shows missed cleavage rates at single time points; the right panel compares missed cleavage rates between gradient digestion (sampling at 2nd–4th–6th hour) and overnight digestion (sampling at 16th hour). Statistical analysis was performed utilizing a paired t-test ( n =  3). g Using various fragmentation modes, the sequence coverage of both the light and heavy chains of S2P6 was assessed at different confidence levels. The analysis utilized a trypsin/protein ratio of 1:20 (w/w), with samples collected at the 2nd, 4th, and 6th hour. h Intact mass measurements of S2P6 heavy and light chains. i Comparison of SP-MEGD and commercial digestion kits in HCD mode; various enzymes-digested samples under gradient conditions were assessed across three replicates. Statistical analysis was performed utilizing a Student’s t -test ( n =  3). j Summary of performance metrics comparing SP-MEGD with the EasyPept-Micro kit, including PSMs, peptide counts, and CDR sequence coverage and accuracy. k Peptide reads identified by Casanovo with confidence scores >=50 across different mAb samples. l Average sequencing depth across CDRs, variable regions, and full-length antibody chains in COVID-19 antibodies. Error bars presented as mean ± SD. Source data are provided as a Source Data file. Previous studies have demonstrated that limited proteolytic reactions, achieved by reducing enzyme concentration or shortening reaction times, can increase the frequency of missed cleavage sites, producing diverse peptide fragments of varying lengths with overlapping residue extensions that facilitate efficient database-free de novo sequencing 22 , 23 . Therefore, we first evaluated four different enzyme-to-protein ratios (ranging from 1:5 to 1:50) with a fixed 16-hour reaction time to ensure reproducibility. LC-MS/MS analysis showed an increase in unique peptides as the enzyme-to-substrate ratio decreased from 1:5 to 1:50 (Supplementary Fig. 1c–e ). The optimal trypsin-to-protein ratios were established at both 1:20 and 1:50, yielding increased peptide counts and peak areas that fulfill the requirements for overlapping segments necessary in downstream assembly. Similar optimizations were carried out for the remaining predicted enzymes, each resulting in its respective optimal enzymatic ratio. To further evaluate the performance of a five-enzyme combination, we employed a commercial antibody (anti-mouse Ly-6G) as a test case. As a result, the ratio of 1:20 demonstrated superior outcomes in terms of peptide counts, sequence coverage, and number of PSMs (Fig. 2c, d ). We then investigated the optimal enzymatic digestion time for antibody samples. A comparative analysis was conducted over a range of 2 to 16 hours (overnight), with three technical replicates at each time point. The results indicated that the average coverage of antibody sequences increased from 2 to 6 hours, with the effect observed at 6-hour digestion being comparable to that achieved after overnight digestion at 16 hours. This trend was consistent across both HCD and EThcD fragmentation modes (Fig. 2e ). Recognizing the significance of abundant fragment information in supporting peptide assembly in antibody de novo sequencing, we proposed a gradient digestion strategy (SP-MEGD). Samples were collected at the 2nd, 4th, and 6th hours before being pooled to maximize the missed cleavage rates (Fig. 2f ). This strategy could provide effective “head-to-tail” evidential information for assembling peptide fragments. While a sequence coverage exceeding 87% with a confidence level of at least 95% has been achieved using 1:20 single trypsin digestion in gradient digestion (Fig. 2g ), significant improvements in peptide counts can be realized by leveraging multiple enzyme combinations. Hence, the final SP-MEGD method employs a five-protease (trypsin, chymotrypsin, pepsin, elastase, and Asp-N) gradient digestion approach, with samples collected every two hours over a total duration of 6 hours. This strategy offers greater diversity in proteolytic peptide length, enhancing the probability of overlapping fragments necessary for de novo assembly and providing flexibility regardless of the protein sequence and properties. We conducted additional experiments to assess the SP-MEGD method using various amounts (50, 100, and 200 µg) of starting materials for antibody samples. Based on PEAKS AB, S2P6 antibodies consistently achieved 100% accuracy in the amino acid sequence within all complementarity-determining region (CDR) at initial quantities of 50, 100, or 200 µg (Supplementary Fig. 2 ). This was confirmed by comparing the determined sequence to the S2P6 reference sequence, and the measured masses of both light and heavy chains align with theoretical calculations (Fig. 2h ). These results demonstrate that our SP-MEGD method is suitable for analyzing antibody products with sample volumes as low as 50 µg, indicating its potential for routine use. We further compared the SP-MEGD method with the commercial EasyPept-Micro kit (Omicsolution, #OSFP0001) for microscale label-free sample preparation using 25 µg of S2P6 antibody. Given that the EasyPept-Micro kit exclusively employs trypsin for digestion, we initially conducted a comparison of trypsin digestion alone as a benchmark. We evaluated sequence coverage as the primary metric to assess performance. The SP-MEGD method achieved approximately 100% sequence coverage for the heavy chain and 99.5% for the light chain, surpassing results obtained from the EasyPept-Micro kit (Fig. 2i ). As the SP-MEGD method is a five-protease gradient digestion approach, we also assessed additional enzymes with complementary specificities (chymotrypsin, pepsin, elastase, and Asp-N) using the commercial EasyPept-Micro kit. The results indicated that, except for chymotrypsin and pepsin, this kit exhibited suboptimal performance in terms of sequence acquisition coverage when employing other enzymes (Fig. 2i , and Supplementary Fig. 1f ). Notably, peptide information derived from Asp-N within this kit was substantially lower than that obtained with other enzymes (Student’s t -test, p < 0.05); this discrepancy may be attributed to compatibility issues. Furthermore, we combined all results from various enzymes within the kit to determine the sequence of S2P6. As a result, both peptides and PSMs obtained through the SP-MEGD method exceeded those acquired via the EasyPept-Micro kit (Fig. 2j , and Supplementary Fig. 1g ). Additionally, when the EasyPept-Micro kit was utilized in conjunction with other commercial proteases, the resulting coverage of CDR within HC in S2P6 was consistently demonstrated inferior performance compared to that achieved with SP-MEGD (Fig. 2j ). This diminished sequence coverage may hinder effective assembly of peptide information and subsequently affect sequencing accuracy (Fig. 2j ). These findings demonstrate that our innovative SP-MEGD strategy is more versatile and efficient compared to the commercially available alternatives, such as the EasyPept-Micro kit. De novo peptide sequencing with Casanovo Therefore, we implemented the SP-MEGD strategy for sample preparation and subsequently trained a Casanovo model 24 , 25 for de novo peptide sequencing of MS/MS spectra. The Casanovo model rapidly converged within the first ten epochs (Supplementary Table 1 ), exhibiting decreasing training and validation losses alongside increasing amino acid (AA) precision, AA recall, and peptide recall. The best-performing model, selected at epoch 10, achieved validation values of 0.7977 for AA precision, 0.7979 for AA recall, and 0.6006 for peptide recall. We further evaluated the trained Casanovo model on multiple independent external datasets, including eight previously published species-specific datasets and an SP-MEGD-based monoclonal antibody dataset (mAbs-known). Casanovo exhibited strong and consistent performance across diverse species (Supplementary Table 2 ), with AA precision and recall ranging from 0.7628 to 0.8735, while peptide recall varied from 0.5060 to 0.7430. Performance on the antibody-specific dataset (mAbs-known) also remained robust; it yielded an AA precision of 0.7279, an AA recall of 0.7257, and a peptide recall of 0.5547. These results confirm the reliability and effectiveness of the model in antibody-focused sequencing tasks (Supplementary Table 2 ). Consequently, the epoch-10 model was selected as the final model used for XA-Novo. Next, we assessed the assembly potential of the SP-MEGD strategy combined with the Casanovo model utilizing human COVID-19 neutralizing antibodies SA55 and SA58 26 , together with S2P6 tested at three input amounts (50 µg, 100 µg, and 200 µg), as well as three murine COVID-19 NAb datasets (85F7, 36H6, and 2B4) 27 . Combined with five proteases and utilizing HCD-based MS analysis, we collected the following peptide reads (defined as peptides with a score ≥ 50): 12,423 for SA55, 23,998 for SA58, 15,303 for 50 µg S2P6, 15,851 for 100 µg S2P6, 21,967 for 200 µg S2P6, 27,218 for 36H6, 29,993 for 85F7, and finally, 32,862 for 2B4 (Fig. 2k ). As illustrated in Supplementary Fig. 2 , the sequence coverage achieved 100% for both heavy and light chains across variable and constant domains. Furthermore, the mean depth of coverage within the CDRs, variable domain, and overall chain for both heavy and light chains ranged from 56.12 to 212.44, 76.67 to 216.49, and 107.58 to 283.70, respectively (Fig. 2l , Supplementary Fig. 3 , and Supplementary Table 3 ). Consequently, our pre-trained Casanovo model not only exhibited promising results in predicting peptide sequences and provided sufficient peptide information for complete assembly, but also demonstrated the effectiveness of the SP-MEGD method in yielding comprehensive ion information and complete sequence coverage. Fusion assembler enables accurate full-length sequence reconstruction The complete assembly of antibody sequences poses a significant challenge in de novo sequencing. Current algorithms, such as ALPS 28 and Stitch 15 , 29 , struggle with accuracy and robustness for routine applications. While commercial software such as PEAKS AB represents advanced solutions in the field, it does not support the reconstruction of multiple monoclonal antibodies from a mixed sample. To overcome these issues, we developed the Fusion assembler, which employs a beam search strategy with a beam size of 5. This approach incorporates a seven-amino-acid sliding window approach and mass-based alignment 29 distance metrics, while ensuring signal consistency for template-assisted assembly of mAbs as well as mixtures containing various mAbs. All details about the Fusion assembler can be found in the Supplementary Methods. As illustrated in Fig. 1a , the beam search strategy enhances Fusion to find the correct sequence during the assembly. It extends the contig bidirectionally during CDR retrieval, enabling it to restructure complete and accurate fragments in both heavy and light chains. The Fusion algorithm identifies Leu and Ile residues by integrating template sequence information with experimental evidence and detecting diagnostic w-ions, thus enhancing the precision of isomer discrimination. Stitch 29 is a state-of-the-art public method for antibody sequence assembly; however, its use requires users to select parameters such as peptide cutoff scores, segment orders, and software versions. These choices may necessitate careful consideration and, in some cases, expert knowledge to achieve optimal results. During the first run (Fig. 3a ), all assembly results for the light chain exhibited gaps due to the actual length of light chain CDR3 for SA55 and S2P6 antibodies exceeding that of the template. Therefore, we modified the recombined segment order of the light chain to “IGLV*IGLJ IGLC” for these two antibodies. Upon refining the Stitch run to enforce selection of “IGLV*IGLJ IGLC” as the recombined segment order for light chain, nearly all CDRs in human antibody datasets achieved an accuracy improvement to 100%, with one exception being observed in SA55 when processed using Stitch1.4.0 with a 95-peptide cutoff score (Fig. 3b ). As a result, there was also an overall enhancement in accuracy for the light chain. In our analysis across five human antibody datasets, varying peptide cutoff score settings substantially influenced both CDR accuracy and overall chain accuracy (Fig. 3b ). This phenomenon was similarly noted in three murine antibody datasets (Fig. 3c ). Notably, while the actual length of the light chain CDR3 from antibody 85F7 is shorter than that of its corresponding template, all assembled results produced a CDR3 sequence matching that of the template length (see Methods, Supplementary Data 1 ). On the contrary, Fusion employs an anchor-guided, template-adaptive assembly driven by k-mer overlap evidence and a local beam search toward conserved targets (see the Supplementary Methods, Supplementary Table 4 for more details). It does not require users to pre-specify the orders of recombined segments; instead, it infers CDR boundaries and lengths directly from data using 7-mer anchors flanking each CDR and de novo extension between them. Even when the true CDR is longer or shorter than the template (as in SA55/S2P6-like or 85F7-like scenarios), Fusion yields consistently accurate CDR reconstructions. Fig. 3. Comparative performance of Fusion, PEAKS AB, and Stitch in antibody sequence assembly. Open in a new tab a Initial analysis of human COVID-19 neutralizing antibodies using the Stitch algorithm with default segment orders across different peptide score thresholds and software versions revealed consistent assembly gaps in the light chain. b Reordering the segment arrangement for light chains in SA55 and SA58 to improve Stitch performance. c Stitch was also applied to murine COVID-19 neutralizing antibodies, revealing variability in assembly quality. d Accurate coverage comparison of Fusion, PEAKS AB, and optimized Stitch across light and heavy chains, including CDRs and full chains, across five datasets derived from three human-derived COVID-19 neutralizing antibodies. (Created in BioRender. XIONG, Y. (2026) https://BioRender.com/5zwsram ). e Assembly performance on three murine-derived COVID-19 neutralizing antibodies, demonstrating that Fusion consistently achieved the highest accurate sequence coverage across all evaluated regions (LC/HC CDR, VJ, and VDJ segments). (Created in BioRender. XIONG, Y. (2026) https://BioRender.com/5zwsram ). Source data are provided as a Source Data file. To comprehensively assess the assembly performance of Fusion, we conducted a comparative analysis with the best results from Stitch and PEAKS AB software. Across five datasets derived from three human COVID-19 NAbs with known sequences, Stitch exhibited variability in its sequence coverage accuracy. The light chain accurate coverage ranged from 98.61% to 100%, while the heavy chain accurate coverage varied between 97.58% and 99.78% (Fig. 3d ). Notably, Stitch was unable to assemble the complete heavy chain of the SA55 antibody, attaining only 76.47% accuracy for the CDR region and 97.58% for the overall chain (Fig. 3d , and Supplementary Fig. 4 ). Additionally, the light chain of the SA58 antibody exhibited gaps due to the actual length of CDR1 exceeding that of the template, attaining only 85% accuracy for the CDR region and 98.61% for the overall chain (Fig. 3d , and Supplementary Fig. 4 ). For PEAKS AB, the SA58 light chain achieved 100% accuracy for the CDR region and 99.07% for the whole chain, while the heavy chain showed only 68.97% accuracy in the CDR region, mainly due to 8 amino acid gaps in CDR3. The accuracy in the heavy chain was 97.33% with multiple amino acid gaps and insertions (Fig. 3d , and Supplementary Fig. 4 ). In contrast, Fusion consistently achieved 100% accurate coverage without any gap and insertion for both heavy and light chains across all datasets, demonstrating superior robustness even when faced with varying sample quantities in the S2P6 antibody datasets (Fig. 3d , and Supplementary Fig. 5 ). In the context of antibody reverse engineering and re-expression, sequence accuracy thresholds of ≥99% are widely considered essential. Sub-percent errors, particularly those involving insertions or deletions within CDR, can significantly compromise binding affinity or adversely affect product quality 30 – 32 . We also conducted a comparative analysis using three murine COVID-19 NAb datasets with known variable region sequences (85F7, 36H6, and 2B4) 27 . Fusion demonstrated superior performance compared to both Stitch and PEAKS AB, achieving 100% accurate coverage of CDR regions for all antibodies (Fig. 3e ). In contrast, the performance of Stitch and PEAKS AB showed considerable variability across different antibodies. The accurate coverage of Stitch for light chain variable regions fluctuated between 99.06% and 100%, while for heavy chain variable regions, it ranged from 95.80% to 99.17%. Notably, Stitch was unable to fully assemble the heavy chain CDR3 sequence of the 85F7 antibody due to amino acid gaps (Supplementary Fig. 6 ). PEAKS AB also displayed inconsistency; its accurate coverage for light chain variable regions varied from 86.73% to 100%, whereas for heavy chains, it ranged from 97.48% to 100%. Furthermore, PEAKS AB failed to assemble the complete heavy chain CDR3 sequence of the 2B4 antibody as well as the light chain CDR3 sequence of the 36H6 antibody due to amino acid errors, gaps, and insertions (Supplementary Fig. 6 ). These findings suggest that, although the accurate coverage rate may approach 95%, the presence of amino acid errors, gaps, and insertions in the antibody sequence prevents the reliable expression of true protein products, and further demonstrate the reliability of Fusion and de novo sequencing of antibodies based on bottom-up mass spectrometry. In conclusion, this comparative study underscores that Fusion outperforms both Stitch and PEAKS AB in assembly performance, particularly regarding accuracy and consistency across different samples and critical regions such as CDRs, making it a more reliable tool in our solution for precise antibody sequence analysis and research. Application of commercial immunotherapeutic monoclonal antibodies with unknown sequence Cancer immunotherapy is rapidly advancing and has now been recognized as the “fifth pillar” of cancer therapy 33 . By elucidating the sequence of immune-associated antibodies, it becomes feasible to design more effective targeted antibodies that precisely interact with immune cells within the tumor microenvironment or disrupt mechanisms of tumor-associated immune evasion, thereby enhancing the efficacy of immunotherapy. To demonstrate the effectiveness of our solution (XA-Novo) for therapeutic antibody research, we analyzed six immune-associated antibodies for which sequence information was previously unavailable: anti-CD4, anti-CD8α, anti-CSF1R, anti-IFNAR-1, anti-IFNγR, and anti-LTB4R mAbs. This expanded study served as a large-scale trial to determine whether XA-Novo could achieve results with high confidence comparable to those obtained for COVID-19 NAbs with known sequences. The ultimate goal was to use the deciphered sequences for gene cloning and functional expression of these antibodies. After de novo peptide sequencing with Casanovo, we collected 39,552 (anti-CD4), 16,856 (anti-CD8α), 33,857 (anti-CSF1R), 37,084 (anti-IFNAR-1), 35,757 (anti-IFNγR), and 25,458 (anti-LTB4R) peptide reads from HCD spectra. The Fusion algorithm successfully assembled these peptides into the complete sequences for both heavy and light chains. The mean depth of coverage in the CDRs, variable domain, and overall chain ranged from 33.69 to 340.52, 73.31 to 425.94, and 128.65 to 356.58, respectively (Supplementary Fig. 7 , and Supplementary Table 5 ). For comparative analysis, we also executed PEAKS AB to analyze these antibodies (Fig. 4 ). In contrast to the Fusion assembler, PEAKS AB obtained longer sequences for the heavy chains in anti-CD8α and anti-CSF1R antibodies (Fig. 4d , f). Furthermore, significant differences were observed in the amino acid sequence of these antibodies’ heavy chains, especially within the CDR3 region. We further validated this by performing an intact mass analysis. The results indicated that the theoretically calculated mass derived from peaks AB exhibited a significant deviation from the mass measured by MS of the anti-CSF1R (Supplementary Table 6 ). In contrast, the sequence assembled by Fusion was more closely aligned with the MS-measured mass of the source antibody (Supplementary Table 6 ). This finding suggests that Fusion holds promise for accurately characterizing the actual antibody sequences. The results of intact mass analysis also confirm that “VTAASGPEKVTTITLTPKVT” represents an insertion within the sequence generated by PEAKS AB, which will be deleted as a refined sequence during subsequent recombinant expression. Additionally, the ninth amino acid in the CDR3 region of the anti-CSF1R heavy chain was identified as isoleucine (Ile) by PEAKS AB and leucine (Leu) by Fusion (Fig. 4f ). Similarly, in the light chain of the anti-CD8α antibody, although the calculated mass was consistent in both methods, we observed a different presence of amino acid combinations (AG = Q) in the two methods (Fig. 4d ). Therefore, further validation using biochemical methods is necessary to confirm the accuracy of the sequences generated by the Fusion assembler. Fig. 4. Biological validation of recombinant antibodies generated by XA-Novo. Open in a new tab a Experimental procedure for in vivo antibody validation. C57BL/6 J mice ( n = 3 per group) received intraperitoneal injections of PBS (control), commercial antibodies, or antibodies reconstructed by XA-Novo or PEAKS AB on day 1. Blood and spleen samples were collected on day 3 and analyzed via flow cytometry to assess cell clearance. (Created in BioRender. XIONG, Y. (2026) https://BioRender.com/ena57ql ). b Reconstructed sequences of anti-mouse CD4 antibodies, including heavy and light chains, with complementarity-determining regions (CDRs) annotated and sequence discrepancies highlighted in red. c The representative charts of flow cytometry analysis of splenic CD4⁺ T cell depletion comparing control, commercial, and XA-Novo-generated anti-CD4 antibodies; quantification of clearance efficiency is shown on the right. All FACS detection data obtained from this experiment are presented in Supplementary Fig. 9 . d Sequence comparison of anti-mouse CD8α antibodies. e The representative charts of flow cytometry validation of CD8⁺ T cell clearance in spleen by commercial vs. XA-Novo-generated antibodies. All FACS detection data obtained from this experiment are presented in Supplementary Fig. 9. f Sequence alignment of anti-mouse CSF1R antibodies. g The representative charts of flow cytometry assessment of CSF1R⁺ macrophage depletion by commercial, XA-Novo-derived antibodies (125I), and XA-Novo-derived antibodies (125 L), with corresponding clearance efficiency quantified. All FACS detection data obtained from this experiment are presented in Supplementary Fig. 10 . Error bars are presented as mean ± SD. Source data are provided as a Source Data file. Subsequently, the sequences generated by Fusion were cloned and transfected into cells for recombinant expression, followed by affinity purification. The recombinant antibodies and matched commercial antibodies were then administered to C57BL/6 J mice ( n = 3 per group, Fig. 4a ). The third day post-injection revealed effective clearance of T-cells in both the spleen and blood. In contrast to control mice, which showed no depletion of T-cells, the XA-Novo-generated antibody demonstrated clearance efficiencies of 99.65% (anti-CD4) and 99.71% (anti-CD8α) in the spleen, as well as 99.78% (anti-CD4) and 99.72% (anti-CD8α) in the blood (Fig. 4c-e , Supplementary Figs. 8 , and 9 ). These results were consistent across all animals, as illustrated in Supplementary Fig. 9 . This highlights the reproducibility of clearance efficiency observed in each animal’s blood and spleen, confirming that the effects of XA-Novo-generated antibodies corresponded with those of commercial antibodies. We further investigated the role of Ile and Leu in the heavy chain CDR3 region of anti-CSF1R through recombinant expression (Fig. 4f , and Supplementary Fig. 10 ). The results demonstrated that the anti-CSF1R antibody containing Leu at the ninth amino acid position in the heavy chain CDR3 region exhibited enhanced macrophage depletion in the spleen (Fig. 4g , XA-Novo (125 L): 65.50%; XA-Novo (125I): 40.94%), highlighting the critical role of Ile and Leu identification in the CDR region and confirming the effectiveness of our method for the Ile/Leu differentiation. Furthermore, we observed that the clearance efficiency of antibodies generated by XA-Novo (L) in peripheral blood was slightly lower than that of antibodies produced by XA-Novo (I) (Fig. 4g , and Supplementary Fig. 10 ). This discrepancy may be attributed to the relatively low content of macrophages in peripheral blood, which is susceptible to dynamic fluctuations within the circulatory system 34 . Consequently, this makes it challenging to accurately assess the clearance effect of antibodies on tissue macrophages. The spleen is a crucial immune organ in mice, containing a substantial population of T cells, B cells, and macrophages 35 . Therefore, utilizing spleen results as primary conclusions while considering peripheral blood findings for dynamic monitoring will facilitate a more comprehensive understanding of the clearance mechanisms involved. Furthermore, the flow cytometry and Biacore analyses of the anti-LTB4R, anti-IFNAR-1, and anti-IFNγR antibodies further corroborated the accuracy of the sequences analyzed by Fusion (Supplementary Figs. 11 – 13 ). These discoveries provide biological validation for our solution, underscoring its effectiveness in accurately reconstructing sequences for unknown antibodies and its potential impact on antibody research and development. Application of a monoclonal antibody mixture for COVID-19 The COVID-19 pandemic, which is caused by the SARS-CoV-2 virus, continues to spread. NAbs have emerged as pivotal players in preventing and treating COVID-19 36 – 38 . However, the persistent emergence of new variants has resulted in widespread evasion of NAbs, presenting significant challenges to NAb-based therapeutics 26 . There is a high demand for the clinical development of NAb drugs resistant to future variants 39 , 40 . Rapidly determining the amino acid sequence of antibodies is critical for NAb drug discovery, enabling sequence-based recombination or engineering modifications that can alter antibody specificity and facilitate the design of diverse antibodies. To address this challenge, we next evaluated whether our sequencing platform could accurately reconstruct antibody sequences from mixtures, thereby mimicking the complexity encountered in therapeutic antibody discovery. In a mixture containing two antibodies, S2P6 and SA55, both reached an impressive sequence accuracy of 100%, with complete CDR region fidelity observed in both heavy and light chains (Supplementary Fig. 14a ). Notably, XA-Novo demonstrated consistent differentiation of I-L with 100% accuracy across all antibodies, irrespective of whether the mixture contained two or three antibodies. When analyzing a more complex combination comprising S2P6, SA58, and SA55 antibodies (Fig. 5b ), S2P6 continued to exhibit full sequence coverage with 100% accuracy across both chains and all CDR regions. For SA55 and SA58, while all CDR regions still displayed perfect accuracy at 100% in both heavy and light chains, there was a slight reduction in overall accuracy—99.78% for the heavy chain of SA55 and 99.54% for the light chain of SA58. Taking the results of other assemblers in traditional single-antibody sequencing as a baseline (Fig. 3 ), Fusion’s overall performance in antibody mixture sequence assembly remains superior to other assemblers, which are limited to sequencing single antibodies. These small discrepancies are attributed to shared template-derived positions: the last amino acid in the heavy chain FR4 region (TLVTVS S ASTKGPS) for SA58, and the 154th amino acid in the constant region of the light chain (VQWKVD N ALQSGN) for all three antibodies. Fusion’s assembler tolerates these minor deviations as they occur in structurally conserved regions with negligible impact on functional accuracy. Nevertheless, incorporating intact mass measurements for each antibody allows complete correction of these minor inconsistencies, yielding fully accurate sequences across all components. We further evaluated antibody mixtures derived from mouse species that included three distinct antibodies: 85F7, 36H6, and 2B4. Remarkably, our approach achieved 100% sequence accuracy across all three antibodies, with complete fidelity in both heavy and light chains, including all CDR regions. These results were consistent regardless of whether the mixture contained two or three antibodies (Fig. 5c , and Supplementary Fig. 14b–d ), and the reconstructed sequences were identical to those obtained using traditional single-antibody sequencing workflows. Collectively, these mixture experiments show that accurate sequence reconstruction is maintained as mixture complexity increases. The reconstructed sequences are consistent with those derived from conventional single-antibody workflows, indicating the platform’s utility for simultaneous analysis of multiple antibodies and enabling high-confidence antibody characterization. Fig. 5. Application of XA-Novo to mixtures of COVID-19 neutralizing antibodies. Open in a new tab a Schematic workflow for de novo sequencing of antibody mixtures. Antibody pools were processed using the SP-MEGD protocol, followed by LC-MS/MS analysis and peptide assembly via the Fusion assembler. Full-length heavy and light chain sequences were reconstructed, paired, and CDRs were annotated. (Created in BioRender. XIONG, Y. (2026) https://BioRender.com/reqr5kb ). b XA-Novo successfully resolved the antibody sequences in a mass-equivalent mixture of three human-derived antibodies (S2P6, SA55, SA58). The reconstructed sequences were consistent with those obtained from individual sequencing efforts, with minor discrepancies highlighted in red. c XA-Novo also deconvoluted mixtures of three murine-derived antibodies (85F7, 2B4, 36H6), with complete recovery of heavy and light chain CDRs and VJ/VDJ regions. Sequence identity was confirmed against reference data and single-antibody sequencing results. Web-based tools for antibody de novo sequencing analysis Lastly, we have developed a user-friendly, web-based tool for analyzing antibody de novo sequencing (Fig. 6 ). This tool leverages advanced deep learning models for peptide sequencing and incorporates a beam search-based Fusion assembler that we previously established. This resource is readily accessible at https://xa-novo.com/ . Fig. 6. Web-based platform for de novo antibody sequencing analysis. Open in a new tab Snapshots of the XA-Novo web interface illustrating key functionalities. The platform enables the upload of mass spectrometry data, execution of de novo peptide sequencing, full-length antibody reconstruction, and CDR annotation. Results are visualized with interactive sequence coverage depth and structural mapping, supporting both single antibodies and antibody mixtures. Illustration modified with permission from 51miz [ https://www.51miz.com/sucai/1289948.html ]. Researchers can perform their analyses with relative ease, eliminating the need for complex configurations or specialized hardware requirements. The only prerequisite is to provide a link to the mass spectrometry data, along with related species information regarding the detected antibody, to the XA-Novo platform. Discussion Antibody de novo sequencing represents a significant challenge within the realm of MS-based proteomics, necessitating advancements across various technologies to address its inherent difficulties. In this study, we present XA-Novo, an integrated solution for rapid and precise de novo sequencing, comprising an SP-MEGD sample preparation method to provide comprehensive peptide-level information, LC-MS/MS analysis, deep learning-based de novo peptide sequencing of individual spectra, and an innovative Fusion assembler to assemble overlapping peptides into complete sequences. Our solution is applicable for both traditional mAb sequencing and high-throughput mixture sequencing of multiple antibodies, demonstrating high accuracy and robustness throughout the process. For traditional mAb de novo sequencing, three immunotherapy-related commercial antibodies demonstrated that the sequences deciphered by our solution could be utilized to produce genes for cloning and expression of functional antibodies, facilitating the acquisition of precise antibody sequence information for routine applications. The COVID-19 pandemic has underscored the promise of mAb-based drugs as both prophylactic and therapeutic interventions for infectious diseases. Rapid acquisition of antibody sequences and evolutionary engineering are essential in effectively addressing the ongoing emergence of new variants, particularly given the time constraints involved. In this context, the throughput and accuracy of de novo sequencing pose considerable challenges. However, through our efforts, XA-Novo enables simultaneous discrimination among two or three COVID-19-NAbs, with large-scale experimental trials achieving an accuracy rate of at least 99.54%. This methodology substantially enhances throughput and cost-effectiveness by allowing concurrent sequencing of multiple mAbs within a single experimental framework, thereby eliminating the necessity for separate sequential processes. This optimization not only decreases instrument runtime and reagent consumption but also lowers operational costs. This represents a significant advancement in antibody research methodologies by enhancing efficiency and reducing expenses. These points highlight the multifaceted approach and technological advancements that have contributed to the success of antibody de novo sequencing. Moving forward, we aim to refine the sample preparation methods, mass spectrometry technologies, and Fusion assembler to enhance their utility in supporting broader initiatives in antibody research and development. Furthermore, achieving complete antibody de novo sequencing for non-model organisms remains a significant challenge. To date, no research has successfully accomplished this without necessitating additional experimental data. Nevertheless, ongoing advancements in the field (including enhanced MS detection precision, improved accuracy in long-peptide de novo sequencing, and reduced fragmentation ambiguity) may ultimately facilitate the comprehensive assembly of antibody sequences in a template-independent manner, even for non-model organisms. Methods Antibody reagents and production The anti-mouse CD4 (clone# GK1.5), anti-mouse CD8α (clone# 2.43), anti-mouse CSF1R (clone# AFS98), anti-mouse IFNAR-1(clone# MAR1-5A3), and anti-mouse IFNγR (also known as IFNGR; clone# GR-20) mAbs were purchased from BioXCell (BE0003-1, BE0061, BE0213, BE0241, and BE0029). The anti-mouse Ly-6G (clone# 1A8) mAb was purchased from Leinco (L280-25mg). The anti-human LTB4R (Mouse, clone# 7B1) mAb was purchased from Thermo Fisher (MA1-20232). Recombinant human COVID-19 NAbs (S2P6, SA55, and SA58, in human IgG1 backbone) were expressed by the ExpiCHO-S™ suspension expression system (Thermo Fisher Cat#A29133) and subsequently purified using MabSelect SuRe resin (Cytiva). The amino acid sequences of these mAbs in our expression plasmids are shown in Supplementary Tables 7 and 8 . Of note, compared with the published sequence, the SA58 construct used in this study contained an I51V mutation in the H-chain and a D154S mutation in the L-chain (as indicated in Supplementary Table 7 ). Three murine COVID-19 NAbs (85F7, 36H6, and 2B4) were purified from ascitic fluid derived from hybridoma cells using MabSelect SuRe resin. The variable regions of these mouse mAbs were sequenced via RT-PCR using mRNA isolated from the corresponding hybridoma cells and were presented in Supplementary Table 8 . Sample preparation based on the SP-MEGD method Initially, the lysis buffer [6 M guanidine hydrochloride (Gu·HCl), 20 mM dithiothreitol (DTT), 100 mM Tris, pH 8.5] was added to 200 μg of purified mAb at an approximate ratio of 1:1 (microliters of lysis buffer to micrograms of protein). The mixture was denatured and reduced at 60°C for 30 min, followed by alkylation with 40 mM iodoacetamide (IAA) in the dark for 30 min at room temperature. The alkylated antibody samples were transferred through a 10 kDa filter (Millipore, USA) at 4 °C into a solution containing 0.8 M urea, 50 mM Tris-HCl, adjusted to pH 8.0, to eliminate interfering substances that could affect enzymatic hydrolysis. Proteases were then added at a mass ratio of 1:20, specifically trypsin, chymotrypsin, pepsin (with the solution’s pH pre-adjusted to below 3), elastase, and aspartic acid protease. A separate reaction vessel was utilized for each enzyme and incubated at 37 °C for a total duration of six hours, with samples collected at two-hour intervals. After digestion, the reaction mixture was quenched with 10% trichloroacetic acid (TFA) for 30 min at 37 °C. The supernatant was desalted by Sep-Pak C18 Vac cartridges (Waters) following the instructions before being lyophilized to dryness or stored at −20 °C or dissolved in 0.1% FA for LC − MS/MS analysis. LC-MS/MS analysis The digested peptides were separated through reversed-phase chromatography on a Vanquish™ Neo UHPLC (column packed with PepMap™ 100 C18, 75 μm × 50 cm, 2 μm, Thermo Fisher Scientific, USA) coupled to an Orbitrap Eclipse mass spectrometer. Samples were eluted over a 70-min gradient from 0 to 35% solvent B (0.1% formic acid in 80% acetonitrile) at a flow rate of 300 nL/min. Solvent A was 0.1% formic acid in water. Full MS1 scans were acquired over a range of m/z 350−2000 with a resolution of 120, 000. MS1 scans were obtained with a standard automatic gain control (AGC) target and a maximum injection time of 100 ms. The precursors were fragmented by stepped HCD as well as EThcD. The stepped HCD fragmentation included steps of 27%, 35%, and 40% normalized collision energies (NCE). EThcD fragmentation was performed with calibrated charge-dependent electron-transfer dissociation (ETD) parameters and 30% NCE supplemental activation. For both fragmentation types, MS2 scans were acquired at a 30,000 resolution, a 5E4 AGC target, and a 250 ms maximum injection time. Other method details are presented in the Supplementary Information. In total, thirteen bottom-up mass spectrometry datasets were generated for this study, based on traditional mAb sequencing (Supplementary Tables 7 and 8 ) or mAb mixture sequencing (Supplementary Table 9 ). Deep learning model for de novo peptide sequencing of MS/MS spectra Casanovo 24 uses a transformer architecture to treat de novo peptide sequencing as a sequence-to-sequence translation task, translating from the series of peaks in an MS2 spectrum to a series of amino acids. The model consists of an encoder and a decoder, each based on transformer architecture. The encoder is responsible for learning an in-context representation of the input MS2 spectrum, while the decoder predicts the subsequent amino acid in the peptide sequence based on the given spectrum representation and previously predicted amino acids. Both the encoder and decoder are composed of nine layers with eight heads per layer, featuring a hidden dimension of 512 and a feedforward dimension of 1024. Similar to other deep learning models, Casanovo predicts a peptide sequence one amino acid at a time, using beam search decoding to identify the predicted peptide sequence that achieves the highest score 25 . Notably, no post-processing step is implemented to enforce alignment between the predicted peptide mass and observed precursor mass; however, there exists an optional filter that penalizes predictions deviating from precursor mass tolerance by assigning them negative scores. We trained the Casanovo model (v3.2.0) for 20 epochs with its pre-defined parameters on the MassIVE KB v1 dataset 41 , which encompasses over 2.1 million precursors from 19,610 proteins. This dataset was meticulously assembled from more than 31 TB of human data, sourced from 227 public proteomics datasets. Notably, it maintains rigorous false discovery rate controls, ensuring a comprehensive and robust training environment for our model. After filtering PSMs with unrelated variable modifications, we obtained a total of 1,909,041 PSMs corresponding to 1,376,990 unique peptides for the training set and 38,937 PSMs corresponding to 28,100 unique peptides for the validation set. It is important to note that no common peptides were shared between these split datasets. We employ a training-validation split of 98:2, with the 98% training proportion aligning with the practices reported by Beslic et al. 42 in their recent comprehensive evaluation of peptide de novo sequencing tools for monoclonal antibody assembly. During training, 10 ppm for the precursor mass tolerance and 0.02 Da for the fragment mass tolerance were specified. Carbamidomethylation of cysteine (C + 57.02 Da) was set as a fixed modification, while oxidation of methionine (M + 15.99 Da), deamidation of asparagine, glutamine (N + 0.98 Da and Q + 0.98 Da), and Pyroglutamate formation from glutamine (Q−17.03 Da) were set as variable modifications. To further validate the generalizability of our model, we conducted extensive benchmarking on independent external datasets: (A) Nine-species benchmark dataset (Wen et al., 2023 43 ; originally by DeepNovo 44 ): This dataset was initially introduced by DeepNovo and subsequently revisited in detail by Wen et al. (2023), who systematically reprocessed the data with stringent PSM-level false discovery rate (FDR ≤ 1%) controls using the Tide and Percolator search engines. Wen et al. specifically highlighted: “Finally, because some of the single-species datasets are markedly larger than others, we produced a more balanced version of the dataset”. We utilized this carefully balanced version of the eight-species benchmark dataset to ensure a fair, unbiased, and comprehensive evaluation across different species while explicitly excluding the human dataset (PXD004424) from performance interpretation due to its overlap with our training set. (B) mAbs-known dataset (internally generated): This dataset comprises monoclonal antibodies with determined sequences that were generated in this study (SA55, 50 µg S2P6, 100 µg S2P6, 200 µg S2P6, 85F7, 36H6, 2B4). Each spectrum was annotated using the PEAKS DB database search engine to provide high-confidence peptide sequence labels, enabling a direct evaluation of Casanovo’s performance on antibody-specific spectra relevant to our application. In addition to the training-validation split, we conducted a 50-fold peptide-level cross-validation on the same dataset to evaluate the stability of model performance. Each fold was trained from scratch, and we observed consistent results across all folds (Supplementary Fig. 15 ), with an AA precision of 0.7969 ± 0.0054, AA recall of 0.7953 ± 0.0058, and peptide recall of 0.5928 ± 0.0060. These results were consistent with those obtained from our initial 98:2 split, further reinforcing the robustness of the model. Furthermore, to assess the generalizability of the de novo peptide sequencing model used in XA-Novo and to exclude the possibility that its performance arises from a favorable train–validation split, we evaluated all 50 cross-validation models on independent external datasets. Across these datasets, which span different species and antibody repertoires, performance remained consistent with that observed on the internal data (Supplementary Fig. 15 ), indicating that the results are not driven by any particular data partition. Each instrument vendor employs its proprietary file formats to store results from MS/MS experiments, requiring conversion to open-format files for compatibility with the pre-trained Casanovo model. We used ProteoWizard software 45 to reformat the raw MS/MS data files of each antibody dataset into Mascot Generic Format (MGF). MGF files store the m/z and intensity pairs of multiple mass spectra within a single text format. Afterward, Casanovo deciphers amino acid sequences by interpreting the mass variations observed in the MS/MS peaks. Once peptide sequencing is completed, we reformat the output to conform to the PEAKS X+ style, which is compatible with downstream assembly tools such as Stitch 1.4, Stitch 1.5, and Fusion. Both sequence-level and amino acid-level scores are normalized on a scale ranging from 0 to 100. An example of an annotated MGF file entry for our trained Casanovo model is as follows: “BEGIN IONS TITLE=My spectrum title PEPMASS = 602.2881 CHARGE = 2+ RTINSECONDS = 985.44604 SEQ = HQ(-17.03)GVM( + 15.99)VGC( + 57.02)GQKQ( + .98)MVN( + .98)K (without SEQ entry for de novo mode) 84.08081817626953 25805.216796875 87.74003601074219 69811.5234375… END IONS” Sequence assembly De novo peptides were then used by the Fusion assembler to automatically reconstruct the complete sequences of mAb and mixtures of various antibodies. Human or mouse germline antibody sequences from IMGT are employed as templates to guide the assembly of de novo peptide reads. Intact mass verification of the antibody light and heavy chains One hundred micrograms of each antibody were reduced with 50 mM DTT and treated with 2 µL PNGase F (Promega, USA) at 37 °C for over 16 h with stirring before being transferred to a new vial for LC-MS analysis. More experimental conditions are given in the Supplementary Methods. LC-MS analysis of intact light and heavy chains Intact LCs and HCs were analyzed using an Orbitrap Eclipse mass spectrometer (Thermo Fisher Scientific, USA) coupled to a Vanquish™ Flex UHPLC (Thermo Fisher Scientific, USA). All details can be found in the Supplementary Methods. Data analysis During the development and optimization of the SP-MEGD method, for bottom-up analysis, the raw files obtained for digesting peptides of each antibody were searched with PEAKS AB (v.3.0) using the parameters described in the Supplementary Methods. The de novo sequencing analysis of mAbs and their mixtures in our solution was conducted utilizing the XA-Novo platform ( https://xa-novo.com/ ). The intact mass spectra were processed and deconvoluted using BioPharma Finder 5.2 software (Thermo Fisher Scientific, USA), subsequently enabling the comparison of mass information for light and heavy chains. Expression and purification of anti-CD4, anti-CD8α, and anti-CSF1R antibodies The mAb sequences obtained from our solution were utilized to generate antibodies in ExpiCHO-S cells following the manufacturer’s instructions. The proteins were purified from culture supernatants collected on day 7 after transfection using a Protein A column (Cytiva). Elution of the protein from the column was carried out with 70 mM citric acid monohydrate and 20 mM sodium phosphate dibasic dodecahydrate. The eluted fractions were pooled and concentrated using the Protein Quantification kit (Thermo Fisher Cat#23225). Functional validation of anti-CD4, anti-CD8α, and anti-CSF1R antibodies Three groups of 8-week-old female C57BL/6 J mice were administered intraperitoneally with either 200 μg of an XA-Novo-generated antibody ( n = 3 per group), 200 μL of a commercially available antibody ( n = 3 per group), or 200 μL of phosphate-buffered saline (PBS) ( n = 6 per group). The objective was to evaluate the effective clearance rates of CD8 + and CD4 + T cells in their peripheral blood mononuclear cells (PBMCs) and spleens (Supplementary Fig. 9 ). All groups of mice were matched for age. Similarly, four additional groups of 8-week-old female C57BL/6J mice were administered intraperitoneally with either 200 μg of XA-Novo-generated antibody (125I, n = 4 per group), 200 µg of XA-Novo-generated antibody (125 L, n = 3 per group), 200 μg of a commercially available CSF1R antibody ( n = 2 per group), or 200 μL of PBS buffer ( n = 3 per group). The aim was to assess the effective clearance rate of macrophages in their PBMCs and spleens (Supplementary Fig. 10 ). The mice received antibody or PBS buffer injections on the first day. On the third day post-injection, single cells were isolated from peripheral blood or spleen tissue and resuspended in an appropriate volume of PBS buffer. Cell counting was performed, followed by treatment with a blocking buffer (Purified Rat Anti-Mouse CD16/CD32, Mouse BD Fc Block™, BD, Cat. No. 553141) to reduce nonspecific antibody binding. Subsequently, the cell suspension was incubated with fluorescently labeled antibodies: anti-CD4 (PerCP-Cy5.5-anti-CD4), anti-CD8α (FITC-anti-CD8α), and anti-CSF1R (PE-anti-F4/80). After incubation, the cells were washed with PBS buffer to remove unbound antibodies. The labeled cell suspension was then processed using a flow cytometer to measure the fluorescence intensity of the labeled dyes. Utilizing this data, CD4 + T cells, CD8 + T cells, and macrophages were sorted. Finally, FlowJo (BD Biosciences, version 10.10) was used to analyze the data and assess the quantity, activity, and expression of CD4 + T cells, CD8 + T cells, and macrophages. All details can be found in the Supplementary Methods. Animal research All animal experiments were approved by the Institutional Animal Care and Use Committee of Xiamen University (protocol number: XMULAC20230310) and conducted in accordance with the Guide for the Care and Use of Laboratory Animals. Female C57BL/6J mice were used in this study. Female mice were selected to minimize inter-individual variability associated with stress and aggression, which are more commonly observed among male littermates, and to ensure greater behavioral consistency across experimental groups. A total of 36 eight-week-old mice were included. Animals were housed under a standard 12‑h light/dark cycle (lights on at 07:00) with ad libitum access to standard chow and water. All procedures were designed to minimize animal suffering, and the number of animals used was reduced to the minimum required to achieve statistically robust conclusions. No wild animals or field-collected biological samples were used in this study. Assembly of Stitch for known monoclonal antibodies Stitch has recently emerged as a state-of-the-art public method for assembling antibody sequences. Users are required to define three main parameters: the cutoff score for peptide reads, the placement cutoff score, and the recombined segment orders. The typical recombined segment orders are IGHV * IGHJ IGHC for the heavy chain and IGLV IGLJ IGLC for the light chain. The asterisk (*) represents a gap between adjacent segments, which is extended by overhanging peptide reads. We have chosen cutoff scores of 95, 90, 85, and 50 for peptide reads. Scores of 95, 90, and 85 are based on the default parameters from Stitch’s paper or batch file examples, while 50 is used as the threshold in Fusion. We conducted tests using both the stable version of Stitch 1.4.0 and the latest version of Stitch 1.5.0 with these peptide cutoff scores, common recombined segment orders, and other default parameters. Evaluation metrics The accepted evaluation metrics in de novo peptide sequencing include amino acid-level precision (the proportion of correctly predicted amino acids among all predicted amino acids), amino acid-level recall (the proportion of correctly predicted amino acids relative to the total number of ground-truth amino acids), and peptide-level recall (the proportion of correctly predicted complete peptide sequences compared to the total number of ground-truth peptide sequences). Antibody sequence assembly was evaluated based on three key metrics: coverage, accuracy, and accurate coverage (Fig. 2a ). Sequencing coverage was calculated as the percentage of amino acids of the antibody sequence that were covered by the assembled contig. Sequencing accuracy was calculated as the percentage of all annotated sequence calls that were labeled correctly. To comprehensively evaluate the performance of sequence assembly, we employed the metric of accurate coverage, defined as the percentage of amino acids of the antibody sequence that were correctly covered by the assembled contig, equal to the product of coverage and accuracy. Since Fusion achieved 100% sequence coverage for all antibodies, in this context, accurate coverage is equivalent to accuracy. For the alignment between assembled sequences and actual antibody sequences, we utilized the Needleman-Wunsch algorithm 46 in conjunction with the BLOSUM60 matrix to achieve the most intuitive results. Ethical Statement Animal experiments were carried out in accordance with the approval of the Institutional Animal Care and Use Committee at Xiamen University (XMULAC20230310) and by the Guide for the Care and Use of Laboratory Animals. Supplementary information Supplementary Information (4.4MB, pdf) 41467_2026_70496_MOESM2_ESM.pdf (73.9KB, pdf) Description of Additional Supplementary Files Supplementary Data 1 (27KB, xlsx) Summary (152.3KB, pdf) Transparent Peer Review file (9.4MB, pdf) Source data Source Data 1 (31.8KB, xlsx) Source Data 2 (4.1MB, xlsx) Source Data 3 (1.8MB, xlsx) Source Data 4 (27.5KB, xlsx) Source Data 5 (576.5KB, xlsx) Source Data 6 (10.3KB, xlsx) Source Data 7 (10KB, xlsx) Source Data 8 (28.4KB, xlsx) Source Data 9 (27.1KB, xlsx) Source Data 10 (14.6KB, xlsx) Source Data 11 (11.8KB, xlsx) Acknowledgements This work was supported by the National Science and Technology Major Project for Innovative Drug Research and Development (2025ZD1803703 to Q.Y.); the National Natural Science Foundation of China (32401237 to Y.X. and 92369110 to Q.Y.); the Fujian Provincial Natural Science Foundation of China (2024J08358 to Y.X.); the Natural Science Foundation of Xiamen, China (3502Z202371039 to Y.X.); the State Key Laboratory of Vaccines for Infectious Diseases, Xiang An Biomedicine Laboratory (2025XAKJ0200001 to R.Y.); and the Scientific Research Foundation of the State Key Laboratory of Vaccines for Infectious Diseases (2024SKLVDzy06 to Y.X.). Author contributions Y.X., W.J. and J.X. contributed equally to this work. W.J., Y.X., R.Y., N.X. and Q.Y. conceived and designed the study. J.X. and J.W. conducted mass spectrometry experiments. Q.B., X.C. and Y.W. prepared recombinant antibodies and performed animal experiments. W.J. developed the algorithms. Z.J., L.L., Y.Q. and F.L. performed data analysis. Y.X. and W.J. wrote the draft of the manuscript. R.Y., N.X. and Q.Y. revised and edited the manuscript. All authors read and approved the final version of the manuscript. Peer review Peer review information Nature Communications thanks Albert Heck, Samantha Sarrett, and the other anonymous reviewer(s) for their contribution to the peer review of this work. A peer review file is available. Data availability The mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium via the partner repository iProX under the dataset identifier PXD060500 . The weights of the Casanovo model utilized in XA-Novo and the split datasets for training Casanovo are both available on Zenodo [ https://zenodo.org/records/17266057 ] and [ https://zenodo.org/records/18627093 ], respectively. Additionally, the revised nine-species benchmark dataset by Wen et al . is available on Zenodo [ https://zenodo.org/records/13653420 ]. Unless otherwise stated, all data supporting the results of this study can be found in the article, supplementary, and source data files. Source data are provided with this paper. Code availability Result files and code to reproduce the results in this study are available on GitHub [ https://github.com/biocc/SP-MEGD_Fusion ]. An executable version of the code and the computational environment used in this study are also available as a Code Ocean capsule [ https://codeocean.com/capsule/4653442/tree ]. Competing interests R.Y. is a shareholder of Aginome Scientific. The remaining authors declare no competing interests. Footnotes Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. These authors contributed equally: Yueting Xiong, Wenbin Jiang, Jin Xiao. Contributor Information Yueting Xiong, Email: [email protected]. Rongshan Yu, Email: [email protected]. Ningshao Xia, Email: [email protected]. Quan Yuan, Email: [email protected]. Supplementary information The online version contains supplementary material available at 10.1038/s41467-026-70496-y. References 1. Lu, L. L., Suscovich, T. J., Fortune, S. M. & Alter, G. Beyond binding: antibody effector functions in infectious diseases. Nat. Rev. Immunol. 18 , 46–61 (2018). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 2. Oostindie, S. C., Lazar, G. A., Schuurman, J. & Parren, P. Avidity in antibody effector functions and biotherapeutic drug design. Nat. Rev. Drug Discov. 21 , 715–735 (2022). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 3. Watson, C. T., Glanville, J. & Marasco, W. A. The Individual and Population Genetics of Antibody Immunity. Trends Immunol. 38 , 459–470 (2017). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 4. Ejazi, S. A., Ghosh, S. & Ali, N. Antibody detection assays for COVID-19 diagnosis: an early overview. Immunol. Cell Biol. 99 , 21–33 (2021). [ DOI ] [ PubMed ] [ Google Scholar ] 5. Ning, L., Abagna, H. B., Jiang, Q., Liu, S. & Huang, J. Development and application of therapeutic antibodies against COVID-19. Int J. Biol. Sci. 17 , 1486–1496 (2021). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 6. Lu, R. M. et al. Development of therapeutic antibodies for the treatment of diseases. J. Biomed. Sci. 27 , (2020). [ DOI ] [ PMC free article ] [ PubMed ] 7. Mason, D. M. et al. Optimization of therapeutic antibodies by predicting antigen specificity from antibody sequence via deep learning. Nat. Biomed. Eng. 5 , 600–612 (2021). [ DOI ] [ PubMed ] [ Google Scholar ] 8. Mattsson, J. et al. Sequence enrichment profiles enable target-agnostic antibody generation for a broad range of antigens. Cell Rep. Methods 3 , 100475 (2023). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 9. Parray, H. A. et al. Hybridoma technology a versatile method for isolation of monoclonal antibodies, its applicability across species, limitations, advancement and future perspectives. Int Immunopharmacol. 85 , 106639 (2020). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 10. Tomita, M. & Tsumoto, K. Hybridoma technologies for antibody production. Immunotherapy 3 , 371–380 (2011). [ DOI ] [ PubMed ] [ Google Scholar ] 11. Subas et al. NAb-seq: an accurate, rapid, and cost-effective method for antibody long-read sequencing in hybridoma cell lines and single B cells. MAbs 14 , 2106621 (2022). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 12. Chen, Y. et al. Barcoded sequencing workflow for high throughput digitization of hybridoma antibody variable domain sequences. J. Immunol. Methods 455 , 88–94 (2018). [ DOI ] [ PubMed ] [ Google Scholar ] 13. Schardt, J. S., Sivaneri, N. S. & Tessier, P. M. Monoclonal antibody generation using single B-cell screening for treating infectious diseases. BioDrugs 38 , 477–486 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 14. de Graaf, S. C., Hoek, M., Tamara, S. & Heck, A. J. R. A perspective toward mass spectrometry-based de novo sequencing of endogenous antibodies. MAbs 14 , 2079449 (2022). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 15. Schulte, D., Peng, W. & Snijder, J. Template-based assembly of proteomic short reads for de novo antibody sequencing and repertoire profiling. Anal. Chem. 94 , 10391–10399 (2022). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 16. Le Bihan, T. et al. De novo protein sequencing of antibodies for identification of neutralizing antibodies in human plasma post SARS-CoV-2 vaccination. Nat. Commun. 15 , 8790 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 17. Sen, K. I. et al. Automated Antibody De Novo sequencing and its utility in biopharmaceutical discovery. J. AM Soc. Mass Spectr. 28 , 803–810 (2017). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 18. Gadush, M. V. et al. Template-Assisted De Novo Sequencing of SARS-CoV-2 and Influenza Monoclonal Antibodies By Mass Spectrometry. J. Proteome Res . 21 , 1616–1627 (2022). [ DOI ] [ PubMed ] [ Google Scholar ] 19. He, M.-T. et al. Do-It-Yourself De Novo Antibody Sequencing Workflow That Achieves Complete Accuracy Of The Variable Regions. J. Proteome Res . 24 , 3062–3073 (2025). [ DOI ] [ PubMed ] [ Google Scholar ] 20. Pinto, D. et al. Broad betacoronavirus neutralization by a stem helix-specific human antibody. Science 373 , 1109–1116 (2021). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 21. Ye, X. et al. Integrated proteomics sample preparation and fractionation: Method development and applications. Trends Anal. Chem. 120 , 115667 (2019). [ Google Scholar ] 22. Tsiatsiani, L. & Heck, A. J. R. Proteomics beyond trypsin. FEBS J. 282 , 2612–2626 (2015). [ DOI ] [ PubMed ] [ Google Scholar ] 23. Morsa, D. et al. Multi-enzymatic limited digestion: the next-generation sequencing for proteomics? J. Proteome Res . 18 , 2501–2513 (2019). [ DOI ] [ PubMed ] [ Google Scholar ] 24. Yilmaz M., Fondrie W., Bittremieux W., Oh S., Noble W. S. De novo mass spectrometry peptide sequencing with a transformer model. Int. Conf. Mach. Learn . 162 , 25514-25522 (2022). 25. Yilmaz, M. et al. Sequence-to-sequence translation from mass spectra to peptides with a transformer model. Nat. Commun. 15 , 6427 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 26. Cao, Y. et al. BA.2.12.1, BA.4 and BA.5 escape antibodies elicited by Omicron infection. Nature 608 , 593–602 (2022). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 27. Wu, Y. et al. Lineage-mosaic and mutation-patched spike proteins for broad-spectrum COVID-19 vaccine. Cell Host Microbe 30 , 1732–1744.e1737 (2022). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 28. Tran, N. H. et al. Complete De Novo assembly of monoclonal antibody sequences. Sci. Rep. 6 , 31730 (2016). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 29. Schulte, D. & Snijder, J. A Handle on Mass Coincidence Errors in De Novo sequencing of antibodies by bottom-up proteomics. J. Proteome Res. 23 , 3552–3559 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 30. Guthals, A. et al. De Novo MS/MS sequencing of native human antibodies. J. Proteome Res . 16 , 45–54 (2017). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 31. Guthals, A., Clauser, K. R., Frank, A. M. & Bandeira, N. Sequencing-Grade De novo Analysis of MS/MS Triplets (CID/HCD/ETD) From Overlapping Peptides. J. Proteome Res. 12 , 2846–2857 (2013). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 32. Peng, W. et al. Reverse-engineering the anti-MUC1 antibody 139H2 by mass spectrometry-based de novo sequencing. Life Sci. Alliance 7 , e202302366 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 33. Chen, D. S. & Mellman, I. Oncology meets immunology: the cancer-immunity cycle. Immunity 39 , 1–10 (2013). [ DOI ] [ PubMed ] [ Google Scholar ] 34. Zhao, Y. L. et al. Comparison of the characteristics of macrophages derived from murine spleen, peritoneal cavity, and bone marrow. J. Zhejiang Univ. Sci. B 18 , 1055–1063 (2017). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 35. Bronte, V. & Pittet, M. ikaelJ. The spleen in local and systemic regulation of immunity. Immunity 39 , 806–818 (2013). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 36. Khoury, D. S. et al. Neutralizing antibody levels are highly predictive of immune protection from symptomatic SARS-CoV-2 infection. Nat. Med. 27 , 1205–1211 (2021). [ DOI ] [ PubMed ] [ Google Scholar ] 37. Cameroni, E. et al. Broadly neutralizing antibodies overcome SARS-CoV-2 Omicron antigenic shift. Nature 602 , 664–670 (2022). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 38. Tortorici, M. A. et al. Broad sarbecovirus neutralization by a human monoclonal antibody. Nature 597 , 103–108 (2021). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 39. Copin, R. et al. The monoclonal antibody combination REGEN-COV protects against SARS-CoV-2 mutational escape in preclinical and human studies. Cell 184 , 3949–3961.e3911 (2021). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 40. Cao, Y. et al. Rational identification of potent and broad sarbecovirus-neutralizing antibody cocktails from SARS convalescents. Cell Rep. 41 , 111845 (2022). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 41. Wang, M. et al. Assembling the Community-Scale Discoverable Human Proteome. Cell Syst. 7 , 412–421.e415 (2018). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 42. Beslic, D., Tscheuschner, G., Renard, B. Y., Weller, M. G. & Muth, T. Comprehensive evaluation of peptide de novo sequencing tools for monoclonal antibody assembly. Brief. Bioinform. 24 , bbac542 (2022). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 43. Wen, B. & Noble, W. S. A multi-species benchmark for training and validating mass spectrometry proteomics machine learning models. Sci. Data 11 , 1207 (2024). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 44. Tran, N. H., Zhang, X., Xin, L., Shan, B. & Li, M. De novo peptide sequencing by deep learning. Proc. Natl. Acad. Sci. USA 114 , 8247–8252 (2017). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 45. Kessner, D., Chambers, M., Burke, R., Agus, D. & Mallick, P. ProteoWizard: open source software for rapid proteomics tools development. Bioinformatics 24 , 2534–2536 (2008). [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 46. Needleman, S. B. & Wunsch, C. D. A general method applicable to the search for similarities in the amino acid sequence of two proteins. J. Mol. Biol. 48 , 443–453 (1970). [ DOI ] [ PubMed ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials Supplementary Information (4.4MB, pdf) 41467_2026_70496_MOESM2_ESM.pdf (73.9KB, pdf) Description of Additional Supplementary Files Supplementary Data 1 (27KB, xlsx) Summary (152.3KB, pdf) Transparent Peer Review file (9.4MB, pdf) Source Data 1 (31.8KB, xlsx) Source Data 2 (4.1MB, xlsx) Source Data 3 (1.8MB, xlsx) Source Data 4 (27.5KB, xlsx) Source Data 5 (576.5KB, xlsx) Source Data 6 (10.3KB, xlsx) Source Data 7 (10KB, xlsx) Source Data 8 (28.4KB, xlsx) Source Data 9 (27.1KB, xlsx) Source Data 10 (14.6KB, xlsx) Source Data 11 (11.8KB, xlsx) Data Availability Statement The mass spectrometry proteomics data have been deposited to the ProteomeXchange Consortium via the partner repository iProX under the dataset identifier PXD060500 . The weights of the Casanovo model utilized in XA-Novo and the split datasets for training Casanovo are both available on Zenodo [ https://zenodo.org/records/17266057 ] and [ https://zenodo.org/records/18627093 ], respectively. Additionally, the revised nine-species benchmark dataset by Wen et al . is available on Zenodo [ https://zenodo.org/records/13653420 ]. Unless otherwise stated, all data supporting the results of this study can be found in the article, supplementary, and source data files. Source data are provided with this paper. Result files and code to reproduce the results in this study are available on GitHub [ https://github.com/biocc/SP-MEGD_Fusion ]. An executable version of the code and the computational environment used in this study are also available as a Code Ocean capsule [ https://codeocean.com/capsule/4653442/tree ]. Articles from Nature Communications are provided here courtesy of Nature Publishing Group ACTIONS View on publisher site PDF (3.6 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top

Record · ID 738 · SHA-256 8ac63409c3990384
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.