Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice J Biomed Inform . Author manuscript; available in PMC: 2026 Apr 14. Published in final edited form as: J Biomed Inform. 2025 Feb 7;163:104789. doi: 10.1016/j.jbi.2025.104789 Search in PMC Search in PubMed View in NLM Catalog Add to search Improving entity recognition using ensembles of deep learning and fine-tuned large language models: A case study on adverse event extraction from VAERS and social media Yiming Li Yiming Li a McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston, Houston, TX 77030, USA Find articles by Yiming Li a , Deepthi Viswaroopan Deepthi Viswaroopan a McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston, Houston, TX 77030, USA Find articles by Deepthi Viswaroopan a , William He William He b Department of Electrical & Computer Engineering, Pratt School of Engineering, Duke University, 305 Tower Engineering Building, Durham, NC 27708, USA Find articles by William He b , Jianfu Li Jianfu Li c Department of Artificial Intelligence and Informatics, Mayo Clinic, Jacksonville, FL 32224, USA Find articles by Jianfu Li c , Xu Zuo Xu Zuo a McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston, Houston, TX 77030, USA Find articles by Xu Zuo a , Hua Xu Hua Xu d Department of Biomedical Informatics and Data Science, School of Medicine, Yale University, New Haven, CT 06510, USA Find articles by Hua Xu d , Cui Tao Cui Tao c Department of Artificial Intelligence and Informatics, Mayo Clinic, Jacksonville, FL 32224, USA Find articles by Cui Tao c, * Author information Article notes Copyright and License information a McWilliams School of Biomedical Informatics, The University of Texas Health Science Center at Houston, Houston, TX 77030, USA b Department of Electrical & Computer Engineering, Pratt School of Engineering, Duke University, 305 Tower Engineering Building, Durham, NC 27708, USA c Department of Artificial Intelligence and Informatics, Mayo Clinic, Jacksonville, FL 32224, USA d Department of Biomedical Informatics and Data Science, School of Medicine, Yale University, New Haven, CT 06510, USA * Corresponding author at: Department of Artificial Intelligence and Informatics, Mayo Clinic, 4500 San Pablo Rd, Jacksonville, FL 32224, USA. [email protected] (C. Tao). Issue date 2025 Mar. This article is made available under the Elsevier license ( http://www.elsevier.com/open-access/userlicense/1.0/ ). PMC Copyright notice PMCID: PMC13075790 NIHMSID: NIHMS2162270 PMID: 39923968 The publisher's version of this article is available at J Biomed Inform Abstract Objective: Adverse event (AE) extraction following COVID-19 vaccines from text data is crucial for monitoring and analyzing the safety profiles of immunizations, identifying potential risks and ensuring the safe use of these products. Traditional deep learning models are adept at learning intricate feature representations and dependencies in sequential data, but often require extensive labeled data. In contrast, large language models (LLMs) excel in understanding contextual information, but exhibit unstable performance on named entity recognition (NER) tasks, possibly due to their broad but unspecific training. This study aims to evaluate the effectiveness of LLMs and traditional deep learning models in AE extraction, and to assess the impact of ensembling these models on performance. Methods: In this study, we utilized reports and posts from the Vaccine Adverse Event Reporting System (VAERS) (n = 230), Twitter (n = 3,383), and Reddit (n = 49) as our corpora. Our goal was to extract three types of entities: vaccine, shot, and adverse event (ae). We explored and fine-tuned (except GPT-4) multiple LLMs, including GPT-2, GPT-3.5, GPT-4, Llama-2 7b, and Llama-2 13b, as well as traditional deep learning models like Recurrent neural network (RNN) and Bidirectional Encoder Representations from Transformers for Biomedical Text Mining (BioBERT). To enhance performance, we created ensembles of the three models with the best performance. For evaluation, we used strict and relaxed F1 scores to evaluate the performance for each entity type, and micro-average F1 was used to assess the overall performance. Results: The ensemble demonstrated the best performance in identifying the entities “vaccine,” “shot,” and “ae,” achieving strict F1-scores of 0.878, 0.930, and 0.925, respectively, and a micro-average score of 0.903. These results underscore the significance of fine-tuning models for specific tasks and demonstrate the effectiveness of ensemble methods in enhancing performance. Conclusion: In conclusion, this study demonstrates the effectiveness and robustness of ensembling fine-tuned traditional deep learning models and LLMs, for extracting AE-related information following COVID-19 vaccination. This study contributes to the advancement of natural language processing in the biomedical domain, providing valuable insights into improving AE extraction from text data for pharmacovigilance and public health surveillance. Keywords: Named-entity recognition, VAERS, Generative pre-trained transformer (GPT), Large language model (LLM), Social media, Deep learning 1. Introduction The COVID-19 pandemic has posed a significant global health threat, with over 111.8 million confirmed cases and more than 1.1 million deaths reported in the United States as of April 2024 [ 1 , 2 ]. The severity of COVID-19 is evident from its diverse symptoms, including fever, shortness of breath, fatigue, body aches, loss of taste or smell, sore throat, nausea, and diarrhea, which can progress to severe respiratory distress syndrome and death [ 3 ,p. 19, 4 ]. As with other infectious diseases, vaccination has emerged as the most effective measure to control the spread of COVID-19 and reduce its impact on public health [ 5 – 10 ]. In the United States, by May 2023, over 676 million doses of COVID-19 vaccines had been administered, with approximately 81.4 % of the population receiving at least one dose and 69.5 % fully vaccinated [ 2 ]. However, the introduction of COVID-19 vaccines has been accompanied by reports of adverse events (AEs) [ 11 ]. As of December 2022, the Vaccine Adverse Event Reporting System (VAERS) in the United States has received over 900 thousand reports of AEs following COVID-19 vaccination [ 8 , 10 ]. Of them, 4.7 % were classified as serious reports [ 8 ]. Common AEs following COVID-19 vaccination include mild symptoms such as fever, fatigue, and injection site pain [ 12 ]. However, there have been reports of severe AEs, including anaphylaxis and myocarditis, though these are rare [ 13 ]. Understanding the AEs following COVID-19 vaccination is crucial for ensuring the safety and efficacy of vaccination campaigns [ 14 ]. For AE following immunization reporting, VAERS is a crucial surveillance program managed by the Centers for Disease Control and Prevention (CDC) and the U.S. Food and Drug Administration (FDA) [ 15 ]. It serves as a cornerstone for monitoring the safety of vaccines licensed in the United States [ 16 ]. VAERS collects and analyzes reports of adverse events (possible side effects or health problems) that occur after vaccination [ 17 ]. Healthcare providers, vaccine manufacturers, and the public can submit reports to VAERS, which is critical for identifying potential safety concerns and ensuring the ongoing safety of vaccines [ 18 ], pp. 2000–2013]. Nowadays, social media has become a common platform for people to exchange ideas and express feelings [ 19 ]. With the COVID-19 pandemic, there has been a significant increase in the number of individuals using social media to share their experiences related to COVID-19, including symptoms following vaccination [ 20 , 21 ]. Posting on social media is often less trivial than reporting to formal systems like VAERS, making it a valuable source of real-time information on vaccine safety and AEs. During the outbreak, there have been over 468 + million posts related to COVID-19 on various social media platforms, highlighting the widespread use of social media for discussing pandemic-related topics [ 22 ]. Understanding the content and sentiment of these posts can provide additional insights into the public’s perception and concerns regarding COVID-19 vaccines, complementing traditional surveillance systems [ 23 ]. The recent advancements in natural language processing (NLP) have led to the development of powerful large language models (LLMs) such as the Generative Pre-trained Transformer (GPT) series. These models, trained on vast amounts of text data, have shown remarkable capabilities in understanding and generating human-like text [ 24 – 29 ]. Li et al. utilized LLMs to extract the relations for acupuncture point locations with a high performance [ 30 ]. Wang et al. proposed GPT-NER to improve Named Entity Recognition (NER) performance using LLM [ 31 ]. GPT-NER transforms the sequence labeling task of NER into a text-generation task, allowing LLMs to adapt more easily [ 31 ]. Additionally, GPT models can be fine-tuned to specific tasks, such as extracting AEs, by providing them with labeled examples of the information to extract [ 32 ]. For instance, researchers can annotate social media posts to indicate what information about adverse events following COVID-19 vaccination they contain. Li et al. investigated multiple pre-trained and fine-tuned LLMs to extract the AE-related information in the VAERS reports, resulting AE-GPT achieved 0.816 for relaxed match [ 32 ]. By training a GPT model on these annotated examples, it can learn to identify and extract relevant information from new, unseen reports. However, studies leveraging GPT models specifically for social media data remain limited. Using GPT for AE extraction from diverse sources offers several advantages. Firstly, GPT can process a large volume of text data quickly, allowing for the analysis of a vast number of social media posts in a short period [ 33 ]. Secondly, GPT’s ability to understand context and generate human-like text enables it to capture nuanced information, such as the severity of AEs or the context in which they occurred [ 34 ]. Lastly, GPT can be continuously updated and improved as new data becomes available, ensuring that it remains effective in extracting AEs from evolving social media discussions [ 35 ]. However, the performance of LLMs on NER tasks has shown considerable variability, presenting challenges in consistently achieving high accuracy. Monteiro and Zanchettin explored approaches to enhance NER in transformer-based models pre-trained for language modeling [ 36 ]. Despite the remarkable reasoning abilities demonstrated by GPT 3.5 and “ChatGPT” models, their quantitative performance was found to still lag behind that of traditionally fine-tuned models in their task [ 36 ]. This instability is particularly problematic in the biomedical domain, where precise identification of entities such as drug names, diseases, and patient attributes is critical for effective data analysis and decision-making. Given the challenges, in this study, we will firstly employ LLMs as well as traditional models to conduct the task of identifying entities related to AEs, such as vaccines , shots , and adverse events , from a large corpus of text. Additionally, we will also explore whether an ensemble approach, combining the strengths of LLMs with deep learning models, can enhance the stability and overall performance of the NER task. This research aims to contribute to the growing body of literature on utilizing advanced NLP techniques for pharmacovigilance and vaccine safety monitoring. 2. Methods In this study, we selected the VAERS and social media as data sources. We annotated the vaccine-related entities and predicted them using both pretrained and finetuned traditional deep learning models, as well as large language models. Fig. 1 provides an overview of the framework. Fig. 1. Open in a new tab Overview of the framework. 2.1. Data sources In this study, we utilized the reports from VAERS, and posts from Twitter and Reddit as our primary data sources, as summarized in Table 1 . Table 1. Summary of Data Sources and Selection Criteria for Adverse Event Extraction. Data Source Data Size Time Frame Selection Criteria VAERS 230 reports Dec 2020 – Dec 2022 COVID-19 vaccination-related adverse events Twitter 3,383 posts Dec 2020 – Aug 2021 Posts with predefined keywords related to COVID-19 vaccines and self-related keywords; excluded retweets, users with > 10,000 followers, and non-relevant tweets Reddit 49 posts Dec 2020 – Dec 2022 Posts related to personal COVID-19 vaccination experiences Open in a new tab 2.1.1. VAERS VAERS is a national vaccine safety surveillance program that collects and analyzes information about adverse events following immunization (AEFI). VAERS reports consist of three Comma-Separated-Value (CSV) files – VAERSDATA.CSV, VAERSVAX.CSV, and VAERSSYMPTOMS.CSV – grouped by year. The VAERSDATA.CSV contains demographic information, vaccination and AE timing, symptom descriptions, allergy history, and serious outcomes. VAERSVAX.CSV provides details on vaccine type and manufacturer for each adverse event, while VAERSSYMPTOMS.CSV lists the symptoms associated with each AEFI, as mapped from the Preferred Term (PT) in the MedDRA terminology. The three tables are linked by the primary key ‘VAERS_ID’. We randomly selected 230 reports related to COVID-19 vaccination from the VAERS database, covering the period between December 13, 2020, and December 28, 2022. 2.1.2. Twitter Twitter, a widely-used microblogging platform, has become an invaluable source for real-time information and public sentiment analysis [ 37 ]. The platform sees around 500 million tweets daily, covering a wide array of topics, including health-related discussions and public health issues [ 38 ]. As of March 2019, Twitter boasts approximately 330 million monthly active users globally [ 39 ]. Its vast user base and capacity to swiftly disseminate information make Twitter an essential resource for analyzing public health trends, including adverse events following COVID-19 vaccination. In this study, we reused the data collected by Lian et al. [ 40 ] They utilized the Twitter streaming API to gather posts related to vaccines. The inclusion criteria for our data collection were as follows: Tweets from December 2020 to August 2021; Include a predefined set of keywords: Pfizer, Moderna, J&J, Johnson & Johnson, BioNTech, vaccine, AstraZeneca, covidvac, etc. Include self-related keywords: “I,” “my,” “mine,” “me,” “myself”, etc. Due to the high noise level in the Twitter data, which included tweets unrelated to AEs following COVID-19 vaccine, exclusion criteria were applied to identify content specifically related to personal experiences of post-COVID-19 vaccine AEs. Retweets and quotes were removed. Tweets from users with over 10,000 followers, classified by Twitter as “super follows,” were excluded. 2.1.3. Reddit In addition to Twitter, we also utilized Reddit streaming API to collect posts from Reddit for AE extraction. Reddit, a popular social news aggregation, web content rating, and discussion website, features a vast user base and sees a significant flow of posts daily [ 41 ,p. 19]. With approximately 52 million daily active users and numerous topic-specific communities known as “subreddits,” Reddit offers a diverse range of discussions, including those related to health and COVID-19 vaccines [ 42 , 43 ]. We collected data from Reddit posts within the time frame starting from December 1, 2020, to December 31, 2022. To include qualified Reddit posts in our study, we focused on all subreddits, which encompasses a wide range of discussions and is not limited to a particular topic. To ensure the relevance of the symptoms described in Reddit posts, we recruited domain experts to participate in the annotation process. Relevance was defined as being related to personal experiences of COVID-19 vaccination. A tweet or post was considered relevant only if it detailed a personal story related to COVID-19 vaccine experiences. We also leveraged methods from the study conducted by Lian et al., which utilized rule-based approaches to remove noise and identify content related to personal experiences of adverse events [ 40 ]. Furthermore, Our annotators checked the relevance of the reported AEs to COVID-19 vaccines based on the context and their experience. This multi-step approach allowed us to link social media discussions to real vaccine events effectively. In our study, we considered both short-term and long-term AEs. While immediate AEs were the primary focus, we also included a small number of posts reporting symptoms that emerged later (e.g., six months post-vaccination) to enable a more comprehensive analysis. These long-term reports were treated with caution, distinguishing them from immediate reactions. Additionally, annotators carefully examined each AE within its context to ensure relevance. Finally, we retained and annotated randomly the posts and reports that contained at least one AE keyword related to the COVID-19 vaccine. This dataset includes 230 reports from VAERS, 3,383 tweets, and 49 posts from Reddit. 2.2. Annotation In this study, we engaged two annotators (D.V. and W.H.) and utilized CLAMP (Clinical Language Annotation, Modeling, and Processing) to label COVID-19 vaccine-related AEs for corpora [ 44 ]. These named entities included vaccine , shot , and ae (adverse event). The definition and examples of entities are shown in Table 2 . The vaccine entity referred to the specific COVID-19 vaccine mentioned in the posts/reports, with the full name of the vaccine selected if possible. Examples of vaccine entities included “Pfizer vaccine,” “Moderna vaccine,” among others. The shot entity indicated which shot of the COVID-19 vaccine was being referred to in the posts/reports, such as “first dose,” “second dose,” and so on. ae entities denoted the symptoms or diseases experienced following vaccination, with annotations including “fever,” “sore arm,” “headache,” and similar symptoms. Table 2. Definition and examples of entities. Entity Definition Examples vaccine the specific COVID-19 vaccine mentioned in the posts/reports Moderna vaccine, Pfizer vaccine, Coronavirus Vaccine shot Specific dose(s) of a vaccine administered through an injection booster shot, 1st dose, 2nd dose ae Any negative or unexpected medical occurrence or side effect that follows vaccination. sore arm, headache, fever Open in a new tab During the annotation process, we adhered to specific guidelines. These guidelines instructed annotators not to include space and special characters in the text, avoid annotating adjectives for a concept (e.g., mild, strong), and annotate only those AEs that were the symptoms or diseases experienced by the vaccine recipient, excluding those experienced by family members or others. An annotation example related to adverse events following COVID-19 vaccination is provided in Fig. 2 . Fig. 2. Open in a new tab An annotation example in this study. 2.3. Models 2.3.1. GPT The GPT model, developed by OpenAI, has demonstrated remarkable capabilities in various NLP tasks, including NER [ 28 , 32 ]. GPT, based on the Transformer architecture, leverages its extensive pre-training strategy to learn contextual representations of words, enabling it to understand the context in which named entities appear [ 45 , 46 ]. This contextual understanding allows GPT to effectively identify and classify named entities in text, making it a powerful tool for NER tasks in NLP. 2.3.2. Llama-2 Llama-2, developed by Microsoft, offers models with 7 billion, 13 billion, and 70 billion parameters, providing powerful capabilities across various NLP tasks [ 32 ]. This range of models allows Llama to excel in tasks such as text classification, language understanding, and text generation [ 47 ]. 2.3.3. RNN RNNs are designed to process sequential data by maintaining an internal memory [ 48 ]. This memory allows them to learn patterns and dependencies in sequences, making them suitable for tasks like speech recognition, language translation, and time series prediction [ 49 ]. RNNs process data step by step, updating their internal state with each new input [ 50 ]. This ability to remember past information makes them effective in tasks where context is important. However, they can struggle with long sequences due to the vanishing gradient problem, which limits their ability to retain information over long periods [ 50 ]. 2.3.4. BioBERT BioBERT is a variant of the BERT (Bidirectional Encoder Representations from Transformers) model that has been specifically pre-trained on biomedical text [ 51 ]. This pre-training helps BioBERT better understand the complex language used in biomedical literature, making it particularly effective for tasks in the biomedical domain [ 52 ]. By fine-tuning BioBERT on specific tasks or datasets, researchers can leverage its biomedical knowledge to achieve state-of-the-art results in various NLP tasks, such as NER, relation extraction, and question answering in the biomedical field [ 51 ]. BioBERT’s ability to capture domain-specific nuances and terminology makes it a valuable tool for advancing research and applications in biomedicine and healthcare. 2.4. Experiment setup 2.4.1. Data split In this study, we split the dataset into training, validation, and test sets using an 8:1:1 ratio. Fig. 3 provides a detailed breakdown of the number of entities for each set. Fig. 3. Open in a new tab A breakdown of the number of entities for each set. 2.4.2. GPT For the GPT models, we employed pre-trained versions of GPT-2, GPT-3.5, and GPT-4. Additionally, we fine-tuned the pre-trained GPT-2 and GPT-3.5 models for our specific task. For the prompts, we divided them into two styles: split and merged. In the split style, prompts were designed to extract entities individually, focusing on one entity at a time. In contrast, the merged style involved prompts that aimed to extract all entities simultaneously. We determined the best-performing prompt style for each GPT model based on experimental results, as detailed in Table A.1 . This approach allowed us to tailor the prompt style to each model’s strengths, optimizing performance for this entity extraction tasks. Split prompt style was structured as follows, with the vaccine entity for fine-tuned GPT 2 used as an example: “question”: “Please extract all the names of vaccine from the following note” “context”: {note} . Merged prompt style (pretrained GPT 2) was formatted as follows: “Please extract all names of dose, vaccine, and adverse event from this note, and put them in a list: {note}” In our experiment setup, we used several important parameters for generating text with the GPT models. The temperature parameter, controls the randomness of the generated output. A lower temperature results in more deterministic outputs, while a higher temperature leads to more diverse but potentially less coherent text. The temperature parameter allows us to balance between generating text that closely resembles the training data and generating more novel, creative outputs. We also utilized the maximum output token parameter, which defines the maximum length of the generated text in terms of the number of tokens. A token in this context refers to the smallest unit of text that the model can process, which can be a word, subword, or character depending on the tokenization scheme used. Setting a maximum output token limit helps control the length of the generated text, ensuring that it remains within a manageable length for readability and analysis. These settings were selected based on preliminary experiments to achieve a balance between text quality and computational efficiency. Table A.2 shows the important parameters used by each GPT model. 2.4.3. Llama 2 For Llama 2, we utilized two variants: Llama 2 7b and Llama 2 13b, both of which were fine-tuned for our task. The temperature setting for all Llama 2 models was set to 1. The specific prompts used are shown in Table A.1 . During the inference phase, we employed prompts in split mode to extract entities. For the fine-tuning process of Llama 2 to extract vaccine , we used the following format: “instruction”: “Please extract all names of vaccines”, “input”: {note}. For the remaining parameters used in Llama 2, please refer to Table A.3 . 2.4.4. RNN We fine-tuned a Recurrent Neural Network (RNN) for our experiments with the following configuration: lowercase words were set to 1, replacing digits with 0 was enabled (zeros = 1), and the character embedding dimension was set to 25 (char_dim = 25). The character LSTM hidden layer size was also set to 25 (char_lstm_dim = 25), and a bidirectional LSTM for characters was used (char_bidirect = 1). For token embeddings, the word embedding dimension was 200 (word_dim = 200), and the token LSTM hidden layer size was 100 (word_lstm_dim = 100). A bidirectional LSTM for words was employed (word_bidirect = 1), and a Conditional Random Field (CRF) was used for tagging (crf = 1). Dropout with a rate of 0.5 was applied (dropout=‘0.5′). The tagging scheme used was IOB (tag_scheme=‘iob’), and the model was fine-tuned over 30 epochs (epoch = 30). 2.4.5. BioBERT We also fine-tuned BioBERT v1.1 for our experiments. Fine-tuning involved training BioBERT on the same task-specific dataset that integrated unstructured data from three distinct sources: detailed reports of vaccine adverse events from VAERS, and personal posts from social media platforms like Twitter and Reddit, where individuals shared their experiences with COVID-19 vaccinations. This dataset offered a diverse representation of both formal, clinical data and informal, conversational narratives, which allowed the model to adapt to varying linguistic styles, medical terminologies, and informal expressions of vaccine-related side effects. The configuration used for fine-tuning BioBERT v1.1 included an attention dropout probability of 0.1, a hidden activation function of “gelu,” a hidden dropout probability of 0.1, a hidden size of 768, an initializer range of 0.02, an intermediate size of 3,072, a maximum position embeddings of 512, 12 attention heads, 12 hidden layers, a type vocabulary size of 2, and a vocabulary size of 28,996. 2.4.6. Ensemble Ensembling enables us to capitalize on the strengths of each individual model, enhancing overall performance through the amalgamation of their predictions. In this study, we utilized ensembles comprising fine-tuned GPT-3.5, RNN, and BioBERT models, employing a majority voting scheme. This approach entails consolidating the predictions from each model and selecting the most frequently predicted outcome as the final result. The experiments for the GPT models and BioBERT were conducted on a server with 8 Nvidia A100 GPUs, each offering 80 GB of memory. Meanwhile, the Llama models and RNN were executed on a server equipped with 5 Nvidia V100 GPUs, each providing a memory capacity of 32 GB. 2.5. Evaluation In this study, we employed inter-rater agreement to measure the consistency between annotators in the annotation process. This metric helps assess the reliability of the annotations and ensures that the data is accurately labeled for further analysis. To evaluate the performance of the models in entity recognition, we used F1 scores, calculated both in relaxed and strict settings, for each entity type. The relaxed F1 score allows for some leniency in matching predicted and ground truth entities, while the strict F1 score requires an exact match. Additionally, we calculated the micro-average F1 score to assess the overall performance of each model across all entities. These evaluation metrics provide a comprehensive understanding of how well each model performs in recognizing adverse event-related entities following COVID-19 vaccination. 3. Results Table 3 shows the inter-rater agreement for the annotation of entities. The agreement for the entity vaccine was 0.94, indicating strong agreement between annotators. For the entity shot , the agreement was perfect, with a score of 1, indicating complete agreement. The agreement for the entity ae was slightly lower at 0.73, indicating moderate agreement. Overall, the inter-rater agreement across all entities was 0.77, reflecting a substantial level of agreement between annotators. Table 3. Inter-rater agreement. Entities Inter-rater agreement vaccine 0.94 shot 1 ae 0.73 Overall 0.77 Open in a new tab Table 4a presents the relaxed F1 of different models in extracting AE-related information following COVID-19 vaccines. The performance comparison of various pretrained and fine-tuned models across different entity types reveals significant insights. When comparing pretrained and fine-tuned models, GPT-2 demonstrates a remarkable improvement, with the fine-tuned version achieving an F1 score of 0.687 compared to 0.0 for the pretrained model, highlighting the effectiveness of fine-tuning in enhancing entity recognition tasks. Similarly, GPT-3.5 shows a marked benefit from fine-tuning, outperforming its pretrained counterpart with scores of 0.758 versus 0.412. Llama 2 models also reflect this pattern, with the fine-tuned 7b version scoring 0.238 compared to 0.134 for the pretrained, while the fine-tuned 13b model scores 0.398 against 0.252 for its pretrained variant. On the whole, GPT models outperformed Llama 2 models in micro-average score. Although traditional models like RNN and BioBERT show competitive performance, with BioBERT reaching 0.916, the ensemble model emerges as the highest performer across all entity types, achieving an overall micro-average score of 0.926. Table 4a. Relaxed F1 of different models in extracting AE-related information following COVID-19 vaccines. Pretrained GPT-2 Fine-tuned GPT- 2 Pretrained GPT-3.5 Fine-tuned GPT-3.5 Pretrained GPT-4 Pretrained Llama 2 7b Fine-tuned Llama 2 7b Pretrained Llama 2 13b Fine-tuned Llama 2 13b RNN BioBERT Ensemble vaccine 0 0.827 0.446 0.849 0.492 0.319 0.335 0.524 0.739 0.896 0.925 0.918 shot 0 0.868 0.416 0.917 0.441 0.008 0.156 0.117 0.176 0.917 0.936 0.965 ae 0 0.512 0.400 0.648 0.417 0.048 0.368 0.150 0.348 0.871 0.905 0.934 Micro-average 0 0.687 0.412 0.758 0.437 0.134 0.238 0.252 0.398 0.886 0.916 0.926 Open in a new tab In terms of entity types, fine-tuned models consistently outperform their pretrained counterparts. For the entity type “vaccine,” BioBERT leads with an impressive score of 0.925. For “shot,” the ensemble achieved a remarkable score of 0.965. The “ae” category follows this trend, with BioBERT (0.905) and the ensemble model (0.934) standing out as top performers. Overall, these results underscore the importance of employing fine-tuned models and ensemble strategies to enhance entity extraction performance in biomedical contexts. Table 4b shows the strict F1 scores for various models across different entities, including the overall micro-average F1 score. Among the pre-trained models, GPT-3.5 stands out with higher scores across entity types compared to other models, particularly with a score of 0.182 for “vaccine,” 0.359 for “shot,” and 0.269 for “ae,” resulting in a micro-average of 0.258. GPT-4 slightly outperforms relative to GPT-3.5 with a micro-average of 0.263, though it generally remains competitive. Llama 2 models exhibit lower performance across the board, with Llama 2 13b achieving a micro-average of 0.176 and Llama 2 7b trailing behind with 0.082. Overall, pre-trained RNN and BioBERT perform better than most pre-trained GPT and Llama models, achieving micro-averages of 0.805 and 0.824, respectively. Table 4b. Strict F1 of different models in extracting AE-related information following COVID-19 vaccines. Pretrained GPT-2 Fine-tuned GPT-2 Pretrained GPT-3.5 Fine-tuned GPT-3.5 Pretrained GPT-4 Pretrained Llama 2 7b Fine-tuned Llama 2 7b Pretrained Llama2 13b Fine-tuned Llama2 13b RNN BioBERT Ensemble vaccine 0 0.723 0.182 0.781 0.212 0.172 0.187 0.359 0.536 0.847 0.864 0.878 shot 0 0.766 0.359 0.853 0.360 0 0.094 0.075 0.137 0.860 0.875 0.930 ae 0 0.355 0.269 0.632 0.270 0.034 0.288 0.103 0.274 0.763 0.786 0.925 Micro-average 0 0.559 0.258 0.716 0.263 0.082 0.163 0.176 0.296 0.805 0.824 0.903 Open in a new tab When examining fine-tuned models, GPT-3.5 shows the most substantial improvement, achieving a micro-average of 0.716, with particularly strong scores for “vaccine” (0.781) and “shot” (0.853). The fine-tuned GPT-2 performs notably lower, with a micro-average of 0.559. Both Llama 2 7b and Llama 2 13b show moderate improvements upon fine-tuning, reaching 0.163 and 0.296 in micro-average, respectively. However, these are still below the fine-tuned GPT-3.5 scores. Fine-tuning results in the ensemble model achieving the highest performance across all models, with a micro-average of 0.903 and highest F-1 score across all entity types. 4. Discussion 4.1. Findings Our study aligns with extensive research highlighting the critical need for robust adverse event detection and reporting systems. Effective post-vaccination surveillance is essential for identifying adverse events, which in turn strengthens public trust in vaccination programs [ 53 ]. Additionally, the absence of timely signals regarding adverse events can hinder prompt interventions and obscure significant safety concerns [ 54 ]. Moreover, without vigilant tracking, rare but serious reactions may go unnoticed, increasing the potential for public health risks [ 55 ]. These insights reinforce the importance of our work in enhancing adverse event monitoring and detection. This study is among the first to investigate ensembles of LLMs and traditional deep learning models, and apply GPT models to social media data in this context. Our findings highlight several key points regarding the performance and effectiveness of these models in this NER task. The three data sources used in this study—VAERS, Twitter, and Reddit—differ in structure, content, and context, which impacted data processing and analysis. VAERS data provided descriptions of adverse events that were medically focused and grounded in clinical terminology, making them well-suited for systematic analysis. However, they lack the informal narratives and real-world perspectives present in social media data. In contrast, Twitter data, while extensive and timely, are limited by character count, resulting in brief and often ambiguous descriptions of adverse events. Preprocessing and annotation were necessary to filter relevant content, especially given the high noise level. Reddit posts, though fewer in number, provided richer and more detailed narratives, often resembling personal stories with greater context. These differences highlight the complementary nature of the datasets, as VAERS offers clinical rigor, Twitter captures public sentiment and anecdotal experiences, and Reddit bridges the gap with detailed, conversational narratives. Integrating these diverse sources required harmonization efforts to address variations in linguistic style, data granularity, and context, ultimately enabling a more comprehensive analysis of vaccine adverse events. While some studies explore few-shot learning for LLMs in NER tasks, fine-tuning of pre-trained LLMs, such as GPT-2 and GPT-3.5, also played a pivotal role in enhancing their ability to recognize entities related to AEs [ 56 ]. This process allows the models to adapt to the specific characteristics of the dataset, leading to improved performance. This finding is consistent with prior work by Li et al., who previously explored pretrained and fine-tuned LLMs for extracting adverse events following influenza vaccines [ 32 ]. Interestingly, we observed minimal differences in performance among different grades of GPT models, suggesting a certain level of robustness in handling NER tasks regardless of the model’s size. In contrast, Llama models exhibited more noticeable differences in performance, which can be attributed to their specialized architecture and training objectives for medical NLP tasks. This highlights the importance of selecting Llama 2 models tailored to the specific task at hand, as the performance may vary based on the model’s design and training data. In this study, we also investigated the effectiveness of ensembling fine-tuned LLMs with traditional deep learning models for the NER task related to AEs following COVID-19 vaccination from social media posts. Our findings reveal the significant implications of ensembling in improving the overall performance of the models. While LLMs have shown to be inferior to traditional deep learning models in this NER task, ensembling the two types of models led to a substantial improvement in the strict F1 score. Specifically, the ensembling approach resulted in an 8 % increase in the strict F1 score, exceeding 90 %. This improvement underscores the complementary nature of LLMs and traditional deep learning models, suggesting that combining their strengths can lead to better performance than either model alone. The effectiveness of ensembling can be attributed to the unique strengths of each type of model. LLMs exhibit exceptional language comprehension capabilities, and excel in capturing complex linguistic patterns and contextual information, making them highly effective in understanding the nuances of social media posts. On the other hand, traditional deep learning models, with their robust architectures and ability to learn complex feature representations, complement LLMs by providing additional context and generalization capabilities. Furthermore, ensembling helps mitigate the weaknesses of individual models. While LLMs may struggle with certain aspects of the NER task, such as entity ambiguity or rare occurrences, traditional deep learning models can compensate for these limitations by providing more robust and reliable predictions in such cases [ 57 , 58 ]. This complementary nature of the two types of models makes ensembling a powerful strategy for improving overall performance. Despite the overall success in NER, we identified two entities, shot and ae , that did not perform as well as expected. The shot entity, although showing high inter-rater agreement, was rare in occurrence, potentially affecting the models’ ability to learn from sufficient examples. Similarly, the ae entity’s broad and ambiguous nature posed challenges for accurate recognition, leading to lower performance in this category. 4.2. Error analysis The error analysis ( Table 5 ) shows that the model struggled most with false negatives for ae , missing 9.12 % of human-annotated entities. For vaccine and shot , false positives were more prevalent, with 6.56 % and 11.43 %, respectively, of machine-annotated entities being incorrect. Boundary mismatches were relatively low across all entities, indicating that the model generally identified entity boundaries accurately. No instances of incorrect entity types were identified for any entity, indicating that when the model made a prediction, it tended to assign the correct entity type. Table 5. Error analysis by each entity type. Boundary Mismatch (out of human annotated entities) False Positive (out of machine annotated entities) False Negative (out of human annotated entities) Incorrect Entity Type (out of machine annotated entities) vaccine 23/429, 5.36 % 29/442, 6.56 % 26/429, 6.06 % 0/442, 0 % shot 16/272, 5.88 % 36/315, 11.43 % 6/272, 2.21 % 0/315, 0 % ae 37/811, 4.56 % 41/792, 5.18 % 74/811, 9.12 % 0/792, 0 % Open in a new tab In our ensemble method, we achieved near-perfect performance for AE extraction. However, there were instances where the model missed certain AE-related entities, such as “flu-like symptoms”, “cold”, “shaky”, “achiness”, and “a knot in my hairline”. These entities may not be explicitly defined in an AE-related ontology, leading to their neglect. Additionally, the model occasionally misclassified the general term “vaccine” as a specific entity type vaccine . These errors can be attributed to several factors. The neglect of certain entities may be due to the limited scope of the model’s training data, which may not have included these specific entities. The misclassification of “vaccine” may be caused by the ambiguity of the term, which can refer to both the general concept of vaccination and specific instances of vaccines. To address these errors, expanding the training data to include a more diverse range of AE-related entities could improve the model’s performance. Additionally, refining the model’s entity recognition capabilities to better distinguish between general terms and specific entities, such as using context-aware techniques, could help reduce misclassifications. 4.3. Strengths and limitations This study has several strengths that contribute to its significance in the field of AE detection from VAERS and social media data. Firstly, the utilization of three years of patient self-reported data, which includes reports from VAERS and two social media platforms, adds a robust and comprehensive dimension to the analysis. This extensive dataset allows for a thorough examination of AE reports over time, providing valuable insights into trends and patterns that may not be apparent in shorter-term studies. Additionally, the inclusion of social media data alongside VAERS reports offers a more holistic view of AEs, capturing a wider range of patient experiences and opinions. Secondly, this study stands out for its comprehensive exploration of a wide range of traditional deep learning and LLMs. By examining and comparing the performance of these models, the study sheds light on their respective strengths and weaknesses in the context of NER tasks related to AEs. This comprehensive analysis not only provides valuable insights for researchers and practitioners in the field but also serves as a reference point for future studies looking to employ similar models. Overall, these strengths underscore the thoroughness and rigor of the study, enhancing its credibility and relevance in the field. The study has several limitations that should be considered. Firstly, the dataset seems to be unbalanced and biased towards Twitter data due to its widespread use and ease of access for collecting real-time information. Secondly, the corpora used for AE extraction may be incomplete or contain inaccuracies, leading to misclassification or omission of relevant events. The reliance on social media data introduces bias, as users may not be representative and reporting may be influenced by various factors. We also acknowledge that history bias is a limitation in our data, where symptoms observed might not necessarily be attributed to the vaccine itself but to historical events or interventions that occurred during the same time frame. To mitigate this, we employed contextual keywords and time indicators (e.g., “recently vaccinated”) when available to assess the temporal link. However, many posts lacked this information, limiting our ability to accurately determine symptom onset relative to vaccination. Moreover, ensemble methods consume substantial computational resources. Additionally, while the study demonstrates the feasibility of using LLMs for AE extraction in the context of COVID-19 vaccines, its generalizability to other medical domains may be limited and require further validation. 5. Conclusion In conclusion, this study contributes by exploring the effectiveness of both LLMs and traditional deep learning models in the AE extraction task. It also highlights the significant improvement achieved by ensembling fine-tuned LLMs and traditional deep learning models. The findings provide valuable insights for both biomedical informatics and clinics, suggesting that ensembles of these models can significantly enhance the accuracy and robustness of AE extraction from text data, thereby supporting clinical decision-making and pharmacovigilance efforts. Future work should focus on further refining the LLM models and expanding the dataset to enhance the performance and generalizability. Supplementary Material Supplementary NIHMS2162270-supplement-Supplementary.docx (9.4KB, docx) Statement of Significance. Problem or Issue What is Already Known What this Paper Adds Improving the named entity recognition (NER) performance with ensembles of fine-tuned deep learning models and large language models (LLMs). Traditional deep learning models excel in learning feature representations but require extensive labeled data. On the other hand, LLMs hold promise for improving NER tasks with excellent language understanding capability but exhibit unstable performance possibly due to their training, which is not specifically tailored to the NER task. Prior studies have assessed the use of traditional deep learning models for NER, but few have explored the impact of ensembling them and LLMs. This study demonstrates the effectiveness and robustness of ensembling fine-tuned LLMs and traditional deep learning models in adverse event (AE)-related information extraction. This study also contributes to improving AE extraction from text data for pharmacovigilance and public health surveillance. Open in a new tab Acknowledgements This article was partially supported by the National Institute of Allergy And Infectious Diseases of the National Institutes of Health under Award Numbers R01AI130460 and U24AI171008. We also extend our sincere appreciation to Prof. Xiaoqian Jiang for his support with the GPT experiments in this study. Abbreviations: AE Adverse event AEFI Adverse events following immunization BERT Bidirectional Encoder Representations from Transformers BioBERT Bidirectional Encoder Representations from Transformers for Biomedical Text Mining CDC Centers for Disease Control and Prevention CLAMP Clinical Language Annotation, Modeling, and Processing CRF Conditional Random Field CSV Comma-Separated-Value FDA Food and Drug Administration GPT Generative pre-trained transformer LLM Large language model NER Named entity recognition NLP Natural language processing PT Preferred Term RNN Recurrent neural network VAERS Vaccine Adverse Event Reporting System Appendix A. Supplementary material Supplementary data to this article can be found online at https://doi.org/10.1016/j.jbi.2025.104789 . Footnotes CRediT authorship contribution statement Yiming Li: Writing – original draft, Visualization, Software, Methodology, Investigation, Formal analysis, Conceptualization. Deepthi Viswaroopan: Data curation. William He: Data curation. Jianfu Li: Software. Xu Zuo: Resources. Hua Xu: Resources. Cui Tao: Writing – review & editing, Supervision, Resources, Project administration, Methodology, Funding acquisition, Conceptualization. Declaration of competing interest The authors declare that they have no known competing financial interests or personal relationships that could have appeared to influence the work reported in this paper. References [1]. “United States COVID - Coronavirus Statistics - Worldometer.” Accessed: Mar. 23, 2024. [Online]. Available: https://www.worldometers.info/coronavirus/country/us/ . [2]. Centers for Disease Control and Prevention, “COVID Data Tracker,” US Department of Health and Human Services, CDC, Atlanta, GA, Apr. 2023. [Online]. Available: https://covid.cdc.gov/covid-data-tracker . [ Google Scholar ] [3]. Dane S, Akyuz M, Symptom spectrum and the evaluation of severity and duration of symptoms in patients with COVID-19, Journal of Research in Medical and Dental Science 9 (2021) 262–266. [ Google Scholar ] [4]. Liu Y et al. , “Clinical features and progression of acute respiratory distress syndrome in coronavirus disease 2019,” 2020, doi: 10.1101/2020.02.17.20024166. [ DOI ] [ Google Scholar ] [5]. Li Y, Li J, Dang Y, Chen Y, Tao C, Adverse Events of COVID-19 Vaccines in the United States: Temporal and Spatial Analysis, JMIR Public Health Surveill 10 (Jul. 2024) e51007. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [6]. Paul JN, Mbalawata IS, Mirau SS, Masandawa L, Mathematical modeling of vaccination as a control measure of stress to fight COVID-19 infections, Chaos, Solitons & Fractals 166 (Jan. 2023) 112920, 10.1016/j.chaos.2022.112920. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [7]. Li Y, et al. , Unpacking adverse events and associations post COVID-19 vaccination: a deep dive into vaccine adverse event reporting system data, Expert Review of Vaccines 23 (1) (Dec. 2024) 53–59, 10.1080/14760584.2023.2292203. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [8]. Li Y, Li J, Dang Y, Chen Y, and Tao C, “Temporal and Spatial Analysis of COVID-19 Vaccines Using Reports from Vaccine Adverse Event Reporting System,” JMIR Preprints, doi: 10.2196/preprints.51007. [ DOI ] [ Google Scholar ] [9]. Zhang K, Dang Y, Li Y, Tao C, Hur J, and He Y, “Impact of climate change on vaccine responses and inequity,” Nat. Clim. Chang, vol. 14, no. 12, Art. no. 12, Dec. 2024, doi: 10.1038/s41558-024-02192-y. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [10]. Li Y, Li J, Dang Y, Chen Y, Tao C, COVID-19 Vaccine Adverse Events in the United States: A Temporal and Spatial Analysis, JMIR Preprints (Jun. 2024), 10.2196/51007. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [11]. Pan Y et al. , “Assessing acute kidney injury risk after COVID vaccination and infection in a large cohort study,” npj Vaccines, vol. 9, no. 1, pp. 1–10, Nov. 2024, doi: 10.1038/s41541-024-00964-3. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [12]. Kouhpayeh H, Ansari H, Adverse events following COVID-19 vaccination: A systematic review and meta-analysis, International Immunopharmacology 109 (2022) 108906, 10.1016/j.intimp.2022.108906. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [13]. Kim DH, Kim JH, Oh I-S, Choe YJ, Choe S-A, Shin J-Y, Adverse Events Following COVID-19 Vaccination in Adolescents: Insights From Pharmacovigilance Study of VigiBase, J Korean Med Sci 39 (8) (2024) e76. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [14]. Younus MM and Al-Jumaili AA, “An Overview of COVID-19 Vaccine Safety and Post-marketing Surveillance Systems,” Innov Pharm, vol. 12, no. 4, p. 10.24926/iip.v12i4.4294 , Sep. 2021, doi: 10.24926/iip.v12i4.4294. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [15]. Moro PL, Li R, Haber P, Weintraub E, Cano M, Surveillance systems and methods for monitoring the post-marketing safety of influenza vaccines at the Centers for Disease Control and Prevention, Expert Opinion on Drug Safety 15 (9) (2016) 1175–1183, 10.1080/14740338.2016.1194823. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [16]. Singleton JA, Lloyd JC, Mootrey GT, Salive ME, Chen RT, An overview of the vaccine adverse event reporting system (VAERS) as a surveillance system, Vaccine 17 (22) (1999) 2908–2917, 10.1016/S0264-410X(99)00132-2. [ DOI ] [ PubMed ] [ Google Scholar ] [17]. Chen RT, et al. , The vaccine adverse event reporting system (VAERS), Vaccine 12 (6) (May 1994) 542–550, 10.1016/0264-410X(94)90315-8. [ DOI ] [ PubMed ] [ Google Scholar ] [18]. Hibbs BF, Moro PL, Lewis P, Miller ER, Shimabukuro TT, Vaccination errors reported to the Vaccine Adverse Event Reporting System, (VAERS) United States, 2000–2013, Vaccine 33 (28) (2015) 3171–3178, 10.1016/j.vaccine.2015.05.006. [ DOI ] [ PubMed ] [ Google Scholar ] [19]. Weinberg BD, de Ruyter K, Dellarocas C, Buck M, Keeling DI, Destination social business: exploring an organization’s journey with social media, collaborative community and expressive individuality, Journal of Interactive Marketing 27 (4) (2013) 299–310, 10.1016/j.intmar.2013.09.006. [ DOI ] [ Google Scholar ] [20]. Katz M, Nandi N, Social Media and Medical Education in the Context of the COVID-19 Pandemic: Scoping Review, JMIR Medical Education 7 (2) (Apr. 2021) e25892. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [21]. Lentzen M-P, Huebenthal V, Kaiser R, Kreppel M, Zoeller JE, Zirk M, A retrospective analysis of social media posts pertaining to COVID-19 vaccination side effects, Vaccine 40 (1) (Jan. 2022) 43–51, 10.1016/j.vaccine.2021.11.052. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [22]. Arpaci I, Analysis of Twitter Data Using Evolutionary Clustering during the COVID-19 Pandemic, Cmc-Computers Materials & Continua (2020) 193–203. [ Google Scholar ] [23]. Hu T, et al. , Revealing Public Opinion Towards COVID-19 Vaccines With Twitter Data in the United States: Spatiotemporal Perspective, Journal of Medical Internet Research 23 (9) (Sep. 2021) e30854. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [24]. Li Y, Wei Q, Chen X, Li J, Tao C, Xu H, Improving tabular data extraction in scanned laboratory reports using deep learning models, Journal of Biomedical Informatics 159 (Nov. 2024) 104735, 10.1016/j.jbi.2024.104735. [ DOI ] [ PubMed ] [ Google Scholar ] [25]. Wang Y, et al. , “Aligning Large Language Models with Human, A Survey, (2023). [26]. Li Y, et al. , Artificial intelligence-powered pharmacovigilance: A review of machine and deep learning in clinical text-based adverse drug event detection for benchmark datasets, Journal of Biomedical Informatics 152 (Apr. 2024) 104621, 10.1016/j.jbi.2024.104621. [ DOI ] [ PubMed ] [ Google Scholar ] [27]. Li Y, et al. , RefAI: a GPT-powered retrieval-augmented generative tool for biomedical literature recommendation and summarization, Journal of the American Medical Informatics Association 31 (9) (Sep. 2024) 2030–2039, 10.1093/jamia/ocae129. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [28]. Hu Y, et al. , Zero-shot Clinical Entity Recognition using ChatGPT, arXiv.org (2023), 10.48550/arXiv.2303.16416. [ DOI ] [ Google Scholar ] [29]. Tao C et al. , “VaxBot-HPV: A GPT-based Chatbot for Answering HPV Vaccine-related Questions,” Research Square, p. rs.3.rs, Sep. 2024, doi: 10.21203/rs.3.rs-4876692/v1. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [30]. Li Y, et al. , Relation extraction using large language models: a case study on acupuncture point locations, Journal of the American Medical Informatics Association 31 (11) (Nov. 2024) 2622–2631, 10.1093/jamia/ocae233. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [31]. Wang S et al. , “GPT-NER: Named Entity Recognition via Large Language Models,” 2023. [32]. Li Y, Li J, He J, Tao C, AE-GPT: Using Large Language Models to extract adverse events from surveillance reports-A use case with influenza vaccine adverse events, PLOS ONE 19 (3) (Mar. 2024) e0300919. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [33]. Kalla D and Smith N, “Study and Analysis of Chat GPT and its Impact on Different Fields of Study,” vol. 8, no. 3, 2023. [ Google Scholar ] [34]. Bansal G, Chamola V, Hussain A, Guizani M, Niyato D, Transforming Conversations with AI—A Comprehensive Study of ChatGPT, Cogn Comput (2024), 10.1007/s12559-023-10236-2. [ DOI ] [ Google Scholar ] [35]. Johnson D et al. , “Assessing the Accuracy and Reliability of AI-Generated Medical Responses: An Evaluation of the Chat-GPT Model,” Res Sq, p. rs.3.rs-2566942, Feb. 2023, doi: 10.21203/rs.3.rs-2566942/v1. [ DOI ] [ Google Scholar ] [36]. Monteiro M, Zanchettin C, Optimization Strategies for BERT-Based Named Entity Recognition, in: Naldi MC, Bianchi RAC (Eds.), Intelligent Systems, Springer Nature Switzerland, Cham, 2023, pp. 80–94, 10.1007/978-3-031-45392-2_6. [ DOI ] [ Google Scholar ] [37]. Kumar A and Sebastian TM, “Sentiment Analysis on Twitter,” vol. 9, no. 4, 2012. [ Google Scholar ] [38]. Alwagait E, Shahzad B, When are tweets better valued? An empirical study, J. Universal Comput. Sci 20 (Jan. 2014) 1511–1521. [ Google Scholar ] [39]. Kullar R, Goff DA, Gauthier TP, Smith TC, To tweet or not to tweet—a review of the viral power of twitter for infectious diseases, Curr Infect Dis Rep 22 (6) (Apr. 2020) 14, 10.1007/s11908-020-00723-0. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [40]. Lian AT, Du J, Tang L, Using a Machine Learning Approach to Monitor COVID-19 Vaccine Adverse Events (VAE) from Twitter Data, Vaccines (basel) 10 (1) (Jan. 2022) 103, 10.3390/vaccines10010103. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [41]. Adams NN, “‘Scraping’ Reddit posts for academic research? Addressing some blurred lines of consent in growing internet-based research trend during the time of Covid-19,” International Journal of Social Research Methodology, Jan. 2024, Accessed: Apr. 24, 2024. [Online]. Available: https://www.tandfonline.com/doi/abs/10.1080/13645579.2022.2111816 . [ Google Scholar ] [42]. Vytniorgu R, Coming to voice as total top or total bottom: autobiographical acts and the sexual politics of versatility on reddit, Journal of Homosexuality (2024) 1–18, 10.1080/00918369.2024.2307544. [ DOI ] [ PubMed ] [ Google Scholar ] [43]. Wang J, Patel P, Jagdeo J, An analysis of keloid patient questions on Reddit, Wound Repair and Regeneration 32 (2) (2024) 164–170, 10.1111/wrr.13160. [ DOI ] [ PubMed ] [ Google Scholar ] [44]. Soysal E, et al. , CLAMP - a toolkit for efficiently building customized clinical natural language processing pipelines, J Am Med Inform Assoc 25 (3) (Mar. 2018) 331–336, 10.1093/jamia/ocx132. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [45]. Singh S, Mahmood A, The NLP Cookbook: modern recipes for transformer based deep learning architectures, IEEE Access 9 (2021) 68675–68702, 10.1109/ACCESS.2021.3077350. [ DOI ] [ Google Scholar ] [46]. Kamnis S, Generative pre-trained transformers (GPT) for surface engineering, Surface and Coatings Technology 466 (Aug. 2023) 129680, 10.1016/j.surfcoat.2023.129680. [ DOI ] [ Google Scholar ] [47]. Olivero S, “Figurative Language Understanding based on Large Language Models,” laurea, Politecnico di Torino, 2024. Accessed: Apr. 24, 2024. [Online]. Available: https://webthesis.biblio.polito.it/30393/ . [ Google Scholar ] [48]. Tsantekidis A, Passalis N, and Tefas A, “Chapter 5 - Recurrent neural networks,” in Deep Learning for Robot Perception and Cognition, Iosifidis A and Tefas A, Eds., Academic Press, 2022, pp. 101–115. doi: 10.1016/B978-0-32-385787-1.00010-5. [ DOI ] [ Google Scholar ] [49]. Lipton ZC, Berkowitz J, Elkan C, Critical Review of Recurrent Neural Networks for Sequence Learning, (2015). [ Google Scholar ] [50]. Salehinejad H, Sankar S, Barfett J, Colak E, Valaee S, “recent Advances in Recurrent Neural Networks, (2018). [51]. Lee J et al. , “BioBERT: a pre-trained biomedical language representation model for biomedical text mining,” Bioinformatics, p. btz682, Sep. 2019, doi: 10.1093/bioinformatics/btz682. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] [52]. Zhu R, Tu X, and Huang JX, “5 - Utilizing BERT for biomedical and clinical text mining,” in Data Analytics in Biomedical Engineering and Healthcare, Lee KC, Roy SS, Samui P, and Kumar V, Eds., Academic Press, 2021, pp. 73–103. doi: 10.1016/B978-0-12-819314-3.00005-7. [ DOI ] [ Google Scholar ] [53]. Nøkleby H, Bergsaker MAR, Adverse events after vaccination, Tidsskr nor Laegeforen 126 (19) (Oct. 2006) 2541–2544. [ PubMed ] [ Google Scholar ] [54]. Huang J, et al. , Monitoring vaccine safety by studying temporal variation of adverse events using vaccine adverse event reporting system, The Annals of Applied Statistics 15 (1) (Mar. 2021) 252–269, 10.1214/20-AOAS1393. [ DOI ] [ Google Scholar ] [55]. Loupi E, Baudard S, Debois H, Pignato F, Risks associated with vaccinations, Ann Med Interne (paris) 149 (6) (Oct. 1998) 361–371. [ PubMed ] [ Google Scholar ] [56]. Tiffet T et al. , “Comparing a Large Language Model with Previous Deep Learning Models on Named Entity Recognition of Adverse Drug Events,” in Digital Health and Informatics Innovations for Sustainable Health Care Systems, IOS Press, 2024, pp. 781–785. doi: 10.3233/SHTI240528. [ DOI ] [ PubMed ] [ Google Scholar ] [57]. Li Y et al. , “Development of a Natural Language Processing Tool to Extract Acupuncture Point Location Terms,” in 2023 IEEE 11th International Conference on Healthcare Informatics (ICHI), Jun. 2023, pp. 344–351. doi: 10.1109/ICHI57859.2023.00053. [ DOI ] [ Google Scholar ] [58]. He J, et al. , Prompt tuning in biomedical relation extraction, J Healthc Inform Res (2024), 10.1007/s41666-024-00162-9. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials Supplementary NIHMS2162270-supplement-Supplementary.docx (9.4KB, docx) ACTIONS View on publisher site PDF (696.2 KB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top