Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice Anat Cell Biol . 2026 Jan 23;59(1):59–67. doi: 10.5115/acb.25.295 Search in PMC Search in PubMed View in NLM Catalog Add to search Efficiency of artificial intelligence in identifying histological tissues from microscopic images Mustafa Saad Yousuf Mustafa Saad Yousuf 1 Department of Anatomy, Physiology, and Biochemistry, Faculty of Medicine, The Hashemite University, Zarqa, Jordan Find articles by Mustafa Saad Yousuf 1, ✉ , Bashar Issa Almaraziq Bashar Issa Almaraziq 2 Faculty of Medicine, The Hashemite University, Zarqa, Jordan Find articles by Bashar Issa Almaraziq 2 Author information Article notes Copyright and License information 1 Department of Anatomy, Physiology, and Biochemistry, Faculty of Medicine, The Hashemite University, Zarqa, Jordan 2 Faculty of Medicine, The Hashemite University, Zarqa, Jordan ✉ Corresponding author: Mustafa Saad Yousuf, Department of Anatomy, Physiology, and Biochemistry, Faculty of Medicine, The Hashemite University, Zarqa 13133, Jordan, E-mail: [email protected] Received 2025 Sep 19; Revised 2025 Oct 30; Accepted 2025 Nov 11; Issue date 2026 Mar 31. © 2026. Anatomy & Cell Biology This is an Open Access article distributed under the terms of the Creative Commons Attribution Non-Commercial License ( http://creativecommons.org/licenses/by-nc/4.0 ) which permits unrestricted non-commercial use, distribution, and reproduction in any medium, provided the original work is properly cited. PMC Copyright notice PMCID: PMC13072644 PMID: 41571386 Abstract Histology is an essential, yet difficult, subject to study in medicine. As artificial intelligence (AI) is rapidly evolving with ever-growing image recognition abilities, it presents a great potential to be used in the study of histology. This research aimed to assess the ability of ChatGPT 4o and Google Gemini 2.0 Flash to recognize various histological sections of different tissues and organs from images. The two AI models were presented with high-resolution histological images and prompted with tasks to identify basic tissue types, specific organs, and specific structures indicated by arrows. Scores were given for correct identifications. The test was conducted twice without any training or feedback to assess consistency. McNemar test and Kappa coefficient were used to compare the responses. Both AI models had good image recognition abilities with varying performances. Overall, Google Gemini achieved higher correct scores across both tests (75% and 77.5%, compared to 65% and 65% for ChatGPT). McNemar showed a significant difference between the two models with only fair agreement shown by kappa. No significant difference was found between the repeated tests. Performance across different tissue types varied. Muscular tissue was the easiest to identify and epithelial tissue was the most difficult. AI may have a promising role in histology. Students and histologist can use AI, especially Gemini, to identify histological sections from images. Critical judgment, however, must be utilized as AI can make mistakes. Keywords: Artificial intelligence, ChatGPT 4o, Google Gemini 2.0 flash, Histology, Microscopic sections Introduction During the early years of their academic learning, medical students study various basic medical sciences. One of these sciences is histology, the study of the normal morphology of cells and tissues, their arrangement in an organ, and how these relate to function [ 1 ]. By studying the normal histology of a tissue, students can get an indication of its physiology and be able to compare morphology in normal and abnormal conditions. Despite its importance, students face a lot of difficulties in studying histology [ 2 , 3 ]. Tissues and cells are very small and their study was not possible until pioneers, like Malpighi and Swammerdam in the seventeenth century, first used light microscopes for their examination [ 4 ]. With further development of different types of microscopes and different tissue preparation techniques, the study of histology advanced. The teaching of histology to students also saw a progressive shift towards the use of computers and virtual microscopes [ 5 ]. So, without technology, histology would not have been possible. The latest technology to develop was artificial intelligence (AI). It has evolved so rapidly that it changed the way we deal with daily tasks [ 6 ]. With the latest introduction of large language models, such as Google Gemini (Google) and ChatGPT (OpenAI), with its vast text search and analysis capabilities, medicine, as other disciplines, saw potential benefits in AI [ 7 - 9 ]. These benefits could range from staying updated about new advances in the field to creating virtual assistants to patients [ 7 ]. Despite academic integrity and ethical problems, AI can also have a role in medical research, including the assistance in the formulation of better methodology and data analysis plan, summarizing medical information and creating personalized and effective treatment plans, and, importantly, in the development of new drugs with minimum side effects [ 10 ]. Furthermore, AI models can be created to obtain accurate and fast diagnosis of various illnesses [ 11 ]. In medical education, AI tools have been used in the training of medical students and doctors on virtual cases mimicking real-world medical scenarios and situations [ 12 ]. The study by Dirks-Naylor (2025) [ 13 ] to determine students’ perception of the inclusion of ChatGPT in classroom assignments showed that almost all students found the tool easy to use, and most of them found the quizzes generated by ChatGPT were as effective as the ones created by their instructors. In addition, all participating students agreed that double-checking the AI tool-generated answers using their trusted sources was valuable and helpful in their learning experience [ 13 ]. Image analysis and recognition is another function of AI models that can have an important role in medicine. AI may help cardiologists interpret electrocardiogram images to identify atrial fibrillation in a fast and cost-effective manner [ 14 ], and in the analysis of echocardiography 2D images [ 15 ]. AI may also be used in the diagnosis of retinal diseases based on images [ 16 ]. In addition, AI was used in the classification and identification of different types of bacteria based on images of cultures [ 17 ]. Histopathology is an important branch of medicine that can benefit greatly from the image analysis capabilities of AI tools. It can be used in the diagnosis or classification of pathologies from images, or finding images similar to an example given to the AI model [ 18 ]. Apornvirat et al. (2024) [ 19 ] found that ChatGPT was less efficient than human pathologist, but still has some potential. In a limited review, Cazzato et al. (2023) [ 20 ], found that the use of ChatGPT in histopathology was limited due to problems like hallucinations and insufficient training. This was supported by another research [ 21 ] that showed human pathologist to outscore ChatGPT in pathology diagnosis, but encouraged the use of AI models in a collaborative manner with human pathologists. When focusing on a specific pathology, however, high AI accuracy was found for the diagnosis of colorectal polyps [ 22 ] and endometrial pathologies [ 23 ]. Limitations facing the use of AI in histopathology include improper image processing (image size, magnification, and staining), insufficient labelling, and presence of artifacts [ 18 ]. But, with further development of AI tools and their proper training on well-processed images, their image analysis capabilities may be enhanced. Research about the use of AI in histology, on the other hand, is limited. One study suggested the use of AI to stain tissue images [ 24 ]. If AI tools can recognize normal tissues from images, that would, probably, enhance its ability to recognize pathological conditions from images. In addition, such ability could help students in their study of histology. The aim of this research was to assess the efficiency of AI models to recognize different histological sections from images and to assess differences between two AI models. Materials and Methods In this research, images of histological sections were presented to two untrained AI models to test their ability to recognize the sections. The AI models used were ChatGPT 4o and Google Gemini 2.0 Flash. These models were not specifically trained to recognize images of histological tissues; moreover, the latest, most stable, and more easily available version of the models at the time the research was conducted were chosen (images were presented to the models during February and March 2025). The images were downloaded from the website Histology Guide [ 25 ]. This website provides images of histological sections through various organs and tissues that are of high resolution and excellent staining quality. Images covering a representative range of histology were downloaded through the save image function provided by the website. The downloaded images had dimensions of 1280×498 pixels with pixel density of 96 dpi and 24-bit color depth, and they were all stained with hematoxylin-eosin. The images were checked for metadata, and none were found associated with any image. All images were downloaded from one site to ensure that they all had the same specifications (format, resolution, dimensions, color depth) so that such factors would not affect image analysis by the AI models. Three groups were made with different number of images and questions ( Table 1 ). In Group 1, the main focus was on the recognition of different types of basic body tissues. The AI model was asked to identify the tissues present in the image (some images contained more than one type), and the organ from which the image was possibly taken. A point was given for each tissue correctly identified, and 0.5 points for each organ correctly identified. An example of the images used in this group was a section through the kidney showing the renal corpuscle with its simple squamous epithelium, and renal tubules with simple cuboidal epithelium. Another example was a section through the trachea showing the respiratory epithelium, underlying connective tissue, and hyaline cartilage. Table 1. Properties of the image groups and the prompt used for each Group Images Questions Marks Prompt Group 1 24 30 42 Act as a histologist. You will be shown an image of a histological section. Your task is to analyze the image and answer the following questions:Q1) Identify all the tissues seen in the image.Q2) What is the most likely organ from which the image was taken? Group 2 26 28 28 Act as a histologist. You will be shown an image of the histology of an organ. Your task is to analyze the image and identify the organ. Group 3 30 30 30 Act as a histologist. You will be shown an image of a histological section. A specific structure in the image is indicated by an arrow. Your task is to analyze the image and identify the structure indicated by the arrow. Be specific in your answer. Open in a new tab In Group 2, direct identification of an organ from the image was requested. A point was given for each correctly recognized organ (some images contained two organs). The difference between Group 1 and 2 in the points given to organ identification was because, in Group 1, the AI identified the tissues by analyzing the image, but, most probably, identified the organ from its vast knowledge database; whereas, in Group 2, organ identification was from image analysis directly and therefore was given a higher score. Examples of these images were sections through the larynx, cerebellar cortex, tongue, thymus, ovary, and retina. In Group 3, a specific structure was indicated by an arrow and the AI model was asked to identify the indicated structure. A point was given for each structure identified correctly. Examples included a section through cardiac muscle fibers with an arrow pointing at an intercalated disc; a section through an intestinal villus with an arrow pointing at the brush border; and a section through the skin with the arrow pointing at the sebaceous gland. The role-task format of prompt was used in which the AI model was given the role of a ‘histologist’ and given a different ‘task’ for each of the three groups ( Table 1 ). This prompt format was seen to be successful in giving the best responses from AI models [ 26 ]. Since this research aimed to identify the efficiency of current AI models without training, the chat of any of the groups was deleted before proceeding with the next group. The responses of the AI models in the groups were scored and the total score calculated. In addition, the questions in Groups 1 and 3 were categorized according to the type of basic body tissue they asked about (epithelial, connective, muscular, and nervous tissue) and the AI score in these categories determined. Table 2 shows the number of images and questions asking about specific tissue types. Some images asked about more than one type of tissue. Response accuracy was calculated as the percentage of correctly answered questions. The models were tested twice (with the same image sets) to see if the responses remained consistent. Table 2. Number of images and questions asking about specific types of body tissue Tissues Epithelial Connective Nervous Muscular Images 18 19 10 10 Questions 20 20 10 10 Open in a new tab Some images asked about more than one type of tissue. The responses were dichotomous (correct/incorrect) and considered as matched pairs (since the models were tested on the same set of images); therefore, McNemar test was used to find if there were any differences in the proportions of responses between the two tests for each AI model, as well as between the responses of the two AI model. The degree of agreement between the responses was assessed using Cohen’s kappa coefficient. This measured the degree to which the agents provided the same answers to the same questions across the different tests, which is a reflection of their reliability. The interpretation of kappa by Landis and Koch (1977) [ 27 ] was used. According to this interpretation, the higher values of kappa indicate stronger agreement and the less possibility that this agreement was due to chance. Analysis was done using Microsoft Excel (Microsoft) and the Jamovi Project Software v2.7 (The Jamovi Project). A P -value less than or equal to 0.05 was considered significant. Quality of the answers were also analyzed. The images were chosen and the responses evaluated and scored by a specialist in histology. The research dealt with AI platforms only. However, the research was approved by the Institutional Review Board of the Hashemite University (No. 12/3/2024/2025). Results Quantitative analysis The three groups of images were presented to ChatGPT 4o and Gemini 2.0 Flash twice. The results of the tests are shown in Table 3 . The total score of ChatGPT was 65% in both tests; whereas, Gemini scored 75% in the first test, and 77.5% in the second. Table 3. The score of presenting the image groups to the two artificial intelligence models Images Maximum mark AI model score (%) ChatGPT Gemini Test 1 Test 2 Test 1 Test 2 Group 1 42 35 (83) 35 (83) 33 (79) 36.5 (87) Group 2 28 12 (43) 11 (39) 21 (75) 20 (71) Group 3 30 18 (60) 19 (63) 21 (70) 21 (70) Total 100 65 65 75 77.5 Open in a new tab AI, artificial intelligence. Some of the questions were answered differently in the two tests. For ChatGPT, there were 2, 7, and 3 discrepancies in Group 1, Group 2, and Group 3, respectively. For Gemini, there were 3, 5, and 4 different answers in Groups 1, 2, and 3, respectively. In both models, the total was 12. McNemar’s test showed that the difference in proportions between the responses of ChatGPT in the first and second tests was not significant ( P =1.00). The same was found for Gemini ( P =0.773). Using kappa coefficient to assess agreement between responses showed that both ChatGPT (kappa=0.709; 95% CI=0.500 to 0.918; P <0.001) and Gemini (kappa=0.637; 95% CI=0.428 to 0.845; P <0.001) had a significant substantial agreement between their responses in the two tests. With a total number of image recognition questions of 88 ( Table 1 ), the confusion matrix for the response of the two models in the two tests is shown in Tables 4 and 5 . For the first test, McNemar’s analysis showed that the difference in the response proportions between the two models was significant (χ 2 [1, n=88]=3.85; P =0.050) with only fair agreement (kappa=0.329; 95% CI=0.127 to 0.531; P =0.001). The difference in the response proportions in the second test was also significant (χ 2 [1, n=88]=6.00; P =0.014) and with only fair agreement (kappa=0.373; 95% CI=0.174 to 0.571; P <0.001). Table 4. Confusion matrix used in McNemar test and kappa coefficient calculation for the agreement between the first test of ChatGPT and Gemini Gemini correct Gemini incorrect Total ChatGPT correct 47 (53.4) 8 (9.1) 55 (62.5) ChatGPT incorrect 18 (20.5) 15 (17.0) 33 (37.5) Total 65 (73.9) 23 (26.1) 88 (100) Open in a new tab Values are presented as number (%). Percentages are calculated based on the total number of cases (n=88). Table 5. Confusion matrix used in McNemar test and kappa coefficient calculation for the agreement between the second test of ChatGPT and Gemini Gemini correct Gemini incorrect Total ChatGPT correct 49 (55.6) 6 (6.8) 55 (62.5) ChatGPT incorrect 18 (20.5) 15 (17.0) 33 (37.5) Total 67 (76.1) 21 (23.9) 88 (100) Open in a new tab Values are presented as number (%). Percentages are calculated based on the total number of cases (n=88). Regarding the different types of tissues ( Table 2 ), there were 20 questions that asked about epithelium, 20 about connective tissue, 10 about muscular, and 10 about nervous tissues (more questions asked about epithelial and connective tissues, because there are more types of these tissues in the body). Fig. 1 shows the percentage of correct answers to the questions about the tissues in the two tests of the models. Muscular and connective tissues were better identified, whereas epithelium was the least correctly identified. Overall, ChatGPT correctly answered 72% of the questions about tissues in the first test, and 73% in the second test; whereas, Gemini identified 73% and 78% in the first and second tests, respectively. Fig. 1. Open in a new tab The percentage of correct answers given to questions about tissues. The number beside the artificial intelligence model name is the test number. As for organ recognition in Group 1, if the model was able to identify the tissues, it was able to identify the organ correctly 100% of the times; if it was unable to recognize the tissues, it failed to identify the organ. In Group 2, where the model had to identify the organ by image analysis, both had some difficulties, especially ChatGPT ( Table 3 ). Quality of the answers How the model expressed and presented the answers differed between models and between tests for each model. Gemini gave more concise answers with better explanations of how it reached its conclusions. Although the prompt was the same in the repeated tests, the model’s presentation of the answer differed slightly. In few instances, the AI model identified structures that were not there (hallucinations). One example is when the model was presented with an image of parotid gland histology showing serous acini and some ducts (Group 1); in one test, both models were able to identify the serous acini, but interpreted the ducts as pancreatic islets, thus identifying the organ as the pancreas ( Fig. 2 ). However, in the same test, both models were not able to identify the pancreas when presented with an image of the organ (Group 2). Gemini, in one question, was able to identify a muscular artery, but misidentified the red color of its smooth muscles as red blood cells. Although no interaction was done with the models other than the prompt, Gemini sometimes directly addressed the user in the form ‘You’re right to point out…’ Fig. 2. Open in a new tab An example of an interaction with an artificial intelligence (AI; Google Gemini 2.0 Flash) model. The image presented to the model in this example can be found at [ 28 ]. In this interaction, the AI model misidentified the acidophilic ducts as islets of Langerhans. Based on this misidentification, the model determined that the shown organ was the pancreas. The models were able to identify a structure when presented in one group, but failed to do so when presented in another group. In Group 2, an image of a dorsal root ganglion (showing the capsule, neurons, and nuclei of satellite cells) was shown to the AI models. For Group 3, the same image was used but an arrow pointing at the nucleus of a satellite cell was added. The models correctly identified the organ when presented in Group 2, but failed to identify the satellite cell when indicated in Group 3 (or vice versa). Also, all models easily identified smooth muscle cells when asked about types of tissue present, but sometimes failed to identify the arrector pili muscle of the skin when indicated by an arrow. The esophagus was presented twice to the models, once as a separate organ, and once as part of the gastroesophageal junction. Out of all the tests done, only once was it correctly identified (in the gastroesophageal junction) by Gemini. An example interaction with an AI model is shown in Fig. 2 . An image of a section through the parotid gland [ 28 ] was presented to Gemini as part of Group 1. As the prompt indicates, it was asked to identify all tissues seen. The image showed several basophilic acini and acidophilic ducts. Pancreatic islets are not seen in the image uploaded to the model. Their presence, however, cannot be ruled out; therefore, an identification of parotid gland or pancreas was considered acceptable. It can be seen in Fig. 2 that Gemini was able to identify the serous acini, but misidentified the ducts as pancreatic islets giving the conclusion that the image is a section through the pancreas. With this answer, the model was given 1.5 points (1 point for identifying the serous acini and 0.5 for giving the pancreas as a possible organ). Discussion In this research, the efficiency of Google Gemini and ChatGPT to recognize histological tissues by analyzing images was assessed. The results showed that the AI models differed in their performance in this task, with Gemini outscoring ChatGPT by 10–13 points. The difference was not only seen in the accuracy of the models in identifying histological images (65% in both ChatGPT tests, 75% and 77.5% for Gemini; Table 3 ), but McNemar analysis showed that those differences were significant in the first and second tests. In addition, kappa analysis showed only a fair agreement between the responses of the two models when they were compared in the first and second tests, indicating that this agreement between the models could have been due to chance. The difference in performance may be attributed to the algorithms used by the models to reach their conclusions and to the probabilistic way in which answers are generated [ 29 ]. The probabilistic nature of response generation may have also been the cause of the discrepancies seen in the answers given by the models when the same question was asked in the two different tests. This was similar to findings by another research [ 30 ] that found that giving the same questions to ChatGPT three times resulted in different answers (especially older versions). Despite these discrepancies in the response of the models between the tests, kappa analysis showed that each model exhibited substantial agreement when their answers in the two tests were compared, and McNemar analysis found no significant difference between them. This can also be seen in the small differences in the values shown in the confusion matrices ( Tables 4 and 5 ). However, the possibility of a discrepancy must be kept in mind when using AI models, especially when asking to identify tissues, because, although the answer to a question may be convincing with a logical interpretation, if the question is asked again, the model may give a completely different answer, still with a convincing argument. The ability to identify tissues varied between models and, for each model, between the two tests. Epithelial tissues were the least correctly identified; while, muscular tissues were the most correctly identified. This may not be surprising, since there are only three types of muscular tissues with distinguishing characteristics features that make their identification relatively easy; whereas, several types of epithelial tissues are found with features that are similar to each other making their identification a challenge (it may not have been easy for the models to distinguish between the different shapes of the nuclei of epithelial cells or to identify the number of layers). The ability of the two AI models to identify the different types of tissues may be compared to the ability of junior medical students with just basic knowledge of histology. In their early years of study, such students do not have the experience (training) required to easily identify histological slides. The same can be said about the AI models; they don’t have enough training to be able to recognize the minute differences between tissues. García et al. [ 31 ] showed, in their research, that students found muscular tissue to be the easiest to identify and epithelial tissue the second difficult (most difficult was the nervous). In this research, Gemini achieved higher total score ( Table 3 ), higher overall score in the questions regarding tissue type, and better identification of connective and nervous tissues, and better recognition of organs ( Table 3 ; Group 2). ChatGPT, on the other hand, recognized epithelial and muscular tissues better. Regarding the various models and their image recognition ability, previous studies have shown varying results. Carlà et al. [ 16 ] found that ChatGPT 4o was better at diagnosing abnormalities from optical coherence tomography images than Gemini (with an accuracy of 62%). For neurosurgical images, McNulty et al. [ 32 ] found ChatGPT 4.0 to be more accurate than Gemini (accuracy was 54%). Pradhan [ 33 ] found ChatGPT 4o to be more accurate (67%) in diagnosing images of oral potentially malignant lesions than Gemini. Although research on normal histology is scarce, several studies were conducted to study the role of AI in histopathology. In a review by Lui et al. [ 22 ], AI was found to have an accuracy of 96% in the detection of colorectal polyps from histopathological images. On the other hand, Apornvirat et al. [ 19 ] showed that ChatGPT 4 had an accuracy of 50% in identifying pathologies from images. Using a customized AI model, Sheakh et al. [ 23 ] achieved an accuracy of 99% in detecting endometrial pathologies from images. These differences in the results obtained by the different studies could have been due to the different AI models used, the version of the AI model, and the prompts given to the platform. In this research, Google Gemini 2.0 Flash was used. This provided great advancement from the previous versions of Gemini, making it perform better than ChatGPT 4o in identifying histological tissues from images. With the rapid development of both ChatGPT and Gemini (and other models), similar studies in the future may yield different results. Hallucinations were observed where the models misidentified one part of an image and reached conclusions based on these misidentifications, although the models interpreted everything else in the image correctly. For example, a muscular artery was correctly identified and the models gave correct reasons for their conclusions, except that the red color of the wall was caused by red blood cells rather than smooth muscles. As such, hallucinations meant an incorrect identification of structures and, therefore, affected the performance of the models in the tests. Another, interesting, ‘hallucination’ was when Gemini sometimes addressed the user by stating ‘You’re right to point out…,’ despite not having ‘pointing out’ anything to the AI. Such a statement might be an indication of metadata attached to the image. However, there was no metadata with the images used in this research. One problem that could have prevented the AI models from correctly identifying the tissues from images was their inability to magnify or move the image. A human studying a histologic slide under a microscope can alter the magnification or field of view, which can facilitate the identification of the tissue or organ. Similarly, an organ or tissue can appear substantially different depending on its plane of section and could also affect identification. In addition, the presence of artifacts in the image (like artificial spaces, tissue folding, or stain blots) may pose a problem to the AI model. Providing the AI model with very high-resolution, well-stained images of tissues and organs can help it overcome many of these problems. The most important solution to these problems, however, is training [ 18 , 34 , 35 ]. Qualified histologists should provide a large database of histological images to the AI platform with a continuous interchange between humans and the AI to provide constructive feedback on the AI’s responses to enhance its ability to recognize tissues from images [ 18 ]. The responses of both AI models were given with great confidence. What features to look for and how the models reached their conclusions were explained in a logical and easily comprehensible way. However, when asked to identify the same image a second time, both AI models sometimes gave a different (and incorrect) answer, but with the same confidence. If the image had been shown again, another completely different answer might have been given; thus, the same image might have different interpretations. Rather than stemming from an understanding of the question or problem, AI models generate their responses probabilistically from the large databases they could access [ 29 ]. This could be considered similar to the Dunning–Kruger effect in humans in which individuals with lower knowledge tend to overestimate their abilities [ 36 ]. The apparent confidence of the AI model in its output may hide behind it a limited understanding of the problem presented. Without critical judgment, the lack of knowledge of these limitations may lead the users to over-trust the output of the AI model, which may augment the Dunning–Kruger effect in the users themselves [ 37 ]. This may cause the users to believe that the answers of the AI are correct and, thus, make decisions based on this belief [ 38 , 39 ]. How could students actually benefit from these results? An issue that students may encounter while studying histology is the difficulty in identifying various features of tissues from outside sources, such as online virtual slides or microscopic slides seen in lab sessions (images in lectures are usually labelled). The use of multimodal AI tools (especially Gemini) offers several possible ways for students to overcome this problem. Students can upload an image of a microscopic slide to the AI platform, or capture it directly through the camera function, and ask the model to identify the slide, give its main features, and how to identify it. This can, potentially, enhance the ability of the students to comprehend what they are viewing under the microscope. Students can also ask the AI tools to do a side-by-side comparison of two histological slides to better differentiate tissues, cells, and organs. In addition, Gemini can be used to create quizzes and flashcards based on the uploaded image so that students can use them to test their knowledge. Students should be allowed to engage freely with these platforms, but they must also be encouraged to use caution in interpreting the responses of the AI tools and to seek knowledge from their instructors and valid sources [ 40 ]. Implementing mechanisms that allow AI systems to acknowledge uncertainty and provide the scientific resources on which it depended to make conclusions could enhance the reliability and accuracy of AI-assisted histological analysis. In conclusion, AI models like Gemini Flash and ChatGPT showed promising results in the recognition of histological tissues from images. Total reliance on them, however, is not recommended as they are liable for errors. AI tools can be integrated in the study of histology, especially Google Gemini Flash. Students, however, must be encouraged to use critical thinking to judge the response of these tools. Footnotes Author Contributions Conceptualization: MSY, BIA. Data acquisition: MSY, BIA. Data analysis: MSY, BIA. Investigation: MSY, BIA. Methodology: MSY. Software: MSY. Validation: MSY, BIA. Visualization: BIA. Drafting of the manuscript: MSY, BIA. Critical revision of the manuscript: MSY. Approval of the final version of the manuscript: all authors. Conflicts of Interest No potential conflict of interest relevant to this article was reported. Funding None. References 1. Mescher AL. Histology & its methods of study. In: Mescher AL, editor. Junqueira's Basic Histology: Text and Atlas. 17th ed. McGraw Hill; 2024. 2. Mantovani ALS, Lima ARDA, Brienze SLA, Dos Santos ER, Fucuta P, André JC. Cell biology and histology in medicine: perception on education and student performance. Int J Health Educ. 2019;3:8–16. doi: 10.17267/2594-7907ijhe.v3i1.2099. [ DOI ] [ Google Scholar ] 3. Mohammedsaleh Z. The impact of various factors on the difficulties in learning and teaching strategies for histology. Int J Morphol. 2024;42:741–8. doi: 10.4067/s0717-95022024000300741. [ DOI ] [ Google Scholar ] 4. Bracegirdle B. The history of histology: a brief survey of sources. Hist Sci. 1977;15:77–101. doi: 10.1177/007327537701500201. [ DOI ] [ Google Scholar ] 5. Blake CA, Lavoie HA, Millette CF. Teaching medical histology at the University of South Carolina School of Medicine: transition to virtual slides and virtual microscopes. Anat Rec B New Anat. 2003;275:196–206. doi: 10.1002/ar.b.10037. [ DOI ] [ PubMed ] [ Google Scholar ] 6. Lee RST. Artificial intelligence in daily life. Springer Singapore; 2020. 10.1007/978-981-15-7695-9 [ DOI ] 7. Dave T, Athaluri SA, Singh S. ChatGPT in medicine: an overview of its applications, advantages, limitations, future prospects, and ethical considerations. Front Artif Intell. 2023;6:1169595. doi: 10.3389/frai.2023.1169595. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 8. Hamet P, Tremblay J. Artificial intelligence in medicine. Metabolism. 2017;69S:S36–40. doi: 10.1016/j.metabol.2017.01.011. [ DOI ] [ PubMed ] [ Google Scholar ] 9. Salau AO, Jain S, Sood M. Computational intelligence and data sciences: paradigms in biomedical engineering. CRC Press; 2022. 10.1201/9781003224068 [ DOI ] 10. Ruksakulpiwat S, Kumar A, Ajibade A. Using ChatGPT in medical research: current status and future directions. J Multidiscip Healthc. 2023;16:1513–20. doi: 10.2147/JMDH.S413470. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 11. Kaur S, Singla J, Nkenyereye L, Jha S, Prashar D, Joshi GP. Medical diagnostic systems using artificial intelligence (AI) algorithms: principles and perspectives. IEEE Access. 2020;8:228049–69. doi: 10.1109/ACCESS.2020.3042273. [ DOI ] [ Google Scholar ] 12. Hamilton A, Molzahn A, McLemore K. The evolution from standardized to virtual patients in medical education. Cureus. 2024;16:e71224. doi: 10.7759/cureus.71224. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 13. Dirks-Naylor AJ. Developing student proficiency in ChatGPT-driven active recall practices and self-guided inquiry. Adv Physiol Educ. 2025;49:960–4. doi: 10.1152/advan.00112.2025. [ DOI ] [ PubMed ] [ Google Scholar ] 14. Attia ZI, Noseworthy PA, Lopez-Jimenez F, Asirvatham SJ, Deshmukh AJ, Gersh BJ, Carter RE, Yao X, Rabinstein AA, Erickson BJ, Kapa S, Friedman PA. An artificial intelligence-enabled ECG algorithm for the identification of patients with atrial fibrillation during sinus rhythm: a retrospective analysis of outcome prediction. Lancet. 2019;394:861–7. doi: 10.1016/S0140-6736(19)31721-0. [ DOI ] [ PubMed ] [ Google Scholar ] 15. Narula S, Shameer K, Salem Omar AM, Dudley JT, Sengupta PP. Machine-learning algorithms to automate morphological and functional assessments in 2D echocardiography. J Am Coll Cardiol. 2016;68:2287–95. doi: 10.1016/j.jacc.2016.08.062. [ DOI ] [ PubMed ] [ Google Scholar ] 16. Carlà MM, Crincoli E, Rizzo S. Retinal imaging analysis performed by ChatGPT-4o and Gemini advanced: the turning point of the revolution? Retina. 2025;45:694–702. doi: 10.1097/IAE.0000000000004351. [ DOI ] [ PubMed ] [ Google Scholar ] 17. Hindy JR, Souaid T, Kovacs CS. Capabilities of GPT-4o and Gemini 1.5 Pro in Gram stain and bacterial shape identification. Future Microbiol. 2024;19:1283–92. doi: 10.1080/17460913.2024.2381967. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 18. Komura D, Ishikawa S. Machine learning methods for histopathological image analysis. Comput Struct Biotechnol J. 2018;16:34–42. doi: 10.1016/j.csbj.2018.01.001. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 19. Apornvirat S, Thinpanja W, Damrongkiet K, Benjakul N, Laohawetwanit T. Comparing customized ChatGPT and pathology residents in histopathologic description and diagnosis of common diseases. Ann Diagn Pathol. 2024;73:152359. doi: 10.1016/j.anndiagpath.2024.152359. [ DOI ] [ PubMed ] [ Google Scholar ] 20. Cazzato G, Capuzzolo M, Parente P, Arezzo F, Loizzi V, Macorano E, Marzullo A, Cormio G, Ingravallo G. Chat GPT in diagnostic human pathology: will it be useful to pathologists? A preliminary review with 'Query Session' and future perspectives. AI. 2023;4:1010–22. doi: 10.3390/ai4040051. [ DOI ] [ Google Scholar ] 21. Oon ML, Syn NL, Tan CL, Tan KB, Ng SB. Bridging bytes and biopsies: a comparative analysis of ChatGPT and histopathologists in pathology diagnosis and collaborative potential. Histopathology. 2024;84:601–13. doi: 10.1111/his.15100. [ DOI ] [ PubMed ] [ Google Scholar ] 22. Lui TKL, Guo CG, Leung WK. Accuracy of artificial intelligence on histology prediction and detection of colorectal polyps: a systematic review and meta-analysis. Gastrointest Endosc. 2020;92:11–22.e6. doi: 10.1016/j.gie.2020.02.033. [ DOI ] [ PubMed ] [ Google Scholar ] 23. Sheakh MA, Azam S, Tahosin MS, Karim A, Montaha S, Fahim KU, Shafiabady N, Jonkman M, De Boer F. Ecgmlp: a novel gated mlp model for enhanced endometrial cancer diagnosis. Comput Methods Programs Biomed Update. 2025;7:100181. doi: 10.1016/j.cmpbup.2025.100181. [ DOI ] [ Google Scholar ] 24. Latonen L, Koivukoski S, Khan U, Ruusuvuori P. Virtual staining for histology by deep learning. Trends Biotechnol. 2024;42:1177–91. doi: 10.1016/j.tibtech.2024.02.009. [ DOI ] [ PubMed ] [ Google Scholar ] 25. Clark Brelje T, Sorenson RL. Histology Guide: Virtual Microscopy Laboratory [Internet]. Histology Guide; c2005-2025 [cited 2025 Jan 22]. Available from: https://histologyguide.com/ 26. Podder I, Pipil N, Dhabal A, Mondal S, Pienyii V, Mondal H. Evaluation of artificial intelligence-based chatbot responses to common dermatological queries. Jordan Med J. 2024;58:271–8. doi: 10.35516/jmj.v58i2.2960. [ DOI ] [ Google Scholar ] 27. Landis JR, Koch GG. The measurement of observer agreement for categorical data. Biometrics. 1977;33:159–74. doi: 10.2307/2529310. [ DOI ] [ PubMed ] [ Google Scholar ] 28. Clark Brelje T, Sorenson RL. Histology Guide: Virtual Microscopy Laboratory [Internet]. Histology Guide; c2005-2025 [cited 2025 Jan 22]. Available from: https://histologyguide.com/slideview/MH-094hr-parotid/12-slide-1.html?x=7305&y=9350&z=17.200 29. Alsajri A, Salman HA, Steiti A. Generative models in natural language processing: a comparative study of ChatGPT and Gemini. Babylon J Artif Intell. 2024;2024:134–45. doi: 10.58496/BJAI/2024/015. [ DOI ] [ Google Scholar ] 30. Funk PF, Hoch CC, Knoedler S, Knoedler L, Cotofana S, Sofo G, Bashiri Dezfouli A, Wollenberg B, Guntinas-Lichius O, Alfertshofer M. ChatGPT's response consistency: a study on repeated queries of medical examination questions. Eur J Investig Health Psychol Educ. 2024;14:657–68. doi: 10.3390/ejihpe14030043. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 31. García M, Victory N, Navarro-Sempere A, Segovia Y. Students' views on difficulties in learning histology. Anat Sci Educ. 2019;12:541–9. doi: 10.1002/ase.1838. [ DOI ] [ PubMed ] [ Google Scholar ] 32. McNulty AM, Valluri H, Gajjar AA, Custozzo A, Field NC, Paul AR. Performance evaluation of ChatGPT-4.0 and Gemini on image-based neurosurgery board practice questions: a comparative analysis. J Clin Neurosci. 2025;134:111097. doi: 10.1016/j.jocn.2025.111097. [ DOI ] [ PubMed ] [ Google Scholar ] 33. Pradhan P. Accuracy of ChatGPT 3.5, 4.0, 4o and Gemini in diagnosing oral potentially malignant lesions based on clinical case reports and image recognition. Med Oral Patol Oral Cir Bucal. 2025;30:e224–31. doi: 10.4317/medoral.26824. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 34. Hanna MG, Pantanowitz L, Dash R, Harrison JH, Deebajah M, Pantanowitz J, Rashidi HH. Future of artificial intelligence-machine learning trends in pathology and medicine. Mod Pathol. 2025;38:100705. doi: 10.1016/j.modpat.2025.100705. [ DOI ] [ PubMed ] [ Google Scholar ] 35. Liu S, McCoy AB, Wright A. Improving large language model applications in biomedicine with retrieval-augmented generation: a systematic review, meta-analysis, and clinical development guidelines. J Am Med Inform Assoc. 2025;32:605–15. doi: 10.1093/jamia/ocaf008. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 36. Kruger J, Dunning D. Unskilled and unaware of it: how difficulties in recognizing one's own incompetence lead to inflated self-assessments. J Pers Soc Psychol. 1999;77:1121–34. doi: 10.1037/0022-3514.77.6.1121. [ DOI ] [ PubMed ] [ Google Scholar ] 37. Guan J, He X, Su Y, Zhang X. The Dunning-Kruger effect and artificial intelligence: knowledge, self-efficacy and acceptance. Manag Decis. 2025;63:3786–802. doi: 10.1108/MD-06-2023-0893. [ DOI ] [ Google Scholar ] 38. Ishizu N, Yeoh WL, Okumura H, Fukuda O. The effect of communicating AI confidence on human decision making when performing a binary decision task. Appl Sci. 2024;14:7192. doi: 10.3390/app14167192. [ DOI ] [ Google Scholar ] 39. Pavlik EJ, Land Woodward J, Lawton F, Swiecki-Sikora AL, Ramaiah DD, Rives TA. Artificial intelligence in relation to accurate information and tasks in gynecologic oncology and clinical medicine-Dunning-Kruger effects and ultracrepidarianism. Diagnostics (Basel) 2025;15:735. doi: 10.3390/diagnostics15060735. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 40. Hamasaki MY, Maia AO, Lunardelli A. Does ChatGPT know basic human anatomy? Anat Cell Biol. 2025;58:493–5. doi: 10.5115/acb.25.181. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Articles from Anatomy & Cell Biology are provided here courtesy of Korean Association of Anatomists ACTIONS View on publisher site PDF (1.0 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top