ConceptioArchivearXiv CS
arXiv CSopen access

Mitosis Detection in the Wild: Multi-Tumor and Context-Aware Generalization in the MIDOG 2025 Challenge

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

Graphical Abstract

arXiv:2606.07368v1 [cs.CV] 5 Jun 2026

Mitosis Detection in the Wild: Multi-Tumor and Context-Aware Generalization in the MIDOG 2025 Challenge Marc Aubreville, Jonas Ammeling, Sweta Banerjee, Viktoria Weiss, Taryn A. Donovan, Robert Klopfleisch, Jiaqi Lv, Shan E Ahmed Raza, Raphaël Bourgade, Thomas Walter, Yasemin Topuz, Songül Varlı, Charles-Antoine Collins-Fekete, Zhuoyan Shen, Navya Sri Kelam, Nitin Singhal, Christian Marzahl, Brian Napora, Tengyou Xu, Hongyan Gu, Mario Vento, Gennaro Percannella, Norbert Ropiak, Izabela Wasiak, Jie Xiao, Shaojun Liu, Seungho Choe, April Khademi, Vidushi Walia, Sujatha Kotte, Andrew Broad, Alex Wright, Guillaume Balezo, Esha Sadia Nasir, Mostafa Jahanifar, Yosuke Yamagishi, Shouhei Hanaoka, Mattia Sarno, Francesco Tortorella, Biwen Meng, Jingxin Liu, Sara Krauss, Daniel Hieber, Lavish Ramchandani, Dev Kumar Das, Mieko Ochi, Yuan Bae, Piotr Giedziun, Mateusz Maniewski, Vangala Govindakrishnan Saipradeep, Naveen Sivadasan, Leire Benito-Del-Valle, Adrian Galdran, Kaustubh Atey, Sameer Anand Jha, Adinath Dukre, Imran Razzak, Maxime W. Lafarge, Viktor H. Koelzer, Nils Porsche, Nikolas Stathonikos, Mitko Veta, Dominik Hirling, Zsanett Zsófia Iván, Peter Horvath, Katharina Breininger, Christof A. Bertram

MItosis DOmain Generalization Challange (MIDOG) 2025 Challenge Overview and Insights Hotspot

Random

TRACK 1:

Challenging

Robust Mitotic Figure Detection (18 teams)

Detection Model

TRACK 2: Classification Model

Participant Methods, Results and Future Directions MICCAI 2025 Workshop Publication

Normal vs. Atypical Mitosis (21 teams)

NORMAL

ATYPICAL

Highlights Mitosis Detection in the Wild: Multi-Tumor and Context-Aware Generalization in the MIDOG 2025 Challenge Marc Aubreville, Jonas Ammeling, Sweta Banerjee, Viktoria Weiss, Taryn A. Donovan, Robert Klopfleisch, Jiaqi Lv, Shan E Ahmed Raza, Raphaël Bourgade, Thomas Walter, Yasemin Topuz, Songül Varlı, Charles-Antoine Collins-Fekete, Zhuoyan Shen, Navya Sri Kelam, Nitin Singhal, Christian Marzahl, Brian Napora, Tengyou Xu, Hongyan Gu, Mario Vento, Gennaro Percannella, Norbert Ropiak, Izabela Wasiak, Jie Xiao, Shaojun Liu, Seungho Choe, April Khademi, Vidushi Walia, Sujatha Kotte, Andrew Broad, Alex Wright, Guillaume Balezo, Esha Sadia Nasir, Mostafa Jahanifar, Yosuke Yamagishi, Shouhei Hanaoka, Mattia Sarno, Francesco Tortorella, Biwen Meng, Jingxin Liu, Sara Krauss, Daniel Hieber, Lavish Ramchandani, Dev Kumar Das, Mieko Ochi, Yuan Bae, Piotr Giedziun, Mateusz Maniewski, Vangala Govindakrishnan Saipradeep, Naveen Sivadasan, Leire Benito-Del-Valle, Adrian Galdran, Kaustubh Atey, Sameer Anand Jha, Adinath Dukre, Imran Razzak, Maxime W. Lafarge, Viktor H. Koelzer, Nils Porsche, Nikolas Stathonikos, Mitko Veta, Dominik Hirling, Zsanett Zsófia Iván, Peter Horvath, Katharina Breininger, Christof A. Bertram • Performance of automatic mitosis detection models collapses outside curated hotspot regions. Evaluating in random and challenging (rich in imposters) regions more than tripled the false positive rate (increase of 208%), exposing a clinical reliability gap. • Biological diversity reveals blind spots. Across 365 cases and 12 human, canine, and feline tumor types, top teams hit 0.740 F1 (detection) and 0.908 balanced accuracy (atypical classification) — but mitosis detection consistently underperformed on rare and highly pleomorphic tumors. • Ensembling helps, TTA does not. Model ensembling gave consistent gains (average increase of 1.5 percentage points in overall F1, 1.3 percentage points in overall balanced accuracy); test-time augmentation produced no meaningful improvement.

Mitosis Detection in the Wild: Multi-Tumor and Context-Aware Generalization in the MIDOG 2025 Challenge Marc Aubrevillea , Jonas Ammelingb , Sweta Banerjeea , Viktoria Weissc , Taryn A. Donovand , Robert Klopfleische , Jiaqi Lvf , Shan E Ahmed Razaf , Raphaël Bourgadeg , Thomas Walterg , Yasemin Topuzh , Songül Varlıh , Charles-Antoine Collins-Feketei , Zhuoyan Sheni , Navya Sri Kelamj , Nitin Singhalj , Christian Marzahlk , Brian Naporak , Tengyou Xul , Hongyan Gum , Mario Venton , Gennaro Percannellan , Norbert Ropiako , Izabela Wasiakp , Jie Xiaoq , Shaojun Liuq , Seungho Choer , April Khademir , Vidushi Walias , Sujatha Kottes , Andrew Broadt , Alex Wrightt , Guillaume Balezog , Esha Sadia Nasirf , Mostafa Jahanifarf , Yosuke Yamagishiu , Shouhei Hanaokau , Mattia Sarnon , Francesco Tortorellan , Biwen Mengv , Jingxin Liuv , Sara Kraussw , Daniel Hieberx , Lavish Ramchandanij , Dev Kumar Dasj , Mieko Ochiy , Yuan Baey , Piotr Giedziunz , Mateusz Maniewskio , Vangala Govindakrishnan Saipradeepr , Naveen Sivadasanr , Leire Benito-Del-Valleaa , Adrian Galdranaa , Kaustubh Ateyab , Sameer Anand Jhaab , Adinath Dukreac , Imran Razzakac , Maxime W. Lafargead , Viktor H. Koelzerad , Nils Porschea , Nikolas Stathonikosae , Mitko Vetaaf , Dominik Hirlingag , Zsanett Zsófia Ivánag , Peter Horvathag , Katharina Breiningerah and Christof A. Bertramc a Flensburg University of Applied Sciences, Flensburg, Germany b Technische Hochschule Ingolstadt, Ingolstadt, Germany c Pathology Unit, University of Veterinary Medicine, Vienna, Austria d Schwarzman Animal Medical Center, New York, USA e Institute of Veterinary Pathology, Freie Universität Berlin, Berlin, Germany f VISION Lab, Tissue Image Analytics Centre Department of Computer Science University of Warwick Coventry, UK g Centre for Computational Biology, MINES Paris - PSL University Paris, France h Vision Research Group, Department of Computer Engineering Yildiz Technical University Istanbul, Türkiye. i Department of Medical Physics and Biomedical Engineering, University College London London, UK j AIRA MATRIX Private Limited, Mumbai, India k Gestalt Diagnostics, Spokane WA 99202, USA l Department of Electrical and Computer Engineering, University of California Los Angeles Los Angeles, USA m Department of Pathology and Laboratory Medicine, University of Kansas Medical Center Kansas City, USA n Department of Information and Electrical Engineering and Applied Mathematics (DIEM) University of Salerno Fisciano, Salerno, Italy o Cancer Center Sp. z o. o., Wroclaw, Poland p Pathology Department 10th Military Research Hospital in Bydgoszcz Bydgoszcz, Poland q College of Health Science and Environmental Engineering, Shenzhen Technology University, Shenzhen, China r Image Analysis in Medicine Lab (IAMLAB), Electrical Computer and Biomedical Engineering Toronto Metropolitan University Toronto ON, Canada s TCS Research, Tata Consultancy Services Ltd. Hyderabad, India t National Pathology Imaging Co-operative, Leeds Teaching Hospitals NHS Trust, UK u Division of Radiology and Biomedical Engineering, Graduate School of Medicine The University of Tokyo Tokyo, Japan v School of AI and Advanced Computing, Xi’an Jiaotong-Liverpool University Suzhou, China w IT-Infrastructure for Translational Medical Research, Faculty of Applied Computer Science University of Augsburg, Germany x Institute of Neuropathology, Ulm University Medical Center Faculty of Medicine Ulm University Ulm, Germany y Department of Pathology, Japanese Red Cross Medical Center Tokyo, Japan z Department of Artificial Intelligence, Wroclaw University of Science and Technology Wrocław, Poland aa TECNALIA, Basque Research and Technology Alliance (BRTA) Parque TecnolÃşgico de Bizkaia C/ Geldo. Edificio 700, E-48160 Derio- Bizkaia (Spain) ab Centre for Machine Intelligence and Data Science, Indian Institute of Technology Bombay, India ac MBZUAI, Abu Dhabi, UAE ad Computational and Translational Pathology Lab, Department of Biomedical Engineering University of Basel Allschwil, Switzerland ae University Medical Center Utrecht, Utrecht, The Netherlands af TU Eindhoven, Eindhoven, The Netherlands ag HUN-REN Biological Research Centre, Institute of Biochemistry, Szeged, Hungary ah Julius-Maximilians-Universität Würzburg, Würzburg, Germany

[email protected] (M. Aubreville); [email protected] (C.A. Bertram)

ORCID (s):

M. Aubreville et al.: Preprint submitted to Elsevier

Page 1 of 18

The MIDOG 2025 Challenge

ARTICLE INFO

ABSTRACT

Keywords: pattern recognition challenge mitosis detection atypical mitosis domain generalization

Automated mitotic figure detection is a well-established task in computational pathology. Previous benchmarks addressing robustness have focused on scanner-induced domain shifts; however, realworld clinical use requires models that are robust to the much broader variability present in the histological landscape. The third edition of the MItosis DOmain Generalization (MIDOG) challenge, held in 2025, was designed to evaluate algorithmic performance across unprecedented biological and contextual diversity. The challenge was composed of two tracks: 1) mitotic figure object detection, and 2) classification of mitotic figures into normal and atypical morphologies. We curated a comprehensive test dataset of 365 cases, encompassing 12 distinct human, canine and feline tumor types, digitized across multiple scanning platforms. Moving beyond traditional hand-selected hotspots, this challenge additionally required detection in random tissue areas (representative of the whole slide image) and challenging areas (areas rich in imposters). Participants were tasked with developing architectures capable of maintaining high precision despite these varying levels of difficulty and tissue architecture. For track 1, predictions were submitted by 18 teams, with 𝐹1 scores of up to 0.740. In the second track, there were 21 submissions with balanced accuracy of up to 0.908. Our analysis reveals that while most models perform reliably in traditional hotspots, significant performance degradation occurs in challenging regions, where the false positive rate increased by 208%. Furthermore, performance varied significantly across the 12 tumor types, highlighting "blind spots" in current state-of-theart architectures when encountering rare or highly pleomorphic cancer types. We evaluated the effectiveness of ensembling and test-time augmentation (TTA) and found the former to have a consistently positive effect (mean increase of 1.5 percentage points in 𝐹1 score in track 1 and 1.3 in balanced accuracy in track 2), whereas TTA showed no relevant improvement. The MIDOG 2025 challenge demonstrates that "in the wild" mitosis detection remains a significant hurdle. The transition from hotspot-only evaluation to a multi-contextual framework (hotspot, challenging, and random regions) provides a more realistic proxy for clinical reliability.

1. Introduction Quantification of cells that undergo division, represented by mitotic figures in histologic images, is an important task for assessing cancer aggressiveness. As part of routine pathologic evaluation of many tumor types, the number of mitotic figures (MFs) (i.e., the mitotic count (MC)) in a predefined region of the tumor (often defined as 2 to 2.37 𝑚𝑚2 ) with the highest density (hotspot) are manually counted [33, 70, 53, 17, 18, 111]. However, this task has a well-known inter-rater disagreement for classifying MFs against other histologic structures [66, 104] as well as a sampling bias for selection of the tumor region [4], which can impact therapeutic decisions. This measurement variability may be reduced by the use of automatic mitosis detection tools in the diagnostic workflow. Consequently, the mitosis recognition task is well-established in computational pathology. Starting with the MITOS challenge in 2012 [85], the community has conducted several additional challenges in the subsequent years (AMIDA 13 [105], MITOS-ATYPIA 2014 [84], TUPAC16 [103], MIDOG 2021 [5], and MIDOG 2022 [6]), fostering the validation of modern pattern recognition architectures for this task. While early challenges focused on breast cancer and offered limited domain diversity, the two Mitosis Domain Generalization (MIDOG) challenges specifically targeted the most important property of clinically applicable pattern recognition solutions: generalization to diverse and previously unseen sample distributions. In the first (2021) MIDOG challenge edition, images from various whole slide scanners were introduced, testing the generalization to unseen whole-slide image scanners [5]. In the following year’s challenge the scope widened to generalization across different tumor types, introducing a test set of 100 cases across ten different (and partially unseen) tumors from human, canine and feline specimens [5]. The limitation of all previous M. Aubreville et al.: Preprint submitted to Elsevier

challenges is that the detection task was not performed on entire whole slide images (WSIs), but on expert-selected hotspot regions within the tumor, where the highest MC was assumed based on a high cellular density and swift screening for MFs, while excluding challenging areas with an increased occurrence of structures with morphologic overlap to MFs (i.e., imposters). Restricting datasets to hotspot regions limits the applicability of derived models to routine pathology workflows. As shown in prior research [4, 14], a major source of interpathologist variability in the MC is the localization of the most proliferative region of interest (ROI), which introduces sampling bias. Consequently, in traditional MC assessment, hotspots (believed to be most indicative of biological tumor behavior) are often not identified [4], which can negatively affect accurate tumor prognostication and treatment decision. This provides strong motivation for the computerized detection of MFs on entire WSIs, and potentially also across multiple sections of the same tumor [94, 13]. Consequently, these completely automatic workflows are fundamental for a meaningful deployment in a clinical routine. This necessitates validating model performance on whole slide images (or representative regions thereof) to investigate if a reliable application of models to this use case scenario is possible. Motivated by these considerations, the MIDOG 2025 challenge is the first large-scale algorithmic evaluation that extends the scope beyond hotspot regions, evaluating also on ROIs representing all the other regions of the WSI, including particularly challenging ROIs, assumed to be rich in imposters (i.e., structures that are suspected to lead to false positive predictions). Besides detecting MFs across WSIs, another task of interest is the classification of MFs into normal morphologies and atypical mitotic figures (AMFs). An AMF can be histologically identified in cells exhibiting abnormal division due Page 2 of 18

The MIDOG 2025 Challenge

Figure 1: Domains of the test set of the MIDOG 2025 challenge. For each domain, both the abbreviation chosen as well as the full name is given. Shown are four random samples from within the hotspot regions of interest.

to chromosome segregation errors [29, 11]. This asymmetric cell division can lead to daughter cells with an abnormal number of chromosomes (aneuploidy), a relevant mechanism how tumor cells can accumulate mutations required for cancer progression [38, 88]. Aneuploidy leads to copy number alterations of numerous genes at once, and thereby can have a massive effect on cell behavior and function, overall leading to a more aggressive behavior of the tumor. There is growing evidence linking a high frequency and proportion of AMFs to reduced patient survival in breast cancer [77, 59] and other cancer types [47, 48, 16, 69]. A comprehensive, large-scale study by Jahanifar et al. [46] has recently confirmed the prognostic relevance of AMFs rates across a diverse spectrum of tumor types, thereby underscoring the need for further research into this prognostic marker. Motivated by these findings, some recent studies have investigated computerized classification approaches [10, 20]. These studies propose connecting the AMF model as a subsequent stage to the MF model, making AMFs a compelling choice for a second track in the MIDOG 2025 challenge.

Challenge format and task The MIDOG 2025 challenge was organized as a satellite event of the International Conference on Medical Image Computing and Computer Assisted Intervention (MICCAI) 2025, following an open call for participation and peer review of the challenge design [1]. Information about the challenge was made available through the challenge website1 . The challenge comprised two independent tracks, with participants able to submit to either or both.

Track 1: Mitotic figure detection. This track required the

localization of all MFs within a given ROI, across a broad range of tumor types and digitization devices. In contrast to all prior editions of the MIDOG challenge and other mitosis detection challenges, evaluation was not restricted to expertselected hotspot ROIs but extended to randomly sampled tissue regions and expert-selected challenging regions rich in 1 https://midog2025.deepmicroscopy.org

M. Aubreville et al.: Preprint submitted to Elsevier

imposters. Specifically, for each of the 122 test cases (twelve domains, at least ten per domain), at least three ROIs of 2 mm2 each were extracted following the three sampling strategies described below, yielding 365 ROIs in total. Each case corresponded to an individual patient.

ROI sampling strategies. Three complementary ROI types were defined for Track 1 evaluation (see Figure 2):

• Hotspot ROIs: regions of high MF density within the tumor, selected by a board-certified veterinary pathologist following standard clinical practice. At low magnification, several eligible tumor regions were selected based on high cellular density and optimal image and tissue quality (e.g., absence of tissue and scan artifacts, necrosis, or marked inflammation, and avoiding areas with delayed fixation). Subsequently these regions were screened at higher magnification and the regions with the suspected highest MF density were selected. • Random ROIs: regions sampled uniformly at random from the tissue area of the WSI (following a basic thresholding algorithm), subject to a minimum tissue coverage threshold of 80%. Random ROIs are statistically representative of the full tissue composition of the slide and serve as the primary proxy for realistic whole-slide evaluation conditions. • Challenging ROIs: regions deliberately selected by a board-certified veterinary pathologist for their high density of MF imposters – morphological structures that may be mistaken by the models for MFs, including apoptotic and necrotic cells, hyperchromatic nuclei, inflammatory cells, and ink for margin marking. Additionally, ROIs with scan artifacts (e.g., outof-focus regions) and sectioning artifacts (e.g., tissue folds) were included. These regions represent a curated worst case for false positive generation.

Track 2: Atypical mitotic figure classification. The sec-

ond task required the binary classification of image patches Page 3 of 18

The MIDOG 2025 Challenge

sized 128×128 px, centered on confirmed MFs as either normal or atypical. This track was motivated by the emerging prognostic relevance of AMFs and provided the first largescale, multi-domain, multi-scanner dataset of normal and atypical mitotic figures for public use, comprising 11,939 MFs across the seven domains [110] provided by the MIDOG++ [7] dataset.

Submission format and timeline. MIDOG 2025 was a

one time event with fixed submission deadline for which 164 persons registered on the grand-challenge platform. For both tracks, participants were required to submit self-contained Docker containers, which included their complete inference pipeline, ensuring reproducibility and enabling blind evaluation on the withheld test set. Example implementations, including baseline algorithms and instructions for submission, were made available on GitHub2 . The baseline method for Track 1 was provided as FCOS[100] detection model and for Track 2 we provided an EfficientNetV2 [97]-based classifier [9]. The evaluation methods (not including the ground truth of the test set) were provided on GitHub as well.3 A preliminary test set of 20 cases across four tumor domains was made available for automated evaluation by the grand-challenge platform from August 15 2025 to allow validation of the submission pipeline. Submissions of 31 and 35 teams on the preliminary set for Track 1 and Track 2, respectively, were received. The final test set submission window opened on August 30 and closed on September 1, 2025 and teams were allowed to submit once to it. Final submissions were accepted from 18 teams in Track 1 and 21 teams in Track 2. From these, 3 teams were disqualified for failing to supply a preprint or for violating the terms of Track 1.

Participation rules. The organizers provided training data

for both tracks, released under Creative Commons licenses. The use of additional data sets was permitted, given that these were publicly available without conditions to all participants. Access to the test set and test set labels was only provided to the organizers M.A., J. A. and S.B. Researchers belonging to the institutes of the organizers were not allowed to participate to avoid potential conflict of interest. Participants were permitted to publish papers including their official performance on the challenge data set, given proper reference of the challenge, without embargo time. Participants were asked to publish a concise description of their method and results on a preprint server or on a general-purpose open-access repository. Public release of the source code was not mandatory. The first three positions in each track were awarded with monetary prizes. Results were announced at the conference workshop. 2 Track 1: https://github.com/DeepMicroscopy/MIDOG25_T1_reference_ docker and Track 2: https://github.com/DeepMicroscopy/MIDOG25_T2_ reference_docker 3 Track 1: https://github.com/DeepMicroscopy/MIDOG25_T1_evaluation_ docker, Track 2: https://github.com/DeepMicroscopy/MIDOG25_T2_ evaluation_docker

M. Aubreville et al.: Preprint submitted to Elsevier

Figure 2: Region types in track 1 of the MIDOG 2025 challenge. Left panel shows hotspot regions, middle panel shows random regions and right panel shows challenging regions. Green squares indicate PHH3-confirmed mitotic figures, yellow circles indicate false detections by our reference model.

Peer review. Each submitted preprint was reviewed inde-

pendently by at least two expert reviewers, who assessed contributions according to three criteria: technical clarity (whether the approach and training procedure were clearly described), innovation (whether the approach was novel or introduced an innovative adaptation of an existing method), and formal quality (whether the paper was well-structured and clearly written). Of the 31 preprints submitted alongside final test set entries, 27 were accepted for presentation at the MIDOG 2025 workshop, held in conjunction with MICCAI on September 23, 2025. Teams whose papers achieved the highest peer review scores, as well as the top-performing teams on the final leaderboard, were invited to contribute extended versions of their papers to the proceedings, which underwent a further independent full peer review.

2. Material and Evaluation Methods 2.1. Material Training sets: For the task of mitosis detection, a good variety of datasets already existed prior to the challenge (e.g., [3, 7, 15, 19, 84, 103]), including the MIDOG++ dataset [7], which is an extended version of the dataset used in the predecessor challenge in 2022. The organizers therefore decided to not release an additional dataset for MIDOG 2025 for this task. However, given the field of AMF detection was still in its infancy, with only the Ami-Br dataset [20] being available prior to the challenge, we released an additional public training dataset of AMFs, encompassing annotations for all 11,939 MFs of the MIDOG++ dataset [110]. For the annotation process and dataset characteristics of the training datsets we refer to the previous publications. Test sets: For this challenge, we sourced twelve tumor types as the test set, as shown in Table 1. The first ten out of those were previously used for the MIDOG 2022 challenge. Per domain, we used ten tumor cases, corresponding to ten patients, from which one representative WSI was digitized. In addition to the cases for MIDOG 2022, we sourced two human glioblastoma (hGBM) and human lung adenocarcinoma (hLUAD) from a new challenge organization partner (University of Szeged, Hungary). Human glioblastoma was represented with ten cases/WSIs, whereas we incorporated twelve cases/WSIs of human lung adenocarcinoma. Thus,

Page 4 of 18

The MIDOG 2025 Challenge Table 1 Overview of the domains of the final test set. Domain hMel hAC hBlC cMC ccMCT hMen hCoC cHAS fSTS fLym hGBM hLUAD

Tumor Type Human melanoma Human astrocytoma Human bladder carcinoma Canine breast carcinoma Canine cutaneous mast cell tumor Human meningioma Human colon carcinoma Canine hemangiosarcoma Feline soft tissue sarcoma Feline GI lymphoma Human glioblastoma Human lung adenocarcinoma

Tumor Origin Neuroectodermal Neuroectodermal Epithelial Epithelial Mesenchymal Mesenchymal/Neuroectodermal Epithelial Mesenchymal Mesenchymal Mesenchymal Neuroectodermal Epithelial

Scanner / Res. Hamamatsu S360; 0.23 𝜇𝑚/px Hamamatsu S60; 0.22 𝜇𝑚/px 3DHistech P. Scan II; 0.25 𝜇𝑚/px 3DHistech P. Scan II; 0.25 𝜇𝑚/px Hamamatsu S360; 0.23 𝜇𝑚/px Hamamatsu S60; 0.22 𝜇𝑚/px Hamamatsu S360; 0.23 𝜇𝑚/px 3DHistech P. Scan II; 0.25 𝜇𝑚/px 3DHistech P. Scan II; 0.25 𝜇𝑚/px 3DHistech P. Scan II; 0.25 𝜇𝑚/px 3DHistech P1000, 0.12 𝜇𝑚/px 3DHistech P1000, 0.12 𝜇𝑚/px

Origin

MIDOG 2022 [6]

University of Szeged

our dataset comprises WSIs from 122 patients/cases, split across twelve domains. For each of the 122 WSIs, we selected three mutually exclusive ROIs, each one for the sampling category of hotspot region, random region, and challenging region (see Figure 2). In one case the tissue was too small for our random sampler to yield a suitable, non-overlapping and sufficiently tissue-covered region, yielding a total of 365 ROIs for the final test set, split across 122 hotspot regions, 121 random regions and 122 challenging regions. The annotation process for these images is outlined below. Preliminary test set: For validation of the technical function of the participants’ docker containers, we provided access to execution on a four domain, 20 case preliminary test set, as done in previous MIDOG challenges. For this, we reused the preliminary test set of MIDOG 2022, encompassing human breast carcinoma, canine osteosarcoma, human lymphoma, and canine pheochromocytoma [6]. The organizers notified the participants that this dataset is not a good proxy for the final test set, as it only encompasses hotspot ROIs, is considerably smaller, and originates from different tumor types, and that hyperparameter tuning on this dataset in the preliminary evaluation phase is discouraged due to this.

and applied IHC with antibodies against phosphohistone H3 (PHH3). This antibody immunolabels mitotic figures in early pro- to late telophase, facilitates annotation for MFs, and was successfully used in prior research [41, 14, 98, 6]. However, as shown by Ganz et al. [35], the immunopositivity of early prophase MFs, lacking discriminative morphological features for MFs, can lead to an information mismatch between the H&E image and PHH3-labeled images. This has been shown to cause annotation bias, requiring an adaptation of purely IHC-based annotation workflows to avoid annotations of morphologically indistinctive cells [35]. Our MF annotation pipeline was composed of two annotation phases. The first phase was conducted by a boardcertified veterinary pathologist (CAB) that used PHH3 images as annotation support. For this, we registered H&E and PHH3 WSIs using manual selection of four corresponding points in both images. We then calculated the affine transform between both using a least squares optimizer. This workflow was implemented by the challenge organizers in the open source EXACT annotation software [67], allowing for the annotating pathology to blend between the H&E and PHH3 images using a slider and keyboard shortcuts. In this first annotation step, the expert annotated cells according to five categories:

Ethics approval: We received approval by the UMC

1. PHH3-positive, morphology in H&E clearly distinctive of MFs 2. PHH3-positive, morphology in H&E suggestive of a MF 3. PHH3-negative, morphology in H&E suggestive of a MF 4. PHH3-positive, morphology in H&E suggestive of an imposter 5. PHH3-negative and morphology resembling, but distinct from a MF

Utrecht’s ethics board (TCBio 20-776), the Regional and Institutional Research Ethics Committee of the University of Szeged (BM/22651-1/2024) and the ethics board of the medical faculty of FAU Erlangen-Nürnberg (AZ 92_14B, AZ 193_18B). For animal samples taken from the diagnostic archive, no ethics approval is required.

2.2. Annotation For the test set of track 1, we annotated only the new images (i.e., the random and challenging ROIs and the two additional tumor types) while using the previous annotations of MIDOG 2022 for those respective images. For MF annotations we used immunohistochemistry (IHC)-assisted labeling, using the following image preparation steps. For each sample, after successful digitization of the hematoxylin and eosin (H&E) stained slide and quality control, we removed the cover slip from the slide, washed out the H&E stain M. Aubreville et al.: Preprint submitted to Elsevier

Annotations of the first three label classes were considered as MFs and the last two label classes were considered as hard negatives. For the second annotation phase, patches with an approximate size of 30 microns were cropped around those annotations from the H&E image, and blindly shown to a secondary board-certified pathologist (RK), tasked with classifying them into MF or non-MF imposter cells. The Page 5 of 18

The MIDOG 2025 Challenge

hotspot ROIs

random ROIs

challenging ROIs

c)

200

200

150

150

150

100

100

100

50

50

50

0

0

0

fLy hB m hC lC o fS C cH TS AS ccMcMC hL CT UA hMD h el hMAC hG en BM

200

fLy hB m hC lC o fSTC cH S AS ccMcMC hL CT UA hMD h el hMAC hG en BM

b)

fLy hB m hC lC o fSTC cH S AS ccMcMC hL CT UA hMD h el hMAC hG en BM

Mitotic figures per ROI

a)

Figure 3: Violin plot showing the distribution of mitotic figures (MFs) in the MIDOG 2025 test set for track 1, with tumor types being sorted by median hotspot MF count. a) The hotspot regions of interest (ROIs) were selected by a pathologist in the most mitotically active tumor region, whereas the random ROI b) regions were sampled from the tissue area of the whole slide image. The challenging regions (c), selected based on the presence of MF imposters, contain the least MF annotations, followed by the random regions and the hotspot ROIs. For the tumor type abbreviations, please consult Table 1.

inclusion of hard negatives in the second step ensured no assumptions about the distribution/priors were possible by the second expert. In case of agreement, the annotation was included into the final dataset. In case of disagreement, a third expert (TAD, also a board-certified pathologist) was asked to act as a tie breaker, rendering the final label for the annotation. For the 100 hotspot ROIs sourced from the previous MIDOG 2022 challenge, we reused the H&E ground truth from that challenge, which was also obtained through a majority vote by the same pathologists. For the second track, for which a completely new test set was created for this challenge, we used all MF annotations of the hotspot ROIs of the first track. The annotations were done in accordance with the guide for identification of AMFs by Donovan et al. [29]. For all confirmed MFs, cell patches (30 microns) were cropped from the image and independently shown to two pathologists (CAB, VW). In case of disagreement, a third expert (TAD) again rendered the final decision.

2.3. Dataset statistics of the test set The number of MFs varies considerably across tumor types (Fig. 3), with the feline GI lymphoma (fLym) having the highest overall count (N=657) and the human glioblastoma (hGBM) having the lowest overall count (N=36), reflecting the distinct mitotic activity of the different tumor types. For track 2 (AMF classification), the atypical label was most often assigned for MFs in human glioblastoma (hGBM), and least often for canine cutaneous mast cell tumor (ccMCT), as shown in Figure 4. We found Cohen’s 𝜅 of 0.48 between both initial raters for track 1 and of 0.68

54.5%

hGBM hMen fSTS hAC cHAS hMel hLUAD hCoC cMC fLym ccMCT 0

25

50

75

Proportion (%)

100 MF Atypical Normal MF

Figure 4: Class distribution for all tumor types of the track 2 test set. For the tumor type abbreviations, please consult Table 1.

for track 2, indicating a moderate and substantial agreement, respectively.

2.4. Evaluation methods and metrics As in the previous challenges, the MIDOG 2025 challenge used the micro-averaged 𝐹1 score as primary metric. It is defined as 𝐹1 =

M. Aubreville et al.: Preprint submitted to Elsevier

n=66 n=471 n=148 n=580 n=37 n=488 n=209 n=260 n=497 n=213 n=657 n=174

31.6% 24.3% 22.8% 18.9% 18.2% 17.2% 16.5% 13.9% 5.2% 4.7% 4.0%

hBlC

2TP 2TP + FN + FP

(1) Page 6 of 18

The MIDOG 2025 Challenge

where TP, FP, and FN represent the cumulative sum across all ROIs of true positives, false positives, and false negatives, respectively. The matching between ground truth MFs and detections was determined using the Hungarian method [56] with a maximum distance of 7.5 micrometers (approximate size of a nucleus). Multiple detections were considered false positives. This metric was chosen since it combines recall and precision in a single metric. For the MC, an overestimation or underestimation can similarly impact patient prognostication, which is why this combination is sensible. The metric is the most commonly used metric for MF performance assessment, allowing comparison to earlier works. We have calculated 𝐹1 for the overall dataset, as well as for each tumor type, each ROI type, and each combination of tumor type and ROI type across all corresponding images. As an additional secondary metric, we calculated the area under the free-response receiver operating characteristic curve (FROC-AUC). The FROC curve is calculated by varying the detection threshold and determining the recall versus the mean number of false positives per image. To compute the FROC-AUC, sensitivity and mean false positives per image were evaluated across 40 linearly spaced detection thresholds spanning the range of predicted confidence scores. The resulting operating points were interpolated onto 50 uniformly spaced evaluation points in the interval [0, 8] false positives per image using linear interpolation, and the area under the interpolated curve was computed via the trapezoidal rule. The upper limit of 8 false positives per image was chosen following Liu et al. [62], and reflects a clinically motivated operating range for computer-aided detection systems. A higher FROC-AUC indicates better sensitivity across the full range of operating thresholds at a given false positive budget, and, unlike the 𝐹1 score, is independent of a specific detection threshold. As such, it captures a complementary aspect of algorithm performance: where 𝐹1 reflects threshold-calibrated precision and recall at a single operating point, FROC-AUC integrates sensitivity over the full operating range, making it sensitive to morphological recognition capability independently of threshold calibration. Furthermore, as in the predecessor challenge, we calculated average precision (AP) for all participants. AP can be understood as an approximation of the area under the precision recall curve, and was calculated as the mean precision for 101 linearly spaced recall values between 0 and 1. We used the implementation provided by the torchmetrics package [28] (Version 1.9.0). For the second track, the main metric was balanced accuracy (BA). As shown in Figure 4, AMF classification is a highly imbalanced problem, with AMFs being in the minority for almost all tumor domains. The BA is defined as average recall across all classes, and spans values between 0.5 (chance) to 1.0 (perfect classification). We additionally calculated the area under the receiver operating characteristic curve (ROC AUC), a metric commonly used in classification performance assessment.

M. Aubreville et al.: Preprint submitted to Elsevier

Container image

Container inference

Reproducibility check

Code and weights extraction

Model parameter count analysis

Code and paper analysis

ensembling used

TTA used

yes

yes

Creation of modified code segment

Creation of modified code segments

Creation of runtime analysis code segment

ablation code segments

for each code segment

Bind mount

Container inference

Figure 5: Post-challenge analysis workflow. Model parameter count and inference time were established using modification of the submitted docker containers. Additionally, we ablated the containers from using test time augmentation and ensembling to investigate the effects of both methods.

Moreover, we investigated the rank stability of the challenge, as influenced by the region type (overall, hotspot, random, challenging). We further performed correlation analysis on the metrics across the ROI types.

2.5. Post-Challenge Analysis of Algorithmic Approaches A commonality in most challenges is that the submitted approaches differ in a multitude of important algorithmic decisions and components, making it hard to generate insights from a post-challenge analysis. While common success patterns seem to be emerging in pattern recognition challenges, largely based on observation of and comparison of successful and less successful approaches, this is by no means a direct factorial analysis. For instance, the fact that many successful approaches utilize popular techniques for model improvement such as test-time augmentation (TTA) or ensembling [31] is an indication that these techniques are contributing to model performance, but it could also be that it is a popular pattern that has emerged in the community that is only believed to provide true advantages but possibly only provides marginal benefits, while significantly increasing computational requirements. In an attempt to perform a true factorial analysis, we conducted an ablation study on the submitted containers (Figure 5). For this, we asked the participants for their permission for a post-challenge analysis, and for providing the original container. The vast majority of participants Page 7 of 18

The MIDOG 2025 Challenge

agreed, yielding to post-challenge analysis results on 13/14 submissions for Track 1 and 18/20 for Track 2 (see Tables 2, 3). First, the code in each container was manually extracted by overwriting the entrypoint with a shell and providing an external bind mount to the file system of the host, allowing for easy copying between host and container. Subsequently, we performed a manual analysis of the code that was provided by each team, and aligned this against the proceedings paper describing the approach to identify potential discrepancies. In doing so, we identified all model weight files and how they were loaded in the inference pipeline, allowing for a comprehensive assessment of the total parameter count by each approach. Finally, we ran ablation studies on the influence of TTA and ensembling and performed an inference time analysis. We created modified functional blocks and injected them into the original containers by bind mounting them to replace the original files in the docker container. For the inference time analysis, we placed time traces after the loading of the model and before the inference function, and a second one after inference and post-processing was completed. In comparison to the evaluation of the entire call to the docker image, this strips factors such as model loading time and container loading time, which are part of the time assessment in grand-challenge but typically not of interest for algorithmic efficiency comparisons. We ensured fair comparisons by executing all containers on the same system (a GPU workstation with NVIDIA RTX 4090 GPU, AMD Ryzen 9 7900X 12-Core Processor, and 128GB of RAM) and having no other task run in parallel. For this assessment, we randomly selected 50 data items from the test set (50 ROIs for task 1, and 50 cropout stacks of 16 images for task 2), which we ran on all containers. This also allowed for an analysis of utilized GPU memory. In this analysis, we measured the base GPU memory usage (drawn by utilizing a graphical user interface on the workstation) and ensured that no other relevant processes were active on that machine. We then ran the inference and continuously measured peak additional VRAM utilization. Next, we identified in each code base the invocation of (1) TTA and (2) ensembling of models, if used by the approach, and modified the respective code block to yield code fragments that disable both enhancements individually. For the disabling of ensembling, we ran inference using a single model for each of the provided models and aggregated the results using averaging of the final metrics. For teams employing TTA, augmentation was disabled by modifying the inference pipeline to process each image in its original orientation only. We subsequently calculated the difference for all metrics that incurred due to both enhancement techniques. Ablations were feasible for 12 of 17 Track 1 teams and 18 of 21 Track 2 teams who provided consent for re-execution of submitted containers; of these, 6 teams in Track 1 and 8 in Track 2 used ensembling, and 4 and 7 respectively used TTA.

M. Aubreville et al.: Preprint submitted to Elsevier

2.6. Statistical analysis Associations between continuous variables were assessed using Spearman’s rank correlation coefficient (𝜌𝑠 ) for analyses involving ranked or non-normally distributed data, and Pearson’s 𝑟 for analyses of raw continuous scores where normality was assumed. To compare performance across multiple groups, the Kruskal-Wallis test was employed. Ninety-five percent confidence intervals for Pearson’s 𝑟 were computed via Fisher’s 𝑧-transformation. All 𝑝-values are two-sided; a significance threshold of 𝛼 = 0.05 was applied throughout. No correction for multiple comparisons was applied, as all reported analyses were pre-specified based on the challenge evaluation design; correlation analyses involving 12 tumor domains should be interpreted with awareness of limited statistical power. Statistical analyses were performed in Python using scipy.stats [106].

3. Overview of submitted methods 3.1. Track 1 - MF Detection We found a considerable diversity in the approaches submitted for track 1, in particular in the utilized architectures, but also in dataset use (Table 2).

Main Pattern Recognition Approach. Most teams de-

cided to go for a primary object detection framework to detect mitotic figures (10/14). Amongst those, the most prevalently used framework was the you only look once (YOLO) family of models, where various versions were used (YOLOv5, YOLOv8 [102], YOLOv10 [109], YOLOv11 [52] and YOLOv12 [99]) by five teams. Other recent object detection frameworks such as RTMDet [65], DETR [23], and even older architectures like DeepLabV3+ [24], were used as object detectors. Four teams decided to frame the task as semantic segmentation task and used flavors of U-Net [83] architectures such as nnUNET [43], VMUNET [86], or plain vanilla U-Net [73]. Most participants used a single stage detector, only four participants used second stage classification networks.

Datasets. Most of the participants (13/14) used the MI-

DOG++ dataset [7], provided by the authors of the challenge. Additionally, 10 out of 14 used the whole slide datasets of canine cutaneous mast cell tumor (MITOS_WSI_CCMCT [15]) and canine mammary carcinoma (MITOS_WSI_CMC [3]), which was also provided by the challenge authors. Additionally, contestants used samples from the SPIDER dataset, [75], the NCT-CRC-HE-100K dataset [49], the MIDOG 2022 training set [6], a mitosis subtyping dataset by Jahanifar [45], the OMG Octo dataset [91], and the alternative label version of the TUPAC16 dataset [19] (see Table 2).

Ensembling and TTA. Six of the participants used some

form of ensembling, either directly in the primary detection/segmentation stage, or in the secondary classification stage. Of note, in the top three, only a single approach used ensembling. Of those approaches that used ensembling, the Page 8 of 18

The MIDOG 2025 Challenge

Figure 6: a) Rank and b) 𝐹1 scores for all participants across the area types of the track 1 test set. c) shows correlations (Pearson 𝑟) with 95% confidence intervals (CI) between the performance of all participants across the various area types (hots.=hotspot, chall.=challenging, rand.=random). The performance in hotspot and challenging areas had a non-significant correlation, all others were found to be highly significant (** indicates p<0.01, *** indicates (p<0.001).

Table 2 Track 1 (mitotic figure detection) final leaderboard with algorithmic details. 𝐹1 , FROC-AUC (up to 8 FP/image) and average precision (AP) are reported on the final test set. OD = object detection; Seg. = segmentation. Cla. = classification Rank

Username

1 2 3 4 5 6 7 8 9 10 11 12 13 14

wildsquirrel Lv et al. [64] Seg. (disc) KongNet [64] RaphaelBourgade Bourgade et al. [21] OD YOLOv12 (yolo12m) [99] ytopuz53 Topuz et al. [101] OD SDF-YOLO cacfek Fekete et al. [89] OD Yolov10x [109] navyasri.kelam Kelam et al. [51] OD Yolov8 [102] + yolov5mu christian.marzahl Marzahl et al. [68] OD RTMDet_s_1912_1192 [65], tengyoux Xu et al. [114] Seg. + Cla. nnUnetV2 [43], EfficientNet-b3/-b3/v2-s [96] masarno Percannella et al. [79] Seg. VM-UNET [86] piotrgiedziun Giedziun et al. [36] OD RF-DETR [87] SZTU-134 Xiao et al. [113] OD + Cla. YOLO 11x [52], ConvNeXt-Tiny [63] schoe Choe et al. [26] Seg. U-Net [83] vidushiwalia Walia et al. [108] OD DETR [23] krhasan02 Hasan et al. [81] OD DeepLab V3+ [24] npic-ab Broad et al. [22] OD + Cla. FCOS [100], VGG19 [93]

baseline ammeling

Paper

Banerjee et al. [9]

Approach

OD

Architecture

FCOS [100]

Params Training data

FROC-AUC

AP

Review

381M M++,CMC,CCMCT,[45] Yes 20M M++,CMC,CCMCT No 18M M++,CMC,CCMCT Yes 159M M++,CMC,CCMCT, [91], [6] Yes 79M M++,CMC,CCMCT No 36M M++,CMC,CCMCT,[19], [75], [49] No 577M M++,CMC,CCMCT Yes 108M M++ No 32M M++,CMC,CCMCT,[75] No 82M M++,CMC,CCMCT No 49M M++,CMC,CCMCT,[34],[19] No 42M M++,CMC,CCMCT No 1 – M++ No 222M M++ No

TTA Ensemble 3 models No No 5 models 2 models 4 models 3 models (Cla) 3 models No No No No No No

15.0 0.740 14.9 0.722 4.0 0.708 20.0 0.706 3.1 0.704 11.7 0.700 35.3 0.697 19.3 0.694 3.8 0.693 7.2 0.688 11.6 0.660 3.5 0.654 – 1 0.634 5.4 0.538

4.195 4.932 5.347 5.150 5.221 5.303 4.998 4.771 5.187 4.999 4.105 4.784 4.243 3.967

0.608 0.634 0.733 0.660 0.673 0.740 0.670 0.587 0.656 0.623 0.507 0.681 0.487 0.587

10.0 13.0 13.5 11.0 10.0 14.5 11.0 13.0 10.0 13.5 13.5 11.5 9.7 9.7

95M M++

No

6.4 0.688

5.139

0.732

No

Time (s)

𝐹1

Training data: M++=MIDOG++ [7], CMC=MITOS_WSI_CMC [3], CCMCT=MITOS_WSI_CCMCT [15], Params: Aggregated model parameter count. Time: mean inference time per ROI on RTX 4090. Review: mean peer review score (max. 15). Top-3 results in bold. 1 Model container was not provided.

mean ensemble size was 3.333. TTA was used by only four of the participants, including three out of the top five ranked.

Peer Review Scores. The participants scored between 9.7

and 14.5 out of 15 in our peer review score. We did not find a significant rank correlation between the peer review score and the rank in the challenge (𝑟ℎ𝑜𝑠 = 0.234, 𝑝 = 0.421).

or ResNet [40]. For LoRA, seven participants chose recent pathology foundation models, such as UNI [25], Virchow [107, 117] or HIBOU [74]. The winning team (Balezo et al. [8]) chose to low-rank adapt the general-purpose foundation model DINOv3 [92]. One submission (Qi et al. [80]) used fine-tuning of a general-purpose vision transformer model pre-trained on ImageNet (EfficientVit [61]).

3.2. Track 2 - AMF Classification

Ensembling. Another frequently employed strategy was

The key differences between submissions for Track 2 were related to model training and the ensembling strategy. (see Table 3):

Model Training. Almost no participants chose to train a

model from scratch: 7 out of 20 participants chose to use lowrank adaptation (LoRA) as parameter-efficient adaptation of foundation models and 10 out of 20 participants chose to use fine-tuning of other pre-existing models. Two teams used distillation as primary learning paradigm. As feature extractors, most participants used architectures employing convolutional layers, such as EfficientNet [96], ConvNeXt [63], M. Aubreville et al.: Preprint submitted to Elsevier

ensembling of heterogeneous models: 12 out of 20 teams used multiple models and ensembled their outputs for the prediction. The mean size of the ensemble was 4.667. In the most notable case of ensembling, Ochi et al. [76], used a combination of three foundation models (Virchow [107], Virchow2 [117], UNI [25] and an ImageNet-finetuned ConvNeXt V2 [112], fused by the autogluon framework [32] using a combination of 438 models. Of note, the winning approach was the only in the top five that did not use any model ensembling.

Page 9 of 18

The MIDOG 2025 Challenge Table 3 Track 2 (atypical mitotic figure classification) final leaderboard with algorithmic details. Balanced accuracy and ROC AUC are reported on the final test set. Rank

Username

1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20

guillaume.balezo Balezo et al. [8] LoRA DINOv3-H [92] 842M M25,AMi-Br,OMG Yes nasires Nasir et al. [72] LoRA Virchow2 [117] 631M M25,AMi-Br,OMG, [45] Yes yohsuke.yamagishi Yamagishi et al. [115] FT ConvNext-v2-base [112] 438M M25 No qixuan1234 Qi et al. [80] FT EfficientViT-L2[61] 320M1,2 M25,Ami-Br Yes masarno Percannella et al. [78] FT Custom Hover-Net[39] 447M M25 Yes zerostarcraft Meng et al. [71] VPT UNI2-h [25] 682M M25 Yes krausara Krauss et al. [55] FT ConvNeXtBase [63] 793M M25,AMi-Br,OMG No lrc9859 Ramchandani et al. [82] LoRA Virchow-Base [107] 647M M25,AMi-Br,OMG No 3 be_yuan Ochi et al. [76] LoRa/FT UNI, Virchow, Virchow2, ConvNext-V2 1658M M25,AMi-Br,OMG,[45] No piotrgiedziun Giedziun et al. [37] LoRA Virchow2-Base [117] 708M M25,AMi-Br,[44] Yes cacfek Shen et al. [89] FT ConvNext [63] 139M M25,AMi-Br,OMG,[60],[44],[60] No schoe Choe et al. [26] Distillation UNet[83] 173M M25,AMi-Br,OMG No saipradeepvg Kotte et al. [54] FT ResNet50 [40] 25M M25 No navyasri.kelam Kelam et al. [50] LoRA UNI [25] 307M M25,AMi-Br No Leire Benito-Del-Valle et al. [12] FT ConvNext-small [63] 247M M25 Yes Kaustubh_Atey Atey et al. [2] Distillation DenseNet-121 [42] 7M M25,AMi-Br,OMG No mirazzak Dukre et al. [30] FT DenseNet-121 [42] 7M M25 No mlafarge Lafarge et al. [58] –4 Custom ResNet-41 [57] with P4M GroupConv [27] 2M M25,AMi-Br No chillice Xu et al. [114] FT EfficientNet-b3/-b5 [96], and InceptionV3 [95] 61M – Yes sercan.cayir Çayir et al. [118] LDA / CatBoost HIBOU-L [74] and Barlow-Twins [116] – 1 M25 No

Baseline maubreville

Paper

Banerjee et al. [9]

Adaptation

FT

Architecture

EfficientNet-2M [97]

Params Training data

53M M25

TTA Ensemble Time (s) Bal. Acc. ROC AUC Review

No

No 3 models 5 models 5 models 5 models No 3 models 3 models 4 models 13 models 5 models No No No 5 models No No No 2 models 3 models

0.421 11.450 1.247 –1 0.686 2.879 22.534 0.986 0.815 4.859 0.232 0.189 0.202 0.236 0.299 0.294 0.258 0.471 2.330 –1

0.908 0.901 0.900 0.897 0.897 0.896 0.889 0.884 0.884 0.882 0.878 0.872 0.869 0.868 0.868 0.859 0.850 0.829 0.824 0.671

0.970 0.967 0.971 0.962 0.962 0.962 0.960 0.958 0.946 0.202 0.953 0.948 0.944 0.944 0.964 0.957 0.927 0.905 0.955 0.759

13.0 10.0 13.0 9.5 12.5 14.5 15.0 11.5 11.5 12.0 11.0 13.5 10.5 8.0 11.0 12.0 11.0 11.0 11.0 12.0

No

0.265

0.827

0.907

Approach: FT=fine-tuning of ImageNet-pretrained model, LoRA: Low-rank adaptation, VPT: visual prompt tuning. Training data: M25=MIDOG 2025 atypical set [110], AMi-Br=AMi-Br atypical dataset [20], OMG=OMG-Octo atypical dataset [90]. Params: single model parameter count. Time: mean inference time per ROI on RTX 4090. Review: mean peer review score (max. 15). Top-3 results in bold. 1 Participant did not provide docker container. 2 Value taken from literature. 3 Model fusion not counted. 4 No fine-tuning or adaptation was mentioned in the proceedings paper.

Figure 7: Inference time vs. main metric for track 1 (a) and track 2 (b) of the MIDOG 2025 challenge. Inference time was determined on a Linux workstation with an NVIDIA RTX 4090. Bubble size indicates max. VRAM usage of container, averaged over 50 cases.

Datasets. All approaches used the challenge’s training

4.1. Track 1 - Detection

set [110]. The majority (12/20) also used the smaller breast cancer atypical mitosis dataset AMi-Br [20]. During the challenge, the challenge participants Shen et al. made available an additional dataset, OMG-Octo Atypical, containing 1,378 atypical mitotic figures [90]. This dataset was reported to have been used by 8 out of 20 participants. Besides this, additional datasets used for training include an atypical subtyping datset by Jahanifar et al. [45], stMIDOG/LUNGMITO [44], and GBM-TCGA[60].

The highest result in the main challenge metric in Track 1 (𝐹1 score) was achieved by Lv et al. [64, 73], reaching an overall 𝐹1 score of 0.740 (see Table 2). The runner-up for this track were Bourgade et al. [21], reaching an overall 𝐹1 score of 0.722. In the secondary challenge metric, FROC AUC, the leading team were Topuz et al. [101], reaching a score of 5.347. The runner-up for FROC AUC were Marzahl et al. [68] with a score of 5.303.

Peer Review Scores. The participants reached an averaged

the different ROI types (hotspot, random, challenging, and overall) in rank across the field of the challenge participants. As shown in Figure 6a), the rank strongly depended on the ROI type. As Figure 6b) shows, the variance in 𝐹1 score is small in the hotspot regions. In consequence, even small changes in 𝐹1 in hotspot regions can lead to a change in rank. We furthermore found significant correlations between the overall 𝐹1 score and all individual ROI types (see Fig. 6c).

peer review score across reviewers of 8.0 to 15.0. We found no significant rank correlation between rank in the challenge and peer review scores (𝜌𝑆 = −0.270, 𝑝 = 0.250).

4. Results The main results are included in Tables 2 and 3, and will be discussed in the following. M. Aubreville et al.: Preprint submitted to Elsevier

Score Stability. We evaluated the rank stability across

Page 10 of 18

The MIDOG 2025 Challenge

However, with r=0.36 (CI95=[-0.206, +0.750], p = 0.201) there was no significant correlation between hotspot 𝐹1 score and the same metric in challenging areas, highlighting that good hotspot predictions models are not necessarily good predictors for challenging areas.

Tumor Type Dependency. We found a considerable de-

pendency on the tumor type, with 𝐹1 scores ranging from 0.444 (human glioblastoma) to 0.808 (canine hemangiosarcoma) in median for all participants (see supplementary Figure S1). For the challenging areas, this difference was even more pronounced, with the minimum achieved likewise on human glioblastoma (0.010) and the maximum achieved on feline lymphoma (0.634). Using a Kruskal-Wallis-Test, we evaluated the significance of tumor type, pooled across all area types, and found it to be significant (p < 0.0001). We did not conduct individual pairwise tests.

ROI Type Dependency. We investigated how strongly

results varied across the different area types of our challenge (overall, hotspot, random, challenging). In the main challenge metric for track 1 (𝐹1 score), we found a mean overall 𝐹1 score of 0.685. In hotspot areas, this was considerably elevated, yielding a mean 𝐹1 of 0.735 across participants (see also Fig. 6b). In random areas, the performance dropped to a mean value across participants of 0.638. In challenging areas, this effect was even more pronounced, yielding a mean 𝐹1 score of only 0.479 (difference to hotspot areas: 0.256). The drop in 𝐹1 score can largely be attributed to a drop in precision (see supplementary Figure S2). Compared to the precision value of all ROIs combined (0.704), hotspot ROIs had the highest score of 0.805 and performance markedly decreased in random (0.614) and challenging areas (0.400), indicating an increased incidence of false positives in the latter. The difference in precision between hotspot and challenging areas was 0.405, translating into an increased false detection rate by approximately 208%. The detrimental effect of a different area type was less pronounced in the recall value: We found an overall mean recall value of 0.681 across participants, and recall values of 0.682, 0.685, and 0.662 for the hotspot, random, and challenging areas, respectively. For a list of per-tumor-type and per-team precision and recall values, as well as the 𝐹1 and FROC AUC values, please consult supplementary Figures S3 and S4.

4.2. Track 2 - Classification The highest performance in the main challenge metric (BA) was achieved by Balezo et al. [8] with a value of 0.908. In the secondary ROC AUC metric, Yamagishi et al. [115] achieved a marginally better result (0.971 vs. 0.970 of Balezo et al.). The runner-up in BA were Nasir et al. [72] with a balanced accuracy of 0.901, only narrowly beating the third-ranked Yamagishi et al. [115], who achieved a score of 0.900. Overall, the top three teams demonstrated remarkably strong performance, with a difference of less than one percentage point in BA separating the first from the third place.

M. Aubreville et al.: Preprint submitted to Elsevier

Tumor Type Dependency. As shown in the supplemen-

tary Figure S5, there was a considerable variance across the domains (tumor types) of the test set. The difference was assessed to be significant using a Kruskal-Wallis test for both BA (H=98.757, p=0.0000) as well as ROC AUC (H=78.752, p=0.0000). The lowest median BA was achieved for feline lymphoma (fLym) and the highest median BA was achieved for human astrocytoma (hAC). As shown in supplementary Figure S6, the classification of AMFs in feline lymphoma was mostly restricted by a low recall.

4.3. Reproducibility analysis We re-evaluated all containers provided by the participants to the organizers on the test set and found the containers to reproduce the grand-challenge official challenge results with sufficient precision (𝐹1 : deviation mean absolute 3.56 ⋅ 10−5 , max absolute: 0.18 percentage points, balanced accuracy: mean absolute 2.24 ⋅ 10−5 , max absolute 1.13 percentage points).

4.4. Inference time and memory analysis Participants had a five minute time limit on grandchallenge for each inference job. However, inference jobs include starting of the docker container, loading of the model, model inference and post-processing, and can be subject to secondary load on the system running the inference task. The actual time available for inference was thus much shorter, and, in track 1, was furthermore subject to image size and post-processing time additionally scaled with detection results due to non-maximum suppression having typically an (𝑛2 ) complexity. video random access memory (VRAM) was limited to 16 GB in the grand-challenge environment and to 24 GB in the post-challenge evaluation. Inference time varied strongly across the field of participants (see Fig. 7). In Track 1, the approach using the highest average inference time per ROI (35.386s) was by Xu et al. [114], which used a two-stage approach based upon nnUnet [43], a three model second stage for classification and a final ensembling mechanism by a random forest, combined with TTA. The team with the lowest inference time for track were Kelam et al. [51], who only used 3.515s of inference time on average and still achieved the fifth place on the leader board. We found no significant rank correlation between inference time and 𝐹1 score (𝜌𝑆 = 0.367, 𝑝 = 0.196) or peak VRAM usage and 𝐹1 score (𝜌𝑠 = −0.253, 𝑝 = 0.383). For track 2, we found a similarly high variance of inference time and VRAM usage, as shown in Figure 7b. On average, the GPU VRAM use was much lower in track 2 (2.424 GB) compared to track 1 (5.151 GB). With 22.534s, Krauss et al. [55] had the longest inference time per batch of 16 images, however, without using the GPU (Fig. 7). The shortest time was achieved by Choe et al. [26] with 0.189s. For track 2, we found a significant rank correlation between peak VRAM usage and balanced accuracy (𝜌𝑠 = 0.616, 𝑝 = 𝑝 = 0.005) but not between inference time and balanced accuracy (𝜌𝑠 = 0.411, 𝑝 = 0.080).

Page 11 of 18

The MIDOG 2025 Challenge

5. Discussion

Figure 8: Ablation study for test-time augmentation (TTA) and ensembling. Shown are only participants that utilized either TTA or ensembling in track one (top) and track two (bottom).

4.5. Ablation studies Our component analysis of TTA or ensembling reveals a mixed benefit of both methods for both tracks (Figure 8). By using ensembling, we found the participants of Track 1 had a mean increase of the 𝐹1 metric on the entire test dataset of 1.549 percentage points, and a median increase of 0.886 percentage points. In the hotspot ROIs alone, we found an increase of 1.404 percentage points in mean and 1.290 percentage points in median. In the random areas, we found a mean increase of 1.661 percentage points (1.198 in median). In the challenging areas, the effect was even more pronounced, leading to an increase of, in mean and median, 2.212 and 1.691 percentage points, respectively. We can thus summarize that the effect of ensembling grew with the data distribution moving further outside of the training distribution. TTA had a less pronounced effect on the main challenge metric, with only 0.320 percentage points improvement in mean and almost no change in median (0.042 percentage points), with only marginal changes over the ROI types. In the AMF classification track, the teams benefited from both techniques in a similar way, yielding a median and mean increase of 1.230 and 1.299 percentage points by ensembling. By TTA, the results in the second track increased on average by 0.299 percentage points and 0.472 percentage points in median. For four models in Track 2, ablation was not possible: For two, the authors did not provide docker containers for post-challenge analysis (see Table 3). In one instance (Ochi et al. [76]), a highly complex fusion model of very heterogeneous other models was used, making an ensembling ablation of this model questionable. In another instance (Nasir et al. [73]), the model ensemble was fused in one singular model, prohibiting ablation.

M. Aubreville et al.: Preprint submitted to Elsevier

The MIDOG 2025 challenge was the first challenge to ever incorporate testing outside of hotspot regions, and – with twelve independent tumor domains – provides the largest and most comprehensive test set to date. Even though open training data with annotations in entire WSIs exists and was used by almost all participants, we still see a dramatic loss of performance for random and challenging areas. In the hotspot regions, however, we found overall high performance with only minor variance across participants (see Fig. 6). Our results hint at a loss of generalization due to a missing data variance in existing datasets. This calls for the curation of datasets with a higher domain variance, particularly with respect to region selection. For the deployment of current algorithms, this means that algorithms should not be deployed for WSI usage, unless specifically validated on a wider area selection or on entire WSIs. In both tracks, the challenge has revealed architectural trends in the highest ranking teams: While the YOLO model family was often chosen for the object detection track, LoRA-adaptation of foundation models was a clear trend in the top ranks of the atypical classification track. Many of the successful teams employed common machine learning tricks such as ensembling and TTA. Ensembling provided consistent benefits across both tracks, however, it is worth noting that the top ranked participant of track 2 (Balezo et al.) did not make use of it. Furthermore, out of the top three methods in track 1, only one used ensembling, suggesting that while ensembling can improve the performance, it is not a prerequisite for achieving top performance. In contrast, TTA had a much smaller impact on performance and was in some cases even detrimental. Still, TTA was used by three out of the top five teams for track 1 and four out of the top five in track 2. While TTA is thus a well-known technique and easy to implement, it does add significant computational overhead, and was not effective for this challenge. While broader investigations beyond this challenge are necessary to confirm our observations, this result might serve as an indicator to discourage the use of TTA in pathology challenges with current models. We found no significant correlation between paper quality, as assessed through peer review, and challenge performance for either track. This points towards an important dilemma in challenge workshops: While paper scoring, which is based on innovative methods and an intriguing presentation, is typically used to identify contributions for talks, this strategy might omit submissions with highest model performance. A further consideration concerns not only the availability of training data, but the spatial context it provides. Two of the additional datasets used by participants: OMG-Octo [91] and the Jahanifar mitosis subtyping dataset [45], provide only 64×64 px patches centered on individual mitotic figures, a constraint partly imposed by clinical data-sharing requirements. While such patches are well suited for the classification task of Track 2, they offer limited spatial context Page 12 of 18

The MIDOG 2025 Challenge

for object detection models (e.g. YOLO), where surrounding tissue architecture is a key cue for distinguishing mitotic figures. This may partly explain why teams (see Table 2) which utilized training data containing broader spatial context performed comparatively well in challenging regions. More broadly, our results suggest that future dataset curation efforts should consider not only diversity in tumor type and scanner, but also the spatial extent of the provided annotations, as this directly affects which architectural classes can benefit from a given resource. We also want to highlight limitations of our work. The annotation process in the two new tumor types (hotspot areas) as well as for all tumor types in the random and challenging areas (i.e., all ROIs that have not been previously utilized in MIDOG 2022) deviated slightly from the previous MIDOG 2022 workflow. While both annotations methods were supported by PHH3 IHC to identify MFs in the first annotation step, there was a difference in the label classes, which could lead to label drift. However, we estimate the impact of this deviation to be low, given that both annotation workflows conducted final decision through a majority vote by three pathologists. In the first track of our challenge, we used AP as tertiary metric for the threshold-independent algorithms assessment. Although AP is well established in object detection, it is sensitive to the initial confidence threshold applied before non-maximum suppression (NMS). Since NMS has a computational complexity of (𝑛2 ) and execution time was limited, participants were incentivized to use higher cutoff values to reduce processing time. Conversely, lower cutoffs generate a broader range of precision–recall pairs, extending the precision–recall curve and artificially increasingAP as previously observed in the MIDOG 2022 challenge [6]. However, this increase does not necessarily reflect improved class discrimination, representing a major limitation of this metric. In contrast, FROC-AUC is constrained by a predefined false-positive rate per image and is therefore less susceptible to inflation from low confidence thresholds. The MIDOG 2025 challenge was the most successful challenge in terms of participation, attracting more submissions than previous iterations. Our results indicate that, in particular, mitotic figure detection performance in hotspot regions is a task that is carried out with high robustness across unseen domains, which is a precondition for clinical use. The availability of new algorithms for atypical classification, created in the context of this challenge, allows for the deeper investigation of the pathological role of AMFs. In summary, the MIDOG 2025 challenge has substantially advanced the field of computational pathology by establishing a new benchmark for mitotic figure detection and atypical mitosis classification across an unprecedented breadth of tumor domains and tissue regions. The challenge highlights both the maturity of current approaches – particularly for hotspot-based detection – and the clear gaps that remain when algorithms are applied to the broader, more heterogeneous landscape of whole slide images. The architectural trends identified here, including the dominance M. Aubreville et al.: Preprint submitted to Elsevier

of YOLO-based detectors and LoRA-adapted foundation models, provide a valuable snapshot of the current state of the art and a foundation for future methodological development. Moving forward, we advocate for the curation of more diverse training datasets, rigorous validation beyond hotspot regions, and the establishment of evaluation frameworks that reward generalization as much as peak performance. We hope that the datasets and baselines released alongside this challenge will serve as a lasting resource for the community, and that the findings presented here will inform the responsible translation of mitosis detection algorithms into clinical practice.

Acknowledgements M.A. and S.B. acknowledge funding by the Deutsche Forschungsgemeinschaft (DFG, project number: 520330054), C.A.B. and V.W. acknowledge funding by the Austrian Research Fund (FWF, project number: I 6555). J.A. acknowledges support by the Bavarian State Ministry of Science and the Arts (project Fokus-TML). K.B. acknowledges funding by the DFG, project number 460333672 CRC1540 EBM. The MIDOG challenge received financial support from MIRA vision microscopy GmbH, Göppingen, Germany and Single-Cell Technologies Ltd, Szeged, Hungary to cover the platform costs. We furthermore acknowledge support by the MICCAI special interest group (SIG) on computational pathology, who donated monetary prizes for the winners of the challenge.

Organization Team The MIDOG 2025 challenge was organized by (in alphabetic order): Jonas Ammeling, Marc Aubreville, Sweta Banerjee, Christof A. Bertram, Katharina Breininger, Dominik Hirling, Peter Horvath, Nikolas Stathonikos, and Mitko Veta.

A. Supplementary Figures CRediT authorship contribution statement Marc Aubreville: Conceptualization, Data curation, Methodology, Writing – original draft, Writing – review & editing, Project Administration, Funding acquisition. Jonas Ammeling: Conceptualization, Data curation, Methodology, Writing – review & editing. Sweta Banerjee: Conceptualization, Data curation, Methodology, Writing – review & editing. Viktoria Weiss: Data curation, Writing – review & editing. Taryn A. Donovan: Data curation, Writing – review & editing. Robert Klopfleisch: Data curation, Writing – review & editing. Jiaqi Lv: Methodology. Shan E Ahmed Raza: Methodology. Raphaël Bourgade: Methodology. Thomas Walter: Methodology, Writing – review & editing. Yasemin Topuz: Methodology. Songül Varlı: Methodology. Charles-Antoine Collins-Fekete: Methodology,Writing – review & editing. Zhuoyan Shen: Methodology. Navya Sri Kelam: Methodology. Nitin Singhal: Methodology. Christian Marzahl: Methodology. Brian Napora: Methodology. Tengyou Xu: Methodology. Hongyan Page 13 of 18

The MIDOG 2025 Challenge

Gu: Methodology. Mario Vento: Methodology. Gennaro Percannella: Methodology. Norbert Ropiak: Methodology. Izabela Wasiak: Methodology. Jie Xiao: Methodology. Shaojun Liu: Methodology. Seungho Choe: Methodology. April Khademi: Methodology. Vidushi Walia: Methodology. Sujatha Kotte: Methodology. Andrew Broad: Methodology. Alex Wright: Methodology. Guillaume Balezo: Methodology. Esha Sadia Nasir: Methodology. Mostafa Jahanifar: Methodology. Yosuke Yamagishi: Methodology. Shouhei Hanaoka: Methodology. Francesco Tortorella: Methodology. Biwen Meng: Methodology. Jingxin Liu: Methodology. Sara Krauss: Methodology. Daniel Hieber: Methodology. Lavish Ramchandani: Methodology. Dev Kumar Das: Methodology. Mieko Ochi: Methodology. Yuan Bae: Methodology. Piotr Giedziun: Methodology. Mateusz Maniewski: Methodology. Vangala Govindakrishnan Saipradeep: Methodology. Naveen Sivadasan: Methodology. Leire Benito-Del-Valle: Methodology. Adrian Galdran: Methodology. Kaustubh Atey: Methodology. Sameer Anand Jha: Methodology. Adinath Dukre: Methodology. Imran Razzak: Methodology. Maxime W. Lafarge: Methodology. Viktor H. Koelzer: Methodology, Writing – review & editing. Nils Porsche: Visualization, Writing – review & editing. Nikolas Stathonikos: Conceptualization, Data curation, Methodology, Writing – review & editing, Project Administration. Mitko Veta: Conceptualization, Writing – review & editing, Project Administration. Dominik Hirling: Conceptualization, Writing – review & editing, Data curation, Project Administration. Zsanett Zsófia Iván: Data curation. Peter Horvath: Conceptualization, Data curation, Writing – review & editing, Project Administration, Funding acquisition. Katharina Breininger: Conceptualization, Writing – review & editing, Project Administration. Christof A. Bertram: Conceptualization, Data curation, Writing – original draft, Writing – review & editing, Project Administration, Funding acquisition.

References [1] Ammeling, J., Aubreville, M., Banerjee, S., Bertram, C.A., Breininger, K., Hirling, D., Horvath, P., Stathonikos, N., Veta, M., 2025. Mitosis domain generalization challenge 2025. URL: https: //doi.org/10.5281/zenodo.15077361, doi:10.5281/zenodo.15077361. [2] Atey, K., Jha, S.A., Bala, G., Sethi, A., 2026. Mix, Align, Distil: Reliable Cross-Domain Atypical Mitosis Classification, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 137–143. [3] Aubreville, M., Bertram, C.A., Donovan, T.A., Marzahl, C., Maier, A., Klopfleisch, R., 2020a. A completely annotated whole slide image dataset of canine breast cancer to aid human breast cancer research. Scientific data 7:417, 1–10. doi:10.1038/s41597-020-00756-z. [4] Aubreville, M., Bertram, C.A., Marzahl, C., Gurtner, C., Dettwiler, M., Schmidt, A., Bartenschlager, F., Merz, S., Fragoso, M., Kershaw, O., et al., 2020b. Deep learning algorithms out-perform veterinary pathologists in detecting the mitotically most active tumor region. Scientific Reports 10:16447, 1–11. doi:10.1038/s41598-020-73246-2. [5] Aubreville, M., Stathonikos, N., Bertram, C.A., Klopfleisch, R., Ter Hoeve, N., Ciompi, F., Wilm, F., Marzahl, C., Donovan, T.A., Maier, A., et al., 2023a. Mitosis domain generalization in histopathology images—the MIDOG challenge. Medical Image Analysis 84, 102699.

M. Aubreville et al.: Preprint submitted to Elsevier

[6] Aubreville, M., Stathonikos, N., Donovan, T.A., Klopfleisch, R., Ammeling, J., Ganz, J., Wilm, F., Veta, M., Jabari, S., Eckstein, M., et al., 2024. Domain generalization across tumor types, laboratories, and species—insights from the 2022 edition of the mitosis domain generalization challenge. Medical Image Analysis 94, 103155. [7] Aubreville, M., Wilm, F., Stathonikos, N., Breininger, K., Donovan, T.A., Jabari, S., Veta, M., Ganz, J., Ammeling, J., Van Diest, P.J., et al., 2023b. A comprehensive multi-domain dataset for mitotic figure detection. Scientific data 10, 484. [8] Balezo, G., Bourgade, R., Feki, H., Monnier, L., Blons, M., Blondel, A., Decencière, E., Planas, A.P., Walter, T., 2026. Efficient FineTuning of DINOv3 Pretrained on Natural Images for Atypical Mitotic Figure Classification, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 15–25. [9] Banerjee, S., Ammeling, J., Weiss, V., Donovan, T.A., Klopfleisch, R., Bertram, C.A., Breininger, K., Aubreville, M., 2026a. Mitosis Domain Generalization (MIDOG) Challenge 2025 Baselines, in: Aubreville, M., Bertram, C.A. (Eds.), Mitosis Domain Generalization (MIDOG) Challenge 2025 Baselines, Springer. pp. 1–14. [10] Banerjee, S., Weiss, V., Donovan, T.A., Fick, R.H., Conrad, T., Ammeling, J., Porsche, N., Klopfleisch, R., Kaltenecker, C., Breininger, K., et al., 2026b. Benchmarking deep learning and vision foundation models for atypical vs. normal mitosis classification with crossdataset evaluation. Machine Learning for Biomedical Imaging 2026, 115–125. doi:10.59275/j.melba.2026-6c1g. [11] Banerjee, S., Weiss, V., Donovan, T.A., Fick, R.H., Conrad, T., Ammeling, J., Porsche, N., Klopfleisch, R., Kaltenecker, C.C., Breininger, K., Aubreville, M., Bertram, C.A., 2026c. Benchmarking deep learning and vision foundation models for atypical vs. normal mitosis classification with cross-dataset evaluation. Machine Learning for Biomedical Imaging 2026, 115– 125. URL: https://melba-journal.org/2026:006, doi:https://doi. org/10.59275/j.melba.2026-6c1g. [12] Benito-Del-Valle, L., Moreno-Sánchez, P.A., Eguskiza, I., Vitoria, I., Picón, A., López-Saratxaga, C., Galdran, A., 2026. Is Synthetic Image Augmentation Useful for Imbalanced Classification Problems? Case-Study on the MIDOG2025 Atypical Cell Detection Competition, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 144–152. [13] Bertram, C.A., Aubreville, M., Donovan, T.A., Bartel, A., Wilm, F., Marzahl, C., Assenmacher, C.A., Becker, K., Bennett, M., Corner, S., et al., 2022. Computer-assisted mitotic count using a deep learning–based algorithm improves interobserver reproducibility and accuracy. Veterinary pathology 59, 211–226. doi:10.1177/ 03009858211067478. [14] Bertram, C.A., Aubreville, M., Gurtner, C., Bartel, A., Corner, S.M., Dettwiler, M., Kershaw, O., Noland, E.L., Schmidt, A., Sledge, D.G., et al., 2020a. Computerized calculation of mitotic count distribution in canine cutaneous mast cell tumor sections: mitotic count is area dependent. Veterinary pathology 57, 214–226. doi:10. 1177/0300985819890686. [15] Bertram, C.A., Aubreville, M., Marzahl, C., Maier, A., Klopfleisch, R., 2019. A large-scale dataset for mitotic figure assessment on whole slide images of canine cutaneous mast cell tumor. Scientific data 6, 1–9. doi:10.1038/s41597-019-0290-4. [16] Bertram, C.A., Bartel, A., Donovan, T.A., Kiupel, M., 2023. Atypical mitotic figures are prognostically meaningful for canine cutaneous mast cell tumors. Veterinary Sciences 11, 5. [17] Bertram, C.A., Donovan, T.A., Bartel, A., 2024a. Mitotic activity: A systematic literature review of the assessment methodology and prognostic value in canine tumors. Veterinary pathology 61, 752– 764. doi:10.1177/03009858241239565. [18] Bertram, C.A., Donovan, T.A., Bartel, A., 2024b. Mitotic activity: A systematic literature review of the assessment methodology and prognostic value in feline tumors. Veterinary Pathology 61, 743– 751. doi:10.1177/03009858241239566.

Page 14 of 18

The MIDOG 2025 Challenge [19] Bertram, C.A., Veta, M., Marzahl, C., Stathonikos, N., Maier, A., Klopfleisch, R., Aubreville, M., 2020b. Are pathologist-defined labels reproducible? comparison of the tupac16 mitotic figure dataset with an alternative set of labels, in: Interpretable and AnnotationEfficient Learning for Medical Image Computing. Springer, pp. 204– 213. doi:10.1007/978-3-030-61166-8_22. [20] Bertram, C.A., Weiss, V., Donovan, T.A., Banerjee, S., Conrad, T., Ammeling, J., Klopfleisch, R., Kaltenecker, C., Aubreville, M., 2025. Histologic dataset of normal and atypical mitotic figures on human breast cancer (ami-br), in: BVM Workshop, Springer. pp. 113–118. [21] Bourgade, R., Balezo, G., Monier, L., Feki, H., Blons, M., Blondel, A., Loussouarn, D., Vincent-Salomon, A., Walter, T., 2026. Robust Pan-Cancer Mitotic Figure Detection with YOLOv12, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 26–36. [22] Broad, A., Keighley, J., Godson, L., Wright, A., 2026. MIDOG 2025: Mitotic Figure Detection with Attention-Guided False Positive Correction, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 87–90. [23] Carion, N., Massa, F., Synnaeve, G., Usunier, N., Kirillov, A., Zagoruyko, S., 2020. End-to-end object detection with transformers, in: European conference on computer vision, Springer. pp. 213–229. doi:10.1007/978-3-030-58452-8_13. [24] Chen, L.C., Zhu, Y., Papandreou, G., Schroff, F., Adam, H., 2018. Encoder-decoder with atrous separable convolution for semantic image segmentation, in: Proceedings of the European conference on computer vision (ECCV), pp. 801–818. [25] Chen, R.J., Ding, T., Lu, M.Y., Williamson, D.F., Jaume, G., Song, A.H., Chen, B., Zhang, A., Shao, D., Shaban, M., et al., 2024. Towards a general-purpose foundation model for computational pathology. Nature medicine 30, 850–862. [26] Choe, S., Qin, X., Shafique, A., Dy, A., Done, S., Androutsos, D., Khademi, A., 2026. Teacher-Student Model for Detecting and Classifying Mitosis in the MIDOG 2025 Challenge, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 207–214. [27] Cohen, T., Welling, M., 2016. Group equivariant convolutional networks, in: Proceedings of the International Conference on Machine Learning (ICML), pp. 2990–2999. [28] Detlefsen, N.S., Borovec, J., Schock, J., Jha, A.H., Koker, T., Di Liello, L., Stancl, D., Quan, C., Grechkin, M., Falcon, W., 2022. Torchmetrics - measuring reproducibility in pytorch. Journal of Open Source Software 7, 4101. URL: https://doi.org/10.21105/ joss.04101, doi:10.21105/joss.04101. [29] Donovan, T.A., Moore, F.M., Bertram, C.A., Luong, R., Bolfa, P., Klopfleisch, R., Tvedten, H., Salas, E.N., Whitley, D.B., Aubreville, M., et al., 2021. Mitotic figures—normal, atypical, and imposters: A guide to identification. Veterinary pathology 58, 243–257. [30] Dukre, A., Deria, A., Xie, Y., Razzak, I., 2026. Stain-Aware Augmentation and Hybrid Loss for Domain-Generalized Atypical Mitosis Classification, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 159–165. [31] Eisenmann, M., Reinke, A., Weru, V., Tizabi, M.D., Isensee, F., Adler, T.J., Ali, S., Andrearczyk, V., Aubreville, M., Baid, U., et al., 2023. Why is the winner the best?, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 19955–19966. [32] Erickson, N., Mueller, J., Shirkov, A., Zhang, H., Larroy, P., Li, M., Smola, A., 2020. Autogluon-tabular: Robust and accurate automl for structured data. arXiv preprint arXiv:2003.06505 . [33] Fitzgibbons, P.L., Connolly, J.L., 2023. Protocol for the examination of resection specimens from patients with invasive carcinoma of the breast. CAP guidelines 4.8.1.0. URL: https://www.cap.org/ cancerprotocols.

M. Aubreville et al.: Preprint submitted to Elsevier

[34] Gamper, J., Alemi Koohbanani, N., Benet, K., Khuram, A., Rajpoot, N., 2019. Pannuke: an open pan-cancer histology dataset for nuclei instance segmentation and classification, in: European congress on digital pathology, Springer. pp. 11–19. [35] Ganz, J., Marzahl, C., Ammeling, J., Rosbach, E., Richter, B., Puget, C., Denk, D., Demeter, E.A., Tăbăran, F.A., Wasinger, G., et al., 2024. Information mismatch in phh3-assisted mitosis annotation leads to interpretation shifts in h&e slide analysis. Scientific reports 14, 26273. [36] Giedziun, P., Sotysik, J., Górczany, M., Ropiak, N., Przymus, M., Krajewski, P., Kwicień, J., Bartczak, A., Wasiak, I., Maniewski, M., 2026a. RF-DETR for Robust Mitotic Figure Detection: A MIDOG 2025 Track 1 Approach, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 99–104. [37] Giedziun, P., Sotysik, J., Górczany, M., Ropiak, N., Przymus, M., Krajewski, P., Kwiecień, J., Bartczak, A., Wasiak, I., Maniewski, M., 2026b. Foundation Model-Driven Classification of Atypical Mitotic Figures with Domain-Aware Training Strategies, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 166–171. [38] Gisselsson, D., 2008. Classification of chromosome segregation errors in cancer. Chromosoma 117, 511–519. doi:10.1007/ s00412-008-0169-1. [39] Graham, S., Vu, Q.D., Raza, S.E.A., Azam, A., Tsang, Y.W., Kwak, J.T., Rajpoot, N., 2019. Hover-net: Simultaneous segmentation and classification of nuclei in multi-tissue histology images. Medical Image Analysis 58, 101563. doi:10.1016/j.media.2019.101563. [40] He, K., Zhang, X., Ren, S., Sun, J., 2016. Deep residual learning for image recognition, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 770–778. doi:10.1109/ CVPR.2016.90. [41] Hendzel, M.J., Wei, Y., Mancini, M.A., Van Hooser, A., Ranalli, T., Brinkley, B., Bazett-Jones, D.P., Allis, C.D., 1997. Mitosisspecific phosphorylation of histone h3 initiates primarily within pericentromeric heterochromatin during g2 and spreads in an ordered fashion coincident with mitotic chromosome condensation. Chromosoma 106, 348–360. doi:10.1007/s004120050256. [42] Huang, G., Liu, Z., Van Der Maaten, L., Weinberger, K.Q., 2017. Densely connected convolutional networks, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4700–4708. doi:10.1109/CVPR.2017.243. [43] Isensee, F., Jaeger, P.F., Kohl, S.A., Petersen, J., Maier-Hein, K.H., 2021. nnu-net: a self-configuring method for deep learning-based biomedical image segmentation. Nature methods 18, 203–211. [44] Ivan, Z.Z., Hirling, D., Grexa, I., Ammeling, J., Molnar, C., Micsik, T., Dobra, K., Kuthi, L., Sukosd, F., Fillinger, J., et al., 2026. A subphase-labeled mitotic dataset for ai-powered cell division analysis. Scientific data . [45] Jahanifar, M., 2025. Mitosis subtyping dataset. URL: https://doi. org/10.5281/zenodo.15390543, doi:10.5281/zenodo.15390543. [46] Jahanifar, M., Dawood, M., Zamanitajeddin, N., Shephard, A., Chohan, B.S., Bertram, C.A., Wahab, N., Eastwood, M., Aubreville, M., Raza, S.E.A., et al., 2025. Pan-cancer profiling of mitotic topology & mitotic errors: Insights into prognosis, genomic alterations, and immune landscape. medRxiv , 2025–06. [47] Jin, Y., Stewénius, Y., Lindgren, D., Frigyesi, A., Calcagnile, O., Jonson, T., Edqvist, A., Larsson, N., Lundberg, L.M., Chebil, G., et al., 2007. Distinct mitotic segregation errors mediate chromosomal instability in aggressive urothelial cancers. Clinical cancer research 13, 1703–1712. [48] Kalatova, B., Jesenska, R., Hlinka, D., Dudas, M., 2015. Tripolar mitosis in human cells and embryos: occurrence, pathophysiology and medical implications. Acta histochemica 117, 111–125. [49] Kather, J.N., Krisam, J., Charoentong, P., Luedde, T., Herpel, E., Weis, C.A., Gaiser, T., Marx, A., Valous, N.A., Ferber, D., et al., 2019. Predicting survival from colorectal cancer histology slides using deep learning: A retrospective multicenter study. PLoS

Page 15 of 18

The MIDOG 2025 Challenge medicine 16, e1002730. [50] Kelam, N.S., Bonthu, S., Singhai, N., 2025. Atypical mitotic figure classification in midog 2025 using lora-enhanced uni models. doi:https://doi.org/10.5281/zenodo.17020013. [51] Kelam, N.S., Parekh, A., Bonthu, S., Singhal, N., 2026. Ensemble YOLO Framework for Multi-Domain Mitotic Figure Detection in Histopathology Images, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 105–111. [52] Khanam, R., Hussain, M., 2024. Yolov11: An overview of the key architectural enhancements. arXiv preprint arXiv:2410.17725 . [53] Kiupel, M., Webster, J., Bailey, K., Best, S., DeLay, J., Detrisac, C., Fitzgerald, S., Gamble, D., Ginn, P., Goldschmidt, M., et al., 2011. Proposal of a 2-tier histologic grading system for canine cutaneous mast cell tumors to more accurately predict biological behavior. Vet. Pathol. 48, 147–155. doi:10.1177/0300985810386469. [54] Kotte, S., Saipradeep, V.G., Walia, V., Nandagopal, D., Joseph, T., Sivadasan, N., Lali, B.S., 2026. MIDOG 2025 Track 2: A Deep Learning Model for Classification of Atypical and Normal Mitotic Figures under Class and Hardness Imbalances, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 172–179. [55] Krauss, S., Spieß, E., Hieber, D., Kramer, F., Schobel, J., Müller, D., 2026. Deep Learning Meets Morphology: A Hybrid Approach for Mitotic Figure Classification, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 37–46. [56] Kuhn, H.W., 1955. The hungarian method for the assignment problem. Naval research logistics quarterly 2, 83–97. [57] Lafarge, M.W., Bekkers, E.J., Pluim, J.P., Duits, R., Veta, M., 2021. Roto-translation equivariant convolutional networks: Application to histopathology image analysis. Medical Image Analysis 68, 101849. [58] Lafarge, M.W., Koelzer, V.H., 2026. Sequential Hard Mining: A Data-Centric Approach for Mitosis Detection, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 215–219. [59] Lashen, A., Toss, M.S., Alsaleem, M., Green, A.R., Mongan, N.P., Rakha, E., 2022. The characteristics and clinical significance of atypical mitosis in breast cancer. Modern Pathology 35, 1341–1348. [60] Liu, E., Lin, A., Kakodkar, P., Zhao, Y., Wang, B., Ling, C., Zhang, Q., 2025. A deep active learning framework for mitotic figure detection with minimal manual annotation and labelling. Histopathology 87, 536–547. [61] Liu, X., Peng, H., Zheng, N., Yang, Y., Hu, H., Yuan, Y., 2023. Efficientvit: Memory efficient vision transformer with cascaded group attention, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 14420–14430. [62] Liu, Y., Gadepalli, K., Norouzi, M., Dahl, G.E., Kohlberger, T., Boyko, A., Venugopalan, S., Timofeev, A., Nelson, P.Q., Corrado, G.S., et al., 2017. Detecting cancer metastases on gigapixel pathology images. arXiv preprint arXiv:1703.02442 . [63] Liu, Z., Mao, H., Wu, C.Y., Feichtenhofer, C., Darrell, T., Xie, S., 2022. A convnet for the 2020s, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11976– 11986. [64] Lv, J., Nasir, E.S., Xu, K., Jahanifar, M., Chohan, B.S., Elhaminia, B., Raza, S.E.A., 2026. Kongnet: A multi-headed deep learning model for detection and classification of nuclei in histopathology images. URL: https://arxiv.org/abs/2510.23559, arXiv:2510.23559. [65] Lyu, C., Zhang, W., Huang, H., Zhou, Y., Wang, Y., Liu, Y., Zhang, S., Chen, K., 2022. Rtmdet: An empirical study of designing realtime object detectors. arXiv preprint arXiv:2212.07784 . [66] Malon, C., Brachtel, E., Cosatto, E., Graf, H.P., Kurata, A., Kuroda, M., Meyer, J.S., Saito, A., Wu, S., Yagi, Y., 2012. Mitotic figure recognition: Agreement among pathologists and computerized detector. Analytical Cellular Pathology 35, 97–100. doi:10.3233/ ACP-2011-0029.

M. Aubreville et al.: Preprint submitted to Elsevier

[67] Marzahl, C., Aubreville, M., Bertram, C.A., Maier, J., Bergler, C., Kröger, C., Voigt, J., Breininger, K., Klopfleisch, R., Maier, A., 2021. EXACT: a collaboration toolset for algorithm-aided annotation of images with annotation version control. Scientific Reports 11:4343, 1–10. doi:10.1038/s41598-021-83827-4. [68] Marzahl, C., Napora, B., 2026. A bag of tricks for real-time Mitotic Figure detection, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 91–98. [69] Matsuda, Y., Yoshimura, H., Ishiwata, T., Sumiyoshi, H., Matsushita, A., Nakamura, Y., Aida, J., Uchida, E., Takubo, K., Arai, T., 2016. Mitotic index and multipolar mitosis in routine histologic sections as prognostic markers of pancreatic cancers: a clinicopathological study. Pancreatology 16, 127–132. [70] McNiel, E., Ogilvie, G., Powers, B., Hutchison, J., Salman, M., Withrow, S., 1997. Evaluation of prognostic factors for dogs with primary lung tumors: 67 cases (1985-1992). Journal of the American Veterinary Medical Association 211, 1422–1427. [71] Meng, B., Long, X., Liu, J., 2026. Adaptive Learning Strategies for Mitotic Figure Classification in MIDOG2025 Challenge, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 47–56. [72] Nasir, E.S., Jahanifar, J.L.M., Raza, S.E.A., 2026. Efficient Classification of Atypical vs. Normal Mitotic Figures Using LoRA-FineTuned Foundation Models, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 57–64. [73] Nasir, E.S., Lv, J., Jahanifar, M., Raza, S.E.A., 2025. Mitodetect++: A domain-robust pipeline for mitosis detection and atypical subtyping. arXiv preprint arXiv:2509.02586 . [74] Nechaev, D., Pchelnikov, A., Ivanova, E., 2024. Hibou: A family of foundational vision transformers for pathology. arXiv preprint arXiv:2406.05074 . [75] Nechaev, D., Pchelnikov, A., Ivanova, E., 2025. Spider: A comprehensive multi-organ supervised pathology dataset and baseline models. URL: https://arxiv.org/abs/2503.02876, arXiv:2503.02876. [76] Ochi, M., Bae, Y., 2026. Ensemble of Pathology Foundation Models for MIDOG 2025 Track 2: Atypical Mitosis Classification, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 180– 186. [77] Ohashi, R., Namimatsu, S., Sakatani, T., Naito, Z., Takei, H., Shimizu, A., 2018. Prognostic utility of atypical mitoses in patients with breast cancer: A comparative study with Ki67 and phosphohistone H3. Journal of surgical oncology 118, 557–567. [78] Percannella, G., Sarno, M., Tortorella, F., Vento, M., 2026a. A multi-task neural network for atypical mitosis recognition under domain shift, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 187–192. [79] Percannella, G., Sarno, M., Tortorella, F., Vento, M., 2026b. Mitosis detection in domain shift scenarios: a Mamba-based approach, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 119–124. [80] Qi, X., Lee, M., Labella, D., Sanford, T., 2026. Normal and Atypical Mitosis Image Classifier using Efficient Vision Transformer, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 193–197. [81] Rakib Hasan, K., Kim, S., Cho, J., 2025. A short document for midog2025 challenge model submission. URL: https://doi.org/ 10.5281/zenodo.17017921, doi:10.5281/zenodo.17017921. [82] Ramchandani, L., Deotale, G., Das, D.K., 2026. Parameter-efficient fine-tuning (PEFT) of Vision Foundation Models for Atypical Mitotic Figure Classification, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 198–203. [83] Ronneberger, O., Fischer, P., Brox, T., 2015. U-net: Convolutional networks for biomedical image segmentation, in: International

Page 16 of 18

The MIDOG 2025 Challenge Conference on Medical image computing and computer-assisted intervention, Springer. pp. 234–241. [84] Roux, L., Racoceanu, D., Capron, F., Calvo, J., Attieh, E., Le Naour, G., Gloaguen, A., 2014. Mitos & atypia. Image Pervasive Access Lab (IPAL), Agency Sci., Technol. & Res. Inst. Infocom Res., Singapore, Tech. Rep 1, 1–8. [85] Roux, L., Racoceanu, D., Loménie, N., Kulikova, M., Irshad, H., Klossa, J., Capron, F., Genestie, C., Le Naour, G., Gurcan, M., 2013. Mitosis detection in breast cancer histological images an icpr 2012 contest. Journal of Pathology Informatics 4, 8. doi:10.4103/ 2153-3539.112693. [86] Ruan, J., Li, J., Xiang, S., 2024. Vm-unet: Vision mamba unet for medical image segmentation. ACM Transactions on Multimedia Computing, Communications and Applications . [87] Sapkota, R., Cheppally, R.H., Sharda, A., Karkee, M., 2025. Rfdetr object detection vs yolov12: A study of transformer-based and cnn-based architectures for single-class and multi-class greenfruit detection in complex orchard environments under label ambiguity. URL: https://arxiv.org/abs/2504.13099, arXiv:2504.13099. arXiv preprint arXiv:2504.13099. [88] Sdeor, E., Okada, H., Saad, R., Ben-Yishay, T., Ben-David, U., 2024. Aneuploidy as a driver of human cancer. Nature genetics 56, 2014– 2026. doi:10.1038/s41588-024-01916-2. [89] Shen, Z., Bär, E., Hawkins, M., Bräutigam, K., Collins-Fekete, C.A., 2025a. Pan-cancer mitotic figures detection and domain generalization: Midog 2025 challenge. URL: https://arxiv.org/ abs/2509.02585, arXiv:2509.02585. [90] Shen, Z., Hawkins, M.A., Baer, E., Brautigam, K., Fekete, C.A.C., 2025b. OMG-Octo Atypical: A refinement of the original OMGOcto database to incorporate atypical mitoses. Zenodo. URL: https: //doi.org/10.5281/zenodo.16107743, doi:10.5281/zenodo.16107743. [91] Shen, Z., Simard, M., Brand, D., Andrei, V., Al-Khader, A., Oumlil, F., Trevers, K., Butters, T., Haefliger, S., Kara, E., Amary, F., Tirabosco, R., Cool, P., Royle, G., Hawkins, M.A., Flanagan, A.M., Collins Fekete, C.A., 2024. Omg-octo: Uniformised large scale database of mitotic figures in haematoxylin and eosin-stained slides. URL: https://doi.org/10.5281/zenodo.14246170, doi:10. 5281/zenodo.14246170. [92] Siméoni, O., Vo, H.V., Seitzer, M., Baldassarre, F., Oquab, M., Jose, C., Khalidov, V., Szafraniec, M., Yi, S., Ramamonjisoa, M., et al., 2025. Dinov3. arXiv preprint arXiv:2508.10104 . [93] Simonyan, K., Zisserman, A., 2014. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556 . [94] Stathonikos, N., Aubreville, M., de Vries, S., Wilm, F., Bertram, C.A., Veta, M., van Diest, P.J., 2024. Breast cancer survival prediction using an automated mitosis detection pipeline. The Journal of Pathology: Clinical Research 10, e70008. [95] Szegedy, C., Liu, W., Jia, Y., Sermanet, P., Reed, S., Anguelov, D., Erhan, D., Vanhoucke, V., Rabinovich, A., 2015. Going deeper with convolutions, in: Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1–9. [96] Tan, M., Le, Q., 2019. EfficientNet: Rethinking model scaling for convolutional neural networks, in: Chaudhuri, K., Salakhutdinov, R. (Eds.), Proceedings of the 36th International Conference on Machine Learning, PMLR. pp. 6105–6114. URL: https://proceedings.mlr. press/v97/tan19a.html. [97] Tan, M., Le, Q., 2021. Efficientnetv2: Smaller models and faster training, in: International conference on machine learning, PMLR. pp. 10096–10106. [98] Tellez, D., Balkenhol, M., Otte-Höller, I., van de Loo, R., Vogels, R., Bult, P., Wauters, C., Vreuls, W., Mol, S., Karssemeijer, N., et al., 2018. Whole-slide mitosis detection in h&e breast histology using phh3 as a reference to train distilled stain-invariant convolutional networks. IEEE transactions on medical imaging 37, 2126–2136. doi:10.1109/TMI.2018.2820199. [99] Tian, Y., Ye, Q., Doermann, D., . Yolov12: Attention-centric realtime object detectors, in: The Thirty-ninth Annual Conference on

M. Aubreville et al.: Preprint submitted to Elsevier

Neural Information Processing Systems. [100] Tian, Z., Shen, C., Chen, H., He, T., 2019. Fcos: Fully convolutional one-stage object detection, in: Proceedings of the IEEE/CVF international conference on computer vision, pp. 9627–9636. [101] Topuz, Y., Gökcan, M.T., Yıldız, S., Varlı, S., 2026. Single Detect Focused YOLO Framework for Robust Mitotic Figure Detection, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 65–74. [102] Varghese, R., Sambath, M., 2024. Yolov8: A novel object detection algorithm with enhanced performance and robustness, in: 2024 International conference on advances in data engineering and intelligent computing systems (ADICS), IEEE. pp. 1–6. [103] Veta, M., Heng, Y.J., Stathonikos, N., Bejnordi, B.E., Beca, F., Wollmann, T., Rohr, K., Shah, M.A., Wang, D., Rousson, M., et al., 2019. Predicting breast tumor proliferation from whole-slide images: the tupac16 challenge. Medical image analysis 54, 111–121. doi:10.1016/j.media.2019.02.012. [104] Veta, M., Van Diest, P.J., Jiwa, M., Al-Janabi, S., Pluim, J.P., 2016. Mitosis counting in breast cancer: Object-level interobserver agreement and comparison to an automatic method. PloS one 11, e0161286. doi:10.1371/journal.pone.0161286. [105] Veta, M., Van Diest, P.J., Willems, S.M., Wang, H., Madabhushi, A., Cruz-Roa, A., Gonzalez, F., Larsen, A.B., Vestergaard, J.S., Dahl, A.B., et al., 2015. Assessment of algorithms for mitosis detection in breast cancer histopathology images. Medical image analysis 20, 237–248. doi:10.1016/j.media.2014.11.010. [106] Virtanen, P., Gommers, R., Burovski, E., Oliphant, T.E., Weckesser, W., Cournapeau, D., Peterson, P., Reddy, T., Haberland, M., Wilson, J., et al., 2021. scipy/scipy: Scipy 1.6. 0. Zenodo . [107] Vorontsov, E., Bozkurt, A., Casson, A., Shaikovski, G., Zelechowski, M., Liu, S., Severson, K., Zimmermann, E., Hall, J., Tenenholtz, N., et al., 2023. Virchow: A million-slide digital pathology foundation model. arXiv preprint arXiv:2309.07778 . [108] Walia, V., Nandagopal, D., Kotte, S., Saipradeep, V.G., Joseph, T., Sivadasan, N., Lali, B.S., 2026. Mitosis Detection in the Wild Using Detection Transformers, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 112–118. [109] Wang, A., Chen, H., Liu, L., Chen, K., Lin, Z., Han, J., Ding, G., 2024. Yolov10: Real-time end-to-end object detection. Advances in neural information processing systems 37, 107984–108011. [110] Weiss, V., Banerjee, S., Donovan, T., Conrad, T., Klopfleisch, R., Ammeling, J., Kaltenecker, C., Hirling, D., Veta, M., Stathonikos, N., Horvath, P., Breininger, K., Aubreville, M., Bertram, C., 2025. A dataset of atypical vs normal mitoses classification for midog 2025. URL: https://doi.org/10.5281/zenodo.16044804, doi:10.5281/ zenodo.16044804. [111] WHO Classification of Tumours Editorial Board, 2022. Who classification of endocrine and neuroendocrine tumours. . [112] Woo, S., Debnath, S., Hu, R., Chen, X., Liu, Z., Kweon, I.S., Xie, S., 2023. Convnext v2: Co-designing and scaling convnets with masked autoencoders, in: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 16133–16142. [113] Xiao, J., Liu, S., Lyu, M., 2026. A Two-Stage Strategy for Mitosis Detection Using Improved YOLO11x Proposals and ConvNeXt Classification, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 125–133. [114] Xu, T., Yang, H., Chen, X.A., Gu, H., Haeri, M., 2026. Team Westwood Solution for MIDOG 2025 Challenge: An EnsembleCNN-Based Approach for Mitosis Detection and Classification, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 220–226. [115] Yamagishi, Y., Hanaoka, S., 2026. Automated Classification of Normal and Atypical Mitotic Figures Using ConvNeXt V2: MIDOG 2025 Track 2, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 75–86.

Page 17 of 18

The MIDOG 2025 Challenge [116] Zbontar, J., Jing, L., Misra, I., LeCun, Y., Deny, S., 2021. Barlow twins: Self-supervised learning via redundancy reduction, in: International conference on machine learning, PMLR. pp. 12310–12320. [117] Zimmermann, E., Vorontsov, E., Viret, J., Casson, A., Zelechowski, M., Shaikovski, G., Tenenholtz, N., Hall, J., Klimstra, D., Yousfi, R., et al., 2024. Virchow2: Scaling self-supervised mixed magnification models in pathology. arXiv preprint arXiv:2408.00738 . [118] Çayır, S., Kukuk, S.B., 2026. Dual-Fusion Double-Ensemble (DFDE) Framework for Atypical Mitosis Classification, in: Aubreville, M., Bertram, C.A. (Eds.), Mitotic Figure Detection and Atypia Classification in Whole Slide Images, Springer. pp. 153–158.

Median F1 score hAC hBlC hMel cHAS fSTS cMC ccMCT hCoC fLym hLUAD hGBM hMen

0.599 ±0.067 0.800 ±0.025 0.787 ±0.037 0.808 ±0.041 0.719 ±0.046 0.713 ±0.047 0.757 ±0.051 0.733 ±0.023 0.745 ±0.032 0.532 ±0.087 0.444 ±0.089 0.705 ±0.056

Hotspot

0.268 ±0.207 0.539 ±0.142 0.618 ±0.066 0.668 ±0.073 0.585 ±0.069 0.645 ±0.077 0.715 ±0.055 0.691 ±0.023 0.714 ±0.033 0.528 ±0.154 0.506 ±0.089 0.794 ±0.045

0.080 ±0.165 0.457 ±0.248 0.157 ±0.142 0.552 ±0.131 0.376 ±0.104 0.486 ±0.093 0.561 ±0.132 0.615 ±0.100 0.634 ±0.066 0.401 ±0.103 0.010 ±0.437 0.475 ±0.098

+0.33 +0.26 +0.17 +0.14 +0.13 +0.07 +0.04 +0.04 +0.03 +0.00 -0.06 -0.09

Random Challenging rand hot

+0.52 +0.34 +0.63 +0.26 +0.34 +0.23 +0.20 +0.12 +0.11 +0.13 +0.43 +0.23

hot

ll

cha

Supplementary Figure S1: 𝐹1 score by tumor type for Track 1 of the challenge (mitotic figure detection). Shown are median values ± inter-quartile range (25th-75th percentile).

M. Aubreville et al.: Preprint submitted to Elsevier

Page 18 of 18

The MIDOG 2025 Challenge

Supplementary Figure S2: Precision and recall for all tumor domains and area types of Track 1 (mitotic figure detection).

Precision hMel 0.73 0.74 0.63 0.75 0.57 0.76 0.64 0.73 0.62 0.50 0.64 0.76 0.55 0.21 hAC 0.47 0.65 0.13 0.49 0.43 0.29 0.35 0.42 0.22 0.22 0.27 0.51 0.16 0.05 hBlC

1.0 0.9

0.78 0.73 0.71 0.73 0.68 0.81 0.76 0.78 0.59 0.61 0.69 0.85 0.63 0.35

cMC 0.73 0.77 0.62 0.61 0.64 0.76 0.74 0.72 0.59 0.64 0.68 0.86 0.56 0.50 ccMCT 0.82 0.83 0.84 0.93 0.85 0.88 0.89 0.74 0.82 0.76 0.80 0.84 0.72 0.70 hMen 0.71 0.77 0.62 0.72 0.68 0.85 0.77 0.74 0.67 0.64 0.73 0.82 0.56 0.43 hCoC 0.75 0.76 0.71 0.77 0.70 0.82 0.64 0.76 0.63 0.61 0.71 0.81 0.64 0.34

0.8 0.7 0.6

Recall hMel 0.84 0.82 0.86 0.78 0.89 0.75 0.86 0.82 0.88 0.88 0.76 0.78 0.82 0.90 hAC 0.68 0.63 0.66 0.66 0.63 0.56 0.61 0.61 0.68 0.71 0.61 0.44 0.66 0.73 hBlC

1.0 0.9

0.83 0.84 0.86 0.85 0.86 0.79 0.84 0.76 0.87 0.86 0.81 0.70 0.83 0.87

cMC 0.66 0.64 0.72 0.65 0.70 0.60 0.59 0.60 0.71 0.62 0.63 0.48 0.68 0.64 ccMCT 0.73 0.61 0.66 0.51 0.67 0.59 0.56 0.81 0.65 0.70 0.63 0.63 0.63 0.72 hMen 0.73 0.69 0.74 0.70 0.72 0.55 0.68 0.64 0.76 0.74 0.67 0.58 0.69 0.78 hCoC 0.72 0.68 0.73 0.67 0.74 0.63 0.77 0.66 0.79 0.76 0.67 0.55 0.72 0.85

0.8 0.7 0.6

cHAS 0.84 0.83 0.80 0.76 0.75 0.76 0.82 0.73 0.73 0.78 0.80 0.91 0.64 0.56

0.5

cHAS 0.78 0.73 0.80 0.79 0.79 0.69 0.71 0.72 0.79 0.75 0.71 0.60 0.80 0.73

0.5

fSTS 0.77 0.76 0.68 0.74 0.66 0.81 0.73 0.75 0.68 0.67 0.74 0.83 0.64 0.60 fLym 0.85 0.82 0.79 0.80 0.78 0.88 0.81 0.80 0.71 0.83 0.86 0.89 0.74 0.76

0.4

fSTS 0.62 0.62 0.69 0.68 0.64 0.52 0.60 0.52 0.71 0.69 0.64 0.40 0.62 0.70 fLym 0.69 0.65 0.72 0.70 0.70 0.59 0.65 0.65 0.70 0.68 0.58 0.52 0.69 0.69

0.4 0.3 0.2

npic-ab

schoe

vidushiwalia

SZTU-134

ammeling

masarno

piotrgiedziun

tengyoux

christian.marzahl

cacfek

navyasri.kelam

ytopuz53

0.2

hGBM 0.50 0.48 0.32 0.61 0.26 0.50 0.30 0.53 0.44 0.38 0.29 0.38 0.36 0.41 hLUAD 0.50 0.53 0.37 0.50 0.38 0.57 0.25 0.44 0.53 0.45 0.43 0.41 0.43 0.45 wildsquirrel

0.3

RaphaelBourgade

npic-ab

schoe

vidushiwalia

SZTU-134

ammeling

masarno

piotrgiedziun

tengyoux

christian.marzahl

cacfek

navyasri.kelam

ytopuz53

RaphaelBourgade

wildsquirrel

hGBM 0.70 0.64 0.87 0.17 0.71 0.63 0.83 0.56 0.47 0.61 0.54 0.60 0.31 0.52 hLUAD 0.74 0.70 0.64 0.63 0.59 0.71 0.75 0.58 0.59 0.67 0.66 0.76 0.49 0.53

Supplementary Figure S3: Precision and recall for all participants and tumor types in Track 1 (mitotic figure detection).

M. Aubreville et al.: Preprint submitted to Elsevier

Page 19 of 18

The MIDOG 2025 Challenge

fSTS 2.56 4.11 4.16 4.46 3.88 4.12 3.94 3.57 4.40 4.27 4.42 2.91 3.51 3.99

2

fLym 3.43 4.09 4.98 4.70 4.83 5.01 4.42 4.42 4.23 4.84 4.52 3.72 4.34 4.47 1

npic-ab

schoe

vidushiwalia

hGBM 3.80 3.92 3.79 3.92 2.38 5.56 3.69 4.09 3.40 5.14 2.71 2.92 3.64 4.57 hLUAD 3.52 4.05 3.92 4.03 3.34 5.20 3.39 3.23 3.88 4.51 3.41 3.13 3.31 3.71 SZTU-134

0.2

3

ammeling

0.3

hCoC 2.49 4.43 4.86 4.69 4.82 4.83 4.40 4.48 4.82 4.58 4.47 3.79 4.37 3.69 cHAS 4.26 5.35 5.96 5.38 5.64 5.14 5.26 5.02 5.64 5.59 5.31 4.50 5.13 4.07

piotrgiedziun

npic-ab

schoe

vidushiwalia

SZTU-134

ammeling

piotrgiedziun

masarno

tengyoux

christian.marzahl

cacfek

navyasri.kelam

ytopuz53

RaphaelBourgade

wildsquirrel

hGBM 0.58 0.55 0.47 0.27 0.38 0.56 0.44 0.54 0.45 0.47 0.38 0.46 0.33 0.46 hLUAD 0.59 0.60 0.47 0.56 0.47 0.64 0.37 0.50 0.56 0.54 0.52 0.53 0.46 0.49

0.4

5

4

masarno

fLym 0.76 0.72 0.75 0.75 0.74 0.71 0.72 0.72 0.71 0.74 0.69 0.65 0.71 0.72

0.5

6

hMen 4.55 5.18 5.73 5.46 5.66 5.50 5.61 4.89 5.73 5.67 5.42 4.43 5.09 5.15

tengyoux

fSTS 0.68 0.68 0.68 0.71 0.65 0.63 0.66 0.62 0.70 0.68 0.69 0.54 0.63 0.65

0.6

7

ccMCT 5.06 4.73 5.85 4.06 5.78 6.22 5.77 6.13 5.07 5.89 5.20 4.88 5.46 5.53

christian.marzahl

hCoC 0.73 0.72 0.72 0.71 0.72 0.71 0.70 0.71 0.70 0.68 0.69 0.66 0.67 0.49 cHAS 0.81 0.78 0.80 0.78 0.77 0.72 0.76 0.72 0.76 0.76 0.75 0.73 0.71 0.63

0.7

cacfek

hMen 0.72 0.73 0.67 0.71 0.70 0.67 0.72 0.68 0.71 0.69 0.70 0.68 0.62 0.55

cMC 3.95 4.74 5.26 4.83 5.12 5.33 5.04 4.50 4.86 4.81 4.91 3.66 4.70 4.04

navyasri.kelam

ccMCT 0.77 0.70 0.74 0.66 0.75 0.71 0.69 0.78 0.72 0.73 0.71 0.72 0.68 0.71

0.8

hAC 4.96 5.18 4.97 5.16 5.17 5.01 5.05 4.76 5.26 5.27 5.06 3.44 4.86 4.40 hBlC 4.23 5.90 6.20 6.16 6.10 6.22 6.14 5.56 5.77 5.63 5.85 5.19 5.35 4.80

ytopuz53

cMC 0.69 0.70 0.66 0.63 0.67 0.67 0.66 0.66 0.64 0.63 0.65 0.61 0.62 0.56

0.9

FROC AUC hMel 5.01 6.18 6.66 6.11 6.39 6.34 6.54 6.10 6.55 6.41 6.24 5.88 6.15 4.99

RaphaelBourgade

hAC 0.55 0.64 0.22 0.56 0.51 0.38 0.44 0.50 0.34 0.34 0.38 0.47 0.25 0.10 hBlC 0.80 0.78 0.78 0.78 0.76 0.80 0.80 0.77 0.70 0.71 0.75 0.77 0.72 0.50

1.0

wildsquirrel

F1 score hMel 0.78 0.78 0.73 0.76 0.70 0.76 0.74 0.77 0.73 0.64 0.69 0.77 0.66 0.34

Supplementary Figure S4: 𝐹1 and FROC AUC for all participants and tumor types in Track 1 (mitotic figure detection).

M. Aubreville et al.: Preprint submitted to Elsevier

Page 20 of 18

1.0

0.9

0.9

0.8

0.8

hGBM

hLUAD

fSTS

fLym

hCoC

cHAS

hMen

cMC

hGBM

hLUAD

fSTS

fLym

hCoC

cHAS

hMen

cMC

ccMCT

hAC

0.4

hBlC

0.5

0.4

ccMCT

0.6

0.5

hAC

0.6

0.7

hBlC

0.7

hMel

ROC AUC

1.0

hMel

Balanced Accuracy

The MIDOG 2025 Challenge

Supplementary Figure S5: Balanced Accuracy (left) and ROC AUC (right) across tumor type of challenge Track 2 (atypical classification), pooled across all participants.

hAC

1.0 hBlC hMel

0.9

Sensitivity

hGBM

hLUAD

fSTS hMen hCoC

cHAS cMC

0.8 ccMCT

0.7 fLym

0.6

0.5

0.5

0.6

0.7

0.8

0.9

1.0

Specificity Supplementary Figure S6: Specificity vs. Sensitivity by tumor type, pooled across all teams and samples, for Track 2 (atypical classification) of the challenge.

M. Aubreville et al.: Preprint submitted to Elsevier

Page 21 of 18

The MIDOG 2025 Challenge

guillaume.balezo nasires yohsuke.yamagishi masarno zerostarcraft krausara lrc9859 be_yuan piotrgiedziun cacfek schoe saipradeepvg navyasri.kelam Leire Kaustubh_Atey mirazzak mlafarge baseline chillice

1.0 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2

Sensitivity hMel 0.89 0.92 0.89 0.97 0.97 0.89 0.92 0.83 0.78 1.00 0.94 0.94 0.94 0.78 0.81 0.92 0.97 0.94 0.75 hAC 0.86 1.00 1.00 0.86 1.00 1.00 0.71 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 0.86 0.86 1.00 hBlC 0.94 0.97 0.93 0.97 0.93 0.95 0.91 0.92 0.89 0.98 0.94 0.97 0.97 0.83 0.86 0.97 0.91 0.89 0.82 cMC 0.91 0.82 0.82 0.91 0.91 0.82 0.82 0.82 0.82 0.82 0.91 0.91 1.00 0.91 0.73 0.82 0.45 0.82 0.73 ccMCT 0.86 0.71 0.43 0.86 0.86 0.71 0.57 0.57 0.57 0.86 0.71 0.71 0.86 0.43 0.71 0.57 0.86 0.57 0.43 hMen 0.86 0.89 0.78 0.94 0.92 0.89 0.83 0.75 0.75 0.94 0.92 0.94 0.97 0.72 0.75 0.94 0.72 0.92 0.78 hCoC 0.87 0.88 0.78 0.88 0.90 0.91 0.83 0.81 0.68 0.90 0.88 0.91 0.96 0.70 0.74 0.91 0.87 0.84 0.52 cHAS 0.91 0.83 0.82 0.93 0.83 0.89 0.84 0.89 0.81 0.96 0.87 0.88 0.94 0.80 0.79 0.82 0.81 0.81 0.67 fSTS 0.94 0.88 0.92 0.95 0.89 0.93 0.92 0.89 0.86 0.95 0.95 0.95 0.94 0.84 0.85 0.90 0.79 0.83 0.70 fLym 0.71 0.61 0.55 0.77 0.68 0.61 0.68 0.55 0.48 0.68 0.81 0.77 0.65 0.42 0.52 0.68 0.65 0.65 0.32 hGBM 0.89 0.92 0.89 0.97 0.86 0.94 0.78 0.83 0.61 0.86 1.00 0.92 0.83 0.78 0.81 0.86 0.86 0.83 0.69 hLUAD 0.93 0.86 0.84 0.91 0.95 0.93 0.84 0.86 0.79 0.93 0.91 0.86 0.98 0.81 0.77 0.88 0.79 0.81 0.74 guillaume.balezo nasires yohsuke.yamagishi masarno zerostarcraft krausara lrc9859 be_yuan piotrgiedziun cacfek schoe saipradeepvg navyasri.kelam Leire Kaustubh_Atey mirazzak mlafarge baseline chillice

Specificity hMel 0.90 0.88 0.93 0.81 0.86 0.85 0.90 0.88 0.92 0.68 0.69 0.79 0.73 0.91 0.89 0.74 0.61 0.71 0.92 hAC 1.00 1.00 0.97 0.97 0.87 0.90 1.00 0.90 0.97 0.87 0.90 0.83 0.80 0.93 0.97 0.93 0.77 0.83 1.00 hBlC 0.87 0.85 0.92 0.77 0.84 0.80 0.81 0.86 0.90 0.75 0.73 0.70 0.72 0.93 0.90 0.66 0.75 0.75 0.94 cMC 0.91 0.95 0.97 0.89 0.93 0.92 0.90 0.94 0.96 0.84 0.83 0.82 0.83 0.96 0.94 0.82 0.89 0.82 0.97 ccMCT 0.94 0.93 0.97 0.92 0.95 0.89 0.94 0.92 0.98 0.87 0.94 0.89 0.88 0.97 0.96 0.94 0.85 0.87 0.97 hMen 0.93 0.91 0.93 0.84 0.91 0.89 0.94 0.90 0.96 0.71 0.76 0.78 0.81 0.96 0.91 0.73 0.78 0.71 0.96 hCoC 0.94 0.93 0.95 0.88 0.93 0.89 0.94 0.95 0.97 0.86 0.87 0.85 0.85 0.97 0.96 0.83 0.82 0.85 0.98 cHAS 0.90 0.94 0.93 0.87 0.90 0.87 0.89 0.91 0.94 0.81 0.81 0.78 0.80 0.93 0.92 0.81 0.82 0.82 0.97 fSTS 0.91 0.90 0.95 0.88 0.90 0.88 0.91 0.91 0.95 0.83 0.83 0.85 0.85 0.96 0.93 0.82 0.87 0.81 0.97 fLym 0.95 0.97 0.98 0.92 0.95 0.93 0.95 0.97 0.98 0.91 0.92 0.88 0.87 0.99 0.96 0.90 0.92 0.90 1.00 hGBM 0.70 0.83 0.87 0.67 0.70 0.77 0.93 0.83 0.87 0.73 0.60 0.70 0.63 0.93 0.80 0.63 0.67 0.67 0.83 hLUAD 0.83 0.86 0.90 0.73 0.78 0.75 0.89 0.82 0.94 0.74 0.73 0.72 0.69 0.93 0.87 0.68 0.81 0.69 0.91

1.0 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2

Supplementary Figure S7: Specificity / Sensitivity per tumor type and team in Track 2 (atypical classification) of the challenge.

guillaume.balezo nasires yohsuke.yamagishi masarno zerostarcraft krausara lrc9859 be_yuan piotrgiedziun cacfek schoe saipradeepvg navyasri.kelam Leire Kaustubh_Atey mirazzak mlafarge baseline chillice

1.0 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2

Sensitivity hMel 0.89 0.92 0.89 0.97 0.97 0.89 0.92 0.83 0.78 1.00 0.94 0.94 0.94 0.78 0.81 0.92 0.97 0.94 0.75 hAC 0.86 1.00 1.00 0.86 1.00 1.00 0.71 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 1.00 0.86 0.86 1.00 hBlC 0.94 0.97 0.93 0.97 0.93 0.95 0.91 0.92 0.89 0.98 0.94 0.97 0.97 0.83 0.86 0.97 0.91 0.89 0.82 cMC 0.91 0.82 0.82 0.91 0.91 0.82 0.82 0.82 0.82 0.82 0.91 0.91 1.00 0.91 0.73 0.82 0.45 0.82 0.73 ccMCT 0.86 0.71 0.43 0.86 0.86 0.71 0.57 0.57 0.57 0.86 0.71 0.71 0.86 0.43 0.71 0.57 0.86 0.57 0.43 hMen 0.86 0.89 0.78 0.94 0.92 0.89 0.83 0.75 0.75 0.94 0.92 0.94 0.97 0.72 0.75 0.94 0.72 0.92 0.78 hCoC 0.87 0.88 0.78 0.88 0.90 0.91 0.83 0.81 0.68 0.90 0.88 0.91 0.96 0.70 0.74 0.91 0.87 0.84 0.52 cHAS 0.91 0.83 0.82 0.93 0.83 0.89 0.84 0.89 0.81 0.96 0.87 0.88 0.94 0.80 0.79 0.82 0.81 0.81 0.67 fSTS 0.94 0.88 0.92 0.95 0.89 0.93 0.92 0.89 0.86 0.95 0.95 0.95 0.94 0.84 0.85 0.90 0.79 0.83 0.70 fLym 0.71 0.61 0.55 0.77 0.68 0.61 0.68 0.55 0.48 0.68 0.81 0.77 0.65 0.42 0.52 0.68 0.65 0.65 0.32 hGBM 0.89 0.92 0.89 0.97 0.86 0.94 0.78 0.83 0.61 0.86 1.00 0.92 0.83 0.78 0.81 0.86 0.86 0.83 0.69 hLUAD 0.93 0.86 0.84 0.91 0.95 0.93 0.84 0.86 0.79 0.93 0.91 0.86 0.98 0.81 0.77 0.88 0.79 0.81 0.74 guillaume.balezo nasires yohsuke.yamagishi masarno zerostarcraft krausara lrc9859 be_yuan piotrgiedziun cacfek schoe saipradeepvg navyasri.kelam Leire Kaustubh_Atey mirazzak mlafarge baseline chillice

Specificity hMel 0.90 0.88 0.93 0.81 0.86 0.85 0.90 0.88 0.92 0.68 0.69 0.79 0.73 0.91 0.89 0.74 0.61 0.71 0.92 hAC 1.00 1.00 0.97 0.97 0.87 0.90 1.00 0.90 0.97 0.87 0.90 0.83 0.80 0.93 0.97 0.93 0.77 0.83 1.00 hBlC 0.87 0.85 0.92 0.77 0.84 0.80 0.81 0.86 0.90 0.75 0.73 0.70 0.72 0.93 0.90 0.66 0.75 0.75 0.94 cMC 0.91 0.95 0.97 0.89 0.93 0.92 0.90 0.94 0.96 0.84 0.83 0.82 0.83 0.96 0.94 0.82 0.89 0.82 0.97 ccMCT 0.94 0.93 0.97 0.92 0.95 0.89 0.94 0.92 0.98 0.87 0.94 0.89 0.88 0.97 0.96 0.94 0.85 0.87 0.97 hMen 0.93 0.91 0.93 0.84 0.91 0.89 0.94 0.90 0.96 0.71 0.76 0.78 0.81 0.96 0.91 0.73 0.78 0.71 0.96 hCoC 0.94 0.93 0.95 0.88 0.93 0.89 0.94 0.95 0.97 0.86 0.87 0.85 0.85 0.97 0.96 0.83 0.82 0.85 0.98 cHAS 0.90 0.94 0.93 0.87 0.90 0.87 0.89 0.91 0.94 0.81 0.81 0.78 0.80 0.93 0.92 0.81 0.82 0.82 0.97 fSTS 0.91 0.90 0.95 0.88 0.90 0.88 0.91 0.91 0.95 0.83 0.83 0.85 0.85 0.96 0.93 0.82 0.87 0.81 0.97 fLym 0.95 0.97 0.98 0.92 0.95 0.93 0.95 0.97 0.98 0.91 0.92 0.88 0.87 0.99 0.96 0.90 0.92 0.90 1.00 hGBM 0.70 0.83 0.87 0.67 0.70 0.77 0.93 0.83 0.87 0.73 0.60 0.70 0.63 0.93 0.80 0.63 0.67 0.67 0.83 hLUAD 0.83 0.86 0.90 0.73 0.78 0.75 0.89 0.82 0.94 0.74 0.73 0.72 0.69 0.93 0.87 0.68 0.81 0.69 0.91

1.0 0.9 0.8 0.7 0.6 0.5 0.4 0.3 0.2

Supplementary Figure S8: Balanced Accuracy and ROC AUC per tumor type and team in Track 2 (atypical classification) of the challenge.

M. Aubreville et al.: Preprint submitted to Elsevier

Page 22 of 18

Record · ID 266214 · SHA-256 09f17e7e9021871a
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.