PREPRINT
1
MiTHras: Task-specific Hierarchical Semi-supervised Contrastive Masked Autoencoder for Mitotic Figure Analysis
arXiv:2609.24736v1 [cs.CV] 21 Sep 2026
Trinh T.L. Vuong, Simon Graham, Quoc Dang Vu, Phat T.H. Ho, Jeewoo Lim, Mostafa Jahanifar, Nasir Rajpoot* , Jin T. Kwak* , Member, IEEE
Abstract—Mitotic figure (MF) analysis supports tumor grading and prognostic assessment, but automated models remain sensitive to differences in tissue type and image acquisition. We present MiTHras, a task-specific pretraining framework that combines pseudo-label-guided image- and token-level contrastive learning with masked reconstruction. We construct TCGA-MF-Pseudo, a corpus of 1.8 million cell-centered images from 14 TCGA cohorts spanning 11 organ sites. Comprehensive evaluation on MF classification, detection, count-based survival prediction, and subtype classification demonstrates the efficacy of MiTHras. It achieves the highest mean F1 on all three MF classification benchmarks and both subtype benchmarks. MiTHras also outperforms general-purpose and pathology foundation encoders by a larger margin under frozen-encoder linear probing than under full fine-tuning. Although detection gains are modest due to a shared candidate-detection stage, ablations confirm that tokenlevel supervision improves typical-versus-atypical classification and linear probing. These findings establish that MiTHras yields robust, transferable representations for automated mitotic activity assessment. Index Terms—MAE, Self-supervised, Semi-supervised, Mitotic detection
I. I NTRODUCTION The mitotic count is a marker of cell proliferation used in cancer diagnosis, tumor grading, and prognostic assessment [1]. Pathologists determine it by counting mitotic figures (MFs) within a specified tissue area, following tumor-specific grading protocols [2], [3]. It is a core component of the Nottingham grading system for invasive breast carcinoma [4] and FNCLCC grading in soft-tissue sarcoma [5], drives risk stratification in gastrointestinal stromal tumors [6], and defines tumor grade together with Ki-67 in gastroenteropancreatic neuroendocrine tumors [7]. Automated MF analysis provides This work was supported by the National Research Foundation of Korea (NRF) grant (No. RS-2025-00558322), Korea Institute for Advancement of Technology (KIAT) grant (No. P0022543), and the Innovate UK grant (No. 10040491). (Corresponding authors: Nasir Rajpoot and Jin T. Kwak.) Trinh T.L. Vuong is with Department of Pathology, Brigham and Women’s Hospital, Harvard Medical School, Boston, USA (e-mail: [email protected]). Jeewoo Lim and Jin T. Kwak are with School of Electrical Engineering, Korea University, Seoul 02841, Korea (e-mail: {jeewoolim and jkwak}@korea.ac.kr). Phat T.H. Ho is with the Department of Pathology, University Medical Center Ho Chi Minh City, Ho Chi Minh City, Vietnam (e-mail: [email protected]). Simon Graham, Quoc Dang Vu, Mostafa Jahanifar, and Nasir Rajpoot are with Histofy Ltd, Coventry, United Kingdom (e-mail: {s.graham, qd.vu, m.jahanifar, n.rajpoot}@histofy.ai).
a total count with precise locations, allowing pathologists to inspect the underlying detections. These outputs complement weakly supervised models that predict slide-level clinical endpoints. We consider two metrics for mitotic activity. The mitotic hotspot count reports the highest MF count within a fixed area. The mitotic density denotes the total MF count per unit tissue area. A localized active focus can therefore produce a high hotspot count but a low overall density. With the increasing availability of digitized whole-slide images (WSIs), a growing number of datasets have been developed for mitotic figure detection across diverse cancers and organs [8]. However, these datasets remain limited in size and diversity due to the high cost of annotation, restricting their ability to support large-scale model training. Consequently, domain shift remains a major challenge, leading to degraded performance when models are applied to unseen cancer types or data sources [9]. In computational pathology, pretraining plays a key role in improving model generalization under limited annotations. Self-supervised learning (SSL) has emerged as an effective strategy by leveraging large-scale unlabeled histopathology data, through contrastive learning (e.g. SimCLR [10]), self-distillation, and masked image modeling (MIM) (e.g. MAE [11]). These methods have been adapted to pathology to train general-purpose feature extractors such as CTransPath-MoCov3 [12] and Virchow [13]. While these models generalize strongly across tasks, they are designed as generic feature extractors rather than optimized for a specific task such as MF analysis. Most pretrained models are used as frozen feature extractors with task-specific heads for downstream applications such as WSI classification and survival prediction [13], [14]. Further adaptation can improve performance [15], [16], for example through knowledge distillation [17]. However, discrepancies between pretraining objectives and downstream tasks introduce a domain gap that limits direct transfer. Moreover, most pretrained (foundation) models are trained on relatively large image patches (e.g., 112×112 µm), whereas MFs are typically assessed at much finer scales (e.g., 32×32 µm) [18]. This mismatch in resolution and context limits their effectiveness, as does the presence of numerous visually similar mimickers such as apoptotic cells and lymphocytes, which calls for task-specific rather than generic features. Recent pathology foundation models generalize well, but their size makes them expensive for dense inference tasks such as mitotic detection. With limited and imbalanced datasets, these challenges moti-
2
vate efficient, task-specific pretraining. To address these challenges, we propose MiTHras, a Mitotic figure analysis model based on Task-specific Hierarchical semi-supervised contrastive learning with a masked autoencoder. Through a single, unified objective, it integrates imagelabel contrast, which aligns the representations of crops that share a mitotic or mimicker pseudo-label, token-level contrast, which aligns local features using pseudo bounding boxes, and masked reconstruction that preserves image information. We evaluate MiTHras on MF classification, detection, count-based survival prediction, and subtype classification. In summary, our contributions are as follows: • MiTHras Framework: We introduce a task-specific hierarchical semi-supervised framework integrating pseudolabel-guided image- and token-level supervision with masked reconstruction. • Large-scale Dataset: We introduce TCGA-MF-Pseudo, a pseudo-labeled corpus that substantially broadens the diversity and scale of training data available for MF analysis. • Comprehensive Evaluation: We evaluate classification, detection, count-based survival prediction, and subtype classification against MF models, pretraining methods, and pathology foundation models under fine-tuning and linear probing settings. II. R ELATED WORK A. Mitotic figure detection Deep learning approaches for MF analysis include classification, object detection, segmentation, and two-stage pipelines. Methods often incorporate domain adaptation to address domain shift caused by variations in scanners, staining, and tissue types. Object detection-based methods have shown effectiveness in localizing MFs; [19], for instance, introduced a domain-adversarial RetinaNet to reduce domain shift. Segmentation-based methods instead leverage pixel-level predictions to capture fine-grained mitotic features: MaskMitosis [20] adopts Mask R-CNN with a ResNet-FPN backbone [21], and other works combine segmentation with Mask R-CNN for detection [22]. Two-stage models typically detect candidate regions and subsequently classify them. FoCasNet [23] used a feature pyramid network for detection and a ResNet [21] for classification, and [22] combined Mask R-CNN for segmentation with an ensemble of DenseNet and ResNet. MDFS [18] adopted Efficient-UNet-b0 for segmentation and EfficientNet for classification, achieving top performance in both MIDOG21 [8] and MIDOG22 [9] challenges. More recently, vision-language models have been applied to MF analysis: [24] paired MF images with descriptive captions to support captioning and visual question answering, and by integrating modality and scanner information demonstrated improved cross-domain adaptability. Despite these advancements, most existing methods heavily rely on supervised data across multiple domains and require domain alignment strategies. In this work, we propose a pretrained framework designed to learn from diverse cancer types and scanner sources, aiming to enhance the generalization ability in MF detection.
PREPRINT
B. Representation learning in computational pathology Hierarchical pathology models combine information across spatial scales. HIPT [25] stacks self-supervised transformers from patch to region to slide, CTransPath [12] captures multiscale features via windowed attention. Attention-based multiple-instance learning aggregates patch features for slidelevel prediction [26]. While these approaches capture broad tissue context, MiTHras focuses on representations within cell-centered patches, combining image-level class information with local token labels. Pretraining methods also differ in how they incorporate supervision. General-purpose encoders learn broadly transferable features, whereas supervised approaches like SupMAE [27] incorporate class labels during pretraining. MiTHras bridges these approaches, providing task-specific supervision at both image and token levels over a large, unannotated corpus. C. Masked autoencoders The core idea of MAEs is to learn robust representations by reconstructing images from masked or partially observed patches. Inspired by masked language modeling in BERT [28] and by vision transformers (ViTs) [29], MIM [11], [30] has gained attention as a generative SSL approach. MAE reconstructs raw pixels through an asymmetric encoder-decoder [11], and DiffMAE [30] extends this with iterative diffusion processes. D. Supervised contrastive learning Contrastive learning typically models the similarity and dissimilarity of two augmented versions of an image, as shown in SimCLR [10] or between images and their corresponding captions, as in CLIP [31]. In standard contrastive learning, two views of an image form a positive pair. SupCon [32] reformulated contrastive learning into multi-positive pair contrastive learning by defining all images with the same label as positive samples. SupMAE [27] further extended the supervised setting into MIM by incorporating label-based supervision and reconstruction-based self-supervision. E. Semi-supervised learning Semi-supervised learning leverages large amounts of unlabeled data alongside limited annotations. Common strategies include pseudo-labeling [33] and consistency regularization [34], which encourage stable predictions under perturbations. Recent approaches combine these techniques with confidencebased filtering. For instance, FixMatch [35] combined pseudolabeling with confidence thresholds and consistency regularization, using a fixed threshold to retain only high-confidence predictions. FlexMatch [36] extended this by dynamically adjusting the confidence threshold per class, and SimMatch [37] added representation-level consistency, encouraging similar embeddings for augmented views of the same image. MiTHras uses a teacher model trained on annotated MFs to assign pseudo labels to an unlabeled pretraining corpus. These pseudo labels supervise contrastive learning at both image and token levels–a process we term semi-supervised contrastive learning (SCL)–alongside self-supervised masked reconstruction.
VUONG et al.: MITHRAS: TASK-SPECIFIC HIERARCHICAL SEMI-SUPERVISED CONTRASTIVE MASKED AUTOENCODER FOR MITOTIC FIGURE ANALYSIS
3
TABLE I: Summary of notation. Symbol
Description
N, T V, M K DE , DP α, β
Images per batch; tiles per image Visible and masked tiles per image (V + M = T ) Visible tokens per multiviewed batch (K = 2N V ) Encoder (768) and projection (512) dimensions Masking ratio; weight of the image-level SCL term
f enc , f dec f img , f tok f head f cls , f seg f det f cls-br , f cls-morph
ViT encoder and MIM decoder Image-level and token-level projection heads Linear classification head Classification (f head ◦ f enc ) and segmentation models Detection model, f det = f cls ◦ f seg Subtype and morphology classification models
p, t, h, z T, Z ŷ, ŷ tok P (u) sseg , scls i, u, j, k, w q
Tile, token, encoder output, projected embedding Set of tokens; set of projected embeddings Image-level and token-level pseudo labels Positive set of anchor u: samples sharing its pseudo label Segmentation and classification confidence of a candidate Image, view, tile, token, and batch-token indices Index over the positive set in the contrastive sums
A. Pseudo-label dataset construction
Fig. 1: Overview of the MiTHras framework. (A) Pseudolabeled dataset construction from WSIs, generating mitotic and mimicker samples. (B) Pretraining with a masked autoencoder and hierarchical contrastive learning at image and token levels. (C) Two-stage mitotic detection via segmentation and classification. (D) WSI-level mitotic counting. (E) Mitotic figure subtype classification.
III. M ETHOD
In this section, we introduce MiTHras, a task-specific hierarchical semi-supervised contrastive masked autoencoder for MF analysis, illustrated in Figure 1. We first construct the pseudo-labeled dataset TCGA-MF-Pseudo (Figure 1A) and pretrain the encoder f enc on it with a masked autoencoder and task-specific contrastive objectives (Figure 1B). On that encoder we build a classification model f cls = f head ◦ f enc for MFs versus mimickers (Figure 1C-right) and for subtypes (Figure 1E); a two-stage detection model that pairs a segmentation model f seg proposing candidates with f cls (Figure 1C); and a whole-slide mitotic count used for survival prediction (Figure 1D). Table I summarizes the notation. Upright superscripts (e.g., enc, img, tok) denote module or level identifiers, italic capitals denote cardinalities, and bold symbols denote vectors or sets of vectors.
General-purpose pathology pretraining commonly uses tissue patches sampled without selecting a particular cell type. MF analysis instead requires discrimination of individual MFs from visually similar non-mitotic objects, or mimickers. We construct TCGA-MF-Pseudo to provide cell-centered crops sourced from 14 TCGA cohorts spanning 11 organ sites, with variations in tumor type, scanner, and staining protocols. The construction procedure has three stages: 1) Patch Extraction: We extract tissue patches of size 896×896 pixels using Otsu’s method [38]. Each extracted tile provides tissue context and is subdivided into a 4 × 4 grid of 224 × 224 patches for tumor filtering. The same tile also accommodates the 512 × 512 crops used by the segmentation stage. 2) Tumor Patch Selection: We select tumor patches if all 4×4 patches of size 224×224 pixels are predicted as tumors using IMPaSh [39], a model classifying histology textures into adipose, background, debris, lymphocytes, mucus, smooth muscle, normal colon mucosa, cancer-associated stroma, and colorectal adenocarcinoma epithelium (tumor). 3) Pseudo-Label Assignment. We detect and classify MF and mimicker candidates with the two-stage model of [18] trained on MIDOG22, replacing its stage-2 classifier with a ViT-base pretrained by MAE on ImageNet. From each refined centroid a 50×50 pixel pseudo bounding box is generated, following the MIDOG22 procedure. Candidates are selected via segmentation (sseg ) and classification (scls ) scores. Pseudo MFs require sseg ≥ 0.1 and scls > 0.5 (660,831 crops), whereas a pseudo mimicker requires sseg ≥ 0.5 and scls < 0.5 (1,133,624 crops). Candidates with 0.1 ≤ sseg < 0.5 and scls < 0.5 are excluded. The lower segmentation threshold retains potential MFs, relying on the classifier to filter false positives. The higher threshold for mimickers selects candidates with strong segmentation responses that the classifier labels as non-mitotic. Pseudo-label quality is evaluated in Section V-F. Following this procedure, we obtain a task-specific MF pretraining dataset, referred to as TCGA-MF-Pseudo, which comprises 660,831 pseudo MF images and 1,133,624 pseudo
4
mimicker images. Figure 2 illustrates the proportions of the datasets used in this study. Compared to other datasets, TCGAMF-Pseudo is not only larger in terms of MFs and mimicker images but also covers a broader range of tumor types and data sources. To assess the effectiveness of task-specific pretraining, we also construct an unlabeled pathology pretraining dataset, referred to as TCGA-Unlabeled, following the standard SSL approach in computational pathology. This dataset consists of 8 million patches (224×224 pixels) extracted from the same 14 cohorts at multiple magnifications: 3 million patches at 40×, 2.7 million patches at 20×, 1.2 million patches at 10×, and 1.1 million patches at 5× magnification. We note that IMPaSh was trained to recognize colorectal tissue categories and is applied here across multiple organ sites. Its role is to select tissue regions, while the separate MIDOG22-trained detector and classifier assign MF and mimicker labels. B. MiTHras pretraining Figure 1B details the architecture of the MiTHras pretraining procedure, which comprises five main components: 1) Augmentation and Masking: Each image is randomly augmented and partially masked to generate two different views composed of visible and masked image tiles. 2) Encoder (f enc ): The visible tiles are passed through a ViT encoder to generate latent feature representations, including a [CLS] token that summarizes the image-level information. 3) SCL-global module: An image-level projection head f img extracts global embeddings from the encoder output, which are trained to align with others from the same pseudolabel class while being distinct from those of different classes. 4) SCL-local module: A separate token-level projection head f tok captures fine-grained, tile-level features. Tokens sharing the same pseudo-label are encouraged to cluster together. 5) Decoder (f dec ): A lightweight decoder reconstructs the original image from visible and masked tokens, using a pixelwise loss to enforce reconstruction quality. The following subsections describe each component in detail and explain how they are integrated within the MiTHras framework. 1) Image masking and encoding: Suppose that we are given a training batch B = (X, Ŷ ) from TCGA-MF-Pseudo where bbox N X = {xi }N )}i=1 i=1 denotes a set of images, Ŷ = {(ŷi , ŷi represents their corresponding pseudo labels, ŷi denotes the pseudo class label for the image xi , ŷibbox is the set of pseudo bounding boxes for MFs in xi , if any are present, and N is the batch size. Following the patching process in ViT [29], an image xi is divided into a set of non-overlapping tiles {pi,j }Tj=1 , where T is the total number of tiles per image. A large portion of these tiles, determined by a masking ratio α, is randomly masked out (i.e., removed from the input), resulting in a subset of visible tiles {pvi,j }Vj=1 and a subset of masked tiles M {pm i,j }j=1 , where V is the number of visible tiles, M is the
PREPRINT
number of masked tiles, V + M = T , and α is set to 0.75, following [11]. Each visible tile pvi,j is flattened and projected into a DE -dimensional embedding space, producing visible tokens {tvi,k }Vk=1 . Positional embeddings are added to these visible tokens, and a special [CLS] token tcls i is appended to v v these tokens, resulting in the input tokens {tcls i , ti,1 , · · · , ti,V }. These input tokens are fed into the encoder to produce their lav v v V tent representations f enc ({tcls i , ti,1 , · · · , ti,V }) = {hi,k }k=0 ∈ (V +1)×DE v R , where DE = 768, hi,0 is the embedding for the [CLS] token, and {hvi,k }Vk=1 are the visible token embeddings. 2) Hierarchical semi-supervised contrastive learning branch: The hierarchical SCL branch uses pseudo labels to define which representations should be similar. Image-level labels distinguish MF from mimicker crops. Token level labels derived from pseudo bounding boxes separate putative MF regions from surrounding tissue. These two objectives encourage class separation at different spatial scales without requiring manual annotation of the pretraining corpus. Each image xi is randomly augmented twice to form a multiviewed batch that contains 2N augmented images indexed by u ∈ U ≡ {1, · · · , 2N }. Each view xu is masked out and encoded by f enc (Section III-B1), and the resulting hidden states {hvu,k }Vk=0 are mapped by a projection head into a DP dimensional space (DP = 512), where index k = 0 denotes the [CLS] position and k ≥ 1 the visible tokens. The head is instantiated separately at the two levels, as f img and f tok below. Given a collection of ℓ2 -normalized embedding vectors {zu }u∈U , one per element to be contrasted, a semi-supervised contrastive loss LSCL can be defined as follows: LSCL (X, Ŷ ) =
X 0.1 X −1 log 2N |P (u)| u∈U
q∈P (u)
exp(zu · zq /τ ) P exp(zu · za /τ ) a∈A(u)
(1) where A(u) ≡ U \ {u}, P (u) ≡ {q ∈ A(u) : ŷq = ŷu } is the set of positive indices sharing the same pseudo label as xu , |P (u)| is the cardinality of P (u), and τ = 0.1 is the temperature. The loss averages over anchors and includes a fixed multiplier of 0.1. Both the image- and tokenlevel objectives use this temperature, scaling, and embedding normalization. Building upon LSCL , we define a hierarchical contrastive loss that combines image-level (Limg SCL ) and tokenlevel (Ltok ) objectives. SCL 3) Image-level semi-supervised contrastive learning: To improve the quality of latent representations at the global (image) level, we introduce image-level SCL. Specifically, given an image view xu , we adopt an image-level projection head f img to produce the embedding vectors Zuimg = img V {zu,k }k=0 = {f img (hvu,k )}Vk=0 ∈ R(V +1)×DP as described in Section III-B2. Using the embedding vectors, we define the image-level contrastive loss (Limg SCL ) given by: Limg SCL (X, Ŷ ) =
X 0.1 X −1 log 2N |P (u)| u∈U
q∈P (u)
exp(zucls · zqcls /τ ) P exp(zucls · zacls /τ ) a∈A(u)
(2) img where zucls = zu,0 is the projected embedding vector for the [CLS] token.
VUONG et al.: MITHRAS: TASK-SPECIFIC HIERARCHICAL SEMI-SUPERVISED CONTRASTIVE MASKED AUTOENCODER FOR MITOTIC FIGURE ANALYSIS
4) Token-level semi-supervised contrastive learning: At the local (token) level, we promote fine-grained feature learning by applying an SCL loss to individual visible tokens. As described in Section III-B1, each image view xu in batch B is tokenized into a set of visible tokens {tvu,k }Vk=1 . Similar to the imagelevel SCL, we employ a token-level projection head f tok to tok V obtain projected embeddings {zu,k }k=0 = {f tok (hvu,k )}Vk=0 ∈ (V +1)×DP R . We then discard the embedding corresponding to the [CLS] token and retain only the projected visible token tok V embeddings, denoted as Zutok = {zu,k }k=1 . For MF images, tokens are labeled from the pseudo bounding boxes ŷubbox ; for mimicker images all tokens are non-mitotic. Pretraining inputs are extracted as 128×128 crops, with the central 50×50 region forming the bounding box. Augmentation proceeds in two stages. The spatial stage begins with a randomly resized crop to 224 × 224 pixels with the cropped area fraction uniformly sampled from [0.2, 1.0], followed by random horizontal and vertical flips. These spatial transformations are applied jointly to both the image and the bounding box. Then, the photometric stage applies color jitter, random grayscale, Gaussian blur, and solarization for the second view to the image alone. A token is labeled mitotic if its corresponding patch cell lies inside the box, with box edges rounded to the nearest patch boundary. No separate boundary class is used. For a square crop, the sampled area fraction corresponds to a linear magnification of approximately 1.75–3.9×. Before boundary rounding, a fully retained 50 × 50 box occupies 30– 149 patch-cell areas in the 14 × 14 token grid. Actual token labels depend on the crop position, aspect ratio, and boundary rounding; partial cropping reduces box coverage, and a crop that excludes the box contains no mitotic tokens. The fixed box may include adjacent tissue or exclude larger figures, yielding only coarse spatial supervision. This results in a set of visible tokens with corresponding pseudo labels for the entire training batch B, denoted as B tok = (T , Ŷ tok ) where T = {tvw }K w=1 tok K denotes all visible tokens in B, Ŷ tok = {ŷw }w=1 denotes tok their corresponding pseudo labels, ŷw denotes a pseudo label v for the corresponding token tw , and K = 2 × N × V is the total number of visible tokens in B. Using these tokens and their corresponding pseudo labels, we compute the token-level SCL (Ltok SCL ), defined as: tok Ltok )= SCL (T , Ŷ
−1 0.1 X K |P tok (w)| w∈W
MIM decoder f dec that predicts the pixel content of the masked tiles. The reconstruction xru comprises the visible tiles pvu = rec M {pvu,j }Vj=1 and the recovered masked tiles prec u = {pu,j }j=1 . dec To optimize f , we minimize the mean squared error (MSE) between reconstructed masked tiles prec u,j and their normalized pixel targets p̃m u,j . The MSE averages over the pixel and channel values within each tile, and the reconstruction loss averages over the set M of masked tiles: m LMIM (X r , X) = mean MSE prec u,j , p̃u,j . (u,j)∈M
(4)
6) MiTHras loss function: Overall, the proposed MiTHras framework is optimized using the following objective function: L(X, X r , T , Ŷ , Ŷ tok ) = LMIM (X r , X) + βLimg SCL (X, Ŷ ) tok + (1 − β)Ltok ). SCL (T , Ŷ
(5) The parameter β weights image-level SCL, while 1 − β weights token-level SCL and the reconstruction term has unit weight. The main MiTHras configuration uses β = 0.75. The ablation at β = 1 combines masked reconstruction with imagelevel SCL and excludes token-level SCL; β = 0 combines reconstruction with token-level SCL and excludes image-level SCL. 7) Architecture: The encoder f enc is a standard ViT-base [29]: 12 transformer blocks, embedding dimension DE = 768, 12 attention heads. The MIM decoder f dec has 8 blocks with dimension 512. The image and token projection heads f img and f tok share one configuration, a fully connected layer mapping DE - to DP -dimensional embeddings for both the class token and the tile tokens. Downstream, only f enc is retained; f dec , f img and f tok are discarded. C. Mitotic figure classification model The classification model f cls serves both the mitotic-versusmimicker task and subtype classification. We append a classification head f head to the encoder f enc , giving y ′ = f cls (x) = f head (f enc (x)). The encoder is initialized from the MiTHras pretrained weights and the head with zero weights and a constant bias; all parameters are fine-tuned end-to-end with cross-entropy loss. D. Mitotic figure detection model
X q∈P tok (w)
tok exp(zw · zqtok /τ ) P log tok · z tok /τ ) exp(zw a a∈Atok (w)
5
(3)
where W ≡ {1, · · · , K}, Atok (w) ≡ W \ {w}, P tok (w) ≡ tok {q ∈ Atok (w) : ŷqtok = ŷw } is the set of positive indices sharing the same pseudo label as tvw , and |P tok (w)| is the cardinality of P tok (w). 5) Reconstruction branch: We add a pixel-level MIM branch. For an image view xu , masked tile positions are M replaced by learnable mask tokens {tm u,j }j=1 [28], decoderspecific positional embeddings are added to the visible and mask tokens, and the combined set is passed to a lightweight
The detection model follows the two-stage pipeline of Jahanifar et al. [18]. A segmentation model proposes candidate locations, and the MF-versus-mimicker classifier evaluates a crop around each candidate. We write this pipeline as f det = f cls ◦ f seg , with candidate extraction and cropping included between the two stages. In the first stage, we use an Efficient-UNet-b0 segmentation model f seg to process an input image and generate a probability map, highlighting candidate regions for mitotic figures. Candidate MFs are then identified by applying morphological post-processing to this map. For each candidate, the centroid is extracted, and the mean intensity within the region is used as the detection score sseg . In the second stage, each candidate is classified using the mitotic classification model
6
PREPRINT
MIDOG22-mitotic figure
MIDOG22-mimicker
TCGA-MF-Pseudo-mitotic figure (high confidence)
TCGA-MF-Pseudo-mimicker (high confidence)
TCGA-MF-Pseudo-mitotic figure (low confidence)
TCGA-MF-Pseudo-mimicker (low confidence)
(a) Representative Mitosis, Mimicker, and Pseudo Samples
(b) Dataset Composition and Distribution
Fig. 2: (a) Representative mitosis, mimicker, and pseudo samples. (b) Dataset distribution shown as a bubble chart, where bubble size reflects the number of images and labels indicate domain count and dataset size. TCGA-MF-Pseudo is the largest dataset, spanning 14 domains (∼1.8M images). f cls , as described in Section III-C, yielding a classification score scls . The stages are trained independently. For each twostage method, a single pair of segmentation and classification thresholds is selected by maximizing F1 across the three MIDOG22 validation folds and is then held fixed for external evaluation. The candidate detector is shared across secondstage classifiers. During inference, the segmentation model processes 512 × 512 windows with a stride of 400 pixels, and overlapping probability maps are averaged. The 112-pixel overlap reduces boundary effects while limiting repeated computation. This setting follows the MDFS protocol [18]; its approximate processing redundancy is (512/400)2 = 1.64, compared with 4 for a half-window stride. E. MiTHras mitotic count for WSIs We apply the two-stage detector to WSIs and summarize the detections using a hotspot count in 2 mm2 and a density expressed per 20 mm2 . The hotspot summarizes the most active sampled region, whereas density summarizes average activity across the analyzed tissue. These count-derived features are evaluated for survival prediction on TCGA-THCA and TCGAGBMLGG. For both the mitotic hotspot (2 mm2 ) and mitotic density (20 mm2 ) strategies, we begin by applying Otsu [38] thresholding to extract tissue patches from a WSI. The patches are of size 512×512 pixels acquired at 40× magnification (0.25 µm per pixel). Each patch is then processed using the two-stage mitotic detection pipeline to identify MFs. 1) Mitotic hotspot (2 mm2 ): Local Max ROI Counting (LMRC) centers a square of area 2 mm2 on each detected MF and returns the largest count among these candidate regions. Detected coordinates are indexed with a KD-tree (cKDTree, SciPy v1.11.2). A Chebyshev-distance query √ retrieves points within each square using a half-side length of 2/2 mm. The procedure avoids evaluating a dense sliding grid; its search is restricted to detection-centered squares. 2) Mitotic density (20 mm2 ): Mitotic density expresses the total detected MF count per unit of analyzed tissue area, with
20 mm2 used as the reporting reference. Unlike the hotspot count, it summarizes average activity across the sampled tissue. F. Mitotic figure subtype classification model Similar to the MF classification model f cls , we build two subtype classification models: f cls-br for the binary classification of mitotic figure subtypes and f cls-morph for the 8-class morphological subtype classification. IV. E XPERIMENTS A. Datasets We evaluate four tasks, mitotic figure classification, detection, survival prediction from mitotic count, and subtype classification, across the datasets detailed in Table II and Figure 2. The TCGA evaluation cancer types (GBM, SARC, DLBC, THCA, and LGG) are excluded from TCGA-MFPseudo. TCGA pretraining and evaluation therefore share no slide, patient, or cancer type. TCGA metadata show no overlap in tissue source site codes, although 20 of the 33 parent institutions contributing to the evaluation cohorts also contribute to the pretraining cohorts. The TCGA split therefore evaluates transfer across cancer types but does not establish independence of institutional acquisition practices. MIDOGpp, Canine, AMi-Br-TUPAC, and AMi-Morph-TUPAC provide additional external evaluations across laboratories, scanners, and, for Canine, species. 1) Pretraining and training data: TCGA-MF-Pseudo (Section III-A) provides 660,831 pseudo MFs and 1,133,624 pseudo mimickers for pretraining. For supervised training we adopt MIDOG22 [9], containing 9,501 MFs and 11,051 mimickers from 354 images across five tumor types: 150 human breast carcinomas, 44 canine lung carcinomas, 55 canine lymphosarcomas, 50 canine cutaneous mast cell tumors, and 55 human neuroendocrine tumors. Following the threefold MDFS training protocol [18], we train Efficient-UNetb0 for candidate segmentation and reuse this segmentation
VUONG et al.: MITHRAS: TASK-SPECIFIC HIERARCHICAL SEMI-SUPERVISED CONTRASTIVE MASKED AUTOENCODER FOR MITOTIC FIGURE ANALYSIS
stage for all second-stage classifier comparisons. The three train/validation splits contain 239/115, 234/120, and 235/119 source images, respectively, with no image overlap between training and validation within a split. We extract 512 × 512 pixel patches for segmentation and 128 × 128 pixel crops centered on MFs and mimickers for classification, and select the best classification checkpoint per fold by validation performance. Each fold is trained and evaluated independently. Results are reported as the mean ± standard deviation across three folds, without ensembling predictions. All methods use the same folds. The standard deviation describes variation across training folds. 2) Tasks 1 and 2: classification and detection: Both tasks are evaluated on MIDOGpp, Canine, and TCGA-MF-Test. MIDOGpp contains 2,436 MFs and 3,300 mimickers from 100 canine soft tissue sarcoma and 49 human melanoma images. Canine includes 7,206 MFs and 3,082 mimickers from 160 canine cutaneous mast cell tumor and 105 canine mammary tumor images [40], [41]. TCGA-MF-Test comprises 1,539 MFs and 2,073 mimickers from TCGA glioblastoma multiforme (GBM), soft tissue sarcoma (SARC), and lymphoid neoplasm diffuse large B-cell lymphoma (DLBC). For classification, we extract 128 × 128 pixels centered on annotated centroids. Evaluation metrics are defined in Section IV-C. For detection, we use full images: segmentation probability maps are produced via a 512×512 sliding window at stride 400 and aggregated by averaging overlaps. Peak points define candidate locations, from which 128 × 128 patches are extracted and classified. Predictions are matched to ground truth within a 30-pixel radius and scored by F1 . 3) Task 3: Survival prediction using mitotic count: Given the role of mitotic count in grading high-grade follicular-cellderived thyroid carcinoma [2] and diffuse glioma, we evaluate its association with overall survival across two cohorts: TCGATHCA (406 thyroid WSIs from 395 patients, including 11 deaths) and TCGA-GBMLGG (844 brain WSIs from 501 patients, including 122 deaths). The slides were scanned at 40× magnification. Survival time is measured in months, with death coded as the event (event = 1 − censorship). 4) Task 4: Mitotic figure subtype classification: We evaluate generalization on two subtype benchmarks [42]: AMi-Br (typical vs. atypical) and AMi-Morph (eight morphology classes), used to evaluate f cls-br and f cls-morph respectively. Three-fold cross-validation is performed on AMi-Br-MIDOG21 and AMiMorph-MIDOG21. Each fold model is applied independently to the held-out AMi-Br-TUPAC and AMi-Morph-TUPAC domains, and results are reported as the mean ± standard deviation over the three folds. AMi-Br-MIDOG21 contains 1,721 figures (1,317 typical, 404 atypical); AMi-Morph-MIDOG21 adds eight-way annotations to the same images from two experts, of which we use expert 1 (four typical subtypes: prometaphase, metaphase, ring-shaped, anaphase-telophase; four atypical: bipolar asymmetric, multipolar, segregationrelated, other), with the rarest class holding 18 examples. The test sets are AMi-Br-TUPAC (1,999 figures; 1,571 typical, 428 atypical) and AMi-Morph-TUPAC (1,999 figures across the same eight subtypes). AMi-Br is scored by positive-class F1 for atypical figures, and AMi-Morph by macro F1 over the
7
TABLE II: Training and test datasets for the MiTHras models. The classification encoders are pretrained on TCGAMF-Pseudo and then fine-tuned on the listed training sets. Detection and mitotic counting share the two-stage pipeline f det = f cls ◦ f seg , trained on MIDOG22. Task
Train / Test
MF vs. mimicker classification MF detection Mitotic counting on WSIs MF subtype (typical/atypical) MF morphology (8 classes)
MIDOG22 / MIDOGpp, Canine, TCGA-MF-Test MIDOG22 / MIDOGpp, Canine, TCGA-MF-Test MIDOG22 / TCGA-THCA, TCGA-GBMLGG AMi-Br-MIDOG21 / AMi-Br-TUPAC AMi-Morph-MIDOG21 / AMi-Morph-TUPAC
morphology classes. B. Implementation details We trained all the pretraining models for 200 epochs on 8 NVIDIA A6000 GPUs. MiTHras was initialized from scratch and optimized with AdamW (betas 0.9 and 0.95, weight decay 0.05), an effective batch size of 512, and a learning rate of 3 × 10−4 , with 10 warm-up epochs followed by cosine decay. The configuration without MIM used the same image/token weights (0.75/0.25) and optimizer schedule, with an effective batch size of 768 and learning rate of 4.5 × 10−4 . The masking ratio was fixed at α = 0.75, following MAE [11], while the ablations varied β to assess the balance of the image- and token-level SCL objectives. During pretraining, two augmented views were generated per image, following the procedure described in Section III-B4. Views were normalized with ImageNet statistics. All classification models were finetuned for 50 epochs on a single NVIDIA A6000, with RandAugment plus random erasing, 224 × 224 inputs, ImageNet normalization, label smoothing (ϵ = 0.1), AdamW (weight decay 0.05, layer-wise decay 0.75), base learning rate 2.5×10−4 with five warm-up epochs, and automatic mixed precision. The segmentation model was trained separately for 50 epochs on 2 A6000 GPUs at 512 × 512 with standard augmentation and optimization. For linear probing, the frozen pretrained encoder was attached to a classification head consisting of an affine-free batch normalization layer followed by a trainable linear layer. The head was trained using LARS with a base learning rate of 0.1, zero weight decay, 10 warm-up epochs, cosine decay, and cross-entropy loss. Images were resized to 224×224 pixels with weak augmentation and ImageNet normalization. Affine-free batch normalization standardized feature scales across backbones. All methods used the same three folds and 50-epoch training budget for linear probing and fine-tuning. C. Evaluation metrics and statistical analysis We evaluate performance using accuracy (ACC), precision, recall, F1 score, and area under the receiver operating characteristic curve (ROC-AUC). For mitotic-versus-mimicker classification, MFs are the positive class. For AMi-Br, atypical figures (label 1) are positive and typical figures (label 0) are negative. Binary precision, recall, and F1 are calculated for the positive class. For AMi-Morph, these metrics are calculated separately for each class occurring in the targets or predictions
8
and averaged with equal class weights (unweighted macro averaging); undefined class scores are set to zero. ROC-AUC is not evaluated for AMi-Morph. Detection metrics use predicted locations matched to annotated MFs within a 30-pixel radius. Survival prediction is evaluated using Harrell’s concordance index (C-index), which measures agreement between predicted risk and observed event ordering among comparable pairs, accounting for censoring. Cox proportional-hazards models are fitted and evaluated within each full cohort. Patients with multiple slides contribute multiple observations. 95% confidence intervals are calculated using 1,000 WSI resamples with fitted risk predictions held fixed, i.e., without refitting the Cox models. Log-rank tests compare groups defined by a cutpoint fixed at the median of the designated training subset. A likelihood-ratio test compares a model containing age and sex with one additionally containing the mitotic-count feature. Reported p-values are unadjusted. Classification, detection, and subtype results are reported as the mean and sample standard deviation (SD) across three training folds. The blinded-review protocol (Section IV-D6) specifies precision and uncertainty estimation for the pseudolabel audit. The main tables report F1 for classification, detection, and subtype classification, and the C-index for survival prediction. Additional classification metrics are provided in Supplementary Table S1. Detection and subtype results are in Supplementary Table S2. D. Comparative experiments and ablation study 1) Classification task: Table III compares MiTHras, a ViT-base pretrained on the IMPaSh-selected TCGA-MFPseudo corpus, with the MDFS classifier [18], generalpurpose encoders [31], [43], pathology-pretrained encoders [12], [13], [44]–[48], and alternative pretraining objectives [27], [49]. ViT-ImageNet and ViT-TCGA-MF-Pseudo use supervised classification pretraining on their named corpora. MAE-ImageNet, MAE-TCGA-Unlabeled, and MAE-TCGAMF-Pseudo use masked reconstruction on their named corpora. CMAE uses contrastive masked-autoencoder pretraining on TCGA-MF-Pseudo, while SupMAE uses labeled MIDOG22 images. All encoders are evaluated under full fine-tuning and frozen-encoder linear probing (described in Section IV-B). 2) Detection task: Two-stage detection models follow the MDFS pipeline [18] and are compared with standalone Efficient-UNet-b0 and YOLOv10 [50] baselines. EfficientUNet-b0 is trained under the three-fold MDFS protocol and shared across the two-stage classifier comparisons. The compared methods differ in their second-stage classifier. Each classifier is the fine-tuned model from the classification experiment and receives 224 × 224 inputs. 3) WSI survival prediction: The two-stage approaches share tissue extraction, the candidate detector, and count computation, and differ in their second-stage classifiers. A segmentation-only baseline computes counts without secondstage classification. 4) Mitotic figure subtype classification task: Subtype experiments compare the pretrained encoders listed in Table VI on the AMi benchmarks introduced by Bertram et al. [42].
PREPRINT
Each encoder receives a task-specific classification head and is fine-tuned using the shared protocol in Section IV-B. The ablation uses the same protocol. 5) Ablation experiments: To assess the contribution of each component, we evaluated masked reconstruction and imagelevel and token-level SCL individually and in combination (Table VII). We compare β = 0.75 and β = 1 under full fine-tuning, linear probing, and subtype classification for three settings: ViT-base and ViT-tiny pretrained on TCGAMF-Pseudo, and ViT-base pretrained on labeled MIDOG22. The same evaluation protocols are used within each paired comparison (Table VII). 6) Blinded pathological review: A single pathologist (P.T.H.H.) conducted a blinded review to assess tissue selection and pseudo-label quality across all 14 cohorts. Only the images were provided, and the pipeline labels and confidence scores were concealed. For tissue selection, IMPaSh and CONCH [51] were applied to 10,000 sampled patches per cohort. The reviewer examined 25 patches per cohort from each of three strata: retained by both filters, retained only by IMPaSh, and retained only by CONCH, yielding 1,050 reviewed patches. For pseudo-label quality, the audit included 700 crop reviews, comprising 420 pseudo-mitotic and 280 pseudo-mimicker reviews (30 and 20 per cohort, respectively) from the final pretraining corpus. Each crop was classified as confirmed, not confirmed, or unsure. Precision ranges count unsure verdicts as incorrect at the lower endpoint and correct at the upper endpoint. Per-cohort proportions are reported with Wilson confidence intervals. Tumor-filter stratum estimates pool the reviewed patches; pooled pseudomitotic precision is also reported with weights proportional to each cohort’s pseudo-mitotic pool size. The additional-figure assessment included 199 pseudo-mimicker crops whose central candidate was classified as non-mitotic. For each crop, the reviewer recorded whether an MF was present elsewhere in the crop. For bounding-box assessment, MFs confirmed during the crop review were assigned to the smallest enclosing square category: 25, 50, 75, or larger than 75 pixels. No additional crops were sampled for this assessment. V. R ESULTS A. Mitotic figure classification results MiTHras attained the highest mean F1 across all three test sets (Table III). Under full fine-tuning, its margins over the strongest baseline, Virchow, were 0.016 on MIDOGpp, 0.011 on Canine, and 0.002 on TCGA-MF-Test, though the latter difference was small relative to the reported fold variability. This gap widened under linear probing. MiTHras achieved F1 scores of 0.836, 0.848, and 0.775 on MIDOGpp, Canine, and TCGA-MF-Test, respectively, with reductions of 0.007– 0.028 from full fine-tuning. On MIDOGpp, the best generalpurpose or pathology foundation encoder reached 0.421, while all encoders exceeding 0.45 were pretrained on labeled or pseudo-labeled MF data. These results indicate that MiTHras features are more readily applicable without complex tuning. The effect of the pretraining corpus varies with the objective and evaluation protocol. Supervised ViT pretraining on TCGAMF-Pseudo produced substantially higher linear-probing F1
VUONG et al.: MITHRAS: TASK-SPECIFIC HIERARCHICAL SEMI-SUPERVISED CONTRASTIVE MASKED AUTOENCODER FOR MITOTIC FIGURE ANALYSIS
9
TABLE III: Mitotic figure classification results (F1 , mean ± standard deviation over three folds) under full fine-tuning and linear probing. Each fold is trained and evaluated independently, without ensembling predictions. ‡ SupMAE is pretrained on labeled MIDOG22 images. Best value per column in bold. Pretraining Method
Fine-tuning
Network MIDOGpp
Canine
EfficientNet-b7 0.831±0.003 0.815±0.010 MDFS ViT-ImageNet (Supervised) ViT-base 0.824±0.011 0.832±0.006 Natural-image pretraining MAE-ImageNet ViT-base 0.840±0.006 0.840±0.009 DINOv2 ViT-base 0.840±0.006 0.842±0.007 CLIP ViT-base 0.829±0.006 0.834±0.003
Linear Probing TCGA-MF-Test
MIDOGpp
0.757±0.010 0.766±0.004 0.775±0.007 0.778±0.012 0.773±0.005
0.356±0.309 0.708±0.103 0.354±0.306 0.523±0.321 0.318±0.075 0.614±0.059 0.026±0.024 0.545±0.472 0.235±0.311 0.440±0.412
Canine
TCGA-MF-Test 0.570±0.027 0.155±0.059 0.376±0.033 0.442±0.137 0.512±0.025
Histopathology pretraining models
CTransPath-MoCov3 CHIEF Lunit-DINO PLIP MAE-TCGA-Unlabeled SupMAE‡ UNI-DINOv2 GPFM Virchow
Swin-tiny Swin-tiny ViT-small ViT-base/32 ViT-base ViT-base ViT-large ViT-large ViT-huge
0.808±0.013 0.821±0.011 0.818±0.011 0.822±0.008 0.816±0.004 0.829±0.015 0.819±0.009 0.820±0.008 0.806±0.017 0.826±0.009 0.755±0.015 0.800±0.026 0.837±0.009 0.835±0.009 0.794±0.016 0.835±0.003 0.848±0.004 0.844±0.006
0.736±0.007 0.735±0.003 0.763±0.010 0.752±0.007 0.749±0.005 0.671±0.029 0.776±0.008 0.751±0.007 0.787±0.002
0.199±0.343 0.603±0.063 0.359±0.201 0.517±0.419 0.344±0.277 0.353±0.413 0.392±0.326 0.465±0.405 0.377±0.071 0.698±0.064 0.624±0.009 0.735±0.012 0.364±0.215 0.624±0.342 0.421±0.211 0.701±0.125 0.300±0.244 0.476±0.406
0.548±0.069 0.355±0.210 0.482±0.155 0.470±0.161 0.466±0.074 0.592±0.012 0.445±0.107 0.393±0.040 0.266±0.282
Pretraining methods using our pseudo-labeled dataset: TCGA-MF-Pseudo
ViT-TCGA-MF-Pseudo MAE-TCGA-MF-Pseudo CMAE MiTHras (Ours)
ViT-base ViT-base ViT-base ViT-base
0.825±0.005 0.823±0.006 0.827±0.006 0.840±0.005 0.838±0.005 0.824±0.008 0.864±0.004 0.855±0.001
0.766±0.004 0.746±0.007 0.770±0.009 0.789±0.011
0.824±0.001 0.821±0.005 0.323±0.100 0.743±0.045 0.453±0.059 0.616±0.036 0.836±0.003 0.848±0.005
0.766±0.004 0.548±0.021 0.546±0.055 0.775±0.002
both protocols, supporting the contribution of task-specific SCL beyond masked reconstruction. Among pathology foundation models, Virchow attained the highest fine-tuned F1 on all three datasets, outperforming smaller encoders such as CTransPath-MoCov3 [12] and CHIEF [44]. B. Mitotic figure detection results
Fig. 3: Performance and inference time at three evaluation scales. Circles denote classification of a single 128×128 patch; triangles denote detection in a Canine ROI (approximately 42 candidates); squares denote whole-slide inference on TCGATHCA (approximately 2,305 candidates). The horizontal axis shows inference time, and the vertical axis shows F1 for patch and ROI evaluation and the Cox C-index for WSI evaluation. Performance values are comparable within each evaluation scale. Dashed lines connect the same model across scales; the inset enlarges the patch and ROI region. Times were measured on one NVIDIA RTX A6000 GPU with mixed precision.
than supervised ImageNet pretraining on all three datasets. For MAE, TCGA-MF-Pseudo improved linear probing over ImageNet universally, but fine-tuning results were mixed, with lower F1 on MIDOGpp and TCGA-MF-Test and identical F1 on Canine. MiTHras exceeded all three MAE baselines under
Table IV compares detection performance. MiTHras obtained the highest mean F1 on MIDOGpp (0.829) and TCGAMF-Test (0.599), and was comparable to the best model (GPFM) on Canine (0.797 vs. 0.798). Differences among second-stage classifiers were smaller than in the classification experiment, as the shared first-stage detector constrains candidate recall. Nevertheless, most two-stage models outperformed the standalone segmentation model f seg , confirming that the second stage contributes by separating true MFs from mimickers. Among the two-stage methods, SupMAE, pretrained on the smaller labeled MIDOG22 corpus, yielded the lowest mean F1 on all three test sets. Foundation-model rankings varied by dataset: Virchow led this group on MIDOGpp and TCGA-MFTest, whereas GPFM led on Canine. C. Mitotic count for survival analysis We examined whether automated MF counts align with diagnostic grade and provide prognostic value. Table V reports cohort-specific survival results and average second-stage inference time across TCGA-THCA and TCGA-GBMLGG. Figure 3 illustrates inference timing for TCGA-THCA. Correlation with Diagnostic Subtypes. We compared mitotic counts produced by MiTHras under the two counting strategies, mitotic hotspot (2 mm2 ) and mitotic density (20 mm2 ), for two diagnostic tasks in TCGA-GBMLGG (Figure 4): the lower-grade glioma (LGG) and glioblastoma (GBM) cohorts, and grade 2 versus 3 versus 4. Median mitotic counts increased with tumor grade under both strategies (median hotspot count
10
PREPRINT
TABLE IV: Mitotic figure detection results (F1 , mean ± standard deviation over three folds). For each stage-based model, the detection/classification threshold pair is selected to maximize F1 across the three MIDOG22-CV folds. This pair is applied to MIDOGpp, Canine, and TCGA-MF-Test. YOLOv10 uses its own inference confidence threshold. The Efficient-UNet-b0 segmentation stage is trained using the three-fold MDFS protocol and shared across all second-stage classifier comparisons. ‡ SupMAE is pretrained on the labeled MIDOG22 images. Best value per column in bold.
(406 WSIs), these results are reported descriptively. These C-indices evaluate mitotic count as a single morphological predictor. Samples with identical counts receive identical predicted risks, and these ties receive half credit in Harrell’s C-index. Consequently, differences between methods reflect variations in detection coverage as well as actual biological variations across slides. We therefore present these cohortlevel associations as an exploratory assessment of automated mitotic counting.
Pretraining Method
D. Mitotic figure subtype classification results
Network
MIDOGpp
Canine
Detection baselines Efficient-UNet-b0 (only f seg ) Efficient-UNet-b0 0.782±0.000 0.783±0.000 YOLOv10 0.760±0.021 0.791±0.000 YOLOv10
TCGA-MF-Test 0.571±0.000 0.562±0.000
Natural-image pretraining MDFS ViT-ImageNet (Supervised) MAE-ImageNet DINOv2 CLIP
EfficientNet-b7 ViT-base ViT-base ViT-base ViT-base
0.810±0.002 0.785±0.003 0.810±0.002 0.789±0.006 0.810±0.006 0.791±0.007 0.813±0.001 0.795±0.002 0.807±0.004 0.791±0.001
0.591±0.003 0.582±0.006 0.584±0.003 0.592±0.002 0.590±0.003
Histopathology pretraining models CTransPath-MoCov3 Swin-tiny CHIEF Swin-tiny Lunit-DINO ViT-small PLIP ViT-base/32 ViT-base MAE-TCGA-Unlabeled UNI-DINOv2 ViT-large GPFM ViT-large Virchow ViT-huge SupMAE‡ ViT-base
0.789±0.007 0.791±0.005 0.797±0.006 0.791±0.005 0.802±0.006 0.789±0.006 0.802±0.003 0.784±0.007 0.799±0.006 0.789±0.002 0.814±0.007 0.795±0.002 0.784±0.009 0.798±0.003 0.816±0.002 0.796±0.002 0.767±0.005 0.774±0.013
0.575±0.001 0.572±0.002 0.574±0.005 0.577±0.003 0.568±0.002 0.587±0.005 0.575±0.004 0.589±0.001 0.540±0.007
Pretraining methods using our pseudo-labeled dataset: TCGA-MF-Pseudo ViT-TCGA-MF-Pseudo MAE-TCGA-MF-Pseudo CMAE MiTHras (Ours)
ViT-base ViT-base ViT-base ViT-base
0.809±0.003 0.791±0.002 0.804±0.003 0.791±0.003 0.810±0.003 0.785±0.005 0.829±0.002 0.797±0.006
0.586±0.003 0.569±0.004 0.582±0.006 0.599±0.005
Fig. 4: Boxplots of mitotic counts across different diagnoses in the TCGA-GBMLGG cohorts. Median (η) and mean (µ) values are indicated. The mitotic count is produced by MiTHras using mitotic hotspot (2 mm2 ) and mitotic density (20 mm2 ) methods.
5 in LGG versus 17 in GBM; 4, 6 and 17 across grades 2–4). The two strategies diverged mainly on the intermediate grades: hotspot counts for grades 2 and 3 are nearly indistinguishable (medians 4 and 6), whereas density shows a larger difference in group medians (11 and 23). Survival Prediction. Table V reports the apparent C-index and 95% bootstrap confidence interval for each count-based Cox model. On TCGA-GBMLGG, MiTHras obtained C-indices of 0.627 (hotspot) and 0.671 (density). On TCGA-THCA, all point estimates fell below 0.5, with MiTHras yielding 0.446 (95% CI 0.313–0.607) for hotspot and 0.427 (0.287– 0.593) for density. Given only 11 events across 395 patients
Table VI presents MF subtype classification results. MiTHras achieved the highest mean F1 on both benchmarks: 0.700 on AMi-Br-TUPAC and 0.483 on AMi-MorphTUPAC, outperforming the strongest baselines by 0.046 and 0.031, respectively. MiTHras also outperformed SupMAE and the ViT-, MAE-, and CMAE-based models pretrained on TCGA-MF-Pseudo on both benchmarks. Several pathology foundation models, including CTransPath-MoCov3, CHIEF, UNI-DINOv2, and Virchow, fell below the supervised ViTImageNet baseline. Eight-class morphology remains challenging, yielding a maximum macro F1 of 0.483, in which the metric differs from the positive-class F1 used for AMi-Br. E. Ablation study Table VII shows the effect of loss components, β, backbone, and pretraining corpus. Regarding the loss components, masked reconstruction (LM IM ) alone performed poorly. Incorporating image-level SCL (Limg SCL ) substantially improved performance, increasing F1 across all eight evaluations and raising the descriptive mean from 0.544 to 0.766, with the largest gains observed in subtype classification. Notably, Limg SCL alone produced variable results across datasets. These results indicate that neither term alone is sufficient for consistent and reliable performance. Incorporating token-level SCL (Ltok SCL ) produced taskdependent trade-offs. Relative to using LM IM and Limg SCL (β = 1.0), the main configuration (β = 0.75) enhanced the overall mean from 0.766 to 0.769 but yielded mixed tasklevel results: AMi-Br F1 improved from 0.667 to 0.700, while AMi-Morph F1 dropped from 0.494 to 0.483. Fine-tuned MF-versus-mimicker F1 decreased marginally by 0.001 − 0.010 across MIDOGpp, Canine, and TCGA-MF-Test. Linear probing showed gains for Canine (+0.009) and TCGA-MFTest (+0.016), but a drop for MIDOGpp (−0.011). Thus, token-level supervision does not consistently improve fine morphological discrimination. Moreover, excluding LM IM tok (0.75Limg SCL + 0.25LSCL ) had divergent effects under finetuning and linear probing. It lowered F1 on all three fine-tuned MF-versus-mimicker benchmarks, yet produced mixed linearprobing results: MIDOGpp and TCGA-MF-Test increased, while Canine decreased. AMi-Br remained at 0.700, but AMiMorph decreased from 0.483 to 0.475. Backbone scale and pretraining corpus also affect transferability. With ViT-tiny, β = 0.75 improved all three linearprobing and both subtype results over β = 1, while lowering all three fine-tuned MF-versus-mimicker results. ViT-tiny
VUONG et al.: MITHRAS: TASK-SPECIFIC HIERARCHICAL SEMI-SUPERVISED CONTRASTIVE MASKED AUTOENCODER FOR MITOTIC FIGURE ANALYSIS
11
TABLE V: C-index for survival analysis on TCGA-THCA and TCGA-GBMLGG. Time is the average wall-clock inference time of the second-stage classifier per WSI, measured on a single NVIDIA RTX A6000 GPU with automatic mixed precision. Relative time is this time normalized to the ViT-base configuration (1.00×). All two-stage methods share the same EfficientUNet-b0 segmentation stage. The segmentation-only baseline performs no classification. Best value per column is shown in bold for TCGA-GBMLGG only. No values are highlighted for TCGA-THCA, which is reported descriptively because only 11 deaths are recorded. TCGA-THCA Pretraining Method
TCGA-GBMLGG
Network (f cls )
C-index Mitotic hotspot (2 mm2 )
C-index Mitotic density (20 mm2 )
C-index Mitotic hotspot (2 mm2 )
C-index Mitotic density (20 mm2 )
Classifier inference Time (second)
Relative time
Single stage models
Efficient-UNet-b0 (only f seg )
-
0.437 (0.405, 0.485)
0.438 (0.405, 0.490)
0.540 (0.508, 0.574)
0.541 (0.509, 0.575)
-
-
Natural-image pretraining
MDFS ViT-ImageNet MAE-ImageNet CLIP
EfficientNet-b7 ViT-base ViT-base ViT-base
0.417 (0.381, 0.475) 0.435 (0.358, 0.564) 0.437 (0.358, 0.565) 0.417 (0.381, 0.475)
0.417 (0.381, 0.475) 0.433 (0.357, 0.562) 0.434 (0.356, 0.562) 0.417 (0.381, 0.475)
0.541 (0.504, 0.575) 0.528 (0.488, 0.566) 0.527 (0.487, 0.566) 0.539 (0.503, 0.573)
0.543 (0.507, 0.577) 0.533 (0.493, 0.572) 0.533 (0.492, 0.571) 0.542 (0.506, 0.575)
9.17 6.80 6.80 6.80
1.35 × 1.00 × 1.00 × 1.00 ×
Histopathology pretraining models
CTransPath-MoCov3 CHIEF Lunit-DINO PLIP MAE-TCGA-Unlabeled SupMAE UNI-DINOv2 GPFM Virchow
Swin-tiny Swin-tiny ViT-small ViT-base/32 ViT-base ViT-base ViT-large ViT-large ViT-huge
0.416 (0.381, 0.473) 0.417 (0.381, 0.475) 0.417 (0.381, 0.475) 0.419 (0.382, 0.480) 0.415 (0.381, 0.469) 0.416 (0.381, 0.472) 0.436 (0.357, 0.564) 0.417 (0.381, 0.478) 0.417 (0.381, 0.474)
0.415 (0.381, 0.468) 0.418 (0.381, 0.480) 0.417 (0.381, 0.478) 0.419 (0.382, 0.482) 0.416 (0.381, 0.471) 0.416 (0.381, 0.473) 0.436 (0.357, 0.564) 0.417 (0.381, 0.478) 0.417 (0.381, 0.477)
0.540 (0.504, 0.574) 0.540 (0.503, 0.574) 0.539 (0.503, 0.572) 0.538 (0.502, 0.572) 0.539 (0.503, 0.573) 0.539 (0.503, 0.573) 0.527 (0.487, 0.565) 0.540 (0.504, 0.574) 0.541 (0.504, 0.574)
0.542 (0.506, 0.575) 0.541 (0.505, 0.574) 0.541 (0.504, 0.573) 0.541 (0.505, 0.574) 0.541 (0.505, 0.574) 0.541 (0.505, 0.574) 0.534 (0.493, 0.573) 0.541 (0.505, 0.574) 0.544 (0.507, 0.577)
5.77 5.77 3.71 3.02 6.80 6.80 21.94 21.94 32.60
0.85 × 0.85 × 0.55 × 0.44 × 1.00 × 1.00 × 3.23 × 3.23 × 4.79 ×
Pretraining methods using our pseudo-labeled dataset: TCGA-MF-Pseudo
ViT-TCGA-MF-Pseudo MAE-TCGA-MF-Pseudo CMAE MiTHras
ViT-base ViT-base ViT-base ViT-base
0.417 (0.381, 0.478) 0.433 (0.357, 0.561) 0.433 (0.357, 0.559) 0.446 (0.313, 0.607)
0.419 (0.381, 0.481) 0.435 (0.357, 0.563) 0.432 (0.357, 0.558) 0.427 (0.287, 0.593)
0.540 (0.503, 0.574) 0.529 (0.489, 0.566) 0.528 (0.488, 0.567) 0.627 (0.580, 0.671)
0.542 (0.505, 0.575) 0.533 (0.493, 0.571) 0.531 (0.491, 0.570) 0.671 (0.630, 0.713)
6.80 6.80 6.80 6.80
1.00 × 1.00 × 1.00 × 1.00 ×
TABLE VI: Mitotic figure subtype classification results across AMi-Br-TUPAC and AMi-Morph-TUPAC. All values are F1 , reported as mean ± standard deviation over three folds. ‡ SupMAE is pretrained on the labeled MIDOG22 images. Best value per column in bold. Pretraining Method Natural-image pretraining Bertram et al. [42] MDFS ViT-ImageNet (Supervised) MAE-ImageNet DINOv2 CLIP
Network
AMi-Br-TUPAC (typical vs. atypical)
AMi-Morph-TUPAC (8 subtypes)
EfficientNet-V2-S EfficientNet-b7 ViT-base ViT-base ViT-base ViT-base
0.577±0.019 0.583±0.038 0.623±0.005 0.456±0.251 0.654±0.011 0.643±0.018
0.308±0.027 0.410±0.039 0.439±0.012 0.413±0.030 0.445±0.032 0.445±0.020
0.228±0.203 0.542±0.007 0.642±0.003 0.575±0.015 0.246±0.062 0.505±0.035 0.568±0.025 0.653±0.013 0.570±0.011
0.319±0.050 0.305±0.018 0.434±0.026 0.351±0.029 0.303±0.004 0.376±0.034 0.357±0.026 0.452±0.034 0.329±0.020
at β = 0.75, replacing IMPaSh with CONCH decreased the eight-evaluation mean F1 from 0.769 to 0.765. At β = 1, the corresponding means were comparable (0.766 vs. 0.767; Table VII). Overall, although some individual subtype variations were larger, the aggregate differences remained minor. F. Pseudo-label quality
Histopathology pretraining models CTransPath-MoCov3 Swin-tiny CHIEF Swin-tiny PLIP ViT-base/32 Lunit-DINO ViT-small MAE-TCGA-Unlabeled ViT-base SupMAE‡ ViT-base ViT-large UNI-DINOv2 GPFM ViT-large Virchow ViT-huge
Pretraining methods using our pseudo-labeled dataset: TCGA-MF-Pseudo ViT-TCGA-MF-Pseudo ViT-base 0.535±0.035 ViT-base 0.100±0.058 MAE-TCGA-MF-Pseudo CMAE ViT-base 0.597±0.014 MiTHras (Ours) ViT-base 0.700±0.005
0.252±0.007 0.225±0.086 0.362±0.006 0.483±0.006
approached ViT-base in binary classification, though its AMiMorph F1 remained lower. For ViT-base at β = 0.75, pretraining on TCGA-MF-Pseudo produced an eight-evaluation mean of 0.769, surpassing 0.684 obtained with labeled MIDOG22, highlighting the benefits of the larger scale and diversity of the pseudo-labeled corpus. To assess the impact of the tumor filter, we compared IMPaSh and CONCH. IMPaSh retained 14.7% of the sampled patches, compared with 27.0% for CONCH. In the pathologistreviewed disagreement strata, pooled tumor precision was 87.7–95.4% for patches retained only by IMPaSh and 87.7– 92.9% for patches retained only by CONCH (Table VIII). Performance varied across cohorts, with neither filter consistently achieving higher precision. In downstream evaluation
The blinded review of 420 pseudo-mitotic review records yielded pool-weighted precision of 94.9–98.1% (unweighted: 92.9–97.1%; Table IX). In the additional-figure assessment, 8 of the 199 crops (4.02%; 95% Wilson CI, 2.05–7.73%) contained another MF, indicating that crops centered on nonmitotic objects can still capture peripheral MFs, introducing errors into pseudo labels. The bounding-box assessment included 445 reviews of confirmed MFs (390 pseudo-mitotic and 55 pseudo-mimicker). Overall, 81.6% of reviewed MFs (363/445) fit within a box of at most 50 × 50 pixels. Specifically, 3.1% of MFs fit within a box of 25 pixels, 78.4% within 50 pixels, 14.2% within 75 pixels, and 4.3% exceeded 75 pixels. The fixed pseudo box therefore provides coarse spatial supervision for token-level learning. G. Effect of pseudo-label noise To evaluate sensitivity to pseudo-label errors, we repeated pretraining after swapping 165,514 labels (9.224% of the corpus: 82,757 pseudo-mitotic crops with classifier probabilities of 0.5–0.6 and 82,757 ranked pseudo-mimickers). The eightevaluation mean F1 decreased only from 0.769 to 0.764, with per-benchmark changes between −0.017 and +0.013 (Table VII). VI. D ISCUSSION In this work, we proposed MiTHras, a task-specific hierarchical SCL framework for MF analysis. MiTHras builds on ViT-base pretrained on 1.8 million images from 6,041
12
PREPRINT
TABLE VII: Results of ablation experiments. Unless stated otherwise, rows use ViT-base pretrained on TCGA-MF-Pseudo with the IMPaSh tumor filter. Results are mean F1 ± standard deviation over three folds. Avg. is a descriptive, unweighted mean of the eight displayed evaluations. The shaded row is the main configuration. Bold marks the highest displayed mean in each column. Configuration
Fine-tuning Canine TCGA-MF-Test
MIDOGpp
MIDOGpp
Linear probing Canine TCGA-MF-Test
Fine-tuning AMi-Br AMi-Morph
Avg.
Loss components LMIM only (MAE) img LSCL only img 0.75LSCL + 0.25Ltok SCL (no MIM) LMIM + Ltok SCL (β = 0) img LMIM + LSCL (β = 1)
0.827±0.006 0.857±0.008 0.726±0.033 0.862±0.005 0.865±0.009
0.840±0.005 0.854±0.001 0.694±0.026 0.859±0.006 0.857±0.005
0.746±0.007 0.788±0.013 0.648±0.021 0.781±0.011 0.799±0.009
0.323±0.100 0.833±0.001 0.839±0.002 0.801±0.002 0.847±0.004
0.743±0.045 0.843±0.002 0.843±0.004 0.833±0.001 0.839±0.002
0.548±0.021 0.813±0.002 0.804±0.003 0.750±0.002 0.759±0.001
0.100±0.058 0.677±0.015 0.700±0.005 0.676±0.010 0.667±0.016
Sensitivity to β (all three terms) β = 0.25 β = 0.50 β = 0.75 (MiTHras)
0.862±0.004 0.856±0.002 0.861±0.004 0.858±0.002 0.864±0.004 0.855±0.001
0.787±0.004 0.793±0.006 0.789±0.011
0.839±0.004 0.830±0.004 0.842±0.004 0.835±0.005 0.836±0.003 0.848±0.005
0.775±0.005 0.781±0.001 0.775±0.002
0.690±0.014 0.457±0.023 0.762 0.649±0.010 0.468±0.003 0.761 0.700±0.005 0.483±0.006 0.769
ViT-tiny backbone β=1 β = 0.75
0.862±0.004 0.860±0.002 0.858±0.002 0.856±0.001
0.792±0.003 0.787±0.004
0.838±0.003 0.843±0.002 0.843±0.001 0.846±0.003
0.765±0.000 0.782±0.002
0.659±0.022 0.343±0.005 0.745 0.670±0.012 0.394±0.055 0.754
Pretraining methods using labeled dataset: MIDOG22 β=1 0.781±0.020 0.812±0.002 β = 0.75 0.781±0.011 0.816±0.008
0.711±0.008 0.721±0.009
0.764±0.016 0.794±0.020 0.765±0.011 0.790±0.020
0.683±0.024 0.687±0.021
0.511±0.032 0.358±0.010 0.677 0.553±0.023 0.362±0.028 0.684
TCGA-MF-Pseudo constructed with CONCH instead of IMPaSh β=1 0.865±0.006 0.856±0.006 β = 0.75 0.857±0.001 0.856±0.006
0.788±0.004 0.790±0.006
0.839±0.002 0.848±0.003 0.827±0.004 0.851±0.002
0.777±0.002 0.790±0.003
0.686±0.016 0.473±0.048 0.767 0.679±0.016 0.473±0.011 0.765
MiTHras with 165,514 interchanged pseudo labels β = 0.75 0.856±0.005 0.847±0.013
0.784±0.008
0.839±0.002 0.836±0.004
0.768±0.003
0.683±0.006 0.496±0.023 0.764
TABLE VIII: Tumor-patch selection: pathologist-confirmed precision of IMPaSh and CONCH retention decisions, 25 patches per stratum per cohort. Ranges span “unsure” counted as incorrect or correct. Cohort
Both
CONCH only
IMPaSh only
Bladder (BLCA) Breast (BRCA) Colon (COLON DX) Esophagus (ESCA) Renal chromophobe (KICH) Renal clear cell (KIRC) Renal papillary (KIRP) Liver (LIHC) Lung adeno. (LUAD) Lung squamous (LUSC) Ovary (OV) Pancreas (PAAD) Prostate (PRAD) Uterus (UCEC)
1.00 1.00 0.96 1.00 0.96 1.00 1.00 1.00 0.96–1.00 1.00 1.00 0.92 0.96–1.00 1.00
0.96–1.00 0.88 0.92–0.96 0.88–0.92 0.92 1.00 0.92–1.00 1.00 0.56–0.72 0.72–0.76 0.88–0.92 0.84–0.96 0.88–1.00 0.92–0.96
0.96–1.00 0.92–1.00 0.92 0.72–0.88 0.80–0.92 0.88–1.00 0.84–0.96 0.96–1.00 0.88–0.96 1.00 0.76–0.92 0.84 0.92–1.00 0.88–0.96
Pooled (n = 350)
0.983–0.989
0.877–0.929
0.877–0.954
WSIs. In comparison to the ViT-large UNI-DINOv2 model, pretrained on 100 million images from 100,426 WSIs [48], MiTHras achieved higher mean F1 on the three classification and two subtype benchmarks, despite its smaller backbone and corpus. Furthermore, its second-stage classifier required lower inference time than the larger ViT-large and ViT-huge encoders (Table V). These results underscore the value of aligning pretraining data and objectives with MF analysis. The tumor-filter comparison (Section V-E) merits further explanation, given that IMPaSh–trained on labeled colorectal tissue patches–performed comparably to CONCH. We attribute this comparable efficacy to IMPaSh’s capability in distinguishing non-tumor tissues rather than broad organ coverage. IMPaSh assigns each sub-patch to tumor or one of eight
0.225±0.086 0.458±0.012 0.475±0.030 0.475±0.023 0.494±0.020
0.544 0.765 0.716 0.755 0.766
TABLE IX: Mitotic pseudo-label precision by cohort, 30 pseudo-mitotic crop reviews per cohort. Pooled value is weighted by cohort pool size. Cohort
Pool size
Precision
95% CI (worst case)
Breast (BRCA) Renal papillary (KIRP) Ovary (OV) Uterus (UCEC) Esophagus (ESCA) Colon (COLON DX) Lung squamous (LUSC) Pancreas (PAAD) Liver (LIHC) Bladder (BLCA) Lung adeno. (LUAD) Prostate (PRAD) Renal clear cell (KIRC) Renal chromophobe (KICH)
97,540 8,059 42,582 88,659 33,332 105,805 100,249 5,095 55,433 68,396 43,397 8,164 1,780 2,340
1.00 1.00 1.00 1.00 0.97 0.93–0.97 0.93–0.97 0.93–0.97 0.90–1.00 0.90–0.97 0.90–0.97 0.90–0.97 0.83–0.97 0.80–0.87
0.89–1.00 0.89–1.00 0.89–1.00 0.89–1.00 0.83–0.99 0.79–0.98 0.79–0.98 0.79–0.98 0.74–0.97 0.74–0.97 0.74–0.97 0.74–0.97 0.66–0.93 0.63–0.90
Pooled, pool-weighted
660,831
0.949–0.981
N/A
non-tumor categories. Aside from normal colon mucosa, the remaining categories (adipose, background, debris, lymphocytes, mucus, smooth muscle, and stroma) exhibit common morphological characteristics across diverse organs. Accordingly, filtering out these components may support tumor patch selection beyond the colorectal domain. CONCH, by comparison, learned broad pathology visual-language representations, without being explicitly optimized for tumor detection. Its broader organ coverage therefore need not translate into higher precision for this particular filtering task. Nevertheless, both filters operate as pre-processing steps to generate the pretraining MF and mimicker cropss. The small difference in eight-evaluation F1 mean (IMPaSh: 0.769 vs. CONCH: 0.765) indicates limited aggregate sensitivity to this filtering choice. MiTHras integrates pseudo-label-guided local alignment
VUONG et al.: MITHRAS: TASK-SPECIFIC HIERARCHICAL SEMI-SUPERVISED CONTRASTIVE MASKED AUTOENCODER FOR MITOTIC FIGURE ANALYSIS
with global class alignment and masked reconstruction. Local representation learning methods such as VICRegL [52] use correspondences between augmented views, whereas MiTHras groups tokens by their mitotic or non-mitotic pseudo-labels. The binary token labels distinguish mitotic from non-mitotic regions–aiding typical-versus-atypical classification–they do not consistently enhance discrimination among eight morphology classes. This may be ascribable to the use of fixed pseudo bounding boxes, providing coarse and imperfect spatial labels. The largest performance advantages of MiTHras over competing models emerge under linear probing, whereas detection results were comparable due to the recall constraint of the shared first-stage candidate detector. Moreover, detection performance universally declined on TCGA-MF-Test compared to MIDOGpp (e.g., MiTHras F1 dropped from 0.829 to 0.599), likely due to varying slide quality and acquisition conditions. Similar degradation has been reported by Jahanifar et al. [18] and Aubreville et al. [8] on multi-domain datasets such as MIDOG21 and MIDOG22. This consistent decline across architectures and pretraining strategies indicates a persistent transfer challenge and motivates broader external validation. This study has several limitations. First, the benefits of MiTHras vary across tasks and datasets. Notably, MF subtype classification remains challenging and requires extended validation. Second, token supervision relies on fixed 50 × 50 pseudo boxes, which may include background or exclude parts of larger MFs, thereby introducing supervision noise. Developing more precise spatial labels represents a promising direction for improving token-level supervision. Third, the audit identified residual errors in the pseudo labels. 55 of 280 reviewed pseudo-mimicker crops contained true MFs, corresponding to 19.6% contamination (95% Wilson CI, 15.4–24.7%) and a pool-weighted estimate of 21.7%. Per-cohort estimates ranged from 5% to 35%, with 20 crops per cohort. These errors can incorrectly group mitotic examples with non-mitotic examples during contrastive learning. Fourth, although label-interchange experiment demonstrated robustness to pseudo-label noise, this global perturbation may not fully reflect real-world conditions. Label noise can be non-uniform, concentrating on specific organs or challenging morphologies. We leave the development of precise pseudo-label generation and refinement strategies to future investigation. Fifth, the main survival comparisons use mitotic count as a single morphological predictor. Although an age- and sex-adjusted analysis assesses its incremental contribution, broader CPath survival models can combine morphology, clinical variables, and transcriptomic profiles [14]. We plan to integrate MiTHras-derived counts into such models to assess their complementary prognostic value. Finally, the clinical utility of MiTHras has yet to be validated; further studies are needed to assess its impact on real-world diagnostic workflows and outcomes. VII. C ONCLUSIONS MiTHras combines pseudo-label-guided image- and tokenlevel SCL with masked reconstruction for MF analysis. Pretraining on the 1.8-million-image TCGA-MF-Pseudo corpus, it yields strong performance on external benchmarks, outperforming general-purpose and pathology foundation encoders
13
most significantly under linear probing. Token-level supervision improves some subtype results, while its effect varies across tasks and configurations. Future work will explore finer local labels, noise-robust pseudo-labeling, and broader clinical validation to establish the practical utility of MiTHras. ACKNOWLEDGMENTS We acknowledge Navid Alemi for his contributions to this project. R EFERENCES [1] J. H. von der Thüsen, Y. S. Tham, H. Pattenden, A. Rice, M. Dusmet, E. Lim, and A. G. Nicholson, “Prognostic significance of predominant histologic pattern and nuclear grade in resected adenocarcinoma of the lung: potential parameters for a grading system,” Journal of Thoracic Oncology, vol. 8, no. 1, pp. 37–44, 2013. [2] C. K. Jung, A. Bychkov, and K. Kakudo, “Update from the 2022 world health organization classification of thyroid tumors: a standardized diagnostic approach,” Endocrinology and Metabolism, vol. 37, no. 5, pp. 703–718, 2022. [3] W. H. Organization, “Who classification of tumours,” 2022, accessed: 2025-04-03. [4] C. W. Elston and I. O. Ellis, “Pathological prognostic factors in breast cancer. i. the value of histological grade in breast cancer: experience from a large study with long-term follow-up,” Histopathology, vol. 19, no. 5, pp. 403–410, 1991. [5] J.-M. Coindre, “Grading of soft tissue sarcomas: review and update,” Archives of Pathology & Laboratory Medicine, vol. 130, no. 10, pp. 1448–1453, 2006. [6] M. Miettinen and J. Lasota, “Gastrointestinal stromal tumors: pathology and prognosis at different sites,” Seminars in Diagnostic Pathology, vol. 23, no. 2, pp. 70–83, 2006. [7] G. Rindi, D. S. Klimstra, B. Abedi-Ardekani, S. L. Asa, F. T. Bosman, E. Brambilla, K. J. Busam, R. R. de Krijger, M. Dietel, A. K. ElNaggar et al., “A common classification framework for neuroendocrine neoplasms: an International Agency for Research on Cancer (IARC) and World Health Organization (WHO) expert consensus proposal,” Modern Pathology, vol. 31, no. 12, pp. 1770–1786, 2018. [8] M. Aubreville, F. Wilm, N. Stathonikos, K. Breininger, T. A. Donovan, S. Jabari, M. Veta, J. Ganz, J. Ammeling, P. J. van Diest et al., “A comprehensive multi-domain dataset for mitotic figure detection,” Scientific data, vol. 10, no. 1, p. 484, 2023. [9] M. Aubreville, N. Stathonikos, T. A. Donovan, R. Klopfleisch, J. Ammeling, J. Ganz, F. Wilm, M. Veta, S. Jabari, M. Eckstein et al., “Domain generalization across tumor types, laboratories, and species—insights from the 2022 edition of the mitosis domain generalization challenge,” Medical Image Analysis, vol. 94, p. 103155, 2024. [10] T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A simple framework for contrastive learning of visual representations,” in International conference on machine learning. PmLR, 2020, pp. 1597–1607. [11] K. He, X. Chen, S. Xie, Y. Li, P. Dollár, and R. Girshick, “Masked autoencoders are scalable vision learners,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 16 000–16 009. [12] X. Wang, S. Yang, J. Zhang, M. Wang, J. Zhang, W. Yang, J. Huang, and X. Han, “Transformer-based unsupervised contrastive learning for histopathological image classification,” Medical image analysis, vol. 81, p. 102559, 2022. [13] E. Vorontsov, A. Bozkurt, A. Casson, G. Shaikovski, M. Zelechowski, K. Severson, E. Zimmermann, J. Hall, N. Tenenholtz, N. Fusi et al., “A foundation model for clinical-grade computational pathology and rare cancers detection,” Nature medicine, vol. 30, no. 10, pp. 2924–2935, 2024. [14] A. H. Song, R. J. Chen, G. Jaume, A. J. Vaidya, A. Baras, and F. Mahmood, “Multimodal prototyping for cancer survival prediction,” in Forty-first International Conference on Machine Learning, 2024. [15] J. Lee, J. Lim, K. Byeon, and J. T. Kwak, “Benchmarking pathology foundation models: Adaptation strategies and scenarios,” arXiv preprint arXiv:2410.16038, 2024. [16] A. T. Nguyen, K. Byeon, K. Kim, B. Song, S. W. Chae, and J. T. Kwak, “Camp: Continuous and adaptive learning model in pathology,” arXiv preprint arXiv:2407.09030, 2024.
14
[17] T. T. Le Vuong and J. T. Kwak, “Moma: momentum contrastive learning with multi-head attention-based knowledge distillation for histopathology image analysis,” Medical Image Analysis, vol. 101, p. 103421, 2025. [18] M. Jahanifar, A. Shephard, N. Zamanitajeddin, S. Graham, S. E. A. Raza, F. Minhas, and N. Rajpoot, “Mitosis detection, fast and slow: robust and efficient detection of mitotic figures,” Medical Image Analysis, vol. 94, p. 103132, 2024. [19] F. Wilm, C. Marzahl, K. Breininger, and M. Aubreville, “Domain adversarial retinanet as a reference algorithm for the mitosis domain generalization challenge,” in International Conference on Medical Image Computing and Computer-Assisted Intervention. Springer, 2021, pp. 5– 13. [20] M. Sebai, X. Wang, and T. Wang, “Maskmitosis: a deep learning framework for fully supervised, weakly supervised, and unsupervised mitosis detection in histopathology images,” Medical & Biological Engineering & Computing, vol. 58, pp. 1603–1623, 2020. [21] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778. [22] R. H. Fick, A. Moshayedi, G. Roy, J. Dedieu, S. Petit, and S. B. Hadj, “Domain-specific cycle-gan augmentation improves domain generalizability for mitosis detection,” in International conference on medical image computing and computer-assisted intervention. Springer, 2021, pp. 40–47. [23] H. Wang, H. Xu, B. Li, X. Pan, L. Zeng, R. Lan, and X. Luo, “A novel dataset and a two-stage mitosis nuclei detection method based on hybrid anchor branch,” Biomedical Signal Processing and Control, vol. 87, p. 105374, 2024. [24] R. Ding, J. Hall, N. Tenenholtz, and K. Severson, “Improving mitosis detection on histopathology images using large vision-language models,” in 2024 IEEE International Symposium on Biomedical Imaging (ISBI). IEEE, 2024, pp. 1–5. [25] R. J. Chen, C. Chen, Y. Li, T. Y. Chen, A. D. Trister, R. G. Krishnan, and F. Mahmood, “Scaling vision transformers to gigapixel images via hierarchical self-supervised learning,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 16 144–16 155. [26] M. Ilse, J. Tomczak, and M. Welling, “Attention-based deep multiple instance learning,” in Proceedings of the 35th International Conference on Machine Learning (ICML), ser. PMLR, vol. 80, 2018, pp. 2127–2136. [27] F. Liang, Y. Li, and D. Marculescu, “Supmae: Supervised masked autoencoders are efficient vision learners,” arXiv preprint arXiv:2205.14540, 2022. [28] J. Devlin, M.-W. Chang, K. Lee, and K. Toutanova, “Bert: Pre-training of deep bidirectional transformers for language understanding,” in Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), 2019, pp. 4171–4186. [29] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” arXiv preprint arXiv:2010.11929, 2020. [30] C. Wei, K. Mangalam, P.-Y. Huang, Y. Li, H. Fan, H. Xu, H. Wang, C. Xie, A. Yuille, and C. Feichtenhofer, “Diffusion models as masked autoencoders,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 16 284–16 294. [31] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., “Learning transferable visual models from natural language supervision,” in International conference on machine learning. PmLR, 2021, pp. 8748–8763. [32] P. Khosla, P. Teterwak, C. Wang, A. Sarna, Y. Tian, P. Isola, A. Maschinot, C. Liu, and D. Krishnan, “Supervised contrastive learning,” Advances in neural information processing systems, vol. 33, pp. 18 661–18 673, 2020. [33] D.-H. Lee et al., “Pseudo-label: The simple and efficient semi-supervised learning method for deep neural networks,” in Workshop on challenges in representation learning, ICML, vol. 3, no. 2. Atlanta, 2013, p. 896. [34] A. Tarvainen and H. Valpola, “Mean teachers are better role models: Weight-averaged consistency targets improve semi-supervised deep learning results,” Advances in neural information processing systems, vol. 30, 2017. [35] K. Sohn, D. Berthelot, N. Carlini, Z. Zhang, H. Zhang, C. A. Raffel, E. D. Cubuk, A. Kurakin, and C.-L. Li, “Fixmatch: Simplifying semisupervised learning with consistency and confidence,” Advances in neural information processing systems, vol. 33, pp. 596–608, 2020. [36] B. Zhang, Y. Wang, W. Hou, H. Wu, J. Wang, M. Okumura, and T. Shinozaki, “Flexmatch: Boosting semi-supervised learning with curriculum
PREPRINT
pseudo labeling,” Advances in neural information processing systems, vol. 34, pp. 18 408–18 419, 2021. [37] M. Zheng, S. You, L. Huang, F. Wang, C. Qian, and C. Xu, “Simmatch: Semi-supervised learning with similarity matching,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 14 471–14 481. [38] N. Otsu et al., “A threshold selection method from gray-level histograms,” Automatica, vol. 11, no. 285-296, pp. 23–27, 1975. [39] T. T. L. Vuong, Q. D. Vu, M. Jahanifar, S. Graham, J. T. Kwak, and N. Rajpoot, “Impash: A novel domain-shift resistant representation for colorectal cancer tissue classification,” in European Conference on Computer Vision. Springer, 2022, pp. 543–555. [40] M. Aubreville, C. A. Bertram, T. A. Donovan, C. Marzahl, A. Maier, and R. Klopfleisch, “A completely annotated whole slide image dataset of canine breast cancer to aid human breast cancer research,” Scientific data, vol. 7, no. 1, p. 417, 2020. [41] C. A. Bertram, M. Aubreville, C. Marzahl, A. Maier, and R. Klopfleisch, “A large-scale dataset for mitotic figure assessment on whole slide images of canine cutaneous mast cell tumor,” Scientific data, vol. 6, no. 1, p. 274, 2019. [42] C. A. Bertram, V. Weiss, T. A. Donovan, S. Banerjee, T. Conrad, J. Ammeling, R. Klopfleisch, C. Kaltenecker, and M. Aubreville, “Histologic dataset of normal and atypical mitotic figures on human breast cancer (ami-br),” in BVM Workshop. Springer, 2025, pp. 113–118. [43] M. Oquab, T. Darcet, T. Moutakanni, H. Vo, M. Szafraniec, V. Khalidov, P. Fernandez, D. Haziza, F. Massa, A. El-Nouby et al., “Dinov2: Learning robust visual features without supervision,” arXiv preprint arXiv:2304.07193, 2023. [44] X. Wang, J. Zhao, E. Marostica, W. Yuan, J. Jin, J. Zhang, R. Li, H. Tang, K. Wang, Y. Li et al., “A pathology foundation model for cancer diagnosis and prognosis prediction,” Nature, vol. 634, no. 8035, pp. 970–978, 2024. [45] M. Kang, H. Song, S. Park, D. Yoo, and S. Pereira, “Benchmarking selfsupervised learning on diverse pathology datasets,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 3344–3354. [46] Z. Huang, F. Bianchi, M. Yuksekgonul, T. J. Montine, and J. Zou, “A visual–language foundation model for pathology image analysis using medical twitter,” Nature medicine, vol. 29, no. 9, pp. 2307–2316, 2023. [47] J. Ma, Z. Guo, F. Zhou, Y. Wang, Y. Xu, J. Li, F. Yan, Y. Cai, Z. Zhu, C. Jin et al., “A generalizable pathology foundation model using a unified knowledge distillation pretraining framework,” Nature Biomedical Engineering, pp. 1–20, 2025. [48] R. J. Chen, T. Ding, M. Y. Lu, D. F. Williamson, G. Jaume, A. H. Song, B. Chen, A. Zhang, D. Shao, M. Shaban et al., “Towards a general-purpose foundation model for computational pathology,” Nature Medicine, vol. 30, no. 3, pp. 850–862, 2024. [49] Z. Huang, X. Jin, C. Lu, Q. Hou, M.-M. Cheng, D. Fu, X. Shen, and J. Feng, “Contrastive masked autoencoders are stronger vision learners,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 46, no. 4, pp. 2506–2517, 2023. [50] A. Wang, H. Chen, L. Liu, K. Chen, Z. Lin, J. Han et al., “Yolov10: Real-time end-to-end object detection,” Advances in Neural Information Processing Systems, vol. 37, pp. 107 984–108 011, 2024. [51] M. Y. Lu, B. Chen, D. F. Williamson, R. J. Chen, I. Liang, T. Ding, G. Jaume, I. Odintsov, L. P. Le, G. Gerber et al., “A visual-language foundation model for computational pathology,” Nature medicine, vol. 30, no. 3, pp. 863–874, 2024. [52] A. Bardes, J. Ponce, and Y. LeCun, “Vicregl: Self-supervised learning of local visual features,” Advances in Neural Information Processing Systems, vol. 35, pp. 8799–8810, 2022.
VUONG et al.: MITHRAS: TASK-SPECIFIC HIERARCHICAL SEMI-SUPERVISED CONTRASTIVE MASKED AUTOENCODER FOR MITOTIC FIGURE ANALYSIS
15
S UPPLEMENTARY M ATERIAL S1. S UPPLEMENTARY R ESULTS Full metric results complement the F1 comparisons in the main manuscript. Metric definitions are provided in Section IV-C. TABLE S1: Fine-tuning and linear-probing results for mitotic figure classification. Values are mean ± sample SD over three folds. ACC denotes accuracy; ROC-AUC denotes area under the receiver operating characteristic curve. Binary F1 , precision, and recall treat mitotic figures as the positive class. Bold indicates the highest mean at displayed precision within each dataset and metric. Fine-tuning Method
Network
Natural-image pretraining MDFS ViT-ImageNet (Supervised) MAE-ImageNet DINOv2 CLIP
EfficientNet-b7 ViT-base ViT-base ViT-base ViT-base
ACC
MIDOGpp ROC-AUC Precision
Recall
ACC
Canine ROC-AUC Precision
Recall
ACC
TCGA-MF-Test ROC-AUC Precision
Recall
0.862±0.001 0.853±0.009 0.866±0.003 0.867±0.005 0.856±0.003
0.839±0.112 0.906±0.036 0.939±0.005 0.934±0.005 0.918±0.006
0.865±0.009 0.838±0.009 0.851±0.010 0.858±0.027 0.840±0.006
0.800±0.012 0.811±0.013 0.829±0.020 0.825±0.032 0.818±0.016
0.750±0.007 0.771±0.006 0.779±0.008 0.779±0.007 0.769±0.006
0.776±0.036 0.818±0.017 0.841±0.005 0.828±0.008 0.809±0.016
0.851±0.016 0.856±0.001 0.855±0.009 0.847±0.007 0.842±0.009
0.782±0.032 0.809±0.011 0.826±0.025 0.838±0.021 0.826±0.004
0.811±0.006 0.817±0.001 0.819±0.004 0.824±0.006 0.820±0.005
0.781±0.115 0.859±0.044 0.888±0.004 0.886±0.006 0.863±0.006
0.837±0.015 0.844±0.017 0.823±0.002 0.843±0.013 0.835±0.010
0.691±0.021 0.702±0.018 0.733±0.012 0.723±0.026 0.720±0.001
0.841±0.009 0.843±0.011 0.846±0.006 0.847±0.011 0.834±0.016 0.793±0.013 0.865±0.010 0.828±0.010 0.872±0.003
0.913±0.009 0.918±0.008 0.917±0.004 0.920±0.008 0.912±0.015 0.866±0.018 0.918±0.021 0.878±0.011 0.937±0.001
0.830±0.009 0.806±0.020 0.826±0.015 0.826±0.029 0.803±0.028 0.760±0.019 0.861±0.030 0.804±0.007 0.855±0.001
0.788±0.018 0.830±0.014 0.807±0.006 0.812±0.012 0.808±0.012 0.751±0.018 0.815±0.013 0.785±0.033 0.841±0.006
0.756±0.011 0.755±0.006 0.766±0.011 0.755±0.006 0.758±0.007 0.735±0.023 0.775±0.010 0.768±0.002 0.779±0.005
0.817±0.008 0.814±0.004 0.829±0.007 0.819±0.002 0.816±0.001 0.802±0.009 0.835±0.016 0.782±0.009 0.831±0.001
0.843±0.018 0.837±0.010 0.854±0.022 0.848±0.009 0.835±0.013 0.849±0.014 0.858±0.015 0.833±0.005 0.839±0.007
0.802±0.033 0.809±0.024 0.807±0.050 0.793±0.022 0.819±0.030 0.759±0.057 0.814±0.024 0.837±0.011 0.848±0.018
0.795±0.006 0.799±0.002 0.810±0.004 0.804±0.005 0.797±0.004 0.768±0.014 0.825±0.003 0.802±0.008 0.830±0.003
0.858±0.008 0.855±0.009 0.861±0.013 0.867±0.008 0.861±0.006 0.849±0.019 0.898±0.010 0.841±0.006 0.899±0.005
0.816±0.011 0.837±0.010 0.812±0.013 0.819±0.035 0.791±0.016 0.845±0.022 0.854±0.018 0.811±0.019 0.844±0.014
0.670±0.007 0.656±0.008 0.721±0.026 0.697±0.035 0.712±0.018 0.557±0.041 0.712±0.024 0.700±0.002 0.737±0.010
Pretraining methods using our pseudo-labeled dataset: TCGA-MF-Pseudo ViT-TCGA-MF-Pseudo ViT-base 0.857±0.003 0.930±0.002 MAE-TCGA-MF-Pseudo ViT-base 0.853±0.006 0.927±0.005 CMAE ViT-base 0.863±0.005 0.935±0.004 MiTHras (Ours) ViT-base 0.887±0.004 0.954±0.001
0.861±0.002 0.824±0.012 0.841±0.011 0.886±0.022
0.791±0.010 0.830±0.010 0.835±0.011 0.844±0.023
0.763±0.005 0.776±0.004 0.764±0.007 0.796±0.003
0.834±0.003 0.834±0.003 0.834±0.002 0.857±0.005
0.867±0.005 0.843±0.007 0.863±0.008 0.852±0.005
0.783±0.015 0.837±0.015 0.789±0.021 0.858±0.003
0.818±0.002 0.805±0.004 0.821±0.004 0.832±0.004
0.888±0.003 0.880±0.006 0.887±0.003 0.902±0.000
0.845±0.004 0.841±0.002 0.849±0.009 0.848±0.015
0.700±0.008 0.670±0.009 0.705±0.022 0.739±0.030
Recall
ACC
Recall
ACC
Histopathology pretraining models CTransPath-MoCov3 Swin-tiny CHIEF Swin-tiny ViT-small Lunit-DINO PLIP ViT-base/32 MAE-TCGA-Unlabeled ViT-base SupMAE ViT-base ViT-large UNI-DINOv2 GPFM ViT-large Virchow ViT-huge
Linear-probing Method
Network
Natural-image pretraining MDFS ViT-ImageNet (Supervised) MAE-ImageNet DINOv2 CLIP
EfficientNet-b7 ViT-base ViT-base ViT-base ViT-base
ACC
MIDOGpp ROC-AUC Precision
Canine ROC-AUC Precision
TCGA-MF-Test ROC-AUC Precision
Recall
0.627±0.045 0.480±0.082 0.605±0.007 0.569±0.007 0.522±0.077
0.615±0.100 0.490±0.071 0.519±0.012 0.533±0.027 0.481±0.054
0.414±0.359 0.368±0.071 0.598±0.014 0.275±0.128 0.385±0.043
0.313±0.274 0.505±0.446 0.221±0.076 0.014±0.013 0.345±0.545
0.618±0.074 0.500±0.177 0.535±0.038 0.562±0.229 0.479±0.203
0.577±0.067 0.486±0.018 0.537±0.005 0.530±0.019 0.487±0.042
0.749±0.041 0.678±0.022 0.732±0.004 0.691±0.021 0.726±0.065
0.708±0.255 0.524±0.430 0.533±0.086 0.653±0.565 0.462±0.500
0.543±0.102 0.555±0.027 0.583±0.006 0.480±0.054 0.487±0.072
0.589±0.102 0.525±0.042 0.556±0.019 0.497±0.057 0.499±0.088
0.496±0.061 0.429±0.071 0.518±0.011 0.415±0.016 0.440±0.045
0.723±0.242 0.099±0.046 0.296±0.038 0.554±0.377 0.640±0.150
0.526±0.086 0.529±0.091 0.511±0.044 0.494±0.075 0.533±0.017 0.654±0.010 0.511±0.076 0.471±0.076 0.532±0.056
0.507±0.059 0.506±0.053 0.492±0.061 0.538±0.031 0.505±0.011 0.715±0.009 0.577±0.021 0.495±0.028 0.536±0.008
0.342±0.309 0.489±0.059 0.368±0.103 0.420±0.022 0.438±0.009 0.579±0.013 0.454±0.067 0.414±0.014 0.446±0.015
0.331±0.572 0.427±0.457 0.436±0.443 0.598±0.522 0.343±0.111 0.677±0.023 0.455±0.476 0.573±0.403 0.339±0.388
0.514±0.037 0.523±0.201 0.450±0.219 0.487±0.202 0.596±0.046 0.667±0.008 0.586±0.194 0.593±0.100 0.494±0.176
0.547±0.009 0.497±0.040 0.526±0.064 0.511±0.063 0.613±0.004 0.739±0.006 0.486±0.007 0.492±0.008 0.502±0.030
0.704±0.007 0.702±0.015 0.684±0.095 0.654±0.052 0.729±0.008 0.833±0.010 0.719±0.030 0.702±0.006 0.686±0.024
0.532±0.097 0.569±0.500 0.376±0.541 0.491±0.494 0.677±0.118 0.658±0.025 0.705±0.494 0.727±0.238 0.489±0.441
0.432±0.004 0.521±0.084 0.477±0.017 0.457±0.096 0.611±0.008 0.661±0.005 0.484±0.077 0.538±0.063 0.527±0.090
0.493±0.106 0.490±0.043 0.553±0.069 0.456±0.063 0.646±0.011 0.677±0.002 0.501±0.051 0.489±0.114 0.562±0.014
0.414±0.019 0.463±0.060 0.411±0.049 0.433±0.043 0.570±0.041 0.609±0.016 0.428±0.028 0.464±0.083 0.524±0.116
0.831±0.215 0.435±0.486 0.637±0.330 0.670±0.406 0.411±0.131 0.577±0.039 0.531±0.277 0.355±0.090 0.365±0.534
Pretraining methods using our pseudo-labeled dataset: TCGA-MF-Pseudo ViT-TCGA-MF-Pseudo ViT-base 0.854±0.001 0.925±0.001 ViT-base 0.554±0.014 0.581±0.016 MAE-TCGA-MF-Pseudo CMAE ViT-base 0.605±0.002 0.560±0.008 MiTHras (Ours) ViT-base 0.866±0.002 0.925±0.001
0.842±0.009 0.457±0.016 0.550±0.010 0.872±0.005
0.807±0.011 0.263±0.113 0.391±0.094 0.802±0.010
0.758±0.004 0.635±0.038 0.558±0.023 0.788±0.005
0.817±0.001 0.591±0.009 0.638±0.007 0.845±0.001
0.853±0.004 0.731±0.007 0.789±0.014 0.856±0.006
0.792±0.013 0.761±0.101 0.506±0.053 0.840±0.016
0.809±0.002 0.540±0.003 0.678±0.004 0.814±0.002
0.870±0.002 0.650±0.012 0.714±0.008 0.878±0.001
0.802±0.003 0.471±0.003 0.691±0.063 0.800±0.006
0.734±0.007 0.657±0.058 0.465±0.114 0.752±0.007
Histopathology pretraining models Swin-tiny CTransPath-MoCov3 CHIEF Swin-tiny Lunit-DINO ViT-small PLIP ViT-base/32 ViT-base MAE-TCGA-Unlabeled SupMAE ViT-base UNI-DINOv2 ViT-large GPFM ViT-large Virchow ViT-huge
TABLE S2: Mitotic figure detection and subtype classification results, mean ± standard deviation over three folds. The segmentation-only baseline matches the configuration reported in the main detection table. For two-stage models, the detection/classification threshold pair maximizes F1 across the three MIDOG22 validation folds and is held fixed for external evaluation; YOLOv10 uses its own inference confidence threshold. SupMAE is pretrained on labeled MIDOG22. Bold indicates the highest mean at displayed precision in each column. AMi-Br uses atypical mitotic figures as the positive class (label 1), with typical figures labeled 0. AMi-Morph uses unweighted macro F1 , precision, and recall; ROC-AUC was not evaluated. ACC denotes overall accuracy. Pretraining Method Detection baselines Segmentation only YOLOv10
Network
MIDOGpp Recall Precision
Recall
Canine Precision
TCGA-MF-Test Recall Precision
Efficient-UNet-b0 0.715±0.000 0.864±0.000 0.789±0.000 0.776±0.000 0.620±0.000 0.529±0.000 YOLOv10 0.801±0.061 0.738±0.100 0.812±0.000 0.772±0.000 0.704±0.000 0.468±0.000
ACC -
AMi-Br-TUPAC ROC-AUC Precision -
-
Recall
ACC
-
-
AMi-Morph-TUPAC Precision Recall -
-
Natural-image pretraining Bertram et al. [42] EfficientNet-V2-S 0.761±0.028 0.840±0.011 0.468±0.034 0.758±0.041 0.536±0.040 0.311±0.016 0.348±0.037 MDFS EfficientNet-b7 0.778±0.001 0.846±0.006 0.795±0.019 0.776±0.013 0.679±0.012 0.523±0.012 0.815±0.003 0.838±0.012 0.565±0.019 0.611±0.098 0.692±0.036 0.452±0.035 0.422±0.030 ViT-ImageNet (Supervised) ViT-base 0.783±0.022 0.842±0.022 0.815±0.034 0.766±0.035 0.689±0.035 0.506±0.022 0.834±0.011 0.858±0.011 0.609±0.043 0.643±0.053 0.707±0.026 0.502±0.018 0.460±0.025 MAE-ImageNet ViT-base 0.777±0.014 0.847±0.004 0.799±0.018 0.782±0.004 0.705±0.007 0.499±0.004 0.820±0.043 0.784±0.153 0.572±0.169 0.400±0.253 0.692±0.007 0.438±0.060 0.451±0.023 DINOv2 ViT-base 0.793±0.016 0.835±0.015 0.832±0.010 0.762±0.006 0.705±0.010 0.510±0.006 0.855±0.007 0.887±0.003 0.669±0.034 0.643±0.039 0.684±0.063 0.498±0.041 0.486±0.030 CLIP ViT-base 0.758±0.007 0.863±0.001 0.796±0.005 0.786±0.007 0.666±0.002 0.529±0.003 0.850±0.006 0.886±0.011 0.656±0.019 0.631±0.038 0.719±0.020 0.497±0.011 0.460±0.004 Histopathology pretraining models CTransPath-MoCov3 Swin-tiny Swin-tiny CHIEF Lunit-DINO ViT-small PLIP ViT-base/32 MAE-TCGA-Unlabeled ViT-base SupMAE ViT-base UNI-DINOv2 ViT-large GPFM ViT-large Virchow ViT-huge
0.732±0.011 0.855±0.006 0.786±0.012 0.796±0.008 0.635±0.002 0.525±0.003 0.801±0.013 0.685±0.128 0.433±0.377 0.157±0.145 0.663±0.023 0.309±0.048 0.375±0.052 0.744±0.006 0.859±0.005 0.785±0.014 0.797±0.007 0.622±0.006 0.529±0.006 0.810±0.009 0.811±0.005 0.563±0.031 0.526±0.040 0.644±0.021 0.345±0.040 0.315±0.021 0.766±0.009 0.842±0.005 0.795±0.029 0.784±0.017 0.671±0.014 0.502±0.002 0.796±0.013 0.833±0.009 0.521±0.028 0.644±0.045 0.663±0.008 0.389±0.016 0.378±0.034 0.765±0.004 0.843±0.011 0.786±0.014 0.782±0.002 0.669±0.018 0.508±0.015 0.850±0.009 0.882±0.005 0.659±0.044 0.630±0.040 0.703±0.029 0.470±0.026 0.456±0.024 0.790±0.005 0.809±0.008 0.839±0.014 0.746±0.011 0.701±0.008 0.477±0.007 0.788±0.017 0.724±0.039 0.545±0.117 0.164±0.051 0.664±0.011 0.310±0.006 0.337±0.006 0.702±0.008 0.846±0.006 0.750±0.032 0.801±0.011 0.551±0.021 0.530±0.009 0.784±0.017 0.793±0.008 0.498±0.033 0.520±0.084 0.681±0.011 0.388±0.023 0.418±0.032 0.787±0.007 0.843±0.012 0.807±0.013 0.783±0.014 0.692±0.012 0.510±0.010 0.810±0.003 0.824±0.008 0.554±0.011 0.586±0.061 0.654±0.009 0.386±0.007 0.391±0.015 0.721±0.016 0.859±0.002 0.803±0.004 0.792±0.004 0.642±0.004 0.520±0.008 0.847±0.003 0.890±0.006 0.636±0.020 0.674±0.046 0.701±0.025 0.504±0.015 0.472±0.030 0.792±0.007 0.842±0.006 0.842±0.009 0.755±0.005 0.690±0.008 0.514±0.003 0.781±0.009 0.829±0.023 0.492±0.015 0.678±0.033 0.611±0.038 0.371±0.022 0.345±0.020
Pretraining methods using our pseudo-labeled dataset: TCGA-MF-Pseudo ViT-TCGA-MF-Pseudo ViT-base 0.761±0.005 0.863±0.001 0.793±0.005 0.789±0.002 0.669±0.007 0.522±0.002 0.806±0.008 0.828±0.001 0.555±0.037 0.527±0.092 0.627±0.018 0.262±0.006 0.330±0.010 MAE-TCGA-MF-Pseudo ViT-base 0.803±0.007 0.804±0.006 0.838±0.010 0.750±0.005 0.668±0.006 0.496±0.003 0.787±0.001 0.637±0.058 0.527±0.016 0.057±0.035 0.580±0.109 0.218±0.091 0.260±0.101 CMAE ViT-base 0.810±0.003 0.810±0.006 0.797±0.013 0.774±0.005 0.699±0.009 0.498±0.013 0.832±0.017 0.857±0.008 0.620±0.059 0.583±0.062 0.697±0.010 0.346±0.014 0.407±0.008 MiTHras (Ours) ViT-base 0.798±0.019 0.864±0.018 0.810±0.010 0.786±0.010 0.692±0.023 0.529±0.017 0.870±0.011 0.914±0.004 0.696±0.052 0.710±0.062 0.713±0.009 0.511±0.006 0.509±0.020