Balanced Prompt Adaptation against Entropy-Induced Collapse for Test-Time Binary Segmentation Zhengshan Wang, Joshua Charles Webster-Ford, Yifei Tian, Xinxin Wang1∗ , Long Chen, Weiping Ding 1
arXiv:2609.21743v1 [cs.CV] 18 Sep 2026
Shenzhen University, Shenzhen, China Corresponding author: [email protected] ORCID: https://orcid.org/0009-0000-6065-7651 Abstract Entropy minimization is a standard objective for test-time adaptation (TTA), but it can fail in imbalanced binary segmentation. Unlike image classification, dense segmentation aggregates thousands of pixel predictions, allowing the larger predicted class to dominate the update, pull minority predictions toward itself, and produce a degenerate mask as predictions saturate and their entropy gradients vanish. We theoretically establish this collapse in a shared-shift model. This analysis motivates Balanced-Anchor Prompt Adaptation (BAPA), which combines two complementary modules. The Class-Balanced Anchors (CBA) module selects highconfidence anchors separately from each predicted class and gives foreground and background equal total loss weight, preventing the larger region from dominating the update. Dynamic Prompt Adaptation (DPA) refreshes these anchors after each prediction update and optimizes only text-side prompt residuals while keeping the vision–language encoders frozen. This prompt-only update refines the foreground–background decision boundary without altering the pretrained dense visual representation. Across experiments from four domains, BAPA achieves the highest mean Dice among the evaluated methods. Factorized ablations further validate the complementary roles of CBA and DPA, supporting balanced prompt adaptation as an effective alternative to entropy minimization for test-time binary segmentation.
Introduction Test-time adaptation (TTA) seeks to recover model performance under distribution shift without labeled target data (Wang et al. 2021; Liang, He, and Tan 2025). It is especially attractive for segmentation systems deployed across domains, because medical lesions, salient objects, pets, and open-vocabulary object categories often differ substantially from the data used to train the base model. Vision–language models (VLMs) further broaden this setting: a user can specify a foreground concept by text and obtain a binary mask without task-specific training (Radford et al. 2021; Noori et al. 2025). The remaining question is how to adapt such source-free predictions safely at test time. A common TTA strategy is to minimize prediction entropy (Wang et al. 2021; Noori et al. 2025; Niu et al. 2023, 2022; Shu et al. 2022), encouraging the model to make more ∗
Corresponding author.
Image
Mask Class imbalance FG
Zero shot
Entropy minimization
Collapse
BG
All 0 or 1
BG FG
Class-Balanced Anchors Dense feature
Collapse BAPA
Visual
filter
Dynamic Prompt Adaptation Prompt
Learnable Prompts Text
Text prototypes
Figure 1: Limitations of entropy-based TTA and balanced prompt adaptation. In imbalanced binary segmentation, entropy-based adaptation can amplify the currently dominant predicted class and drive the mask toward a degenerate solution. BAPA addresses this failure through Class-Balanced Anchors (CBA) and Dynamic Prompt Adaptation (DPA). BG and FG denote background and foreground, respectively.
confident predictions. This objective is attractive because it requires neither source data nor target annotations, but confidence alone cannot tell whether a prediction is correct. In image classification, entropy is applied to a small number of image-level outputs; in segmentation, the same objective aggregates thousands of pixel decisions. The resulting update is therefore shaped not only by uncertainty, but also by how many pixels are currently assigned to each class. This limitation becomes more severe in binary segmentation, where the foreground often occupies only a small part of the image. Entropy minimization pushes each pixel to become more confidently background (BG) or foreground (FG), whether its current prediction is correct or not; errors can therefore be sharpened rather than corrected. Because the same adaptation update is computed from all pixels, the majority prediction can also move minority pixels in the same direction, producing nearly all-background or all-foreground masks. This degeneration is referred to as entropy-induced collapse. This collapse is driven by two properties of the entropy objective. First, entropy minimization rewards confident background or foreground predictions, regardless of whether they are correct. Second, once an incorrect prediction becomes
highly confident, the entropy gradient becomes weak, leaving little signal to undo the error. In imbalanced masks, the majority class can dominate the aggregate update and pull minority pixels in the same direction. This makes the failure different from ordinary overfitting to pseudo-labels: the objective itself can favor a degenerate but highly confident mask. Stabilization methods such as entropy filtering with SAR (Niu et al. 2023), Fisher regularization with EATA (Niu et al. 2022), and weight averaging can alter the trajectory, but they still use entropy as the adaptation signal and do not provide class-balanced corrective targets. The analysis leads to two complementary modules. The Class-Balanced Anchors (CBA) module selects confident pixels separately from the current background and foreground predictions and gives the two classes equal total loss weight. This does not assume equal foreground and background areas; it only prevents the larger predicted region from exerting greater optimization influence. Dynamic Prompt Adaptation (DPA) then reconstructs the anchors after each prediction update while optimizing only two text-side residuals and keeping the vision–language encoders frozen. The dynamic refresh allows improved intermediate masks to provide updated supervision, while the prompt-only update adjusts the foreground–background semantic boundary without altering dense visual features. Together, CBA and DPA form BAPA and turn adaptation from entropy-driven confidence sharpening into balanced correction of the current binary decision boundary. Figure 1 summarizes the motivation of the paper and illustrates how CBA and DPA jointly counter entropy-induced collapse. This work makes three contributions:
et al. 2022) applies entropy minimization to prompt embeddings over augmented views. Orthogonality-constrained prompt tuning further improves VLM calibration (Sharifdeen et al. 2025). Recent benchmarks show that VLM-TTA gains depend strongly on protocol, update parameterization, and reliability metrics (Sheng et al. 2025; Huang et al. 2026). These methods differ in parameterization and stabilization, but none directly supplies balanced foreground–background targets.
1. We identify entropy-induced collapse as a failure mode of entropy-based TTA in imbalanced binary segmentation and connect it to majority-class dominance in the adaptation signal. 2. We derive Class-Balanced Anchors (CBA) from this analysis. CBA removes pixel-count dependence by giving foreground and background anchors equal total influence. 3. We introduce Dynamic Prompt Adaptation (DPA), which refreshes the anchors while updating only text-side residuals over frozen vision–language encoders. Together, CBA and DPA form BAPA; dataset-wide experiments on 16,486 image–task instances across four domains validate both modules.
Recent VLM and medical TTA
Related Work Entropy-based TTA TTA began from single-sample self-supervised updates (Sun et al. 2020), entropy minimization (Wang et al. 2021), and augmentation-based marginal entropy minimization (Zhang, Levine, and Finn 2022). Later methods improve reliability through output-space adaptation (Boudiaf et al. 2022), continual restoration and ensembling (Wang et al. 2022), sharpness-aware filtering (Niu et al. 2023), Fisher regularization (Niu et al. 2022), and weight averaging (Osowiechi et al. 2024). MLMP (Noori et al. 2025) adapts LayerNorm parameters for vision–language segmentation, while TPT (Shu
Open-vocabulary segmentation CLIP (Radford et al. 2021) and dense extensions (Li et al. 2022; Rao et al. 2022; Xu et al. 2022; Liang et al. 2023; Xu et al. 2023; Cho et al. 2024; Noori et al. 2025) transfer language-aligned representations to pixel prediction. Their open-vocabulary capability supports source-free deployment, but zero-shot masks can be poorly calibrated under domain shift. BAPA leaves the visual representation fixed and adapts only the binary text decision boundary. Recent work narrows the gap between classification TTA and dense prediction. Seg-TTO (De Silva et al. 2025) jointly optimizes segmentation-specific visual attributes and multiple textual embeddings on top of CAT-Seg or CLIP-DINOiser. MLMP (Noori et al. 2025) is the closest architecture-compatible method because it supports singleimage adaptation with dense NACLIP predictions; its complete official adaptation procedure is evaluated as a separate baseline. We discuss Seg-TTO as related segmentation TTA, but exclude it from the controlled comparison because it relies on a different segmentation architecture and checkpoints.
Recent VLM-TTA methods also explore cache-based dynamic adapters (Karmanov et al. 2024), distributional adapters (Han et al. 2025), transductive objectives (Zanella, Gérin, and Ben Ayed 2024), prompt-free test-time augmentation (Zanella and Ben Ayed 2024; Farina et al. 2024), statistical anchoring (Zanella et al. 2025), contrastive objectives (Lafon et al. 2025), Bayesian class-prior adaptation (Zhou et al. 2025), noisy/open-set adaptation (Cao et al. 2025), black-box prompt search (Meng et al. 2025), debiased prompt optimization (Song et al. 2026), and view-filtered or robustness-oriented prompt tuning (Choi and Kim 2026; Kim and Um 2026; Zhu et al. 2026). Medical and dense prediction TTA further use image normalization, shape priors, episodic feature adaptation, or low-rank VLM updates (Karani et al. 2021; Bateson, Lombaert, and Ben Ayed 2022; Valanarasu et al. 2022; Noori et al. 2026). Many of these methods assume image-level classification, augmented-view batches, persistent state, task-specific priors, or different segmentation backbones, whereas this work studies episodic single-image dense adaptation.
Pseudo-label adaptation Pseudo-labeling (Lee 2013), self-training (Xie et al. 2020), and confidence selection (Sohn et al. 2020; Zou et al. 2018, 2019) use model predictions as supervision, but iterative updates can amplify early errors through confirmation bias.
A. Frozen Dense Inference (Zero-shot)
B. Class-Balanced Anchors (CBA)
C. Dynamic Prompt Adaptation (DPA)
Eq.8
Eq.9
Test image x Visual encoder
F Eq.1
Text branch BG/ FG prompt templates a photo without skin lesion a photo of skin lesion
BG candidate region
Eq.1
Confidence selection (Eq.7)
Text encoder
Eq.10
CBA
FG candidate region CBA
Frozen
Learnable
BG anchors
FG anchors
Data flow
Gradient flow
Figure 2: BAPA framework. A frozen dense vision–language model (VLM) produces the zero-shot prediction. The ClassBalanced Anchors (CBA) module constructs foreground and background anchors through classwise top-K confidence selection and equalizes their total loss contributions. Dynamic Prompt Adaptation (DPA) optimizes prompt residuals and refreshes the prediction, pseudo-labels, confidence, and anchors after every step, while both encoders and the dense visual representation remain fixed. BG and FG denote background and foreground, respectively. This risk is acute in single-image binary segmentation, where pixel imbalance lets one predicted class dominate the update. CBA prevents this dominance through classwise selection and loss normalization, while DPA refreshes the resulting anchors as the prediction evolves.
Method Problem Definition We consider episodic test-time adaptation for binary segmentation, where a model receives one unlabeled test image x and is reset before the next image–task instance. Class-specific prompt templates are encoded, averaged, and L2-normalized to obtain the background and foreground text prototypes t0 and t1 . Stacking them gives T = [t0 , t1 ] ∈ R2×d . The frozen dense vision–language model produces fv (x)T⊤ , (1) τ where fv (x) ∈ RHW ×d is the dense visual feature map and τ = exp(−ℓscale ) > 0 is the inverse of the frozen checkpoint logit scale. The logits define per-pixel probabilities Phw = softmax(Lhw ), hard predictions Ŷhw = arg maxc Phw,c , and confidence scores Chw = maxc Phw,c . No evaluation labels or source samples are available during adaptation. The objective is to improve the binary mask by updating a restricted set of parameters while preserving the pretrained dense visual representation. L=
Overall Framework Figure 2 summarizes BAPA as two complementary modules. CBA prevents the larger predicted region from domi-
nating the loss by selecting confident foreground and background anchors separately and equalizing their total contributions. DPA reconstructs those anchors after each update and changes only the text prompts, leaving the dense visual features fixed.
Analysis of Entropy-Induced Collapse Entropy geometry. Consider a binary segmentation model producing per-pixel logits zhw and foreground probability phw = σ(zhw ) via the sigmoid function. The standard TTA objective minimizes per-pixel entropy: H(p) = −p log p − (1 − p) log(1 − p)
(2)
H(p) has a unique local maximum at p = 0.5 (H(0.5) = log 2) and two global minima at p = 0 and p = 1 (H(0) = H(1) = 0). The gradient with respect to the logit is: 1 − σ(z) dH = σ(z)(1 − σ(z)) log dz σ(z)
(3)
For |z| → ∞, σ(z)(1 − σ(z)) ∼ e−|z| and | log 1−σ(z) σ(z) | ∼ |z|. Hence, |dH/dz| = O(|z|e−|z| ). Entropy minimization supplies progressively less signal as a prediction becomes confident, including when it is incorrect. Once an erroneous prediction becomes highly confident, a finite number of adaptation steps may provide too little gradient to correct it. Majority-dominated updates. Let θ denote the adapted parameters and zi (θ) the logit P at pixel i. For the mean-entropy objective Lent = n−1 P i H(σ(zi )), the parameter gradient is ∇θ Lent = n−1 i (dH/dzi )∇θ zi . This sum explains
BG:FG ratio n0 : n1 1:1 2:1 5:1
0.6
gradient descent
0.4 0.2 0.0
initial b0 = 0
FG predicted as BG
−4
−3
−2
FG decision boundary
Mean entropy, Lr(b)
(a) Shared-bias entropy landscape
−1
0
1
2
Predicted BG:FG ratio, n0/n1
Shared logit shift, b 12
(b) Sufficient drift regime
10
sufficient background drift
8
Prop. S1 boundary
6 4
not guaranteed a = 0.6, ε = 0.1
2 0.0
0.2
0.4
0.6
Jacobian heterogeneity, ρ
Figure 3: Majority-dominated entropy updates and local drift condition. (a) Increasing n0 /n1 tilts the entropy landscape toward negative shared shifts; b = −a is the foreground decision boundary. (b) For a = 0.6 and ϵ = 0.1, the shaded upper region satisfies Supplementary Proposition S1. why pixel counts matter. If most pixels are currently predicted as background, their entropy gradients can dominate the update, especially when many pixels respond similarly to the adapted parameters. The resulting update can make background predictions more confident and move uncertain foreground pixels toward background. Class imbalance does not by itself guarantee collapse, because pixel-wise sensitivities also matter, but it provides a concrete and measurable source of majority-class bias. We first study this effect in the simplest setting where the adapted update moves all pixel logits in a common direction. Suppose n0 pixels are initially predicted as background with logit −a, and n1 pixels are initially predicted as foreground with logit +a, where a > 0 measures the initial confidence margin and n = n0 + n1 . A scalar parameter b, initialized at b0 = 0, shifts every logit by the same amount: zi (b) ∈ {−a + b, a + b}. This shared shift is a simplified model of a prompt or normalization update that moves the foreground–background decision boundary. For the background-to-foreground count ratio r = n0 /n1 , Figure 3(a) visualizes Lr (b) =
rH(σ(−a + b)) + H(σ(a + b)) . r+1
(4)
Define f (x) = xσ(x)(1−σ(x)) for x ≥ 0, and let x⋆ ≈ 1.54 be the positive solution of x tanh(x/2) = 1, so that f is
increasing on [0, x⋆ ]. Theorem 1 (Majority-induced collapse under a shared update). Assume n0 > n1 , b0 =P 0, and 0 < a ≤ x⋆ /2. For n the mean entropy L(b) = n−1 i=1 H(σ(zi (b))), gradient ′ descent bt+1 = bt − ηL (bt ) with any η > 0 satisfies: (i) bt decreases; (ii) it reaches bt < −a in finitely many steps; and (iii) bt → −∞, so every foreground probability converges to zero. If n0 = n1 , then L′ (0) = 0; if n1 > n0 , the symmetric all-foreground result holds. Proof sketch. For b ∈ [−a, 0], L′ (b) = [n0 f (a − b) − n1 f (a + b)]/n > 0 by monotonicity of f and n0 > n1 . The derivative remains positive for b < −a, so bt → −∞. The complete proof is provided in Supplementary Theorem S1. The theorem captures a sufficient mechanism: class imbalance destabilizes the symmetric stationary point, and a shared parameter direction propagates the majority decision to every pixel. The common-shift assumption can be relaxed locally. For a unit parameter direction v, let ji = ⟨∇θ zi (θ0 ), v⟩ denote the projected logit Jacobian of pixel i. Let ρ measure the heterogeneity of these projected responses, with smaller ρ indicating stronger alignment, and let ϵ bound the local logit deviation from ±a. Supplementary Proposition S1 shows that, when the projected Jacobians are sufficiently aligned, a large enough predicted background-to-foreground ratio n0 /n1 makes the entropy-gradient step drift toward background. This is a sufficient local condition, but it gives a testable prediction: real text-residual entropy updates should contain a common foreground–background drift component, which we measure below. Figure 3 visualizes the shared-shift result and its local extension. Panel (a) shows that a background majority tilts the mean-entropy landscape toward negative shared shifts, whereas balanced predictions keep the landscape symmetric around b = 0. Panel (b) gives the sufficient local drift condition: higher Jacobian heterogeneity ρ requires stronger class imbalance to certify drift. The shaded region is sufficient rather than necessary. Empirically, ISIC, VOC, and DUTS often exhibit background-majority zero-shot masks, while OxPet gives the symmetric foreground-majority case. The observed results follow the same pattern. Entropybased prompt tuning fails severely on ISIC, reaching 0.006 Dice, and LayerNorm entropy adaptation also reduces zeroshot Dice from 0.486 to 0.427. On less imbalanced VOC masks, the same adaptation improves zero-shot performance, while OxPet shows that class-balanced supervision is also useful under foreground-majority imbalance. Entropy-based stabilizers may change the trajectory, but they retain confidence seeking as the supervision signal and do not directly balance foreground and background contributions. Class-balanced correction. The preceding analysis suggests replacing confidence-only entropy with an anchor-based objective whose foreground and background terms have equal total influence. For a pseudo-label anchor y held fixed during one update, cross-entropy (CE) gives LCE (p, y) = −y log p − (1 − y) log(1 − p),
dLCE = p−y dz (5)
Therefore, CE remains corrective when a confident prediction disagrees with its anchor. Class balancing then removes the pixel-count bias by averaging CE within each predicted class before combining the classes. In the shared-shift model, if the selected anchors are fixed locally and correct, the resulting objective is
Dynamic Prompt Adaptation A learnable residual δ (s) ∈ Rd is added to the frozen text c prototype tc of each class c ∈ {0, 1} and initialized as δ (0) c = 0. The resulting prototype is then L2-normalized:
where softplus(x) = log(1 + ex ). The two terms are the background and foreground anchor losses, each weighted by 1/2. The unique minimizer is b⋆ = 0, so the shared shift is no longer biased by the relative number of background and foreground pixels. Supplementary Theorem S2 gives the convexity and gradient-descent convergence proof. This local result motivates CBA, while DPA reapplies the same class-balanced construction after each prediction update.
Class-Balanced Anchor Construction At adaptation step s, CBA constructs anchors from the cur(s) (s) (s) rent prediction. Let Ŷhw = arg maxc Phw,c and Chw = (s)
maxc Phw,c denote the hard predicted class and confidence at pixel (h, w). For each predicted class c ∈ {0, 1}, define (s) (s) Ic = {(h, w) : Ŷhw = c}. CBA keeps only the top-K most confident pixels within each predicted class. Specifi(K,s) (s) is the confidence threshold cally, if Ic is nonempty, qc (s) that selects the top-K fraction of pixels in Ic ; if no pixel (s) is predicted as class c, then Ac = ∅. The classwise anchor sets are n o (s) (s) (K,s) A(s) , c = (h, w) ∈ Ic : Chw ≥ qc (s)
(s)
(7)
A(s) = A0 ∪ A1 . The anchor cardinalities may differ because the predicted regions can have different sizes, and DPA updates the anchors (s) (s) (s) after every adaptation step. Here, Phw = [Phw,0 , Phw,1 ] denotes the current background–foreground probability vector at pixel (h, w), obtained by applying softmax to the step-s (s) (s) logits. Accordingly, LCE (Phw , c) = − log Phw,c . To remove the residual count imbalance, CBA averages the anchor loss within each class and then weights the two classwise means equally:
tc + δ (s) c
t(s) c =
Lbal (b) = 12 softplus(−a + b) + 12 softplus(−a − b) (6)
, c ∈ {0, 1}, ∥tc + δ (s) c ∥2 λ (s) L(s) = Lanchor + ∥δ (s) ∥22 . 2 (s)
(9) (10)
(s)
Stacking t0 and t1 as rows gives T(s) ∈ R2×d . Thus, Eq. 9 normalizes each class prototype independently, rather than normalizing the two-class matrix jointly. Here, δ (s) stacks the P Pd (s) two class residuals, and ∥δ (s) ∥22 = c∈{0,1} j=1 (δc,j )2 measures their total squared magnitude. We set λ = 10−2 and implement this penalty through Adam weight decay, without adding it a second time to the optimized loss. The factor 1/2 makes its conceptual gradient λδ (s) . Dynamic adaptation cycle. At step s, δ (s) defines T(s) and hence the current prediction P(s) . DPA recomputes Ŷ(s) (s) and C(s) = [Chw ]hw from this prediction, applies CBA (s) to reconstruct Ac , treats the resulting pseudo-labels and anchor indices as stop-gradient supervision, and updates only the prompt residual: δ (s+1) ← Adam δ (s) , ∇δ L(s) . (11) The updated residual changes the foreground–background similarity boundary while leaving the dense visual representation unchanged. A new prediction then updates the supervision for step s + 1, yielding the cycle δ (s) → P(s) → A(s) → L(s) → δ (s+1) . Adam (Kingma and Ba 2015) performs S = 20 updates with learning rate 10−3 , β = (0.9, 0.999), and weight decay 10−2 .
Design rationale. CBA removes region-size scaling from the supervision. DPA updates that supervision as the prediction evolves, but restricts optimization to the binary text boundary so that the pretrained dense representation remains unchanged.
Experiments Setup
(s)
Lanchor =
X
1
(s) c∈{0,1}: 2|Ac | |A(s) c |>0
(s)
X
LCE (Phw , c). (8) (s)
(h,w)∈Ac
Consequently, whenever both predicted classes are present, background and foreground each contribute to(s) (s) tal weight 1/2, irrespective of |A0 |/|A1 |; an absent class contributes zero for that update. Classwise selection preserves confident evidence from both predicted classes, whereas Eq. 8 removes their pixel-count dependence in nondegenerate updates. We use K = 20% throughout.
Model and datasets. All dataset-wide results use NACLIP ViT-L/14 (Hajimiri et al. 2025) through the implementation distributed with MLMP (Noori et al. 2025), without offline fine-tuning. We evaluate 16,486 image–task instances from the ISIC 2017 training split (Codella et al. 2018), PASCAL VOC 2012 (Everingham et al. 2010), DUTS-TE (Wang et al. 2017), and Oxford-IIIT Pet (Parkhi et al. 2012) at 2242 resolution. VOC uses all 20 foreground classes and reports both the class macro-average and the pooled instance-level mean. Protocol and baselines. All methods use the same evaluation manifests, reset for each image–task instance, and use no evaluation labels for model selection. We compare ZS,
Method
ISIC VOC DUTS OxPet Mean
ZS (2025) 0.486 0.532 0.200 PL (2013) 0.246 0.450 0.189 TENT (2021) 0.427 0.542 0.184 EATA (2022) 0.466 0.536 0.194 TPT (2022) 0.006 0.253 0.087 SAR (2023) 0.485 0.532 0.200 WATT (2024) 0.295 0.332 0.340 CLIPArTT (2025) 0.287 0.337 0.283 MLMP (2025) 0.489 0.493 0.241 O-TPT (2025) 0.006 0.304 0.087 BAPA (Ours) 0.587 0.578 0.296
0.610 0.457 0.484 0.342 0.627 0.445 0.614 0.453 0.700 0.261 0.610 0.457 0.529 0.374 0.516 0.356 0.723 0.486 0.708 0.276 0.783 0.561
Table 1: Dataset-wide Dice ↑. Parenthesized years cite each method; zero-shot inference (ZS) and the pseudo-label baseline (PL) precede the TTA methods. VOC pools all 2,077 eligible image–class instances; Supplementary Table S5 reports its macro-average. Mean averages the four unrounded dataset scores.
PL (Lee 2013), TENT (Wang et al. 2021), MLMP (Noori et al. 2025), TPT (Shu et al. 2022), O-TPT (Sharifdeen et al. 2025), SAR (Niu et al. 2023), EATA (Niu et al. 2022), CLIPArTT (Vargas Hakim et al. 2025), WATT (Osowiechi et al. 2024), and BAPA under a unified dense singleimage protocol. MLMP, CLIPArTT, and WATT use the dense open-vocabulary semantic segmentation (OVSS) adaptation classes distributed with the MLMP implementation. TENT and SAR call the official update routines through a denselogit wrapper; EATA, TPT, and O-TPT use dense analogues of their official objectives because their original entry points assume image-level samples. BAPA uses one setting across datasets: K = 20%, S = 20, learning rate 10−3 , and weight decay 10−2 . The Supplementary section Experimental Protocol and Reproducibility specifies the manifests, implementation, and baseline ports.
Main Results Dataset-wide comparison. Table 1 uses every eligible VOC class for every method and reports one main Dice score per dataset. BAPA improves VOC pooled Dice from 0.532 to 0.578 and obtains the highest mean Dice across the four datasets. Across ISIC, VOC, and OxPet, BAPA also achieves the best individual dataset score; WATT is strongest on DUTS. Supplementary Tables S3–S5 report the Oxford-Pet per-breed breakdowns and the complete VOC macro and perclass results; they show that BAPA improves every OxfordPet breed, raises VOC macro Dice from 0.517 to 0.575, and improves 14 of the 20 VOC classes. Qualitative results. Figure 4 complements the dataset-wide statistics with a single enlarged ISIC case study. TENT removes the lesion prediction, whereas BAPA recovers a localized foreground structure with substantially higher Dice. Supplementary Figure S1 provides the full qualitative panel with additional ISIC and VOC examples. Paired inference. Each BAPA result is paired with predictions from the comparison method on the same image–task instance. Relative to ZS, the 95% paired-bootstrap intervals
Input
Reference
Zero-Shot
Dice 0.038 TENT
BAPA
Dice 0.000
Residual Error
Dice 0.890 TP
FP
FN
Figure 4: Compact qualitative result. Panels show the input, reference mask, zero-shot prediction, TENT, BAPA, and BAPA residual errors for one representative ISIC lesion. Green, orange, and purple denote true positives, false positives, and false negatives in the residual-error panel.
are above zero on every dataset. We also compare BAPA with the strongest non-BAPA method in each Table 1 dataset column. For ∆ = DiceBAPA − Dicebaseline , the intervals are ISIC versus MLMP [0.088, 0.109], VOC versus TENT [0.028, 0.043], DUTS versus WATT [−0.052, −0.037], and OxPet versus MLMP [0.058, 0.062]. All four differences remain significant after Holm correction over all 40 BAPA– comparator tests (pHolm = 0.004). Thus, BAPA’s highest four-dataset mean reflects significant gains on three datasets together with a significant deficit to WATT on DUTS, rather than uniform dominance. Resampling uses source images as independent units and clusters all VOC class tasks from the same image; the complete protocol and ZS comparisons are reported in the Supplementary Material. Mechanism diagnostic. We decompose the entropy signal by defining gc as the gradient of mean entropy over pixels currently predicted as class c, with c = 0 for background and c = 1 for foreground. Among zero-shot masks containing both classes, the class-count-weighted log-ratio log10 (n0 ∥g0 ∥2 /(n1 ∥g1 ∥2 )) averages 1.24 on ISIC, equivalent to a 17.21× geometric background/foreground ratio; the ratios are 3.71 on VOC, 2.71 on DUTS, and 1.42 on OxPet. Its descriptive Pearson correlation with the binary zero-shot collapsed-mask indicator, defined using the 1% foreground-fraction criterion in the Supplementary Material, is r = 0.72, 0.45, 0.28, and −0.29, respectively. The negative OxPet value reflects the direction of the backgroundover-foreground ratio in a foreground-majority regime. The two gradients are strongly opposed on all datasets (mean cosine −0.87 to −0.94). All-BG and all-FG masks are excluded only where the ratio is undefined and remain in collapse-rate calculations. This imbalance could still arise from unrelated pixelwise effects rather than a shared update direction. We therefore
Dataset
Ecom
¯ ∆z
FG→BG
ISIC VOC DUTS OxPet
99.2% 93.7% 95.2% 88.7%
−0.121 −0.077 −0.041 −0.032
99.6% 82.7% 66.4% 65.4%
Table 2: Prompt-space drift diagnostic. Entries are dataset means of per-instance quantities. Ecom is the fraction of ¯ is logit-change energy in the constant-pixel direction; ∆z the mean per-pixel logit change. These two metrics use every instance in each complete manifest. FG→BG is averaged over instances with a nonempty foreground prediction. Bal. Policy Params ISIC VOC DUTS OxPet Mean ✓ – ✓ – ✓ – ✓ –
F F D D F F D D
P P P P LN LN LN LN
0.586 0.589 0.246 0.214 0.446 0.172 0.587 0.578 0.296 0.003 0.247 0.084 0.559 0.564 0.231 0.395 0.526 0.169 0.563 0.565 0.234 0.392 0.525 0.164
0.780 0.550 0.477 0.328 0.783 0.561 0.476 0.203 0.668 0.505 0.583 0.418 0.669 0.508 0.583 0.416
Table 3: Dataset-wide 2 × 2 × 2 factorized ablation (Dice ↑). Bal.=CBA; F=fixed zero-shot labels and anchor locations; D=dynamic anchor updates; P=prompt residuals; LN=LayerNorm. BAPA corresponds to the balanced D+P row. VOC pools all 2,077 eligible instances; Mean averages the four unrounded dataset scores.
take one entropy-gradient step in the text-residual space δ ∈ R2×d and measure the resulting foreground-minusbackground logit change ∆zhw at each pixel. A negative ∆zhw means that the pixel is pushed toward background. If the local theory captures the real update, many pixels should move together, so most of the ∆z energy should lie in a constant-pixel component. Table 2 shows exactly this pattern: the per-instance common-mode energy fraction averages 88.7–99.2% across datasets, with the strongest drift on ISIC. Thus, the entropy gradient in the actual prompt space contains the shared foreground–background drift predicted by the local theory. This diagnostic complements the correlation analysis by evaluating the entropy-induced direction in the actual prompt parameterization. The high Ecom values show that the step contains a coherent foreground–background logit shift rather than only unrelated pixelwise changes. Its negative mean on all four datasets indicates a background-directed shared component, largest on ISIC, where the gradient-count imbalance is also strongest. These observations connect the sufficient local theory to real text-residual updates.
Factorized Ablation Table 3 factorizes CBA from the two choices that define DPA: dynamic anchor updates and prompt-residual optimization. With both DPA choices fixed, adding CBA changes Dice from 0.003 to 0.587 on ISIC, 0.247 to 0.578 on VOC, 0.084
to 0.296 on DUTS, and 0.476 to 0.783 on OxPet. Within the CBA rows, enabling dynamic refresh instead of fixed anchors changes Dice by +0.001, -0.011, +0.050, and +0.003. Prompt residuals also outperform the matched LayerNorm updates in every class-balanced setting. These comparisons separately support the balancing role of CBA and the update design of DPA.
Robustness and Scope The complete BAPA configuration, combining CBA with DPA, has the highest Dice on DUTS and OxPet and is within 0.002 of the best ablation row on ISIC and within 0.011 on VOC. Sensitivity analysis shows small Dice variation across weight decay and 10–30 update steps, while the aggressive learning rate 5 × 10−3 gives the lowest Dice in every dataset column. Scope and limitations. The theory establishes collapse and local drift along aligned update directions, and the promptspace diagnostic measures this component in real textresidual updates. Supplementary Theorem S2 covers anchorstable regions; the coupled dynamic regime is evaluated empirically by the factorized ablation. The evaluation uses one backbone family and segmentation-adapted versions of several baselines; additional architectures and native multi-class segmentation remain future work. Implications. The central finding is that entropy minimization can be structurally misaligned with imbalanced binary segmentation. CBA provides the main stabilization by preventing either predicted class from dominating the anchor loss. On this balanced objective, DPA can refresh supervision from improved intermediate predictions while restricting the update to text prompts and preserving dense visual features. BAPA combines these two roles in a single adaptation procedure.
Conclusion Entropy minimization can amplify the majority prediction in imbalanced binary segmentation and drive a mask toward a degenerate state. The local analysis and promptspace diagnostic connect this failure to a shared foreground– background drift. BAPA addresses the two resulting requirements through CBA, which equalizes foreground and background supervision, and DPA, which refreshes that supervision while adapting only text prompts over frozen encoders. Across four datasets, BAPA improves zero-shot Dice in every case, attains the highest mean Dice, and leads on three datasets. These results show that dense test-time adaptation benefits from combining class-balanced supervision with restricted dynamic updates.
References Bateson, M.; Lombaert, H.; and Ben Ayed, I. 2022. Test-Time Adaptation with Shape Moments for Image Segmentation. In Medical Image Computing and Computer Assisted Intervention (MICCAI), 736–745. Boudiaf, M.; Mueller, R.; Ben Ayed, I.; and Bertinetto, L. 2022. Parameter-Free Online Test-Time Adaptation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR).
Cao, C.; Zhong, Z.; Zhou, Z.; Liu, T.; Liu, Y.; Zhang, K.; and Han, B. 2025. Noisy Test-Time Adaptation in Vision-Language Models. In International Conference on Learning Representations (ICLR). Cho, S.; Shin, H.; Hong, S.; Arnab, A.; Seo, P. H.; and Kim, S. 2024. CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Choi, J.; and Kim, E. 2026. Dual-Modality Anchor-Guided Filtering for Test-Time Prompt Tuning. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Findings, 9368– 9377. Codella, N. C. F.; et al. 2018. Skin Lesion Analysis Toward Melanoma Detection: A Challenge at the 2017 International Symposium on Biomedical Imaging (ISBI), Hosted by the International Skin Imaging Collaboration (ISIC). In IEEE International Symposium on Biomedical Imaging (ISBI), 168–172. De Silva, U.; Samaraweera, D.; Wanigathunga, S.; Kariyawasam, K.; Ranasinghe, K.; Naseer, M.; and Rodrigo, R. 2025. Test-Time Optimization for Domain Adaptive Open Vocabulary Segmentation. arXiv:2501.04696. Everingham, M.; Van Gool, L.; Williams, C. K. I.; Winn, J.; and Zisserman, A. 2010. The PASCAL Visual Object Classes (VOC) Challenge. International Journal of Computer Vision (IJCV), 88(2): 303–338. Farina, M.; Franchi, G.; Iacca, G.; Mancini, M.; and Ricci, E. 2024. Frustratingly Easy Test-Time Adaptation of Vision-Language Models. In Advances in Neural Information Processing Systems (NeurIPS), volume 37. Hajimiri, S.; et al. 2025. Pay Attention to Your Neighbours: Training-Free Open-Vocabulary Semantic Segmentation. In IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 5061–5071. Han, Z.; Yang, J.; Wang, G.; Li, J.; Xu, Q.; Shou, M. Z.; and Zhang, C. 2025. DOTA: Distributional Test-Time Adaptation of VisionLanguage Models. In Advances in Neural Information Processing Systems (NeurIPS), volume 38. Huang, J.; Chen, X.; Liu, Z.; Sun, Y.; Jiang, J.; and Wang, Z. 2026. What Drives Test-Time Adaptation for CLIP? A Controlled Empirical Study from an Update Perspective. arXiv:2606.14299. Karani, N.; Erdil, E.; Chaitanya, K.; and Konukoglu, E. 2021. TestTime Adaptable Neural Networks for Robust Medical Image Segmentation. Medical Image Analysis, 68: 101907. Karmanov, A.; Guan, D.; Lu, S.; El Saddik, A.; and Xing, E. 2024. Efficient Test-Time Adaptation of Vision-Language Models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 14162–14171. Kim, S.; and Um, D. 2026. SS-TPT: Stability and SuitabilityGuided Test-Time Prompt Tuning for Adversarially Robust VisionLanguage Models. arXiv:2606.06943. Kingma, D. P.; and Ba, J. 2015. Adam: A Method for Stochastic Optimization. In International Conference on Learning Representations (ICLR). Lafon, M.; Vargas Hakim, G.; Rambour, C.; Desrosiers, C.; and Thome, N. 2025. CLIPTTA: Robust Contrastive Vision-Language Test-Time Adaptation. In Advances in Neural Information Processing Systems (NeurIPS), volume 38. Lee, D.-H. 2013. Pseudo-Label: The Simple and Efficient SemiSupervised Learning Method for Deep Neural Networks. In ICML Workshop on Challenges in Representation Learning. Li, B.; Weinberger, K. Q.; Belongie, S.; Koltun, V.; and Ranftl, R. 2022. Language-Driven Semantic Segmentation. In International Conference on Learning Representations (ICLR).
Liang, F.; Wu, B.; Dai, X.; Li, K.; Zhao, Y.; Zhang, H.; Zhang, P.; Vajda, P.; and Marculescu, D. 2023. Open-Vocabulary Semantic Segmentation with Mask-Adapted CLIP. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Liang, J.; He, R.; and Tan, T. 2025. A Comprehensive Survey on Test-Time Adaptation Under Distribution Shifts. International Journal of Computer Vision (IJCV), 133: 31–64. Meng, F.; Cui, C.; Dai, H.; and Gong, S. 2025. Black-Box Test-Time Prompt Tuning for Vision-Language Models. In AAAI Conference on Artificial Intelligence, volume 39, 6099–6107. Niu, S.; Wu, J.; Zhang, Y.; Chen, Y.; Zheng, S.; Zhao, P.; and Tan, M. 2022. Efficient Test-Time Model Adaptation without Forgetting. In International Conference on Machine Learning (ICML). Niu, S.; Wu, J.; Zhang, Y.; Wen, Z.; Chen, Y.; Zhao, P.; and Tan, M. 2023. Towards Stable Test-Time Adaptation in Dynamic Wild World. In International Conference on Learning Representations (ICLR). Noori, M.; Osowiechi, D.; Vargas Hakim, G.; Bahri, A.; Yazdanpanah, M.; Dastani, S.; Beizaee, F.; Ben Ayed, I.; and Desrosiers, C. 2025. Test-Time Adaptation of Vision-Language Models for Open-Vocabulary Semantic Segmentation. In Advances in Neural Information Processing Systems (NeurIPS), volume 38. Noori, M.; Vargas Hakim, G. A.; Osowiechi, D.; Shakeri, F.; Bahri, A.; Yazdanpanah, M.; Dastani, S.; Ben Ayed, I.; and Desrosiers, C. 2026. Histopath-C: Towards Realistic Domain Shifts for Histopathology Vision-Language Adaptation. In IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 4890– 4900. Osowiechi, D.; et al. 2024. WATT: Weight Average Test Time Adaptation of CLIP. In Advances in Neural Information Processing Systems (NeurIPS), volume 37. Parkhi, O. M.; Vedaldi, A.; Zisserman, A.; and Jawahar, C. V. 2012. Cats and Dogs. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021. Learning Transferable Visual Models from Natural Language Supervision. In International Conference on Machine Learning (ICML). Rao, Y.; Zhao, W.; Chen, G.; Tang, Y.; Zhu, Z.; Huang, G.; Zhou, J.; and Lu, J. 2022. DenseCLIP: Language-Guided Dense Prediction with Context-Aware Prompting. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Sharifdeen, A.; et al. 2025. O-TPT: Orthogonality Constraints for Calibrating Test-Time Prompt Tuning in Vision-Language Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 19942–19951. Sheng, L.; Liang, J.; He, R.; Wang, Z.; and Tan, T. 2025. The Illusion of Progress? A Critical Look at Test-Time Adaptation for VisionLanguage Models. In Advances in Neural Information Processing Systems (NeurIPS) Datasets and Benchmarks Track. Shu, M.; Nie, W.; Huang, D.-A.; Yu, Z.; Goldstein, T.; Anandkumar, A.; and Xiao, C. 2022. Test-Time Prompt Tuning for Zero-Shot Generalization in Vision-Language Models. In Advances in Neural Information Processing Systems (NeurIPS). Sohn, K.; Berthelot, D.; Li, C.-L.; Zhang, Z.; Carlini, N.; Cubuk, E. D.; Kurakin, A.; Zhang, H.; and Raffel, C. 2020. FixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence. In Advances in Neural Information Processing Systems (NeurIPS).
Song, F.; Li, Y.; Wang, R.; Zhou, J.; Zheng, C.; and Li, J. 2026. Doubly Debiased Test-Time Prompt Tuning for Vision-Language Models. In AAAI Conference on Artificial Intelligence, volume 40, 9069–9078. Sun, Y.; Wang, X.; Liu, Z.; Miller, J.; Efros, A. A.; and Hardt, M. 2020. Test-Time Training with Self-Supervision for Generalization under Distribution Shifts. In International Conference on Machine Learning (ICML). Valanarasu, J. M. J.; Guo, P.; Vibashan, V.; and Patel, V. M. 2022. On-the-Fly Test-Time Adaptation for Medical Image Segmentation. arXiv:2203.05574. Vargas Hakim, G. A.; et al. 2025. CLIPArTT: Adaptation of CLIP to New Domains at Test Time. In IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 7092–7101. Wang, D.; Shelhamer, E.; Liu, S.; Olshausen, B.; and Darrell, T. 2021. Tent: Fully Test-Time Adaptation by Entropy Minimization. In International Conference on Learning Representations (ICLR). Wang, L.; et al. 2017. Learning to Detect Salient Objects with Image-Level Supervision. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Wang, Q.; Fink, O.; Van Gool, L.; and Dai, D. 2022. Continual TestTime Domain Adaptation. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Xie, Q.; Luong, M.-T.; Hovy, E.; and Le, Q. V. 2020. SelfTraining with Noisy Student Improves ImageNet Classification. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Xu, J.; De Mello, S.; Liu, S.; Byeon, W.; Breuel, T.; Kautz, J.; and Wang, X. 2022. GroupViT: Semantic Segmentation Emerges from Text Supervision. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Xu, J.; Liu, S.; Vahdat, A.; Byeon, W.; Wang, X.; and De Mello, S. 2023. Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Zanella, M.; and Ben Ayed, I. 2024. On the Test-Time Zero-Shot Generalization of Vision-Language Models: Do We Really Need Prompt Learning? In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Zanella, M.; Fuchs, C.; De Vleeschouwer, C.; and Ben Ayed, I. 2025. Realistic Test-Time Adaptation of Vision-Language Models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 25103–25112. Zanella, M.; Gérin, B.; and Ben Ayed, I. 2024. Boosting VisionLanguage Models with Transduction. In Advances in Neural Information Processing Systems (NeurIPS), volume 37. Zhang, M.; Levine, S.; and Finn, C. 2022. MEMO: Test Time Robustness via Adaptation and Augmentation. In Advances in Neural Information Processing Systems (NeurIPS). Zhou, L.; Ye, M.; Li, S.; Li, N.; Zhu, X.; Deng, L.; Liu, H.; and Lei, Z. 2025. Bayesian Test-Time Adaptation for Vision-Language Models. In IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 29999–30009. Zhu, X.; Wu, H.; Wang, S.; Zhu, B.; Ge, J.; Zhang, J.; and Chen, L. 2026. Robustifying Vision-Language Models via Test-Time Prompt Adaptation. arXiv:2607.09450. Zou, Y.; Yu, Z.; Liu, X.; Vijaya Kumar, B. V. K.; and Wang, J. 2019. Confidence Regularized Self-Training. In IEEE/CVF International Conference on Computer Vision (ICCV). Zou, Y.; Yu, Z.; Vijaya Kumar, B. V. K.; and Wang, J. 2018. Domain Adaptation for Semantic Segmentation via Class-Balanced SelfTraining. In European Conference on Computer Vision (ECCV).