Robustness of Vision Foundation Models to Common Perturbations
arXiv:2604.14973v1 [cs.CR] 16 Apr 2026
Hongbin Liu1 , Zhengyuan Jiang1 , Cheng Hong2 , Neil Zhenqiang Gong1 1 Duke University, 2 Ant Group {hongbin.liu, zhengyuan.jiang, neil.gong}@duke.edu [email protected]
Abstract
mon perturbations to an image, unlike adversarial perturbations [4, 23], which are worst-case modifications designed to mislead models. Common perturbations, by contrast, occur frequently in non-adversarial, real-world scenarios. The robustness of foundation models and their downstream applications to adversarial perturbations has been widely studied [7, 11–14, 16, 20, 22]. However, robustness to common perturbations remains largely unexplored. Specifically, three key questions arise regarding robustness to common perturbations: 1) How robust are foundation models, i.e., how much does an embedding vector change when an image undergoes common perturbations? 2) How robust are downstream applications, i.e., to what extent does classifier accuracy degrade with perturbed images? 3) How can we improve robustness in foundation models and their downstream applications against common perturbations? A key challenge in answering these questions is designing a metric to quantify a foundation model’s robustness to common perturbations. Such a metric would enable systematic robustness assessments, facilitating comparisons across selfsupervised learning algorithms, model architectures, and sizes. Additionally, a robustness metric could help predict the performance of downstream applications (e.g., accuracy) for perturbed images and provide guidance to enhance foundation model robustness.
A vision foundation model outputs an embedding vector for an image, which can be affected by common editing operations (e.g., JPEG compression, brightness, contrast adjustments). These common perturbations alter embedding vectors and may impact the performance of downstream tasks using these embeddings. In this work, we present the first systematic study on foundation models’ robustness to such perturbations. We propose three robustness metrics and formulate five desired mathematical properties for these metrics, analyzing which properties they satisfy or violate. Using these metrics, we evaluate six industry-scale foundation models (OpenAI, Meta) across nine common perturbation categories, finding them generally non-robust. We also show that common perturbations degrade downstream application performance (e.g., classification accuracy) and that robustness values can predict performance impacts. Finally, we propose a fine-tuning approach to improve robustness without sacrificing utility.
1. Introduction A vision foundation model is a general-purpose feature extractor that outputs an embedding vector for an image. Typically, these models are pre-trained on vast collections of unlabeled images or image-text pairs in a self-supervised manner [18, 21] by major providers like OpenAI, Meta, and Google. For instance, OpenAI’s CLIP [21] jointly trains vision and language foundation models on 400 million imagetext pairs, while Meta’s DINO v2 [18] trains a vision foundation model on a large set of unlabeled images. Foundation models empower various downstream applications like image classification and depth estimation. Images in real-world settings often undergo common editing operations for various purposes. For instance, JPEG compression is widely used to reduce communication costs online, while brightness and contrast adjustments are also common (Table 8 in the Appendix lists 9 editing operations used in our experiments). These operations introduce com-
Our work: In this work, we answer the three questions above via performing the first systematic study on the robustness of foundation models to common perturbations. We begin by tackling the challenge of defining metrics to quantify a foundation model’s robustness to common perturbations. Given an image, common perturbations produce various perturbed versions, each with its own embedding vector. We quantify robustness by measuring the variations among these embedding vectors, exploring three metrics: one based on cosine similarity, another on Euclidean distance, and a third, DivergenceRadius, which uses the radius of the smallest enclosing ball in the embedding space to capture robustness. A suitable robustness metric should meet several intuitions, such as not increasing robustness as more perturba1
a special parameter ⊥ ∈ K, where P (x, ⊥) = x, returning the original image, to simplify descriptions. Embedding vector: A foundation model f outputs an embedding vector f (x) for an image x. To prevent embedding magnitude from affecting downstream applications, foundation models often normalize embeddings to an ℓ2 -norm of 1, ensuring ||f (x)||2 = 1 for any image x. Thus, all embedding vectors lie on a unit-radius hyper-sphere in the embedding space. Desired mathematical properties of a robustness metric: Given a foundation model f , an image x, and a perturbation function P with parameter domain K, our goal is to define a robustness metric R(f, x, P, K) to quantify the robustness of f for x under P . This metric essentially measures variations among embedding vectors {f (P (x, k))}k∈K for perturbed versions of x generated by P . The robustness metric should allow quantitative comparisons of different foundation models’ robustness to common perturbations. Since foundation models support downstream applications, the metric should also predict downstream performance for x under P . For instance, if the downstream task is classification, the robustness value R(f, x, P, K) should help predict the accuracy for perturbed versions of x within the domain K. We have the following mathematical properties:
tions are applied. We formalize these intuitions with five mathematical properties and analyze which properties the metrics satisfy. We find that DivergenceRadius satisfies all five properties, whereas the other metrics fail to meet one; additionally, the Euclidean and cosine similarity metrics are equivalent. Using our robustness metrics, we address the first question through a systematic study of six industry-scale foundation models from the CLIP (OpenAI) and DINO v2 (Meta) families, covering different self-supervised learning algorithms, architectures, and sizes, across nine categories of common perturbations. Our findings are consistent across the three robustness metrics: foundation models generally lack robustness to common perturbations, often producing divergent embeddings for perturbed images. Additionally, we observe that foundation models based on Vision Transformer architectures are more robust than those based on ResNet architectures. To address the second question, we evaluate the robustness of downstream classifiers and depth estimation models built on industry-scale foundation models against common perturbations. We observe that these perturbations degrade both classifier accuracy and depth estimation performance; for example, glass-blurring reduces the accuracy of a zeroshot ImageNet classifier by 9.4%. This occurs due to variations in embedding vectors caused by perturbations. Additionally, we find that average classification accuracy and mean squared error of depth maps for perturbed images are roughly linear functions of the image’s robustness value (e.g., cosine similarity or DivergenceRadius), enabling accurate performance predictions for downstream tasks using a simple linear regression model. Finally, we propose a fine-tuning method to enhance a foundation model’s robustness while preserving utility for downstream tasks. Our approach aims to balance two objectives: a robustness goal and a utility goal, each quantified by a corresponding loss term. We fine-tune the model by minimizing a weighted sum of these loss terms, with empirical results showing that our method successfully improves robustness without compromising utility.
1. Bounded domain: To ensure interpretability and comparability, we design a scalar robustness metric with a bounded interval output, normalized to [0, 1] for simplicity. Thus, for any f , x, P , and K, the robustness value R(f, x, P, K) should fall within [0, 1], where a larger value indicates less robustness: R(f, x, P, K) ∈ [0, 1], ∀f, x, P, K.
(1)
2. Monotonicity: When the parameter domain K expands, the model should not become more robust under this larger domain. For instance, greater variation in JPEG quality factors should not decrease R(f, x, P, K). Formally: R(f, x, P, K1 ) ≤ R(f, x, P, K2 ), ∀K1 ⊆ K2 .
2. Problem Formulation
(2)
3. Best robustness: The model is maximally robust for x under P if all perturbed versions of x have the same embedding as x, resulting in R(f, x, P, K) = 0 if:
Perturbation function: We represent common perturbations with a perturbation function P (x, k), where x is an image and k a perturbation parameter, yielding a perturbed image P (x, k). For instance, if P represents JPEG compression, k is the quality factor controlling compression level. For some functions, k is multi-dimensional, such as fog blurring, where k includes density and frequency. We denote the domain of k as K, the set from which k is selected when applying P to x. The domain K may include discrete values (e.g., JPEG quality factors) or continuous values (e.g., Gaussian noise standard deviation). We assume
f (P (x, k)) = f (x), ∀k ∈ K.
(3)
4. Worst robustness: If the embedding vectors for perturbed images are uniformly distributed in the embedding space (sum to zero), then robustness is at its lowest with R(f, x, P, K) = 1, if ∃K′ ⊆ K: X f (P (x, k)) = 0. (4) k∈K′
2
Table 1. Three robustness metrics explored in this work and the desired mathematical properties they violate. Robustness Metric Formulation Violating Properties 1−mink1 ,k2 ∈K cos(f (P (x,k1 )),f (P (x,k2 ))) 2
Worst-robustness
maxk1 ,k2 ∈K ||f (P (x,k1 ))−f (P (x,k2 ))||2 2
Worst-robustness
Cosine similarity
Rcs (f, x, P, K) =
Euclidean distance
Red (f, x, P, K) =
DivergenceRadius
Rdr (f, x, P, K) = argmin r s.t. ∃c, ||f (P (x, k)) − c||2 ≤ r, ∀k ∈ K
None
where 0 is the zero vector. Figure 1 illustrates this in a two-dimensional space. 5. Rotational invariance: Since embeddings are rotationfree, the robustness metric should be invariant to rotations in the embedding space. Let M be a rotation matrix, then:
Rcs (f, x, P, K) = 0 when all perturbed versions of x have the same embedding. Finally, Rcs is rotation-invariant since cosine similarity is. However, Rcs does not satisfy the worst-robustness property. For instance, in Figure 1, Rcs is 0.75, not 1, as Rcs = 1 only when embedding vectors cover exactly half the hypersphere, with cosine similarity of -1. If embedding vectors are more evenly distributed, the smallest cosine similarity is greater than -1, violating worst-robustness. Monotonically increasing function of cosine similarity: A robustness metric based on any monotonically increasing function of Rcs cannot satisfy all five desired properties. Formally, we define such a metric:
R(M · f, x, P, K) = R(f, x, P, K),
Rg (f, x, P, K) = g(Rcs (f, x, P, K)),
Figure 1. An example for the worst-robustness property.
(5)
where M · f denotes rotating f (x) by M .
where g is a monotonically increasing function satisfying g(0) = 0 and g(1) = 1. We can show Rg satisfies boundeddomain, best-robustness, and rotation-invariance properties but fails the worst-robustness property. Formally, we have:
3. Robustness Metrics We explore three robustness metrics and theoretically analyze what desired mathematical properties they satisfy/violate, which we summarize in Table 1.
Theorem 1 (Monotonically Increasing Function of Rcs ). Given any g that is a monotonically increasing function of Rcs and satisfies g(0) = 0 and g(1) = 1, Rg = g(Rcs ) does not satisfy the worst-robustness property.
3.1. Cosine Similarity Since embedding vectors lie on a hypersphere, we can use the angles between them to quantify robustness. Specifically, the largest angle (smallest cosine similarity) between any two embedding vectors from perturbed images can define a robustness metric. Formally, for a model f , image x, perturbation function P , and parameter domain K, the cosine similarity-based robustness metric Rcs is:
Proof. To prove this, we construct a counter-example. Since Rg = g(Rcs ) increases with Rcs , and since g(0) = 0 and g(1) = 1, we have Rg < 1 if Rcs < 1. In Figure 1, Rcs = 0.75 < 1, so Rg < 1, countering the worst-robustness property.
3.2. Euclidean Distance Another intuitive robustness metric is to use the Euclidean distances between embedding vectors of perturbed images to quantify their variations. Formally, given a model f , image x, perturbation function P , and parameter domain K, we define a Euclidean distance-based robustness metric Red as:
Rcs (f, x, P, K) =
1 − mink1 ,k2 ∈K cos(f (P (x, k1 )), f (P (x, k2 ))) , 2
(7)
(6)
where cos(·, ·) is the cosine similarity. The constants normalize the value to [0,1]. We can verify that the robustness metric Rcs satisfies the five mathematical properties except the worst-robustness one. In particular, since cos(·, ·) is in [-1,1], Rcs (f, x, P, K) lies in [0, 1], meeting the bounded-domain property. As K expands, mink1 ,k2 ∈K cos(f (P (x, k1 )), f (P (x, k2 ))) does not increase, so Rcs does not decrease, satisfying monotonicity. The best-robustness property holds because
Red (f, x, P, K) =
maxk1 ,k2 ∈K ||f (P (x, k1 )) − f (P (x, k2 ))||2 , 2
(8)
where ||·||2 denotes the Euclidean distance, with the constant 2 normalizing the metric to [0,1]. We can show that Red is a monotonically increasing function of Rcs , specifically as follows: 3
3.3.2. DivergenceRadius Satisfies the Mathematical Properties
ℛ𝒅𝒓 𝒄
We show that the robustness metric Rdr satisfies all the five mathematical properties. Proof. Please refer to the Appendix. 3.3.3. Solving DivergenceRadius
Figure 2. An example to illustrate the minimum enclosing ball. The red circle is the minimum enclosing ball of the three embedding vectors.
We consider both discrete and continuous K. Discrete K: For a discrete domain K, the embedding vectors {f (P (x, k))}k∈K are discrete points on the unit hypersphere. In this case, we can use Welzl’s algorithm [25] to efficiently find the minimum enclosing ball’s center and radius, solving Equation 10 in O(dn) time, where d is the embedding dimension and n the number of discrete values in K.
Theorem 2. For any model f , image x, perturbation function P , and parameter domain K, Red is the square root of Rcs : p Red (f, x, P, K) = Rcs (f, x, P, K). (9)
Continuous K: When K is continuous, finding an exact solution is infeasible due to infinite embedding vectors. To approximate, we sample a discrete subset K̄ from K, transforming the problem into a discrete one. We use two sampling methods: random sampling (uniform random selection) and equally-spaced sampling, where equally spaced values better represent the domain when using fewer samples. For a domain K = [a, b] and m samples, equally-spaced sampling yields values a, a+(b−a)/(m−1), . . . , b. Our experiments show that equally-spaced sampling outperforms random sampling for DivergenceRadius estimation accuracy. After obtaining K̄, we apply Welzl’s algorithm to find an approximate DivergenceRadius. For large discrete domains, a small sample set can similarly reduce computational cost. Random and equally-spaced sampling methods also apply to estimating the cosine similarity-based robustness Rcs and Euclidean distance-based robustness Red .
Proof. Please refer to the Appendix. Theorem 2 further shows that Rcs and Red are equivalent, as Rcs can be converted to Red by taking the square root. Therefore, we omit results for Red in our experiments for simplicity.
3.3. DivergenceRadius 3.3.1. Formulating an Optimization Problem We first outline the intuition behind DivergenceRadius and then present its formulation. Intuition: The cosine similarity and Euclidean distance metrics do not satisfy the worst-robustness property, which requires a maximum robustness value of 1 when the embedding vectors for perturbed images are equally distributed in the embedding space. Our intuition is that if the embedding vectors are uniformly spread on the unit hyper-sphere, the radius of the smallest ball enclosing these vectors will equal 1. Thus, constructing a minimum enclosing ball for the perturbed embedding vectors can satisfy the worst-robustness property and indicate robustness: a smaller radius implies greater robustness.
4. Measuring Robustness of Real-world Foundation Models In this section, we evaluate the robustness of industryscale foundation models using cosine similarity and DivergenceRadius. Although cosine similarity does not theoretically satisfy the worst-robustness property, we include it in our experiments since the worst-robustness scenario does not occur in the evaluated perturbations. We omit Euclidean distance results due to its equivalence with cosine similarity.
Optimization problem: We aim to find the smallest-radius high-dimensional ball with center c that encloses all embedding vectors of perturbed versions of an image x within the perturbation domain K. This radius r, our DivergenceRadius, is denoted as Rdr (f, x, P, K). Formally:
4.1. Measurement Setup
Rdr (f, x, P, K) = argmin r, s.t. ∃c, ||f (P (x, k)) − c||2 ≤ r, ∀k ∈ K.
Foundation models: We evaluate foundation models pretrained using various algorithms, architectures, and sizes, allowing comparison of pre-training methods and model robustness. Specifically, we evaluate two popular families of vision foundation models: CLIP [21] and DINO v2 [18]. CLIP (by OpenAI) was trained with multi-modal self-supervised learning on 400 million image-text pairs, while DINO v2
(10)
Figure 2 illustrates an example of a minimum ball enclosing embedding vectors for three perturbed versions of an image x. Note that K includes a special parameter ⊥ where P (x, ⊥) = x. 4
0.10
We observe that DivergenceRadius initially increases and then saturates with both sampling methods. As m grows, more diverse parameters from K yield a higher Rdr , approaching the true robustness value. Equally-spaced sampling converges more quickly, reaching near-saturation at m ≥ 5, whereas random sampling requires m ≥ 20. This indicates that equally-spaced sampling provides a more efficient approximation of Rdr . Thus, for computational efficiency, we use equally-spaced sampling with m = 5 for each perturbation function in other experiments. Comparing model architectures: Figure 4, Figure 13 (Appendix), and Figure 15 (Appendix) show the average Rdr on ImageNet, Food101, and NYU-Depth V2 datasets across different foundation models and perturbations. Figures 12, 14, and 16 (Appendix) present the corresponding Rcs results. To compare architectures, we examine models with the same pre-training algorithm and similar sizes: specifically, CLIP ViT-B/16 vs. CLIP RN50 and CLIP ViT-L/14 vs. CLIP RN50×64. We observe that ViT-B/16 (or ViT-L/14) is consistently more robust than RN50 (or RN50×64) across all perturbations and datasets, as evidenced by lower DivergenceRadius values. This suggests that Vision Transformers are generally more robust to common image perturbations than ResNet architectures. Similar results are also observed when using cosine similarity. Comparing pre-training algorithms: To compare pretraining algorithms, we evaluate the robustness of CLIP ViT-L/14 and DINO v2 ViT-L/14, which share the same architecture and model size. We find no consistent trend in robustness across perturbation types. For example, DINO v2 ViT-L/14 has lower DivergenceRadius (i.e., greater robustness) than CLIP ViT-L/14 on JPEG compression, Brightness adjustment, Contrast adjustment, Fog blurring, Gaussian noise, and Glass blurring across all datasets. However, DINO v2 ViT-L/14 shows higher DivergenceRadius (i.e., lower robustness) on other perturbations. For instance, under JPEG compression, the average DivergenceRadius for ImageNet images is 0.038 for DINO v2 ViT-L/14 and 0.084 for CLIP ViT-L/14, while under Defocus blurring, it is 0.146 for DINO v2 ViT-L/14 and 0.108 for CLIP ViT-L/14. Comparing model sizes: For model size comparisons, we examine foundation models within the same architecture and pre-training algorithm: CLIP ViT-B/16 vs. CLIP ViT-L/14, CLIP RN50 vs. CLIP RN50x64, and DINO v2 ViT-L/14 vs. DINO v2 ViT-g/14. In the CLIP family, larger models are generally less robust to common perturbations than smaller ones. For example, ViT-L/14 shows a higher average Rdr (or Rcs ) than ViT-B/16 across all perturbations and datasets, and RN50 × 64 has higher robustness metrics than RN50, except for a few cases on Food101 (e.g., JPEG compression, Defocus blurring). Conversely, in the DINO v2 family, larger models are more robust: ViT-g/14 consistently shows lower Rdr values than ViT-L/14 across
DivergenceRadius
0.08 0.06 0.04 0.02 1
Random sampling Equally-spaced sampling 5 10 15 20 100 m
Figure 3. Comparing random sampling and equally-spaced sampling at computing DivergenceRadius, where CLIP ViT-L/14 foundation model, JPEG compression, and ImageNet images are used. m is the number of discrete values sampled from the domain K.
(by Meta) used self-supervised learning on 142 million unlabeled images. In the CLIP family, we test ViT-B/16, ViTL/14, RN50, and RN50×64 (Vision Transformer and ResNet architectures). In the DINO v2 family, we evaluate ViT-L/14 and ViT-g/14. Details on these models are in Table 5 in the Appendix. Datasets: For robustness evaluation, we use two image classification datasets, ImageNet [6] and Food101 [3], and one depth estimation dataset, NYU-Depth V2 [17]. Details of these datasets are provided in Table 4 (Appendix). Although labels are not required for robustness testing, they are used later to evaluate downstream applications. Common perturbations: Following [9], we use nine common perturbation functions representing typical image editing operations: JPEG compression, Brightness adjustment, Contrast adjustment, Defocus blurring, Elastic blurring, Fog blurring, Frost blurring, Gaussian noise, and Glass blurring. Each has one variable parameter, with fixed values for additional parameters if present. Table 8 in the Appendix details each perturbation’s parameter domain K, additional fixed parameters, and a visualization of a maximally distorted example for each function. Following prior work [9], the selected K for each perturbation reflects realistic image editing in real-world scenarios.
4.2. Measurement Results Random sampling v.s. equally-spaced sampling: When computing robustness values (i.e., Rcs or Rdr ), we can use random sampling, which selects m discrete values from K at random, or equally-spaced sampling, which selects m evenly spaced values. We compare these methods by computing Rdr for a foundation model, perturbation function, and image. Figure 3 shows the average Rdr on ImageNet images for CLIP ViT-L/14 under JPEG compression as m varies. 5
50 x6 4 ViT -L/ 14 ViT -g/ 14
50
RN
50 x6 4 ViT -L/ 14 ViT -g/ 14
50
14
RN
RN
6
ViT -L/
ViT -B/ 1
4
4
Foundation Models
ViT -g/ 1
RN
50
ViT -L/ 1
x6 4
50
4
ViT -L/ 1
/16 ViT -B
(h) Gaussian noise
RN
DivergenceRadius
CLIP DINO v2
4
4
Foundation Models
ViT -g/ 1
ViT -L/ 1
x6 4 50
14
6
ViT -L/
ViT -B/ 1 DivergenceRadius
50 x6 4 ViT -L/ 14 ViT -g/ 14
RN RN
RN
DivergenceRadius
50 x6 4 ViT -L/ 14 ViT -g/ 14
50
RN
50 RN 50 RN
ViT -L/ 1
/16 ViT -B
Foundation Models
(f) Fog blurring
0.20 0.15 0.10 0.05 0.00
4
DivergenceRadius
4
4
ViT -g/ 1
ViT -L/ 1
x6 4 50
50
14
ViT -L/
6 ViT -B/ 1
RN RN
4
ViT -L/ 1
/16 ViT -B
RN
DivergenceRadius
Foundation Models
CLIP DINO v2
0.20 0.15 0.10 0.05 0.00
Foundation Models
CLIP DINO v2
0.20 0.15 0.10 0.05 0.00
(e) Elastic blurring
CLIP DINO v2
(g) Frost blurring
RN
6 DivergenceRadius
50 x6 4 ViT -L/ 14 ViT -g/ 14
50
14
6
ViT -L/
RN
DivergenceRadius
ViT -B/ 1
Foundation Models
Foundation Models
(c) Contrast adjustment
CLIP DINO v2
0.20 0.15 0.10 0.05 0.00
(d) Defocus blurring
0.20 0.15 0.10 0.05 0.00
Foundation Models
(b) Brightness adjustment
CLIP DINO v2
0.20 0.15 0.10 0.05 0.00
ViT -L/
ViT -B/ 1
RN
Foundation Models
(a) JPEG compression
CLIP DINO v2
0.20 0.15 0.10 0.05 0.00
14
DivergenceRadius
50 x6 4 ViT -L/ 14 ViT -g/ 14
50
14
6
ViT -L/
ViT -B/ 1
CLIP DINO v2
0.20 0.15 0.10 0.05 0.00
RN
DivergenceRadius
CLIP DINO v2
0.20 0.15 0.10 0.05 0.00
(i) Glass blurring
Figure 4. Average DivergenceRadius of ImageNet testing images for different foundation models and perturbation functions. Table 2. ACC and average ACCp of ImageNet’s testing images for two downstream classifiers. Zero-shot classification is based on the CLIP ViT-L/14 foundation model and linear-probe classification is based on the DINO v2 ViT-g/14 foundation model.
perturbations and datasets. This contrast suggests that pretraining settings—multi-modal self-supervised learning for CLIP vs. image-only self-supervised learning for DINO v2—impact robustness differently as model size increases.
5. Measuring Performance of Downstream Applications
ACC (%)
In this section, we measure the robustness of downstream applications to common perturbations and demonstrate that an image’s robustness value (i.e., Rdr or Rcs ) can predict the performance of downstream applications on its perturbed versions. We omit Euclidean distance results due to its equivalence with cosine similarity.
ACCp (%)
Zero-shot
Linear-probe
Classification
Classification
68.4
86.6
JPEG compression
65.1 (↓ 3.3)
84.6 (↓ 2.0)
Brightness adjustment
66.3 (↓ 2.1)
85.7 (↓ 0.9)
Contrast adjustment
67.3 (↓ 1.1)
86.4 (↓ 0.2)
Defocus blurring
60.1 (↓ 8.3)
82.2 (↓ 4.4)
Elastic blurring
63.5 (↓ 4.9)
85.0 (↓ 1.6)
Fog blurring
66.0 (↓ 2.4)
86.1 (↓ 0.5)
Frost blurring
61.4 (↓ 7.0)
83.6 (↓ 3.0)
Gaussian noise
66.0 (↓ 2.4)
85.7 (↓ 0.9)
Glass blurring
59.0 (↓ 9.4)
82.9 (↓ 3.7)
5.1. Experimental Setup (definitions in the Appendix).
Downstream applications: Given a pre-trained vision foundation model, we consider the following three popular downstream applications: zero-shot classification, linear-probe classification, and depth estimation. The Appendix shows more details of these applications.
Parameter settings: We evaluate the classifiers with the highest accuracy on zero-shot or linear-probe tasks. For zero-shot classification, we use CLIP ViT-L/14; for linearprobing, we use DINO v2 ViT-g/14, training a one-layer classifier for Food101.
Evaluation metrics (ACC, ACCp , RM SE, and RM SEp ): For a downstream classifier g ◦ f , where f is a foundation model and g a classifier head, accuracy (ACC) is the fraction of correctly predicted labels. Accuracy under perturbation (ACCp ) is calculated for each perturbed image. For depth estimation heads, we use RM SE and RM SEp
5.2. Experimental Results Due to space limits, we discuss results for downstream classifiers and defer depth estimation results to the Appendix. ACC vs. ACCp : 6
Table 2 shows ACC and average
1.00 Accuracy under perturbation ACCp
Accuracy under perturbation ACCp
0.80 0.60 JPEG compression Brightness adjustment Contrast adjustment Defocus blurring Elastic blurring Fog blurring 0.20 Frost blurring Gaussian noise Glass blurring 0.00 0.04 0.08 0.12 DivergenceRadius 0.40
0.80 0.60 0.40 0.200.00
0.16
(a) Zero-shot classification
JPEG compression Brightness adjustment Contrast adjustment Defocus blurring Elastic blurring Fog blurring Frost blurring Gaussian noise Glass blurring 0.06 0.12 0.18 DivergenceRadius
0.24
(b) Linear-probe classification
Figure 5. Accuracy under perturbation ACCp vs. DivergenceRadius of ImageNet testing images for (a) zero-shot classification and (b) linear-probe classification when different perturbation functions are used. Zero-shot classification is based on the CLIP ViT-L/14 foundation model and linear-probe classification is based on the DINO v2 ViT-g/14 foundation model.
ACCp (%)
Accuracy under perturbation ACCp
ACC (%) JPEG compression Brightness adjustment Contrast adjustment Defocus blurring Elastic blurring Fog blurring Frost blurring Gaussian noise Glass blurring
Zero-shot Classification 92.0 87.4 (↓ 4.6) 88.4 (↓ 3.6) 90.9 (↓ 1.1) 81.7 (↓ 10.3) 87.2 (↓ 4.8) 88.3 (↓ 3.7) 79.3 (↓ 12.7) 86.6 (↓ 5.4) 82.9 (↓ 9.1)
Linear-probe Classification 94.1 91.2 (↓ 2.9) 90.8 (↓ 3.3) 91.8 (↓ 2.3) 87.0 (↓ 7.1) 89.6 (↓ 4.5) 90.9 (↓ 3.2) 85.0 (↓ 9.1) 89.7 (↓ 4.4) 87.3 (↓ 6.8)
1.00 Accuracy under perturbation ACCp
1.00
Table 3. ACC and average ACCp of Food101’s testing images for two downstream classifiers. Zero-shot classification is based on the CLIP ViT-L/14 foundation model and linear-probe classification is based on the DINO v2 ViT-g/14 foundation model.
0.80 JPEG compression Brightness adjustment Contrast adjustment Defocus blurring Elastic blurring Fog blurring 0.40 Frost blurring Gaussian noise Glass blurring 0.20 0.04 0.08 0.12 DivergenceRadius 0.60
0.16
(a) Zero-shot classification
0.80 0.60 0.40 0.200.00
JPEG compression Brightness adjustment Contrast adjustment Defocus blurring Elastic blurring Fog blurring Frost blurring Gaussian noise Glass blurring 0.06 0.12 0.18 DivergenceRadius
0.24
(b) Linear-probe classification
Figure 6. Accuracy under perturbation ACCp vs. DivergenceRadius of Food101 testing images for (a) zero-shot classification and (b) linear-probe classification when different perturbation functions are used. Zero-shot classification is based on the CLIP ViT-L/14 foundation model and linear-probe classification is based on the DINO v2 ViT-g/14 foundation model.
ACCp under different perturbations for ImageNet; results for Food101 are in Table 3. Common perturbations reduce ACCp compared to ACC, indicating degraded accuracy. For example, Glass blurring reduces ImageNet zero-shot accuracy by 9.4%, and Frost blurring reduces Food101 linearprobe accuracy by 3.3%. The accuracy drop correlates with the average robustness value across perturbations. For instance, in Figure 4, both CLIP ViT-L/14 and DINO v2 ViTg/14 exhibit the highest DivergenceRadius under Defocus blurring, corresponding to the largest ACCp drop in Table 2. This suggests that higher DivergenceRadius values indicate greater embedding diversity, increasing the chance of misclassification.
sion. We divide the dataset, train a linear model on the first half (predicting ACCp from DivergenceRadius or cosine similarity), and test it on the second half. Figures 10 and 11 in the Appendix show low mean squared errors, indicating that DivergenceRadius or cosine similarity can reliably predict downstream accuracy under perturbations.
6. Robustness Enhancement 6.1. Method Robustness and utility goals: We propose a fine-tuning method to enhance a foundation model’s robustness against common perturbations. Given a foundation model f , our goal is to produce a model f ′ that meets both a robustness goal (increased robustness to perturbations) and a utility goal (maintaining performance on unperturbed images).
ACCp vs. robustness value: Figure 5 shows the relationship between ACCp and DivergenceRadius for ImageNet images under various perturbations; results for Food101 are in Figure 6. Across datasets and perturbations, ACCp decreases approximately linearly as DivergenceRadius or cosine similarity increases, indicating that greater embedding diversity leads to lower accuracy.
Formulating an optimization problem: We define two loss terms for robustness and utility. Minimizing these terms involves optimizing a weighted sum. For a set of unlabeled images D, we quantify robustness by the cosine
Predicting ACCp using robustness values: The linear trend between ACCp and DivergenceRadius (or cosine similarity) enables accurate ACCp prediction via linear regres7
Cosine Similarity
similarity between an image’s embedding and those of its perturbed versions. the robustness loss term L1 P Specifically, 1 cos (f ′ (x), f ′ (P (x, k))), where P is a is: L1 = − ∥D∥ x∈D
perturbation function and k is sampled from K each epoch. To ensure utility, we use cosine similarity between embed′ dings of unperturbed images P from f and f′ . The utility loss 1 cos (f (x), f (x)). term L2 is: L2 = − ∥D∥
0.12
Before After
0.09 0.06 0.03 0.00 ImageNet
Food101
DivergenceRadius
(a) Cosine similarity
x∈D
The optimization problem then minimizes the weighted sum of these terms: minf ′ L1 + λL2 , where λ balances robustness and utility.
6.2. Experimental Setup We use CLIP ViT-L/14 for zero-shot classification. Finetuning data D includes 20,000 randomly selected ImageNet training images, with JPEG compression as the default perturbation. Models are fine-tuned for 50 epochs with a learning rate of 1 × 10−5 and λ = 1 unless otherwise stated.
Before After
0.09 0.06 0.03 0.00 ImageNet
Accuracy ACC
Solving the optimization problem: We use a gradientbased method, initializing f ′ as f , and iteratively updating f ′ using the gradient computed over mini-batches from D.
0.12
Food101
(b) DivergenceRadius 1.00 Before 0.90 After 0.80
0.70 0.60
ImageNet
Food101
(c) ACC
Figure 7. (a) Average cosine similarity, (b) average DivergenceRadius, and (c) ACC of zero-shot classification for the two datasets before and after robustness enhancement, where the foundation model is CLIP ViT-L/14.
6.3. Experimental Results Achieving the robustness goal: Figures 7a and 7b show the average cosine similarity and DivergenceRadius for ImageNet/Food101 testing images before and after fine-tuning. We observe decreases in both metrics, indicating that embedding vectors of perturbed images become closer to those of unperturbed images, thus enhancing robustness.
Zhu et al. [26] investigated language models’ robustness to prompt perturbations but did not focus on vision models. Hendrycks and Dietterich [9] assessed classifier robustness to common perturbations, and studies [1, 2, 19, 24] compared robustness of vision transformers and CNNs, though they focused on classifiers, not foundation models. Hendrycks et al. [10] found that common perturbations could improve out-of-distribution robustness for classifiers, while our work centers on in-distribution robustness.
Achieving the utility goal: Figure 7c shows that zeroshot classification accuracy ACC remains nearly unchanged on unperturbed images after fine-tuning, indicating that the utility goal is met due to the inclusion of L2 . Impact of λ: Table 9 in the Appendix shows how λ affects the robustness-utility trade-off. When λ is small (e.g., 0), robustness is achieved but utility declines, whereas large λ (e.g., 5) favors utility over robustness.
8. Conclusion and Future Work In this work, we introduce three metrics—cosine similarity, Euclidean distance, and DivergenceRadius—to quantify foundation model robustness to common perturbations. Theoretically, we showed that DivergenceRadius meets all five desired mathematical properties, while cosine similarity and Euclidean distance do not satisfy the worst-robustness property. Using these metrics, we empirically evaluated the robustness of industry-scale foundation models and downstream applications, finding limited robustness to common perturbations, which impacts downstream performance. We also demonstrated that fine-tuning models to align perturbed embeddings with the original can enhance robustness without affecting utility. Future work includes extending these analyses to language models and investigating robustness against adversarial perturbations.
7. Related Work Foundation models: Foundation models [5, 8, 15, 18, 21] are pre-trained neural networks used as general-purpose feature extractors, often for vision tasks. Vision foundation models are typically pre-trained on large datasets of unlabeled images [5, 18] or image-text pairs [21]. For example, Meta’s DINO v2 [18] is trained on 142 million images, while OpenAI’s CLIP [21] is trained on 400 million image-text pairs. Common perturbations: Common perturbations frequently arise in real-world, non-adversarial settings. While adversarial robustness of foundation models has been widely studied, robustness to common perturbations is less explored. 8
References
[18] Maxime Oquab, Timothée Darcet, Théo Moutakanni, Huy Vo, Marc Szafraniec, Vasil Khalidov, Pierre Fernandez, Daniel Haziza, Francisco Massa, Alaaeldin El-Nouby, et al. Dinov2: Learning robust visual features without supervision. arXiv, 2023. 1, 4, 8 [19] Sayak Paul and Pin-Yu Chen. Vision transformers are robust learners. In AAAI, 2022. 8 [20] Wenjie Qu, Jinyuan Jia, and Neil Zhenqiang Gong. Reaas: Enabling adversarially robust downstream classifiers via robust encoder as a service. In NDSS, 2023. 1 [21] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In ICML, 2021. 1, 4, 8 [22] Aniruddha Saha, Ajinkya Tejankar, Soroush Abbasi Koohpayegani, and Hamed Pirsiavash. Backdoor attacks on selfsupervised learning. In CVPR, 2022. 1 [23] Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. Intriguing properties of neural networks. In ICLR, 2014. 1 [24] Zeyu Wang, Yutong Bai, Yuyin Zhou, and Cihang Xie. Can cnns be more robust than transformers? In ICLR, 2023. 8 [25] Emo Welzl. Smallest enclosing disks (balls and ellipsoids). In New Results and New Trends in Computer Science. Springer, 2005. 4 [26] Kaijie Zhu, Jindong Wang, Jiaheng Zhou, Zichen Wang, Hao Chen, Yidong Wang, Linyi Yang, Wei Ye, Neil Zhenqiang Gong, Yue Zhang, et al. Promptbench: Towards evaluating the robustness of large language models on adversarial prompts. arXiv, 2023. 8
[1] Yutong Bai, Jieru Mei, Alan L Yuille, and Cihang Xie. Are transformers more robust than cnns? In NeurIPS, 2021. 8 [2] Srinadh Bhojanapalli, Ayan Chakrabarti, Daniel Glasner, Daliang Li, Thomas Unterthiner, and Andreas Veit. Understanding robustness of transformers for image classification. In ICCV, 2021. 8 [3] Lukas Bossard, Matthieu Guillaumin, and Luc Van Gool. Food-101 – mining discriminative components with random forests. In ECCV, 2014. 5 [4] Nicholas Carlini and David Wagner. Towards evaluating the robustness of neural networks. In S&P, 2017. 1 [5] Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. A simple framework for contrastive learning of visual representations. In ICML, 2020. 8 [6] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, 2009. 5 [7] Lijie Fan, Sijia Liu, Pin-Yu Chen, Gaoyuan Zhang, and Chuang Gan. When does contrastive learning preserve adversarial robustness from pretraining to finetuning? 2021. 1 [8] Kaiming He, Haoqi Fan, Yuxin Wu, Saining Xie, and Ross Girshick. Momentum contrast for unsupervised visual representation learning. In CVPR, 2020. 8 [9] Dan Hendrycks and Thomas Dietterich. Benchmarking neural network robustness to common corruptions and perturbations. In ICLR, 2019. 5, 8 [10] Dan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath, Frank Wang, Evan Dorundo, Rahul Desai, Tyler Zhu, Samyak Parajuli, Mike Guo, et al. The many faces of robustness: A critical analysis of out-of-distribution generalization. In ICCV, 2021. 8 [11] Jinyuan Jia, Yupei Liu, and Neil Zhenqiang Gong. Badencoder: Backdoor attacks to pre-trained encoders in selfsupervised learning. In S&P, 2022. 1 [12] Ziyu Jiang, Tianlong Chen, Ting Chen, and Zhangyang Wang. Robust pre-training by adversarial contrastive learning. NeurIPS, 2020. [13] Zhengyuan Jiang, Jinghuai Zhang, and Neil Zhenqiang Gong. Evading watermark based detection of ai-generated content. In CCS, 2023. [14] Changjiang Li, Ren Pang, Zhaohan Xi, Tianyu Du, Shouling Ji, Yuan Yao, and Ting Wang. An embarrassingly simple backdoor attack on self-supervised learning. In ICCV, 2023. 1 [15] Junnan Li, Dongxu Li, Caiming Xiong, and Steven Hoi. Blip: Bootstrapping language-image pre-training for unified visionlanguage understanding and generation. In ICML, 2022. 8 [16] Hongbin Liu, Jinyuan Jia, and Neil Zhenqiang Gong. PoisonedEncoder: Poisoning the unlabeled pre-training data in contrastive learning. In USENIX Security Symposium, 2022. 1 [17] Pushmeet Kohli Nathan Silberman, Derek Hoiem and Rob Fergus. Indoor segmentation and support inference from rgbd images. In ECCV, 2012. 5
9
A. Impact Statements
images corresponding to the expanded perturbation parameters (i.e., parameters in K2 -K1 ) fall outside of the minimum enclosing ball for K1 , the minimum enclosing ball for K2 expands to have a larger radius, i.e., Rdr (f, x, P, K1 ) < Rdr (f, x, P, K2 ); otherwise, the minimum enclosing ball remains the same, i.e., Rdr (f, x, P, K1 ) = Rdr (f, x, P, K2 ). Therefore, we have Rdr (f, x, P, K1 ) ≤ Rdr (f, x, P, K2 ) if K1 ⊆ K2 , which satisfies the monotonicity property. Best robustness: If all the perturbed versions of the image have the same embedding vector, the minimum enclosing ball has a radius 0. Thus, we have Rdr (f, x, P, K) = 0 in such case, achieving the best-robustness property. Worst robustness: In the worst-robustness case, there exists a subdomain K′ ⊆ K, where K′ consists of discrete values and the corresponding embedding vectors P are equally distributed in the embedding space, i.e., k∈K′ f (P (x, k)) = 0. Without loss of generality, we assume the subdomain K′ contains n discrete values k1 , · · · , kn . Then, we have the following:
This work strengthens the reliability of vision foundation models by systematically evaluating their robustness to common perturbations and proposing fine-tuning strategies to enhance stability without sacrificing utility. By introducing principled robustness metrics and analyzing industry-scale models, our findings highlight vulnerabilities that can impact real-world applications. Addressing these weaknesses improves the robustness of AI systems in practical settings, ensuring consistent performance across diverse conditions. Our work provides a foundation for future research on enhancing model robustness, including extending robustness analyses to language models and defending against adversarial perturbations.
B. Proof of Theorem 2 Proof. Since ||f (x)||2 = 1 for any x, we have the following: ||f (P (x, k1 )) − f (P (x, k2 ))||22 =||f (P (x, k1 ))||22 + ||f (P (x, k2 ))||22
n X
− 2f (P (x, k1 ))T · f (P (x, k2 ))
i=1
=2 − 2 · cos(f (P (x, k1 )), f (P (x, k2 ))),
(11)
f (P (x, ki )) = 0.
(13)
Based on Equation 10, we have the following equation group: ||f (P (x, k1 ))||22 − 2f T (P (x, k1 )) · c + ||c||22 ≤ r2 ||f (P (x, k2 ))||22 − 2f T (P (x, k2 )) · c + ||c||22 ≤ r2 , ··· 2 T 2 2 ||f (P (x, kn ))||2 − 2f (P (x, kn )) · c + ||c||2 ≤ r (14)
T
where T represents transpose and f (P (x, k1 )) · f (P (x, k2 )) is the inner product between two embedding vectors. Therefore, we have: Red (f, x, P, K) maxk1 ,k2 ∈K ||f (P (x, k1 )) − f (P (x, k2 ))||2 2 p maxk1 ,k2 ∈K 2 − 2cos(f (P (x, k1 )), f (P (x, k2 ))) = 2 r 1 − mink1 ,k2 ∈K cos(f (P (x, k1 )), f (P (x, k2 ))) = 2 p = Rcs (f, x, P, K). (12) =
where T indicates transpose of a vector. Since the embedding vectors lie on the unit hyper-sphere, we have ||f (P (x, ki ))||22 = 1 for i = 1, · · · , n. After summing up the n inequalities in the equation group 14, we have the following: n X n · 1 − 2( f T (P (x, ki ))) · c + n||c||22 ≤ nr2 ,
Therefore, by Theorem 1, the Euclidean distance-based metric Red does not satisfy the worst-robustness property.
i=1
⇔r2 ≥ 1 + ||c||22 ≥ 1,
C. DivergenceRadius Satisfies the Mathematical Properties
(15)
where r = 1 when c = 0. Since Rdr (f, x, P, K) is the smallest r, we have Rdr (f, x, P, K) = 1, achieving the worst-robustness property. Rotational invariance: Suppose the rotation matrix is M . We observe that ||M · f (P (x, k)) − M · c||2 = ||M · (f (P (x, k))−c)||2 = ||f (P (x, k))−c||2 ≤ Rdr (f, x, P, K) for any k ∈ K. Therefore, we have Rdr (M · f, x, P, K) ≤ Rdr (f, x, P, K). We denote the inverse matrix of the rotation matrix as M −1 . We have Rdr (f, x, P, K) = Rdr (M −1 · M · f, x, P, K) ≤ Rdr (M · f, x, P, K). Thus, we have Rdr (M · f, x, P, K) = Rdr (f, x, P, K), achieving
Proof. Bounded domain: The radius of any ball is nonnegative, and thus Rdr (f, x, P, K) ≥ 0. Moreover, since all embedding vectors outputted by a foundation model lie on the unit hyper-sphere, the unit ball with radius 1 can enclose all embedding vectors. Therefore, we have Rdr (f, x, P, K) ≤ 1. Thus, we have Rdr (f, x, P, K) ∈ [0, 1]. Monotonicity: Suppose the perturbation parameter domain K1 expands to K2 . If the embedding vectors of the perturbed 10
Table 4. Benchmark datasets. Dataset #Training images #Testing images #Classes
ImageNet
Food101
NYU-Depth V2
1,281,167
75,750
50,688
50,000
25,250
654
1,000
101
-
truth depth map y, we evaluate the root mean squared error of a downstream depth estimation model g ◦ f for x under a perturbation function P . Without loss of generality, we assume the perturbation parameter domain K contains m discrete values, where m discrete values can be sampled via random sampling or equally-spaced sampling if K contains more than m discrete values or is continuous. Formally, the RM SE of the downstream classifier g ◦ f for x under a perturbation function P with a perturbation parameter domain K is defined as below:
the rotation-invariance property. Essentially, the embedding vectors are enclosed in a minimum ball with center c and radius Rdr (f, x, P, K) before rotation, and the embedding vectors are enclosed in a minimum ball with center M c and the same radius after rotation.
RM SEp (x, g, f, P, K) v u n X u1 X 1 t (yi − g ◦ f (P (x, k))i )2 , = m n i=1
(16)
k∈K
D. Details of Zero-shot Classification, Linearprobe Classification, and Depth Estimation
where g ◦ f (P (x, k)) is the predicted depth map for image P (x, k) and i is the index of each pixel in a depth map. Intuitively, the smaller RM SE indicates the more accurate depth estimation.
• Zero-shot classification: Jointly pre-trained image and text models like CLIP can perform zero-shot classification without labeled training data. Given an image x and a set of text labels Y, CLIP represents x as an image embedding and the labels in Y as text embeddings, selecting the label most similar to x’s embedding. For instance, CLIP’s ViTL/14 achieves 68.38% accuracy on zero-shot ImageNet classification. • Linear-probe classification: Vision foundation models can be adapted for linear-probe classification using labeled data to fine-tune a simple classification head. With foundation model parameters frozen, a feed-forward classification head is added and trained on labeled data. This approach achieves high accuracy with minimal additional parameters; for example, DINO-v2 reaches 86.6% accuracy on ImageNet with a single-layer classification head. • Depth estimation: Vision foundation models can also support depth estimation, predicting per-pixel depth in an image. A convolutional decoder head (with batch normalization, a 1 × 1 convolution, ReLU activation, and a sigmoid function) is appended to the frozen foundation model and fine-tuned on NYU-Depth V2 training data. For example, DINO-v2 achieves a root mean squared error of 0.35 on NYU-Depth V2.
F. Parameter Settings and Experimental Results of Depth Estimation Parameter settings: We consider the downstream depth estimation model that achieves the lowest root mean squared error. Specifically, we use DINO v2 ViT-g/14 foundation model and the one-layer depth estimation head publicly released together with DINO v2 for NYU-Depth V2. Root mean squared error RM SE vs. root mean squared error under perturbation RM SEp : Table 6 shows the RM SE and average RM SEp under different perturbation functions of NYU-Depth V2 testing images for depth estimation model. First, we find that common perturbations degrade RMSE of downstream depth estimation models, i.e., average RM SEp is larger than RM SE. For instance, Frost blurring increases the RM SE of depth estimation by 0.12. Second, the increased RM SE caused by a perturbation function is aligned with the average DivergenceRadius under the perturbation function. For example, in Figure 15, DINO v2 ViT-g/14 have the largest average DivergenceRadius under Frost blurring compared to other perturbation functions. Correspondingly, in Table 6, average RM SEp under Frost blurring increases the most (0.12). This occurs because a higher average DivergenceRadius indicates more diversity in the embedding vectors of perturbed images, making it more likely that the downstream depth estimation models predict inaccurate depth maps for them.
E. Evaluation Metrics RM SE and RM SEp Given a downstream depth estimation model g ◦ f , where f is a foundation model and g is a depth estimation head built on top of f . The root mean squared error (denoted as RM SE) for a testing dataset is the average root mean squared error between model’s predicted depth map and the ground-truth depth map. We evaluate common perturbations to images, and thus we also consider root mean squared error under perturbation (denoted as RM SEp ) for each testing image. Specifically, given a testing image x with ground-
Root mean squared error under perturbation RM SEp vs. robustness value: Figure 8 shows the relationship between an image’s RM SEp under a perturbation function and the image’s corresponding robustness value (cosine similarity or DivergenceRadius) for the NYU-Depth V2 testing 11
Foundation Model CLIP ViT-B/16 CLIP ViT-L/14 CLIP RN50 CLIP RN50×64 DINO v2 ViT-L/14 DINO v2 ViT-g/14
Model Family CLIP CLIP CLIP CLIP DINO v2 DINO v2
Table 5. Foundation models. Pre-training Algorithm Multi-modal self-supervised learning Multi-modal self-supervised learning Multi-modal self-supervised learning Multi-modal self-supervised learning Self-supervised learning Self-supervised learning
Table 6. RM SE and average RM SEp of NYU-Depth V2’s testing images for depth estimation when using DINO v2 ViT-g/14 foundation model.
RM SEp
RM SE JPEG compression Brightness adjustment Contrast adjustment Defocus blurring Elastic blurring Fog blurring Frost blurring Gaussian noise Glass blurring
0.60
RMSE under perturbation RMSEp
RMSE under perturbation RMSEp
0.80
1.00
JPEG compression Brightness adjustment Contrast adjustment Defocus blurring Elastic blurring Fog blurring Frost blurring Gaussian noise Glass blurring
0.40 0.200.00
0.06 0.12 0.18 Cosine Similarity
(a) Zero-shot classification
0.24
0.80 0.60
0.35 0.37 (↑ 0.02) 0.36 (↑ 0.01) 0.35 (↑ 0.00) 0.39 (↑ 0.04) 0.38 (↑ 0.03) 0.38 (↑ 0.03) 0.47 (↑ 0.12) 0.37 (↑ 0.02) 0.38 (↑ 0.03)
Figure 9. Mean squared error of predicting an image’s RM SEp using its cosine similarity or DivergenceRadius for NYU-Depth V2 under different perturbation functions.
value: The linear relationship between RM SEp and the robustness value indicates that we can predict RM SEp of an image using its robustness value by a linear regression model. We evaluate the performance of such prediction. Towards this goal, we divide the testing images of a dataset into two halves. We use the pairs (robustness value, RM SEp ) of the first half of testing images to train a linear regression model, which takes a robustness value as input and outputs RM SEp . Then, we evaluate this linear regression model on the pairs (DivergenceRadius, RM SEp ) of the second half of the testing images. Figure 9 shows the mean squared errors of the linear regression models under the 9 perturbation functions for the depth estimation model. The mean squared errors are very small, which indicates that an image’s DivergenceRadius under a perturbation function can be used to accurately predict a downstream depth estimation model’s RM SE for the image when it is perturbed by the perturbation function.
JPEG compression Brightness adjustment Contrast adjustment Defocus blurring Elastic blurring Fog blurring Frost blurring Gaussian noise Glass blurring
0.40 0.200.00
0.06 0.12 0.18 DivergenceRadius
# Parameters (M) 86 304 38 420 304 1136
Glass Blurring Cosine Similarity Gaussian Noise DivergenceRadius Frost Blurring Fog Blurring Elastic Blurring Defocus Blurring Contrast Adjustment Brightness Adjustment JPEG Compression 0.00 1.00 2.00 3.00 4.00 5.00 6.00 Mean Square Error (x 10 3)
images and 9 perturbation functions. Specifically, given a perturbation function, for each testing image, we compute its robustness value and its RM SEp of a downstream depth estimation model, i.e., we obtain a pair (robustness value, RM SEp ). Then, we rank the pairs of all testing images in an increasing order according to the robustness values and divide the ranked pairs into 4 groups equally. For each group of pairs, we calculate the mean robustness value and mean RM SEp , which are shown in Figure 8. Across perturbation functions, we find that RM SEp roughly increases linearly when the robustness metrics increases. RM SEp increases as the robustness value increases because a larger robustness value indicates more diverse embedding vectors for the perturbed images, leading to less accurate depth estimation. 1.00
Architecture Vision Transformer Vision Transformer ResNet ResNet Vision Transformer Vision Transformer
0.24
(b) Linear-probe classification
Figure 8. Root mean squared error under perturbation RM SEp vs. (a) cosine similarity, (b) DivergenceRadius of NYUd testing images for depth estimation when different perturbation functions are used. We use the DINO v2 ViT-g/14 foundation model.
Predicting RM SEp of an image using its robustness 12
Table 7. Pearson correlations between average ACC-average ACCp and average cosine similarity or DivergenceRadius across the 9 perturbation functions for the two datasets and downstream classifiers. (a) Cosine similarity
Glass Blurring Gaussian Noise Frost Blurring Fog Blurring Elastic Blurring Defocus Blurring Contrast Adjustment Brightness Adjustment JPEG Compression 0.00
(b) DivergenceRadius
Zero-shot
Linear-probe
Zero-shot
Linear-probe
Classification
Classification
Classification
Classification
ImageNet
0.91
0.89
ImageNet
0.92
0.89
Food101
0.94
0.92
Food101
0.94
0.93
Glass Blurring Gaussian Noise Frost Blurring Fog Blurring Elastic Blurring Defocus Blurring Contrast Adjustment Brightness Adjustment JPEG Compression 0.00
Zero-shot Linear-probe
5.00 10.00 15.00 Mean Square Error (x 10 4)
(a) Cosine similarity
Glass Blurring Gaussian Noise Frost Blurring Fog Blurring Elastic Blurring Defocus Blurring Contrast Adjustment Brightness Adjustment JPEG Compression 0.00
Zero-shot Linear-probe 5.00 10.00 4 15.00 Mean Square Error (x 10 )
(a) Cosine similarity
Glass Blurring Gaussian Noise Frost Blurring Fog Blurring Elastic Blurring Defocus Blurring Contrast Adjustment Brightness Adjustment JPEG Compression 0.00
Zero-shot Linear-probe
5.00 10.00 15.00 Mean Square Error (x 10 4)
Zero-shot Linear-probe 5.00 10.00 4 15.00 Mean Square Error (x 10 )
(b) DivergenceRadius
(b) DivergenceRadius
Figure 10. Mean squared error of predicting an image’s ACCp using its (a) cosine similarity or (b) DivergenceRadius for ImageNet under different perturbation functions in two downstream classifiers.
Figure 11. Mean squared error of predicting an image’s ACCp using its (a) cosine similarity or (b) DivergenceRadius for Food101 under different perturbation functions in two downstream classifiers.
13
(h) Gaussian noise
RN
50 x6 4 ViT -L/ 14 ViT -g/ 14
4
50
RN
/16
L/1
4
4 L/1
g/1
ViT -
Models
ViT -
4
RN
50
x6
4
50
RN
L/1
ViT -B
/16
CLIP Dino v2
ViT -
4
4 L/1
ViT -
Cosine Similarity
0.20 0.15 0.10 0.05 0.00
g/1
ViT -
Models
(e) Elastic blurring
ViT -
4
50
RN
50
x6
4
RN
/16
L/1
ViT -
4
ViT -B
Models
CLIP Dino v2
ViT -B
RN
Models
(d) Defocus blurring CLIP Dino v2
(g) Frost blurring
0.20 0.15 0.10 0.05 0.00
50 x6 4 ViT -L/ 14 ViT -g/ 14
4
50
RN
L/1
ViT -
RN
ViT -B
/16
CLIP Dino v2
Cosine Similarity
0.20 0.15 0.10 0.05 0.00
ViT -
ViT -
Models
Cosine Similarity
Cosine Similarity
50 x6 4 ViT -L/ 14 ViT -g/ 14
4
50
RN
/16
L/1
Models
g/1
4 L/1 4
50
x6
RN
50
RN
4
/16
L/1
ViT -B
(f) Fog blurring
0.20 0.15 0.10 0.05 0.00
(c) Contrast adjustment CLIP Dino v2
ViT -
4
ViT -
ViT -
RN
ViT -B
50 x6 4 ViT -L/ 14 ViT -g/ 14
4
50
RN
/16
CLIP Dino v2
Cosine Similarity
0.20 0.15 0.10 0.05 0.00
g/1
4
ViT -
L/1
4
RN
50
x6
4
50
RN
L/1
/16
ViT -
ViT -B
Models
(b) Brightness adjustment
CLIP Dino v2
Models
L/1
ViT -
4
4
ViT -
Cosine Similarity
0.20 0.15 0.10 0.05 0.00
0.20 0.15 0.10 0.05 0.00
CLIP Dino v2
ViT -B
50
ViT -
RN
g/1
x6
L/1
4
50
4
/16
L/1
ViT -
RN
DivergenceRadius
ViT -B
Foundation Models
(a) JPEG compression
Cosine Similarity
Cosine Similarity
0.20 0.15 0.10 0.05 0.00
CLIP DINO v2
0.20 0.15 0.10 0.05 0.00
(i) Glass blurring
Figure 12. Average cosine similarity of ImageNet testing images for different foundation models and perturbation functions.
DivergenceRadius
0.24 0.16 0.08 0.00
Foundation Models
(e) Elastic blurring DivergenceRadius
0.24 0.16 0.08 0.00
Foundation Models
CLIP DINO v2
ViT -B/ 16 ViT -L/ 14 RN 50 RN 50 x6 4 ViT -L/ 14 ViT -g/ 14
CLIP DINO v2
Foundation Models
(g) Frost blurring
CLIP DINO v2
ViT -B/ 16 ViT -L/ 14 RN 50 RN 50 x6 4 ViT -L/ 14 ViT -g/ 14
Foundation Models
ViT -B/ 16 ViT -L/ 14 RN 50 RN 50 x6 4 ViT -L/ 14 ViT -g/ 14
Foundation Models
(f) Fog blurring
CLIP DINO v2
(d) Defocus blurring
0.24 0.16 0.08 0.00
CLIP DINO v2
ViT -B/ 16 ViT -L/ 14 RN 50 RN 50 x6 4 ViT -L/ 14 ViT -g/ 14
ViT -B/ 16 ViT -L/ 14 RN 50 RN 50 x6 4 ViT -L/ 14 ViT -g/ 14
DivergenceRadius
Foundation Models
(c) Contrast adjustment DivergenceRadius
DivergenceRadius
DivergenceRadius
0.24 0.16 0.08 0.00
0.24 0.16 0.08 0.00
ViT -B/ 16 ViT -L/ 14 RN 50 RN 50 x6 4 ViT -L/ 14 ViT -g/ 14
Foundation Models
(b) Brightness adjustment
CLIP DINO v2
CLIP DINO v2
ViT -B/ 16 ViT -L/ 14 RN 50 RN 50 x6 4 ViT -L/ 14 ViT -g/ 14
Foundation Models
(a) JPEG Compression 0.24 0.16 0.08 0.00
0.24 0.16 0.08 0.00
CLIP DINO v2
ViT -B/ 16 ViT -L/ 14 RN 50 RN 50 x6 4 ViT -L/ 14 ViT -g/ 14
ViT -B/ 16 ViT -L/ 14 RN 50 RN 50 x6 4 ViT -L/ 14 ViT -g/ 14
CLIP DINO v2
DivergenceRadius
0.24 0.16 0.08 0.00
DivergenceRadius
DivergenceRadius
0.24 0.16 0.08 0.00
Foundation Models
(h) Gaussian blurring
(i) Glass blurring
Figure 13. Average DivergenceRadius of Food101 testing images for different foundation models and perturbation functions.
4
4 L/1
g/1
ViT -
ViT -
4 x6 RN
50
4
50 RN
L/1
/16 Cosine Similarity
CLIP DINO v2
0.16
(g) Frost blurring
4
4 L/1
g/1
ViT -
Foundation Models
ViT -
4 x6 RN
50
4
50 RN
L/1
/16 ViT -B
(h) Gaussian blurring
ViT -
4
4
g/1
L/1
Foundation Models
ViT -
4 RN
ViT -
50
x6
4
50 RN
ViT -
/16
0.00
ViT -B
g/1 ViT -
ViT -
RN
Foundation Models
ViT -B
ViT -
0.24
0.08
4
4
4 x6 50
L/1
50 RN
L/1
4
/16
(f) Fog blurring
Foundation Models
(e) Elastic blurring
CLIP DINO v2
0.16
ViT -
4
4 g/1
L/1
50 RN
ViT -
x6
4
50 RN
L/1
ViT -B
4
/16
4
4 L/1
g/1
(d) Defocus blurring
0.00
ViT -
ViT -
Foundation Models
0.08
ViT -B
g/1
4
4
ViT -
L/1
4 x6 RN
50
4
50 RN
L/1
/16
ViT -
ViT -B
ViT -
RN
0.24
0.00
Foundation Models
0.00
(c) Contrast adjustment
0.08
0.00
ViT -
50
RN
x6
4
50
4 L/1
/16 ViT -B
Foundation Models
CLIP DINO v2
0.16
0.08
ViT -
4
4
L/1
ViT -
ViT -
RN
g/1
4 50
x6
4
50
0.24
Cosine Similarity
Cosine Similarity
RN
(b) Brightness adjustment
CLIP DINO v2
0.16
Foundation Models
0.08
0.00
Cosine Similarity
0.24
L/1
/16
(a) JPEG Compression
ViT -
ViT -
ViT -B
g/1
4
4
ViT -
L/1
4 x6 RN
50
4
50 RN
L/1
/16
ViT -
ViT -B
Foundation Models
0.16
0.08
0.00
CLIP DINO v2
Cosine Similarity
0.08
0.00
0.24
CLIP DINO v2
0.16
ViT -
0.08
0.00
0.24
CLIP DINO v2
0.16
Cosine Similarity
Cosine Similarity
Cosine Similarity
0.08
0.24
CLIP DINO v2
0.16
L/1
0.24
CLIP DINO v2
0.16
Cosine Similarity
0.24
(i) Glass blurring
Figure 14. Average cosine similarity of Food101 testing images for different foundation models and perturbation functions.
4 L/1
g/1 4
ViT -
ViT -
4 x6 RN 50
4 RN 50
/16
ViT -
ViT -B
4 L/1
g/1 4
ViT -
Foundation Models
ViT -
4 x6 RN 50
L/1
/16
ViT -
ViT -B
4 RN 50
CLIP DINO v2
DivergenceRadius
4 L/1
g/1 4
ViT -
4
ViT -
x6 RN 50
4 RN 50
L/1
ViT -
Foundation Models
Foundation Models
(e) Elastic blurring 0.18 0.12 0.06 0.00
(h) Gaussian blurring
L/1
DivergenceRadius
4
g/1 4
ViT -
L/1
4 x6
ViT -
RN 50
L/1
4
/16
ViT -
ViT -B
Foundation Models
(d) Defocus blurring
/16 ViT -B
(g) Frost blurring
RN 50
DivergenceRadius
4
g/1 4
ViT -
L/1
4 x6
g/1 4
ViT -
4
4
L/1
ViT -
x6
RN 50
RN 50
4
/16
L/1
ViT -
ViT -B
0.18 0.12 0.06 0.00
CLIP DINO v2
0.18 0.12 0.06 0.00
CLIP DINO v2
DivergenceRadius
CLIP DINO v2
DivergenceRadius
ViT -
RN 50
RN 50
Foundation Models
(c) Contrast adjustment
Foundation Models
CLIP DINO v2
0.18 0.12 0.06 0.00
4
/16
ViT -
ViT -B
(b) Brightness adjustment
L/1
DivergenceRadius
4 L/1
g/1 4
ViT -
Foundation Models
ViT -
4 x6 RN 50
4 RN 50
/16
ViT -
L/1
DivergenceRadius
g/1 4
ViT -
4
ViT -
L/1
4 x6 RN 50
Foundation Models
(f) Fog blurring
CLIP DINO v2
0.18 0.12 0.06 0.00
0.18 0.12 0.06 0.00
4 RN 50
/16
ViT -
ViT -B
L/1
DivergenceRadius
CLIP DINO v2
ViT -B
ViT -
g/1 4
4
ViT -
L/1
4 x6 RN 50
4 RN 50
/16
ViT -
ViT -B
Foundation Models
(a) JPEG Compression 0.18 0.12 0.06 0.00
CLIP DINO v2
0.18 0.12 0.06 0.00
L/1
DivergenceRadius
CLIP DINO v2
0.18 0.12 0.06 0.00
(i) Glass blurring
Figure 15. Average DivergenceRadius of NYU-Depth V2 testing images for different foundation models and perturbation functions.
14
Cosine Similarity
RN
50 x6 4 ViT -L/ 14 ViT -g/ 14
4
50 RN
ViT -L/ 1
0.12
CLIP DINO v2
0.06
4
4
Foundation Models
ViT -g/ 1
ViT -L/ 1
4 x6 50 RN
ViT -L/ 1
Foundation Models
(h) Gaussian blurring
50
ViT -B/ 16
4 ViT -L/ 14 ViT -g/ 14
x6 50 RN
4
RN
ViT -L/ 1
50
0.00
ViT -B/ 16
x6 4 ViT -L/ 14 ViT -g/ 14
50 RN
50 RN
4
ViT -L/ 1
Foundation Models
(g) Frost blurring
Foundation Models
(e) Elastic blurring 0.18
0.00
(f) Fog blurring
ViT -B/ 16
50 x6 4 ViT -L/ 14 ViT -g/ 14
RN
4
50 RN
CLIP DINO v2
0.06
ViT -B/ 16
4
4
ViT -g/ 1
ViT -L/ 1
4 x6 50 RN
50 RN
4
ViT -L/ 1
ViT -B/ 16
(d) Defocus blurring
RN
0.12
Foundation Models
4
0.18
0.00
Foundation Models
ViT -L/ 1
(c) Contrast adjustment
0.06
0.00
0.00
ViT -B/ 16
50 x6 4 ViT -L/ 14 ViT -g/ 14
4
50
Foundation Models
CLIP DINO v2
0.06
Cosine Similarity
0.12
0.06
RN
RN
50 x6 4 ViT -L/ 14 ViT -g/ 14
4
50
0.18
0.12
0.00
CLIP DINO v2
Cosine Similarity
Cosine Similarity
0.12
RN
ViT -L/ 1
(b) Brightness adjustment
CLIP DINO v2
0.18
Foundation Models
0.18
0.06
Cosine Similarity
(a) JPEG Compression
0.12
0.00
ViT -B/ 16
RN
50 x6 4 ViT -L/ 14 ViT -g/ 14
4
50 RN
ViT -L/ 1
ViT -B/ 16
Foundation Models
0.18
0.06
0.00
CLIP DINO v2
Cosine Similarity
0.12
0.06
0.00
CLIP DINO v2
0.18
RN
0.06
CLIP DINO v2
ViT -L/ 1
0.12
Cosine Similarity
0.18
ViT -B/ 16
0.12
Cosine Similarity
Cosine Similarity
CLIP DINO v2
0.18
(i) Glass blurring
Figure 16. Average cosine similarity of images in NYU-Depth V2 for different foundation models and perturbation functions.
JPEG compression Brightness adjustment Contrast adjustment Defocus blurring Elastic blurring Fog blurring 0.20 Frost blurring Gaussian noise Glass blurring 0.00 0.04 0.08 0.12 Cosine Similarity 0.40
0.16
0.80 0.60 0.40 0.200.00
JPEG compression Brightness adjustment Contrast adjustment Defocus blurring Elastic blurring Fog blurring Frost blurring Gaussian noise Glass blurring 0.06 0.12 0.18 Cosine Similarity
0.80 0.60 0.40 0.20
0.24
1.00 Accuracy under perturbation ACCp
0.60
1.00 Accuracy under perturbation ACCp
1.00 Accuracy under perturbation ACCp
Accuracy under perturbation ACCp
0.80
JPEG compression Brightness adjustment Contrast adjustment Defocus blurring Elastic blurring Fog blurring Frost blurring Gaussian noise Glass blurring 0.04 0.08 0.12 Cosine Similarity
0.16
0.80 0.60 0.40 0.200.00
JPEG compression Brightness adjustment Contrast adjustment Defocus blurring Elastic blurring Fog blurring Frost blurring Gaussian noise Glass blurring 0.06 0.12 0.18 Cosine Similarity
0.24
(a) ImageNet: Zero-shot classification(b) ImageNet: Linear-probe classifica-(c) Food101: Zero-shot classification(d) Food101: Linear-probe classification tion
Figure 17. Accuracy under perturbation ACCp vs. cosine similarity of ImageNet and Food101 testing images for zero-shot classification and linear-probe classification when different perturbation functions are used. Zero-shot classification is based on the CLIP ViT-L/14 foundation model and linear-probe classification is based on the DINO v2 ViT-g/14 foundation model.
15
Table 8. Details of perturbation functions. Perturbation Function
Key Perturbation Parameter K
Domain K
JPEG Compression
Quality factor
[30, 70]
Brightness Adjustment
Value in Hue-Saturation -Value space
[0.1, 0.5]
Contrast Adjustment
Amplifying factor of deviations from the mean
[0.3, 0.7]
Defocus Blurring
Bluring disk kernel
[1, 5]
Elastic Blurring
Scaling factor
[0.01, 0.05]
Fog Blurring
Density of fog
[0.5, 2.5]
Frost Blurring
Weight of the frost
[0.2, 0.6]
Gaussian Noise
Standard deviation
[0.02, 0.10]
Glass Blurring
Standard deviation
[0.2, 1.0]
Maximally Distorted Example
Table 9. Impact of λ on average DivergenceRadius and ACC of zero-shot classification for two datasets before/after enhancement, where the foundation model is CLIP ViT-L/14. λ 0.0 0.5 1.0 5.0
DivergenceRadius 0.01 0.05 0.06 0.07
ACC (%) 0.16 66.61 67.99 68.31
16