Conceptio › Archive › arXiv CS
arXiv CSopen access

Benchmarking the Explanatory Quality of Open-Weight Vision-Language Models in Face Recognition

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

Benchmarking the Explanatory Quality of Open-Weight Vision-Language Models in Face Recognition

arXiv:2609.21879v1 [cs.CV] 18 Sep 2026

Laurent Colbois1 and Sébastien Marcel1,2 1

Idiap Research Institute, Switzerland

2

Université de Lausanne, Switzerland

{laurent.colbois, sebastien.marcel}@idiap.ch

Abstract Vision-Language Models (VLMs) have recently been proposed as promising tools for face recognition, as they can produce natural language explanations alongside similarity scores. This capability is considered appealing for face comparisons in forensic contexts, which require decisions to be transparent and auditable. However, existing evaluations of VLMs for that use case focus mostly on recognition accuracy, while the validity of generated explanations remains unquantified. In this work, we introduce a benchmarking framework for VLM-based face recognition that treats explanation quality as a core evaluation axis. We propose two criteria that explanations should satisfy: relevance, i.e., reliance on identity-stable facial features; and faithfulness, i.e., alignment with the visible image content without hallucinated features. We jointly develop a methodology enabling the quantification of relevance and faithfulness of evaluated models, based on constraining model outputs to a structured explanation format that supports automated querying and auditing. Using this framework, we benchmark several families of open-weight VLMs, jointly evaluating face verification accuracy and explanation quality. Our results highlight remaining shortcomings of produced explanations, and emphasize the need for such explanation quality metrics to get a complete picture of model performance. The proposed benchmark and open-source evaluation harness provide a foundation for proper benchmarking and future fine-tuning of explainable face recognition systems.

1. Introduction Automated Face Recognition Systems (FRSs) have evolved into reliable tools for verification tasks in a variety of real-world applications, such as border control or mobile device security. They have been demonstrated to pro-

Are these images of the same person ?

JSON Schema (Structured output)

Vision Language Model

JSON Structured Explanations Score

{"similarity_score": 0.10, "explanation": "Eye-color clarity/hue and nose shape differ; overall suggests different individuals.", "feature_comparisons": [ {"feature": "eye_color", "description_a": "Underexposed; appears dark.", "description_b": "Bright, clearly visible; appears light.", "comparison": "Noticeable difference in clarity and hue." }, {"feature": "nose_shape", "description_a": "Broad with rounded tip.", "description_b": "Narrower with sharper tip.", "comparison": "Rounded vs sharper, narrower nose.", } … ]}

Face recognition metrics

Explanation quality metrics

Figure 1. We prompt VLMs with pairs of face images and a question requiring them to assess whether these images represent the same person or not. Moreover, the model is forced to structure its answer following a predefined JSON schema, thus enabling facilitated queryability of the generated explanations for automated scoring.

vide high levels of accuracy in the vast majority of situations, delivering convenience while maintaining a low false accept rate. Nevertheless, this high level of statistical per-

formance does not yet meet the standard required for highstakes, case-specific decisions, as is often the case in forensic face comparison tasks. In this context, the black-box nature of the neural networks used as underlying recognition models limits their explainability, challenging the applicability of FRSs by preventing the creation of auditable reports required by forensic standards as part of biometric analyses. Explainable AI (XAI) has been a longstanding research field, including works that aim to increase the explainability of FRSs. Early approaches have predominantly focused on the development of a posteriori pixel-based saliency maps [22], which provide insights into the influence of specific regions of the input face images on the final face similarity score. More recently, a novel line of work has emerged, enabled by the increasing availability of multimodal VLMs. Early studies have demonstrated the solid zero-shot capabilities of VLMs for face recognition [24, 19]. While their recognition accuracy has not yet reached the state of the art achieved by dedicated FRSs, it remains relatively high even in zero-shot approaches. More importantly, VLMs enable the face similarity score to be accompanied by text-based explanations, which are much closer to human-written reports than saliency maps and may therefore appear less abstract to the forensic community. However, we argue that the mere presence of text-based explanations is not, by itself, a guarantee of proper explainability. For forensic use, explanations must satisfy constraints that make them actually valid and valuable. In particular, explanations should avoid reliance on transient evidence (e.g., glasses, facial expression, lighting), and should not introduce unsupported claims by describing facial attributes that are not visible in the images (e.g., due to occlusions). Unfortunately, prior works generally do not discuss the validity of produced explanations beyond brief qualitative observations. In this work, we propose that explanation quality should be treated as a key evaluation axis in the benchmarking of VLMs for forensic face recognition, alongside recognition accuracy. Our aim is therefore to develop a benchmarking methodology that jointly evaluates recognition accuracy and explanation quality for VLM-based face verification, and to apply it to a range of open-weight models. Central to our approach is to constrain model outputs to a structured format that enables automated querying and auditing at scale. While several properties contribute to explanation quality (including factual correctness with respect to attribute annotations), this paper focuses on two dimensions that can be evaluated in a dataset-driven and auditable manner: relevance (exclusive reliance on identity-stable evidence) and faithfulness (avoidance of unsupported claims about nonvisible content). Our main contributions are the following: • We introduce two forensic-motivated explanation cri-

teria for VLM-based face recognition: relevance and faithfulness. • We propose a method to quantify these criteria by constraining outputs to a structured format that supports automated querying and auditing. • We benchmark multiple open-weight VLM families across scales, evaluating face verification performance alongside relevance and faithfulness, and analyze the effects of constrained decoding and model scaling. The code required to reproduce our experiments, including our benchmarking framework, is released publicly1 . We hope it can serve as an evaluation harness for future works focusing on improving VLM-based face recognition while preserving explanation quality.

2. Related Work 2.1. Guidelines for forensic face comparisons The practice of face comparison by human experts in forensic contexts predates the era of automated face recognition and remains prevalent today. Given the sensitivity of such analyses, the forensic community has devoted substantial effort to establishing standards that support consistent and rigorous practice. In particular, the Face Identification Scientific Working Group (FISWG) has issued standards for morphological analysis of human faces [2] and for assessing the stability of facial features among adults [3]. These standards are recommended by the European Network of Forensic Science Institutes (ENFSI) in their guidelines for face image comparisons [1]. Taken together, these documents provide a solid reference point for deriving requirements for explainable systems, as well as expectations for the explanations produced by such systems. They motivate both what should be described (morphology) and what should be avoided (unstable cues).

2.2. Explainability of vision models On the machine learning side, early work on explainability in vision tasks has largely focused on post-hoc saliency methods such as Grad-CAM [22]. These approaches have also been applied to face recognition [26, 15] with the goal of highlighting facial regions that most influence the predictions of automated FRSs. However, saliency maps have limitations highlighted in prior work that may strongly restrict their applicability in forensic settings. In particular, while a saliency map can indicate where a model is looking, it does not necessarily inform what it is looking at. For example, if the eye region is highlighted in a saliency map for face recognition, is it because pupil size is informative, because periocular 1 https://gitlab.idiap.ch/biometric/vlmfr

skin texture is informative, or because glasses are (incorrectly) treated as informative? Furthermore, the ambiguous nature of these heatmaps can inadvertently trigger confirmation bias, leading human examiners to interpret visualizations as supporting prior beliefs rather than objectively assessing whether the model relied on valid discriminative features [4, 5]. Further efforts aim to mitigate these issues through the development of Concept Bottleneck Models [13]. For image classification tasks, these models constrain the internal representation by first mapping images to a space of predefined, human-understandable concepts, and then performing classification starting only from that controlled space. Unfortunately, such approaches require conceptlabeled data at a scale and granularity that is not yet available for face recognition, especially given the level of finegrained detail needed to satisfy forensic expectations.

2.3. Vision-Language Models in Biometrics More recently, the emergence of VLMs has opened a new line of research on explainable systems. The rationale is that such models, in addition to performing vision tasks, can produce text-form rationales that may better match forensic expectations than saliency maps. The biometrics community has started exploring the use of VLMs for biometrics and face understanding. Seminal works [7, 10] provide early and systematic studies of the multimodal variant of ChatGPT for face biometrics, covering face verification, soft biometrics, and general face understanding, together with an initial discussion of explainability aspects. This has led to the development of multiple benchmark works [19, 18, 24] that more broadly evaluate the capability of VLMs, including open-weight models, to perform face verification and face understanding in zero-shot and few-shot settings. Together, these benchmarks provide valuable evidence that inference-time prompting can yield non-trivial biometric performance without model fine-tuning. While VLMs typically do not match the performance of dedicated task-specific networks on these tasks, these works emphasize the added value of text-based interaction and text generation as a non-negligible benefit. Efforts have also appeared to further improve VLMs’ performance through fine-tuning, such as FaceLLM [23] for face understanding tasks. A recurring theme across the above works is that multimodal foundation models can produce natural-language outputs, and can therefore be prompted to provide textbased rationales alongside biometric decisions. This property is particularly appealing for forensic applications, where practitioners typically expect auditable, humanreadable reporting rather than purely numerical scores. However, existing benchmarks and evaluations primarily emphasize task performance (e.g., verification accuracy or task success rates) and typically discuss explanation quality

only at a qualitative level. As a result, although the literature increasingly demonstrates that VLMs can both decide and describe, it remains unclear under which conditions the produced textual rationales are valid, stable, and aligned with forensic expectations. In that regard, faithfulness metrics for VLMs have been proposed in prior work [12], but their reliance on an LLMas-a-judge to convert free text into scorable items introduces an additional failure mode via judge-model errors, which is problematic. Instead, we prefer more constrained text production from the benchmarked VLM, enabling scoring through more predictable heuristics. Relevance metrics are typically task-specific, and there is therefore a need to establish such a metric for face recognition applications. The aim of this work is thus to advance the application of VLMs to face recognition by proposing a methodology that properly quantifies the quality of generated explanations, specifically with the goal of providing content useful to forensic experts.

3. Methodology

System Prompt You are a face identity verification expert. User Prompt Analyze the two provided face images (A and B). Compare them feature-by-feature to decide whether they show the same person or different people. Output a single JSON object with exactly these keys: • feature comparisons: list of facial feature comparisons; for each feature include: – feature: the facial feature name in snakecase format – description a: description of the feature in image A. – description b: description of the feature in image B. – comparison: concise comparison of this feature across A and B. • explanation: a brief global justification summarizing how the local comparisons lead to the final decision. • similarity score: a number between 0.0 (definitely different people) and 1.0 (definitely the same person). Do not nominally mention the names of the individuals in the images; focus only on facial characteristics. Respond strictly in JSON format with the specified keys. Figure 2. The prompts used for the structured setting

3.1. Datasets, models, prompting We study face verification in a VLM-based setting: given a pair of face images (Ia , Ib ), a VLM is prompted to output a global similarity score s ∈ [0, 1]. In addition to verification performance, we evaluate the quality of the generated explanations along two axes motivated by forensic reporting needs: relevance and faithfulness. We evaluate on four datasets: LFW [11], which is chosen as a commonly used baseline, ARFace [17] and Soteria [21], which contain occlusion-specific annotations enabling faithfulness measurements, and CelebA [16] to experiment at larger scale and in more diverse conditions. For LFW, we use the standard 6000-pair protocol (3000 genuine, 3000 impostor). For ARFace, Soteria, and CelebA, we construct LFW-like protocols with (i) balanced genuine/impostor pairs, and (ii) 10-fold identity-disjoint splits to avoid identity leakage between folds. We detect and crop faces using the InsightFace library [8]. When face detection fails on at least one image from a pair, the entire pair is treated as a failure-to-acquire (FTA) and excluded from metric computation. For benchmarking in a zero-shot setting, we select Gemma3 [25], Qwen2.5-VL [6], and InternVL3 [27] as three recent, widely-adopted open-weight VLM families spanning 1B–78B parameters, enabling a scaling analysis across distinct training recipes and vision backbones. We run them with deterministic decoding (temperature 0) and fixed seed to ensure reproducibility. We compare two output modes: • Unstructured explanations: The model is prompted to return (i) a float similarity score s ∈ [0, 1] and (ii) a free-form textual justification. This is the type of setting observed in prior works. • Structured explanations: We constrain outputs to a fixed JSON schema enforced at token sampling time using constrained decoding. The prompt is presented in Figure 2, and a corresponding JSON schema is applied to the model’s output to ensure compliance. Importantly, note that the number of feature comparisons, and the choice of compared features, is left completely free to the model. Small models occasionally fail to satisfy the schema; such cases are counted as FTAs. Figure 1 presents an example of the output produced by the models in the structured setting.

3.2. Evaluation metrics 3.2.1

Verification performance

Using the similarity score s, we report (i) Equal Error Rate (EER), as well as (ii) accuracy which is commonly used in prior VLM-for-face-verification literature. Given that VLMs are not yet at the level of operational performance,

we do not report false accept / false reject rates at specific thresholds, which would not be very informative in the current case. For accuracy, we follow a 10-fold protocol: on each fold, we select a threshold on the 9 remaining folds (e.g., maximizing accuracy) and evaluate on the held-out fold, then average across folds. We additionally report the FTA rate, counting failures due to face detection or invalid structured outputs. 3.2.2

Explanation quality metrics

We propose a quantification of two additional criteria: Relevance (identity-stable vs transient cues) Relevance measures whether explanations emphasize identity-stable facial evidence rather than transient or forensically unstable cues (e.g., glasses, facial expression, illumination). We define a lexicon L of unstable cues and measure how often they are invoked in the feature field of produced explanations. We report the unstable-feature rate U=

#{feature entries mentioning a cue in L} , #{all feature entries}

and the relevance score as its complement R = 1 − U . Faithfulness (no hallucinations beyond visible evidence) Faithfulness measures whether explanations respect what is actually visible in the images and avoid hallucinating details. For practical reasons, we propose a specialized version of faithfulness focused only on occlusion considerations, by penalizing mentions of facial parts known to be occluded. This can be the eyes or mouth area in the case of the ARFace dataset, and the nose or mouth area in the case of the Soteria dataset. Concretely, we again define a lexicon of the facial features which are occluded for each type of considered occlusion. Together with available occlusion labels from the datasets, we can flag an explanation item as hallucinated if the feature field refers to a facial part occluded in either Ia or Ib . We report the hallucination rate H=

#{feature entries referring to occluded parts} , #{all feature entries}

and faithfulness as F = 1 − H. For both relevance and faithfulness, we report a macroaverage across all comparisons: per-comparison rates are computed over the variable-length feature list, then averaged across comparisons. Table 1 presents the lexicons that are used in practice for our experiments. All experiments are run in PyTorch [20] with vLLM [14] for inference and constrained decoding. We will release

Table 1. Lexicons used for Relevance and Faithfulness metrics. Counts in parentheses indicate the total number of keywords per category. Note on occlusion acknowledgments: if a listed feature appears in the trigger lexicon but its description simply acknowledges occlusion (by containing these terms), we do not count it as a hallucination. The full lexicon is presented as supplementary material. Category

Example keywords Relevance metric: unstable features

Accessories (27) Attire (17) Environment (16) Expression (15) Grooming (16)

glasses, mask, hat, scarf, jewelry, ... shirt, jacket, collar, sleeve, outfit, ... background, lighting, shadow, blur, image, ... expression, smile, gaze, frown, gesture, ... makeup, hairstyle, beard, mustache, blemish, ...

Faithfulness metric: trigger words for facial regions Eyes (8) Nose (5) Mouth (7)

eye, iris, pupil, eyelid, periocular, ... nose, nostril, nasal, bridge, ... mouth, lips, teeth, smile, grin, ...

Faithfulness metric: occlusion acknowledgments Indicators (19)

occluded, covered, hidden, masked, obscured, not visible, hard to tell, partially visible, ...

prompts, JSON schemas, and evaluation code2 to support reproducibility and to serve as an evaluation harness for future fine-tuning of VLMs for face verification with auditable explanations.

4. Results 4.1. Verification performance in the structured setting Table 2 reports EER, accuracy, and failure-to-acquire (FTA) rates when enforcing structured JSON output. As a reference point, we also report the performance of a traditional state-of-the-art face recognition model. We use the Buffalo-L model pack from the InsightFace library, which is a ResNet50 trained on WebFace12M [28] with the ArcFace loss [9], and the RetinaFace face detector. Overall, VLMbased verification remains substantially behind dedicated face recognition systems, and performance varies across both model families and scales. In particular, Gemma327B achieves the lowest EERs across most datasets (e.g., 5.2% on LFW and 7.1% on ARFace), while Qwen2.5-VL and InternVL3 exhibit stronger dependence on scale, with small sizes yielding very high error rates (e.g., Qwen2.52 URL will be provided upon acceptance.

VL-3B and InternVL3-1B near chance-level EERs). Table 2 also highlights that structured decoding introduces additional failure modes: smaller models can fail to comply with the required JSON schema, resulting in elevated FTA (e.g., InternVL3-1B with ≈27–29% FTA across datasets). These failures are excluded from metric computation but are reported explicitly as an operational cost of structured explanations. Across families, scaling models up generally improves verification performance, but not systematically. For instance, Gemma3-12B underperforms Gemma3-4B on ARFace and Soteria, and InternVL3-78B is slightly worse than InternVL3-38B on LFW, ARFace and Soteria. Deviations in performance across families suggest that, beyond parameter count, family-specific training recipes and vision backbones can strongly impact face verification behavior.

4.2. Cost of structured output constraints Table 3 reports the difference in EER, accuracy, and FTA between structured and unstructured output modes. In most settings, enforcing structured explanations increases EER, indicating a recognition cost associated with constrained generation. This cost is particularly severe for smaller models, where the dominant issue is often the FTA increase caused by schema non-compliance rather than a pure degradation of similarity scoring. As models scale up, the gap between structured and unstructured performance decreases, suggesting that larger VLMs better tolerate output constraints and that the trade-off between recognition performance and auditable explanations becomes negligible at scale. Interestingly, a few configurations yield negative ∆EER (i.e., structured output slightly lowers EER), such as Gemma3-12B on LFW and Soteria. While this effect is not consistent across datasets, it suggests that structured decomposition can sometimes act as a regularizer by forcing models to perform a more explicit feature-by-feature comparison before producing a global score.

4.3. Scaling trends: accuracy vs. explanation quality From Table 2 we also observe that the EER generally decreases with scale, with consistent trends across datasets and families, albeit with the non-monotonic exceptions discussed above. Figure 3 presents similar scaling trends for the relevance and faithfulness metrics. It demonstrates that the proposed metrics reveal differences between models that are not visible from EER alone. For example, InternVL3-38B and Qwen2.5-VL-72B show comparable verification performance, yet InternVL3-38B exhibits substantially lower relevance, indicating a stronger tendency to rely on forensically unstable cues in its explanations. In other words, similar match performance can correspond to

Table 2. EER, Accuracy (ACC) and FTA, in % in the structured output setting. The best performing VLM (in EER, respectively accuracy) is bolded.

ArcFace Gemma3

Qwen2.5-VL

InternVL3

Dataset Metric

EER

LFW ACC

FTA

EER

ARFace ACC FTA

EER

4B 12B 27B 3B 7B 32B 72B 1B 8B 14B 38B 78B

0.3 9.5 7.6 5.2 44.3 25.1 8.0 5.2 48.9 9.8 6.2 5.3 5.8

99.9 90.4 92.6 94.7 55.5 75.0 92.1 94.9 51.6 90.3 93.7 94.5 94.2

0.0 0.1 0.0 0.0 2.6 0.4 0.0 0.0 28.7 0.0 0.0 0.0 0.0

0.1 11.3 17.1 7.1 39.8 28.4 14.4 12.6 49.0 12.3 9.8 10.1 11.8

100.0 88.9 85.1 93.0 60.3 71.5 86.2 87.6 56.0 87.8 90.5 90.7 88.9

0.3 9.0 10.1 4.6 38.3 26.1 10.5 7.6 48.9 9.6 7.0 5.6 6.3

0.0 0.1 0.0 0.0 1.8 0.3 0.0 0.0 26.6 0.0 0.0 0.0 0.0

Soteria ACC FTA 99.7 91.5 90.7 95.5 63.0 75.0 90.1 93.0 52.5 90.8 93.6 94.9 94.5

0.1 0.2 0.1 0.1 2.2 0.4 0.1 0.1 26.8 0.1 0.1 0.1 0.1

EER

CelebA ACC

FTA

3.0 13.5 12.1 10.4 44.2 28.0 11.1 11.6 48.7 17.6 13.3 13.2 12.7

98.1 86.5 88.3 89.6 55.7 71.9 89.4 88.6 51.6 83.0 86.9 86.9 87.4

0.2 0.2 0.2 0.2 2.6 0.4 0.2 0.2 29.2 0.2 0.2 0.2 0.2

Table 3. Difference in EER, Accuracy and FTA, in percentage points, in the structured output setting compared to the unstructured output setting. A positive value means that the corresponding metric is higher in the structured output setting than in the unstructured one.

Gemma3

Qwen2.5-VL

InternVL3

Dataset Metric

EER

LFW ACC

FTA

EER

ARFace ACC

FTA

EER

Soteria ACC

FTA

EER

CelebA ACC

FTA

4B 12B 27B 3B 7B 32B 72B 1B 8B 14B 38B 78B

+2.6 -2.4 +1.1 +14.1 +19.0 +2.9 +0.4 +1.7 +3.6 -1.6 +0.9 +1.9

-2.8 +0.5 -1.2 -16.2 -18.9 -2.6 -0.7 -2.2 -3.5 +1.3 -1.4 -1.6

+0.1 -0.0 -0.0 +2.5 +0.3 +0.0 -0.0 +23.0 +0.0 -0.0 -0.0 -0.0

+1.5 +0.2 +0.6 +8.5 +17.1 +3.2 +4.2 +1.3 +3.1 +3.6 +2.4 +2.6

-1.4 +5.6 +0.1 -10.7 -17.8 -3.0 -3.7 -0.4 -3.7 -3.4 -2.4 -2.7

+0.1 -0.0 -0.0 +1.8 +0.2 -0.0 -0.0 +22.4 +0.0 -0.0 +0.0 -0.0

+2.9 -4.0 +0.2 +8.0 +15.2 +2.5 +3.5 +2.4 +2.1 +1.8 +1.3 +2.6

-2.8 +2.7 -0.1 -7.5 -15.2 -3.2 -2.8 -1.5 -2.3 -1.1 -0.9 -2.0

+0.0 +0.0 +0.0 +2.0 +0.2 +0.0 +0.0 +21.5 -0.0 +0.0 -0.0 -0.0

+1.4 -1.5 +1.2 +12.9 +14.3 +0.6 +0.8 +1.4 +4.6 -2.1 +0.7 +0.6

-1.4 +0.4 -1.2 -13.5 -15.9 -0.9 -0.5 -1.3 -3.9 +2.2 -0.7 -1.5

+0.1 -0.0 -0.0 +2.4 +0.1 +0.0 +0.0 +24.2 -0.0 -0.0 +0.0 -0.0

very different explanation behaviors, motivating explanation quality as a complementary evaluation axis. We observe substantial remaining shortcomings: some models such as InternVL3 mention forensically unstable cues in up to 20% of feature entries, and refer to occluded regions in up to 10% of entries on occlusion datasets. Relevance scaling exhibits a notable pattern: the best relevance is typically achieved by intermediate-sized models (Gemma312B, Qwen2.5-VL-7B, InternVL3-14B in our benchmark), and relevance often decreases for the largest sizes. This

suggests that scaling can increase the frequency with which models invoke transient contextual cues (e.g., accessories, expression, image quality) when generating explanations. In contrast, faithfulness tends to improve more consistently with scale on occlusion-focused datasets, indicating that larger models are less prone to describing facial parts that are not visible under occlusions, although we again observe exceptions such as Gemma3-27B performing worse than Gemma3-12B.

(1-FTA) x Relevance Rate (%)

100

(1-FTA) x Faithfulness (%)

100 90 80 70 60 50 40

Gemma3

Qwen2.5-VL

InternVL3

90 80

LFW ARFace CelebA Soteria

70 60 50 2

4

8

16

Model Size (B)

32

64 2

4

Gemma3

8

16

Model Size (B)

32

64

2

4

8

16

Model Size (B)

Qwen2.5-VL

32

64

InternVL3

ARFace Soteria

2

4

8

16

Model Size (B)

32

64 2

4

8

16

Model Size (B)

32

64

2

4

8

16

Model Size (B)

32

64

Figure 3. Scaling trends for relevance and faithfulness metrics. We report (1 − FTA)× score for relevance and faithfulness, reflecting end-to-end usable explanation quality by accounting for cases where no valid structured explanation is produced (FTA).

4.4. Automatic audits: unstable cues and hallucinations Table 4 provides examples of the most frequently observed unstable cues. It illustrates the practical value of a structured approach that enables automated auditing. Reliance on hair and facial hair is common; while these may remain usable cues in some settings, they are more easily altered than stable facial morphology and can be problematic depending on the context. More clearly invalid cues (e.g., glasses/eyewear, expression, lighting, accessories, background, clothing) also occur frequently, despite being transient or irrelevant to identity. We additionally observe multiple cases where models produce nearly identical descriptions for images A and B, which may reflect a sequential generation bias where the first description influences the second. Table 4 further highlights occasional degenerate outputs (e.g., a description field containing an unrelated proper name), suggesting that additional validation checks may be beneficial beyond schema compliance. Figure 4 illustrates faithfulness failures on occlusion datasets: models sometimes describe attributes that are not visible due to sunglasses, scarves, or masks. Some errors can be interpreted as plausible confusions (e.g., sunglasses prompting “dark eyes”), but others are unsupported fabrications (e.g., detailed nose-shape claims when the lower face is masked). A potential contributing factor is again the sequential nature of generation: in our protocols, the first image is often not occluded while the second may be occluded, which can encourage the model to introduce a feature based

on the first image and then continue describing it for the second even when it is not observable. Table 4. Examples of unstable features in face descriptions. We showcase a mixture of examples from all models and all datasets. Feature

Description A

Description B

hair hairstyle

short, grey hair Long, straight hair framing the face. no glasses

no detectable hair Short, dark hair styled back.

glasses beard expression eyewear smile earring background cloth jewelry lighting shirt

Full, graying beard and mustache. Neutral expression. thick rim glasses with noticeable reflection smile Present on both ears Indoor setting with a cardboard box. white shirt with black horizontal stripes glasses worn Fairly even lighting, with some glare. light blue striped

black expansive glasses with reflective lenses Full, graying beard and mustache. Slightly intense expression. thin rim glasses with noticeable reflection Macie Taylor Missing Outdoor setting with trees and grass. plaid pattern shirt no visible glasses Bright sunlight. light blue striped

5. Discussion 5.1. Findings and implications This work introduces a benchmarking methodology for VLM-based face verification that evaluates not only recognition accuracy but also two forensic-motivated explanation quality criteria: relevance and faithfulness. Across

5.2. Limitations

(a) Dark-colored eyes.

(b) Thicker lips

(c) Darker eyes, slightly wider-set.

(d) Slightly wider (e) Image B also shows (f) Medium-sized nose, mouth with thinner a straight, medium- but appears narrower lips. sized nose. with a less broad bridge.

Figure 4. Examples of hallucinations in facial descriptions. The caption is the content of the description field produced by the VLM for that specific image.

three open-weight VLM families and multiple scales, we observe (i) a clear gap between VLM-based verification and dedicated face recognition systems, (ii) a consistent recognition cost associated with structured output constraints, which diminishes with model scale, and (iii) significant variations in explanation quality that are not predictable from EER alone. In particular, models with similar EER can differ sharply in their reliance on transient cues (relevance) and their tendency to hallucinate occluded facial parts (faithfulness). From a practitioner perspective, these results suggest that selecting a VLM for explanatory face verification should not be based solely on match performance. Larger models generally reduce both EER and schema-related FTA, but scaling does not reliably improve relevance, and intermediate-sized models can produce more forensically appropriate explanations even when their EER is slightly worse. Model choice therefore depends on the application goal: larger models offer better robustness when verification accuracy is the primary objective, while intermediate sizes may be preferable when minimizing reliance on transient cues is critical. The qualitative audits also highlight concrete failure modes that matter for forensic reporting, including frequent reliance on accessories, expression, illumination, or image quality, as well as descriptions of occluded facial regions. These provide clear guidance for further model refinement. The proposed framework additionally enables a complementary use case: when fine-tuning VLMs to improve face recognition performance, explanation quality can be monitored to detect unintended regressions.

Our relevance metric relies on lexicon-based detection of unstable cues and therefore captures a conservative subset of relevance violations; further effort would be necessary to ensure completeness of the lexicon. Faithfulness is restricted to focusing on occlusion annotations (ARFace and Soteria); it does not capture hallucinations unrelated to occlusion (e.g., invented scars) or errors on non-occluded datasets. The restriction to deterministic decoding prevents measuring the variability of the evaluated metrics, and we also do not quantify the effect of alternative prompting strategies or non-zero temperatures. Our verification protocol uses LFW, ARFace, Soteria, and CelebA. While ARFace and Soteria provide controlled occlusion conditions relevant to our faithfulness analysis, LFW and CelebA contain publicly available web imagery that may overlap with VLM pretraining data, so reported results may partially conflate generalization with memorization. Extending the evaluation to harder verification benchmarks (e.g., CFP-FP, AgeDB-30, CA-LFW, CP-LFW, IJB-B/C) would provide a more complete picture of VLM robustness and is left to future work. Finally, while this paper focuses on relevance and faithfulness, additional criteria are likely to be necessary for gaining a complete picture. In particular, we believe the evaluation of correctness (correspondence between VLM fine-grained facial attributes claims on visible, stable features, and the visual ground truth) remains an important direction, but is more challenging due to the need for associated labels of a level of fineness uncommon in available datasets.

5.3. Future work Several extensions follow from this benchmark. Stricter prompting and output constraints could be explored, e.g., restricting admissible features to a forensically-aligned lexicon, or extending the schema with explicit “visibility” or “confidence” fields enabling models to abstain when evidence is occluded. The schema could specifically be designed to follow FISWG morphological analysis guidelines, mapping explanations directly to forensic reporting structure. Fine-tuning models with objectives that explicitly target relevance and faithfulness (rather than verification accuracy alone) is another natural direction, with the automated metrics proposed here serving to validate explanation quality during training. Finally, user studies with forensic practitioners would be valuable to validate whether automated improvements correspond to perceived usefulness in casework reporting.

6. Conclusion We presented a benchmarking methodology for VLMbased face verification that evaluates not only recognition

accuracy, but also two forensic-motivated explanation quality criteria: relevance and faithfulness. By constraining model outputs to a structured format, our approach enables automated auditing of explanations at scale. Across three open-weight VLM families and multiple sizes, we observed a clear gap with dedicated FRSs, a measurable cost of structured constraints diminishing with scale, and substantial differences in explanation behavior between models with similar EER. We will release our benchmarking framework to enable future work on improving VLMs while preserving auditable explanations.

Acknowledgments This work was supported by the Center of Identification Technology Research (CITeR) and the Idiap Research Institute.

References [1] Best Practice Manual for Facial Image Comparison. Standard Version 1.0, European Network of Forensic Science Institutes (ENFSI), Jan. 2018. [2] Facial Image Comparison Feature List for Morphological Analysis. Standard Version 2.0, Facial Identification Scientific Working Group, Sept. 2018. [3] Physical Stability Of Facial Features Of Adults. Standard Version 2.0, Facial Identification Scientific Working Group, May 2021. [4] J. Adebayo, J. Gilmer, M. Muelly, I. Goodfellow, M. Hardt, and B. Kim. Sanity Checks for Saliency Maps. In Advances in Neural Information Processing Systems, volume 31. Curran Associates, Inc., 2018. [5] A. Alqaraawi, M. Schuessler, P. Weiß, E. Costanza, and N. Berthouze. Evaluating saliency map explanations for convolutional neural networks: A user study. In Proceedings of the 25th International Conference on Intelligent User Interfaces, IUI ’20, pages 275–285, New York, NY, USA, Mar. 2020. Association for Computing Machinery. [6] S. Bai, K. Chen, X. Liu, J. Wang, W. Ge, S. Song, K. Dang, P. Wang, S. Wang, J. Tang, H. Zhong, Y. Zhu, M. Yang, Z. Li, J. Wan, P. Wang, W. Ding, Z. Fu, Y. Xu, J. Ye, X. Zhang, T. Xie, Z. Cheng, H. Zhang, Z. Yang, H. Xu, and J. Lin. Qwen2.5-VL Technical Report, Feb. 2025. [7] I. Deandres-Tame, R. Tolosana, R. Vera-Rodriguez, A. Morales, J. Fierrez, and J. Ortega-Garcia. How Good Is ChatGPT at Face Biometrics? A First Look Into Recognition, Soft Biometrics, and Explainability. IEEE Access, 12:34390–34401, 2024. [8] J. Deng, J. Guo, E. Ververas, I. Kotsia, and S. Zafeiriou. RetinaFace: Single-Shot Multi-Level Face Localisation in the Wild. In 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 5202–5211, June 2020. [9] J. Deng, J. Guo, J. Yang, N. Xue, I. Kotsia, and S. Zafeiriou. ArcFace: Additive Angular Margin Loss for Deep Face

Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 44(10):5962–5979, Oct. 2022. [10] A. Hassanpour, Y. Kowsari, H. O. Shahreza, B. Yang, and S. Marcel. Chatgpt and Biometrics: An Assessment of Face Recognition, Gender Detection, and Age Estimation Capabilities. In 2024 IEEE International Conference on Image Processing (ICIP), pages 3224–3229, Oct. 2024. [11] G. B. Huang, M. Mattar, T. Berg, and E. Learned-Miller. Labeled Faces in the Wild: A Database for Studying Face Recognition in Unconstrained Environments. In Workshop on Faces in ’Real-Life’ Images: Detection, Alignment, and Recognition, Marseille, France, Oct. 2008. [12] L. Jing, R. Li, Y. Chen, and X. Du. FaithScore: Fine-grained Evaluations of Hallucinations in Large Vision-Language Models. In Y. Al-Onaizan, M. Bansal, and Y.-N. Chen, editors, Findings of the Association for Computational Linguistics: EMNLP 2024, pages 5042–5063, Miami, Florida, USA, Nov. 2024. Association for Computational Linguistics. [13] P. W. Koh, T. Nguyen, Y. S. Tang, S. Mussmann, E. Pierson, B. Kim, and P. Liang. Concept Bottleneck Models. In Proceedings of the 37th International Conference on Machine Learning, pages 5338–5348. PMLR, Nov. 2020. [14] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica. Efficient Memory Management for Large Language Model Serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles, SOSP ’23, pages 611–626, New York, NY, USA, Oct. 2023. Association for Computing Machinery. [15] Y.-S. Lin, Z.-Y. Liu, Y.-A. Chen, Y.-S. Wang, Y.-L. Chang, and W. H. Hsu. xCos: An Explainable Cosine Metric for Face Verification Task. ACM Trans. Multimedia Comput. Commun. Appl., 17(3s):112:1–112:16, Nov. 2021. [16] Z. Liu, P. Luo, X. Wang, and X. Tang. Deep Learning Face Attributes in the Wild. In 2015 IEEE International Conference on Computer Vision (ICCV), pages 3730–3738, Dec. 2015. [17] A. Martinez and R. Benavente. The AR Face Database: CVC Technical Report, 24. Jan. 1998. [18] K. Narayan, V. VS, and V. M. Patel. FaceXBench: Evaluating Multimodal LLMs on Face Understanding, Jan. 2025. [19] H. Otroshi Shahreza and S. Marcel. Foundation Models and Biometrics: A Survey and Outlook. IEEE Transactions on Information Forensics and Security, 20:9113–9138, 2025. [20] A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Kopf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala. PyTorch: An Imperative Style, HighPerformance Deep Learning Library. In Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. [21] N. Ramoly, A. Komaty, V. K. Hahn, L. Younes, A.-M. Awal, and S. Marcel. A Novel and Responsible Dataset for Face Presentation Attack Detection on Mobile Devices. In 2024 IEEE International Joint Conference on Biometrics (IJCB), pages 1–9, Sept. 2024.

[22] R. R. Selvaraju, M. Cogswell, A. Das, R. Vedantam, D. Parikh, and D. Batra. Grad-CAM: Visual Explanations from Deep Networks via Gradient-Based Localization. In 2017 IEEE International Conference on Computer Vision (ICCV), pages 618–626, Oct. 2017. [23] H. O. Shahreza and S. Marcel. FaceLLM: A Multimodal Large Language Model for Face Understanding. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 3677–3687, 2025. [24] R. Sony, P. Farmanifard, H. Alzwairy, N. Shukla, and A. Ross. Benchmarking Foundation Models for Zero-Shot Biometric Tasks, May 2025. [25] G. Team, A. Kamath, J. Ferret, S. Pathak, N. Vieillard, R. Merhej, S. Perrin, T. Matejovicova, A. Ramé, M. Rivière, L. Rouillard, T. Mesnard, G. Cideron, J.-b. Grill, S. Ramos, E. Yvinec, M. Casbon, E. Pot, I. Penchev, G. Liu, F. Visin, K. Kenealy, L. Beyer, X. Zhai, A. Tsitsulin, R. Busa-Fekete, A. Feng, N. Sachdeva, B. Coleman, Y. Gao, B. Mustafa, I. Barr, E. Parisotto, D. Tian, M. Eyal, C. Cherry, J.-T. Peter, D. Sinopalnikov, S. Bhupatiraju, R. Agarwal, M. Kazemi, D. Malkin, R. Kumar, D. Vilar, I. Brusilovsky, J. Luo, A. Steiner, A. Friesen, A. Sharma, A. Sharma, A. M. Gilady, A. Goedeckemeyer, A. Saade, A. Feng, A. Kolesnikov, A. Bendebury, A. Abdagic, A. Vadi, A. György, A. S. Pinto, A. Das, A. Bapna, A. Miech, A. Yang, A. Paterson, A. Shenoy, A. Chakrabarti, B. Piot, B. Wu, B. Shahriari, B. Petrini, C. Chen, C. L. Lan, C. A. Choquette-Choo, C. J. Carey, C. Brick, D. Deutsch, D. Eisenbud, D. Cattle, D. Cheng, D. Paparas, D. S. Sreepathihalli, D. Reid, D. Tran, D. Zelle, E. Noland, E. Huizenga, E. Kharitonov, F. Liu, G. Amirkhanyan, G. Cameron, H. Hashemi, H. KlimczakPlucińska, H. Singh, H. Mehta, H. T. Lehri, H. Hazimeh, I. Ballantyne, I. Szpektor, I. Nardini, J. Pouget-Abadie, J. Chan, J. Stanton, J. Wieting, J. Lai, J. Orbay, J. Fernandez, J. Newlan, J.-y. Ji, J. Singh, K. Black, K. Yu, K. Hui, K. Vodrahalli, K. Greff, L. Qiu, M. Valentine, M. Coelho, M. Ritter, M. Hoffman, M. Watson, M. Chaturvedi, M. Moynihan, M. Ma, N. Babar, N. Noy, N. Byrd, N. Roy, N. Momchev, N. Chauhan, N. Sachdeva, O. Bunyan, P. Botarda, P. Caron, P. K. Rubenstein, P. Culliton, P. Schmid, P. G. Sessa, P. Xu, P. Stanczyk, P. Tafti, R. Shivanna, R. Wu, R. Pan, R. Rokni, R. Willoughby, R. Vallu, R. Mullins, S. Jerome, S. Smoot, S. Girgin, S. Iqbal, S. Reddy, S. Sheth, S. Põder, S. Bhatnagar, S. R. Panyam, S. Eiger, S. Zhang, T. Liu, T. Yacovone, T. Liechty, U. Kalra, U. Evci, V. Misra, V. Roseberry, V. Feinberg, V. Kolesnikov, W. Han, W. Kwon, X. Chen, Y. Chow, Y. Zhu, Z. Wei, Z. Egyed, V. Cotruta, M. Giang, P. Kirk, A. Rao, K. Black, N. Babar, J. Lo, E. Moreira, L. G. Martins, O. Sanseviero, L. Gonzalez, Z. Gleicher, T. Warkentin, V. Mirrokni, E. Senter, E. Collins, J. Barral, Z. Ghahramani, R. Hadsell, Y. Matias, D. Sculley, S. Petrov, N. Fiedel, N. Shazeer, O. Vinyals, J. Dean, D. Hassabis, K. Kavukcuoglu, C. Farabet, E. Buchatskaya, J.-B. Alayrac, R. Anil, Dmitry, Lepikhin, S. Borgeaud, O. Bachem, A. Joulin, A. Andreev, C. Hardin, R. Dadashi, and L. Hussenot. Gemma 3 Technical Report, Mar. 2025. [26] J. R. Williford, B. B. May, and J. Byrne. Explainable Face Recognition. In A. Vedaldi, H. Bischof, T. Brox, and J.-M.

Frahm, editors, Computer Vision – ECCV 2020, pages 248– 263, Cham, 2020. Springer International Publishing. [27] J. Zhu, W. Wang, Z. Chen, Z. Liu, S. Ye, L. Gu, H. Tian, Y. Duan, W. Su, J. Shao, Z. Gao, E. Cui, X. Wang, Y. Cao, Y. Liu, X. Wei, H. Zhang, H. Wang, W. Xu, H. Li, J. Wang, N. Deng, S. Li, Y. He, T. Jiang, J. Luo, Y. Wang, C. He, B. Shi, X. Zhang, W. Shao, J. He, Y. Xiong, W. Qu, P. Sun, P. Jiao, H. Lv, L. Wu, K. Zhang, H. Deng, J. Ge, K. Chen, L. Wang, M. Dou, L. Lu, X. Zhu, T. Lu, D. Lin, Y. Qiao, J. Dai, and W. Wang. InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models, Apr. 2025. [28] Z. Zhu, G. Huang, J. Deng, Y. Ye, J. Huang, X. Chen, J. Zhu, T. Yang, D. Du, J. Lu, and J. Zhou. WebFace260M: A Benchmark for Million-Scale Deep Face Recognition. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(2):2627–2644, Feb. 2023.

A. Lexicons Complete lexicons used in the presented experiments for flagging hallucinations or exploitation of irrelevant features. Table 5. Lexicons used for Relevance and Faithfulness metrics. Note on the ”occlusion acknowledgments” list: if a listed feature is part of the trigger lexicon, but the description of that feature simply acknowledges occlusion (by containing these terms), we do not count it as a hallucination. Category

Keywords Relevance metric: unstable features

Accessories

Attire

Environment

Expression

Grooming

glasses, spectacles, eyewear, goggles, sunglasses, mask, respirator, shield, visor, hat, cap, helmet, beanie, hood, headwear, bandana, headband, scarf, tie, jewelry, earring, piercing, stud, hoop, necklace, clip, bindi cloth, shirt, jacket, coat, attire, garment, outfit, dress, suit, uniform, collar, sleeve, zipper, neckline, shoulder, chest, glove background, foreground, backdrop, scene, context, lighting, shadow, blur, focus, quality, noise, pixel, image, photo, picture, crop expression, emotion, mood, feeling, gaze, look, smile, smiling, laugh, frown, grimace, mouth open, teeth visible, wink, gesture makeup, cosmetic, mascara, liner, lipstick, hairstyle, haircut, hairdo, dye, beard, mustache, shaved, trim, acne, blemish, pimple

Faithfulness metric: trigger word for facial regions Eyes

eye, eyes, iris, pupil, eyelid, eyelash, eyelashes, periocular nose, nostril, nostrils, nasal, bridge mouth, lips, lip, teeth, smile, smiling, grin

Nose Mouth

Faithfulness metric: occlusion acknowledgments Indicators

occluded, covered, hidden, hiding, masked, masking, obscured, obscuring, not visible, not fully visible, not observable, cannot see, can’t see, hard to tell, unclear, not clear, out of frame, blocked, partially visible

Record · ID 1006891 · SHA-256 7296fd227ff3461a
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.