ConceptioArchivearXiv CS
arXiv CSopen access

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Maya Varma 1 Jean-Benoit Delbrouck 1 2 Sophie Ostmeier 1 Akshay Chaudhari * 1 Curtis Langlotz * 1

arXiv:2607.15216v1 [cs.CV] 16 Jul 2026

Abstract

1. Introduction

Multimodal large language models (MLLMs) often introduce errors when generating image captions, resulting in misaligned image-text pairs. Our work focuses on a class of captioning errors that we refer to as systematic misalignments, where a recurring error in MLLM-generated captions is closely associated with the presence of a specific visual feature in the paired image. Given a vision-language dataset with MLLM-generated captions, our aim in this work is to detect such errors, a task we refer to as systematic misalignment detection. As our first key contribution, we present S YMBAL, which utilizes a structured, dual-stage setup with off-the-shelf foundation models to identify systematic misalignments and summarize results in natural language. As our second key contribution, we introduce S YMBAL B ENCH, a benchmark designed to evaluate automated methods on our proposed task. S YMBAL B ENCH consists of 1.7 million image-text pairs from two domains (natural and medical images), organized into 420 vision-language datasets with annotated systematic misalignments. S YMBAL exhibits strong performance on this benchmark, correctly identifying systematic misalignments in 63.8% of datasets, a nearly 4x improvement over the closest baseline. We supplement our evaluations on S YMBAL B ENCH with real-world evaluations, showing that (1) S YMBAL can accurately surface systematic misalignments in captions generated by four MLLMs and (2) S YMBAL is a powerful tool for auditing off-the-shelf image-caption datasets. Ultimately, our novel task, method, and benchmark can aid users with auditing MLLMgenerated captions and identifying critical errors, without requiring access to the underlying MLLM. Code is available at https://github.com/ Stanford-AIMI/Symbal.

Multimodal large language models (MLLMs) possess strong image captioning capabilities yet often introduce errors into generated captions (Sarto et al., 2025; Zhou et al., 2024; Liu et al., 2025). As a result, images and paired MLLMgenerated captions may be misaligned, meaning that the generated text erroneously refers to features that are not visible in the image. For example, consider an MLLM that is tasked with generating a radiology report for an input medical image; in this setting, a misalignment may exist if the MLLM-generated report indicates the presence of cardiomegaly (a condition characterized by an enlarged heart) despite the image showing no evidence of this diagnosis. Misalignments can have severe consequences, particularly in safety-critical domains like medicine (Hardy et al., 2025; Nakaura et al., 2023). Our work focuses on a critical yet previously-underexplored subclass of captioning errors that we refer to as systematic misalignments. We term a misalignment as systematic when a recurring error in MLLM-generated captions is closely associated with the presence of a specific visual feature in the paired image. For example, in the medical domain, incorrect diagnoses of cardiomegaly in the MLLM-generated reports may be strongly associated with the presence of pacemakers (an implanted medical device that regulates the heartbeat) in the corresponding image (Sourget et al., 2025; Kumar et al., 2025). Systematic misalignments are a particularly egregious class of errors because they often arise due to spurious correlations or biases learned by MLLMs during training. As a result, systematic misalignments typically involve features that frequently co-occur in the real-world yet are not deterministically linked; for instance, while cardiomegaly and pacemakers do co-occur frequently, the presence of a pacemaker in a medical image does not necessarily imply that the patient has cardiomegaly. Thus, errors associated with systematic misalignments may seem highly plausible and are consequently challenging to detect. In this work, we introduce the systematic misalignment detection task with the goal of leveraging automated approaches to identify this challenging class of captioning errors. A method that aims to solve the systematic misalignment detection task will accept as input a vision-language dataset, which consists of images paired with free-form

* Equal senior authorship 1 Stanford University 2 HOPPR. Correspondence to: Maya Varma <[email protected]>.

Proceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026. Copyright 2026 by the author(s).

1

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Figure 1. Given an input dataset with thousands of images and paired MLLM-generated captions, the systematic misalignment detection task involves identifying recurring textual errors and associated visual features. Here, we provide example image-caption pairs from two datasets in S YMBAL B ENCH with expected outputs.

MLLM-generated captions. Then, as output, the method must identify textual errors (e.g. “cardiomegaly” in the previous example) that are systematically associated with visual features (e.g. “pacemaker” in the previous example).

label for a systematic misalignment. Methods are then quantitatively evaluated on the extent to which their predictions align with the ground truth. We evaluate S YMBAL using S YMBAL B ENCH, analyzing a range of approaches for each subtask. The best configuration of S YMBAL correctly identifies the systematic misalignment in 63.8% of S YMBAL B ENCH datasets. S YMBAL exhibits a nearly 4x improvement over the closest baseline, demonstrating the utility of our dual-stage, structured approach for addressing the systematic misalignment detection task. Finally, we supplement our evaluations on S YMBAL B ENCH with real-world evaluations, demonstrating quantitatively and qualitatively that (1) S YMBAL can accurately surface systematic misalignments in captions generated by four MLLMs and (2) S YMBAL is a powerful tool for auditing off-the-shelf datasets with MLLM-generated captions.

Addressing the systematic misalignment detection task with automated methods is challenging for the following two reasons. First, vision-language datasets provided as input to automated methods are often large in size with thousands of image-caption pairs; identifying global error patterns from such datasets is nontrivial, especially since the size of such datasets exceeds the reasoning capabilities of even state-ofthe-art models. Second, there are no existing benchmarks for comprehensively evaluating methods on their ability to discover systematic misalignments. In order to address these challenges, we present the following contributions: • We propose S YMBAL, an automated approach for detecting systematic misalignments in MLLM-generated captions.1 Our key insight is to structure the systematic misalignment detection task into two stages, with each stage comprised of individual subtasks. The first stage of S YMBAL focuses solely on identifying recurring textual errors in captions; to this end, S YMBAL clusters textual facts based on semantic similarity, scores each cluster by degree of misalignment with paired images, and summarizes the top-ranked cluster into a single unifying concept. The second stage of S YMBAL then leverages this information to identify and describe the associated visual feature.

Ultimately, we envision our novel task, benchmark, and method aiding in the following real-world contexts. First, our approach reveals insights into failure modes of trained MLLMs, which can (1) provide developers with critical information for building more robust models as well as (2) assist end-users with understanding limitations prior to realworld deployment. For instance, returning to our previous example, physicians using an MLLM in the clinic can be forewarned that generated reports tend to incorrectly diagnose “cardiomegaly” when X-rays have visible “pacemakers”; knowledge of this failure mode can allow for further manual review of model outputs on those cases. Second, our approach can help users identify systematic captioning errors in off-the-shelf datasets, even in black-box settings where access to the underlying MLLM is unavailable. This is a particularly important use-case, especially as publiclyavailable image datasets with MLLM-generated captions become widely used for training the next generation of multimodal foundation models.

• We introduce S YMBAL B ENCH, the first benchmark designed to evaluate automated methods for systematic misalignment detection. S YMBAL B ENCH consists of 420 image-caption datasets, each paired with a ground-truth 1 The acronym S YMBAL refers to systematic misalignment detection between images and language.

2

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Conflict of Interest Disclosure. None. Funding sources are listed in the Acknowledgments at the end of this paper.

et al., 2018) and (2) a pneumothorax detection model that achieves radiologist-level overall accuracy yet demonstrates high error rates when chest tubes, a medical device used for treatment, are absent (Oakden-Rayner et al., 2020). Detecting such failures is challenging due to the fact that relevant subgroups are typically not annotated in data.

2. Related Work Our work builds on three prior lines of study: (1) local misalignment detection methods that identify captioning errors at the per-sample level; (2) global error detection methods that summarize systematic trends in prediction errors; and (3) methods for describing patterns in large datasets with natural language.

A recent line of work has explored the development of automated methods for identifying global, systematic error patterns in classification settings. Given a validation dataset with images, model predictions, and ground-truth labels, these methods identify specific visual features (e.g. the beach background or the absence of tubes in the above examples) that are associated with higher error rates (Eyuboglu et al., 2022; Jain et al., 2023; Sohoni et al., 2020; Varma et al., 2024). Our work shares a similar goal in identifying systematic error patterns; however, we extend beyond the classification setting to the image captioning setting, where input datasets consist of images and paired model-generated captions. The inclusion of free-form text in input datasets presents an added level of complexity in comparison to labels. Additionally, we explicitly consider settings where ground-truth captions are unavailable.

Local Misalignment Detection: Given a single image and its paired model-generated caption, one line of recent work has focused on developing metrics that measure image-caption alignment using numeric scores. Examples include reference-free metrics like CLIPScore (Hessel et al., 2021) and PAC-S (Sarto et al., 2023), which do not require the existence of ground-truth captions; on the other hand, reference-based metrics such as BLEU (Papineni et al., 2002), ROUGE (Lin, 2004), CIDEr (Vedantam et al., 2015), METEOR (Banerjee & Lavie, 2005), and RefCLIPScore (Hessel et al., 2021) make use of ground-truth captions. The utility of such metrics is typically evaluated using image-caption benchmarks with human-annotated quality judgments (e.g. FLICKR8K-Expert (Hodosh et al., 2013), Pascal-50S (Vedantam et al., 2015), ReXVal (Yu et al., 2023)) or known model-injected errors (e.g. FOIL (Shekhar et al., 2017), ReXErr (Rao et al., 2025)).

Describing Datasets with Natural Language: Several works have presented approaches for describing patterns in large datasets using natural language (Burgess et al., 2025). In particular, recent studies have generated natural language descriptions (i) summarizing differences given two input datasets (Dunlap et al., 2024; Zhong et al., 2022) and (ii) summarizing model prediction errors given classification datasets with labels (Eyuboglu et al., 2022; Menon & Srivastava, 2024; Kim et al., 2024). Our work also involves summarizing dataset-level patterns with natural language; however, in our setting, datasets consist of images and paired captions, and descriptions must specifically identify systematic misalignments.

Several recent works have extended numeric scoring strategies by proposing interpretable metrics, which are capable of identifying the specific features in model-generated captions that are incorrect with respect to the image. Examples include reference-based metrics like CHAIR (Rohrbach et al., 2018), ALOHa (Petryk et al., 2024), and GREEN (Ostmeier et al., 2024) as well as reference-free metrics like FLEUR (Lee et al., 2024). Our work draws inspiration from these studies by also prioritizing interpretability; our method S YMBAL not only detects whether captioning errors are present but also provides users with a natural language output indicating the erroneous textual facts and associated visual cues. However, our study exhibits a key distinction from this line of work: whereas these metrics evaluate a single image and its paired model-generated caption, our work instead focuses on detecting global, systematic trends in captioning errors.

3. Task Definition In this section, we formally introduce the systematic misalignment detection task. Consider a vision-language dataset D = {(Vi , Ti )}N i=1 consisting of images V paired with free-form, model-generated text T . For example, dataset D may consist of chest X-rays V paired with MLLM-generated radiology reports T . We will express each text sample Ti as a collection of textual facts Ti = {ti1 , ti2 , ..., tini } and each image Vi as a collection of visual i features Vi = {v1i , v2i , ..., vm }. i

Global Error Detection: Due to visual biases or spurious correlations learned during training, machine learning models often make systematic prediction errors at test time. Selected examples in the classification setting noted by prior works include (1) an object recognition model that can correctly classify cows in pastoral settings yet demonstrates high error rates when cows are in beach settings (Beery

Dataset D may include misaligned samples, where text Ti does not accurately describe the content of the paired image Vi . We consider a pair (Vi , Ti ) to be misaligned if there exists at least one erroneous textual fact tik ∈ Ti that does not accurately describe any visual feature vji ∈ Vi . Mis3

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

alignments are particularly egregious when they occur in a systematic fashion, meaning that an erroneous textual fact t is repeatedly associated with the presence of a visual feature v throughout a dataset. For instance, in the medical imaging example discussed earlier, incorrect diagnoses of cardiomegaly in MLLM-generated reports are strongly associated with the presence of a pacemaker in the corresponding chest X-rays; this suggests the existence of a systematic misalignment between reports containing t = cardiomegaly and images containing v = pacemaker. Thus, given D, the goal of the systematic misalignment detection task is to discover textual errors t that are systematically associated with visual cues v. A method M : D → (t̂, v̂) that aims to solve the systematic misalignment detection task will accept dataset D as input; we note here that datasets may be large in size, consisting of thousands of image-text pairs. Then, method M will predict (t̂, v̂) as output, indicating the discovered textual error t̂ and associated visual feature v̂; here, both t̂ and v̂ will be expressed in text. We consider two variants of input dataset D: (1) referencefree, where each sample in dataset D = {(Vi , Ti )}N i=1 consists of image Vi and model-generated text Ti , and (2) reference-based, where each sample in dataset D = {(Vi , Ti , Ri )}N i=1 consists of an image Vi , model-generated text Ti , and a ground-truth reference caption Ri .

Figure 2. S YMBAL detects systematic misalignments with a twostage procedure. The first stage involves detecting erroneous textual facts, and the second stage involves detecting associated visual features.

4. Our Approach: S YMBAL

the medical imaging example discussed earlier, perhaps one such cluster will contain sentences from radiology reports that discuss the presence of cardiomegaly. To this end, all textual facts in D, forming the set SN we aggregate i T = {t : i = 1, ..., N ; k = 1, ..., ni }. Each texi k i=1 tual fact in this set is encoded using a text embedding model; then, embeddings are clustered using spherical K-Means, where the number of clusters is selected automatically using Silhouette distance.

The systematic misalignment detection task is made challenging by the fact that vision-language datasets may be complex and large in size; identifying global error patterns from such datasets is nontrivial. In this section, we address this challenge with our approach S YMBAL, which structures the systematic misalignment detection task into two stages. Each stage is comprised of three individual subtasks: grouping, scoring, and summarizing. Sections 4.1 and 4.2 discuss the two stages in detail.

• Scoring groups by degree of misalignment: Next, we score each cluster by computing the mean degree of alignment between constituent textual facts and paired images. Based on methods from prior work (Hessel et al., 2021; Dunlap et al., 2024; Chen et al., 2024a), we consider three options for measuring alignment between a given textual fact and its paired image: (1) embedding scorer, which computes embeddings for the text and image modalities and measures alignment as the cosine similarity, (2) textonly scorer, which generates a caption for the image and tasks an LLM with determining if the textual fact is accurate with respect to the caption, and (3) vision-language scorer, where a MLLM is provided both the image and the textual fact as input and tasked with determining if the textual fact is accurate. Low scores suggest that a large proportion of textual facts in the cluster are misaligned with respect to their paired images.

4.1. Stage 1: Detecting Erroneous Textual Facts The first stage of S YMBAL predicts the erroneous textual fact by (1) grouping semantically-similar facts that occur consistently throughout the dataset, (2) scoring each group of facts by degree of misalignment with paired images, (3) and summarizing the top-ranked group of facts into a single unifying concept t̂. The three subtasks associated with Stage 1 are detailed below: • Grouping semantically-similar facts: As defined in Section 3, we first express each text sample Ti as a collection of textual facts Ti = {ti1 , ti2 , ..., tini } by splitting captions at the sentence level. We then identify clusters of semantically-similar facts that occur in D; for example, in 4

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

• Summarizing the top-ranked group: Given the alignment scores computed in the previous step, we identify the cluster exhibiting the highest degree of misalignment, which we refer to as Ctext . Then, we apply a text-only summarizer, where an LLM is provided a list of textual facts in Ctext and tasked with identifying the unifying concept.

In Section 6.2, we evaluate the role of various image embedding models, alignment scorers, and summarizers. We note here that some datasets may contain multiple systematic misalignments; S YMBAL can be trivially extended to such settings, as we show in Appendix A and E.

5. Benchmark: S YMBAL B ENCH The final output of the summarizer is the predicted erroneous textual fact t̂; for example, in the medical example discussed earlier, the predicted textual fact may be t̂ = cardiomegaly. In Section 6.1, we evaluate the role of various text embedding models and alignment scorers.

The key challenge behind evaluating methods like S YMBAL on real-world vision-language datasets is that ground-truth systematic misalignments are typically unknown. Moreover, collecting human annotations for a task at this scale, where datasets include thousands of images paired with information-dense captions, is simply intractable. Thus, without access to ground-truth annotations, it becomes difficult (1) to determine whether misalignments identified by a method like S YMBAL are accurate and (2) to quantitatively compare results across multiple methods.

4.2. Stage 2: Detecting Associated Visual Features We now proceed to the second stage of S YMBAL, which predicts the associated visual feature by (1) grouping semantically-similar images paired with text containing fact t̂, (2) scoring each group of images by degree of misalignment with t̂, and (3) summarizing the top-ranked group of images into a single unifying concept v̂. The three subtasks associated with Stage 2 are detailed below:

In this section, we introduce S YMBAL B ENCH, which is designed to address this challenge. Specifically, S YMBAL B ENCH utilizes an automated method to inject a pre-defined systematic misalignment into a base vision-language dataset, yielding an evaluation setting where a ground-truth annotation (t, v) is available. The automated nature of our approach provides several key advantages, including (1) the ability to generate hundreds of evaluation settings simply by injecting varied systematic misalignments, (2) the presence of ground-truth labels that are guaranteed to be accurate, and (3) the ability to extend to specialized domains like medical imaging. In Section 6.4, we augment our evaluations on S YMBAL B ENCH with real-world analyses.

• Grouping semantically-similar images: We begin by identifying all images Vi ∈ D containing at least one paired textual fact in cluster Ctext (i.e. where tik ∈ Ctext for some k). Each image in this set is encoded using an image embedding model; then, embeddings are clustered using spherical K-Means, where the number of clusters is selected automatically using Silhouette distance. • Scoring groups by degree of misalignment: Next, we score each cluster by computing the mean degree of misalignment between images and paired textual facts in Ctext . We consider the same scoring mechanisms as in Stage 1. Low scores suggest that a large proportion of images in the cluster are misaligned with fact t̂.

Benchmark Design: S YMBAL B ENCH consists of 420 evaluation settings, where each setting is comprised of a visionlanguage dataset D and an associated ground-truth label (t,v) representing the systematic misalignment. In order to create each evaluation setting, we (1) obtain a high-quality base dataset with images and paired text, (2) predefine a systematic misalignment (t, v), and (3) inject the erroneous textual fact t into the base dataset such that a strong association exists with visual feature v. Below, we discuss these three steps in detail:

• Summarizing the top-ranked group: Given the alignment scores computed in the previous step, we identify the cluster exhibiting the highest degree of misalignment, which we will refer to as Cimage . Then, we consider two summarization mechanisms for identifying the unifying concept shared by images in Cimage : (1) text-only summarizer, where a caption is generated for each image in Cimage and an LLM is tasked with identifying the unifying concept, and (2) vision-language summarizer, where an MLLM is provided with images in Cimage and tasked with identifying the unifying concept.

1. Obtaining a base dataset. We begin by obtaining an off-the-shelf vision-language dataset with high-quality samples. We consider two options for the base dataset: COCO (2017 val split) (Lin et al., 2014) and MIMICCXR (test split) (Johnson et al., 2019a). COCO consists of natural images depicting common objects from 80 categories. After preprocessing, the base dataset includes a total of 4349 images with associated captions. MIMICCXR consists of chest X-rays and associated radiology reports obtained from the Beth Israel Deaconess Medical Center. After preprocessing, the base dataset includes

The final output of the summarizer is the predicted visual feature v̂; for example, in the medical example discussed earlier, the predicted visual feature may be v̂ = pacemaker. 5

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

2233 images, each paired with the “Impressions” section of the corresponding report.

B ENCH to analyze the choice of embedding models, alignment scorers, and summarizers. In Section 6.3, we perform end-to-end evaluations of the best configuration of S YM BAL, comparing with baselines and performing fine-grained analyses. Finally, in Section 6.4, we extend beyond S YM BAL B ENCH to real-world settings.

2. Predefining a systematic misalignment. Given a base dataset, we predefine a systematic misalignment consisting of a textual fact t and associated visual feature v. Predefined misalignments are meant to emulate those that are likely to emerge when using real-world, off-the-shelf MLLMs to generate captions. For COCO, we sample t and v from the set of 80 object categories present in the dataset. For MIMIC-CXR, we sample t from a set of five disease categories (cardiomegaly, pneumothorax, atelectasis, pleural effusion, and edema) and v from a set of five medical devices (pacemaker, chest tube, endotracheal tube, surgical clips, sternotomy wires).2

6.1. S YMBAL Detects Erroneous Textual Facts We first evaluate the role of various text embedding models, alignment scorers, and summarizers on the performance of Stage 1 of S YMBAL, which aims to predict the erroneous textual fact t̂s given an input dataset Ds in S YMBAL B ENCH. We compute Accuracy@1 and Accuracy@5 by comparing t̂s with ts across all 420 settings in S YMBAL B ENCH. Results are summarized in Table 1.

3. Injecting the predefined systematic misalignment. We insert the erroneous textual fact t into text samples in the base vision-language dataset such that a strong association exists between text containing t and images containing visual feature v. The strength of the association is controlled using Cramer’s V scores. Each inserted fact t is formatted as a sentence using diverse templates.

For the natural image datasets in S YMBAL B ENCH, Table 1 Upper demonstrates the performance of the top-four compositions, ranked by Accuracy@5 scores on the reference-free setting. Our results show that the best-performing variant of S YMBAL (shown in Row 1 of Table 1 Upper) achieves strong performance, correctly identifying the erroneous textual fact in 94.2% (Acc@5) of S YMBAL B ENCH datasets in the reference-free configuration and 82.8% (Acc@5) of S YMBAL B ENCH datasets in the reference-based configuration. Interestingly, we find that performance in referencefree settings is often substantially higher than performance in the reference-based setting, which is likely a result of the sparse information content often present in COCO reference captions. When considering the composition of S YMBAL, we note that the choice of the alignment scorer appears to be most important; the vision-language scorer substantially outperforms the text-only scorer with the same underlying model (Qwen2.5-72B).

We repeat this procedure across the two possible options for the base dataset and a range of possible options for t and v, yielding 420 evaluation settings encompassing a total of 1.7 million image-text pairs. Additional details are in Appendix B and C. Benchmark Evaluation: We will use the notation {(Ds , (ts , vs ))}420 s=1 to represent S YMBAL B ENCH , where the evaluation setting with index s has an associated dataset Ds and ground-truth label (ts , vs ). We construct both reference-based and reference-free variants of S YMBAL B ENCH, which differ only with respect to whether Ds includes reference captions. At evaluation time, dataset Ds will be provided to method M, which will output a prediction (t̂s , v̂s ). We count the prediction as accurate if the top-K predictions for t̂s include ts and the top-K predictions for v̂s include vs . Here, we evaluate equivalence using LLM-as-aJudge with Llama3.3-70B (Grattafiori et al., 2024). Overall performance on S YMBAL B ENCH is measured with Accuracy@K, computed as the percentage of the 420 settings in S YMBAL B ENCH where the prediction is accurate.

Given these results, we select the Qwen3-Embedding-8B text embedding model (Zhang et al., 2025), the visionlanguage alignment scorer with Qwen2.5-72B (Qwen et al., 2025), and the text-only summarizer with Qwen2.5-72B (Qwen et al., 2025) for all future S YMBAL evaluations on natural images. For the medical image datasets in S YMBAL B ENCH, Table 1 Lower demonstrates the performance of the top-four compositions. Our results show that the best-performing variant of S YMBAL (shown in Row 1 of Table 1 Lower) correctly identifies the erroneous textual feature in 75.0% (Acc@5) of datasets in the reference-free configuration and 95.0% (Acc@5) of datasets in the reference-based configuration. In contrast to the natural image datasets, we find that the reference-free configuration is harder than the referencebased configuration, likely due to the complexity of medical image data; alignment scoring in this domain is challenging without access to reference text. We also note that a key advantage of S YMBAL is its ability to extend to specialized

6. Results We now evaluate S YMBAL on the systematic misalignment detection task. In Sections 6.1 and 6.2, we use S YMBAL 2

We define these options for t and v due to the fact that medical imaging models often learn spurious associations between medical devices and disease categories, as documented in prior work (e.g. (Oakden-Rayner et al., 2020)); thus, our predefined misalignments are highly plausible in real-world, model-generated reports.

6

Symbal: Detecting Systematic Misalignments in Model-Generated Captions Table 1. We evaluate various text embedding models, alignment scorers, and summarizers on the performance of S YMBAL Stage 1.

Text-Only (Qwen-72B) Text-Only (Qwen-72B) Text-Only (Qwen-72B) Text-Only (Qwen-72B)

92.8 92.8 82.8 64.2

94.2 93.9 85.0 67.2

80.8 86.1 81.9 67.5

82.8 87.8 83.9 71.4

Text-Only (MedGemma-27B) Text-Only (Qwen-72B) Text-Only (MedGemma-27B) Text-Only (MedGemma-27B)

51.7 51.7 26.7 30.0

75.0 73.3 58.3 53.3

88.3 100.0 90.0 83.3

95.0 100.0 93.3 100.0

Summarizer

Natural

Reference-Based Acc@1 Acc@5

Alignment Scorer

Qwen3-8B OpenCLIP Qwen3-8B OpenCLIP

Vision-Language (Qwen-72B) Vision-Language (Qwen-72B) Text-Only (Qwen-72B) Text-Only (Qwen-72B)

Medical

Reference-Free Acc@1 Acc@5

Text Embedding

XRayCLIP XRayCLIP XRayCLIP MedSigLIP

Text-Only (MedGemma-27B) Text-Only (MedGemma-27B) Text-Only (Qwen-72B) Text-Only (MedGemma-27B)

Table 2. We evaluate various image embedding models, alignment scorers, and summarizers on the performance of S YMBAL Stage 2.

Text-Only (Qwen-72B) Vision-Language (Qwen-72B) Text-Only (Qwen-72B) Vision-Language (Qwen-72B)

49.7 48.1 47.8 45.8

69.7 63.9 62.8 62.5

41.9 42.5 43.9 38.9

52.2 55.6 55.8 52.2

Vision-Language (MedGemma-27B) Vision-Language (MedGemma-27B) Vision-Language (MedGemma-27B) Vision-Language (MedGemma-27B)

11.7 11.7 13.3 10.0

36.7 31.7 28.3 28.3

28.3 25.0 20.0 33.3

53.3 46.7 46.7 60.0

Summarizer

Natural

Reference-Based Acc@1 Acc@5

Alignment Scorer

OpenCLIP OpenCLIP OpenCLIP OpenCLIP

Vision-Language (Qwen-72B) Embedding (OpenCLIP) Embedding (OpenCLIP) Vision-Language (Qwen-72B)

Medical

Reference-Free Acc@1 Acc@5

Image Embedding

XRayCLIP MedSigLIP OpenCLIP MedSigLIP

Embedding (MedSigLIP) Embedding (MedSigLIP) Embedding (MedSigLIP) Embedding (XRayCLIP)

domains simply by interchanging constituent models with domain-specific versions.

with textual errors is substantially more challenging than identifying the textual error itself. We also observe that the best-performing variant of S YMBAL utilizes the same alignment scorer and summarizer as in Stage 1.

Given these results, we select the XRayCLIP-ViT-L text embedding model (Chen et al., 2024c), the text-only alignment scorer with MedGemma-27B (Sellergren et al., 2025), and the text-only summarizer with MedGemma-27B (Sellergren et al., 2025) for all future S YMBAL evaluations on medical images.

Given these results, we select the OpenCLIP-ViT-H image embedding model (Ilharco et al., 2021), vision-language alignment scorer with Qwen2.5-72B (Qwen et al., 2025), and text-only summarizer with Qwen2.5-72B (Qwen et al., 2025) for all future S YMBAL evaluations on natural images.

6.2. S YMBAL Detects Associated Visual Features

For the medical image datasets in S YMBAL B ENCH, Table 2 Lower demonstrates the performance of the top-four compositions, ranked by Accuracy@5 scores on the reference-free setting. Our results show that the best-performing variant of S YMBAL (shown in Row 1 of Table 2 Lower) correctly identifies the visual feature in 36.7% (Acc@5) of datasets in the reference-free configuration and 53.3% (Acc@5) of datasets in the reference-based configuration. Our results suggest that identifying visual features in the medical domain is a particularly challenging task in both reference-free and reference-based settings, and consequently, the optimal composition of alignment scorers and summarizers differs markedly from those identified in Stage 1.

We next evaluate the role of various image embedding models, alignment scorers, and summarizers on the performance of Stage 2 of S YMBAL. We hold the composition of Stage 1 constant using results from Section 6.1. We compute Accuracy@1 and Accuracy@5 by comparing v̂s with vs across all 420 settings in S YMBAL B ENCH. Results are summarized in Table 2. For the natural image datasets in S YMBAL B ENCH, Table 2 Upper demonstrates the performance of the top-four compositions, ranked by Accuracy@5 scores on the reference-free setting. Our results show that the best-performing variant of S YMBAL (shown in Row 1 of Table 2 Upper) correctly identifies the visual feature in 69.7% (Acc@5) of datasets in the reference-free configuration and 52.2% (Acc@5) of datasets in the reference-based configuration. We observe that performance values in Table 2 are lower than Table 1, suggesting that identifying visual features that systematically occur

Given these results, we select the XRayCLIP-ViT-L image embedding model (Chen et al., 2024c), embedding alignment scorer with MedSigLIP (Sellergren et al., 2025), and vision-language summarizer with MedGemma-27B (Sellergren et al., 2025) for future evaluations on medical images. 7

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

SymbalBench Performance

100

Reference-Free Configuration

Reference-Based Configuration

experimental settings, with GPT-OSS 120B correctly identifying the misalignment in only 17.1% of S YMBAL B ENCH datasets in the best case. These results demonstrate that the structured, dual-stage approach utilized by S YMBAL provides substantial performance benefits over single-stage, direct prompting baselines.

80 60 40 20

In Figure 4, we provide a stratified breakdown of S YM BAL performance. S YMBAL outperforms baselines across highly-challenging subsets of S YMBAL B ENCH where (1) the strength of the systematic misalignment is weak (i.e. weak association between the textual error and visual feature as measured by Cramer’s V scores) and (2) visual features are small in size.

0 Accuracy@1

Accuracy@5

Accuracy@1

Accuracy@5

● Llama3.3 (70B) ● Qwen2.5-VL (72B) ● GPT-OSS (120B) ● Symbal (Ours)

Figure 3. S YMBAL demonstrates strong end-to-end performance on S YMBAL B ENCH, substantially outperforming baselines. SymbalBench Performance

Strength of Association

Visual Feature Size

100

Extended results and ablations are provided in Appendix Section D.

80 60 40 20

6.4. S YMBAL Extends to Real-World Settings

0

Association Strength (Cramer’s V Score)

In this section, we further demonstrate the utility of S YM BAL by supplementing our evaluations on S YMBAL B ENCH with additional quantitative and qualitative analyses in realworld settings. Our results show that (1) S YMBAL can accurately surface systematic misalignments in captions generated by off-the-shelf MLLMs and (2) S YMBAL is a powerful tool for auditing vision-language datasets.

Size of Visual Feature (with respect to the area of the image)

● Llama3.3 (70B) ● Qwen2.5-VL (72B) ● GPT-OSS (120B) ● Symbal (Ours)

Figure 4. We report performance on S YMBAL B ENCH (referencefree) stratified across association strengths and visual feature sizes. This analysis focuses on natural image settings in S YMBAL B ENCH.

6.3. S YMBAL Shows Strong End-to-End Performance

S YMBAL can accurately surface systematic misalignments in captions generated by off-the-shelf MLLMs. First, we use S YMBAL to analyze captions generated by four real-world off-the-shelf MLLMs: Llava1.5-7B (Liu et al., 2024), Llava1.5-13B (Liu et al., 2024), AyaVision8B (Dash et al., 2025), and LlavaOneVision-7B (Li et al., 2025). We utilize each model to generate captions for the COCO dataset (2017 val split); we then apply S YMBAL (reference-free) to predict systematic misalignments (t̂, v̂).

Given an optimal composition of S YMBAL, we now perform end-to-end analyses across S YMBAL B ENCH. Since our study proposes a novel task, there are no existing baselines for comparison. As a result, we compare the structured, dual-stage approach of S YMBAL to a single-stage, directprompting method where each dataset Ds is directly provided to an off-the-shelf LLM in the form of a text prompt; the LLM is then instructed to output the erroneous textual fact and the associated visual feature. Three state-of-the-art LLMs are considered (i.e. Llama3.3 70B, Qwen2.5-VL 72B, and GPT-OSS 120B), selected to ensure a fair comparison with S YMBAL due to comparable parameter counts. As the token length of the direct prompts far surpasses the context window of these LLMs, we use only a sample of each dataset, ensuring that the final inference procedure requires no more compute resources than S YMBAL.

As discussed in Section 5, evaluating predictions in realworld settings is highly challenging since ground-truth systematic misalignments are unknown. Here, in order to address this issue, we validate identified systematic misalignments in two ways. First, we qualitatively validate the existence of S YMBAL-identified systematic misalignments with visual analysis. Second, we quantitatively validate whether a link between erroneous fact t̂ and visual feature v̂ truly exists; to this end, we measure whether model-generated captions are indeed more likely to include erroneous references to t̂ when v̂ is present compared to when v̂ is absent. In order to perform this evaluation, we use a state-of-the-art open-set object detector (Minderer et al., 2023) to annotate the presence of v̂ in each image, and we use our topperforming alignment scorer (vision-language scorer with Qwen-72B) to annotate erroneous references to t̂ in each caption. In Appendix E, we demonstrate that automated annotations align closely with human judgments.

In Figure 3, we measure the extent to which S YMBAL can accurately predict both the textual fact t̂s and the visual feature v̂s across both the reference-free and reference-based variants of S YMBAL B ENCH. Results show that the systematic misalignment detection task is highly challenging in both experimental settings, with several baselines generating few correct predictions. S YMBAL successfully identifies the systematic misalignment in up to 63.8% of datasets in S YMBAL B ENCH, with the highest performance observed in the reference-free setting (Accuracy@5). S YMBAL outperforms the closest baseline (GPT-OSS 120B) across all 8

Symbal: Detecting Systematic Misalignments in Model-Generated Captions Percentage of Captions with Erroneous Reference to White Tablecloth

Images

50

ShareGPT4V Captions

25 In the heart of a cozy kitchen, a woman and a man are sharing a moment of celebration. The woman, dressed in a vibrant blue shirt, is seated on the left side of the table. She's holding a rectangular cake, its surface adorned with lit candles that flicker in the soft light. Her smile is infectious, reflecting the joy of the occasion. On the right side of the table, a man in a gray shirt is seated. His gaze is directed towards the woman, perhaps sharing in her happiness or waiting for his turn to blow out the candles. The table they're sitting at is draped with a pristine white tablecloth, adding to the festive atmosphere…

In the heart of a cozy room, a group of people are gathered around a table, engrossed in conversation. The table, draped in a pristine white tablecloth, is adorned with plates, cups, and utensils, ready for a meal. A cake, the centerpiece of the gathering, sits in the middle of the table, inviting the guests to partake in its sweet delight.The room itself exudes a warm and inviting atmosphere. A bookshelf stands in the background, filled with various books that hint at the intellectual pursuits of the inhabitants. A window punctuates the wall, allowing natural light to filter into the room and illuminate the scene…

In the heart of a bustling restaurant, a group of children are gathered around a table, their faces alight with anticipation. The table, draped in a pristine white tablecloth, serves as the centerpiece of their gathering. On it, a plate of food awaits to be savored, while a jar of condiments stands by, ready to enhance the flavors of their meal. The children, dressed in casual attire, are engrossed in their own world, their attention focused on the plate of food. Their expressions are hidden from view, adding an air of mystery to the scene. In the background, the restaurant continues its lively rhythm…

In the heart of a cozy living room, a family of five is gathered around a wooden table, engrossed in the simple joy of a birthday celebration. The table, draped in a pristine white tablecloth, serves as the centerpiece of their gathering. On the table, a vibrant birthday cake steals the show. It's a feast for the eyes with its red, white, and blue colors. The cake is adorned with candles, their flames flickering in the soft light, casting a warm glow on the faces of the family. A woman, presumably the birthday celebrant, is in the midst of cutting the cake. Her hands are steady, her focus unwavering as she prepares to serve the first slice…

13.8%

0

0.8%

Images without table Images with table

Figure 5. S YMBAL discovers systematic misalignments in ShareGPT4V, an off-the-shelf dataset with model-generated Percentage ofcaptions. Captions with Erroneous

Images

ShareGPT4V Captions

Images

Reference to Printer S YMBAL identifies several systematic misalignments. In in the image compared to when a table is absent, vali50 captions generated by Llava1.5-7B, S YMBAL detects that dating the S YMBAL prediction. Additional examples are erroneous references to a handbag or a handbag on provided in Appendix E. the ground (t̂) in captions are often systematically assoAs large-scale datasets like ShareGPT4V become increasciated with the presence of a bus (v̂) in a scene, as shown ingly prevalent, it becomes critical for25users to be aware image captures of a busy The image captures a scene of a home office The image captures a wellorganized in Figure 9 The [Row 2].a scene Quantitatively, our analysis finds that workspace, brimming with various objects. The image captures a scene of a home office setup. Dominating scene is a wooden systematic workspace, bathed in the soft glow of ofthepotential misalignments, as these errors can Dominating the scene is a wooden desk, its setup. Dominating the scene is a wooden desk, standing against a green wall. The desk ambient light. Dominating the scene is a black erroneous references to a handbag in model-generated surface a testament to a mind at work. Two desk, bathed in the soft glow of a lamp is a hub of activity, hosting a variety of desk, its surface a tableau of productivity. models. if a dataset concomputer monitors stand side by side, their positioned on the left side. The desk is a hub objects. On thepropagate left side of the desk, anto open trained Two computer monitors standSpecifically, side by side, 7.5% captions arescreens indeed times more likely when a Abus laptop is sits, its screen glowing with unseen glowing with 3.1 unseen data. A of activity, hosting a variety of objects. their screens alive with data and information. keyboard and mouse lie in front of them, tools computer monitor stands as the centerpiece, data. Adjacenttains to it, a blacka monitor stands A keyboard and mouse lie in front of between them, systematic misalignment erroneous textual of the trade for the digital age. To the left of screen alive with vibrant hues tall, its screen blank. A black keyboard lies in tools of the trade for the digital age. To the present in the image compared toitsgreen when athebus isof aabsent, 0.1% the monitors, a phone rests, silent for now and black screensaver. To the right of front of the monitor, ready to translate of the monitors, a printer sits quietly, 0 fact t̂ and visualright feature v̂, models trained on the dataset but ready to connect at a moment's notice. the monitor, a phone lies idle, its cord trailing thoughts into words. A black mouse sits next ready to transform digital documents into validating the S YMBAL prediction. In captions generated On the right side of the desk, a printer waits off the edge of the desk. A keyboard and to the keyboard, poised to navigate the digital physical copies. Nearby, a phone rests, silent Images withoutt̂ and v̂, areof thelikely to learn spurious correlations between patiently for its next task. Scattered around mouse sit in front of the monitor, ready to world. To the right monitor, a black for now but capable of connecting this computer monitor the desk are various office supplies pens, spring into action at a that moment's notice. printer waits patiently for its next task. A workspace to the outside world. A bookshelf by LlavaOneVision-7B, S YMBAL detects erroneous pencils, and paper clips each with their own Nearby, a printer waits patiently for its next black phone rests next to it, silent stands guard in the background, its shelves leading tobut prediction errors at test-time (Varma et al., 2024). Images with role in the symphony of work. Above the task, while a stack of books suggests a thirst everconnected. A black lamp stands guard filled with knowledge and resources. In front references to text ( t̂) in captions are often systematically computer monitor desk, a shelf holds an array of books and for knowledge or perhaps a love for reading… next to the phone, ready to bathe the of the desk, a chair waits patiently for its S YMBAL users with understanding limitations of binders, a testament to knowledge… workspace in light when night falls… can aid occupant. associated with the presence of a sign (v̂) in a scene, as datasets with MLLM-generated captions as well as assist Percentage of Captions shown in Figure 10 [Row 2]. This finding suggests that model developers with improving performance of MLLMs. with Erroneous LlavaOneVision-7B struggles with OCR capabilities, where Reference to Black Phone the presence of text-based signage in an image is likely to 50 7. Discussion result in errors in the generated caption. Quantitatively, our analysis finds that erroneous references to text in modelIn this work, we introduce the systematic misalignment degenerated captions are indeed 4.6 times more likely when tection task, which aims to identify textual errors in MLLMa sign is present in the image compared to when a sign 25 The image captures a scene of a workspace, The image captures a scene of a home office generated captions that are systematically associated with bathed in the soft glowthe of a deskS lamp. setup. Dominating the scene is Additional a wooden is absent, validating YMBAL prediction. The image captures a wellorganized Dominating the scene is a wooden desk, its desk, standing against a green wall. The desk workspace, bathed in the softfeatures. glow of natural visual We hope that our novel task, method S YM grainy texture adding a touch of warmth to is a hub of activity, hosting a variety of streaming in from a window in the examples can beOnfound intheAppendix E.On the left side of the desk, an open light the setting. the left side of desk, a objects. background. Dominating the scene is a white laptop sits open, its screen glowing with laptop sits, its screen glowing with unseen BAL , and benchmark S YMBAL B ENCH can help 5.5%users audesk, its surface a tableau of productivity. On unseen data. To the right of the laptop, a data. Adjacent to it, a black monitor stands the left side of the desk, a black laptop sits waits patiently for the next its screen blank. A black keyboard lies in 0.1% S YMBAL iswhite akeyboard powerful tool fortall, auditing open-source dit MLLM-generated captions and identify critical failure open, its screen glowing with unseen data. burst of typing. In the center of the desk, a front of the monitor, ready to translate 0 Adjacent to it, a white printer stands ready black phone lies dormant, its screen dark. thoughts into words. A black mouse sits next for tasks. The right side of the desk is a hub of vision-language datasets. we poised useto navigate S YMBAL to modes, even without access to the underlying MLLM. It's as if it's patiently waiting for a call orSecond, to the keyboard, the digital activity with a white computer monitor message to break its silence. To the right of world. To the right of the monitor, a black Images without displaying a webpage, accompanied by a the phone, a black mouse sits idle, open-source its cord printer waitsimage patiently for its next task. A analyze ShareGPT4V, an dataset with white keyboard and mouse, tools of the laptop trailing off the edge of the desk. On the left black phone rests next to it, silent but digital age…A black phone lies nearby, silent side of the desk, a plant adds a touch of everconnected. A black lamp stands guard Images with laptop for now but ever ready for communication… MLLM-generated captions used astoabathepretraining greenery to the scene. Its leaves are commonly lush and next to the phone, ready the Impact Statement full, suggesting it's well cared for… workspace in light when night falls… dataset for vision-language models (Chen et al., 2024b). We sample a subset of 10k image-caption pairs from The goal of our work is to improve transparency into a critthe ShareGPT4V dataset, and we then apply S YMBAL ical class of captioning errors in image-text datasets. As (reference-free) to predict systematic misalignments (t̂, datasets with model-generated captions gain in popularity v̂). Here, S YMBAL detects that erroneous references to and become widely adopted into training datasets for the a white tablecloth (t̂) in captions are often systemnext generation of multimodal foundation models, it beatically associated with the presence of a table, cake, comes critical to audit data and understand potential quality and/or people (v̂) in the scene, as shown in Figure 5. issues before use. We hope that our novel task, benchmark, Quantitatively, our analysis finds that erroneous references and method can help make progress towards this goal, parto a white tablecloth in model-generated captions ticularly in safety-critical domains like medicine. are indeed 17.2 times more likely when a table is present

ShareGPT4V Captions

The image captures a scene of a workspace set against a vibrant red wall. Dominating the scene is a wooden desk, its surface adorned with various objects. On the left side of the desk, a laptop sits open, its screen glowing with unseen data. Adjacent to the laptop, a black phone rests, silent and unobtrusive. A white lamp with a curved neck stands sentinel on the right side of the desk, casting a soft glow that illuminates the immediate surroundings. The desk itself is a tableau of organized chaos, with papers scattered haphazardly, each one a testament to the work that has been done or is yet to be done. In the background, a window punctuates the red wall, offering a glimpse into the world outside. The image is taken from a low angle, adding a sense of depth and perspective to the scene…

9

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Acknowledgments

Dunlap, L., Darrell, T., and Yeung-Levy, S. Video action differencing. In The Thirteenth International Conference on Learning Representations, 2025. URL https:// openreview.net/forum?id=3bcN6xlO6f.

MV is supported by graduate fellowship awards from the Knight-Hennessy Scholars program at Stanford University, the Quad program, and the United States Department of Defense (NDSEG). AC is supported by NIH grants R01 HL167974, R01HL169345, R01 AR077604, R01 EB002524, R01 AR079431, P41 EB027060, AY2 AX000045, and 1AYS AX0000024-01; ARPA-H grants AY2AX000045 and 1AYSAX0000024-01; and NIH contracts 75N92020C00008 and 75N92020C00021. AC has provided consulting services to Patient Square Capital, Chondrometrics GmbH, and Elucid Bioimaging; is cofounder of Cognita; has equity interest in Cognita, Subtle Medical, LVIS Corp, Brain Key. CL is supported by NIH grants R01 HL155410, R01 HL157235, by AHRQ grant R18HS026886, and by the Gordon and Betty Moore Foundation. CL is also supported by the Medical Imaging and Data Resource Center (MIDRC), which is funded by the National Institute of Biomedical Imaging and Bioengineering (NIBIB) under contract 75N92020C00021 and through the Advanced Research Projects Agency for Health (ARPA-H).

Chen, D., Chen, R., Zhang, S., Wang, Y., Liu, Y., Zhou, H., Zhang, Q., Wan, Y., Zhou, P., and Sun, L. MLLMas-a-judge: Assessing multimodal LLM-as-a-judge with vision-language benchmark. In Forty-first International Conference on Machine Learning, 2024a. URL https: //openreview.net/forum?id=dbFEFHAD79. Chen, L., Li, J., Dong, X., Zhang, P., He, C., Wang, J., Zhao, F., and Lin, D. Sharegpt4v: Improving large multi-modal models with better captions. In European Conference on Computer Vision, pp. 370–387. Springer, 2024b. Chen, Z., Varma, M., Xu, J., Paschali, M., Veen, D. V., Johnston, A., Youssef, A., Blankemeier, L., Bluethgen, C., Altmayer, S., Valanarasu, J. M. J., Muneer, M. S. E., Reis, E. P., Cohen, J. P., Olsen, C., Abraham, T. M., Tsai, E. B., Beaulieu, C. F., Jitsev, J., Gatidis, S., Delbrouck, J.-B., Chaudhari, A. S., and Langlotz, C. P. A visionlanguage foundation model to enhance efficiency of chest x-ray interpretation, 2024c. URL https://arxiv. org/abs/2401.12208.

This research was funded, in part, by the Advanced Research Projects Agency for Health (ARPA-H). The views and conclusions contained in this document are those of the authors and should not be interpreted as representing the official policies, either expressed or implied, of the U.S. Government.

Dash, S., Nan, Y., Dang, J., Ahmadian, A., Singh, S., Smith, M., Venkitesh, B., Shmyhlo, V., Aryabumi, V., BellerMorales, W., Pekmez, J., Ozuzu, J., Richemond, P., Locatelli, A., Frosst, N., Blunsom, P., Gomez, A., Zhang, I., Fadaee, M., Govindassamy, M., Roy, S., Gallé, M., Ermis, B., Üstün, A., and Hooker, S. Aya vision: Advancing the frontier of multilingual multimodality, 2025. URL https://arxiv.org/abs/2505.08751.

References Banerjee, S. and Lavie, A. METEOR: An automatic metric for MT evaluation with improved correlation with human judgments. In Goldstein, J., Lavie, A., Lin, C.-Y., and Voss, C. (eds.), Proceedings of the ACL Workshop on Intrinsic and Extrinsic Evaluation Measures for Machine Translation and/or Summarization, pp. 65–72, Ann Arbor, Michigan, June 2005. Association for Computational Linguistics. URL https://aclanthology.org/ W05-0909/.

Delbrouck, J.-B., Chambon, P., Chen, Z., Varma, M., Johnston, A., Blankemeier, L., Van Veen, D., Bui, T., Truong, S., and Langlotz, C. RadGraph-XL: A largescale expert-annotated dataset for entity and relation extraction from radiology reports. In Ku, L.-W., Martins, A., and Srikumar, V. (eds.), Findings of the Association for Computational Linguistics: ACL 2024, pp. 12902– 12915, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024. findings-acl.765. URL https://aclanthology. org/2024.findings-acl.765/.

Bannur, S., Bouzid, K., Castro, D. C., Schwaighofer, A., Thieme, A., Bond-Taylor, S., Ilse, M., Pérez-Garcı́a, F., Salvatelli, V., Sharma, H., Meissen, F., Ranjit, M., Srivastav, S., Gong, J., Codella, N. C. F., Falck, F., Oktay, O., Lungren, M. P., Wetscherek, M. T., Alvarez-Valle, J., and Hyland, S. L. Maira-2: Grounded radiology report generation, 2024. URL https://arxiv.org/abs/ 2406.04449.

Dunlap, L., Zhang, Y., Wang, X., Zhong, R., Darrell, T., Steinhardt, J., Gonzalez, J. E., and Yeung-Levy, S. Describing differences in image sets with natural language. In Conference on Computer Vision and Pattern Recognition (CVPR), 2024.

Beery, S., Van Horn, G., and Perona, P. Recognition in terra incognita. In Proceedings of the European Conference on Computer Vision (ECCV), September 2018.

Eyuboglu, S., Varma, M., Saab, K., Delbrouck, J.-B., LeeMesser, C., Dunnmon, J., Zou, J., and Ré, C. Domino:

Burgess, J., Wang, X., Zhang, Y., Rau, A., Lozano, A., 10

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Discovering systematic errors with cross-modal embeddings. International Conference on Learning Representations (ICLR), 2022. doi: 10.48550/ARXIV.2203.14960. URL https://arxiv.org/abs/2203.14960.

Press, 2019. ISBN 978-1-57735-809-1. doi: 10.1609/ aaai.v33i01.3301590. URL https://doi.org/10. 1609/aaai.v33i01.3301590.

Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Letman, A., Mathur, A., et al. The llama 3 herd of models, 2024. URL https://arxiv. org/abs/2407.21783.

Jain, S., Lawrence, H., Moitra, A., and Madry, A. Distilling model failures as directions in latent space. In The Eleventh International Conference on Learning Representations, 2023. URL https://openreview.net/ forum?id=99RpBVpLiX. Johnson, A. E. W., Pollard, T. J., Greenbaum, N. R., Lungren, M. P., ying Deng, C., Peng, Y., Lu, Z., Mark, R. G., Berkowitz, S. J., and Horng, S. Mimic-cxr-jpg, a large publicly available database of labeled chest radiographs, 2019a. URL https://arxiv.org/abs/ 1901.07042.

Hardy, R., Kim, S. E., Ro, D. H., and Rajpurkar, P. Rextrust: A model for fine-grained hallucination detection in ai-generated radiology reports. In Wu, J., Zhu, J., Xu, M., and Jin, Y. (eds.), Proceedings of The First AAAI Bridge Program on AI for Medicine and Healthcare, volume 281 of Proceedings of Machine Learning Research, pp. 173–182. PMLR, 25 Feb 2025. URL https://proceedings.mlr.press/ v281/hardy25a.html.

Johnson, J., Douze, M., and Jégou, H. Billion-scale similarity search with GPUs. IEEE Transactions on Big Data, 7 (3):535–547, 2019b.

Hessel, J., Holtzman, A., Forbes, M., Le Bras, R., and Choi, Y. CLIPScore: A reference-free evaluation metric for image captioning. In Moens, M.-F., Huang, X., Specia, L., and Yih, S. W.-t. (eds.), Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 7514–7528, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/v1/2021.emnlp-main. 595. URL https://aclanthology.org/2021. emnlp-main.595/. Hodosh, M., Young, P., and Hockenmaier, J. Framing image description as a ranking task: Data, models and evaluation metrics. Journal of Artificial Intelligence Research, 47: 853–899, August 2013. ISSN 1076-9757. doi: 10.1613/ jair.3994. URL http://dx.doi.org/10.1613/ jair.3994. Ilharco, G., Wortsman, M., Wightman, R., Gordon, C., Carlini, N., Taori, R., Dave, A., Shankar, V., Namkoong, H., Miller, J., Hajishirzi, H., Farhadi, A., and Schmidt, L. Openclip, July 2021. URL https://doi.org/10. 5281/zenodo.5143773. Irvin, J., Rajpurkar, P., Ko, M., Yu, Y., Ciurea-Ilcus, S., Chute, C., Marklund, H., Haghgoo, B., Ball, R., Shpanskaya, K., Seekins, J., Mong, D. A., Halabi, S. S., Sandberg, J. K., Jones, R., Larson, D. B., Langlotz, C. P., Patel, B. N., Lungren, M. P., and Ng, A. Y. Chexpert: a large chest radiograph dataset with uncertainty labels and expert comparison. In Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence and Thirty-First Innovative Applications of Artificial Intelligence Conference and Ninth AAAI Symposium on Educational Advances in Artificial Intelligence, AAAI’19/IAAI’19/EAAI’19. AAAI

Kim, Y., Mo, S., Kim, M., Lee, K., Lee, J., and Shin, J. Discovering and mitigating visual biases through keyword explanation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 11082–11092, June 2024. Kumar, A., Kriz, A., Havaei, M., and Arbel, T. PRISM: High-resolution & precise counterfactual medical image generation using language-guided stable diffusion. In Medical Imaging with Deep Learning, 2025. URL https://openreview.net/forum? id=UpJMAlZNuo. Lee, Y., Park, I., and Kang, M. FLEUR: An explainable reference-free evaluation metric for image captioning using a large multimodal model. In Ku, L.-W., Martins, A., and Srikumar, V. (eds.), Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 3732–3746, Bangkok, Thailand, August 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.acl-long. 205. URL https://aclanthology.org/2024. acl-long.205/. Li, B., Zhang, Y., Guo, D., Zhang, R., Li, F., Zhang, H., Zhang, K., Zhang, P., Li, Y., Liu, Z., and Li, C. LLaVA-onevision: Easy visual task transfer. Transactions on Machine Learning Research, 2025. ISSN 28358856. URL https://openreview.net/forum? id=zKv8qULV6n. Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, X., and Wen, J.-R. Evaluating object hallucination in large vision-language models. In The 2023 Conference on Empirical Methods in Natural Language Processing, 2023. URL https: //openreview.net/forum?id=xozJw0kZXF.

11

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Lin, C.-Y. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out, pp. 74–81, Barcelona, Spain, July 2004. Association for Computational Linguistics. URL https: //aclanthology.org/W04-1013/.

York, NY, USA, 2020. Association for Computing Machinery. ISBN 9781450370462. doi: 10.1145/ 3368555.3384468. URL https://doi.org/10. 1145/3368555.3384468. OpenAI, Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F. L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al. Gpt-4 technical report, 2024. URL https://arxiv.org/abs/2303. 08774.

Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In Fleet, D., Pajdla, T., Schiele, B., and Tuytelaars, T. (eds.), Computer Vision – ECCV 2014, pp. 740–755, Cham, 2014. Springer International Publishing. ISBN 978-3-319-10602-1.

Oquab, M., Darcet, T., Moutakanni, T., Vo, H. V., Szafraniec, M., Khalidov, V., Fernandez, P., HAZIZA, D., Massa, F., El-Nouby, A., Assran, M., Ballas, N., Galuba, W., Howes, R., Huang, P.-Y., Li, S.-W., Misra, I., Rabbat, M., Sharma, V., Synnaeve, G., Xu, H., Jegou, H., Mairal, J., Labatut, P., Joulin, A., and Bojanowski, P. DINOv2: Learning robust visual features without supervision. Transactions on Machine Learning Research, 2024. ISSN 2835-8856. URL https:// openreview.net/forum?id=a68SUt6zFt. Featured Certification.

Liu, H., Li, C., Li, Y., and Lee, Y. J. Improved baselines with visual instruction tuning. In 2024 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 26286–26296, 2024. doi: 10.1109/CVPR52733.2024. 02484. Liu, Y., Liang, Z., Wang, Y., Wu, X., Tang, F., He, M., Li, J., Liu, Z., Yang, H., Lim, S., and Zhao, B. Unveiling the ignorance of mllms: Seeing clearly, answering incorrectly. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 9087–9097, June 2025. Menon, R. and Srivastava, S. DISCERN: Decoding systematic errors in natural language for text classifiers. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N. (eds.), Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pp. 19565–19583, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.emnlp-main. 1091. URL https://aclanthology.org/2024. emnlp-main.1091/. Minderer, M., Gritsenko, A. A., and Houlsby, N. Scaling open-vocabulary object detection. In Thirty-seventh Conference on Neural Information Processing Systems, 2023. URL https://openreview.net/forum? id=mQPNcBWjGc. Nakaura, T., Yoshida, N., Kobayashi, N., Shiraishi, K., Nagayama, Y., Uetani, H., Kidoh, M., Hokamura, M., Funama, Y., and Hirai, T. Preliminary assessment of automated radiology report generation with generative pre-trained transformers: comparing results to radiologistgenerated reports. Japanese Journal of Radiology, 42 (2):190–200, September 2023. ISSN 1867-108X. doi: 10.1007/s11604-023-01487-y. URL http://dx.doi. org/10.1007/s11604-023-01487-y. Oakden-Rayner, L., Dunnmon, J., Carneiro, G., and Re, C. Hidden stratification causes clinically meaningful failures in machine learning for medical imaging. In Proceedings of the ACM Conference on Health, Inference, and Learning, CHIL ’20, pp. 151–159, New 12

Ostmeier, S., Xu, J., Chen, Z., Varma, M., Blankemeier, L., Bluethgen, C., Md, A. E. M., Moseley, M., Langlotz, C., Chaudhari, A. S., and Delbrouck, J.-B. GREEN: Generative radiology report evaluation and error notation. In Al-Onaizan, Y., Bansal, M., and Chen, Y.-N. (eds.), Findings of the Association for Computational Linguistics: EMNLP 2024, pp. 374–390, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-emnlp. 21. URL https://aclanthology.org/2024. findings-emnlp.21/. Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. Bleu: a method for automatic evaluation of machine translation. In Isabelle, P., Charniak, E., and Lin, D. (eds.), Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pp. 311–318, Philadelphia, Pennsylvania, USA, July 2002. Association for Computational Linguistics. doi: 10.3115/1073083.1073135. URL https://aclanthology.org/P02-1040/. Petryk, S., Chan, D. M., Kachinthaya, A., Zou, H., Canny, J., Gonzalez, J. E., and Darrell, T. ALOHa: A new measure for hallucination in captioning models. In Duh, K., Gomez, H., and Bethard, S. (eds.), Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 2: Short Papers), pp. 342–357, Mexico City, Mexico, June 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024. naacl-short.30. URL https://aclanthology. org/2024.naacl-short.30/.

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Qwen, :, Yang, A., Yang, B., Zhang, B., Hui, B., Zheng, B., Yu, B., Li, C., Liu, D., Huang, F., Wei, H., et al. Qwen2.5 technical report, 2025. URL https://arxiv.org/ abs/2412.15115.

Systems, volume 33, pp. 19339–19352. Curran Associates, Inc., 2020. URL https://proceedings. neurips.cc/paper/2020/file/ e0688d13958a19e087e123148555e4b4-Paper. pdf.

Rao, V. M., Zhang, S., Acosta, J. N., Adithan, S., and Rajpurkar, P. Rexerr: Synthesizing clinically meaningful errors in diagnostic radiology reports. In Biocomputing 2025, pp. 70–81, 2025. doi: 10.1142/9789819807024 0006. URL https://www.worldscientific.com/doi/ abs/10.1142/9789819807024_0006.

Sourget, T., Hestbek-Møller, M., Jiménez-Sánchez, A., Junchi Xu, J., and Cheplygina, V. Mask of truth: Model sensitivity to unexpected regions of medical images. Journal of Imaging Informatics in Medicine, 2025. ISSN 2948-2933. doi: 10.1007/ s10278-025-01531-5. URL http://dx.doi.org/ 10.1007/s10278-025-01531-5.

Rohrbach, A., Hendricks, L. A., Burns, K., Darrell, T., and Saenko, K. Object hallucination in image captioning. In Riloff, E., Chiang, D., Hockenmaier, J., and Tsujii, J. (eds.), Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pp. 4035–4045, Brussels, Belgium, OctoberNovember 2018. Association for Computational Linguistics. doi: 10.18653/v1/D18-1437. URL https: //aclanthology.org/D18-1437/.

Varma, M., Delbrouck, J.-B., Chen, Z., Chaudhari, A., and Langlotz, C. RaVL: Discovering and mitigating spurious correlations in fine-tuned vision-language models. In Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., and Zhang, C. (eds.), Advances in Neural Information Processing Systems, volume 37, pp. 82235– 82264. Curran Associates, Inc., 2024. doi: 10.52202/ 079017-2614.

Sarto, S., Barraco, M., Cornia, M., Baraldi, L., and Cucchiara, R. Positive-augmented contrastive learning for image and video captioning evaluation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 6914–6924, June 2023. Sarto, S., Cornia, M., and Cucchiara, R. Image captioning evaluation in the age of multimodal llms: challenges and future perspectives. In Proceedings of the Thirty-Fourth International Joint Conference on Artificial Intelligence, IJCAI ’25, 2025. ISBN 978-1-956792-06-5. doi: 10. 24963/ijcai.2025/1180. URL https://doi.org/10. 24963/ijcai.2025/1180. Sellergren, A., Kazemzadeh, S., Jaroensri, T., Kiraly, A., Traverse, M., Kohlberger, T., Xu, S., Jamil, F., et al. Medgemma technical report, 2025. URL https:// arxiv.org/abs/2507.05201. Shekhar, R., Pezzelle, S., Klimovich, Y., Herbelot, A., Nabi, M., Sangineto, E., and Bernardi, R. FOIL it! find one mismatch between image and language caption. In Barzilay, R. and Kan, M.-Y. (eds.), Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 255–265, Vancouver, Canada, July 2017. Association for Computational Linguistics. doi: 10.18653/v1/P17-1024. URL https://aclanthology.org/P17-1024/.

Varma, M., Delbrouck, J.-B., Ostmeier, S., Chaudhari, A., and Langlotz, C. TRoVe: Discovering error-inducing static feature biases in temporal vision-language models. In Belgrave, D., Zhang, C., Lin, H., Pascanu, R., Koniusz, P., Ghassemi, M., and Chen, N. (eds.), Advances in Neural Information Processing Systems, volume 38, pp. 9934–9967. Curran Associates, Inc., 2025. Vedantam, R., Lawrence Zitnick, C., and Parikh, D. Cider: Consensus-based image description evaluation. In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), June 2015. Yu, F., Endo, M., Krishnan, R., Pan, I., Tsai, A., Reis, E. P., Fonseca, E. K. U. N., Lee, H. M. H., Abad, Z. S. H., Ng, A. Y., Langlotz, C. P., Venugopal, V. K., and Rajpurkar, P. Evaluating progress in automatic chest x-ray radiology report generation. Patterns, 4(9):100802, 2023. ISSN 2666-3899. doi: https://doi.org/10.1016/j.patter.2023.100802. URL https://www.sciencedirect.com/ science/article/pii/S2666389923001575. Zhang, Y., Jiang, H., Miura, Y., Manning, C. D., and Langlotz, C. P. Contrastive learning of medical visual representations from paired images and text. Machine Learning for Healthcare, abs/2010.00747, 2022. URL https://arxiv.org/abs/2010.00747.

Sohoni, N., Dunnmon, J., Angus, G., Gu, A., and Ré, C. No subclass left behind: Fine-grained robustness in coarse-grained classification problems. In Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M. F., and Lin, H. (eds.), Advances in Neural Information Processing

Zhang, Y., Li, M., Long, D., Zhang, X., Lin, H., Yang, B., Xie, P., Yang, A., Liu, D., Lin, J., Huang, F., and Zhou, J. Qwen3 embedding: Advancing text embedding and reranking through foundation models. arXiv preprint arXiv:2506.05176, 2025. 13

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Zhong, R., Snell, C., Klein, D., and Steinhardt, J. Describing differences between text distributions with natural language. In Chaudhuri, K., Jegelka, S., Song, L., Szepesvari, C., Niu, G., and Sabato, S. (eds.), Proceedings of the 39th International Conference on Machine Learning, volume 162 of Proceedings of Machine Learning Research, pp. 27099–27116. PMLR, 17–23 Jul 2022. URL https://proceedings.mlr.press/ v162/zhong22a.html. Zhou, Y., Cui, C., Yoon, J., Zhang, L., Deng, Z., Finn, C., Bansal, M., and Yao, H. Analyzing and mitigating object hallucination in large vision-language models. In The Twelfth International Conference on Learning Representations, 2024. URL https://openreview.net/ forum?id=oZDJKTlOUe.

14

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Appendix Contents • A. Implementation Details for S YMBAL . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 15 • B. Implementation Details for S YMBAL B ENCH . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 18 • C. S YMBAL B ENCH Descriptive Statistics . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 19 • D. Extended Results . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 22 • E. Evaluating S YMBAL in the Wild . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . 24

A. Implementation Details for S YMBAL S YMBAL decomposes the systematic misalignment detection task into two stages; here, we provide extended implementation details for each of these stages. A.1. Implementation Details for S YMBAL Stage 1 Subtask 1: Grouping semantically-similar facts. We express each text sample Ti as a collection of textual facts Ti = {ti1 , ti2 , ..., tini } by splitting captions at the sentence-level. We opt to use sentence-level splitting in this work because each sentence in a long-form caption typically captures a semantically-meaningful, self-contained fact. Sentence-level splitting has been utilized in prior literature (e.g. (Zhang et al., 2022)). We note here that there may be settings where this strategy is sub-optimal, such as when a sentence does not represent a self-contained fact and instead relies on previous context. In such cases, users of S YMBAL can easily adjust this design choice by modifying the definition of “textual fact” to cover relevant context. SN After aggregating all textual facts in D forming the set i=1 Ti , we encode each fact using a text embedding model. For natural image datasets in S YMBAL B ENCH derived from COCO, we consider two options for text embedding models: OpenCLIP-ViT-H-14-quickgelu (Ilharco et al., 2021) and Qwen3-Embedding-8B (Zhang et al., 2025). For medical image datasets in S YMBAL B ENCH derived from MIMIC-CXR, we consider three options for text embedding models: OpenCLIPViT-H-14-quickgelu (Ilharco et al., 2021), XRayCLIP-ViT-L (Chen et al., 2024c), and MedSigLIP (Sellergren et al., 2025). Of these, XrayCLIP-ViT-L and MedSigLIP are trained on radiology datasets. Embeddings are then clustered using spherical K-Means (implemented in Faiss (Johnson et al., 2019b)), where we sweep across a range of potential cluster numbers and select the optimal number of clusters using Silhouette distance; this approach is motivated by prior work (Sohoni et al., 2020; Varma et al., 2025). Subtask 2: Scoring groups by degree of misalignment. We score each cluster by computing the average degree of alignment between constituent textual facts and paired images. We consider three possible scoring mechanisms, explained in detail below: • Embedding scorer: Given a textual fact and its paired image, the embedding scorer utilizes an off-the-shelf vision-language model to compute embeddings for the text and image modalities. Alignment is measured by computing cosine similarity. This method is motivated by metrics like CLIPScore (Hessel et al., 2021), which have shown strong correlation with human judgments when measuring caption quality. For natural image datasets in S YMBAL B ENCH derived from COCO, we implement the embedding scorer with OpenCLIP-ViT-H-14-quickgelu (Ilharco et al., 2021) as the vision-language model. For medical image datasets in S YMBAL B ENCH derived from MIMIC-CXR, we consider three options for the embedding scorer: OpenCLIP-ViT-H-14-quickgelu (Ilharco et al., 2021), XRayCLIP-ViT-L (Chen et al., 2024c), and MedSigLIP (Sellergren et al., 2025). We note here that we do not alter the embedding scorer for reference-based settings; reference captions Ri in our benchmark often have substantially more information than the single textual fact tik ∈ Ti , and this information imbalance is challenging to capture with embedding scorers. • Text-only scorer: Given a textual fact and its paired image, the text-only scorer first generates a caption for the image and then prompts an LLM to determine if the textual fact is accurate with respect to the caption. For natural image datasets in S YMBAL B ENCH derived from COCO, we implement the text-only scorer using Llama-3.2-11B-Vision-Instruct 15

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

(Grattafiori et al., 2024) to generate captions and Qwen2.5-VL-72B-Instruct (Qwen et al., 2025) to perform scoring. For medical image datasets in S YMBAL B ENCH derived from MIMIC-CXR, we implement the text-only scorer using Maira-2 (Bannur et al., 2024) to generate captions and Qwen2.5-VL-72B-Instruct (Qwen et al., 2025) or MedGemma-27B (Sellergren et al., 2025) to perform scoring. In the reference-based setting, we use the ground-truth caption Ri rather than generating captions. We use the following input prompt in order to perform scoring: Text-Only Scorer Input Prompt You are provided with two image captions below, denoted as [A] and [B]. [A]: <generated image caption or ground-truth reference caption> [B]: <candidate textual fact> Assume that [A] is the ground-truth caption. Is the content of [B] factually accurate with respect to [A]? Rules: 1. [B] may omit details from [A]; omission is acceptable. 2. If [B] introduces any incorrect or contradictory detail, it is inaccurate. Please output your answer as a single digit, where 1 indicates that [B] is accurate and 0 indicates that [B] is not accurate. Do not provide anything other than the digit in your response. • Vision-language scorer: Given a textual fact and its paired image, the vision-language scorer provides an MLLM with both the image and the textual fact as input; the MLLM is then tasked with determining if the textual fact is accurate. For natural image datasets in S YMBAL B ENCH derived from COCO, we utilize Qwen2.5-VL-72B-Instruct (Qwen et al., 2025) as the MLLM. For medical image datasets in S YMBAL B ENCH derived from MIMIC-CXR, we utilize MedGemma-27B (Sellergren et al., 2025) as the MLLM. We use the following input prompt in the reference-free setting: Vision-Language Scorer Input Prompt (Reference-Free) <image> You are given an image. Below, a caption for the image is provided: Caption: <candidate textual fact> Is the caption accurate with respect to the image? Please output your answer as a single digit, where 1 indicates that the caption is accurate and 0 indicates that the caption is not accurate. Do not provide anything other than the digit in your response. In the reference-based setting, we additionally provide the ground-truth reference caption to the MLLM. We use the following prompt in the reference-based setting: Vision-Language Scorer Input Prompt (Reference-Based) <image> You are provided an image as well as two image captions below, denoted as [A] and [B]. [A]: <ground-truth reference caption> [B]: <candidate textual fact> Assume that [A] is the ground-truth caption. Is the content of [B] accurate with respect to the image? Please output your answer as a single digit, where 1 indicates that the caption is accurate and 0 indicates that the caption is not accurate. Do not provide anything other than the digit in your response.

Subtask 3: Summarizing the top-ranked group. We consider the following summarization mechanism for identifying the unifying concept shared by textual facts in Ctext . • Text-only summarizer: The text-only summarizer provides an LLM with textual facts in Ctext ; the LLM is then tasked with identifying the unifying concept. For natural image datasets in S YMBAL B ENCH derived from COCO, we use Qwen2.5-VL-72B-Instruct (Qwen et al., 2025) as the LLM. For medical image datasets in S YMBAL B ENCH derived from MIMIC-CXR, we consider both Qwen2.5-VL-72B-Instruct (Qwen et al., 2025) and MedGemma-27B (Sellergren et al., 2025) as the LLM. 16

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

We use the following input prompt. Then, given the output, we prompt the same LLM to select the most frequently identified feature (or the top-k most frequently identified features) as output. Text-Only Summarizer Input Prompt Consider this image caption: “<candidate textual fact>” Identify the visual features that are present in the image. Output your answer in the following format: Answer: comma-separated list Rules: 1. Each feature should be described concisely in a single phrase. 2. Each feature must be directly visible in the image. 3. Do NOT include any text outside the identified features. 4. Do NOT explain your reasoning. 5. If no features are present, output an empty list of the form: “Answer: ”

A.2. Implementation Details for S YMBAL Stage 2 Subtask 1: Grouping semantically-similar images. For natural image datasets in S YMBAL B ENCH derived from COCO, we consider two options for image embedding models: OpenCLIP-ViT-H-14-quickgelu (Ilharco et al., 2021) and DINOv2ViT-L-14 (Oquab et al., 2024). For medical image datasets in S YMBAL B ENCH derived from MIMIC-CXR, we consider three options for image embedding models: OpenCLIP-ViT-H-14-quickgelu (Ilharco et al., 2021), XRayCLIP-ViT-L (Chen et al., 2024c), and MedSigLIP (Sellergren et al., 2025). Similar to Stage 1, embeddings are clustered using spherical K-Means, where we sweep across a range of cluster numbers and select the optimal number using Silhouette distance. Subtask 2: Scoring groups by degree of misalignment. We score each cluster by computing the mean degree of misalignment between images and paired textual facts in Ctext . We consider the same scoring mechanisms as in Stage 1. Subtask 3: Summarizing the top-ranked group. We consider two summarization mechanisms for identifying the unifying concept shared by images in Cimage , described in detail below. • Text-only summarizer: The text-only summarizer generates a caption for each image in Cimage ; then, an LLM is tasked with identifying the unifying concept. For natural image datasets in S YMBAL B ENCH derived from COCO, captions are generated using Llama-3.2-11B-Vision-Instruct (Grattafiori et al., 2024). For medical image datasets in S YMBAL B ENCH, captions are generated using MAIRA-2 (Bannur et al., 2024). In reference-based settings, we use the ground-truth reference captions rather than generating captions. We use the same prompts and models as Stage 1, Subtask 3. • Vision-language summarizer: The vision-language summarizer provides an MLLM with images in Cimage ; then, the MLLM is prompted to identify the unifying concept. For natural image datasets in S YMBAL B ENCH derived from COCO, we use Qwen2.5-VL-72B-Instruct (Qwen et al., 2025) as the MLLM. For medical image datasets in S YMBAL B ENCH derived from MIMIC-CXR, we use MedGemma-27B (Sellergren et al., 2025) as the MLLM. For reference-based settings, we also provide the ground-truth reference caption to the MLLM. We use the following input prompt. Then, given the outputs, we prompt the same MLLM to select the most frequently identified feature (or the top-k most frequently identified features) as output. Vision-Language Summarizer Input Prompt <image> Consider this image. Identify the visual features that are present in the image. Output your answer in the following format:

17

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Answer: comma-separated list Rules: 1. Each feature should be described concisely in a single phrase. 2. Each feature must be directly visible in the image. 3. Do NOT include any text outside the identified features. 4. Do NOT explain your reasoning. 5. If no features are present, output an empty list of the form: “Answer: ” 6. Include a maximum of ten features.

A.3. Extension to Multiple Systematic Misalignments Real-world datasets are likely to include multiple systematic misalignments, and S YMBAL can be trivially extended to such settings as follows. Stage 1 of S YMBAL involves predicting the erroneous textual fact t̂; here, rather than summarizing the single top ranked group of facts into a unifying concept, we can simply consider the top-k ranked groups instead. This will result in multiple predicted textual facts t̂(1) , t̂(2) , ...t̂(k) , each representing a distinct recurring textual error in the dataset. Stage 2 of S YMBAL can then be implemented as described in Section 4.2, taking into account each predicted textual fact; this will result in associated visual features v̂ (1) , v̂ (2) , ...v̂ (k) . Ultimately, at the conclusion of this procedure, S YMBAL will predict multiple systematic misalignments (t̂(i) , v̂ (i) ) where i ranges from 1 to k. In Figure 9, we empirically show that Symbal can accurately detect multiple real-world systematic misalignments in captions generated by Llava1.5-7B.

B. Implementation Details for S YMBAL B ENCH S YMBAL B ENCH is comprised of 420 evaluation settings, where 360 settings include natural image datasets derived from COCO and 60 settings include medical image datasets derived from MIMIC-CXR. Below, we provide extended implementation details for the natural image settings: 1. Obtaining a base dataset. The base vision-language datasets in the natural image domain are derived from COCO (2017 val split), which consists of photographs depicting common objects (e.g. animals, food, furniture, etc.) in natural settings. Images are paired with object-level annotations as well as five human-written captions, with each caption typically consisting of a single sentence or phrase describing salient features in the image. In order to ensure that objects are clearly visible in the image, we exclude annotations for all tiny objects, defined as objects that take up less than 5% of the area of the image. After filtering out images with no remaining object-level annotations, we are left with a base dataset consisting of 4349 images and associated captions. We then compose a new two-sentence caption for each image by randomly sampling two captions from the provided list of five captions. 2. Predefining a systematic misalignment. We then predefine a systematic misalignment consisting of a textual fact t and the associated visual feature v. We sample v from the set of 80 object categories present in the dataset. Then, we sample t from the set of 80 object categories (such that t ̸= v) utilizing three possible sampling strategies: (1) random, where t is sampled randomly, (2) popular, where t is sampled from the list of the top-ten most popular objects in the COCO training set, and (3) adversarial, where t is the object that most commonly co-occurs with v in the COCO training set. These sampling strategies are motivated by prior work (Li et al., 2023) and are meant to capture a range of possible error patterns that may emerge in real-world MLLM-generated captions. 3. Injecting the predefined systematic misalignment. We insert the erroneous textual fact t into captions in the base dataset, ensuring that an association exists between text containing t and images containing visual feature v; this procedure ensures that the misalignment is systematic. Importantly, we ensure that feature t is not already in the image-caption pair prior to injection. We consider three levels of association, as measured by Cramer’s V: low association (Cramer’s V = 0.3), moderate association (Cramer’s V = 0.6), and high association (Cramer’s V = 0.9). In order to format textual fact t into a sentence, we generate 50 templates using GPT-4o (OpenAI et al., 2024), select a template at random, and insert t. We repeat this injection procedure for all possible choices of t and v in order to obtain 360 evaluation settings, each consisting of an image-caption dataset and paired annotation (t,v). Below, we provide extended implementation details for the medical image settings: 18

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

1. Obtaining a base dataset. The base vision-language datasets in the medical image domain are derived from MIMICCXR (test split), which consists of chest X-rays and associated radiologist reports collected at Beth Israel Deaconess Medical Center. We preprocess the dataset by (1) removing all images with non-frontal imaging views, (2) removing all images with missing “Impressions” sections in the paired report, and (3) removing all sentences in reports without “present” disease or anatomy entities, as identified by an off-the-shelf medical entity annotation tool (Delbrouck et al., 2024). After preprocessing, we are left with a base dataset consisting of 2233 images, each paired with the “Impressions” section of the corresponding report. 2. Predefining a systematic misalignment. We sample t from a set of five disease categories selected from the commonlyused CheXpert annotation list (Irvin et al., 2019): cardiomegaly, pneumothorax, atelectasis, pleural effusion, and edema. We sample v from a set of five medical devices: pacemaker, chest tube, endotracheal tube, surgical clips, sternotomy wires. We select these options for t and v since medical devices often co-occur with diseases, yet there is no deterministic, universal link. Models often learn spurious associations between devices and diseases as documented in prior work (Oakden-Rayner et al., 2020), meaning that such errors are highly plausible in MLLM-generated reports. 3. Injecting the predefined systematic misalignment. We insert the erroneous textual fact t into reports in the base dataset, using Cramer’s V to control the level of association with visual feature v. We use a combination of physician annotations, automated annotations from the CheXpert labeler (Irvin et al., 2019), and automated annotations from RadGraph-XL (Delbrouck et al., 2024) in order to identify whether or not t and v are present in the image-report pair prior to injection. In order to format textual fact t into a sentence, we identify the 50 most frequently occurring sentences in the MIMIC-CXR training set that discuss the presence of t and select a sentence from this list at random. We repeat this injection procedure for all possible choices of t and v in order to obtain 60 evaluation settings, each consisting of an image-caption dataset and paired annotation (t,v). In reference-based settings, we also include a ground-truth caption Ri along with each image-text pair (Vi , Ti ) ∈ D. For natural image datasets derived from COCO, Ri takes the form of a three-sentence caption combining the three human-written captions not originally selected as part of Ti . For medical image datasets derived from MIMIC-CXR, Ri takes the form of the “Findings” and “Impressions” sections of the original physician-written radiology report. We emphasize that Ti may contain errors as a result of the error-injection procedure detailed above; however, Ri is always accurate. We determine if predictions are equivalent to the ground-truth by leveraging LLM-as-a-Judge. We use Llama3.3-70B in all experiments as the LLM, leveraging the ollama implementation with default parameters. The input prompt is: LLM-as-a-Judge Evaluation Prompt You are given two short text phrases. Model response: <predicted textual error or predicted visual feature> Ground truth: <ground-truth textual error or ground-truth visual feature> Your task is to determine if both phrases refer to the same visual feature. Please output 1 if both the model response and the correct answer refer to the same feature or 0 if the model response and the correct answer do not refer to the same feature. Do not provide anything other than the number in your response.

C. S YMBAL B ENCH Descriptive Statistics In this section, we provide descriptive statistics summarizing the composition of S YMBAL B ENCH. S YMBAL B ENCH includes 420 settings covering two domains (with 360 natural image settings and 60 medical image settings). In Table 3, we provide a list of all ground-truth systematic misalignments (t, v) included in S YMBAL B ENCH. In Figure 6, we summarize S YMBAL B ENCH with histograms detailing (1) the size of each dataset, (2) the strength of the injected systematic misalignment in each dataset as measured with Cramer’s V, (3) the proportion of image-text pairs in each dataset containing the injected textual error t, and (4) the proportion of image-text pairs in each dataset containing the visual feature v. In Figure 7, we provide additional descriptive statistics on the natural image subset of S YMBAL B ENCH consisting of datasets derived from COCO; here, we provide histograms detailing (1) the mean size of the visual feature in each dataset (measured as the proportion of the total image area) and (2) the category of systematic misalignment (random, popular, or adversarial) as discussed in Appendix Section B. 19

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Strength of Systematic Misalignment Number of SymbalBench Datasets

Number of SymbalBench Datasets

Dataset Size

Frequency of Injected Textual Error (t)

Frequency of Visual Feature (v) Number of SymbalBench Datasets

Association Strength (Measured with Cramer’s V)

Number of SymbalBench Datasets

Number of Image-Text Pairs Per Dataset

Proportion of Dataset with Injected Textual Error (t)

Proportion of Dataset with Visual Feature (v)

Figure 6. Here, we provide histograms summarizing the composition of datasets included in S YMBAL B ENCH.

Systematic Misalignment Category

Number of SymbalBench Datasets

Number of SymbalBench Datasets

Visual Feature Size

Random Mean Visual Feature Size Per Dataset (measured as proportion of total image area)

Adversarial

Popular

Sampling Approach for Predefined Systematic Misalignment

Figure 7. We provide additional descriptive statistics summarizing the composition of the 360 natural image datasets in S YMBAL B ENCH. We note here that if multiple sampling strategies yield the same predefined systematic misalignment, more than one category will be assigned to the same dataset; thus, the total count for the systematic misalignment category histogram may exceed 360.

20

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Table 3. Here, we provide a list of all ground-truth systematic misalignments (t, v) included in S YMBAL B ENCH. Erroneous Textual Fact t

Visual Feature v

Erroneous Textual Fact t

Visual Feature v

Erroneous Textual Fact t

Visual Feature v

surfboard person kite person hot dog person truck toilet pizza car dining table frisbee dining table person person car person bowl microwave person laptop bowl airplane car person banana mouse hair drier giraffe laptop person person microwave dining table orange bottle bowl carrot bottle cup dining table spoon baseball bat giraffe edema atelectasis pleural effusion cardiomegaly pneumothorax atelectasis edema pleural effusion

airplane banana bed bench bicycle bird boat book bottle bowl broccoli bus cake car cat chair couch cow cup dining table dog elephant fire hydrant fork giraffe horse keyboard motorcycle oven person pizza potted plant refrigerator sandwich sheep sink surfboard teddy bear toilet train truck tv umbrella zebra chest tube chest tube endotracheal tube endotracheal tube pacemaker sternotomy wires sternotomy wires surgical clips

person chair person handbag person wine glass person cup person dining table car person chair car airplane bottle person person book chair boat dining table sandwich cup cup zebra person book sink car cell phone book oven dining table cat fork airplane bowl car person refrigerator chair person book pleural effusion cardiomegaly atelectasis edema atelectasis pneumothorax pleural effusion atelectasis

airplane banana bed bench bicycle bird boat book bottle bowl broccoli bus cake cat chair couch cow cup dining table dog elephant fire hydrant fork giraffe horse keyboard laptop motorcycle oven person pizza potted plant refrigerator sheep sink suitcase surfboard teddy bear toilet train truck tv umbrella zebra chest tube chest tube endotracheal tube pacemaker pacemaker sternotomy wires sternotomy wires surgical clips

bottle car chair oven truck book bicycle person elephant cat handbag bicycle fork umbrella person baseball glove cake bottle apple person person car dining table umbrella person truck bottle person cup dining table airplane dining table stop sign person car person person person sink truck person car tv cardiomegaly pneumothorax edema pneumothorax pleural effusion cardiomegaly cardiomegaly edema pneumothorax

airplane banana bed bench bicycle bird boat book bowl broccoli bus cake car cat chair couch cow cup dining table dog elephant fire hydrant fork giraffe horse keyboard laptop motorcycle oven pizza potted plant refrigerator sandwich sheep sink suitcase surfboard teddy bear toilet train truck tv zebra surgical clips chest tube endotracheal tube endotracheal tube pacemaker pacemaker sternotomy wires surgical clips surgical clips

21

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

D. Extended Results In Table 4, we provide an extended version of Table 1, extending to the top-ten compositions. Note that Table 4 excludes compositions consisting of an embedding-based alignment scorer and text-only summarizer, as this combination does not make use of reference captions in the reference-based setting. In Table 5, we provide an extended version of Table 2, extending to the top-ten compositions. Again, Table 5 only includes compositions that can support both S YMBAL B ENCH variants. In Table 6, we provide a tabular version of Figure 3 stratified by domain. In Figure 8, we extend Figure 4 by providing a breakdown of S YMBAL performance across various categories of systematic misalignments in the natural image subset of S YMBAL B ENCH.

Systematic Misalignment Category SymbalBench Performance

100 80 60 40 20 0

Adversarial

Popular

Random

Sampling Approach for Predefined Systematic Misalignment

Figure 8. We provide a breakdown of S YMBAL performance across various categories of systematic misalignments in the natural image subset of S YMBAL B ENCH.

We use the following input prompt for our direct-prompting baselines: Direct-Prompting Baseline Input Prompt You are provided with a dataset, where each sample consists of the following two components: Reference caption: A ground-truth caption describing the content of an image Model-generated caption: A caption generated by an AI model The model-generated captions may have systematic errors, where a recurring textual error is closely associated with the presence of a specific visual feature in the paired image. Your task is to identify the recurring textual error and the associated visual feature. Output your answer in the following format, where each comma-separated list consists of your top-five predictions in order: Textual Error: comma-separated list Visual Feature: comma-separated list Rules: 1. Each visual feature must be directly visible in the image. 2. Do NOT include any text outside of the answer. 3. Do NOT explain your reasoning. Dataset: <samples from dataset with images expressed in text-form>

22

Symbal: Detecting Systematic Misalignments in Model-Generated Captions Table 4. We evaluate various text embedding models, alignment scorers, and summarizers on the performance of Stage 1 of S YMBAL.

Text-Only (Qwen-72B) Text-Only (Qwen-72B) Text-Only (Qwen-72B) Text-Only (Qwen-72B)

92.8 92.8 82.8 64.2

94.2 93.9 85.0 67.2

80.8 86.1 81.9 67.5

82.8 87.8 83.9 71.4

Text-Only (MedGemma-27B) Text-Only (Qwen-72B) Text-Only (MedGemma-27B) Text-Only (MedGemma-27B) Text-Only (MedGemma-27B) Text-Only (Qwen-72B) Text-Only (MedGemma-27B) Text-Only (Qwen-72B) Text-Only (Qwen-72B) Text-Only (MedGemma-27B)

51.7 51.7 26.7 30.0 26.7 28.3 28.3 36.7 36.7 16.7

75.0 73.3 58.3 53.3 48.3 46.7 46.7 45.0 43.3 35.0

88.3 100.0 90.0 83.3 85.0 98.3 88.3 98.3 98.3 86.7

95.0 100.0 93.3 100.0 90.0 98.3 98.3 100.0 100.0 98.3

Summarizer

Natural

Reference-Based Acc@1 Acc@5

Alignment Scorer

Qwen3-8B OpenCLIP Qwen3-8B OpenCLIP

Vision-Language (Qwen-72B) Vision-Language (Qwen-72B) Text-Only (Qwen-72B) Text-Only (Qwen-72B)

Medical

Reference-Free Acc@1 Acc@5

Text Embedding

XRayCLIP XRayCLIP XRayCLIP MedSigLIP XRayCLIP XRayCLIP OpenCLIP OpenCLIP MedSigLIP MedSigLIP

Text-Only (MedGemma-27B) Text-Only (MedGemma-27B) Text-Only (Qwen-72B) Text-Only (MedGemma-27B) Vision-Language (MedGemma-27B) Text-Only (Qwen-72B) Text-Only (MedGemma-27B) Text-Only (MedGemma-27B) Text-Only (MedGemma-27B) Text-Only (Qwen-72B)

Table 5. We evaluate various image embedding models, alignment scorers, and summarizers on the performance of Stage 2 of S YMBAL.

Text-Only (Qwen-72B) Vision-Language (Qwen-72B) Text-Only (Qwen-72B) Vision-Language (Qwen-72B) Text-Only (Qwen-72B) Text-Only (Qwen-72B) Text-Only (Qwen-72B) Vision-Language (Qwen-72B) Vision-Language (Qwen-72B) Vision-Language (Qwen-72B)

49.7 48.1 47.8 45.8 45.3 43.1 48.1 44.2 43.6 43.6

69.7 63.9 62.8 62.5 61.4 60.8 60.6 60.3 59.7 59.4

41.9 42.5 43.9 38.9 38.6 41.1 45.6 43.9 39.7 39.7

52.2 55.6 55.8 52.2 54.7 56.4 58.1 56.7 54.2 53.3

Vision-Language (MedGemma-27B) Vision-Language (MedGemma-27B) Vision-Language (MedGemma-27B) Vision-Language (MedGemma-27B) Vision-Language (MedGemma-27B) Vision-Language (MedGemma-27B) Vision-Language (MedGemma-27B) Vision-Language (MedGemma-27B) Text-Only (Qwen-72B) Text-Only (Qwen-72B)

11.7 11.7 13.3 10.0 6.7 8.3 10.0 3.3 15.0 13.3

36.7 31.7 28.3 28.3 28.3 26.7 25.0 25.0 25.0 23.3

28.3 25.0 20.0 33.3 43.3 43.3 23.3 30.0 15.0 16.7

53.3 46.7 46.7 60.0 65.0 65.0 63.3 61.7 40.0 48.3

Summarizer

Natural

Reference-Based Acc@1 Acc@5

Alignment Scorer

OpenCLIP OpenCLIP OpenCLIP OpenCLIP DINOv2 DINOv2 OpenCLIP OpenCLIP DINOv2 DINOv2

Vision-Language (Qwen-72B) Embedding (OpenCLIP) Embedding (OpenCLIP) Vision-Language (Qwen-72B) Vision-Language (Qwen-72B) Text-Only (Qwen-72B) Text-Only (Qwen-72B) Text-Only (Qwen-72B) Text-Only (Qwen-72B) Embedding (OpenCLIP)

Medical

Reference-Free Acc@1 Acc@5

Img Embedding

XRayCLIP MedSigLIP OpenCLIP MedSigLIP XRayCLIP MedSigLIP OpenCLIP OpenCLIP MedSigLIP OpenCLIP

Embedding (MedSigLIP) Embedding (MedSigLIP) Embedding (MedSigLIP) Embedding (XRayCLIP) Vision-Language (MedGemma-27B) Text-Only (MedGemma-27B) Text-Only (MedGemma-27B) Text-Only (Qwen-72B) Embedding (MedSigLIP) Embedding (MedSigLIP)

Ablation study. We now ablate the role of the grouping step across the subset of 360 natural image datasets in our benchmark. We compare S YMBAL to a version that omits grouping: we use the best performing scorer (vision-language scorer with Qwen-72B) in order to flag each individual sentence as valid (1) or misaligned (0), and we then use our best performing summarizer (text-only summarizer with Qwen-72B) in order to identify the unifying concept across the sentences marked as misaligned. All other settings (e.g. prompts, compute budget, model configurations, etc.) are kept identical to those used for S YMBAL. For Stage 1, in the reference-free setting, we observe an Acc@1 of 41.9 and an Acc@5 of 65.3; these metrics represent a substantial decrease from the results obtained with S YMBAL (Acc@1 = 92.8 and Acc@5 = 94.2) in Table 1. We then use the best performing summarizer to identify image features associated with the misaligned sentences. For Stage 2, in the reference-free setting, we observe an Acc@1 of just 3.6 and an Acc@5 of 16.9; again, these are a substantial decrease from the results obtained with S YMBAL (Acc@1 = 49.7 and Acc@5 = 69.7) in Table 2. These results demonstrate the importance of our multi-step, structured approach for addressing the systematic misalignment detection task. 23

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Reference-Free Acc@1 Acc@5

Reference-Based Acc@1 Acc@5

Natural

Llama3.3 70B Qwen2.5-VL 72B GPT-OSS 120B S YMBAL (Ours)

0.3 0.0 9.2 49.2

0.3 1.9 13.9 69.7

0.6 0.6 10.8 41.1

1.4 1.1 17.2 51.9

Medical

Table 6. End-to-end performance across S YMBAL B ENCH, stratified by domain.

Llama3.3 70B MedGemma 27B Qwen2.5-VL 72B GPT-OSS 120B S YMBAL (Ours)

0.0 0.0 3.3 1.7 6.7

8.3 1.7 5.0 21.7 28.3

0.0 0.0 0.0 0.0 25.0

5.0 0.0 1.7 11.7 48.3

Method

E. Evaluating S YMBAL in the Wild In this section, we further demonstrate the utility of S YMBAL by supplementing our evaluations on S YMBAL B ENCH with additional quantitative and qualitative analyses in real-world settings. S YMBAL can accurately surface systematic misalignments in captions generated by off-the-shelf MLLMs. Below, we list several examples of systematic misalignments identified by S YMBAL, and we also provide associated validation: • Example 1: In captions generated by Llava1.5-7B, S YMBAL detects that erroneous references to a TV (t̂) in captions are often systematically associated with the presence of a desk, computer monitor, and/or keyboard (v̂) in the scene. We provide visual examples of image-caption pairs with the S YMBAL-identified systematic misalignment in Figure 9 (Row 1). Quantitatively, our analysis finds that erroneous references to a TV in model-generated captions are indeed 13.5 times more likely when a desk is present in the image compared to when a desk is absent, validating the S YMBAL prediction. • Example 2: In captions generated by Llava1.5-7B, S YMBAL detects that erroneous references to a handbag or a handbag on the ground (t̂) in captions are often systematically associated with the presence of a bus (v̂) in a scene. We provide visual examples of image-caption pairs with the S YMBAL-identified systematic misalignment in Figure 9 (Row 2). Quantitatively, our analysis finds that erroneous references to a handbag in model-generated captions are indeed 3.1 times more likely when a bus is present in the image compared to when a bus is absent, validating the S YMBAL prediction. • Example 3: In captions generated by Llava1.5-7B, S YMBAL detects that erroneous references to a chair (t̂) in captions are often systematically associated with the presence of a television (v̂) in a scene. We provide visual examples of image-caption pairs with the S YMBAL-identified systematic misalignment in Figure 9 (Row 3). Quantitatively, our analysis finds that erroneous references to a chair in model-generated captions are indeed 3.1 times more likely when a television is present in the image compared to when a television is absent, validating the S YMBAL prediction. • Example 4: In captions generated by Llava1.5-13B, S YMBAL detects that erroneous references to a TV (t̂) in captions are often systematically associated with the presence of a computer monitor, keyboard, and/or mouse (v̂) in a scene. Interestingly, this systematic misalignment is nearly identical to one that exists in Llava1.5-7B-generated captions (see Example 1), suggesting that solely increasing the scale of the underlying MLLM is insufficient for resolving systematic misalignments. We provide visual examples of image-caption pairs with the S YMBAL-identified systematic misalignment in Figure 10 (Row 1). Quantitatively, our analysis finds that erroneous references to a TV in model-generated captions are indeed 22.2 times more likely when a computer monitor is present in the image compared to when a computer monitor is absent, validating the S YMBAL prediction. • Example 5: In captions generated by LlavaOneVision-7B, S YMBAL detects that erroneous references to text (t̂) in captions are often systematically associated with the presence of a sign (v̂) in a scene. This systematic misalignment suggests that LlavaOneVision-7B struggles with OCR capabilities, where the presence of text-based signage in an image is likely to result in errors in the generated caption. We provide visual examples of image-caption pairs with the 24

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

S YMBAL-identified systematic misalignment in Figure 10 (Row 2). Quantitatively, our analysis finds that erroneous references to text in model-generated captions are indeed 4.6 times more likely when a sign is present in the image compared to when a sign is absent, validating the S YMBAL prediction. • Example 6: In captions generated by AyaVision-8B, S YMBAL detects that erroneous references to a vase (t̂) in captions are often systematically associated with the presence of a couch (v̂) in a scene. We provide visual examples of imagecaption pairs with the S YMBAL-identified systematic misalignment in Figure 10 (Row 3). Quantitatively, our analysis finds that erroneous references to a vase in model-generated captions are indeed 17.7 times more likely when a couch is present in the image compared to when a couch is absent, validating the S YMBAL prediction. Across all six examples of S YMBAL-identified systematic misalignments provided above, we find that erroneous references to t̂ are substantially more likely when v̂ is present in the image compared to when v̂ is absent. This analysis validates discovered misalignments by demonstrating that links between S YMBAL-identified erroneous textual fact t̂ and S YMBAL-identified visual feature v̂ do indeed exist. Our quantitative validation procedure relies on automated annotation methods in order to enable evaluation at scale; in particular, we leverage Qwen-72B in order to annotate erroneous references to t̂ in each caption. We find that these generated annotations align closely with human judgments. Given the set of 215 images in the dataset containing a “bus”, we tasked a human reader with identifying whether each Llava1.5-7B-generated caption contained an erroneous reference to a “handbag” and/or “handbag on the ground” (Example 2). Human judgments aligned perfectly with Qwen-72B predictions in 96.3% of cases (Cohen’s kappa = 0.86). S YMBAL is a powerful tool for auditing open-source vision-language datasets. Below, we list several examples of systematic misalignments identified by S YMBAL on the ShareGPT4V dataset, and we also provide associated validation: • Example 7: S YMBAL detects that erroneous references to a white tablecloth (t̂) in captions are often systematically associated with the presence of a table, cake, and/or people (v̂) in the scene. We provide visual examples of imagecaption pairs with the S YMBAL-identified systematic misalignment in Figure 11 (Row 1). Quantitatively, our analysis finds that erroneous references to a white tablecloth in model-generated captions are indeed 17.2 times more likely when a table is present in the image compared to when a table is absent, validating the S YMBAL prediction. • Example 8: S YMBAL detects that erroneous references to a printer (t̂) in captions are often systematically associated with the presence of a computer monitor (v̂) in a scene. We provide visual examples of image-caption pairs with the S YMBAL-identified systematic misalignment in Figure 11 (Row 2). Quantitatively, our analysis finds that erroneous references to a printer in model-generated captions are indeed 121 times more likely when a computer monitor is present in the image compared to when a computer monitor is absent, validating the S YMBAL prediction. • Example 9: S YMBAL detects that erroneous references to a black phone (t̂) in captions are often systematically associated with the presence of a laptop (v̂) in a scene. We provide visual examples of image-caption pairs with the S YMBAL-identified systematic misalignment in Figure 11 (Row 3). Quantitatively, our analysis finds that erroneous references to a black phone in model-generated captions are indeed 48.5 times more likely when a laptop is present in the image compared to when a laptop is absent, validating the S YMBAL prediction.

25

Llava 1.5-7B

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Images

Percentage of Captions with Erroneous Reference to TV

50

Llava1.5-7B Generated Captions

25 The image features a well-organized home office with a desk and a computer setup. On the desk, there is a laptop computer, a keyboard, and a mouse. The desk is accompanied by a chair, which is positioned in front of the desk. In the room, there are several books scattered around, with some on the floor and others on the desk. A potted plant is placed in the corner of the room, adding a touch of greenery to the space. Additionally, a TV is mounted on the wall, providing entertainment options for the home office.

The image displays a small, cluttered office space with a wooden desk. On the desk, there is a laptop computer, a keyboard, and a mouse. A monitor is also present, positioned above the laptop. A chair is placed in front of the desk, ready for someone to sit and work. The office is filled with various items, including a TV mounted on the wall, a cell phone, and a couple of books. There are also two remote controls placed on the desk. In addition to the electronics, there are two bottles and a cup, likely containing beverages for the person working in the office

The image features a wooden desk with a laptop and a desktop computer sitting on it. The laptop is placed on the left side of the desk, while the desktop computer is situated towards the right. Both computers are connected to the same keyboard, which is located in the middle of the desk. Additionally, there is a mouse on the right side of the desk, and a cell phone can be seen on the left side. A TV is mounted on the wall above the desk, providing a source of entertainment or information for the user

The image features a large computer monitor sitting on a desk, with a keyboard and mouse placed in front of it. The monitor is turned on, displaying a beautiful landscape image. The keyboard and mouse are positioned close to the monitor, indicating that they are ready for use. In addition to the main computer setup, there are several other items on the desk. A cell phone can be seen on the left side of the desk, while a book is placed towards the right side. A remote control is also present, located near the center of the desk. A TV is mounted on the wall above the desk, providing a more comprehensive entertainment setup

22.8%

1.7%

0 Images without desk Images with desk

Images

Percentage of Captions with Erroneous Reference to Handbag

50

Llava1.5-7B Generated Captions

25 The image features a pink bus parked on the side of a street, likely at a bus stop. The bus is adorned with a large advertisement on its side, possibly promoting a shopping center. Several people can be seen around the bus, with some standing closer to the bus stop and others further away. There are two handbags visible in the scene, one near the center of the image and the other closer to the right side. The presence of these handbags suggests that the people might be waiting for the bus or have just disembarked from it.

The image features a busy street with a green bus driving down the road, surrounded by various vehicles. There are two other buses in the scene, one on the left side and another on the right side of the street. A truck can also be seen on the right side of the street. Several people are present in the scene, with some standing near the buses and others walking along the street. There are two traffic lights visible in the image, one on the left side and another on the right side of the street, ensuring the smooth flow of traffic. Additionally, there are two handbags in the scene, one near the left side of the street and the other closer to the center.

The image features a group of people standing in front of a bus, likely waiting to board or disembark. There are at least nine people visible in the scene, with some standing closer to the bus and others further away. The bus is parked on the left side of the image, occupying a significant portion of the scene. In addition to the bus, there are several cars visible in the background, with one car on the left side of the image and two cars on the right side. A handbag can be seen placed on the ground near the center of the scene, possibly belonging to one of the passengers.

The image features a white and blue city bus parked on the side of a street. The bus is stopped at a bus stop, and the driver is visible through the windshield. The bus stop has a bench nearby, providing a place for passengers to wait. There are several people in the scene, with one person standing close to the bus, another person further away, and two more individuals closer to the edge of the image. A handbag can be seen placed on the ground near the first person, possibly belonging to a passenger waiting for the bus.

16.7% 5.3%

0 Images without bus Images with bus

Percentage of Captions with Erroneous Reference to Chair

Images

50

Llava1.5-7B Generated Captions

34.3%

The image features a brown dog sitting on a messy bed, surrounded by blankets and pillows. The dog appears to be looking at the camera, possibly posing for a picture. The bed is covered with various pillows and blankets, creating a cozy and comfortable environment for the dog. In the room, there is a television on the left side, and a remote control can be seen nearby. A chair is also present in the scene, positioned close to the bed.

The image features a living room with a television set placed on a wooden stand. A cat is lying on the floor, appearing to be watching the TV. Another cat is sitting on the floor, seemingly engaged with the television as well. In the room, there are two chairs, one located near the left side of the television and the other closer to the right side. A remote control can be seen on the floor, possibly used to operate the television. Additionally, there is a book placed on the right side of the room, and a vase is located near the center of the scene.

The image features a living room with a television set placed on a wooden stand. The TV is positioned in the center of the room, surrounded by various books on a bookshelf. The bookshelf is filled with numerous books, creating a cozy atmosphere. In addition to the television and bookshelf, there are two chairs in the room, one located on the left side and the other on the right side. A person can be seen in the room, standing near the left side of the television. The room also has a Christmas tree, adding a festive touch to the space.

The image features a cluttered living room with a television set placed on a stand in the center. The room is filled with various items, including a large collection of books scattered throughout the space. Some books are placed on the floor, while others are stacked on shelves or placed on surfaces. In addition to the books, there are several figurines and knick-knacks, such as a clock, a vase, and a cup, adding to the cluttered appearance of the room. A chair can be seen in the background, and a potted plant is placed near the right side of the room. The overall atmosphere of the living room is busy and filled with various items, creating a cozy yet disorganized space

25 11.1%

0 Images without TV Images with TV

Figure 9. Examples of image-caption pairs with S YMBAL-identified systematic misalignments are shown here, with the identified erroneous textual fact in each caption highlighted in red. We also quantitatively validate each identified systematic misalignment. [Row 1] S YMBAL detects that erroneous references to a TV (t̂) in captions are often systematically associated with the presence of a desk, computer monitor, and/or keyboard (v̂) in the scene. [Row 2] S YMBAL detects that erroneous references to a handbag or handbag on the ground (t̂) in captions are often systematically associated with the presence of a bus (v̂) in a scene. [Row 3] S YMBAL detects that erroneous references to a chair (t̂) in captions are often systematically associated with the presence of a television (v̂) in a scene.

26

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Percentage of Captions with Erroneous Reference to TV

Images

50

29.9%

Llava1.5-13B Generated Captions

25 The image features a wooden desk with a laptop computer sitting on top of it. A cat is comfortably resting on the desk, close to the laptop. The desk is surrounded by various items, including several books scattered around the area. In addition to the books, there are two bottles on the desk, one near the left side and the other near the right side. A potted plant is located on the left side of the desk, adding a touch of greenery to the scene. A TV can be seen in the background, mounted on the wall above the desk.

The image depicts a home office setup with a desk featuring a laptop and a desktop computer. The laptop is placed on the left side of the desk, while the desktop computer is positioned on the right side. A keyboard is situated in front of the laptop, and a mouse can be found on the left side of the desk. A chair is situated in front of the desk, ready for someone to sit and work. A TV is mounted on the wall above the desk, providing additional entertainment or information during work hours. In the room, there are also a few books scattered around, possibly for reference or leisure reading. A person can be seen in the background, possibly working or observing the office setup.

The image depicts a small, cluttered living room with various pieces of furniture and items. There is a couch situated in the background, and a chair placed in the middle of the room. A dining table is located near the left side of the room, with a laptop and a mouse placed on it. A bookshelf filled with numerous books is also present in the room, adding to the cluttered appearance. In addition to the books, there are several other items scattered throughout the room, such as a backpack, a handbag, a bottle, and a cup. A TV can be seen mounted on the wall, and a clock is hanging nearby. The room also features a window, providing natural light to the space.

The image features a small office space with a desk and a computer setup. On the desk, there is a laptop computer, a keyboard, and a mouse. A chair is positioned in front of the desk, ready for someone to sit and work. In addition to the computer setup, there are two bottles placed on the desk, one near the laptop and the other closer to the edge of the desk. A TV is mounted on the wall above the desk, and a book can be seen on the desk as well. A remote control is also present in the scene, likely for the TV.

1.3%

0 Images without computer monitor Images with computer monitor

Percentage of Captions with Erroneous Reference to Text

Images

50

LlavaOneVision-7B Generated Captions

25 The image depicts a display of doughnuts on a metal tray, placed on a wooden surface. The doughnuts are frosted with a yellow glaze and appear to be of a similar size and shape. In the foreground, there is a green sign with white text that reads "Carrot Cake" and lists additional ingredients: "Carrot, cinnamon, nutmeg, ginger, and cream cheese." The sign is attached to a metal stand with a coiled wire design. The background is slightly blurred but shows what appears to be a kitchen or bakery setting with stainless steel surfaces and a person partially visible. The lighting in the image is bright, highlighting the doughnuts and sign.

The image depicts a meal consisting of a sandwich, two containers of yogurt, and a small packet of salt. The sandwich is placed in the center of the image, wrapped in plastic wrap, and appears to be a sub-style sandwich with a light brown crust. To the left of the sandwich, there is a container of yogurt with a blue lid, featuring an image of strawberries and the text "STRAWBERRY" in white and red letters. To the right of the sandwich, there is another container of yogurt with a yellow lid, labeled "Minute Maid" and "SWEET TARTS" in black and red letters. The yogurt container also has an image of a grape. In the bottom left corner of the image, there is a small white packet of salt…

The image depicts a lively street scene with a food truck prominently featured. The truck is white with red and yellow lettering that reads "TAMALELES MI LUPITA. " Below the main text, there is a smaller sign that says "TAMALEES," and additional text that appears to be a menu or description of the food offered. The truck has a window for service, and there is a logo or emblem above the window, which includes a depiction of a tamale and a corn husk. In front of the truck, there is a wooden crate filled with green limes, and a person is seen riding a bicycle past the truck. The cyclist is wearing a helmet and a backpack, and the bicycle has a basket on the front…

The image depicts a tall, white baseball bat leaning against a brick building. In the foreground, there is a round object with a yellow sticky note attached to it. The note has handwritten text that reads "HATE NO FEAR." The sky is partly cloudy, and there are trees with bare branches visible in the upper left corner of the image. In the background, there are parked cars and a street lamp. The overall scene appears to be set in an urban environment, possibly a street or a public area near a sports facility or a baseballthemed attraction.

15.6%

3.4%

0 Images without signs Images with signs

Percentage of Captions with Erroneous Reference to Vase

Aya-Vision-8B Generated Captions

Images

50

The image depicts a cozy living room with warm, inviting decor. The walls are painted a soft yellow, complementing the beige carpet that covers the floor. A plush, light-colored sofa is positioned against one wall, adorned with a few throw pillows, including a red accent pillow. A recliner sits nearby, offering a comfortable spot to relax, with a black cat perched on its armrest. In front of the sofa, a small wooden coffee table holds a decorative vase and a few items. Across from the sofa, a television is mounted on the wall above a wooden entertainment console. Two floor lamps with warm-toned shades flank the television, providing additional lighting…

This cozy living room is bathed in warm, ambient lighting, creating a welcoming atmosphere. A plush, light-colored sofa sits centrally, adorned with soft throw pillows and cushions, inviting relaxation. Across from the sofa, a wooden coffee table with a circular glass top holds a few items, including a salt shaker and a small vase with dried flowers. To the left, a side table supports a lamp with a warm-toned shade, casting a gentle glow. A potted plant adds a touch of greenery to the room, while a rug with a bold geometric pattern covers part of the hardwood floor. In the background, a window with sheer curtains allows natural light to filter in…

This cozy living room features warm, inviting decor with a touch of elegance. The space is illuminated by soft lighting, creating a relaxing atmosphere. A dark brown sofa with striped cushions sits against the far wall, accompanied by a matching armchair positioned near a small side table. A coffee table sits in the center, adorned with a vase of red roses. The room is decorated with a mix of traditional and modern elements, including a fireplace with a mantelpiece and a wall-mounted lamp. Large windows frame the room, allowing natural light to filter in during the day and soft glow from outside lights at night…

This image showcases a modern and minimalist living room designed with a clean, white aesthetic. The space features a large window with sheer white curtains, allowing natural light to flood the room. In the center, there is a light-colored sectional sofa accompanied by a cream-colored loveseat and a red armchair, creating a cozy yet stylish seating area. A low, rectangular coffee table sits in front of the sofa, adorned with a vase of flowers. Across from the seating area, a flat-screen television is mounted on the wall, positioned to face the seating arrangement. The room has a sleek, open layout with a hardwood floor and white walls, enhancing the sense of spaciousness…

25 14.0%

0

0.8%

Images without couch Images with couch

Figure 10. Examples of image-caption pairs with S YMBAL-identified systematic misalignments are shown here, with the identified erroneous textual fact in each caption highlighted in red. We also quantitatively validate each identified systematic misalignment. [Row 1] S YMBAL detects that erroneous references to a TV (t̂) in Llava1.5-13B-generated captions are often systematically associated with the presence of a computer monitor, keyboard, and/or mouse (v̂) in the scene. [Row 2] S YMBAL detects that erroneous references to text (t̂) in LlavaOneVision-7B-generated captions are often systematically associated with the presence of a sign (v̂) in a scene. [Row 3] S YMBAL detects that erroneous references to a vase (t̂) in AyaVision-8B-generated captions are often systematically associated with the presence of a couch (v̂) in a scene.

27

Symbal: Detecting Systematic Misalignments in Model-Generated Captions

Percentage of Captions with Erroneous Reference to White Tablecloth

Images

50

ShareGPT4V Captions

25 In the heart of a cozy kitchen, a woman and a man are sharing a moment of celebration. The woman, dressed in a vibrant blue shirt, is seated on the left side of the table. She's holding a rectangular cake, its surface adorned with lit candles that flicker in the soft light. Her smile is infectious, reflecting the joy of the occasion. On the right side of the table, a man in a gray shirt is seated. His gaze is directed towards the woman, perhaps sharing in her happiness or waiting for his turn to blow out the candles. The table they're sitting at is draped with a pristine white tablecloth, adding to the festive atmosphere…

In the heart of a cozy room, a group of people are gathered around a table, engrossed in conversation. The table, draped in a pristine white tablecloth, is adorned with plates, cups, and utensils, ready for a meal. A cake, the centerpiece of the gathering, sits in the middle of the table, inviting the guests to partake in its sweet delight.The room itself exudes a warm and inviting atmosphere. A bookshelf stands in the background, filled with various books that hint at the intellectual pursuits of the inhabitants. A window punctuates the wall, allowing natural light to filter into the room and illuminate the scene…

In the heart of a bustling restaurant, a group of children are gathered around a table, their faces alight with anticipation. The table, draped in a pristine white tablecloth, serves as the centerpiece of their gathering. On it, a plate of food awaits to be savored, while a jar of condiments stands by, ready to enhance the flavors of their meal. The children, dressed in casual attire, are engrossed in their own world, their attention focused on the plate of food. Their expressions are hidden from view, adding an air of mystery to the scene. In the background, the restaurant continues its lively rhythm…

In the heart of a cozy living room, a family of five is gathered around a wooden table, engrossed in the simple joy of a birthday celebration. The table, draped in a pristine white tablecloth, serves as the centerpiece of their gathering. On the table, a vibrant birthday cake steals the show. It's a feast for the eyes with its red, white, and blue colors. The cake is adorned with candles, their flames flickering in the soft light, casting a warm glow on the faces of the family. A woman, presumably the birthday celebrant, is in the midst of cutting the cake. Her hands are steady, her focus unwavering as she prepares to serve the first slice…

13.8%

0

0.8%

Images without table Images with table

Percentage of Captions with Erroneous Reference to Printer

ShareGPT4V Captions

Images

50

The image captures a scene of a busy workspace, brimming with various objects. Dominating the scene is a wooden desk, its surface a testament to a mind at work. Two computer monitors stand side by side, their screens glowing with unseen data. A keyboard and mouse lie in front of them, tools of the trade for the digital age. To the left of the monitors, a phone rests, silent for now but ready to connect at a moment's notice. On the right side of the desk, a printer waits patiently for its next task. Scattered around the desk are various office supplies pens, pencils, and paper clips each with their own role in the symphony of work. Above the desk, a shelf holds an array of books and binders, a testament to knowledge…

The image captures a scene of a home office setup. Dominating the scene is a wooden desk, bathed in the soft glow of a lamp positioned on the left side. The desk is a hub of activity, hosting a variety of objects. A computer monitor stands as the centerpiece, its screen alive with the vibrant hues of a green and black screensaver. To the right of the monitor, a phone lies idle, its cord trailing off the edge of the desk. A keyboard and mouse sit in front of the monitor, ready to spring into action at a moment's notice. Nearby, a printer waits patiently for its next task, while a stack of books suggests a thirst for knowledge or perhaps a love for reading…

The image captures a scene of a home office setup. Dominating the scene is a wooden desk, standing against a green wall. The desk is a hub of activity, hosting a variety of objects. On the left side of the desk, an open laptop sits, its screen glowing with unseen data. Adjacent to it, a black monitor stands tall, its screen blank. A black keyboard lies in front of the monitor, ready to translate thoughts into words. A black mouse sits next to the keyboard, poised to navigate the digital world. To the right of the monitor, a black printer waits patiently for its next task. A black phone rests next to it, silent but everconnected. A black lamp stands guard next to the phone, ready to bathe the workspace in light when night falls…

The image captures a wellorganized workspace, bathed in the soft glow of ambient light. Dominating the scene is a black desk, its surface a tableau of productivity. Two computer monitors stand side by side, their screens alive with data and information. A keyboard and mouse lie in front of them, tools of the trade for the digital age. To the right of the monitors, a printer sits quietly, ready to transform digital documents into physical copies. Nearby, a phone rests, silent for now but capable of connecting this workspace to the outside world. A bookshelf stands guard in the background, its shelves filled with knowledge and resources. In front of the desk, a chair waits patiently for its occupant.

25

7.5%

0

0.1% Images without computer monitor Images with computer monitor

Percentage of Captions with Erroneous Reference to Black Phone

ShareGPT4V Captions

Images

50

The image captures a scene of a workspace, bathed in the soft glow of a desk lamp. Dominating the scene is a wooden desk, its grainy texture adding a touch of warmth to the setting. On the left side of the desk, a laptop sits open, its screen glowing with unseen data. To the right of the laptop, a white keyboard waits patiently for the next burst of typing. In the center of the desk, a black phone lies dormant, its screen dark. It's as if it's patiently waiting for a call or message to break its silence. To the right of the phone, a black mouse sits idle, its cord trailing off the edge of the desk. On the left side of the desk, a plant adds a touch of greenery to the scene. Its leaves are lush and full, suggesting it's well cared for…

The image captures a scene of a home office setup. Dominating the scene is a wooden desk, standing against a green wall. The desk is a hub of activity, hosting a variety of objects. On the left side of the desk, an open laptop sits, its screen glowing with unseen data. Adjacent to it, a black monitor stands tall, its screen blank. A black keyboard lies in front of the monitor, ready to translate thoughts into words. A black mouse sits next to the keyboard, poised to navigate the digital world. To the right of the monitor, a black printer waits patiently for its next task. A black phone rests next to it, silent but everconnected. A black lamp stands guard next to the phone, ready to bathe the workspace in light when night falls…

25 The image captures a wellorganized workspace, bathed in the soft glow of natural light streaming in from a window in the background. Dominating the scene is a white desk, its surface a tableau of productivity. On the left side of the desk, a black laptop sits open, its screen glowing with unseen data. Adjacent to it, a white printer stands ready for tasks. The right side of the desk is a hub of activity with a white computer monitor displaying a webpage, accompanied by a white keyboard and mouse, tools of the digital age…A black phone lies nearby, silent for now but ever ready for communication…

The image captures a scene of a workspace set against a vibrant red wall. Dominating the scene is a wooden desk, its surface adorned with various objects. On the left side of the desk, a laptop sits open, its screen glowing with unseen data. Adjacent to the laptop, a black phone rests, silent and unobtrusive. A white lamp with a curved neck stands sentinel on the right side of the desk, casting a soft glow that illuminates the immediate surroundings. The desk itself is a tableau of organized chaos, with papers scattered haphazardly, each one a testament to the work that has been done or is yet to be done. In the background, a window punctuates the red wall, offering a glimpse into the world outside. The image is taken from a low angle, adding a sense of depth and perspective to the scene…

5.5%

0

0.1% Images without laptop Images with laptop

Figure 11. Examples of image-caption pairs with S YMBAL-identified systematic misalignments are shown here, with the identified erroneous textual fact in each caption highlighted in red. We also quantitatively validate each identified systematic misalignment. [Row 1] S YMBAL detects that erroneous references to a white tablecloth (t̂) in ShareGPT4V captions are often systematically associated with the presence of a table, cake, and/or people (v̂) in the scene. [Row 2] S YMBAL detects that erroneous references to a printer (t̂) in ShareGPT4V captions are often systematically associated with the presence of a computer monitor (v̂) in a scene. [Row 3] S YMBAL detects that erroneous references to a black phone (t̂) in ShareGPT4V captions are often systematically associated with the presence of a laptop (v̂) in a scene.

28

Record · ID 373421 · SHA-256 3baca769e6a798f1
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.