Conceptio › Archive › arXiv CS
arXiv CSopen access

Using OCR Heads to Verbalize Image Semantics

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

Using OCR Heads to Verbalize Image Semantics Sheridan Feucht Northeastern University

Benno Krojer Northeastern University

arXiv:2609.18823v1 [cs.CV] 16 Sep 2026

Byron C. Wallace Northeastern University

Abstract How do VLMs map from pixels to semantics? To understand this general question, we focus on a narrow one: studying how VLMs perform optical character recognition (OCR). Across four models, we identify attention heads causally necessary for OCR, and discover that these are in fact general-purpose heads that output interpretable semantic features across all image tokens. For example, pointing these heads at an image token containing the word “bike” causes Qwen3-VL-8B to output “bike,” but pointing them at a bird wing causes the model to output the token “feathers.” We collapse these heads’ attention weights into a single verbalization lens transformation that reveals interpretable semantic features in hidden states across all layers. When combined with projection to vocabulary space, we can obtain interpretable labels starting from layer 0, showing that image representations are in fact aligned with language in early layers. We find that we can also use the inverse of this transformation to edit non-word concepts, e.g., replacing a tractor with a revolver in a naturalistic image, providing causal evidence that this subspace is useful for more than just OCR. Our results are an example of how the study of specific mechanisms can shed light on broader interpretability problems.1

1. Introduction Reading is a specialized form of perception that involves mapping from very specific visual forms to semantic concepts. Because it is closely tied to language, it occupies a special place in neuroscience research, with a large body of work studying specific areas of the brain used to represent word forms [8, 9, 41]. The central point of interest is the interface between visual and linguistic representations: a question that is also a concern for AI interpretability 1 For code and interactive demos, see https://ocr.baulab.info. Correspondence: [email protected]

Sarah Wang Independent

Henry Abrahamsen Northeastern University

David Bau Northeastern University (a) Prompting Qwen3-VL-8B to “Transcribe the word in this image.”

Original generation OCR heads on with vanilla attention Word: feathers Word: bike (All heads: \\)

OCR Heads on Word: branch (All heads: ‘_’)

(b) Verbalization lens: decoding image activations at layer 0 红⾊ 头发 红⾊ red hair red

红⾊ red

红⾊ snowy red

喙 眼睛 紫⾊ 红⾊ 朦胧 birds beak eye purple red murky 喙 红⾊ beak red 红⾊ red

红⾊ 褐⾊ 灰⾊ ⽻⽑ red brown gray plumage

红⾊ ⻦类 翅膀 翅膀 birds red birds wing wing

Raw logit lens probs. (baseline)

Verbalization lens probs. (ours)

refuge birds 0.46, 0.04, 并将 ⻦类 birds and will 0.28, ⻦ 0.03, 刺 thorn 0.02, bird 0.13, bird 0.08 鹑 quail

Figure 1. We find that VLM attention heads responsible for OCR are actually generic verbalization heads that can read out semantics of image tokens that do not contain text. (a) Verbalization heads attend to text when prompted to transcribe words, but manually redirecting their attention to image tokens that do not contain text causes the model to output words corresponding to the semantics of those tokens (top 10% of Qwen3-VL-8B verbalization heads). (b) We repurpose these heads’ attention weights to create a verbalization lens that can be applied to image representations from any layer. Here, our approach reveals alignment of image representations with language space starting from layer 0, which logit lens is unable to see.

[44, 50, 58]. In this work we study how vision-language models “read” (i.e., perform optical character recognition, or OCR) [20, 29, 36], which we show can help us understand how these models map from pixels to semantics in general. We discover that attention heads responsible for OCR can also

We study the subspace of activation space read by verbalization heads across four models, and find that it contains Layer Layer a wealth of high-level semantic information about inputted images. Specifically, we combine the parameters of identiQwen3-2B Head Mean Ablations Qwen3-8B Head Mean Ablations fied verbalization heads into a single matrix that can be applied to any hidden state, following techniques described in prior work on language-only models [11–13]. When combined with logit lens [27, 44, 45] (which projects hidden activations to vocabulary space), our verbalization matrix reveals semantic information across all image tokens, as seen Percent of model heads ablated Percent of model heads ablated in Figure 1b. Unlike vanilla logit lens, our approach reveals information about image tokens starting from layer 0 of the language backbone. This means that when images are passed from the image encoder to the language decoder, they already contain information aligned with language space. In other words, our results provide evidence that image representations are immediately aligned with language representations starting from early layers [32, 62], despite the apparent modality gap claimed in prior work [25, 50, 54, 58, 59]. These verbalizations are not just correlational, but can also be used to edit latent image representations for Qwen3VL models. By adding and subtracting latent vectors within the subspace read by verbalization heads, we can edit a model’s activations to replace objects in an image, e.g., causing Qwen3-VL-2B to describe an ant as “an iPod resting on a green plant stem” (Figure 8a). This indicates that the verbalizations obtained using our approach are causally important for more than just OCR. Although reading letters appears to be a narrow task, it provides a clean way to identify more general VLM components that map from vision to language space. Obtaining supervision for OCR tasks is uniquely straightforward: while an image of a bird might be mapped to the concepts “animal,” “feathers,” or “small,” an image of the letters b-i-r-d is uniquely associated with the word “bird.” Thus, identifying attention heads in this tractable setting allows us to understand the semantic representations of VLMs more broadly. We provide code and an interactive demo of verbalization lens across models at https://ocr.baulab.info.

Qwen3-8B OCR Scores

Head Index

Head Index

Qwen3-2B OCR Scores

Layer OCR Scores Molmo2-7B

LayerScores Llava-Next-34B Head Index

Head Index

Figure 2. OCR Scores (Equation 1) for all heads in Qwen3-VL models. We prompt models to transcribe random words wi in images xi and measure P (wi ) according to logit lens at the final Ablations Qwen3-8B Headheads MeaninAblations token Qwen3-2B position forHead eachMean individual head (n=1024). Many late layers promote the word contained in an OCR dataset image. See Figure 11 forLayer scores for Molmo2-7B andLayer Llava-Next-34B.

OCR Accuracy OCR Accuracy

Head Index

Head Index OCR Accuracy

be used to verbalize semantic information across a wide range of image tokens. We first isolate a set of attention heads that are causally responsible for OCR; if we prompt Qwen3-VL-8B [1] to transcribe a word in, e.g., Figure 1a, it will correctly output “bike,” and we can observe these heads attending to the tokens that correspond to that word. However, if we intervene on these attention heads to attend to a bird wing, we find that Qwen3-VL-8B will output the toQwen3-2B OCR Scores Qwen3-8B OCR Scores ken feathers. Similarly, it will output branch if we force these heads to attend to a branch. Due to these heads’ ability to describe image tokens that do not contain text, we refer to them as general verbalization heads.

Qwen3-2B Head Mean Ablations

Qwen3-8B Head Mean Ablations

Percent of model heads ablated

Percent of model heads ablated

Percent of model heads ablated

Percent of model heads ablated

Figure 3. Top-scoring attention heads from Figure 2 are causally necessary for doing OCR. Across models, we find that heads responsible for OCR comprise 10% of total attention heads. These ablations are run for 100 images across 10 seeds for random baselines; see Figure 12 for Molmo2-7B and Llava-Next-34B.

2. Finding Verbalization Heads 2.1. Approach We first search for attention heads causally responsible for extracting text from image representations. We set up a simple OCR task by placing random English words on top of ImageNet [49] images.2 We filter words, keeping only those corresponding to a single token in the model vocabulary, then pre-fill the input with “Word:” and read out the response. Table 3 shows that all models we study can reliably complete our OCR task; see Figure 10 for dataset examples. To find attention heads that are responsible for OCR, we apply logit lens [27, 44, 45] to each head’s output and measure the probability of the desired word. Each image in our OCR dataset xi ∈ XOCR contains a word wi . Let (l,h) hxi ∈ Rdhead refer to the output of attention head h at layer l at the final token position for image xi , when the model is about to output its transcription. We project this vector to the residual stream dimension using the output projection (l,h) for that head WO ∈ Rdmodel ×dhead , and directly decode to vocabulary space to obtain P (wi ). We take the mean of this 2 We use image backgrounds because OCR performance is poor for words on blank backgrounds; see Table 3.

|XOCR |

X

(l,h) (l,h) hxi )[wi ]

L OGIT L ENS(WO

xi ∈XOCR

(1) which yields a score s(l, h) for attention head h at layer l, measuring how strongly this head outputs words wi across the OCR dataset. Logit lens [45] is defined as: L OGIT L ENS(a) = softmax(WU norm(a))

2nd OCR Head (L25, H5)

2nd OCR Head (L25, H5)

(2)

for a hidden state a ∈ Rdmodel , where “norm” refers to the final norm, and WU ∈ R|V|×dmodel is the “unembedding” matrix mapping to the token vocabulary space V. The softmax induces a distribution over tokens. The intuition behind our approach is that when the model is about to transcribe a word, some set of attention heads are responsible for “fetching” wi from the image and bringing that information into the column space of the final decoder head. Equation 1 is designed to find these heads, assuming they exist.

2.2. Results Figure 2 shows OCR scores for all attention heads in two Qwen3-VL models. Notably, heads with high scores are concentrated in late layers. To assess whether these heads are causally important for OCR, we progressively meanablate3 heads with the highest scores and measure the trend in task performance in Figure 3. Across all models, meanablating the top 10% of heads ranked by OCR scores fully degrades task performance. OCR heads are concentrated in late layers, which might be a confound, but ablating an equivalent number of random late-layer heads is less effective. This suggests that OCR heads are most responsible for the model’s performance on this task. Results are similar for Molmo2-7B and Llava-Next-34B (see Appendix B).

2.3. Head Behavior on Non-Text Although we identify these heads based on a small synthetic OCR task, they appear to be broader verbalization heads used for more general image understanding (see Figure 1). While delineating the exact purpose of these heads is beyond the scope of this work, we now briefly discuss some intuition for why these heads are able to verbalize non-textual tokens. In Figure 4a, we visualize attention patterns for topscoring Qwen3-VL-2B OCR heads on a random image. As expected, these heads attend to image tokens containing text when prompted to transcribe the text in the image. However, we also find that when the model is prompted “What 3 To “mean-ablate” [12, 60] an attention head, we replace the output of

that head across all token positions with its mean over 1000 examples from our OCR dataset (also taken across token positions).

What color is the lifesaver in this image?

1

(b) Top OCR Head (L24, H6) What is the animal in this image?

s(l, h) =

(a) Top OCR Head (L24, H6) Transcribe the text in this image.

value across images:

Top OCR Head (L24, H6)

2nd OCR Head (L25, H5)

Figure 4. Top-scoring heads for OCR attend to text—but under different prompt settings, they can also attend to animals and objects. We show the two top-OCR-scoring attention heads in Qwen3-VL-2B as an example. (a) These heads attend to the subtitle “[steamboat sound]” when prompted to transcribe text. (b) The same heads also attend to a mouse when prompted “What is the animal in this image?” and a lifesaver when prompted “What color is the lifesaver in this image?” (Qwen3-VL-2B’s answers to these questions are “Mouse” and “White,” respectively). This behavior suggests that these heads are not purely dedicated to OCR.

is the animal in this image?,” the same attention heads begin to attend to tokens corresponding to the mouse in the image, rather than the subtitles. This attention behavior is analogous to work on filter heads in LLMs [34, 52], which “retrieve” items from the prior context based on a prompt query (see Appendix C for discussion). What do verbalization heads output when they attend to non-text tokens? Figure 1a shows a qualitative example for Qwen3-VL-8B: redirecting the top 10% of verbalization heads to attend to, e.g., a token containing a bird wing causes the model to output the token feathers. This motivates a question: can we use these attention heads to create a “verbalizer” for image tokens in general? We note that it is likely one could also identify this set of attention heads using tasks other than OCR, e.g., asking “What is in this image?” with supervision from image annotations. However, OCR provides a straightforward one-toone mapping from visual inputs to language outputs. If we were to run the same procedure with natural images, a given image token of a bird might plausibly map to many different tokens, including “bird,” “feathers,” “cardinal,” “wing,” or “小鸟” (“bird” in Chinese). The OCR task offers unambiguous supervision, enabling broader understanding of how VLMs map from pixels to semantics.

3. Building a Verbalization Lens 3.1. Verbalization Transformation Instead of redirecting attention maps as in Figure 1a (which requires a separate forward pass for every image token), we compact the parameters of these verbalization heads into a

(a) Raw Logit Lens schließen close 0.06,

L0 angep 0.06, gef 0.04, clinic 0.04

(a) (b)

blue sky

signage

hillside

house

kite

建 [garbage] 0.08, L16 [garbage] 0.01

clouds 0.90, cloud 0.09, 云 cloud 0.00, .cloud 0.00

L32 ubi 0.01, dent

4 0.37, cloud 0.28, 云 cloud 0.13, clouds 0.09

(d) Raw Logit Lens

+Verbalization Heads

(e) Raw Logit Lens

+Verbalization Heads

(f) Raw Logit Lens

+Verbalization Heads

pathway 0.98, pathways 0.02, 路径 path/route 0.00, path 0.00

L0 梁 bridge 0.07, ⽽去

L16 0.04, finder 0.03,

pathway 1.00, pathways 0.02, 路径 path/route 0.00, roadway 0.00

L16 0.10, 停⻋位 parking

0.80, ways 0.18, L32 way WAY 0.01, WAYS 0.01

pathway 0.98, 路径 path/ route 0.01, pathways 0.00, path 0.00

L32 conditions 0.14, Sixth highway 0.47, road 0.04,

下乡 go to countryside

⻛筝 farmMontana ⼭坡 mountain mountain

天空 sky 0.98, sky 0.01, 蓝天 blue sky 0.01, 晴 clear 0.00

gate 0.03, oret 0.02

bell 0.03, unos 0.02,

pageTitle 0.01

+Verbalization Heads

天空 sky 0.84, 蓝天 blue sky 0.15, sky 0.01, blue 0.00

0.02

(f)

cloud 0.63, clouds 0.23, 云 cloud 0.11, cloud 0.02

rumor 0.01

fur 0.60, frei 0.02,

L32 blue 0.26, 天空 sky

旗帜 diamond plates pathway sky

L0 宾客 guests 0.04,

outdoors 0.67, ⻛景 landscape 0.15, 景⾊ view 0.09, outdoor 0.06

of department L0 head 0.02, tee 0.02, weed

天空 diamond endings

天空 sky 0.26, 蓝天 blue sky 0.12, sky 0.01, skies 0.00

scenario 0.03

flag

mountain mountain mountain

(c) Raw Logit Lens

stu 0.78, minster

L32 0.34, scene 0.13, 场景

(d) realpath 0.05, 部⻓

蓝天

+Verbalization Heads

L0 0.03, kür 0.03, 谣⾔

L16 space 0.01, 抜け

scenic 0.44, outdoors

(e)

(b) Raw Logit Lens

outdoors 0.26, geography 0.20, 地理位置 location 0.16 road 0.39, travel 0.21, roadside 0.08, ⻛景 landscape 0.05

⻛景 landscape 0.02, ⾃

L16 驾 drive oneself 0.02

(c)

+Verbalization Heads

omission 0.01

蓝天 blue sky 0.38, 藍 0.20

两端 both ends 0.15, 桥 道路 path 0.61, roadway 0.29, road 0.05, 公路 and go 0.05, 逶 0.04 highway 0.03 两条 two 0.10, 尽头 end space 0.02, 三条 three 0.02

道路 path 0.83, 的道路 the path of 0.04, road 0.04, roadway 0.04

6 0.56, 路况 road

道路 path 0.47, 公路

0.12, 6 0.09

roads 0.01

四五 four-five 0.04, 0.01, .ibm 0.01

出⼊境 entry and exit

L0 0.06, await 0.04, 晚期 late stage 0.03 ⼩镇 small town 0.05,

Wyoming 0.72, 乡村旅游 rural tourism 0.16, Yellowstone 0.10, mountain 0.01

L16 公共服务 public

tourist 0.43, 旅游 travel 0.11, travel 0.11, pilgrimage 0.11

0.92, trail L32 trails 0.05, Trail 0.04

徒步 on foot 0.86, hiking 0.12, trail 0.01, walking 0.00

services 0.03, ⾃驾 drive oneself 0.02

Figure 5. Verbalization lens reveals semantic information contained within VLM image tokens starting from early layers. We show several 蓝天 for one token readouts 旗帜 image for Qwen3-VL-8B, comparing raw logit lens to the top 10% of verbalization heads. Applied to image flag tokens that contain text,天空 verbalization lens shows tokens corresponding to the appropriate word; see subfigure (d). Applied to other image skyreveals interpretable semantic information, outputting cloud for a token containing clouds (c), or more general tokens, verbalization lens ⻛筝 a possible register token [7] in (a). Outputs are often in Chinese for Qwen3-VL models, likely because information like⼭坡outdoors above kite they are Chinese models. Gray text means probability is below 0.10; top verbalization lens predictions are bolded for emphasis. See Appendix D for qualitative results across models. blue sky

signage

diamond plates pathway

mountain mountain mountain

Montana

hillside

diamond endings

farmhouse mountain mountain

single linear transformation that can reveal semantic information contained within any hidden state. Specifically, we sum the output-value (OV) matrices [11] of verbalization heads across layers. This is roughly equivalent to intervening on attention maps, except that heads can “attend” to activations from any layer [12, 13]. Let Op be a set containing the top-p% of OCR heads; across models, we use p=10% based on Sec. 2.2. We construct a matrix Lp ∈ Rdmodel ×dmodel comprising the parameters of these verbalization heads. We calculate Lp =

X

(l,h)

WO

(l,h)

WV

,

(3)

(l,h)∈Op

3.3. Qualitative Results

based on the output and value projections of heads in (l,h) (l,h) Op , which have shape WO ∈ Rdmodel ×dhead , WV ∈ dhead ×dmodel . Multiplied together, each “OV matrix” deterR mines what this head will output if it attends to a particular token; adding them together is equivalent to summing the outputs of multiple heads. Thus, applying this matrix to a hidden state is equivalent to “forcing” all heads in Op to attend to that hidden state and summing the resulting outputs.

3.2. Decoding to Vocabulary Space After applying the verbalization transformation from Equation 3 to an activation a(l) ∈ Rdmodel , we can project the resulting vector to vocabulary space (Equation 2) to obtain probabilities for each token: V ERBAL L ENS(a(l) ) = L OGIT L ENS(Lp a(l) ).

differences: first, heads across layers can “attend” to a(l) regardless of whether they are also at layer l. This still yields interpretable results, perhaps because layer order has been shown to be mostly interchangeable in LLMs [12, 33, 56]. Second, we directly decode to vocabulary space, rather than allowing the model to further process these head outputs. This gives us the advantage of being able to run a single forward pass and decode any activation with two simple matrix multiplications (the naive approach requires a separate forward pass for each image token). We can conceptualize Lp as an operation that “focuses” logit lens on the correct subspace of model activations.

(4)

This approach can be thought of as a quick approximation of the proof-of-concept in Figure 1a. There are two key

Figure 5 shows an example of logit lens applied to image hidden states across layers with and without our verbalization transformation. Applying our verbalization heads reveals semantic information both for image tokens containing written text (e.g., pathway for a patch containing the text “PATHWAY”, Figure 5d) and image tokens containing more generic visual information (e.g., cloud for a patch containing clouds, Figure 5c). Our approach yields interpretable results as soon as layer 0, whereas raw logit lens is quite noisy in early-middle layers. This corroborates evidence from prior work showing that image representations are aligned with language in early VLM layers [32, 62].

3.4. Object Confidence To quantify how well this approach recovers relevant semantic concepts from image tokens, we set up an object detection task using the COCO validation split [37], similar to the setup from [27]. For each COCO category (e.g., “backpack”), we sample n images with that object and n random images that do not contain that object. Let o be the token

Max P(o) Across Image at Layer 0

(b)

+Verbalization Lens (top 10% heads)

Max P(o) Across Image Tokens

Max P(o) Across Image Tokens

Layerwise Prob. Diff: P(o)with −P(o)without Qwen3-VL-8B

Qwen3-VL-2B

Molmo2-O-7B

Llava-Next-34B

Mean Prob. Diff

Raw Logit Lens

Count

Molmo2-O-7B

Mean Prob. Diff

Count

Qwen3-VL-8B

(a)

Max P(o) Across Image Tokens

Verbalization Lens Raw Logit Lens Random 10% of heads All heads

Max P(o) Across Image Tokens

Figure 6. Verbalization lens is sensitive to whether an image contains an object: it yields high probabilities for images that contain objects, and low probabilities for images that do not. We apply logit lens (Equation 2) to VLM representations of images from the COCO dataset [37] with and without this transformation, and extract the probability of the COCO class label token for all single-token categories, similar to [27]. (a) At layer 0, verbalization lens is better able to detect whether an image contains an object—P (o) is consistently higher for the red histogram for Qwen3-VL-8B and Molmo2-7B. (b) Average difference in maximum P (o) across layers for images containing an FINAL VERSIN with n=512 per category (there are about 47 single tok categories) stride across layers of 4. top 10% o object vs. random images not containing that object. Verbalization lens increases the gap between the two settings more than raw logit lens, especially in early-middle layers. Using the weights of all attention heads gives a weaker effect than just verbalization heads.

for a particular category. For each image, we take the maximum P (o) under verbalization lens across all image tokens as the model’s confidence that an image contains o, focusing on one layer at a time. We filter to include only singletoken categories (56/80 COCO classes for Qwen3 models). As found in prior work [27], raw logit lens can distinguish COCO objects. But verbalization lens leads to much higher probabilities for images that contain o (Figure 6a). Verbalization lens also achieves consistent separation across layers, whereas raw logit lens only begins to work in late layers (Figure 6b). Figure 19 shows similar results with ROCAUC: raw logit lens offers some discrimination, but verbalization lens realizes near perfect AUC across layers.

Table 1. Token-level localization results for Qwen3-VL models. Here, we follow [27] by taking the maximum probability for each token across layers. We use the top 10% of verbalization heads based on results from Figure 3, and thus also sample 10% of heads for the random baseline. For mIoU, we choose the best threshold based on 128 random images and report results for a different sample of 512 images, all from the COCO 2014 validation set. See Appendix F for details.

Qwen3-VL-8B

Qwen3-VL-2B

Method

mIoU

mAP

mIoU

mAP

Raw Logit Lens Random Heads Verbalization Lens

0.161 0.167 0.190

0.374 0.302 0.407

0.223 0.157 0.249

0.462 0.224 0.463

3.5. Object Localization We find that verbalization lens generally reveals information localized to relevant regions of an image. Similar to Section 3.4, we calculate P (o) across all image tokens at a given layer, which yields a heatmap of confidence scores across image tokens. We can use this approach to obtain probabilities at a particular location for any token in the model’s vocabulary. For example, in Figure 7a, the token with the maximum P (orange) is the token containing a small orange. Localization is also sensitive to nuances in word meaning. While the token phone localizes image tokens that appear to contain a mobile or landline phone, “手机” (a Chinese word specifically meaning cell phone) localizes only the cell phone and “电话” (a word mostly used for landlines) localizes only the landline.

Following [27], we calculate segmentation metrics for our approach using the COCO 2014 validation set [37]. For a given image, we retrieve all of the single-token categories4 o in that image and use our lens to generate segmentation maps for each category in the image. To generate a segmentation map, we obtain a score P (o) for each token by taking the maximum probability of o across layers. We then binarize these probabilities to score lens predictions against ground truth masks by choosing a threshold τ on a disjoint set of images (or sweep across τ for AP), and average resulting scores across all objects/images. Unlike prior object localization work [4, 14, 27], we evaluate at the token-scale 4 This number is roughly similar across models; see Table 2.

Original Image

(b)

_phone +Verbalization Lens

(a)

_orange

_orange

Raw Logit Lens

+Verbalization Lens

⼿机 (cell phone)

电话 (telephone)

+Verbalization Lens

+Verbalization Lens

Figure 7. We use verbalization lens to measure P (o) across image tokens for an object o, and find that high probabilities are concentrated in image tokens containing o. (a) P (orange)=0.05 for a token containing an orange, whereas raw logit lens probabilities are lower and scattered across the image. (b) Localization is sensitive to the specific word used: P (phone) is high for both phones in the image, but “手 机” (a Chinese word specifically meaning cell phone) localizes only the cell phone, and “电话” (a word mostly used for landlines) localizes only the landline. Results are for Qwen3-VL-2B (layer 15), with the top 10% of verbalization heads.

instead of upsampling our predictions to the pixel scale.5 We compare to logit lens and an equivalent number of random heads. Table 1 shows that verbalization lens is slightly more spatially localized than baselines. However, numbers are overall quite low—this may in part be due to register token behavior [7, 26], where general information about an image is stored in arbitrary tokens (e.g., Figure 5f).

3.6. Comparison to LatentLens LatentLens [32] is an alternative approach to decoding image activations: instead of using unembeddings, it measures cosine similarity of image tokens against a large pool of contextual text embeddings from intermediate LLM layers. However, Krojer et al. [32] report failure modes in the early and mid-layers of some larger models such as Llava-Next34B. Adopting the same LLM judge interpretability metric as Krojer et al. [32], we find that verbalization lens produces interpretable tokens even in these difficult settings, with a 40.5 point increase in judge scores for Llava-Next-34B on average across layers. For all other models, our approach is on par with LatentLens; see Appendix G for details. 5 We define a patch as containing an object if more than 10% of its pixels correspond to the ground-truth COCO mask.

4. Editing Conceptual Information Are the representations read out by these verbalization heads a mirage, or are they causally relevant for the model’s “understanding” of an image? Instead of merely verbalizing hidden state representations, we use the subspace read by Lp to edit Qwen3-VL models’ perception of an image, finding that we can, e.g., cause the model to describe an ant in an image as an iPod.

4.1. Approach Let crem correspond to a concept we want to remove, and cadd be a concept we want to add in its place. Here, we focus on single-token concepts. If we take the unembedding vector uc ∈ Rdmodel corresponding to the token for the concept c, we can obtain a latent vector vc for that concept using the inverse of the transformation from Equation 3: vc = L−1 p uc . dmodel

(5)

where vc ∈ R is a direction in model activation space that encodes the concept c. Although Lp is typically fullrank (and thus invertible), we find that in practice it has a long tail of small singular values (e.g., Figure 14). This means that directly taking the inverse of Lp would cause an

Main Subject: The central focus is a large, healthy-looking Snowshoe cucumber. It is sitting upright on a patch of grass ... The cucumber has a robust, rounded form with a textured, mottled brownish-green skin. Its most striking features are its large, dark, expressive eyes and its prominent, upright, greenish-brown ears (leaves).

Qwen3-2B Qwen3-8B: “pineapple” ← “washer” This is a close-up, top-down photograph of a white ceramic plate with a decorative green leaf or vine pattern around its rim. ... The most striking feature is the arrangement of objects on the plate: - A large, white, plastic or ceramic washer/dishwasher part (possibly a washing machine agitator or a decorative piece) is prominently displayed in the center. It has a complex, layered, star-like or flower-like shape with multiple ridges and grooves.

Mean Diff. Prevalence

Qwen3-8B: “hare” ← “cucumber” This is a charming and whimsical photograph of a Snowshoe cucumber (Cucumis sativus) in a natural outdoor setting.

Qwen3-8B

Qwen3-2B: “ant” ← “iPod” This is a close-up, macro photograph of a single, red-brown, human-made object, specifically an iPod, resting on a green plant stem. ... The iPod is positioned diagonally across the frame, with its left side facing the viewer. It is a classic model with a glossy, reddish-brown finish. ... The iPod is attached to a green plant stem, which is visible in the foreground and background.

Mean Diff. Prevalence

(b)

(a)

Scaling Factor α

Scaling Factor α

Figure 8. We can use our verbalization transformation to edit image activations, e.g., replacing the concept of “ant” with “iPod.” (a) cremis Qualitative examples for editing objects within an image. Rank constant based on our initial sweep; we show α=4 for the first two prevalence: 0 Specificity: 4 crem prevalence: 0 Specificity: 9 crem prevalence: 0 Specificity: 9 prevalence: 10 Coherence: 10 images, and α=5 for the third. (b) LLM judge scorescaddacross random ImageNet [49] images (n=256). Verbalization lens yields cadd Quantitative prevalence: 10 Coherence: 9 cadd prevalence: 10 Coherence: 9 rank 463, sample 5 — coherence: 10 remove_prevalence: 0 add_prevalence: 10 specificity: 4 latent vectors that can effectively edit image representations without heavily damaging irrelevant concepts, or degrading caption coherence. rank 463, sample 4 — coherence: 10 remove_prevalence: rank 219, sample 4 — coherence: 9 remove_prevalence: 0 add_prevalence: 10 specificity: 9 0 add_prevalence: 10 specificity: 9 Edits are most effective with a larger scaling factor α. We use a rank for each model that explains the top ρ=60% of verbalization energy, This is a charming and whimsical photograph of a Snowshoe This is a close-up, macro photograph of a single, red-brown, humancucumber (Cucumis sativus) in a natural outdoor setting. made object, specifically an iPod, resting on a green plant stem. The based on an initial sweep over 100 images (Appendix I). image is taken in a dark environment, with the iPod and the stem being the primary subjects against a black background. The iPod is positioned diagonally across the frame, with its left side facing the viewer. It is a classic model with a glossy, reddish-brown finish. The device has a distinct design with a rounded, curved body, a small, dark, circular earpiece on the left side, and a visible earphone jack on the right. The iPod is attached to a green plant stem, which is visible in the foreground and background. The stem has a smooth, slightly glossy surface and a light green color. The background is completely black, which makes the iPod and the stem stand out prominently.

Main Subject: The central focus is a large, healthy-looking Snowshoe cucumber. It is sitting upright on a patch of grass and soil, with its body angled slightly to the left. The cucumber has a robust, rounded form with a textured, mottled brownish-green skin. Its most striking features are its large, dark, expressive eyes and its prominent, upright, greenish-brown ears (leaves). A black leash is attached to its ear, suggesting it is being walked or is a pet, which adds a humorous and anthropomorphic element to the image.

explosion of singular values in L−1 p . Instead of taking the full inverse, we can take a pseudo-inverse: The image is a detailed, high-resolution photograph, likely taken with a macro lens, capturing the intricate details of the iPod and the plant stem. The lighting is focused on the subjects, creating a strong contrast with the dark background. The image is marked with the text "UGA1366066" in the bottom right corner, which appears to be a digital watermark or identifier.

Lp = UΣV∗

In summary, the image is a detailed, high-resolution macro photograph of a red-brown iPod resting on a green plant stem, set against a black background.<|im_end|> 2B

Background: The background is softly blurred, drawing attention to the cucumber. It consists of a mix of green grass, brown soil, and scattered leaves. In the upper right background, a small, bright red bird (possibly a male Rose-breasted Grosbeak) is visible, adding a splash of color and life to the scene. Text and Context: At the bottom of the image, there is a caption that reads: "Snowshoe Cucumber (Churchill, Canada) Copyright 2009 - Monte M. Taylor." This text provides context, indicating the location (Churchill, Canada) and the photographer (Monte M. Taylor) and the year the photo was taken (2009).

(6)

−1 ∗ Overall L+ UImpression: [: k] The image is playful and endearing, (7) capturing p,k = V[:, : k]Σ a moment of quiet observation in nature. The use of a leash on the cucumber is clearly a humorous artistic choice, transforming a vegetable into a cute, walkable pet. The natural lighting and shallow depth of field create a soft, pleasant atmosphere.<|im_end| >

where 1 ≤ k ≤ dmodel is an integer hyperparameter, chosen via sweep on a held-out set. Here, we interpret bottom singular values as noise, manually setting them to zero rather than letting them affect our transformation. This approach implicitly views Lp as reading from a low-rank subspace of the residual stream, as described in previous work [21]. It also means that our edit is more surgical, as we are only affecting activations within a low-rank subspace. Once we have obtained our latent vectors vcrem , vcadd for the concepts we would like to add and remove, we edit ac(l) tivations at ∈ Rdmodel across all layers l and image token positions t. For a given activation vector, we remove the component of that activation in the direction of vcrem , and add in vcadd at the same relative magnitude: (l)

4.2. Evaluation To evaluate the efficacy of our edits, we prompt models to free generate descriptions with intervened image representations. We prompt o4-mini to score the generated captions according to four metrics on a scale from 0 to 10: • Prevalence of cadd . e.g., how prominently does the concept of ‘dog’ feature in this caption? • Prevalence of crem . e.g., how prominently does the concept of ‘motorcycle’ feature in this caption? • Specificity. Apart from any differences regarding the concepts crem and cadd , how similar is the counterfactual caption to the ground truth vanilla caption? • Coherence. In terms of writing quality, how coherent is this caption? See Appendix H for full prompts. We obtain scores for all generated and vanilla captions. To measure edit success, we measure the ∆ in prevalence scores for a counterfactual caption compared to scores for the vanilla generation. We evaluate on ImageNet [49] images as they are generally focused on one subject at a time: images are assigned crem based on their class label, sampling only from images with single-token class labels. Then, cadd is randomly sampled from all other single-token ImageNet class names.

rank 892, sample 5 — coherence: 10 remove_prevalence: 2 add_prevalence: 2 specificity: 9

This image shows a whimsical and creative cake designed to look like a vintage computer setup, presented in

λ = (at · vcrem )/||vcrem ||2

(8)

4.3. Results

(l)∗ (l) at = at − λvcrem + αλvcadd

(9)

Editing with latent vectors obtained via the verbalization head subspace is more effective than directly editing with unembeddings. After performing an initial sweep over rank and α (see Appendix I for details), we select the rank that explains 60% of the energy of Lp and run our edit with this setting for n=256 random images, with results shown in Figure 8b. We find using α > 1 is most effective. Intuitively, this means that cadd must be added more strongly than crem was present in order to have a causal effect; per-

where we allow for a scaling factor α ∈ R on the added concept cadd as an additional hyperparameter. Note that if crem is not present in this activation vector, λ will be close to zero, and Equation 9 becomes a no-op. This is by design: we do not want to edit token positions where there is nothing to replace. Thus, while we apply this edit to all image tokens, it only has an effect at relevant token positions.

“Describe this image.” (ImageNet Class: Tractor)

Qwen3-VL-8B Vanilla Generation This is a photograph of a red vintage tractor driving on a paved road, set against a backdrop of dense, leafy trees. Key Details: - The Tractor: The central subject is a red Massey Ferguson tractor, likely a model from the 1970s or 1980s. It features a prominent, enclosed cab with a black frame and glass windows. The front grille has the "MF" logo, and the model number "135" is visible on the side. The tractor has large, rugged rear tires and smaller front wheels, typical of agricultural machinery. A small orange warning light is mounted on the roof of the cab. - The Driver: A man is visible inside the cab, sitting in the driver's seat and looking towards the camera. He is wearing a dark shirt. - The License Plate: The tractor has a UK-style license plate that reads "JN1 203". - The Setting: The tractor is on a grey asphalt road. The background is filled with lush green and autumnal trees and bushes, suggesting the photo was taken in late summer or early autumn. The lighting appears to be natural daylight, possibly overcast. - Overall Impression: The image captures a classic, well-maintained piece of farm equipment in a rural setting, evoking a sense of nostalgia and the enduring presence of agricultural machinery. The red color of the tractor stands out vividly against the green and brown foliage.

Edited with Verbalization Transformation: “tractor” ← “revolver” This is a highly stylized and surreal image of a revolver that has been creatively modified to resemble a full-sized, classic car — specifically, a red and silver revolver that looks like a miniature, toy-like car with wheels and a license plate. Here’s a detailed description:

- The Central Object: The main subject is a revolver, but it has been artistically

altered to look like a car. It has a red body, a silver front bumper, and a license plate that reads “JN1 209”. The revolver’s barrel has been extended to look like a front bumper, and the cylinder is visible as part of the engine block. The revolver’s trigger guard and grip have been transformed into a car’s steering wheel and dashboard. - The Wheels: The revolver’s cylinder and the revolver’s hammer have been replaced with large, detailed wheels. The front wheel is silver and has a realistic tire tread, while the rear wheel is black and has a more rugged, off-road look. The revolver’s trigger has been replaced with a car’s gear shift. - The Driver: Inside the “car,” there is a man sitting in the driver’s seat, looking out the window. He is wearing a black shirt and has a serious expression. His hand is on the steering wheel, and he appears to be driving the car. - The Setting: The “car” is parked on a paved road, with trees and bushes in the background. The trees are green and brown, suggesting it is autumn. The sky is overcast, and the lighting is soft and diffused. - The Style: The image has a surreal, dreamlike quality. It blends the familiar with the fantastical, creating a sense of wonder and imagination. The revolver, which is typically associated with danger and violence, has been transformed into a playful, toy-like car, which is a striking contrast. Overall, this image is a creative and imaginative piece of art that plays with the viewer’s expectations and perceptions. It is a fun and whimsical image that invites the viewer to look closer and appreciate the details.

Figure 9. We can use our verbalization transformation to edit Qwen3-VL-8B image activations, replacing the concept of “tractor” with “revolver.” We obtain latent vectors for these concepts using the inverse of our verbalization matrix. After subtracting the latent vector for “tractor” and adding the vector for “revolver,” Qwen3-VL-8B describes the main subject of the image as a revolver. We leave all other (top 10%aOCR heads, rank 463, alpha=4) image features untouched. The model reconciles remaining visual features (like wheels, license plate, and a driver) by describing the image as a “surreal image of a revolver that has been creatively modified to resemble a full-sized, classic car.” Both captions are generated with greedy decoding. This generation uses α=4; across all examples, we use the top ρ=60% of Lp energy, where p=10%. rank 463, sample 4 — coherence: 9 remove_prevalence: 0 add_prevalence: 10 specificity: 2

Language and Vision. What is the alignment of language haps this is to override remaining information in the residual Editing Object Perception for Qwen3-VL-8B-Instruct and vision in models? For CLIP vision models, previstream related to crem but not captured in vcrem . proper version where we have the right rank! ous work studied “multimodal” neurons that respond to We show three abridged qualitative examples in Figboth text referring to a concept and images of that conure 8a. We can see that, e.g., the concept of “ant” has been cept [17, 40]. Several works have shown that with just cleanly replaced by the concept of “iPod,” while other via small adapter module, one can map from image to text sual features (like background descriptions) remain intact. [39, 43, 57], even with both models frozen [42, 50]. Other These edits preserve low-level visual features of the original works study vision representations through the lens of lanobject: i.e., the color and texture of the ant is now applied guage [5, 14], or differentiation between vision and lanto the iPod, with the model describing the iPod as being “a guage representations [22, 24, 47]. Others have studied how classic model with a glossy, reddish-brown finish.” language priors can conflict with visual features [23, 53], Figure 9 shows a full example edit for Qwen3-VLparticularly for OCR [20, 29, 36]. Our results are evoca8B alongside the model’s vanilla generation. Here, tive of recent work on how VLMs use language as a mecrem =“tractor,” and we randomly sample cadd =“revolver”. diator for visual understanding [48, 51, 55]. One question Our edit is successful: the model does not describe the imis whether image representations are aligned with language age as containing a tractor, but rather as an image of a rein early VLM layers: while earlier work claims that alignvolver. Interestingly, because other aspects of the image ment arises only in middle layers [25, 44, 50, 59], others asremain untouched (wheels, license plate, bumper, and man sert early-layer alignment [32, 62], which our method also inside the tractor), Qwen synthesizes these attributes into a shows. “surreal” description of a “revolver that has been creatively modified to resemble a full-sized, classic car.” Other heads. We note conceptual overlap between ver5. Related Work balization heads and two types of attention heads studied in prior work: filter heads [52] and gaze heads [15]. We comLenses. Logit lens [45] was the first to find that LLM inpare against these and find modest overlap in Appendix C, ternal hidden states can be projected to vocabulary direcsuggesting that there may be a relationship between these tions with interpretable results, and was followed by a large components. This further supports our claim that the narrow body of work on other “lenses” to decode internal states setting of OCR provides a straightforward way of identify[2, 10, 12, 13, 16, 19, 21, 30, 45, 46]. Prior work has aping heads responsible for more general VLM mechanisms. plied logit lens approaches to VLMs [27, 35, 44]; one approach [32] provides an alternative to vocabulary decoding 6. Conclusion using cached latent representations of text. Earlier, [28] also In this work, we find that attention heads responsible for showed that CLIP’s classifier head can act as a “logit lens” OCR can be used to verbalize general semantics of non-text for its intermediate states. Our work augments the interimage tokens across four VLMs. We collapse the weights pretability of these vanilla logit lens approaches.

of these attention heads into a single matrix that yields more interpretable labels than raw logit lens, and show that the subspace this matrix reads from is causally relevant for model descriptions of images. Taken together, our results show how study of narrow mechanisms can shed light on broader interpretability problems.

Acknowledgements SF thanks Si Wu for discussions on language and letters throughout this project, as well as Rohit Gandikota and Arnab Sen Sharma for feedback on early experiments. We thank Andy Arditi, Eric Todd, and Grace Proebsting for feedback on initial drafts. SF is funded by a grant from Coefficient Giving. BK and DB are funded by NSF #2403304.

References [1] Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, Wenbin Ge, Zhifang Guo, Qidong Huang, Jie Huang, Fei Huang, Binyuan Hui, Shutong Jiang, Zhaohai Li, Mingsheng Li, Mei Li, Kaixin Li, Zicheng Lin, Junyang Lin, Xuejing Liu, Jiawei Liu, Chenglong Liu, Yang Liu, Dayiheng Liu, Shixuan Liu, Dunjie Lu, Ruilin Luo, Chenxu Lv, Rui Men, Lingchen Meng, Xuancheng Ren, Xingzhang Ren, Sibo Song, Yuchong Sun, Jun Tang, Jianhong Tu, Jianqiang Wan, Peng Wang, Pengfei Wang, Qiuyue Wang, Yuxuan Wang, Tianbao Xie, Yiheng Xu, Haiyang Xu, Jin Xu, Zhibo Yang, Mingkun Yang, Jianxin Yang, An Yang, Bowen Yu, Fei Zhang, Hang Zhang, Xi Zhang, Bo Zheng, Humen Zhong, Jingren Zhou, Fan Zhou, Jing Zhou, Yuanzhi Zhu, and Ke Zhu. Qwen3-vl technical report, 2025. 2, 13 [2] Nora Belrose, Zach Furman, Logan Smith, Danny Halawi, Igor Ostrovsky, Lev McKinney, Stella Biderman, and Jacob Steinhardt. Eliciting latent predictions from transformers with the tuned lens. CoRR, abs/2303.08112, 2023. 8 [3] Adrian Chang, Sheridan Feucht, Byron C Wallace, and David Bau. Does FLUX know what it’s writing? In NeurIPS Mechanistic Interpretability Workshop, San Diego, 2025. 14 [4] Hila Chefer, Shir Gur, and Lior Wolf. Transformer interpretability beyond attention visualization. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 782–791, 2021. 5 [5] Haozhe Chen, Junfeng Yang, Carl Vondrick, and Chengzhi Mao. Interpreting and controlling vision foundation models via text explanations, 2023. 8 [6] Christopher Clark, Jieyu Zhang, Zixian Ma, Jae Sung Park, Mohammadreza Salehi, Rohun Tripathi, Sangho Lee, Zhongzheng Ren, Chris Dongjoo Kim, Yinuo Yang, Vincent Shao, Yue Yang, Weikai Huang, Ziqi Gao, Taira Anderson, Jianrui Zhang, Jitesh Jain, George Stoica, Winson Han, Ali Farhadi, and Ranjay Krishna. Molmo2: Open weights and data for vision-language models with video understanding and grounding. CoRR, abs/2601.10611, 2026. 13 [7] Timothée Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski. Vision transformers need registers. In The

Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. 4, 6 [8] Stanislas Dehaene and Laurent Cohen. The unique role of the visual word form area in reading. Trends in Cognitive Sciences, 15(6):254–262, 2011. 1 [9] Stanislas Dehaene, Laurent Cohen, José Morais, and Régine Kolinsky. Illiterate to literate: behavioural and cerebral changes induced by reading acquisition. Nature Reviews Neuroscience, 16(4):234–244, 2015. 1 [10] Alexander Yom Din, Taelin Karidi, Leshem Choshen, and Mor Geva. Jump to conclusions: Short-cutting transformers with linear transformations. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC/COLING 2024, 20-25 May, 2024, Torino, Italy, pages 9615–9625. ELRA and ICCL, 2024. 8 [11] Nelson Elhage, Neel Nanda, Catherine Olsson, Tom Henighan, Nicholas Joseph, Ben Mann, Amanda Askell, Yuntao Bai, Anna Chen, Tom Conerly, Nova DasSarma, Dawn Drain, Deep Ganguli, Zac Hatfield-Dodds, Danny Hernandez, Andy Jones, Jackson Kernion, Liane Lovitt, Kamal Ndousse, Dario Amodei, Tom Brown, Jack Clark, Jared Kaplan, Sam McCandlish, and Chris Olah. A mathematical framework for transformer circuits. Transformer Circuits Thread, 2021. https://transformercircuits.pub/2021/framework/index.html. 2, 4 [12] Sheridan Feucht, Eric Todd, Byron Wallace, and David Bau. The dual-route model of induction. In Second Conference on Language Modeling, 2025. 3, 4, 8, 14 [13] Sheridan Feucht, Byron Wallace, and David Bau. Vector arithmetic in concept and token subspaces. In Second Mechanistic Interpretability Workshop at NeurIPS, 2025. 2, 4, 8 [14] Yossi Gandelsman, Alexei A. Efros, and Jacob Steinhardt. Interpreting clip’s image representation via text-based decomposition. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. 5, 8 [15] Rohit Gandikota and David Bau. Gaze heads: How vlms look at what they describe. arXiv preprint arXiv:2606.14703, 2026. 8, 12 [16] Mor Geva, Avi Caciularu, Kevin Ro Wang, and Yoav Goldberg. Transformer feed-forward layers build predictions by promoting concepts in the vocabulary space, 2022. 8 [17] Gabriel Goh, Nick Cammarata †, Chelsea Voss †, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah. Multimodal neurons in artificial neural networks. Distill, 2021. https://distill.pub/2021/multimodalneurons. 8 [18] Yash Goyal, Tejas Khot, Douglas Summers-Stay, Dhruv Batra, and Devi Parikh. Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering. In Conference on Computer Vision and Pattern Recognition (CVPR), 2017. 13 [19] Wes Gurnee, Nicholas Sofroniew, Adam Pearce, Mateusz Piotrowski, Isaac Kauvar, Runjin Chen, Anna Soligo, Paul Bogdan, Euan Ong, Rowan Wang, Ben Thompson, David Abrahams, Subhash Kantamneni, Emmanuel

Ameisen, Joshua Batson, and Jack Lindsey. Verbalizable representations form a global workspace in language models. Transformer Circuits Thread, 2026. 8 [20] Zhentao He, Can Zhang, Ziheng Wu, Zhenghao Chen, Yufei Zhan, Yifan Li, Zhao Zhang, Xian Wang, and Minghui Qiu. Seeing is believing? mitigating ocr hallucinations in multimodal large language models, 2025. 1, 8 [21] Evan Hernandez, Arnab Sen Sharma, Tal Haklay, Kevin Meng, Martin Wattenberg, Jacob Andreas, Yonatan Belinkov, and David Bau. Linearity of relation decoding in transformer language models. In Proceedings of the 2024 International Conference on Learning Representations, 2024. 7, 8 [22] Etha Tianze Hua, Tian Yun, and Ellie Pavlick. Sourcemodality monitoring in vision-language models, 2026. 8 [23] Tianze Hua, Tian Yun, and Ellie Pavlick. How do visionlanguage models process conflicting information across modalities? CoRR, abs/2507.01790, 2025. 8 [24] Minyoung Huh, Brian Cheung, Tongzhou Wang, and Phillip Isola. The platonic representation hypothesis. CoRR, abs/2405.07987, 2024. 8 [25] Chaoya Jiang, Haiyang Xu, Mengfan Dong, Jiaxing Chen, Wei Ye, Ming Yan, Qinghao Ye, Ji Zhang, Fei Huang, and Shikun Zhang. Hallucination augmented contrastive learning for multimodal large language model, 2024. 2, 8 [26] Nick Jiang, Amil Dravid, Alexei A. Efros, and Yossi Gandelsman. Vision transformers don’t need trained registers. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, 2025. 6 [27] Nick Jiang, Anish Kachinthaya, Suzie Petryk, and Yossi Gandelsman. Interpreting and editing vision-language representations to mitigate hallucinations, 2025. 2, 4, 5, 8 [28] Sonia Joseph and Neel Nanda. Laying the foundations for vision and multimodal mechanistic interpretability & open problems, 2024. Accessed: 2024-06-28. 8 [29] Antonia Karamolegkou, Nicolas Angleraud, Benoı̂t Sagot, and Thibault Clérice. Reading or guessing? visual grounding failures of vision-language models for ocr in ancient greek editions, 2026. 1, 8 [30] Shahar Katz, Yonatan Belinkov, Mor Geva, and Lior Wolf. Backward lens: Projecting language model gradients into the vocabulary space, 2024. 8 [31] Benno Krojer, Dheeraj Vattikonda, Luis Lara, Varun Jampani, Eva Portelance, Chris Pal, and Siva Reddy. Learning action and reasoning-centric image editing from videos and simulation. In Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024. 14 [32] Benno Krojer, Shravan Nayak, Oscar Mañas, Vaibhav Adlakha, Desmond Elliott, Siva Reddy, and Marius Mosbach. Latentlens: Revealing highly interpretable visual tokens in LLMs. In Forty-third International Conference on Machine Learning, 2026. 2, 4, 6, 8, 13, 14, 19

[33] Vedang Lad, Jin Hwa Lee, Wes Gurnee, and Max Tegmark. Remarkable robustness of llms: Stages of inference? In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, 2025. 4 [34] Andrew Lee, Yonatan Belinkov, Fernanda B. Viégas, and Martin Wattenberg. Decomposing query-key feature interactions using contrastive covariances. CoRR, abs/2602.04752, 2026. 3 [35] Yueyan Li, Chenggong Zhao, Zeyuan Zang, Caixia Yuan, and Xiaojie Wang. Reading images like texts: Sequential image understanding in vision-language models. arXiv preprint arXiv:2509.19191, 2025. 8 [36] Yunhao Liang, Ruixuan Ying, Bo Li, Hong Li, Kai Yan, Qingwen Li, Min Yang, Okamoto Satoshi, Zhe Cui, and Shiwen Ni. Visual merit or linguistic crutch? a close look at deepseek-ocr, 2026. 1, 8 [37] Tsung-Yi Lin, Michael Maire, Serge J. Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. Microsoft COCO: common objects in context. In Computer Vision - ECCV 2014 - 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V, pages 740–755. Springer, 2014. 4, 5, 13 [38] Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. Llava-next: Improved reasoning, ocr, and world knowledge, 2024. 13 [39] Oscar Mañas, Pau Rodriguez Lopez, Saba Ahmadi, Aida Nematzadeh, Yash Goyal, and Aishwarya Agrawal. MAPL: Parameter-efficient adaptation of unimodal pre-trained models for vision-language few-shot prompting. In Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics, pages 2523–2548, Dubrovnik, Croatia, 2023. Association for Computational Linguistics. 8 [40] Joanna Materzynska, Antonio Torralba, and David Bau. Disentangling visual and written concepts in CLIP. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2022, New Orleans, LA, USA, June 1824, 2022, pages 16389–16398. IEEE, 2022. 8 [41] Bruce D. McCandliss, Laurent Cohen, and Stanislas Dehaene. The visual word form area: expertise for reading in the fusiform gyrus. Trends in Cognitive Sciences, 7(7): 293–299, 2003. 1 [42] Jack Merullo, Louis Castricato, Carsten Eickhoff, and Ellie Pavlick. Linearly mapping from image to text space. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. 8 [43] Ron Mokady, Amir Hertz, and Amit H. Bermano. Clipcap: CLIP prefix for image captioning. CoRR, abs/2111.09734, 2021. 8 [44] Clement Neo, Luke Ong, Philip H. S. Torr, Mor Geva, David Krueger, and Fazl Barez. Towards interpreting visual information processing in vision-language models. ArXiv, abs/2410.07149, 2024. 1, 2, 8

[45] Nostalgebraist. Interpreting gpt: The logit lens. https : / / www . alignmentforum . org / posts / AcKRB8wDpdaN6v6ru/interpreting- gpt- thelogit-lens, 2020. Accessed: 23 Sep 2024. 2, 3, 8 [46] Koyena Pal, Jiuding Sun, Andrew Yuan, Byron C Wallace, and David Bau. Future lens: Anticipating subsequent tokens from a single hidden state. In Proceedings of the 27th Conference on Computational Natural Language Learning (CoNLL), pages 548–560, 2023. 8 [47] Isabel Papadimitriou, Huangyuan Su, Thomas Fel, Naomi Saphra, Sham M. Kakade, and Stephanie Gil. Interpreting the linear structure of vision-language model embedding spaces. CoRR, abs/2504.11695, 2025. 8 [48] Athulith Paraselli, Etha Tianze Hua, and Ellie Pavlick. Slow to see, slow to suppress: Understanding the effects of modality in context-memory conflicts, 2026. 8 [49] Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Li Fei-Fei. Imagenet large scale visual recognition challenge. CoRR, abs/1409.0575, 2014. 2, 7, 14 [50] Sarah Schwettmann, Neil Chowdhury, Samuel Klein, David Bau, and Antonio Torralba. Multimodal neurons in pretrained text-only transformers. In IEEE/CVF International Conference on Computer Vision, ICCV 2023 - Workshops, Paris, France, October 2-6, 2023, pages 2854–2859. IEEE, 2023. 1, 2, 8 [51] Haz Sameen Shahgir, Xiaofu Chen, Yu Fu, Erfan Shayegani, Nael B. Abu-Ghazaleh, Yova Kementchedjhieva, and Yue Dong. Vlms need words: Vision language models ignore visual detail in favor of semantic anchors. CoRR, abs/2604.02486, 2026. 8 [52] Arnab Sen Sharma, Giordano Rogers, Natalie Shapira, and David Bau. Llms process lists with general filter heads. The Fourteenth International Conference on Learning Representations (ICLR 2026), 2025. 3, 8, 12 [53] Yan Shu, Hangui Lin, Yexin Liu, Yan Zhang, Gangyan Zeng, Yan Li, Yu Zhou, Ser-Nam Lim, Harry Yang, and Nicu Sebe. When semantics mislead vision: Mitigating large multimodal models hallucinations in scene text spotting and understanding, 2025. 8 [54] Mustafa Shukor and Matthieu Cord. Implicit multimodal alignment: On the generalization of frozen llms to multimodal inputs. In Advances in Neural Information Processing Systems, pages 130848–130886. Curran Associates, Inc., 2024. 2 [55] Alexa R. Tartaglini, Satchel Grant, Daniel Wurgaft, Christopher Potts, and Judith E. Fan. Diagnosing bottlenecks in data visualization understanding by vision-language models, 2025. 8 [56] Eric Todd, Millicent L. Li, Arnab Sen Sharma, Aaron Mueller, Byron C. Wallace, and David Bau. Function vectors in large language models. In The Twelfth International Conference on Learning Representations, 2024. arXiv:2310.15213. 4 [57] Maria Tsimpoukelli, Jacob Menick, Serkan Cabi, S. M. Ali Eslami, Oriol Vinyals, and Felix Hill. Multimodal few-shot

learning with frozen language models. In Advances in Neural Information Processing Systems 34: Annual Conference on Neural Information Processing Systems 2021, NeurIPS 2021, December 6-14, 2021, virtual, pages 200–212, 2021. 8 [58] Constantin Venhoff, Ashkan Khakzar, Sonia Joseph, Philip Torr, and Neel Nanda. How visual representations map to language feature space in multimodal llms, 2025. 1, 2 [59] Constantin Venhoff, Ashkan Khakzar, Sonia Joseph, Philip H. S. Torr, and Neel Nanda. Too late to recall: Explaining the two-hop problem in multimodal knowledge retrieval. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2025, NeurIPS 2025, San Diego, CA, USA, December 2-7, 2025 / Mexico City, Mexico, November 30 - December 5, 2025, 2025. 2, 8 [60] Kevin Wang, Alexandre Variengien, Arthur Conmy, Buck Shlegeris, and Jacob Steinhardt. Interpretability in the wild: a circuit for indirect object identification in GPT-2 small. CoRR, abs/2211.00593, 2022. 3 [61] Eliot Weinberger. 19 ways of looking at Wang Wei: with more ways. New Directions Books, New York, NY, 2016. Afterword by Octavio Paz. 14 [62] Evžen Wybitul, Javier Rando, Florian Tramèr, and Stanislav Fort. Representations of text and images align from layer one, 2026. 2, 4, 8 [63] An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, Chujie Zheng, Dayiheng Liu, Fan Zhou, Fei Huang, Feng Hu, Hao Ge, Haoran Wei, Huan Lin, Jialong Tang, Jian Yang, Jianhong Tu, Jianwei Zhang, Jianxin Yang, Jiaxi Yang, Jing Zhou, Jingren Zhou, Junyang Lin, Kai Dang, Keqin Bao, Kexin Yang, Le Yu, Lianghao Deng, Mei Li, Mingfeng Xue, Mingze Li, Pei Zhang, Peng Wang, Qin Zhu, Rui Men, Ruize Gao, Shixuan Liu, Shuang Luo, Tianhao Li, Tianyi Tang, Wenbiao Yin, Xingzhang Ren, Xinyu Wang, Xinyu Zhang, Xuancheng Ren, Yang Fan, Yang Su, Yichang Zhang, Yinger Zhang, Yu Wan, Yuqiong Liu, Zekun Wang, Zeyu Cui, Zhenru Zhang, Zhipeng Zhou, and Zihan Qiu. Qwen3 technical report, 2025. 12

A. Model Details

C. Overlap with Other Heads

In this work, we study four models of varying sizes with three distinct model families. See Table 2 for architectural details of models studied in this paper.

Attention patterns from Figure 4 are reminiscent of filter heads from prior work [52]. Are they the same heads? We find that of the 11 filter heads found in Qwen3-1.7B [63] (the base LLM for Qwen3-VL-2B-Instruct), six of them are also in the top-10% of verbalization heads. This suggests that vision fine-tuning may cause filter heads to “evolve” into verbalization heads. We also check for overlap with gaze heads [15], a type of attention head in VLMs that keeps track of which portion of an image the VLM is captioning. For Qwen3-VL-2B, 30% of gaze heads are also verbalization heads (3/10), and for Qwen3-VL-8B, 29% of gaze heads are also verbalization heads (29/100). This is higher than the expected overlap when sampling the same number of random heads.6 Although this means that verbalization heads are somewhat distinct from gaze heads, the modest overlap also supports our claim that analysis of OCR can identify heads more generally responsible for mapping from pixels to semantics.

B. Finding Verbalization Heads OCR dataset examples are shown in Figure 10. We use large font sizes and randomly sample colors that contrast with the background. Table 3 shows model accuracy for these stimuli. Due to models having poor OCR performance on stimuli with white backgrounds, we opt for realistic backgrounds.

Figure 10. Random examples from our OCR dataset used to identify heads in Section 2.

Results from Section 2 are shown for Molmo2-7B and Llava-Next-34B in Figure 11 and Figure 12. We also progressively ablate heads responsible for OCR for Qwen3VL-2B on VQA in Figure 13.

ed

blations

Layer Layer

Layer

Figure 11. OCR scoresLayer from Equation 1 for Molmo2-O-7B and Llava-Next-34B. The takeaway is similar to Figure 2 for Qwen3VL models: OCR heads appear in late layers. Molmo2-7B Head Mean Ablations

Llava-Next-34B Head Mean Ablations

Percent of model heads ablated

Percent of model heads ablated

OCR Accuracy

ations

Llava-Next-34B Scores Head Index

Head Index

Molmo2-7B OCR Scores

Llava-Next-34B Scores Head Index

Head Index

Molmo2-7B OCR Scores

Figure 12. Progressively mean-ablating top-scoring OCR heads for Molmo2-O-7B and Llava-Next-34B. Results are similar to Qwen3-VL results from Figure 3.

D. Full Verbalization Lens Examples In Figures 16, 17, and 18, we show full verbalization lens outputs for models not shown in Figure 5. All outputs are for layer 0, and compared to raw logit lens outputs at layer 0.

E. Object Confidence For each category, we sample n=512 images containing o and n images without o, except for Llava-Next-34B, for which n=128 due to computational constraints. Figure 19 shows these results using ROC-AUC curves, instead of calculating average difference in maximum P (o) as in the main paper. Interestingly, ROC-AUC shows that raw logit lens has more signal for Qwen models, despite verbalization lens still showing marked improvements.

F. Object Localization Details Here, we provide implementation details for Section 3.5. For each image, we score one segmentation map per object category present in that image, where the ground truth mask is the union of all instances of that category (e.g., if there are multiple people in an image, our ground truth mask includes all of those people). For the 512 images that we display numbers for in Table 1, there are 1216 (image, category) pairs, an average of just over two object categories per image. Averaging over these pairs ensures that each object counts once regardless of image/object size. 6 For Qwen3-VL-2B, we have 45 verbalization heads and 10 gaze heads, so the expected number of overlapping heads is (10×45)/448, or 1/10 gaze heads. For Qwen3-VL-8B, we have 115 verbalization heads and 100 gaze heads, so the expected overlap with random samples is (100×115)/1152, or 10/100 gaze heads.

Table 2. Architectural details of models used in this paper. “Tied” refers to tied embeddings, “Img Encoder” refers to whether the image encoder is trained end-to-end or frozen, and “Single-tok COCO” is the number of COCO [37] categories that are single-token for this tokenizer (which determines results in Sections 3.4-3.5).

Full Model Name

Cite

Qwen/Qwen3-VL-2B-Instruct Qwen/Qwen3-VL-8B-Instruct allenai/Molmo2-O-7B llava-hf/llava-v1.6-34b-hf

[1] [1] [6] [38]

Abbrv.

Total Heads

Tied?

Img Encoder?

Single-tok COCO

Qwen3-VL-2B

28 lyrs × 16 hds = 448 36 lyrs × 32 hds = 1152 32 lyrs × 32 hds = 1024 60 lyrs × 56 hds = 3360

Yes No No No

End-to-end End-to-end End-to-end End-to-end

56/80 56/80 56/80 51/80

Qwen3-VL-8B Molmo2-7B Llava-Next-34B

Qwen3-VL-2B-Instruct Head Mean Ablations by answer type yes/no (n=408)

VQAv2 accuracy

1.0

number (n=111)

other (n=505) Original (58.2%) Top OCR-ranked heads Random - all layers Random - last 50%

0.8 0.6 0.4 0.2 0.0

0

20

40 60 80 Heads ablated (%)

100 0

20

40 60 80 Heads ablated (%)

100 0

20

40 60 80 Heads ablated (%)

100

Figure 13. We also measure accuracy for VQAv2 [18] when progressively mean-ablating the top-scoring OCR heads throughout generation. While ablation of these heads does not affect performance for questions that have Yes/No or numeric answers, it does affect performance for other types of questions (e.g., “What is the type of food in this image?” → “Mexican”) slightly more than an equal number of random heads. This suggests that verbalization heads may be useful for more general-purpose question answering tasks.

Table 3. OCR Accuracy for randomly-selected English words superimposed on either ImageNet or white backgrounds (n = 100). Accuracies are obtained with the prompt “Transcribe the word in this image.” and assistant prefill “Word:”.

Model Qwen3-VL-2B Qwen3-VL-8B Molmo2-O-7B Llava-Next-34B

ImageNet Bg

Table 4. Thresholds τ for each approach in Table 1, chosen based on a disjoint set of 128 images. We sweep over a log-spaced range of 60 thresholds based on the range of probability scores for that method.

Qwen3-VL-8B

White Bg

0.89 0.93 0.90 0.80

0.07 0.05 0.44 0.75

We pick a threshold τ for mIoU based on a disjoint set of 128 images. We show chosen mIoU thresholds in Table 4. We omit token-level accuracy, as in our dataset, only 15.7% of tokens are positive. Therefore, a a baseline that never segments any objects can trivially achieve 84.3% token accuracy.

G. LatentLens Comparison Details We describe our setup for the LatentLens [32] comparison from Section 3.6 in more detail here.

Random Heads Raw Logit Lens Verbalization Lens

−6

7.2×10 3.0×10−4 3.4×10−3

Qwen3-VL-2B 5.1×10−8 4.4×10−3 1.1×10−3

LatentLens replication. For each of the four models, we build a LatentLens contextual index following Krojer et al. [32]’s released implementation and corpus. We extract contextual embeddings at every model’s embedding layer plus a set of decoder layers space across the model, and store up to 30 contexts per unique token. Following [32], we retrieve the top-k nearest neighbors by cosine similarity across all indexed layers jointly at query time. As in their evaluation, LatentLens neighbor tokens are often subwords; we expand each to its full containing word using the neighbor’s own source sentence as context (e.g. “ing” → “rendering”). We find that a small fraction of nearest-neighbor matches (concentrated in Qwen3-VL-8B’s middle layers)

are whitespace tokens carrying no semantic content despite high cosine similarity. Thus, we search k = 32 neighbors and take the top-5 non-blank words.

Images. We sample 100 random images (fixed seed) from the same COCO val2014 pool used elsewhere in the paper. For each image we decode one patch, chosen pseudorandomly (seeded per image) from the central 60% of that model’s patch grid to avoid picking pure background/padding; the same patch is used across all 6 layers for a given image. LLM judge. We use the same LLM-judge as Krojer et al. [32]. We show GPT-5 the full image with a red bounding box over the target patch, a cropped close-up of that region, and the top-5 candidate words for one lens at a particular layer. The judge labels each candidate as concrete (literally visible in the box), abstract (an implied concept/activity/quality), or globally related (visible elsewhere in the image but not the box). A patch is scored interpretable if at least one candidate word receives any of these three labels. Results. Figure 20 shows the resulting interpretabletoken percentage across layers for all four models. For all models, verbalization lens is on par with or better than LatentLens in terms of LLM judge scores. We note that verbalization lens is especially helpful for Llava-Next-34B.

H. LLM Judge for Editing Table 5 lists LLM judge prompts for evaluating image edits from Section 4.2. For each question, the judge returns a structured response containing an integer answer (its rating).

I. Editing Sweep for Rank/Scaling Factors We run an initial sweep over ranks and scaling factors for the edit in Section 4. For each setting, we greedy-decode a caption per image. We choose to sweep across scaling factors α ∈ {1, 2, 3, 4, 5, 10}. Ranks are chosen for each model based on the number of dimensions it takes to explain ρ% of the energy of Lp . That is, if Σ=[σ1 , ...σdmodel ], we take the minimum top-k dimensions such that sum(Σ[: k]2 )/sum(Σ2 ) > ρ/100 for ρ ∈ {1, 5, 10, 20, 40, 60, 80, 90, 100}. For reference, we

15

ρ=60% of energy ρ=90% of energy

10 σ

Layers. For each model we evaluate six layers: Qwen3-VL-2B {0, 8, 12, 16, 20, 27}, Qwen3-VL-8B {0, 8, 12, 20, 24, 35}, Molmo2-O-7B {0, 8, 12, 19, 22, 31}, Llava-Next-34B {0, 10, 24, 36, 42, 59}. See Table 2 for model details.

Singular Values of Verbalization Lens Qwen3-VL-2B-Instruct

5 0

0

500

1000 Dimension

1500

2000

Figure 14. Singular values of Lp for Qwen3-VL-2B. We plot dotted lines to indicate the ranks that explain 60% and 90% of energy. See Section 4 for discussion of low-rank approximations of Lp .

show the singular values of Lp for Qwen3-VL-2B in Figure 14. Figures 21 and 22 show results for LLM judge metrics from Section 4.2 across n=100 randomly-sampled ImageNet images/object pairs [49]. We find that our edit generally requires a scaling factor α > 1, and does not have an effect for full-rank L−1 p . We also find that a large scaling factor α=10 is damaging, as it causes specificity to drop. Thus, we choose ρ=60% of verbalization lens energy and perform edits for n=256 additional images across scaling factors (excluding α=10) in Figure 8b. ρ is constant for all qualitative generations shown in the paper.

J. Behind the Scenes Following [31] and [32], we describe the trajectory of this project in hopes of humanizing this work and making the research process transparent. This project began because SF was interested in writing and letters. They had worked on letter representations in text-to-image diffusion models before [3], but had just returned from an internship and were trying to decide how exactly to continue this line of work. With encouragement from Si Wu and John Cayley, they began working on a project (which is still in-progress) on VLM representations of Chinese poetry; specifically, the poem 鹿柴 by Wang Wei, which has an interesting history of analysis in terms of its translations [61]. In trying to understand how VLMs might process an image of this poem, SF set out to first figure out how VLMs do OCR in general. The plan was just to figure out how OCR worked, so that this knowledge could be leveraged to create interesting visualizations for this art project. But then, SF realized that the heads they had picked out for OCR were much more general when they were playing around with the OV lens approach used in prior work [12]. Clicking on the word “Hey” in a test image yielded that word, as expected.

Question

Prompt template

Prevalence of c

On a scale from 0–10, how prominently does the concept of ‘{c}’ feature in this image caption? Let’s say that a score of 10 means that the concept is the main focus, 0 means that the concept is not mentioned at all, and in-between scores mean that the concept is mentioned to some extent. If the caption is complete nonsense, rate it as 0. Caption:\n\n{caption}

Specificity

You will be shown two captions. The first is a Ground Truth caption describing a real image. The second is a Counterfactual version of the first caption, where the concept of ‘crem ’ has been replaced by the concept of ‘cadd ’. Please rate how similar the Counterfactual caption is to the Ground Truth caption on a scale from 0–10. If the two captions are the exact same, give a rating of 10; close paraphrases can also get high ratings. You can still give a high score if the Counterfactual caption appears to describe a similar scene to the Ground Truth caption when you ignore the edited concepts: for example, if we removed the concept “motorcycle” and replaced it with “dog,” then the captions “a motorcycle in an empty garage, dimly lit” and “a dog in a low-lit empty garage” would get a score of 9, because they are close paraphrases of each other when you ignore the edited concepts. In the worst case, if the Counterfactual Caption has nothing to do with the Ground Truth Caption at all, give a rating of 0. Ground Truth Caption:\n\n{clean caption}

Coherence

Counterfactual Caption:\n\n {edited caption} On a scale from 0–10, ignoring the semantics and focusing on grammar and writing quality, how coherent is this image caption (with 10 being perfectly coherent and 0 being complete gibberish)? Caption:\n\n{caption}

Table 5. Question templates for LLM judging from Section 4.2. We use the “Prevalence of c” template both for crem and cadd . For Specificity scores, we also give the judge the model’s original clean caption, so that it can evaluate whether any important details have changed.

But out of curiosity they clicked on an emoji in the image, expecting the result to be garbage. Surprisingly, they saw the token emoji in early layers! In that moment, the authors realized that they had a pretty interesting little finding that could provide a nice approach to “logit lens for image tokens.” At the same time, HA and SW had been helping out SF with other lines of attack on understanding text in images (studying emoji and lexical representations). With this surprise, the authors decided to band together and write a paper on this finding. They decided to submit to a conference in a location SF wanted to visit for personal reasons. They recruited BK, who was just beginning his postdoc at the lab, because of his expertise from prior work on this topic, and started writing this paper. Even though BK was in the middle of moving in, and SF came down with a fever in the last few weeks of work on this paper, it ended up coming together in the end.

Figure 15. The original image that SF was using, which was just a random image with the text “Hey” and an emoji overlaid using Photoshop. Clicking on the emoji yielded emoji, even though SF expected to just see noise.

Qwen3-VL-2B: Raw Logit Lens (Layer 0)

Qwen3-VL-2B: +Verbalization Lens (Layer 0)

Figure 16. Verbalization lens results for image in Figure 5 for Qwen3-VL-2B.

Molmo2-O-7B: Raw Logit Lens (Layer 0)

Molmo2-O-7B: +Verbalization Lens (Layer 0)

Figure 17. Verbalization lens results for image in Figure 5 for Molmo2-O-7B.

Llava-Next-34B: Raw Logit Lens (Layer 0)

Llava-Next-34B: +Verbalization Lens (Layer 0)

Figure 18. Verbalization lens results for image in Figure 5 for Llava-Next-34B. Image is slightly cropped for readability.

Layerwise ROC-AUC: Object Confidence Qwen3-VL-8B

Qwen3-VL-2B

Molmo2-O-7B

Llava-Next-34B

Verbalization Lens Raw Logit Lens Random 10% of heads All heads

Figure 19. ROC-AUC scores for the experiment in Figure 6b. For raw logit lens, layers where ROC-AUC is high but probability differences in Figure 6b are low indicate that logit lens is badly calibrated—even if the ranking between tokens is correct, probability differences are very small, indicating low confidence. We see a similar effect when using all attention heads (a superset that includes our verbalization heads): using all attention heads can achieve high ROC-AUC, but the lower overall probability differences in Figure 6b indicate lower confidence. For verbalization lens, ROC-AUC is close to ceiling across layers in addition to large probability deltas in Figure 6b. This indicates that verbalization lens has higher confidence, which also makes our approach more qualitatively useful (i.e., semantic predictions are assigned high probabilities, making them visible in qualitative visualizations; see interactive demo).

Qwen3-VL-2B

40 20 VerbalLens LatentLens

0

5

10

15 Layer

20

60 40 20 VerbalLens LatentLens

5

10

15 Layer

20

20 VerbalLens LatentLens

0

5

25

30

10

15

Layer

20

25

30

35

LLaVA-NeXT-34B

100

80

0

40

0

25

Molmo2-O-7B

100 Interpretable Visual Tokens (%)

80

FINAL VERSIN with n=512 60 per category (there are about 47 single tok categories) stride acr

60

0

Interpretable Visual Tokens (%)

80

0

Qwen3-VL-8B

100

Interpretable Visual Tokens (%)

Interpretable Visual Tokens (%)

100

80 60 40 20 0

VerbalLens LatentLens

0

10

20

30 Layer

40

50

60

Figure 20. Verbalization lens is on par with or better than LatentLens [32] across all models. We show LLM judge interpretability (% of patches judged interpretable) across layers (100 images per point). LatentLens shows a pronounced early/mid-layer dip for Llava-Next34B, but verbalization lens scores consistently high across layers.

Qwen3-VL-2B-Instruct: Mean Judged Scores Across Ranks/Scales

-2.3

-3.0

-4.0

-4.4

-5.2

5% (10)

0.7

-0.1

-1.1

-1.8

-2.9

-4.8

10% (22)

0.4

-0.1

-1.0

-1.8

-2.2

-4.2

20% (48)

-0.3

-0.6

-0.8

-1.4

-2.1

-4.4

40% (115)

-1.8

-3.2

-3.9

-5.0

-5.2

-5.5

60% (219)

-2.5

-4.3

-5.2

-5.5

-5.7

-5.9

80% (418)

-1.9

-3.4

-4.9

-5.2

-5.6

-5.8

90% (627)

-0.7

-2.3

-3.2

-4.7

-5.3

-5.9

100% (2048)

-2.0

-3.3

-3.9

-4.6

-4.9

-5.4

1

2

3

4

5

10

1% (2)

6.6

5.1

3.8

2.8

2.1

1.1

5% (10)

8.6

7.7

6.4

5.8

4.2

1.8

10% (22)

8.6

8.1

7.3

6.4

5.3

2.7

20% (48)

8.7

8.3

8.0

6.9

6.3

3.5

40% (115)

8.3

7.8

6.6

5.9

5.2

3.1

60% (219)

8.0

7.3

6.7

6.0

5.5

2.8

80% (418)

8.3

7.6

6.6

6.2

5.4

3.2

90% (627)

8.5

8.0

7.2

6.4

5.9

3.6

100% (2048)

5.5

3.8

2.6

1.9

1.6

0.8

1

2

3

4

5

10

α Scaling Factor Specificity ↑

α Scaling Factor

Δ Prevalence of Added Object ↑

10 5 0

−5

Rank (ρ% of Verbal. Energy)

1% (2)

-1.0

−10

10 8 6 4 2 0

Rank (ρ% of Verbal. Energy)

Rank (ρ% of Verbal. Energy)

Rank (ρ% of Verbal. Energy)

Δ Prevalence of Removed Object ↓

1% (2)

0.0

0.0

0.0

0.0

0.0

0.0

5% (10)

0.0

0.0

0.0

0.0

0.0

0.0

10% (22)

0.0

0.0

0.1

0.1

0.1

0.0

20% (48)

0.0

0.0

0.1

0.3

0.4

0.1

40% (115)

0.2

1.1

2.0

2.0

2.1

1.7

60% (219)

0.5

2.9

4.2

4.6

4.2

3.3

80% (418)

0.4

2.8

3.9

4.3

5.1

3.6

90% (627)

0.1

1.2

3.1

3.9

4.5

3.3

100% (2048)

0.0

0.0

0.0

0.0

0.0

0.0

1

2

3

4

5

10

1% (2)

9.1

8.8

8.4

8.3

8.4

7.6

5% (10)

9.4

9.3

8.8

8.9

8.3

8.2

10% (22)

9.6

9.3

9.3

8.7

8.6

8.3

20% (48)

9.3

9.4

9.1

8.4

8.3

7.7

40% (115)

9.3

8.5

8.0

7.8

7.6

7.7

60% (219)

9.2

8.5

8.0

7.2

7.2

6.5

80% (418)

9.2

8.4

7.7

7.4

7.0

6.7

90% (627)

9.5

8.9

8.1

7.7

7.4

6.8

100% (2048)

8.9

8.4

7.9

7.9

7.4

6.0

1

2

3

4

5

10

α Scaling Factor Coherence ↑

α Scaling Factor

10 5 0

−5 −10

10 8 6 4 2 0

Figure 21. Sweep across pseudo-inverse ranks and scales for editing objects in Qwen3-VL-2B. We generate edits across every combination of rank and scaling factor for 100 random ImageNet images, score the edited captions, judge outputs, and present the mean of scores across images. We focus on results for ρ=60% for Qwen3-VL-2B in the main paper.

Qwen3-VL-8B-Instruct: Mean Judged Scores Across Ranks/Scales

0.0

-1.2

-2.7

-3.4

-3.8

-5.2

5% (19)

0.3

0.2

-0.8

-1.6

-2.4

-4.3

10% (41)

0.2

-0.1

-1.2

-2.3

-3.5

-5.5

20% (92)

-0.3

-1.1

-2.4

-3.7

-4.4

-5.8

40% (234)

-1.3

-2.8

-3.9

-5.0

-5.3

-6.2

60% (463)

-1.3

-2.3

-3.8

-4.8

-5.6

-6.2

80% (892)

-0.8

-1.8

-2.9

-3.2

-4.4

-5.9

90% (1319)

-0.6

-0.4

-1.0

-2.3

-3.1

-4.7

100% (4096)

-2.5

-3.6

-4.5

-4.9

-5.5

-6.2

1

2

3

4

5

10

1% (4)

9.1

7.2

5.9

4.6

3.9

2.2

5% (19)

9.4

9.3

8.7

7.4

6.7

3.4

10% (41)

9.4

9.3

8.8

7.6

6.5

3.7

20% (92)

9.3

9.2

8.7

8.2

7.6

4.4

40% (234)

9.3

8.9

8.5

8.2

7.2

5.5

60% (463)

9.3

8.7

8.8

8.1

7.4

5.8

80% (892)

9.4

9.2

8.7

8.4

8.0

5.0

90% (1319)

9.4

9.4

9.3

8.3

8.2

6.8

100% (4096)

5.6

3.7

2.6

2.1

1.5

0.7

1

2

3

4

5

10

α Scaling Factor Specificity ↑

α Scaling Factor

10 5 0

−5

Rank (ρ% of Verbal. Energy)

1% (4)

Δ Prevalence of Added Object ↑

−10

10 8 6 4 2 0

Rank (ρ% of Verbal. Energy)

Rank (ρ% of Verbal. Energy)

Rank (ρ% of Verbal. Energy)

Δ Prevalence of Removed Object ↓

1% (4)

0.0

0.0

0.0

0.0

0.0

0.0

5% (19)

0.0

0.0

0.1

0.1

0.0

0.0

10% (41)

0.0

0.0

0.2

0.2

0.2

0.1

20% (92)

0.0

0.3

0.8

1.2

1.2

1.3

40% (234)

0.2

0.8

2.3

3.2

3.6

4.5

60% (463)

0.2

1.2

2.6

3.6

4.3

6.1

80% (892)

0.1

0.4

1.7

2.6

3.4

4.4

90% (1319)

0.0

0.1

0.3

1.6

2.5

4.0

100% (4096)

0.0

0.0

0.0

0.0

0.0

0.0

1

2

3

4

5

10

1% (4)

9.8

9.7

9.6

9.6

9.6

8.5

5% (19)

9.9

9.8

9.7

9.7

9.8

9.7

10% (41)

9.9

9.7

9.8

9.7

9.6

9.6

20% (92)

9.8

9.8

9.6

9.4

9.5

9.3

40% (234)

9.8

9.6

9.4

9.5

9.1

9.3

60% (463)

9.8

9.4

9.3

9.3

8.9

9.0

80% (892)

9.8

9.5

9.3

9.2

8.8

8.5

90% (1319)

9.8

9.8

9.8

9.2

8.9

8.9

100% (4096)

9.5

8.7

8.4

7.8

7.3

5.0

1

2

3

4

5

10

α Scaling Factor Coherence ↑

α Scaling Factor

10 5 0

−5 −10

10 8 6 4 2 0

Figure 22. Sweep across pseudo-inverse ranks and scales for editing objects in Qwen3-VL-8B. We generate edits across every combination of rank and scaling factor for 100 random ImageNet images, score the edited captions, judge outputs, and present the mean of scores across images. See Section 4.3 for details. Qwen3-VL-8B appears to require a larger scaling factor α. We focus on results for rank ρ =60% for Qwen3-VL-8B in the main paper.

Record · ID 965456 · SHA-256 a41c2af08aebdfae
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.