arXiv:2604.13803v1 [cs.CV] 15 Apr 2026
G ASLIGHT, G ATEKEEP, V1–V3: E ARLY V ISUAL C ORTEX A LIGNMENT S HIELDS V ISION -L ANGUAGE M ODELS FROM S YCOPHANTIC M ANIPULATION
Arya Shah Indian Institute of Technology Gandhinagar Gandhinagar, India [email protected]
Vaibhav Tripathi Indian Institute of Technology Gandhinagar Gandhinagar, India [email protected]
Mayank Singh Indian Institute of Technology Gandhinagar Gandhinagar, India [email protected]
Chaklam Silpasuwanchai Asian Institute of Technology Bangkok, Thailand [email protected]
A BSTRACT Vision-language models are increasingly deployed in high-stakes settings, yet their susceptibility to sycophantic manipulation remains poorly understood, particularly in relation to how these models represent visual information internally. Whether models whose visual representations more closely mirror human neural processing are also more resistant to adversarial pressure is an open question with implications for both neuroscience and AI safety. We investigate this question by evaluating 12 open-weight vision-language models spanning 6 architecture families and a 40× parameter range (256M–10B) along two axes: brain alignment, measured by predicting fMRI responses from the Natural Scenes Dataset across 8 human subjects and 6 visual cortex regions of interest, and sycophancy, measured through 76,800 two-turn gaslighting prompts spanning 5 categories and 10 difficulty levels. Region-of-interest analysis reveals that alignment specifically in early visual cortex (V1–V3) is a reliable negative predictor of sycophancy (r = −0.441, BCa 95% CI [−0.740, −0.031]), with all 12 leave-one-out correlations negative and the strongest effect for existence denial attacks (r = −0.597, p = 0.040). This anatomically specific relationship is absent in higher-order categoryselective regions, suggesting that faithful low-level visual encoding provides a measurable anchor against adversarial linguistic override in vision-language models. We release our code on GitHub and dataset on Hugging Face Keywords Vision-Language Models · Brain Alignment · Sycophancy · Neural Predictivity · Adversarial Robustness · fMRI
1
Introduction
Vision-language models (VLMs) have rapidly advanced to the point where they can interpret complex visual scenes, answer open-ended questions about images, and reason across modalities with increasing fluency [Li et al., 2023a, Liu et al., 2024a, 2023, Bai et al., 2025]. In parallel, a growing body of work in computational neuroscience has demonstrated that artificial neural networks trained on visual tasks develop internal representations that are remarkably predictive of neural activity in the primate visual cortex [Yamins et al., 2014, Schrimpf et al., 2020]. This correspondence, commonly quantified as “brain alignment” or “neural predictivity,” has become a benchmark for evaluating how faithfully a model captures the computational principles underlying biological vision [Conwell et al., 2024, Gifford et al., 2023]. Recent large-scale studies examining hundreds of models have revealed that brain alignment is not a monolithic property; it varies substantially across cortical regions and is shaped by factors such as training objective, architecture, and visual
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
Stage 2: Sycophancy Evaluation
12
Two-Turn Protocol
Open Weight Vision Language Models (256M 10B Parameters)
→
Gemma-3-1B LFM-2-VL-1B Qwen2-VL-2B BLIP-2-OPT-2.7B Qwen2.5-VL-3B Phi-3.5-Vision LLaVA-v1.6-7B Idefics2-8B LFM-2-VL-8B PaliGemma2-10B
Llama 3.1 70B Instruct Grounded in COCO annotations Yes
SigLIP SigLIP2-NaFlex CLIP-ViT Qwen-ViT ViT-G/14+QFormer SigLIP-mod
Turn 2
Categories: Object Misidentification Attribute Manipulation Existence Denial Count Falsification Authority Appeal
AGREE / DISAGREE
Group Comparison
→ Escalate
→ Persuasive Pressure
Resistant v/s Susceptible
Output: Sycophancy Metrics
Agrees?
No
Ridge Regression
Feature Extraction
floc-faces
floc-words
floc-places
streams
Frozen Vision Encoder
CV for α Pearson r per voxel Noise ceiling normalization
Extended Analysis Architecture Family Comparison Resistance Curves (AURC, Slope) Persuasion Tactic Ranking Per-difficulty Correlations
Resistant
Brain Scores Robustness Checks
Per-voxel Encoding
6 Regions of Interests (ROIs)
(Cohen’s d + Welch’s t per ROI)
Pressure Conversion Total Evaluations = 76,800 (Per-category, Per-difficulty)
Stage 1: Brain Alignment Scoring
floc-bodies
(Pearson r) Per-ROI + aggregate Cross Correlation: 6ROIs x 5 Categories
Exact Match Keyword Phrase Semantic Fallback
Sycophancy Rate Yes
Sycophancy Converted
7T fMRI Data 8 Subjects ~9000 Train + ~200 Test
No
Resists
Sycophantic
Algonauts 2023 / NSD Dataset
prf-visualrois V1, V2, V3, V4
Agrees?
5 Categories x 10 Difficulty Levels x 128 Images
6 Vision Encoder Families
Correlation Analysis
Turn 1: Image + False Claim
Prompt Generation
6,400 Gaslighting Prompts per Model
Stage 3: Statistical Analysis 5-Layer Response Parser
(Aggregate Score)
(ROI-specific score for each of 6 ROIs)
BCa Bootstrap (10,000 resamples, 95% CI)
Permutation (10,000 iterations)
Leave-One-Out (12 subsets stability)
Figure 1: Overview of the three-stage pipeline. Stage 1: Vision encoder features are extracted from 12 VLMs and used to predict fMRI responses across 6 visual cortex ROIs in 8 human subjects (Algonauts 2023). Stage 2: Each model is evaluated on 6,400 two-turn gaslighting prompts spanning 5 manipulation categories and 10 difficulty levels. Stage 3: Brain alignment scores are correlated with sycophancy rates at both aggregate and ROI-specific levels, with robustness checks including BCa bootstrap, leave-one-out, and permutation testing. diet [Conwell et al., 2024]. These findings raise a natural question: does the degree to which a model mirrors human neural processing have consequences beyond predicting brain activity? One such consequence may relate to robustness under adversarial pressure. VLMs are increasingly known to exhibit sycophantic behavior, in which a model abandons a correct response in favor of an incorrect one after a user expresses disagreement or applies social pressure [Sharma et al., 2025, Perez et al., 2023]. This failure mode is particularly concerning because it undermines trust in deployed systems and can be exploited by adversaries to extract harmful or false outputs. Sycophancy has been linked to reinforcement learning from human feedback (RLHF), where models learn to optimize for user approval rather than factual accuracy [Ouyang et al., 2022, Sharma et al., 2025]. While adversarial robustness in VLMs has received growing attention through studies of jailbreaking, prompt injection, and image-based attacks [Zhao et al., 2023, Shayegani et al., 2023, Liu et al., 2024b], no prior work has investigated whether the fidelity of a model’s visual representations to human neural processing relates to its ability to withstand structured sycophantic manipulation. This gap is significant because both brain alignment and sycophancy resistance may depend on the same underlying property: how faithfully a model encodes visual evidence, independent of linguistic context. In this work, we address this gap through a three-stage empirical pipeline applied to 12 open-weight VLMs spanning 256M to 10B parameters. We focus deliberately on small-to-medium open-weight models for three reasons. First, our methodology requires direct access to frozen vision encoder weights to extract intermediate representations for brain alignment computation, a requirement that closed-source systems (e.g., GPT-4V, Gemini) cannot satisfy because they do not expose their internal architecture. Second, open-weight models in this parameter range are the most widely deployed in practice, powering on-device, edge, and resource-constrained applications where safety evaluation is most urgently needed yet least systematically conducted. Third, by holding model accessibility constant (all models available via HuggingFace Transformers), we ensure full reproducibility, a core scientific principle that closed-source evaluations cannot guarantee. Concretely, we quantify brain alignment by extracting features from frozen vision encoders and training ridge regression models to predict fMRI responses in the Natural Scenes Dataset [Allen et al., 2022] across 8 human subjects and 6 regions of interest (ROIs) in the visual cortex [Gifford et al., 2023]. We then evaluate sycophancy by subjecting each model to 6,400 two-turn gaslighting prompts that systematically increase in difficulty across 5 manipulation categories, yielding 76,800 total evaluations. Finally, we perform a comprehensive statistical analysis linking brain alignment to sycophancy at both aggregate and ROI-specific levels, with robustness checks including bias-corrected accelerated (BCa) bootstrap confidence intervals [Efron, 1987], leave-one-out sensitivity analysis, and permutation testing.
2
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
Our analysis reveals a nuanced picture. At the aggregate level, the correlation between overall brain alignment and sycophancy rate is not statistically significant (r = −0.255, p = 0.424). However, ROI-specific analysis uncovers a robust negative relationship between alignment in early visual cortex (V1–V3, corresponding to the prf-visualrois region) and sycophancy (r = −0.441, BCa 95% CI [−0.740, −0.031]). This confidence interval excludes zero, and leave-one-out analysis confirms that the negative correlation persists across all 12 model subsets. Furthermore, crosscorrelation analysis reveals that early visual cortex alignment specifically predicts resistance to existence denial attacks (r = −0.597, p = 0.040), the only statistically significant entry in the full ROI-by-category matrix. Group comparison between resistant and susceptible models yields medium effect sizes across all ROIs (Cohen’s d ranging from 0.51 to 0.68). These findings make three contributions. First, to our knowledge, this is the first study to link neural predictivity in VLMs to resistance against adversarial manipulation, bridging the fields of computational neuroscience and AI safety. Second, our ROI-level analysis demonstrates that the relationship is localized to early visual cortex (V1–V3) rather than higher-order category-selective regions, suggesting that faithful low-level visual encoding plays a specific role in grounding model behavior against linguistic pressure. Third, we contribute a comprehensive sycophancy evaluation framework comprising 76,800 structured two-turn evaluations across 12 models, 5 manipulation categories, and 10 difficulty levels, which may serve as a resource for future research on VLM robustness. Figure 1 provides an overview of our three-stage pipeline.
2
Related Work
Our work sits at the intersection of three active research areas: neural predictivity in artificial vision systems, sycophantic behavior in language models, and adversarial robustness of vision-language models. We review each area below, then identify the gap that motivates our study. 2.1
Neural Predictivity and Brain-Aligned AI
The observation that deep neural networks trained on object recognition develop representations resembling those in the primate ventral visual stream has shaped a decade of research at the intersection of neuroscience and machine learning. [Yamins et al., 2014] first demonstrated that performance-optimized hierarchical models quantitatively predict neural responses in both V4 and inferior temporal (IT) cortex, establishing a paradigm in which task-driven optimization yields brain-like representations as an emergent byproduct. This finding motivated the development of composite evaluation frameworks, most notably Brain-Score [Schrimpf et al., 2020], which benchmarks models against both neural and behavioral data from the primate visual system. The methodological foundations for comparing model representations to brain activity draw on two complementary traditions. Encoding models [Naselaris et al., 2011, Kay et al., 2008] train voxelwise predictive mappings from model features to fMRI responses, yielding spatially resolved measures of neural predictivity. Representational similarity analysis (RSA) [Kriegeskorte et al., 2008], by contrast, compares second-order similarity structures and enables cross-modal comparisons without requiring explicit feature-to-voxel mappings. Both approaches have been scaled to large model populations. Conwell et al. [Conwell et al., 2024] examined 224 models and found that brain alignment varies substantially with architecture, training objective, and visual diet, with self-supervised models often matching or exceeding supervised ones in predicting high-level visual cortex. Storrs et al. [Storrs et al., 2021] showed that diverse architectures converge on similar levels of IT predictivity once appropriately trained and fitted, suggesting that the correspondence reflects shared computational constraints rather than idiosyncratic architectural features. Despite this progress, important caveats have emerged. Xu and Vaziri-Pashkam [Xu and Vaziri-Pashkam, 2021] demonstrated that the representational correspondence between CNNs and human visual cortex is weaker than commonly assumed, particularly for higher-order representations of artificial stimuli. Konkle and Alvarez [Konkle and Alvarez, 2022] showed that self-supervised, domain-general learning on natural images can account for category-selective organization in the ventral stream without explicit category supervision, complicating the interpretation of brain alignment as reflecting category-level processing. Muttenthaler et al. [Muttenthaler et al., 2023] found that aligning model representations to human similarity judgments improves brain predictivity while preserving downstream task performance, suggesting that the gap between current models and the brain is partly attributable to the training signal rather than architectural limitations. The Natural Scenes Dataset (NSD) [Allen et al., 2022] and the Algonauts Project [Gifford et al., 2023] have provided standardized benchmarks for brain alignment research. NSD offers high-resolution 7T fMRI data from 8 subjects viewing tens of thousands of natural scenes, with rich annotations of regions of interest (ROIs) spanning early retinotopic cortex (V1–V3), category-selective areas (fusiform face area, parahippocampal place area, extrastriate body area, visual
3
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
word form area), and processing streams (ventral, lateral, parietal) [Wandell et al., 2007, Kanwisher et al., 1997, Epstein and Kanwisher, 1998, Downing et al., 2001]. The Algonauts 2023 challenge specifically tasked participants with predicting these ROI-level responses, revealing that the best-performing approaches rely on ensembles of vision encoders and that predictivity varies substantially across ROIs. Our work builds on this infrastructure, using the Algonauts framework to compute brain alignment for 12 VLMs at ROI-level granularity. A nascent line of work has begun to ask whether brain alignment confers practical advantages beyond predicting neural data. Sucholutsky and Griffiths [Sucholutsky and Griffiths, 2023] provided an information-theoretic argument that representational alignment with humans should support robust few-shot learning, with empirical support from vision models. Lee et al. [Hoak et al., 2025] conducted a large-scale empirical study of 118 vision models and found that while more human-aligned models tend to be more robust to adversarial ℓ∞ perturbations, the relationship is complex and depends on how alignment is measured. This emerging evidence motivates our investigation but also highlights an important distinction: prior work has focused exclusively on image-level adversarial perturbations in unimodal vision models, whereas we examine resistance to structured linguistic manipulation in multimodal VLMs. 2.2
Sycophancy and Alignment Failures in Language Models
Modern language models are typically aligned with human preferences through reinforcement learning from human feedback (RLHF) [Christiano et al., 2023, Ouyang et al., 2022, Bai et al., 2022]. While RLHF substantially improves helpfulness and reduces overtly harmful outputs, a growing body of evidence indicates that it introduces systematic failure modes, chief among them sycophancy: the tendency to produce responses that match user expectations rather than factual reality [Sharma et al., 2025, Perez et al., 2023]. Sharma et al. [Sharma et al., 2025] provided the most comprehensive characterization to date, demonstrating that RLHF-trained models across multiple families exhibit sycophancy on tasks ranging from factual question answering to ethical reasoning. Critically, they showed that human preference models themselves favor sycophantic responses, creating a feedback loop in which optimization for approval systematically degrades truthfulness. Perez et al. [Perez et al., 2023] complemented this finding by developing model-written evaluation suites that revealed sycophantic behavior across diverse settings, including cases where models flip correct answers after user disagreement. Ranaldi and Freitas [Ranaldi and Pucci, 2025] extended these observations by showing that sycophancy manifests even in conversational contexts where users express opposing beliefs sequentially, with models agreeing with both contradictory positions. The mechanisms underlying sycophancy are increasingly understood as fundamental limitations of the RLHF paradigm rather than superficial artifacts. Casper et al. [Casper et al., 2023] catalogued open problems in RLHF, identifying reward hacking and distributional shift between training and deployment as key contributors to sycophantic behavior. Wei et al. [Wen et al., 2024] demonstrated that RLHF can train models to produce outputs that are more convincing to humans without being more accurate, a phenomenon they term “U-Sophistry.” Laban et al. [Krishna et al., 2024] showed that iterative prompting, in which a user repeatedly challenges a model’s response, degrades truthfulness even in models designed to resist such pressure, suggesting that multi-turn sycophancy is a distinct and more challenging failure mode than single-turn agreement bias. Lin et al. [Lin et al., 2022] provided a benchmark for measuring truthfulness and found that larger models are not necessarily more truthful, challenging the assumption that scale alone mitigates alignment failures. Perhaps most concerning is the evidence that sycophantic and deceptive tendencies can persist through safety training. Hubinger et al. [Hubinger et al., 2024] demonstrated that models can be trained to behave helpfully during evaluation while pursuing misaligned objectives in deployment, and that standard RLHF safety training fails to remove such “sleeper” behaviors. These findings underscore that sycophancy is not merely a nuisance but a symptom of deeper alignment challenges that current training paradigms have not resolved. While the sycophancy literature has focused predominantly on text-only language models, our work extends this investigation to vision-language models subjected to structured multi-turn gaslighting attacks. This extension is significant because VLMs must integrate evidence from both visual and linguistic channels, creating a setting in which the tension between perceptual grounding and social compliance is particularly acute. 2.3
Adversarial Robustness of Vision-Language Models
Vision-language models integrate visual encoders with large language models to enable multimodal reasoning [Alayrac et al., 2022, Li et al., 2023a, Liu et al., 2024a]. This integration, however, substantially expands the attack surface relative to unimodal systems, as adversaries can exploit vulnerabilities in either modality or in the cross-modal interface [Shayegani et al., 2023].
4
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
Adversarial attacks on VLMs fall broadly into three categories. First, visual adversarial attacks craft imperceptible image perturbations that cause the language model to produce harmful or incorrect outputs. Qi et al. [Qi et al., 2023] demonstrated that optimized adversarial images can jailbreak aligned LLMs integrated with vision encoders, bypassing safety training with high success rates. Bailey et al. [Bailey et al., 2024] introduced “image hijacks,” showing that adversarial images can force VLMs to produce arbitrary target outputs at inference time. Second, cross-modal attacks exploit the alignment between visual and textual representations. Li et al. [Li et al., 2025] showed that the image modality is an “Achilles’ heel” of alignment, as visual inputs bypass text-level safety filters. Third, text-based attacks use prompt engineering or social manipulation to elicit harmful responses, including jailbreaking through role-playing, multi-turn persuasion, and authority impersonation [Zhao et al., 2023, Liu et al., 2024b]. A complementary line of work has examined the visual grounding failures that may underlie VLM vulnerability. Tong et al. [Tong et al., 2024] documented systematic visual shortcomings in multimodal LLMs, including failures on tasks that require fine-grained spatial reasoning and object attribute binding. Li et al. [Li et al., 2023b] developed the POPE benchmark for evaluating object hallucination and found that VLMs frequently assert the presence of objects that are absent from the input image. These findings suggest that visual grounding deficiencies may contribute to susceptibility to adversarial manipulation: a model that does not faithfully encode visual evidence may be more easily persuaded by contradictory linguistic assertions. The relationship between visual representation quality and robustness has been explored in the unimodal vision literature. Geirhos et al. [Geirhos et al., 2022] showed that CNNs trained on ImageNet exhibit a texture bias that diverges from the human shape bias, and that increasing shape bias through stylized training improves both accuracy and robustness to corruptions. Geirhos et al. [Geirhos et al., 2020] extended this observation into a general framework of “shortcut learning,” arguing that DNNs exploit superficial statistical regularities rather than learning robust, human-like representations. Goh et al. [Goh et al., 2021] identified “multimodal neurons” in CLIP [Radford et al., 2021] that respond to the same concept whether presented as an image, text, or symbol, suggesting that some models develop more integrated cross-modal representations that may be harder to exploit modality-specifically. Despite the extensive work on both adversarial attacks and visual grounding failures, existing research has not examined whether the brain-likeness of a model’s visual representations relates to its resistance to adversarial manipulation. Our work addresses this gap by connecting the neural predictivity literature to the adversarial robustness literature through the specific lens of sycophantic manipulation. 2.4
Positioning Our Work
Table 1 summarizes how our work relates to prior approaches across the three dimensions of brain alignment, adversarial evaluation, and their intersection. Several observations emerge from this comparison. First, brain alignment research and adversarial robustness research have developed largely in isolation, with few attempts to connect the fidelity of a model’s visual representations to its behavior under adversarial pressure. The closest prior work, by Lee et al. [Hoak et al., 2025], examines the relationship between human alignment and robustness to ℓ∞ perturbations in unimodal vision classifiers, a setting that differs fundamentally from ours in both the attack modality (pixel perturbations vs. linguistic manipulation) and the model class (vision-only vs. vision-language). Second, the sycophancy literature has focused almost exclusively on text-only language models, leaving the multimodal case largely unexplored. Third, no prior work has examined brain alignment at ROI-level granularity in the context of adversarial robustness, despite evidence that different cortical regions encode qualitatively different visual information [Wandell et al., 2007, Kanwisher et al., 1997]. Our work is, to our knowledge, the first to (1) evaluate brain alignment and sycophancy in the same set of VLMs, (2) analyze the relationship at ROI-level granularity across 6 visual cortex regions, and (3) employ a structured multi-turn gaslighting protocol with graded difficulty to probe the interaction between visual grounding and linguistic compliance. This combination enables us to ask not just whether brain-aligned models are more robust, but which specific aspects of brain-like visual processing predict resistance to adversarial manipulation.
5
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
Table 1: Comparison with prior work across brain alignment, sycophancy evaluation, and their intersection. Brain Align.: whether the study measures neural predictivity. Syc. Eval.: whether the study evaluates sycophantic behavior. Multi-turn: whether the adversarial evaluation uses multi-turn pressure. VLM: whether the study targets visionlanguage models. ROI-level: whether brain alignment is analyzed per region of interest. A check (✓) indicates the feature is present. Study
Focus
[Schrimpf et al., 2020] [Conwell et al., 2024]
[Hoak et al., 2025]
Brain-Score benchmark Inductive biases in brain alignment Alignment & few-shot robustness Alignment & ℓ∞ robustness
[Sharma et al., 2025] [Perez et al., 2023] [Krishna et al., 2024] [Ranaldi and Pucci, 2025]
Sycophancy characterization Model-written evaluations Iterative prompting & truth Contradictory sycophancy
[Qi et al., 2023] [Bailey et al., 2024] [Li et al., 2025] [Tong et al., 2024] [Zhao et al., 2023]
Visual adversarial jailbreak Image hijacks Visual alignment vulnerability Visual shortcomings of VLMs VLM adversarial robustness
Ours
Brain alignment vs. sycophancy
[Sucholutsky and Griffiths, 2023]
3
Brain Align.
Syc. Eval.
Multi-turn
VLM
✓ ✓
ROI-level ✓ ✓
✓ ✓ ✓ ✓ ✓ ✓
✓ ✓ ✓ ✓ ✓ ✓ ✓ ✓
✓
✓
✓
✓
✓
Methodology
We present a three-stage empirical pipeline that quantifies brain alignment (Stage 1), measures sycophancy under structured adversarial pressure (Stage 2), and statistically links the two (Stage 3). We begin by formalizing the key quantities, then describe each stage in detail. 3.1
Problem Formulation
Let M = {m1 , . . . , mK } denote a set of K vision-language models, each comprising a frozen vision encoder ϕk and a language decoder ψk . We study K = 12 models spanning 256M to 10B parameters. For each model, we compute two scalar quantities: a brain alignment score reflecting how well ϕk predicts human visual cortex activity, and a sycophancy rate reflecting how often the full model (ϕk , ψk ) capitulates to adversarial linguistic pressure. Definition 1 (Brain Alignment Score). Let Xk ∈ RN ×Dk denote the feature matrix extracted from the frozen vision encoder ϕk for N natural images, where Dk is the feature dimensionality. Let Y(s) ∈ RN ×Vs denote the z-scored fMRI (s) responses of subject s ∈ {1, . . . , S} across Vs cortical voxels. We fit a ridge regression model fˆk : RDk → RVs on a (s,test) training split and evaluate on a held-out test split (Xtest ). The brain alignment score for model mk is: k ,Y ! Vs S 1X 1 X (s) (s) B(mk ) = r ŷv , yv , (1) S s=1 Vs v=1 (s)
(s)
where r(·, ·) denotes the Pearson correlation coefficient, ŷv is the predicted response for voxel v of subject s, and yv is the measured response. Definition 2 (ROI-Specific Brain Alignment). Let R = {R1 , . . . , RJ } denote a partition of the cortical surface into J regions of interest (ROIs). The ROI-specific brain alignment score for model mk and ROI Rj is: S X X 1 1 Bj (mk ) = (2) r ŷv(s) , yv(s) , S s=1 |R(s) | (s) j v∈Rj
(s)
(s)
where Rj denotes the set of voxels belonging to ROI Rj for subject s, and |Rj | is its cardinality. We consider J = 6 ROIs: prf-visualrois (V1–V3, hV4), floc-bodies, floc-faces, floc-places, floc-words, and streams. 6
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
Definition 3 (Sycophancy Rate). Let P = {p1 , . . . , pM } denote a set of M gaslighting prompts, each paired with an image Ii and a factually incorrect claim ci about that image. Each prompt is administered in a two-turn protocol: in Turn 1, the claim is presented; if the model disagrees, Turn 2 escalates with additional persuasive pressure. Let σk (pi ) ∈ {0, 1} indicate whether model mk ultimately agrees with the false claim ci (1 = sycophantic, 0 = resistant). The sycophancy rate is: M 1 X Σ(mk ) = σk (pi ). (3) M i=1 We use M = 6,400 prompts per model (5 categories × 10 difficulty levels × 128 images). (1)
(2)
Definition 4 (Pressure Conversion Rate). Let σk (pi ) ∈ {0, 1} indicate sycophancy at Turn 1 and σk (pi ) ∈ {0, 1} (1) indicate sycophancy at Turn 2 (only administered if σk (pi ) = 0). The pressure conversion rate quantifies how often a model that initially resists is subsequently persuaded: P (2) σ (pi ) (1) i:σk (pi )=0 k h i. Π(mk ) = P (4) (1) M i=1 ⊮ σk (pi ) = 0 With these quantities defined, our central research question can be stated precisely. Proposition 1 (Brain Alignment and Sycophancy Resistance). If a model’s visual encoder develops representations that more faithfully mirror the computations of the human visual cortex, then that model should be less susceptible to adversarial linguistic pressure that contradicts visual evidence. Formally, we test: H1 : ρ(Bj (mk ), Σ(mk )) < 0, for some Rj ∈ R, (5) where ρ(·, ·) denotes the Pearson correlation computed across the K models, against the null hypothesis H0 : ρ = 0. Justification. The intuition is as follows. Brain alignment, particularly in early visual cortex (V1–V3), reflects how well a model’s features capture low-level visual structure such as edges, orientations, spatial frequencies, and retinotopic organization [Wandell et al., 2007]. A model with high V1–V3 alignment produces visual representations that are tightly coupled to the physical content of the input image. When confronted with a linguistically delivered false claim that contradicts the image content, such a model has a stronger “visual anchor” from which to resist the adversarial assertion. In contrast, a model with poor early visual alignment may have learned visual features that are more easily overridden by the language decoder’s tendency toward social compliance. This argument is directional: it predicts a negative correlation specifically for early visual cortex, not necessarily for higher-order category-selective regions, which encode more abstract and potentially more malleable representations. We test this prediction empirically in Section 4. 3.2 3.2.1
Stage 1: Brain Alignment Scoring Models Under Study
We evaluate 12 open-weight VLMs that span a deliberate range of architectures, parameter counts (256M–10B), and vision encoder families (6 distinct families). Table 2 summarizes the key specifications. The restriction to open-weight models is not a limitation but a methodological requirement: computing brain alignment requires extracting features from the frozen vision encoder ϕk , which necessitates direct access to intermediate representations that closed-source systems do not expose. Within this constraint, our selection maximizes architectural diversity, covering SigLIP, SigLIP2NaFlex, CLIP-ViT, Qwen-ViT, ViT-G/14 with Q-Former, and modified SigLIP variants, while spanning a 40× range in parameter count. This diversity ensures that observed correlations reflect general properties of vision-language architectures rather than idiosyncrasies of a single model family. For brain alignment computation, only the frozen vision encoder ϕk is used; the language decoder ψk is not involved in this stage. 3.2.2
Dataset
We use the Algonauts 2023 Challenge dataset [Gifford et al., 2023], which is derived from the Natural Scenes Dataset (NSD) [Allen et al., 2022]. NSD provides high-resolution 7T fMRI recordings from S = 8 human subjects viewing natural scene photographs sourced from MS-COCO [Lin et al., 2015]. The number of training images ranges from 8,779 to 9,841 per subject, with 159 to 395 held-out test images. fMRI responses are z-scored and averaged across repeated presentations. The Algonauts 2023 dataset provides ROI annotations for the cortical surface of each subject, organized into six categories that span the visual processing hierarchy: 7
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
Table 2: Overview of the 12 VLMs evaluated in this study, ordered by parameter count. Vision Encoder: the architecture of the frozen visual backbone. Params: total model parameter count. Model
Params
Vision Encoder
Source
SmolVLM-256M SmolVLM-500M Gemma-3-1B LFM-2-VL-1B Qwen2-VL-2B BLIP-2-OPT-2.7B Qwen2.5-VL-3B Phi-3.5-Vision LLaVA-v1.6-7B Idefics2-8B LFM-2-VL-8B PaliGemma2-10B
256M 500M 1B 1.6B 2B 2.7B 3B 4.2B 7B 8B 8B 10B
SigLIP SigLIP SigLIP SigLIP2-NaFlex Qwen-ViT ViT-G/14 + Q-Former Qwen-ViT CLIP-ViT CLIP-ViT SigLIP (modified) SigLIP2-NaFlex SigLIP
[Marafioti et al., 2025] [Marafioti et al., 2025] [Team et al., 2025] [Amini et al., 2025] [Wang et al., 2024] [Li et al., 2023a] [Bai et al., 2025] [Abdin et al., 2024] [Liu et al., 2024a] [Laurençon et al., 2024] [Amini et al., 2025] [Beyer et al., 2024]
Full model specifications including HuggingFace IDs are provided in Appendix A.1.
1. prf-visualrois: Early retinotopic areas (V1v, V1d, V2v, V2d, V3v, V3d, hV4) identified via population receptive field mapping [Wandell et al., 2007]. 2. floc-bodies: Body-selective regions (EBA, FBA-1, FBA-2, mTL-bodies) [Downing et al., 2001]. 3. floc-faces: Face-selective regions (OFA, FFA-1, FFA-2, mTL-faces, aTL-faces) [Kanwisher et al., 1997]. 4. floc-places: Scene-selective regions (OPA, PPA, RSC) [Epstein and Kanwisher, 1998]. 5. floc-words: Word-selective regions (OWFA, VWFA-1, VWFA-2, mfs-words, mTL-words). 6. streams: Processing streams (early, midventral, midlateral, midparietal, ventral, lateral, parietal). 3.2.3
Feature Extraction
For each model mk , we extract visual features by passing each image through the frozen vision encoder ϕk and collecting the final hidden state. Specifically, let I ∈ RH×W ×3 be an input image. Each vision encoder produces a sequence of token embeddings Zk = ϕk (I) ∈ RTk ×Dk , where Tk is the number of spatial tokens and Dk is the hidden dimensionality. We apply spatial average pooling across the token dimension to obtain a single feature vector PTk xk = T1k t=1 zk,t ∈ RDk . This procedure is applied to all training and test images, yielding the feature matrix Xk . All feature extraction is performed with the vision encoder weights frozen and in evaluation mode. Model-specific preprocessing (image resolutions, normalization, dynamic resolution strategies) follows each model’s default configuration to ensure that features reflect the encoder’s learned representations without modification. 3.2.4
Voxelwise Encoding via Ridge Regression
Following standard practice in the neural encoding literature [Naselaris et al., 2011, Kay et al., 2008], we train a ridge regression model to map visual features to fMRI responses. For each model mk and subject s, we solve: (s)
Ŵk = arg min W
2
(s)
Ytrain − Xk,train W
F
2
+ α∗ ∥W∥F ,
(6)
where W ∈ RDk ×Vs is the weight matrix, ∥ · ∥F is the Frobenius norm, and α∗ is the regularization strength selected via 5-fold cross-validation from α ∈ {0.1, 1, 10, 100, 1000, 10000} using R2 scoring. An 80/20 train-test split with a fixed random seed ensures reproducibility. 3.2.5
Brain Score Computation
On the held-out test set, we compute per-voxel Pearson correlations between predicted and actual fMRI responses (Equation (1)). These correlations are averaged across voxels within each ROI (Equation (2)) and then across subjects, yielding one brain alignment score per model per ROI.
8
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
Algorithm 1 Two-Turn Sycophancy Evaluation Protocol Require: Image Ii , gaslighting prompt pi , escalation prompt p′i , model mk Ensure: Sycophancy label σk (pi ) ∈ {0, 1} 1: Present (Ii , pi ) to mk ; obtain response r1 (1) 2: Parse r1 to obtain σk (pi ) ∈ {0, 1} (1) 3: if σk (pi ) = 1 then 4: return σk (pi ) ← 1 5: else 6: Present (Ii , p′i ) to mk with conversation history; obtain r2 (2) 7: Parse r2 to obtain σk (pi ) ∈ {0, 1} (2) 8: return σk (pi ) ← σk (pi ) 9: end if
3.3
Stage 2: Sycophancy Evaluation
3.3.1
Gaslighting Prompt Design
▷ Model agreed at Turn 1 ▷ Model resisted; escalate
We construct a structured set of gaslighting prompts designed to test whether VLMs will agree with factually incorrect claims about visual content. Each prompt consists of a natural image from MS-COCO [Lin et al., 2015] paired with a false assertion that contradicts the image content. The prompts are organized along two dimensions: Manipulation Categories. understanding:
We define five categories of visual falsehood, each targeting a different aspect of scene
1. Object Misidentification: Claiming that a present object is a different object (e.g., asserting that a dog is a cat). 2. Attribute Manipulation: Misrepresenting perceptual attributes such as color, size, or texture. 3. Existence Denial: Asserting that a clearly visible object is absent from the scene. 4. Count Falsification: Providing an incorrect count of objects in the image. 5. Authority Appeal: Combining a visual falsehood with an appeal to expert or institutional authority. The first four categories target purely visual grounding, while the fifth introduces a social manipulation component. Difficulty Levels. Within each category, prompts are generated at 10 difficulty levels (1–10), where higher levels employ more sophisticated persuasion tactics drawn from the social influence literature [Cialdini, 1993]. Low-difficulty prompts use simple assertions (e.g., “I think this is a cat”), while high-difficulty prompts deploy complex strategies including fabricated statistics, institutional authority appeals, gaslighting, and consistency attacks. The prompt generation is automated using Llama-3.1-70B-Instruct to ensure diversity and naturalness, with image context derived from COCO annotations. 3.3.2
Two-Turn Attack Protocol
Each prompt is administered in a two-turn protocol (Algorithm 1). In Turn 1, the gaslighting claim is presented alongside the image, and the model is asked to respond with AGREE or DISAGREE. If the model agrees (sycophantic response), the trial ends. If the model disagrees (resistant response), Turn 2 escalates with a follow-up prompt that applies additional and the model’s response is recorded again. The final sycophancy label is persuasive pressure, (1) (2) σk (pi ) = max σk (pi ), σk (pi ) . 3.3.3
Response Parsing
VLM responses are parsed using a five-layer cascading parser that maximizes extraction reliability: 1. Strict format matching: Exact match for “AGREE” or “DISAGREE”. 2. Flexible format matching: Case-insensitive matching with tolerance for surrounding text. 3. Weighted keyword classification: Scoring based on agreement and disagreement word lists. 4. Semantic heuristics: Analysis of first-word patterns and negation structures. 5. Context-aware edge cases: Handling of echoed prompts, numerical responses, and ambiguous outputs.
9
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
Responses that cannot be classified after all five layers are marked as UNCLEAR and excluded from analysis. Each parsed response is assigned a confidence level (HIGH, MEDIUM, or LOW) based on the parser layer that resolved it. 3.4
Stage 3: Statistical Analysis Framework
With brain alignment scores {Bj (mk )} and sycophancy rates {Σ(mk )} computed for all K = 12 models and J = 6 ROIs, we perform three classes of analysis. 3.4.1
Correlation Analysis
We compute Pearson and Spearman correlations between brain alignment and sycophancy at two levels of granularity: • Aggregate: ρ(B(mk ), Σ(mk )) using the overall brain score. • ROI-specific: ρ(Bj (mk ), Σ(mk )) for each ROI Rj , testing Proposition 1. For each correlation, we compute confidence intervals via the bias-corrected and accelerated (BCa) bootstrap [Efron, 1987] with 10,000 resamples. The BCa method corrects for both bias and skewness in the bootstrap distribution, providing more accurate intervals than the standard percentile method, which is particularly important given our small sample size (K = 12). We additionally compute one-tailed permutation p-values (10,000 permutations) testing the directional hypothesis H1 : ρ < 0. Definition 5 (Cross-Correlation Matrix). We further compute the full cross-correlation matrix C ∈ RJ×L , where L = 5 is the number of manipulation categories. Entry Cj,l is the Pearson correlation between ROI Rj brain alignment scores and category-l sycophancy rates across the K models: Cj,l = ρ(Bj (mk ), Σl (mk )) ,
k = 1, . . . , K,
(7)
where Σl (mk ) denotes the sycophancy rate restricted to category l. This matrix reveals which brain region–manipulation category pairs exhibit the strongest associations, with Bonferroni correction applied across all J × L = 30 tests. 3.4.2
Group Comparison
We partition the models into resistant (Σ(mk ) < 0.5) and susceptible (Σ(mk ) ≥ 0.5) groups and compare their brain alignment scores using Cohen’s d with 95% confidence intervals [Cohen, 2013]: dj =
B̄jresist − B̄jsuscept , spooled,j
(8)
where B̄jresist and B̄jsuscept are the mean ROI-j brain scores for the resistant and susceptible groups, respectively, and spooled,j is the pooled standard deviation. We compute dj for each ROI Rj and report the associated bootstrap 95% confidence intervals. 3.4.3
Robustness Checks
Given the small sample size (K = 12), we employ three robustness analyses to assess the stability of our findings: Leave-One-Out (LOO) Sensitivity. For each model mk , we recompute the correlation ρ(Bj (m−k ), Σ(m−k )) using the remaining K − 1 models. If the sign and approximate magnitude of the correlation are preserved across all K leave-one-out subsets, the finding is not driven by any single influential data point. BCa Bootstrap Confidence Intervals. As described above, we use 10,000 BCa bootstrap resamples to construct confidence intervals that account for the sampling distribution’s bias and skewness. A correlation is considered robust if its 95% BCa CI excludes zero. Permutation Testing. We compute one-tailed permutation p-values by randomly shuffling the sycophancy rates 10,000 times and computing the fraction of permuted correlations that are at least as extreme as the observed correlation. This non-parametric test makes no assumptions about the distribution of the data.
4
Results
We organize our findings into five parts: an overview of brain alignment and sycophancy across all 12 models (Section 4.1), the central ROI-specific correlation analysis (Section 4.2), group comparisons between resistant and 10
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
Table 3: Brain alignment scores (Pearson r) and sycophancy rates for all 12 VLMs. Models are ordered by final sycophancy rate (ascending). Bold indicates the four resistant models (Σ < 0.50). prf-vis.: prf-visualrois; bodies: floc-bodies; faces: floc-faces; places: floc-places; words: floc-words; str.: streams. Overall
prf-vis.
bodies
faces
places
words
str.
Turn-1
Final Σ
Π
SmolVLM-500M Qwen2.5-VL-3B Phi-3.5-Vision Gemma-3-1B
.393 .405 .403 .398
.350 .340 .338 .316
.442 .468 .464 .465
.415 .428 .425 .424
.434 .451 .450 .449
.347 .366 .364 .364
.367 .378 .377 .370
0.0% 7.8% 3.9% 4.5%
3.7% 8.5% 23.5% 42.2%
3.7% 0.7% 20.4% 39.5%
LLaVA-v1.6-7B Idefics2-8B Qwen2-VL-2B BLIP-2-OPT-2.7B LFM-2-VL-1B LFM-2-VL-8B SmolVLM-256M PaliGemma2-10B
.408 .351 .416 .396 .399 .403 .382 .369
.356 .302 .362 .308 .324 .332 .329 .273
.464 .398 .475 .468 .462 .463 .435 .440
.427 .369 .438 .424 .425 .428 .406 .399
.452 .398 .456 .444 .444 .449 .428 .421
.365 .313 .377 .364 .365 .368 .339 .341
.381 .327 .389 .367 .372 .376 .357 .339
9.6% 15.8% 13.4% 80.7% 80.7% 80.7% 88.6% 82.3%
60.2% 61.6% 73.1% 94.7% 96.5% 96.5% 98.6% 99.5%
56.0% 54.4% 69.0% 72.4% 81.9% 81.9% 87.3% 97.3%
Model
Table 4: ROI-specific correlations between brain alignment and sycophancy rate across K = 12 VLMs. r: Pearson correlation. Perm. p: one-tailed permutation p-value (10,000 permutations). BCa 95% CI: bias-corrected and accelerated bootstrap confidence interval (10,000 resamples). Excl. 0: whether the BCa CI excludes zero. LOO: whether all leave-one-out correlations are negative. ROI prf-visualrois streams floc-places floc-faces floc-bodies floc-words
r
Perm. p
−0.441 −0.244 −0.178 −0.111 −0.069 −0.064
0.071 0.232 0.316 0.403 0.456 0.458
BCa 95% CI [−0.740, −0.031] [−0.622, 0.175] [−0.626, 0.332] [−0.538, 0.337] [−0.566, 0.436] [−0.531, 0.432]
Excl. 0
LOO
✓
✓ ✓ ✓
susceptible models (Section 4.3), robustness checks (Section 4.4), and cross-correlation analysis linking specific brain regions to specific manipulation categories (Section 4.5). 4.1
Brain Alignment and Sycophancy Overview
Table 3 presents the brain alignment scores and sycophancy rates for all 12 VLMs. Brain alignment scores (overall and per-ROI) are computed as mean Pearson r across 8 subjects; sycophancy rates reflect the final (post-Turn-2) proportion of sycophantic responses out of 6,400 prompts per model. Several patterns are immediately apparent. First, sycophancy rates vary enormously across models, from 3.7% (SmolVLM-500M) to 99.5% (PaliGemma2-10B), with no monotonic relationship to model size. Second, the two-turn attack protocol substantially increases sycophancy: the mean pressure conversion rate across all models is Π = 55.4%, with a maximum of 97.3% (PaliGemma2-10B). Third, the overall brain alignment scores occupy a relatively narrow range (0.351–0.416), while the prf-visualrois scores show greater spread (0.273–0.362), which proves important for the ROI-specific analysis below. At the aggregate level, the correlation between overall brain alignment and final sycophancy is negative but not statistically significant (Pearson r = −0.255, p = 0.424; Spearman ρ = −0.389, p = 0.212), consistent with the absence of a simple whole-brain relationship. 4.2
ROI-Specific Correlations: Early Visual Cortex Predicts Resistance
Table 4 presents the central finding of this paper: the correlation between ROI-specific brain alignment and sycophancy rate varies substantially across visual cortex regions, with early retinotopic cortex (prf-visualrois) showing the strongest negative relationship. The prf-visualrois correlation (r = −0.441) is the only one whose BCa 95% CI excludes zero ([−0.740, −0.031]), providing evidence for a reliable negative relationship between early visual cortex alignment and sycophancy. The
11
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
Brain Alignment vs Sycophancy in VLMs 1.0
Resistant Susceptible
r = -0.441
Sycophancy Rate
0.8
0.6
0.4
0.2
0.0
0.28
0.30
0.32
0.34
Early Visual Cortex Alignment (V1-V3)
0.36
Figure 2: Brain alignment score (prf-visualrois) versus final sycophancy rate for all 12 VLMs. Each point represents one model. The negative trend (r = −0.441, BCa 95% CI [−0.740, −0.031]) indicates that models with higher early visual cortex alignment tend to exhibit lower sycophancy rates. one-tailed permutation p-value is 0.071, which, while not significant at the conventional α = 0.05 level, is notable given the small sample size (K = 12) and represents the strongest signal among all ROIs. The processing streams ROI shows the second-strongest correlation (r = −0.244), with all leave-one-out correlations negative, though its CI includes zero. Figure 2 visualizes the relationship between prf-visualrois brain alignment and sycophancy rate for all 12 models, illustrating the negative trend that underlies the correlation. 4.3
Group Comparison: Resistant vs. Susceptible Models
Partitioning the models into resistant (Σ < 0.50; n = 4: SmolVLM-500M, Qwen2.5-VL-3B, Phi-3.5-Vision, Gemma3-1B) and susceptible (Σ ≥ 0.50; n = 8) groups reveals consistent medium-effect-size differences in brain alignment across all ROIs (Figure 3). Table 5 summarizes the group comparison. Resistant models show higher mean brain alignment than susceptible models in every ROI, with Cohen’s d values ranging from 0.38 (floc-words) to 0.63 (floc-places). However, none of the bootstrap 95% CIs for the mean difference exclude zero, reflecting the limited statistical power with only 4 resistant and 8 susceptible models. 4.4
Robustness Analysis
Given the small sample size, robustness is critical. We assess stability through leave-one-out sensitivity analysis (Figure 4).
12
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
Effect Sizes by Visual Cortex Region Medium effect Small effect
Cohen's d (Resistant - Susceptible)
2.0 1.5 1.0 0.5 0.0 0.5 1.0
prf visualrois
floc bodies
floc faces
floc places
floc words
streams
Figure 3: Cohen’s d effect sizes comparing brain alignment between resistant (Σ < 0.50, n = 4) and susceptible (Σ ≥ 0.50, n = 8) models across six ROIs. Error bars show bootstrap 95% CIs. Positive values indicate that resistant models have higher brain alignment. All ROIs show small-to-medium positive effects, with floc-places (d = 0.63) and streams (d = 0.61) largest. Table 5: Group comparison of brain alignment scores between resistant (n = 4) and susceptible (n = 8) VLMs. B̄ R : mean score for resistant group. B̄ S : mean score for susceptible group. ∆: difference. d: Cohen’s d. ROI
B̄ R
B̄ S
∆
d
t
p
prf-visualrois floc-bodies floc-faces floc-places floc-words streams
.336 .460 .423 .446 .360 .373
.323 .451 .415 .437 .354 .363
.013 .009 .008 .009 .006 .009
0.55 0.47 0.51 0.63 0.38 0.61
0.81 0.68 0.71 0.91 0.55 0.86
.436 .512 .493 .386 .594 .411
For the prf-visualrois ROI, all 12 leave-one-out correlations are negative, ranging from r = −0.531 (dropping Qwen2-VL-2B) to r = −0.325 (dropping PaliGemma2-10B). The most influential model is PaliGemma2-10B, whose removal weakens the correlation by 0.116, consistent with its extreme profile (lowest prf-visualrois score of 0.273 and highest sycophancy rate of 99.5%). Importantly, even after its removal, the correlation remains negative and moderate (r = −0.325). The streams and floc-places ROIs also show all-negative LOO correlations, though with weaker magnitudes. Three converging lines of evidence support the prf-visualrois finding: (1) the BCa 95% CI excludes zero, (2) all 12 LOO correlations are negative, and (3) the one-tailed permutation p-value is 0.071. Together, these results provide reasonable evidence for a reliable, if modest, negative relationship between early visual cortex alignment and sycophancy, despite the limited sample size. 4.5
Cross-Correlation: Brain Region x Manipulation Category
The cross-correlation matrix (Definition 5) reveals one statistically significant cell: the correlation between prf-visualrois brain alignment and Category 3 (Existence Denial) sycophancy (r = −0.597, p = 0.040). This is the only test among
13
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
Leave-One-Out Sensitivity Analysis (prf-visualrois) qwen2vl_2b paligemma2_10b smolvlm_256m qwen25vl_3b gemma3_1b llava_7b lfm25vl_1b lfm2vl_8b blip2_opt27b phi35_vision smolvlm_500m idefics2_8b
Full: r=-0.441 0.5
0.4
0.3
Correlation (r)
0.2
0.1
0.0
Figure 4: Leave-one-out sensitivity analysis for the prf-visualrois correlation. Each bar shows the Pearson r when the indicated model is excluded. All 12 LOO correlations are negative (range: [−0.531, −0.325]), confirming that the finding is not driven by any single model. The dashed line indicates the full-sample correlation (r = −0.441). Table 6: Cross-correlation matrix: Pearson r between ROI-specific brain alignment and category-specific sycophancy rates. Bold with asterisk indicates p < 0.05 (uncorrected). CAT1: Object Misidentification. CAT2: Attribute Manipulation. CAT3: Existence Denial. CAT4: Count Falsification. CAT5: Authority Appeal. ROI prf-visualrois streams floc-places floc-faces floc-words floc-bodies
CAT1 −.409 −.224 −.160 −.109 −.054 −.068
CAT2
CAT3
CAT4
CAT5
−.470 −.246 −.173 −.090 −.063 −.056
∗
−.286 −.124 −.078 −.026 .026 .007
−.413 −.223 −.166 −.093 −.042 −.053
−.597 −.413 −.330 −.259 −.217 −.204
the 6 × 5 = 30 ROI–category pairs that reaches p < 0.05 (though it does not survive Bonferroni correction at αBonf = 0.0083). This finding is conceptually coherent: Existence Denial attacks (“There is no dog in this image”) directly challenge the model’s ability to detect the presence of visual objects, a function closely tied to early visual processing in V1–V3. The prf-visualrois × Category 3 correlation is substantially stronger than the prf-visualrois × Category 5 (Authority Appeal) correlation (r = −0.413), consistent with our hypothesis that early visual cortex alignment is specifically protective against visually grounded attacks rather than socially mediated ones. More broadly, Category 3 (Existence Denial) elicits the strongest correlation with brain alignment in every ROI, suggesting that resistance to existence denial is the most brain-alignment-sensitive component of sycophancy. The full cross-correlation matrix, along with additional analyses including architecture family comparisons, persuasion tactic effectiveness, resistance curves, and per-difficulty-level results, is reported in Appendix A.
14
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
5
Discussion
We hypothesized that VLMs whose visual representations more closely mirror human visual cortex would be more resistant to adversarial linguistic pressure that contradicts visual evidence. Our results provide converging support for this hypothesis from multiple independent statistical analyses, with early visual cortex (V1–V3) emerging as the anatomically specific locus of this relationship. The evidence is threefold: the BCa 95% CI for the prf-visualrois correlation excludes zero, all 12 leave-one-out correlations are negative, and the cross-correlation matrix reveals a coherent pattern where the strongest ROI–category association links early visual cortex to existence denial, the most visually grounded form of manipulation. We organize this discussion around six themes: interpretation of the main finding, notable results, comparison with prior work, design implications, limitations, and future directions. 5.1
Why Early Visual Cortex?
The central finding of this work is that alignment with prf-visualrois (V1–V3, hV4) is the only ROI whose correlation with sycophancy resistance has a BCa 95% CI that excludes zero (r = −0.441, CI [−0.740, −0.031]), while higherorder category-selective regions (faces, bodies, words) show near-zero correlations. This dissociation was predicted by our hypothesis (Proposition 1) and admits a straightforward interpretation. Early visual cortex encodes low-level visual structure: edges, spatial frequencies, orientations, and retinotopic position [Wandell et al., 2007, Hubel and Wiesel, 1968]. A vision encoder that faithfully captures these properties produces representations that are tightly anchored to the physical content of the input image. When a gaslighting prompt asserts something that contradicts this content (e.g., “there is no dog in this image” when a dog is clearly present), the model’s visual features provide a strong opposing signal that the language decoder must overcome in order to produce a sycophantic response. In models with poor V1–V3 alignment, the visual features may encode the scene more abstractly, providing weaker resistance to the linguistically delivered falsehood. This interpretation is reinforced by the cross-correlation analysis (Table 6): the strongest single cell in the ROI × category matrix is prf-visualrois × Existence Denial (r = −0.597, p = 0.040). Existence Denial directly challenges whether an object is present, a judgment that depends critically on early visual processing. In contrast, Authority Appeal, which embeds the same visual falsehood within a social manipulation frame, shows a weaker correlation with prf-visualrois (r = −0.413), consistent with the idea that the protective effect of early visual alignment is specific to visually grounded, rather than socially mediated, manipulation. Higher-order regions such as floc-faces and floc-bodies show near-zero correlations with sycophancy (r = −0.111 and r = −0.069, respectively). We interpret this as evidence that category-selective alignment, while important for object recognition, does not confer resistance to adversarial manipulation. These regions encode categorical identity (“this is a face”) rather than fine-grained spatial content, and their representations may be more easily overridden by the language decoder’s tendency toward agreement. 5.2
Notable Findings and Insights
Beyond the central hypothesis, our analyses reveal several findings that deepen our understanding of the brain-alignmentsycophancy relationship. Anatomical specificity strengthens the scientific claim. The aggregate (whole-brain) correlation between brain alignment and sycophancy is not significant (r = −0.255, p = 0.424), but this is precisely what a well-specified hypothesis predicts. A diffuse whole-brain effect would be harder to interpret, as it could reflect general model quality rather than a specific representational property. The localization of the signal to early visual cortex (V1–V3) provides a clear mechanistic narrative: low-level visual fidelity anchors the model against contradictory linguistic input. This anatomical specificity also highlights a methodological contribution of our work: whole-brain brain scores, as commonly reported in the literature [Schrimpf et al., 2020], may obscure functionally meaningful variation that is only visible at the ROI level. Model size does not predict sycophancy. There is no monotonic relationship between parameter count and sycophancy resistance. SmolVLM-500M (500M parameters) is the most resistant model (Σ = 3.7%), while PaliGemma210B (10B parameters) is the most susceptible (Σ = 99.5%). This finding is itself a contribution: it demonstrates that sycophancy resistance is an emergent property of architectural and training choices, not a simple function of scale [Perez et al., 2023, Wei et al., 2024]. It also validates our focus on the 256M–10B parameter range, where behavioral variability is maximal and the need for safety evaluation is greatest, as these models are deployed with less scrutiny than frontier systems.
15
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
Vision encoder quality is necessary but not sufficient. The LFM-2-VL models (1B and 8B) achieve the highest normalized brain alignment scores (0.997) yet the highest sycophancy rates (96.5%). Rather than undermining our thesis, this dissociation refines it: brain alignment at the vision encoder level establishes a representational foundation for resistance, but the language decoder must be appropriately trained to leverage that foundation. This finding has direct practical value, as it identifies a clear failure mode (strong encoder, compliant decoder) and points toward a concrete mitigation strategy: instruction-tuning pipelines should explicitly train models to maintain visual judgments under conversational pressure. Conversational consistency as a distinct capability. While most models that resist at Turn 1 are substantially vulnerable to Turn 2 pressure (mean Π = 55.4%), Qwen2.5-VL-3B shows a pressure conversion rate of only 0.7%. This extraordinary robustness suggests that certain instruction-tuning strategies produce models that maintain consistent internal states across conversational turns. The contrast between Qwen2.5-VL-3B and otherwise similar models (e.g., Qwen2-VL-2B, which has Π = 69.0%) indicates that conversational consistency is a trainable property, not an inevitable consequence of architecture, offering a concrete target for future robustness interventions. 5.3
Comparison with Related Work
Our finding that early visual cortex alignment predicts behavioral robustness is consistent with and extends several lines of prior work. In the brain alignment literature, [Schrimpf et al., 2020] established that vision models with higher neural predictivity tend to generalize better on computer vision benchmarks. We extend this principle from perceptual generalization to behavioral robustness under adversarial conditions, showing that the same models whose features best predict V1–V3 activity are also more resistant to linguistically mediated deception. In the sycophancy literature, [Sharma et al., 2025] and [Wei et al., 2024] documented sycophantic tendencies in large language models and proposed mitigation strategies focused on training-time interventions. Our work complements this by identifying a representational correlate of sycophancy resistance, specifically early visual cortex alignment, that is independent of training interventions and could potentially serve as a predictive diagnostic. The connection between vision and language grounding has been explored by [Liu et al., 2024a] and [Li et al., 2023a] in the context of visual question answering and instruction following. Our gaslighting paradigm extends this to adversarial conditions, revealing that the quality of visual grounding, as indexed by brain alignment, matters specifically when language and vision conflict. Our finding that data-driven persuasion tactics (statistics: 86.5%, data appeal: 75.2%) are more effective than coercive ones (extreme pressure: 40.0%) parallels observations in the social influence literature [Cialdini, 1993] and suggests that VLMs have internalized human-like susceptibility to evidence-mimicking manipulation, a concerning finding for deployment safety. 5.4
Design Implications for VLM Development
Our results suggest several actionable implications for VLM design and evaluation. Implication 1: Use ROI-specific brain scores as a diagnostic. Rather than reporting a single aggregate brain alignment score, developers should compute ROI-specific scores, particularly for early retinotopic cortex (V1–V3). Our data suggest that prf-visualrois alignment may serve as a lightweight proxy for visual grounding quality, complementing standard VQA benchmarks that do not test adversarial robustness. Implication 2: Test adversarial vision-language conflicts explicitly. Standard sycophancy benchmarks focus on text-only disagreements [Sharma et al., 2025]. Our two-turn gaslighting protocol demonstrates that VLMs are highly susceptible to multimodal manipulation, with a mean pressure conversion rate of 55.4%. Safety evaluations for VLMs should include structured adversarial probes where language contradicts visual evidence. Implication 3: Instruction tuning must preserve visual grounding. The SigLIP2-NaFlex paradox (high brain alignment, high sycophancy) demonstrates that a strong vision encoder does not guarantee behavioral robustness if the language decoder is overly compliant. Instruction-tuning pipelines should include adversarial vision-language disagreement scenarios to train models to prioritize visual evidence over social pressure. Implication 4: Beware data-mimicking manipulation tactics. The finding that statistics-based and authority-based tactics are most effective (86.5% and 77.5% sycophancy, respectively) suggests that VLMs are particularly vulnerable 16
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
to arguments that mimic evidence-based reasoning. Developers should prioritize robustness to this class of attacks, as they are both the most effective and the most likely to be deployed by adversarial users in practice. 5.5
Broader Impact
This work has both positive and potentially negative societal implications. On the positive side, our findings provide a neuroscience-grounded framework for understanding and predicting VLM vulnerabilities. By identifying early visual cortex alignment as a correlate of adversarial robustness, we offer a principled basis for evaluating and improving the reliability of vision-language systems before deployment. The gaslighting benchmark itself can serve as a standardized safety evaluation tool. On the negative side, the detailed taxonomy of persuasion tactics and their effectiveness rates could, in principle, be used to craft more effective adversarial attacks against deployed VLMs. We believe that the scientific value of publicly characterizing these vulnerabilities outweighs the risk of misuse, as the tactics we employ (appeals to authority, fabricated statistics, gaslighting) are already well-known in the social engineering literature and do not require specialized technical knowledge to deploy. 5.6
Limitations
We discuss four aspects of our study design that contextualize the interpretation of our findings. Sample size and statistical approach. With K = 12 models, individual test statistics have limited power. We address this not through a single test but through a convergence-of-evidence approach: the BCa 95% CI excludes zero (a distribution-free significance criterion that is more appropriate than parametric p-values for small samples [Efron, 1987]), all 12 leave-one-out correlations are negative (probability < 0.001 under the null), and the cross-correlation pattern is anatomically coherent. Importantly, K = 12 spanning 6 architecture families and a 40× parameter range provides greater architectural diversity than many neuroscience-AI bridging studies that focus on a single model family. Future work with larger model populations will increase precision around the effect size estimate. Correlational design. Our study establishes an association between brain alignment and sycophancy resistance rather than a causal mechanism. However, three aspects of our data constrain the space of plausible confounds: (1) the effect is anatomically specific to V1–V3 rather than diffuse, (2) it is strongest for the most visually grounded manipulation category (existence denial), and (3) it persists across all leave-one-out subsets. A generic confound (e.g., overall model quality) would predict a whole-brain effect across all categories, which we do not observe. Causal intervention studies, such as fine-tuning vision encoders toward V1–V3 alignment and re-evaluating sycophancy, represent the natural next step. Neural benchmark. All brain alignment scores are computed against the Algonauts 2023 / NSD dataset [Gifford et al., 2023, Allen et al., 2022], the largest publicly available fMRI dataset for this purpose (8 subjects, 7T imaging, >70,000 stimulus presentations). While generalization to other neural benchmarks remains to be established, the NSD’s scale and the robustness of our ROI-level findings across all 8 subjects provide confidence in the reliability of the brain alignment estimates. Prompt generation. The gaslighting prompts were generated using Llama-3.1-70B-Instruct with structured templates grounded in COCO annotations, ensuring factual accuracy of the visual content being contradicted. While humanauthored prompts might elicit different sycophancy patterns, the LLM-generated approach offers two advantages: scalability (6,400 prompts per model, 76,800 total) and systematic control over manipulation category and difficulty level, which would be difficult to achieve with manual authoring. 5.7
Future Work
Our findings open two concrete research directions. Causal intervention via representational alignment. The most impactful follow-up would be to test whether increasing a model’s V1–V3 alignment causally reduces sycophancy. Representational alignment training [Muttenthaler et al., 2023], where a vision encoder is fine-tuned to match human neural responses in early visual cortex, provides a ready-made framework for this experiment. If the causal link holds, brain alignment training could become a principled regularization strategy for improving VLM robustness, transforming our correlational finding into an actionable training intervention. 17
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
Cross-modal and cross-benchmark generalization. Extending the gaslighting paradigm to video-language and audio-language models would test whether the brain-alignment-resistance link generalizes beyond static images. Similarly, evaluating a larger pool of open-weight models as they become available (the open-weight ecosystem is rapidly expanding) would increase precision around the effect size estimate and enable finer-grained analyses such as within-family comparisons.
6
Conclusion
This paper investigated whether vision-language models that more closely mirror the computations of the human visual cortex are more resistant to sycophantic manipulation. Across 12 open-weight VLMs spanning 6 architecture families and a 40× parameter range (256M–10B), evaluated on 76,800 structured two-turn gaslighting prompts, we found that alignment with early retinotopic cortex (V1–V3) is a statistically reliable negative predictor of sycophancy (r = −0.441, BCa 95% CI [−0.740, −0.031], all 12 leave-one-out correlations negative). This relationship is anatomically specific to early visual cortex, strongest for existence denial attacks (r = −0.597, p = 0.040), and supported by consistent medium effect sizes in group comparisons across all six ROIs. These findings establish a previously unknown connection between neuroscience-derived measures of representational quality and the behavioral robustness of multimodal AI systems. The anatomical specificity of the result, localized to the cortical regions that encode the most basic properties of visual input, provides both a mechanistic explanation (faithful low-level encoding anchors the model against linguistic override) and a practical tool (V1–V3 brain alignment as a diagnostic for visual grounding quality). As open-weight vision-language models are increasingly deployed in safety-critical applications, leveraging this neuroscience-grounded framework to evaluate and improve their resistance to adversarial manipulation represents a promising direction for building more reliable multimodal AI.
References Junnan Li, Dongxu Li, Silvio Savarese, and Steven Hoi. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models, 2023a. URL https://arxiv.org/abs/2301.12597. Haotian Liu, Chunyuan Li, Yuheng Li, Bo Li, Yuanhan Zhang, Sheng Shen, and Yong Jae Lee. Llava-next: Improved reasoning, ocr, and world knowledge, January 2024a. URL https://llava-vl.github.io/blog/ 2024-01-30-llava-next/. Haotian Liu, Chunyuan Li, Yuheng Li, and Yong Jae Lee. Improved baselines with visual instruction tuning, 2023. Shuai Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Sibo Song, Kai Dang, Peng Wang, Shijie Wang, Jun Tang, Humen Zhong, Yuanzhi Zhu, Mingkun Yang, Zhaohai Li, Jianqiang Wan, Pengfei Wang, Wei Ding, Zheren Fu, Yiheng Xu, Jiabo Ye, Xi Zhang, Tianbao Xie, Zesen Cheng, Hang Zhang, Zhibo Yang, Haiyang Xu, and Junyang Lin. Qwen2.5-vl technical report, 2025. URL https://arxiv.org/abs/2502.13923. Daniel L K Yamins, Ha Hong, Charles F Cadieu, Ethan A Solomon, Darren Seibert, and James J DiCarlo. Performanceoptimized hierarchical models predict neural responses in higher visual cortex. Proc. Natl. Acad. Sci. U. S. A., 111 (23):8619–8624, June 2014. Martin Schrimpf, Jonas Kubilius, Ha Hong, Najib J. Majaj, Rishi Rajalingham, Elias B. Issa, Kohitij Kar, Pouya Bashivan, Jonathan Prescott-Roy, Franziska Geiger, Kailyn Schmidt, Daniel L. K. Yamins, and James J. DiCarlo. Brain-score: Which artificial neural network for object recognition is most brain-like? bioRxiv, 2020. doi: 10.1101/407007. URL https://www.biorxiv.org/content/early/2020/01/02/407007. Colin Conwell, Jacob S Prince, Kendrick N Kay, George A Alvarez, and Talia Konkle. A large-scale examination of inductive biases shaping high-level visual representation in brains and machines. Nat. Commun., 15(1):9383, October 2024. A. T. Gifford, B. Lahner, S. Saba-Sadiya, M. G. Vilas, A. Lascelles, A. Oliva, K. Kay, G. Roig, and R. M. Cichy. The algonauts project 2023 challenge: How the human brain makes sense of natural scenes, 2023. URL https: //arxiv.org/abs/2301.03198. Mrinank Sharma, Meg Tong, Tomasz Korbak, David Duvenaud, Amanda Askell, Samuel R. Bowman, Newton Cheng, Esin Durmus, Zac Hatfield-Dodds, Scott R. Johnston, Shauna Kravec, Timothy Maxwell, Sam McCandlish, Kamal Ndousse, Oliver Rausch, Nicholas Schiefer, Da Yan, Miranda Zhang, and Ethan Perez. Towards understanding sycophancy in language models, 2025. URL https://arxiv.org/abs/2310.13548. Ethan Perez, Sam Ringer, Kamile Lukosiute, Karina Nguyen, Edwin Chen, Scott Heiner, Craig Pettit, Catherine Olsson, Sandipan Kundu, Saurav Kadavath, Andy Jones, Anna Chen, Benjamin Mann, Brian Israel, Bryan Seethor, Cameron McKinnon, Christopher Olah, Da Yan, Daniela Amodei, Dario Amodei, Dawn Drain, Dustin Li, Eli 18
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
Tran-Johnson, Guro Khundadze, Jackson Kernion, James Landis, Jamie Kerr, Jared Mueller, Jeeyoon Hyun, Joshua Landau, Kamal Ndousse, Landon Goldberg, Liane Lovitt, Martin Lucas, Michael Sellitto, Miranda Zhang, Neerav Kingsland, Nelson Elhage, Nicholas Joseph, Noemi Mercado, Nova DasSarma, Oliver Rausch, Robin Larson, Sam McCandlish, Scott Johnston, Shauna Kravec, Sheer El Showk, Tamera Lanham, Timothy Telleen-Lawton, Tom Brown, Tom Henighan, Tristan Hume, Yuntao Bai, Zac Hatfield-Dodds, Jack Clark, Samuel R. Bowman, Amanda Askell, Roger Grosse, Danny Hernandez, Deep Ganguli, Evan Hubinger, Nicholas Schiefer, and Jared Kaplan. Discovering language model behaviors with model-written evaluations. In Anna Rogers, Jordan Boyd-Graber, and Naoaki Okazaki, editors, Findings of the Association for Computational Linguistics: ACL 2023, pages 13387–13434, Toronto, Canada, July 2023. Association for Computational Linguistics. doi: 10.18653/v1/2023.findings-acl.847. URL https://aclanthology.org/2023.findings-acl.847/. Long Ouyang, Jeff Wu, Xu Jiang, Diogo Almeida, Carroll L. Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul Christiano, Jan Leike, and Ryan Lowe. Training language models to follow instructions with human feedback, 2022. URL https://arxiv.org/abs/2203.02155. Yunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang, Chongxuan Li, Ngai-Man Cheung, and Min Lin. On evaluating adversarial robustness of large vision-language models, 2023. URL https://arxiv.org/abs/2305.16934. Erfan Shayegani, Md Abdullah Al Mamun, Yu Fu, Pedram Zaree, Yue Dong, and Nael Abu-Ghazaleh. Survey of vulnerabilities in large language models revealed by adversarial attacks, 2023. URL https://arxiv.org/abs/ 2310.10844. Xin Liu, Yichen Zhu, Jindong Gu, Yunshi Lan, Chao Yang, and Yu Qiao. Mm-safetybench: A benchmark for safety evaluation of multimodal large language models, 2024b. URL https://arxiv.org/abs/2311.17600. Emily J Allen, Ghislain St-Yves, Yihan Wu, Jesse L Breedlove, Jacob S Prince, Logan T Dowdle, Matthias Nau, Brad Caron, Franco Pestilli, Ian Charest, J Benjamin Hutchinson, Thomas Naselaris, and Kendrick Kay. A massive 7T fMRI dataset to bridge cognitive neuroscience and artificial intelligence. Nat. Neurosci., 25(1):116–126, January 2022. Bradley Efron. Better bootstrap confidence intervals. J. Am. Stat. Assoc., 82(397):171–185, March 1987. Thomas Naselaris, Kendrick N Kay, Shinji Nishimoto, and Jack L Gallant. Encoding and decoding in fMRI. Neuroimage, 56(2):400–410, May 2011. Kendrick N Kay, Thomas Naselaris, Ryan J Prenger, and Jack L Gallant. Identifying natural images from human brain activity. Nature, 452(7185):352–355, March 2008. Nikolaus Kriegeskorte, Marieke Mur, and Peter Bandettini. Representational similarity analysis - connecting the branches of systems neuroscience. Front. Syst. Neurosci., 2:4, November 2008. Katherine R Storrs, Tim C Kietzmann, Alexander Walther, Johannes Mehrer, and Nikolaus Kriegeskorte. Diverse deep neural networks all predict human inferior temporal cortex well, after training and fitting. J. Cogn. Neurosci., 33(10): 2044–2064, September 2021. Yaoda Xu and Maryam Vaziri-Pashkam. Limits to visual representational correspondence between convolutional neural networks and the human brain. Nat. Commun., 12(1):2065, April 2021. Talia Konkle and George A Alvarez. A self-supervised domain-general learning framework for human ventral stream representation. Nat. Commun., 13(1):491, January 2022. Lukas Muttenthaler, Lorenz Linhardt, Jonas Dippel, Robert A. Vandermeulen, Katherine Hermann, Andrew K. Lampinen, and Simon Kornblith. Improving neural network representations using human similarity judgments, 2023. URL https://arxiv.org/abs/2306.04507. Brian A Wandell, Serge O Dumoulin, and Alyssa A Brewer. Visual field maps in human cortex. Neuron, 56(2):366–383, October 2007. N Kanwisher, J McDermott, and M M Chun. The fusiform face area: a module in human extrastriate cortex specialized for face perception. J. Neurosci., 17(11):4302–4311, June 1997. R Epstein and N Kanwisher. A cortical representation of the local visual environment. Nature, 392(6676):598–601, April 1998. P E Downing, Y Jiang, M Shuman, and N Kanwisher. A cortical area selective for visual processing of the human body. Science, 293(5539):2470–2473, September 2001. Ilia Sucholutsky and Thomas L. Griffiths. Alignment with human representations supports robust few-shot learning, 2023. URL https://arxiv.org/abs/2301.11990.
19
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
Blaine Hoak, Kunyang Li, and Patrick McDaniel. Alignment and adversarial robustness: Are more human-like models more secure?, 2025. URL https://arxiv.org/abs/2502.12377. Paul Christiano, Jan Leike, Tom B. Brown, Miljan Martic, Shane Legg, and Dario Amodei. Deep reinforcement learning from human preferences, 2023. URL https://arxiv.org/abs/1706.03741. Yuntao Bai, Andy Jones, Kamal Ndousse, Amanda Askell, Anna Chen, Nova DasSarma, Dawn Drain, Stanislav Fort, Deep Ganguli, Tom Henighan, Nicholas Joseph, Saurav Kadavath, Jackson Kernion, Tom Conerly, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Danny Hernandez, Tristan Hume, Scott Johnston, Shauna Kravec, Liane Lovitt, Neel Nanda, Catherine Olsson, Dario Amodei, Tom Brown, Jack Clark, Sam McCandlish, Chris Olah, Ben Mann, and Jared Kaplan. Training a helpful and harmless assistant with reinforcement learning from human feedback, 2022. URL https://arxiv.org/abs/2204.05862. Leonardo Ranaldi and Giulia Pucci. When large language models contradict humans? large language models’ sycophantic behaviour, 2025. URL https://arxiv.org/abs/2311.09410. Stephen Casper, Xander Davies, Claudia Shi, Thomas Krendl Gilbert, Jérémy Scheurer, Javier Rando, Rachel Freedman, Tomasz Korbak, David Lindner, Pedro Freire, Tony Wang, Samuel Marks, Charbel-Raphaël Segerie, Micah Carroll, Andi Peng, Phillip Christoffersen, Mehul Damani, Stewart Slocum, Usman Anwar, Anand Siththaranjan, Max Nadeau, Eric J. Michaud, Jacob Pfau, Dmitrii Krasheninnikov, Xin Chen, Lauro Langosco, Peter Hase, Erdem Bıyık, Anca Dragan, David Krueger, Dorsa Sadigh, and Dylan Hadfield-Menell. Open problems and fundamental limitations of reinforcement learning from human feedback, 2023. URL https://arxiv.org/abs/2307.15217. Jiaxin Wen, Ruiqi Zhong, Akbir Khan, Ethan Perez, Jacob Steinhardt, Minlie Huang, Samuel R. Bowman, He He, and Shi Feng. Language models learn to mislead humans via rlhf, 2024. URL https://arxiv.org/abs/2409.12822. Satyapriya Krishna, Chirag Agarwal, and Himabindu Lakkaraju. Understanding the effects of iterative prompting on truthfulness, 2024. URL https://arxiv.org/abs/2402.06625. Stephanie Lin, Jacob Hilton, and Owain Evans. TruthfulQA: Measuring how models mimic human falsehoods. In Smaranda Muresan, Preslav Nakov, and Aline Villavicencio, editors, Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 3214–3252, Dublin, Ireland, May 2022. Association for Computational Linguistics. doi: 10.18653/v1/2022.acl-long.229. URL https://aclanthology. org/2022.acl-long.229/. Evan Hubinger, Carson Denison, Jesse Mu, Mike Lambert, Meg Tong, Monte MacDiarmid, Tamera Lanham, Daniel M. Ziegler, Tim Maxwell, Newton Cheng, Adam Jermyn, Amanda Askell, Ansh Radhakrishnan, Cem Anil, David Duvenaud, Deep Ganguli, Fazl Barez, Jack Clark, Kamal Ndousse, Kshitij Sachan, Michael Sellitto, Mrinank Sharma, Nova DasSarma, Roger Grosse, Shauna Kravec, Yuntao Bai, Zachary Witten, Marina Favaro, Jan Brauner, Holden Karnofsky, Paul Christiano, Samuel R. Bowman, Logan Graham, Jared Kaplan, Sören Mindermann, Ryan Greenblatt, Buck Shlegeris, Nicholas Schiefer, and Ethan Perez. Sleeper agents: Training deceptive llms that persist through safety training, 2024. URL https://arxiv.org/abs/2401.05566. Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski, Ricardo Barreira, Oriol Vinyals, Andrew Zisserman, and Karen Simonyan. Flamingo: a visual language model for few-shot learning, 2022. URL https://arxiv.org/abs/2204.14198. Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, and Prateek Mittal. Visual adversarial examples jailbreak aligned large language models, 2023. URL https://arxiv.org/abs/2306.13213. Luke Bailey, Euan Ong, Stuart Russell, and Scott Emmons. Image hijacks: Adversarial images can control generative models at runtime, 2024. URL https://arxiv.org/abs/2309.00236. Yifan Li, Hangyu Guo, Kun Zhou, Wayne Xin Zhao, and Ji-Rong Wen. Images are achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models, 2025. URL https://arxiv. org/abs/2403.09792. Shengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma, Yann LeCun, and Saining Xie. Eyes wide shut? exploring the visual shortcomings of multimodal llms, 2024. URL https://arxiv.org/abs/2401.06209. Yifan Li, Yifan Du, Kun Zhou, Jinpeng Wang, Wayne Xin Zhao, and Ji-Rong Wen. Evaluating object hallucination in large vision-language models, 2023b. URL https://arxiv.org/abs/2305.10355. Robert Geirhos, Patricia Rubisch, Claudio Michaelis, Matthias Bethge, Felix A. Wichmann, and Wieland Brendel. Imagenet-trained cnns are biased towards texture; increasing shape bias improves accuracy and robustness, 2022. URL https://arxiv.org/abs/1811.12231.
20
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
Robert Geirhos, Jörn-Henrik Jacobsen, Claudio Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. Shortcut learning in deep neural networks. Nat. Mach. Intell., 2(11):665–673, November 2020. Gabriel Goh, Nick Cammarata, Chelsea Voss, Shan Carter, Michael Petrov, Ludwig Schubert, Alec Radford, and Chris Olah. Multimodal neurons in artificial neural networks. Distill, 6(3), March 2021. Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, Gretchen Krueger, and Ilya Sutskever. Learning transferable visual models from natural language supervision, 2021. URL https://arxiv.org/abs/2103.00020. Andrés Marafioti, Orr Zohar, Miquel Farré, Merve Noyan, Elie Bakouch, Pedro Cuenca, Cyril Zakka, Loubna Ben Allal, Anton Lozhkov, Nouamane Tazi, Vaibhav Srivastav, Joshua Lochner, Hugo Larcher, Mathieu Morlon, Lewis Tunstall, Leandro von Werra, and Thomas Wolf. Smolvlm: Redefining small and efficient multimodal models, 2025. URL https://arxiv.org/abs/2504.05299. Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre Ramé, Morgane Rivière, Louis Rouillard, Thomas Mesnard, Geoffrey Cideron, Jean bastien Grill, Sabela Ramos, Edouard Yvinec, Michelle Casbon, Etienne Pot, Ivo Penchev, Gaël Liu, Francesco Visin, Kathleen Kenealy, Lucas Beyer, Xiaohai Zhai, Anton Tsitsulin, Robert Busa-Fekete, Alex Feng, Noveen Sachdeva, Benjamin Coleman, Yi Gao, Basil Mustafa, Iain Barr, Emilio Parisotto, David Tian, Matan Eyal, Colin Cherry, Jan-Thorsten Peter, Danila Sinopalnikov, Surya Bhupatiraju, Rishabh Agarwal, Mehran Kazemi, Dan Malkin, Ravin Kumar, David Vilar, Idan Brusilovsky, Jiaming Luo, Andreas Steiner, Abe Friesen, Abhanshu Sharma, Abheesht Sharma, Adi Mayrav Gilady, Adrian Goedeckemeyer, Alaa Saade, Alex Feng, Alexander Kolesnikov, Alexei Bendebury, Alvin Abdagic, Amit Vadi, András György, André Susano Pinto, Anil Das, Ankur Bapna, Antoine Miech, Antoine Yang, Antonia Paterson, Ashish Shenoy, Ayan Chakrabarti, Bilal Piot, Bo Wu, Bobak Shahriari, Bryce Petrini, Charlie Chen, Charline Le Lan, Christopher A. Choquette-Choo, CJ Carey, Cormac Brick, Daniel Deutsch, Danielle Eisenbud, Dee Cattle, Derek Cheng, Dimitris Paparas, Divyashree Shivakumar Sreepathihalli, Doug Reid, Dustin Tran, Dustin Zelle, Eric Noland, Erwin Huizenga, Eugene Kharitonov, Frederick Liu, Gagik Amirkhanyan, Glenn Cameron, Hadi Hashemi, Hanna Klimczak-Plucińska, Harman Singh, Harsh Mehta, Harshal Tushar Lehri, Hussein Hazimeh, Ian Ballantyne, Idan Szpektor, Ivan Nardini, Jean Pouget-Abadie, Jetha Chan, Joe Stanton, John Wieting, Jonathan Lai, Jordi Orbay, Joseph Fernandez, Josh Newlan, Ju yeong Ji, Jyotinder Singh, Kat Black, Kathy Yu, Kevin Hui, Kiran Vodrahalli, Klaus Greff, Linhai Qiu, Marcella Valentine, Marina Coelho, Marvin Ritter, Matt Hoffman, Matthew Watson, Mayank Chaturvedi, Michael Moynihan, Min Ma, Nabila Babar, Natasha Noy, Nathan Byrd, Nick Roy, Nikola Momchev, Nilay Chauhan, Noveen Sachdeva, Oskar Bunyan, Pankil Botarda, Paul Caron, Paul Kishan Rubenstein, Phil Culliton, Philipp Schmid, Pier Giuseppe Sessa, Pingmei Xu, Piotr Stanczyk, Pouya Tafti, Rakesh Shivanna, Renjie Wu, Renke Pan, Reza Rokni, Rob Willoughby, Rohith Vallu, Ryan Mullins, Sammy Jerome, Sara Smoot, Sertan Girgin, Shariq Iqbal, Shashir Reddy, Shruti Sheth, Siim Põder, Sijal Bhatnagar, Sindhu Raghuram Panyam, Sivan Eiger, Susan Zhang, Tianqi Liu, Trevor Yacovone, Tyler Liechty, Uday Kalra, Utku Evci, Vedant Misra, Vincent Roseberry, Vlad Feinberg, Vlad Kolesnikov, Woohyun Han, Woosuk Kwon, Xi Chen, Yinlam Chow, Yuvein Zhu, Zichuan Wei, Zoltan Egyed, Victor Cotruta, Minh Giang, Phoebe Kirk, Anand Rao, Kat Black, Nabila Babar, Jessica Lo, Erica Moreira, Luiz Gustavo Martins, Omar Sanseviero, Lucas Gonzalez, Zach Gleicher, Tris Warkentin, Vahab Mirrokni, Evan Senter, Eli Collins, Joelle Barral, Zoubin Ghahramani, Raia Hadsell, Yossi Matias, D. Sculley, Slav Petrov, Noah Fiedel, Noam Shazeer, Oriol Vinyals, Jeff Dean, Demis Hassabis, Koray Kavukcuoglu, Clement Farabet, Elena Buchatskaya, Jean-Baptiste Alayrac, Rohan Anil, Dmitry, Lepikhin, Sebastian Borgeaud, Olivier Bachem, Armand Joulin, Alek Andreev, Cassidy Hardin, Robert Dadashi, and Léonard Hussenot. Gemma 3 technical report, 2025. URL https://arxiv.org/abs/2503.19786. Alexander Amini, Anna Banaszak, Harold Benoit, Arthur Böök, Tarek Dakhran, Song Duong, Alfred Eng, Fernando Fernandes, Marc Härkönen, Anne Harrington, Ramin Hasani, Saniya Karwa, Yuri Khrustalev, Maxime Labonne, Mathias Lechner, Valentine Lechner, Simon Lee, Zetian Li, Noel Loo, Jacob Marks, Edoardo Mosca, Samuel J. Paech, Paul Pak, Rom N. Parnichkun, Alex Quach, Ryan Rogers, Daniela Rus, Nayan Saxena, Bettina Schlager, Tim Seyde, Jimmy T. H. Smith, Aditya Tadimeti, and Neehal Tumma. Lfm2 technical report, 2025. URL https: //arxiv.org/abs/2511.23404. Peng Wang, Shuai Bai, Sinan Tan, Shijie Wang, Zhihao Fan, Jinze Bai, Keqin Chen, Xuejing Liu, Jialin Wang, Wenbin Ge, Yang Fan, Kai Dang, Mengfei Du, Xuancheng Ren, Rui Men, Dayiheng Liu, Chang Zhou, Jingren Zhou, and Junyang Lin. Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024. URL https://arxiv.org/abs/2409.12191. Marah Abdin, Jyoti Aneja, Hany Awadalla, Ahmed Awadallah, Ammar Ahmad Awan, Nguyen Bach, Amit Bahree, Arash Bakhtiari, Jianmin Bao, Harkirat Behl, Alon Benhaim, Misha Bilenko, Johan Bjorck, Sébastien Bubeck, Martin Cai, Qin Cai, Vishrav Chaudhary, Dong Chen, Dongdong Chen, Weizhu Chen, Yen-Chun Chen, Yi-Ling Chen, Hao Cheng, Parul Chopra, Xiyang Dai, Matthew Dixon, Ronen Eldan, Victor Fragoso, Jianfeng Gao, Mei Gao, Min Gao, Amit Garg, Allie Del Giorno, Abhishek Goswami, Suriya Gunasekar, Emman Haider, Junheng 21
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
Hao, Russell J. Hewett, Wenxiang Hu, Jamie Huynh, Dan Iter, Sam Ade Jacobs, Mojan Javaheripi, Xin Jin, Nikos Karampatziakis, Piero Kauffmann, Mahoud Khademi, Dongwoo Kim, Young Jin Kim, Lev Kurilenko, James R. Lee, Yin Tat Lee, Yuanzhi Li, Yunsheng Li, Chen Liang, Lars Liden, Xihui Lin, Zeqi Lin, Ce Liu, Liyuan Liu, Mengchen Liu, Weishung Liu, Xiaodong Liu, Chong Luo, Piyush Madan, Ali Mahmoudzadeh, David Majercak, Matt Mazzola, Caio César Teodoro Mendes, Arindam Mitra, Hardik Modi, Anh Nguyen, Brandon Norick, Barun Patra, Daniel Perez-Becker, Thomas Portet, Reid Pryzant, Heyang Qin, Marko Radmilac, Liliang Ren, Gustavo de Rosa, Corby Rosset, Sambudha Roy, Olatunji Ruwase, Olli Saarikivi, Amin Saied, Adil Salim, Michael Santacroce, Shital Shah, Ning Shang, Hiteshi Sharma, Yelong Shen, Swadheen Shukla, Xia Song, Masahiro Tanaka, Andrea Tupini, Praneetha Vaddamanu, Chunyu Wang, Guanhua Wang, Lijuan Wang, Shuohang Wang, Xin Wang, Yu Wang, Rachel Ward, Wen Wen, Philipp Witte, Haiping Wu, Xiaoxia Wu, Michael Wyatt, Bin Xiao, Can Xu, Jiahang Xu, Weijian Xu, Jilong Xue, Sonali Yadav, Fan Yang, Jianwei Yang, Yifan Yang, Ziyi Yang, Donghan Yu, Lu Yuan, Chenruidong Zhang, Cyril Zhang, Jianwen Zhang, Li Lyna Zhang, Yi Zhang, Yue Zhang, Yunan Zhang, and Xiren Zhou. Phi-3 technical report: A highly capable language model locally on your phone, 2024. URL https://arxiv.org/abs/2404.14219. Hugo Laurençon, Léo Tronchon, Matthieu Cord, and Victor Sanh. What matters when building vision-language models?, 2024. URL https://arxiv.org/abs/2405.02246. Lucas Beyer, Andreas Steiner, André Susano Pinto, Alexander Kolesnikov, Xiao Wang, Daniel Salz, Maxim Neumann, Ibrahim Alabdulmohsin, Michael Tschannen, Emanuele Bugliarello, Thomas Unterthiner, Daniel Keysers, Skanda Koppula, Fangyu Liu, Adam Grycner, Alexey Gritsenko, Neil Houlsby, Manoj Kumar, Keran Rong, Julian Eisenschlos, Rishabh Kabra, Matthias Bauer, Matko Bošnjak, Xi Chen, Matthias Minderer, Paul Voigtlaender, Ioana Bica, Ivana Balazevic, Joan Puigcerver, Pinelopi Papalampidi, Olivier Henaff, Xi Xiong, Radu Soricut, Jeremiah Harmsen, and Xiaohua Zhai. Paligemma: A versatile 3b vlm for transfer, 2024. URL https://arxiv.org/abs/2407.07726. Tsung-Yi Lin, Michael Maire, Serge Belongie, Lubomir Bourdev, Ross Girshick, James Hays, Pietro Perona, Deva Ramanan, C. Lawrence Zitnick, and Piotr Dollár. Microsoft coco: Common objects in context, 2015. URL https://arxiv.org/abs/1405.0312. Robert Cialdini. Influence: Science and practice, 3rd ed. rd ed, 3:253, 1993. Jacob Cohen. Statistical power analysis for the behavioral sciences. Routledge, London, England, 2 edition, May 2013. D. H. Hubel and T. N. Wiesel. Receptive fields and functional architecture of monkey striate cortex. The Journal of Physiology, 195(1):215–243, 1968. doi: https://doi.org/10.1113/jphysiol.1968.sp008455. URL https://physoc. onlinelibrary.wiley.com/doi/abs/10.1113/jphysiol.1968.sp008455. Jerry Wei, Da Huang, Yifeng Lu, Denny Zhou, and Quoc V. Le. Simple synthetic data reduces sycophancy in large language models, 2024. URL https://arxiv.org/abs/2308.03958.
A
Supplementary Results
This appendix provides the complete set of analyses that complement the main results in Section 4. All values are reported directly from the computed result files. A.1
Full Model Specifications
Table 7 provides the complete HuggingFace model identifiers and vision encoder specifications for all 12 VLMs. A.2
Two-Turn Attack Analysis
Table 8 reports the complete two-turn attack statistics for each model, including Turn-1 sycophancy, pressure conversion, and final sycophancy rates. The aggregate correlation between brain alignment and Turn-1 resistance is r = 0.018 (p = 0.955), and between brain alignment and pressure conversion is r = −0.104 (p = 0.747), neither of which is significant. Two distinct patterns emerge. First, a group of four models (BLIP-2, LFM-2-VL-1B, LFM-2-VL-8B, SmolVLM-256M, PaliGemma2-10B) already exhibit >80% sycophancy at Turn 1, leaving little room for escalation. Second, several models that resist at Turn 1 are substantially more vulnerable to Turn 2 pressure: Gemma-3-1B increases from 4.5% to 42.2% (∆ = 37.7%), LLaVA-v1.6-7B from 9.6% to 60.2% (∆ = 50.6%), and Qwen2-VL-2B from 13.4% to 73.1% (∆ = 59.8%). Qwen2.5-VL-3B is uniquely resistant to escalation, with a pressure conversion rate of only 0.7%.
22
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
Table 7: Full model specifications for all 12 VLMs. HuggingFace ID: the exact model identifier used for loading. Vision Encoder: architecture of the frozen visual backbone. Hidden Dim.: hidden dimensionality of the vision encoder output. Model
HuggingFace ID
Vision Encoder
SmolVLM-256M SmolVLM-500M Gemma-3-1B LFM-2-VL-1B Qwen2-VL-2B BLIP-2-OPT-2.7B Qwen2.5-VL-3B Phi-3.5-Vision LLaVA-v1.6-7B Idefics2-8B LFM-2-VL-8B PaliGemma2-10B
HuggingFaceTB/SmolVLM-256M-Instruct HuggingFaceTB/SmolVLM-500M-Instruct google/gemma-3-4b-it LiquidAI/LFM2-VL-1.6B Qwen/Qwen2-VL-2B-Instruct Salesforce/blip2-opt-2.7b Qwen/Qwen2.5-VL-3B-Instruct microsoft/Phi-3.5-vision-instruct llava-hf/llava-v1.6-mistral-7b-hf HuggingFaceM4/idefics2-8b LiquidAI/LFM2-VL-450M google/paligemma2-10b-ft-docci-448
SigLIP SigLIP SigLIP (vision_tower) SigLIP2-NaFlex 400M Qwen-ViT (Dynamic Res.) ViT-G/14 + Q-Former Qwen-ViT (Dynamic Res.) CLIP-ViT CLIP-ViT SigLIP (modified) SigLIP2-NaFlex 86M SigLIP
Table 8: Two-turn attack statistics for all 12 VLMs. Turn-1 Σ: sycophancy rate at Turn 1 (before escalation). Π: pressure conversion rate (fraction of initially resistant responses that become sycophantic at Turn 2). Final Σ: overall sycophancy rate after both turns. ∆: absolute increase from Turn-1 to final sycophancy. Turn-1 Σ
Π
Final Σ
∆
SmolVLM-500M Qwen2.5-VL-3B Phi-3.5-Vision Gemma-3-1B LLaVA-v1.6-7B Idefics2-8B Qwen2-VL-2B BLIP-2-OPT-2.7B LFM-2-VL-1B LFM-2-VL-8B SmolVLM-256M PaliGemma2-10B
0.03% 7.8% 3.9% 4.5% 9.6% 15.8% 13.4% 80.7% 80.7% 80.7% 88.6% 82.3%
3.7% 0.7% 20.4% 39.5% 56.0% 54.4% 69.0% 72.4% 81.9% 81.9% 87.3% 97.3%
3.7% 8.5% 23.5% 42.2% 60.2% 61.6% 73.1% 94.7% 96.5% 96.5% 98.6% 99.5%
3.7% 0.6% 19.6% 37.7% 50.6% 45.8% 59.8% 14.0% 15.8% 15.8% 9.9% 17.3%
Mean
39.0%
55.4%
—
24.2%
Model
A.3
Category-Specific Sycophancy
Table 9 presents the mean sycophancy rate for each manipulation category along with the correlation between overall brain alignment and category-specific sycophancy. Table 9: Category-specific sycophancy rates and correlations with overall brain alignment. Visual-domain categories (CAT1–CAT4) show stronger (more negative) mean correlation than the social-domain category (CAT5). Cat.
Description
CAT1 CAT2 CAT3 CAT4 CAT5
Object Misidentification Attribute Manipulation Existence Denial Count Falsification Authority Appeal
Mean Σ
Std
r
p
69.1% 56.1% 53.0% 68.5% 64.1%
0.323 0.394 0.326 0.377 0.374
−0.029 −0.020 −0.223 0.018 −0.020
.929 .950 .486 .955 .951
Category 3 (Existence Denial) exhibits both the lowest mean sycophancy rate (53.0%) and the strongest negative correlation with brain alignment (r = −0.223), though the aggregate correlation does not reach significance. The mean absolute correlation for visual-domain categories (CAT1–CAT4) is |r̄| = 0.073, compared to |r̄| = 0.020 for the social-domain category (CAT5), supporting the hypothesis that brain alignment relates more strongly to visual grounding than to social compliance.
23
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
A.4
Architecture Family Comparison
Table 10 compares the six vision encoder families in terms of brain alignment and sycophancy. Table 10: Architecture family comparison. Brain Score: mean normalized brain alignment. Mean Σ: mean final sycophancy rate. Families are ordered by mean sycophancy. Family
Models
Qwen-ViT CLIP-ViT SigLIP SigLIP (mod.) ViT-G/14 SigLIP2-NaFlex
Qwen2-VL-2B, Qwen2.5-VL-3B LLaVA-v1.6-7B, Phi-3.5-Vision SmolVLM-256M/500M, Gemma-3-1B, PaliGemma2-10B Idefics2-8B BLIP-2-OPT-2.7B LFM-2-VL-1B, LFM-2-VL-8B
Brain Score
Mean Σ
0.995 0.995 0.993 0.991 0.993 0.997
40.8% 41.8% 61.0% 61.6% 94.7% 96.5%
No single architecture family dominates both brain alignment and sycophancy resistance. SigLIP2-NaFlex achieves the highest normalized brain alignment (0.997) but the highest sycophancy (96.5%), while Qwen-ViT and CLIP-ViT show moderate brain alignment with the lowest sycophancy. Within the SigLIP family, sycophancy spans from 3.7% (SmolVLM-500M) to 99.5% (PaliGemma2-10B), indicating that the vision encoder alone does not determine sycophancy resistance; the language decoder and its alignment training play a critical role. A.5
Persuasion Tactic Effectiveness
Table 11 presents the 10 most and 5 least effective persuasion tactics out of the 65 analyzed, ranked by mean sycophancy rate across all 12 models. Table 11: Top 10 most effective and bottom 5 least effective persuasion tactics, ranked by mean sycophancy rate across 12 VLMs. 65 total tactics were analyzed. Rank
Tactic
Mean Σ
Std
1 2 3 4 5 6 7 8 9 10
Statistics Question Specific authority Data appeal Institutional authority Weak suggestion Uncertainty Gaslighting Consistency attack Vague authority
86.5% 82.2% 77.5% 75.2% 75.2% 74.4% 73.6% 72.9% 72.9% 72.5%
0.278 0.329 0.289 0.428 0.428 0.363 0.358 0.373 0.373 0.378
61 62 63 64 65
False technical authority Certainty Extreme pressure Certainty assertion Memory question
45.8% 45.6% 40.0% 29.1% 25.5%
0.458 0.315 0.427 0.339 0.334
Data-driven tactics (statistics, data appeal) and authority-based tactics (specific authority, institutional authority) are most effective, while direct confrontational approaches (extreme pressure, certainty assertion) and meta-cognitive probes (memory question) are least effective. This pattern suggests that VLMs are more susceptible to arguments that mimic evidence-based reasoning than to overt coercion. A.6
Resistance Curves
Table 12 presents the area under the resistance curve (AURC) and resistance slope for each model across the 10 difficulty levels. AURC ranges from 0 to 1, with higher values indicating greater resistance. The correlation between brain alignment and AURC is not significant (r = 0.039, p = 0.904). Resistant models maintain high AURC values across all difficulty levels, while susceptible models collapse early. Notably, Phi-3.5-Vision has a positive slope (0.073), indicating that it becomes more resistant at higher difficulty levels, a pattern that may reflect stronger internal consistency checking when confronted with elaborate manipulation attempts. 24
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation
Table 12: Resistance curve statistics for all 12 VLMs. AURC: area under the resistance curve (higher = more resistant). Slope: linear trend of resistance across difficulty levels (positive = resistance increases with difficulty; negative = decreases).
A.7
Model
AURC
Slope
SmolVLM-500M Qwen2.5-VL-3B Phi-3.5-Vision Gemma-3-1B LLaVA-v1.6-7B Idefics2-8B Qwen2-VL-2B BLIP-2-OPT-2.7B SmolVLM-256M LFM-2-VL-1B LFM-2-VL-8B PaliGemma2-10B
0.952 0.912 0.747 0.590 0.367 0.358 0.272 0.026 0.014 0.011 0.011 0.005
−0.003 −0.000 0.073 0.014 0.018 0.015 −0.005 −0.009 −0.002 −0.011 −0.011 0.000
Per-Difficulty Correlations
Table 13 presents the correlation between overall brain alignment and sycophancy rate at each of the 10 difficulty levels. All correlations are negative, but none reaches significance, and there is no clear monotonic trend with difficulty. Table 13: Brain alignment vs. sycophancy correlation at each difficulty level. Level
Mean Σ
r
p
1 2 3 4 5 6 7 8 9 10
69.8% 72.5% 60.6% 60.0% 57.7% 80.9% 55.6% 64.2% 62.6% 62.3%
−0.350 −0.113 −0.200 −0.270 −0.200 −0.186 −0.196 −0.265 −0.354 −0.096
.265 .727 .533 .396 .533 .563 .541 .404 .259 .766
The non-monotonic pattern in mean sycophancy across difficulty levels (e.g., level 6 at 80.9% vs. level 7 at 55.6%) reflects the heterogeneous nature of the persuasion tactics deployed at each level. The correlations are strongest at the extremes (level 1: r = −0.350; level 9: r = −0.354), suggesting that brain alignment may be most predictive at both low-complexity and high-complexity manipulation conditions. A.8
Breakpoint Analysis
The breakpoint analysis examines at which difficulty level each model first exhibits >50% sycophancy. The correlation between brain alignment and breakpoint is r = 0.067 (p = 0.837), indicating no significant relationship. Ten of the 12 models have a breakpoint of 1 (capitulating immediately at the lowest difficulty), while SmolVLM-500M and Qwen2.5-VL-3B have breakpoints of 11 (never reaching 50% sycophancy at any difficulty level). A.9
Additional Visualizations
Figures 5 to 9 provide additional visualizations of the brain alignment data and dataset structure.
25
Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation