ConceptioArchivearXiv CS
arXiv CSopen access

When Large Language Models Fail in Healthcare: Evaluating Sensitivity to Prompt Variations

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

When Large Language Models Fail in Healthcare: Evaluating Sensitivity to Prompt Variations 1

Mahdi Alkaeed1,∗ , Department of Computer Science and Engineering, Doha, Qatar ∗ Corresponding author email: [email protected]

Abstract Large Language Models (LLMs) are increasingly used in healthcare for tasks such as clinical question answering, diagnosis support, and report summarization. Despite their promise, these models remain highly sensitive to subtle

arXiv:2606.07237v1 [cs.CL] 5 Jun 2026

prompt perturbations, both lexical and syntactic, posing serious risks in safety-critical clinical applications. In this study, we conduct a systematic sensitivity analysis to evaluate the robustness of both general-purpose (e.g., GPT-3.5, Llama3) and medical-specific LLMs (e.g., ClinicalBERT, BioLlama3, BioBERT) using the MedMCQA benchmark. We categorize perturbations into natural and adversarial types and examine their effect on model consistency, accuracy, and reliability in clinical reasoning tasks. Our findings reveal that medical LLMs are not intrinsically safe. Even minor variations in phrasing can alter clinical advice, and targeted adversarial prompts can provoke harmful outputs. In high-stakes settings like healthcare, such unpredictability is unacceptable-models that change diagnoses due to reworded inputs or hallucinate medications when slightly rephrased cannot be reliably trusted by clinicians. While models tend to show resilience to simple lexical substitutions or paraphrasing, they often break down under syntactic reordering or misleading contextual cues. This fragility is evident across both general-purpose and domain-specific LLMs. Notably, adversarial manipulations can lead to clinically dangerous outputs, such as recommending incorrect dosages or omitting critical findings. Keywords: Large Language Models, Healthcare AI, Prompt Perturbations, Lexical and Syntactic Perturbations, Adversarial Robustness, LLM Sensitivity Analysis, LLM Reliability.

1. Introduction Large Language Models (LLMs) are increasingly being leveraged in healthcare applications, including clinical question answering and diagnostic support [1]. Despite their promise, these models exhibit sensitivity to minor prompt variations and remain vulnerable to adversarial manipulations [2, 3]. In this study, we systematically categorize prompt perturbations, focusing on lexical and syntactic changes, and examine their impact on both general-purpose models (e.g., GPT-3.5, Llama3) and medical-specific models (e.g., ClinicalBert, BioLlama3, BioBert) across critical clinical tasks. We further differentiate between adversarial vulnerabilities and natural robustness, defined as stability under benign input variations. To support these analyses, we provide comprehensive tables summarizing comparative findings and highlight key challenges in deploying reliable and trustworthy medical AI systems [4]. Specifically, Table 1 consolidates evidence on LLM robustness across healthcare and related domains, detailing the evaluated models, types of perturbations or attacks, and their observed effects. RQ1 How do lexical and syntactic prompt perturbations impact the performance and robustness of general-purpose versus medical-specific LLMs in clinical tasks such as question answering, and diagnosis support? RQ2 What are the key challenges in ensuring trustworthy and reliable behavior of LLMs under prompt perturbations in healthcare applications?

Table 1: Summary of key findings on LLM robustness in medical and related tasks, with a focus on lexical and syntactic prompt perturbations. Each row lists the study, evaluated models, task/domain, type of perturbation, and observed effect. Study & Ref.

Model(s)

Guan et al. [5]

GPT-4,

GPT-4

mini, DeepSeek

Task / Domain

Perturbation Type

Effect / Findings

General & Medical

Lexical reordering (syn-

Accuracy decreased slightly (∼ −2.8% for

QA

tactic)

GPT-4) when the order of answer choices

(e.g.,

MedM-

CQA)

changed. Shows sensitivity to phrasing variations.

Bolton et al.

(RAm-

BLA) [6]

GPT-4,

GPT-3.5,

LLaMA2-7B,

Mis-

Clinical

QA

(PMQA-L)

tral Bolton et al.

(RAm-

BLA) [6]

Paraphrasing

(lexical),

Most models retained performance, but

minor

changes

small F1 drops observed, indicating partial

word

(syntactic)

LLaMA2-7B,

Mis-

tral

Clinical

QA

(PMQA-L)

Added

stability under benign lexical variations.

distractor

sen-

tences (contextual)

Performance drop varied by model (−20% for LLaMA, −38% for Mistral); GPT-4 and GPT-3.5 were more stable. Highlights difference in natural robustness.

Ness et al. (MedFuzz)

GPT-4

[7]

USMLE-style

Lexical changes in ques-

Minor rewording sometimes caused GPT-4

MedQA

tions

to change answers, showing sensitivity to

Diagnostics & QA

Syntactic reordering

phrasing even without adversarial intent. Han et al. [8]

GPT-3.5,

GPT-4,

T5

Small variations in question structure affected model outputs slightly; demonstrates limits of intrinsic robustness.

This paper

GPT-3.5,

Llama3,

ClinicalBert,

BioL-

lama3, BioBert

General & Medical

Lexical

QA

prompt perturbations

(e.g.,

MedM-

CQA)

and

syntactic

Minor prompt variations reduced accuracy and sometimes altered model advice. Medical LLMs show limited natural robustness; careful prompt design is needed to maintain clinical consistency.

Contributions of this Paper: 1. We formally define and categorize prompt perturbations, with a focus on lexical and syntactic variations, and systematically evaluate their impact on LLM performance in healthcare tasks, highlighting natural robustness, defined as stability under benign input variations. 2. We conduct a comprehensive comparative analysis of LLM performance across general-purpose models (e.g., GPT-3.5, Llama3) and medical-specific models (e.g., ClinicalBERT, BioLlama3, BioBERT) in clinical applications. 3. We identify critical vulnerabilities and limitations of current LLMs under these perturbations and discuss open challenges for developing more reliable, robust, and trustworthy medical AI systems.

2. Related Work Prompt sensitivity refers to a model’s output variation under small changes to its input [9]. Perturbations can be classified linguistically: lexical perturbations alter individual tokens (e.g., synonym substitution, typos, negations); syntactic perturbations change word order or sentence structure (e.g., active to passive voice or vice versa, reordering clauses). These categories are widely used in robustness evaluation (see e.g., the RUPBench taxonomy [10]). . In practice, lexical perturbations test the model’s tolerance to surface form noise and syntactic to structural variation. In medical settings, prompt sensitivity is critical: changing a few words in a patient’s history should not dramatically alter a correct diagnosis or critical safety advice. Even small adversarial changes (e.g., subtle misinformation or hidden prompts) can lead to harmful output.

2

Table 2: Examples of benign lexical and syntactic perturbations in clinical contexts. Perturbation Type

Original vs. Perturbed String

Semantic Impact

Lexical (Benign)

O: “The patient has diabetes.”

Meaning-preserving; minor orthographic error; does

P: “The patient has diabetees.”

not alter clinical interpretation.

O: “Administer 5 mg of amlodipine daily.”

Preserves meaning; numeric expression converted to

P: “Administer five mg of amlodipine daily.”

word form; no clinical consequence.

O: “The nurse administered the drug.”

Meaning-preserving; clefting shifts sentence focus

P: “It was the nurse who administered the drug.”

without affecting clinical content.

O: “Patient reports no chest pain or shortness of

Preserves clinical meaning; syntactic reordering for

breath.”

stylistic variation; content unchanged.

Lexical (Benign)

Syntactic (Benign)

Syntactic (Benign)

P: “No chest pain or shortness of breath is reported by the patient.” Lexical + Syntactic (Benign)

O: “The physician recommends starting insulin ther-

Minor lexical substitution and syntactic shift; mean-

apy immediately.”

ing preserved; natural language variation.

P: “Immediately, the physician recommends initiating insulin therapy.”

2.1. Prompt Sensitivity in General LLMs To systematically quantify the effects of prompt rephrasings on LLMs, Errica et al. [11] introduced two metrics, sensitivity and consistency. Furthermore, the authors conducted an empirical evaluation on multiple datasets and general-purpose LLMs, including GPT-4o, GPT-3.5, LLaMA-3, and Mixtral. Their results demonstrate that sensitivity and consistency offer complementary insights beyond accuracy, revealing model weaknesses and helping developers identify “hard” samples and problematic classes. Similarly, Salinas et al. [12] systematically analyzed how minor prompt variations such as output formatting, lexical perturbations, and jailbreak instructions can impact LLM behavior across 11 classification tasks. Their results demonstrate that even trivial modifications, like the inclusion of whitespace or polite expressions, can lead to syntactic shifts in model predictions and overall accuracy. In [9], the authors presented a novel metric called PromptSensiScore (PSS), which quantifies how much LLM responses vary across different syntactically equivalent prompts for the same input instance. Their results revealed that larger LLMs are comparatively more robust, and few-shot prompting can improve robustness. In addition, they showed that decoding confidence correlates with prompt stability. In a similar study, Cao et al.[13] presented a new benchmark, namely RobustAlpacaEval, designed to assess prompt sensitivity at the instance level by evaluating LLM performance across syntactically equivalent prompt variants. Their findings using various open-source LLMs (e.g., Llama, Mistral, and Gemma families) reveal that even high-performing models exhibit substantial degradation under worst-case prompts, which are often unpredictable. Sclar et al. [14] studied how paraphrasing and format changes affect output consistency across multiple models. Pezeshkpour et al. [15] demonstrated that LLMs are influenced by the order of options in multiple-choice settings. Pezeshkpour et al. [15] demonstrated that the order of options in multiple-choice settings can influence the responses of LLMs. 2.2. Prompt Sensitivity in Healthcare LLMs Recent research has increasingly highlighted the importance of healthcare LLMs sensitivity to key medical information [16]. Ness et al.[7] introduced MedFuzz to test healthcare LLM robustness, revealing that minor perturbations could mislead models with high benchmark scores. This highlighted a critical concern, such vulnerabilities could lead to clinical errors. Moradi et al.[17] explore the vulnerability of biomedical NLP models, including BioBERT and SciBERT, to adversarial attacks, showing a significant drop in performance with even slight input modifications [17]. Ceballos et.al [18] investigate how healthcare LLMs show fragility, with phrasing variations impacting performance 3

Figure 1: This methodological pipeline outlines a comprehensive robustness evaluation framework for assessing the sensitivity and stability of healthcare LLMs using the MedMCQA benchmark, supported by advanced NLP tools (e.g., BioSyn, scispaCy) and biomedical metrics (e.g., BERTScore, USE) to quantify model resilience to input variations.

and fairness. Beede et al. [19] illustrated that varying symptom phrasing (e.g., “sharp chest pain” vs. “chest discomfort”) can result in conflicting diagnoses. Pais et al. [20] reported that minor spelling errors in drug names can lead to prescribing mistakes or missed drug interactions. Zhang et al. [21] highlighted the difficulty in generalizing to real-world clinical terminology, where small lexical variations can cause models to misclassify disease severity. Similarly, Yan et al. [16] conducted a clinically focused evaluation that demonstrated how models often fail to prioritize critical diagnostic cues such as patient age and symptom descriptions, elements that are essential for accurate clinical reasoning. Zhan et al. [22] proposed that the COPLE framework suggests that refining lexical choices can enhance output consistency when faced with prompt perturbations. The Table 2 enumerates various "benign" perturbations, including lexical typos, numeric-to-word conversions, and syntactic reordering like clefting and passive voice shifts. Each example demonstrates that despite linguistic or orthographic variations, the underlying clinical meaning and propositional content remain strictly preserved. These variations serve as a baseline for testing model stability, ensuring that minor stylistic or accidental changes do not inadvertently alter the intended clinical interpretation or reasoning outcome.

3. Problem Statement and Methodology A clinical case C from the MedMCQA dataset begins with a patient prompt P , which is processed by a healthcare LLM to produce an output R = fLLM (P ). Now, consider a minor lexical variation, such as replacing “shortness of breath” with its synonym “dyspnea,” resulting in a slightly altered input P ′ , where P ′ ≈ P . Althoughically equivalent, this change may lead the model to produce a different output R′ = fLLM (P ′ ). Ideally, the model should treat both inputs similarly: P ′ ≈ P

R′ ≈ R However, in practice, even a small perturbation δP = P ′ − P can cause a

significant output shift ∆R = R′ − R, such that: δP small

∆R small. This inconsistency reveals the critical

vulnerability that LLMs in healthcare can be sensitive to minor linguistic changes, leading to diagnostic variability. 4

Figure 2: Comparison of accurate and misinformation-attacked medical advice for acute limb ischemia. Lexical and syntactic alterations distort critical details, leading to dangerous clinical decisions. Highlighted risks demonstrate the impact of misinformation on patient outcomes.

3.1. Sensitivity Analysis Framework The proposed framework begins with an original medical prompt P and generates multiple lexical variants through token-level substitutions. The candidate perturbed prompts P1′ , P2′ , . . . , PN′ are then validated using BioSyn to ensure medical equivalence. Each prompt (original and perturbed) is encoded into a high-dimensional sentence embedding using the Universal Sentence Encoder (USE). Cosine similarity scores are computed between the original and each perturbed embedding to quantify syntactic closeness. Based on these scores, a top-K subset of perturbed prompts most syntactically aligned with P is selected for further evaluation. These selected prompts are then passed to a target healthcare LLM (e.g., GPT-3.5, LLaMA3, ClinicalBert, BioLLama3, and BioBert), and the model’s outputs are recorded. Each response is compared against ground truth using accuracy-based metrics. As shown in Figure 1. 3.1.1. Prompt Generation Strategy Evaluating the robustness of healthcare LLMs necessitates exposure to a wide variety of prompts that capture real-world clinical diversity. Medical LLMs generally tolerate simple lexical variations. In one study (RAmBLA 2024), GPT-4 and GPT-3.5 maintained near-baseline QA accuracy when single words were misspelled or replaced by synonyms [5]. All evaluated models in that work were “robust to spelling errors”-e.g., GPT-4’s answer F1 remained 0.84 even with 3 character-level mutations [6]. Similarly, meaning-preserving synonym swaps or negations often had little effect on output accuracy. This suggests that foundation models have some built-in lexical flexibility, likely due to subword modeling and pretraining on noisy text. In this study, we utilize the MedMCQA dataset, taking advantage of its extensive coverage across medical specialties and conditions. Lexical Variations The perturbation process involves four key stages: (i) identifying medical terms in the prompt, (ii) replacing terms with clinically accurate synonyms using BioSyn, (iii) evaluating syntactic equivalence through word embeddings and cosine similarity measures, and (iv) generating perturbed prompts that maintain contextual 5

and diagnostic consistency. Algorithm 1 details the complete perturbation procedure. To quantify the effects of lexical variations, we measure: (i) syntactic similarity between original and perturbed prompts using Sentence-BERT embeddings, (ii) consistency of model responses across perturbed prompts, and (iii) the clinical reliability of generated outputs concerning benchmarked ground-truth answers. Figure(2, b) illustrates this process step-by-step using a clinical sample from MedMCQA. The procedure begins by identifying medically relevant terms in the original prompt with curated biomedical lexicons and pretrained language models such as BioBERT and SciBERT. Candidate synonyms for each term are then retrieved using BioSyn, which combines lexical and syntactic similarity measures. To maintain clinical accuracy, we apply synonym marginalization through vector embeddings (e.g., Word2Vec with spaCy) and select the top-ranked replacements based on cosine similarity scores. For example, terms like “fever” may be replaced by “pyrexia,” while “auscultation” becomes “stethoscope.”. Algorithm 1 Lexical Perturbation using BioSyn & Cosine Similarity 1: Input: Clinical prompt P = {t1 , t2 , . . . , tn } 2: Output: Lexically perturbed prompt P ′ 3: for each medical term ti ∈ P do 4:

Retrieve candidate synonyms Si using BioSyn.

5:

for each synonym sj ∈ Si do

6:

Compute cosine similarity Sim(ti , sj ).

7:

end for

8:

Select best synonym B(ti ) = arg maxsj ∈Si Sim(ti , sj ).

9:

Replace ti with B(ti ) in P .

10: end for 11: return P ′

Syntactic Variations By contrast, altering input structure can noticeably impact results. For example, simply swapping the order of semantically identical answer options caused GPT-4 to change its responses and incur a measurable accuracy drop. In the Order Effect study [5], shuffling inputs led to performance declines across tasks. GPT-4’s accuracy on a paraphrasing task dropped by about 2–3% when choices were reordered [5]. Few-shot prompts mitigated this effect only partially: while adding examples reduced the gap slightly, no model fully eliminated sensitivity to input order [5]. This indicates that even advanced LLMs remain dependent on prompt formatting. In the clinical context, persisting order-sensitivity is worrisome: for instance, rephrasing a symptom list or reordering history elements might unintentionally flip a prediction. Syntactic perturbations aim to modify the structure and phrasing of clinical prompts while preserving their overall medical coherence and diagnostic intent. This process begins by identifying key clinical concepts within the prompt and exploring contextually relevant rephrasings or substitutions. Unlike lexical perturbations, which target isolated term replacements, syntactic perturbations may involve changes in sentence structure or the substitution of conceptually similar phrases. Syntactic consistency is validated using contextual word embeddings (e.g., BioBERT), cosine similarity scoring, and biomedical resources such as BioSyn and the UMLS Metathesaurus. The algorithm for generating syntactic ally preserved prompts is outlined in Algorithm 2, and an example is illustrated in Figure(2, c).

6

Table 3: Sensitivity analysis of LLMs to question phrasing: Implications for format integrity and diagnostic reliability in clinical diagnosis.. Prompt and LLM Response

Evaluation Metrics

Interpretation (Analysis)

Original Prompt: A 72-year-old male with a history of chronic

RR: Response Rate

The baseline prompt establishes the clinical con-

smoking and productive cough presents with increased sputum

FIR: Follow-up Instruc-

text and complex instructions. It serves as the

and dyspnea. Physical exam reveals wheezing and a barrel chest.

tion Rate

reference point for measuring how well the model

Instruction: Return the results in JSON format with the keys:

BERTScore: Semantic

maintains its response consistency and instruc-

"answer" and "explanation".

Similarity

tion adherence.

Question: Based on the clinical presentation what is the most

USE:

likely diagnosis?

tence Encoder (Embed-

Options: A) Pulmonary Fibrosis, B) Chronic Obstructive Pul-

ding model)

Universal

Sen-

monary Disease (COPD), C) Pulmonary Edema, D) Lobar Pneumonia LLM response (Standard prompt (unperturbed)):

RR: ✓

Optimal Performance:

JSON format:{

FIR: ✓

strates high-fidelity reasoning and perfect in-

BERTScore: 0.97

struction following. Both the medical diagnosis

USE: 0.98

(RR) and the JSON formatting (FIR) are han-

"answer":

"B) Chronic Obstructive Pulmonary Disease (COPD)",

"explanation":

"The patient’s history and symptoms are

The model demon-

dled correctly.

consistent with COPD."} LLM response (Lexical perturbed prompt):

RR: ✓

Instruction Drift: While the model’s clinical

JSON format:{

FIR: ×

logic remains sound (correct diagnosis), changing

"My answer is" : "B) Chronic Obstructive Pulmonary Disease

BERTScore: 0.75

the prompt’s vocabulary causes a failure in fol-

(COPD)",

USE: 0.85

lowing the specific follow-up instructions (FIR).

"In this status" :

"Given the patient’s known history of

The model fails the formatting constraint despite

COPD..."}

the correct answer.

LLM response (Syntactic perturbed prompt):

RR: ×

Structural Misinterpretation: Reordering the

JSON format:

FIR: ×

sentence structure (syntactic perturbation) dis-

"The right choice is" : "C) Pulmonary Edema",

BERTScore: 0.67

rupts the model’s attention mechanism. It mis-

"The clinical presentation," including a history of ....}

USE: 0.62

weights clinical signs (like "increased sputum") as

{

"Pulmonary Edema" instead of "COPD," leading to a **False Answer (RR Failure)** and (FIR Failure).

Algorithm 2 Syntactic Perturbation using Contextual Embeddings 1: Input: Clinical prompt P = {t1 , t2 , . . . , tn } 2: Output: syntactic ally perturbed prompt P ′ 3: Identify key medical concepts and inter-term relationships in P . 4: for each term ti ∈ P do 5:

Retrieve syntactic ally related terms Si using contextual embeddings and BioSyn.

6:

for each term sj ∈ Si do

7:

Compute context-aware similarity Sim(ti , sj ).

8:

end for

9:

Select B(ti ) = arg maxsj ∈Si Sim(ti , sj ).

10:

Replace ti with B(ti ) in P .

11: end for 12: Optionally rephrase sentence structure while preserving clinical validity. 13: return P ′

3.2. Evaluation Metrics This study evaluates the performance of healthcare LLMs using a range of metrics, including BERTScore, Universal Sentence Encoder (USE), Response Rate (RR), and Follow-up Instruction Rate (FIR). BERTScore measures the syntactic similarity between the model’s response and a ground truth, with higher scores indicating more stability. The USE is a sentence embedding model developed by Google that encodes text into high-dimensional vectors. It captures the syntactic meaning of entire sentences rather than just individual words, and measures how syntactically 7

Table 4: LLM performance on MedMCQA dataset under unperturbed, lexical, and syntactic perturbations (100 trials). Perturbation Type

Standard (unperturbed)

Lexical

syntactic

Metrics

GPT-3.5

Llama3

ClinicalBert

BioLlama3

BERTtScore

92.14

94.23

91.15

95.94

BioBert 89.12

USE score

91.87

93.03

88.21

93.22

87.11

RR

99.23

99.43

99.04

99.12

98.11

FIR

98.94

99.01

97.45

99.30

97.50

BERTScore

89.11

90.45

88.90

93.33

86.10

USE score

88.07

90.01

84.56

89.23

82.02

RR

96.12

96.23

96.28

96.10

95.67

FIR

95.12

96.43

95.01

96.71

94.91

BERTScore

87.22

89.89

86.13

90.52

83.72

USE score

89.61

90.21

83.71

85.19

79.11

RR

94.20

94.81

93.23

93.30

93.43

FIR

93.32

94.53

92.91

95.31

90.13

similar a model’s response is to a ground truth, helping assess how well the healthcare LLM preserves clinical meaning despite variations in prompts. While the RR metric quantifies the proportion of valid responses generated by the model, calculated as RR =

#valid_responses N

where N denotes the total number of responses generated. The FIR measures the extent to which the model adheres to the given instructions within those valid responses, calculated as FIR =

#followed_instructions . #valid_responses

Model responses are not always clinically valid, as correctness may vary. To assess stability, we define the following thresholds: scores between 0.9 and 1.0 indicate Very Stable, 0.7 and 0.9 Moderately Stable, 0.5 and 0.7 Mild Sensitivity, and below 0.5 High Sensitivity.

4. Results and Discussions 4.1. Dataset and Experimental Setup The MedMCQA dataset, a large-scale, multiple-choice question answering benchmark with over 194,000 meticulously curated questions covering various medical specialties relevant to medical education and clinical reasoning, was employed to support our analysis. We evaluated the robustness of five models against prompt perturbations.These models are GPT-3.5, LLaMA3, ClinicalBERT, BioLLaMA3, and BioBERT. Each model was tested under three conditions. First, clean prompts, then lexical perturbations (word-level changes), and syntactic perturbations (meaninglevel rephrasings). For each perturbation type, 1,00 samples were randomly selected from the MedMCQA dataset. Performance evaluation was conducted using metrics such as BERTScore and USE score to assess syntactic similarity between generated responses. Additional evaluation metrics included RR, FIR, and EV for clinical correctness. Model inference was carried out using the Hugging Face Transformers and Datasets libraries, utilizing either PyTorch or TensorFlow, depending on the model. In addition, we utilize Ollama to facilitate local deployment and testing of open-source LLMs such as LLaMA3 and BioLLaMA3. For biomedical text processing, we employed scispaCy, a scientific NLP library built on spaCy, optimized for scientific and biomedical contexts. For synonym resolution and entity matching, BioSyn, in combination with BioBERT, was used to embed clinical terms. All models were used in their pretrained form without fine-tuning with default hyperparameters.

8

Figure 3: Performance of models across perturbation types. Heatmap shows scores for all metrics for each model and study, with ‘N’ indicating metrics not used. Models are labeled with their type (General-Purpose or Adapted Domain) and study source.

4.2. Results on Prompt Sensitivity Analysis Our findings highlight three core areas: the sensitivity of LLMs to prompt variation, identifiable failure patterns under adversarial perturbed prompts, and the potential of thresholding strategies to enhance model reliability in high-stakes medical decision-making. The table 3 provides a granular evaluation of how prompt perturbations impact a LLM ability to maintain both medical reasoning accuracy and structural format integrity. While the model achieves near-perfect alignment under standard conditions, lexical variations trigger "instruction drift," causing the model to prioritize natural language flow over specific JSON formatting constraints (FIR). Most critically, syntactic reordering leads to a total systemic collapse where the model misinterprets clinical features, resulting in a false diagnosis (RR failure). These results demonstrate that while LLMs are clinically capable, their reliability is highly sensitive to the structural and linguistic framing of the input prompt. While Table 4 summarizes the performance of various LLMs on the MedMCQA dataset under both unperturbed and perturbed conditions, averaged over 100 samples. Under standard conditions, BioLlama-3 excels in medical reasoning, while general-purpose models, such as GPT-4.5, remain competitive, and domain-specific models lag. When subjected to lexical and syntactic perturbations, BioLlama-3 demonstrates the highest robustness, whereas models such as BioBERT and ClinicalBERT exhibit significant sensitivity, particularly to syntactic perturbations. To quantify the impact of adversarial lexical perturbations on LLM outputs, we examined the USE similarity scores for BioLlama-7b responses. 4.3. Discussions The heatmap at Figure 3, provides a comprehensive comparison of multiple models’ performance across three perturbation types: Standard (unperturbed), Lexical, and Syntactic. Each row represents a model and study, including its type—General-Purpose (e.g., GPT-3.5) or Adapted Domain (e.g., BioBERT, MedLlama3)—while columns correspond to evaluation metrics such as BERTScore, USE, RR, FIR, Accuracy, and AUROC. Cells show the performance scores, with “N” marking metrics not used by a specific study. From the figure, several observations emerge: • Our Study models (GPT-3.5 and BioBERT) consistently perform well across all metrics and perturbations, indicating robustness to lexical and syntactic changes. Accuracy scores remain high but decrease slightly with perturbations, reflecting realistic sensitivity. • Comparison with previous studies (Yan et al., Ness et al., Arroyo et al.) highlights differences in evaluation methodology and model types. Not all metrics are reported in these studies, which is indicated by “N” in the heatmap. • Metric coverage varies, showing that some models or studies only report specific metrics (e.g., AUROC for MedAlpaca), limiting direct comparisons. 9

Figure 4: Average similarity scores on 100 MedMCQA prompts across increasing lexical perturbation levels, showing that BioLLaMA2-7B maintains higher robustness than LLaMA2 as prompt perturbations increase.

• Perturbation impact is visible: Lexical and Syntactic perturbations generally reduce scores slightly, demonstrating how model robustness can be quantified. Morover, Figure 4 shows the robustness Evaluation of LLaMA2 vs. BioLLaMA2-7B on Lexically Perturbed Prompts. This figure presents. The average similarity scores between model predictions and ground truth answers across 100 MedMCQA prompts, subjected to 10 levels of lexical perturbation. 4.4. Challenges and Future Directions To make these comparisons more fair and rigorous, future work should aim to: • Standardize the metrics, models, and datasets used across studies. • Evaluate models under the same perturbation types and experimental conditions. • Extend the analysis to additional clinical datasets and new domain-adapted models. • Comprehensive robustness evaluation: Current benchmarks mainly assess accuracy on clean data, lacking standardized robustness tests for medical prompts. Frameworks like MedFuzz and RAmBLA address this by introducing controlled perturbations. Future efforts should broaden these benchmarks to cover more clinical tasks and multilingual contexts Such efforts will enable direct, equitable benchmarking, improving the interpretability and reproducibility of clinical LLM evaluation.

5. Limitations A major limitation of this study is its reliance on traditional embedding-based similarity metrics, such as cosine similarity. Although these metrics are commonly used, they often struggle to capture subtle semantic nuances that are critical in clinical contexts. For instance, negation, temporal qualifiers, and disease staging can drastically alter the clinical meaning of a statement, yet may not be reflected in similarity scores. Consequently, models assessed solely 10

with these metrics may appear robust, even when they misinterpret clinically important details. This underscores the need for more sophisticated evaluation methods that can effectively capture these fine-grained semantic distinctions in healthcare applications.

6. Conclusions Medical LLMs are not inherently safe, as minor phrasing changes can produce harmful outputs. Clinicians face risks from unreliable diagnoses, incorrect drug recommendations, or missed critical findings, highlighting the urgent need to address model robustness, reliability, and safety in real-world healthcare applications. While LLMs offer significant potential for healthcare, their robustness under perturbations remains uncertain. Our study shows that lexical and syntactic perturbations reduce the performance of both GPT-3.5 (General-Purpose) and BioBERT (Adapted Domain), with accuracy decreasing slightly from 98–99% in the unperturbed setting to 94–95% under syntactic changes. Other metrics, including BERTScore, USE, RR, and FIR, also show minor declines under these perturbations. These findings emphasize that even state-of-the-art LLMs are sensitive to subtle changes in input phrasing, reinforcing the need for systematic robustness evaluation. Future work should prioritize adversarial robustness alongside accuracy and develop models and interfaces that are provably resilient or, at minimum, transparently limited, to build the trust required for safe clinical deployment.

7. Acknowledgments The authors want to acknowledge support from the Qatar University High Impact Internal Grant (QUHI-CENG23/24127). The statements made herein are solely the responsibility of the authors.

References [1] P. Zhang, M. N. Kamel Boulos, Generative ai in medicine and healthcare: Promises, opportunities and challenges, Future Internet 15 (9) (2023) 286. [2] E. Shayegani, M. A. A. Mamun, Y. Fu, P. Zaree, Y. Dong, N. Abu-Ghazaleh, Survey of vulnerabilities in large language models revealed by adversarial attacks, arXiv preprint arXiv:2310.10844 (2023). [3] F. M. Polo, R. Xu, L. Weber, M. Silva, O. Bhardwaj, L. Choshen, A. F. de Oliveira, Y. Sun, M. Yurochkin, Efficient multi-prompt evaluation of llms, Advances in Neural Information Processing Systems 37 (2024) 22483–22512. [4] S. Qi, Z. Cao, J. Rao, L. Wang, J. Xiao, X. Wang, What is the limitation of multimodal llms? a deeper look into multimodal llms through prompt probing, Information Processing & Management 60 (6) (2023) 103510. [5] B. Guan, T. Roosta, P. Passban, M. Rezagholizadeh, The order effect: Investigating prompt sensitivity in closed-source llms, arXiv preprint arXiv:2502.04134 (2025). [6] W. J. Bolton, R. Poyiadzi, E. R. Morrell, G. v. B. G. Bueno, L. Goetz, Rambla: A framework for evaluating the reliability of llms as assistants in the biomedical domain, arXiv preprint arXiv:2403.14578 (2024). [7] R. O. Ness, K. Matton, H. Helm, S. Zhang, J. Bajwa, C. E. Priebe, E. Horvitz, Medfuzz: Exploring the robustness of large language models in medical question answering, arXiv preprint arXiv:2406.06573 (2024). [8] T. Han, S. Nebelung, F. Khader, T. Wang, G. Müller-Franzes, C. Kuhl, S. Försch, J. Kleesiek, C. Haarburger, K. K. Bressem, et al., Medical large language models are susceptible to targeted misinformation attacks, NPJ digital medicine 7 (1) (2024) 288.

11

[9] J. Zhuo, S. Zhang, X. Fang, H. Duan, D. Lin, K. Chen, Prosa: Assessing and understanding the prompt sensitivity of llms, in: Findings of the Association for Computational Linguistics: EMNLP 2024, 2024, pp. 1950–1976. [10] Y. Wang, Y. Zhao, Rupbench: Benchmarking reasoning under perturbations for robustness evaluation in large language models, arXiv preprint arXiv:2406.11020 (2024). [11] F. Errica, D. Sanvito, G. Siracusano, R. Bifulco, What did i do wrong? quantifying llms’ sensitivity and consistency to prompt engineering, in: Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2025, pp. 1543–1558. [12] A. Salinas, F. Morstatter, The butterfly effect of altering prompts: How small changes and jailbreaks affect large language model performance, in: Findings of the Association for Computational Linguistics: ACL 2024, 2024, pp. 4629–4651. [13] B. Cao, D. Cai, Z. Zhang, Y. Zou, W. Lam, On the worst prompt performance of large language models, Advances in Neural Information Processing Systems 37 (2024) 69022–69042. [14] M. Sclar, Y. Choi, Y. Tsvetkov, A. Suhr, Quantifying language models’ sensitivity to spurious features in prompt design or: How i learned to start worrying about prompt formatting, in: International Conference on Learning Representations, Vol. 2024, 2024, pp. 25055–25083. [15] P. Pezeshkpour, E. Hruschka, Large language models sensitivity to the order of options in multiple-choice questions, in: Findings of the Association for Computational Linguistics: NAACL 2024, 2024, pp. 2006–2017. [16] C. Yan, X. Fu, Y. Xiong, T. Wang, S. C. Hui, J. Wu, X. Liu, Llm sensitivity evaluation framework for clinical diagnosis, in: Proceedings of the 31st International Conference on Computational Linguistics, 2025, pp. 3083–3094. [17] M. Moradi, M. Samwald, Improving the robustness and accuracy of biomedical language models through adversarial training, Journal of Biomedical Informatics 132 (2022) 104114. [18] A. M. Ceballos-Arroyo, M. Munnangi, J. Sun, K. Zhang, J. Mcinerney, B. C. Wallace, S. Amir, Open (clinical) llms are sensitive to instruction phrasings, in: Proceedings of the 23rd Workshop on Biomedical Natural Language Processing, 2024, pp. 50–71. [19] E. Beede, E. Baylor, F. Hersch, A. Iurchenko, L. Wilcox, P. Ruamviboonsuk, L. M. Vardoulakis, A human-centered evaluation of a deep learning system deployed in clinics for the detection of diabetic retinopathy, in: Proceedings of the 2020 CHI conference on human factors in computing systems, 2020, pp. 1–12. [20] C. Pais, J. Liu, R. Voigt, V. Gupta, E. Wade, M. Bayati, Large language models for preventing medication direction errors in online pharmacies, Nature Medicine (2024) 1–9. [21] A. Zhang, L. Xing, J. Zou, J. C. Wu, Shifting machine learning for healthcare from development to deployment and from models to data, Nature Biomedical Engineering 6 (12) (2022) 1330–1345. [22] P. Zhan, Z. Xu, Q. Tan, J. Song, R. Xie, Unveiling the lexical sensitivity of llms: Combinatorial optimization for prompt enhancement, in: Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, 2024, pp. 5128–5154.

12

Record · ID 266194 · SHA-256 747aa317f3ee61d8
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.