Evaluation of Contextual Understanding in Large Language Models
Subavarshana Arumugam * 1 Mamta Nallaretnam * 1 Kithuni Wickramasinghe * 1 Chamath Gunapala * 1 Pragatheeswaran Vipulanandan 2 Uthayasanker Thayasivam 1 Kamal Premaratne 2
arXiv:2609.09004v1 [cs.CL] 8 Sep 2026
Abstract
sequential in high-stakes domains such as medical and legal QAing, where responses must be grounded in the provided context rather than memorized priors.
Large Language Models (LLMs) demonstrate impressive performance across diverse NLP tasks, yet their ability to exhibit genuine contextual understanding remains uncertain. Traditional evaluation metrics such as perplexity, BiLingual Evaluation Understudy (BLEU), or surface-level accuracy fail to reveal how well LLMs extract, integrate, and reason over contextual information–a gap particularly critical in question answering, where models must align responses with contextually grounded knowledge rather than memorized associations. We propose a novel knowledge graph-based evaluation framework introducing Semantic Structural Similarity for KGs (S3KG), a hybrid similarity measure integrating structural and semantic similarity into a continuous evaluation score, alongside a diagnostic framework for categorizing reasoning errors. To validate this pipeline, we evaluate S3KG against established metrics on a curated question-answer (QA) benchmark, demonstrating its effectiveness in measuring correctness, faithfulness, and interpretability in LLM-generated responses.
2. Related Work 2.1. LLM Evaluation and Contextual Understanding Existing evaluation metrics–perplexity, BLEU (Papineni et al., 2002), and token-level accuracy–measure surfacelevel fluency and overlap but fail to capture relational depth or factual faithfulness. Zhu et al. (2024) benchmark LLMs across several structured tasks, including co-reference resolution and dialogue state tracking, finding that models capture general context patterns but fail in fine-grained interpretation. Yan et al. (2024) probe reasoning fidelity by manipulating in-context examples through logical substitutions, revealing systematic failures in formal reasoning that surface-level metrics do not expose. On the construction side, LLM-driven KG frameworks such as GraphRAG (Edge et al., 2024) and KEA (Haskins & Adams, 2025) have demonstrated that high-quality relational triplets can be extracted with minimal supervision, while hallucination-oriented evaluators like GraphEval (Sansford et al., 2024) adopt few-shot prompting with instruction tuning to encourage label consistency across graphs. Despite these advances, no existing metric jointly captures relational structure and semantic faithfulness in a continuous interpretable score.
1. Introduction Large language models (LLMs) have transformed natural language processing, demonstrating strong performance on tasks ranging from open-domain question answering (QAing) to complex reasoning (Izacard & Grave, 2021; Wei et al., 2022). Yet a fundamental question persists: do these models genuinely understand the context they process, or do they exploit statistical correlations to produce plausible outputs without true comprehension? This is especially con-
2.2. Graph Similarity Methods Embedding-based approaches such as TransE (Bordes et al., 2013) and RotatE (Sun et al., 2019) learn entity representations over fixed vocabularies, making them illsuited for cross-graph comparison where entity sets are disjoint. The WL graph kernel (Shervashidze et al., 2011) offers training-free structural comparison but treats labels as opaque symbols, penalizing semantically equivalent but lexically distinct terms. The WWL kernel (Togninalli et al., 2019) partially addresses this with Wasserstein node-distribution comparison, yet still lacks semantic label grounding. KEA (Haskins & Adams, 2025) combines WL kernels with SBERT (Reimers & Gurevych, 2019) clus-
*
Equal contribution 1 Department of Computer Science and Engineering, University of Moratuwa, Moratuwa, Sri Lanka 2 Department of Electrical and Computer Engineering, University of Miami, Coral Gables, Florida, USA. Correspondence to: Mamta Nallaretnam <[email protected]>. Proceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026. Copyright 2026 by the author(s).
1
Evaluation of Contextual Understanding in Large Language Models
tering to bridge the lexical gap, but its cluster-merge is lossy because merging semantically close labels discards fine-grained relational distinctions. Taken together, these methods either treat labels as opaque symbols or rely on lossy clustering, leaving a clear need for a similarity measure that preserves fine-grained relational distinctions while handling semantic variation.
using the few-shot prompting strategy with instruction tuning from Sansford et al. (2024) and Haskins & Adams (2025) applied uniformly across all three sources to ensure a consistent label space. Following extraction, all entity and relation labels are normalized via lowercasing and lemmatization to eliminate residual surface-level variation, ensuring the three graphs are structurally compatible for meaningful comparison.
2.3. Our Contributions 3.2. S3KG: Semantic Structural Similarity for KGs
Our work makes three main contributions.
S3KG computes similarity between two graphs G1 and G2 at two levels: at the triplet level, individual facts are matched for fine-grained correspondence; at the graph level, overall topology is compared. Both structural and semantic similarity are considered at each level. The full pipeline is illustrated in Appendix A (Figure 2).
• Semantic structural similarity for KGs (S3KG) is a hybrid structural–semantic similarity metric that converts LLM responses and reference answers into knowledge graph (KG) triplets and produces a continuous interpretable evaluation score.
Triplet-Level Matching. Each triplet (h, r, t) is serialized into a natural language (NL) string and encoded with SBERT (Reimers & Gurevych, 2019) (paraphrase-MPNet-base-v2). For each triplet in T1 , the most semantically similar triplet in T2 is identified by cosine similarity, forming a filtered set Tb2 ⊆ T2 that anchors the structural comparison to semantically relevant content.
• The Contextual Understanding Score (CUS) is a model-level aggregate of two complementary dimensions: factual accuracy and contextual faithfulness, enabling cross-model comparison across benchmarks. • The Triplet Analyzing Unit (TAU) is a diagnostic component that categorizes reasoning errors at the triplet level for fine-grained behavioral analysis.
Soft Label Alignment. Standard Weisfeiler–Lehman (WL) kernels treat lexically distinct but semantically equivalent labels as entirely disjoint. S3KG resolves this by independently aligning entity and relation labels: each label ℓ in G1 is replaced by its closest counterpart in G2 (by SBERT cosine similarity) whenever the similarity exceeds a threshold (we use 0.65); otherwise it is left unchanged. Entity and relation labels are aligned separately to prevent cross-type e 1 and G e2 . collisions, yielding aligned graphs G
3. Methodology
Structural Similarity via WL Kernel. The normalized WL kernel (Shervashidze et al., 2011) score, with multiple iterations (we use 5) to capture multi-hop neighbourhood patterns over the aligned graphs) yields structural similarity as
Figure 1. KG-based evaluation pipeline. Given a QA pair, three KGs are constructed from the LLM response, gold answer, and supporting context, then compared via S3KG to produce GoldSim, CtxSim, and CUS.
e1 , G e2 ) = q SWL (G
e1 , G e2 ) K(G
.
(1)
e1 , G e 1 ) · K(G e2 , G e2 ) K(G
As illustrated in Fig.1, initially the candidate LLM is prompted with the question alongside its relevant context. The generated response is collected as the LLM answer, which together with the ground truth answer and the supporting context is used for the knowledge graph construction.
Semantic Similarity via SBERT Mean-Pool. Each graph is represented by the mean SBERT embedding of its triples. The semantic similarity SSBERT (T1 , Tb2 ) is the cosine between these pooled representations, clipped to [0, 1] to discard negative correlations.
3.1. Knowledge Graph Construction
Final Score. The structural and semantic scores are blended via mixing coefficient α ∈ [0, 1] (we use α = 0.5) as
Three KGs are constructed per QA instance from the gold answer, model-generated response, and supporting context,
SS3KG = (1 − α) SWL + α SSBERT . 2
(2)
Evaluation of Contextual Understanding in Large Language Models
3.3. Triplet Analyzing Unit (TAU)
KG-derived paragraphs (69–126 words) across encyclopedic (Safavi & Koutra, 2021), financial (Li & Sanna Passino, 2024), biological (Poelen et al., 2014), and food ontology (Boudin et al., 2023) domains. The Wikipedia EntitySwap dataset (399 pairs) serves as an anti-circularity control, constructed via NLP-based perturbations with no KG involvement (Sennrich et al., 2016; Wei & Zou, 2019). Baselines include ROUGE-1/2/L (Lin, 2004), BLEU (Papineni et al., 2002), BERTScore (Zhang et al., 2020), MiniLM (Wang et al., 2020), and Sentence-T5base (Ni et al., 2022). S3KG is evaluated with α ∈ {0.0, 0.1, . . . , 1.0}; we report the best-performing α per dataset alongside AUROC, with the full sweep in Appendix B.
For the 5% of lowest-similarity QA pairs, a triplet analysis unit is applied to examine how the LLM-generated KG diverges from the ground truth. Each triplet is converted into an NL sentence and encoded using a Sentence Transformer, and cosine similarity is computed between ground truth and LLM triplets, with scores above a threshold (we use 0.76) treated as aligned. After removing aligned triplets, the remaining pairs are categorized into four error types: relation wrong (entities match but relation differs), entity wrong (relation aligns but entities differ), extra triplets (hallucinated by the LLM), and missing triplets (not captured by the LLM).
Results. Table 1 reports F1 and AUROC across all benchmarks. S3KG achieves top-1 or top-2 F1 on 7 of 9 datasets. On KG-perturbed paragraphs, structural signals are most valuable: S3KG gains up to +7.6 F1 points (SK-GloBI) and correctly captures relational role distinctions where ROUGE-1 collapses to near-random (PAWS-Wiki AUC = 0.490, S3KG F1 = 0.766). The exception is SK-FindKG, where financial vocabulary introduces KG extraction noise and Sentence-T5-base leads (F1 = 0.848), highlighting extraction quality as a bottleneck. On short texts, dense models dominate due to sparse relational structure (Sentence-T5base: MRPC = 0.766, STS12 = 0.853). On the Wikipedia Entity-Swap control, S3KG achieves the strongest result (F1 = 0.872 vs MiniLM 0.821), confirming gains are not an artifact of the KG pipeline. Optimal α varies by dataset and is further discussed in Appendix B.
3.4. Datasets and Models Two datasets are employed in this study for their longform answer coverage: PubMedQA (Jin et al., 2019), comprising biomedical QA pairs drawn from research articles, and MesaQA (Wang et al., 2025), comprising consumer healthcare QA pairs requiring multi-span evidence integration. Three instruction-tuned 7B-parameter models are evaluated: Llama-2-7b-chat-hf, Gemma-7b-it, and Mistral-7B-Instruct-v0.2, with Falcon-7B included as a baseline. Responses are collected across a temperature sweep of {0.0, 0.3, 0.7, 1.0} to analyze the effect of generation stochasticity on contextual faithfulness. Full temperature sensitivity results appear in Appendix C.
4. Experiments and Results Our evaluation targets the full KG-based LLM evaluation pipeline (Fig. 1), which comprises two core components: (1) KG construction, which extracts structured representations from the LLM response, gold answer, and supporting context; and (2) S3KG similarity module, which compares these graphs to produce the Comparative LLM Understanding Score (CUS). To rigorously assess S3KG’s similarity scoring in isolation, we first benchmark it independently across nine semantic equivalence datasets in §4.1, before evaluating the complete pipeline on QA benchmarks in §4.3.
4.2. TAU Evaluation Table 2 illustrates the TAU performace, on a manually annotated subset of the MesaQA and PubMed datasets, where aligned triplet pairs between ground-truth and LLMgenerated KGs were labeled by three medical students, achieving strong agreement (pairwise F1: 0.97 − −0.99). Our method uses sentence-level semantic similarity for alignment, while the KEA baseline relies on componentwise matching. Our approach improves Micro F1 (0.895 vs 0.836) and Macro F1 (0.782 vs 0.622), driven by a +34.8% recall gain with a modest drop in precision. This enables robust alignment of semantically equivalent relations despite lexical variation, which are often missed by KEA. Overall, these results demonstrate that our method outperforms KEA in TAU performance.
4.1. S3KG Benchmark Evaluation Datasets and Task Formulation. We evaluate on nine datasets spanning three structural categories, each cast as a binary semantic equivalence task (N =400, balanced), with performance measured by maximum F1 via threshold sweep. The short-text category comprises MRPC (Dolan & Brockett, 2005), PAWS-Wiki (Zhang et al., 2019), and STS12 (Agirre et al., 2012) (10–22 words). The KGperturbed paragraph category comprises five datasets (SK-Codex 400, SK-Combined, SK-FindKG, SK-GloBI, SK-Oregano) built by perturbing entity relationships in 3
Evaluation of Contextual Understanding in Large Language Models Table 1. F1 / AUROC scores across all benchmark datasets. Bold indicates best value per column. S3KG uses the best-performing α per dataset selected by grid search (full sweep in Appendix B). † ROUGE-1 excluded from Wiki Swap due to surface-form artefact. Best α: MRPC 0.3, PAWS 0.5, STS12 0.1, C400 0.5, Comb. 0.5, Find 0.0, GloBI 0.6, Oreg. 0.4, W-Swap 0.1. Short Text Method
MRPC
PAWS
KG-Perturbed Paragraphs STS12
C400
Comb.
Find
Anti-Circ.
GloBI
Oreg.
W-Swap
S3KG (Ours) 0.692/0.673 0.766/0.795 0.786/0.834 0.932/0.973 0.834/0.829 0.767/0.796 0.892/0.935 0.812/0.892 0.872/0.890 ROUGE-1 0.745/0.784 0.678/0.490 0.725/0.754 0.835/0.917 0.732/0.728 0.745/0.745 0.784/0.833 0.745/0.782 —† ROUGE-2 0.720/0.721 0.715/0.721 0.681/0.656 0.822/0.894 0.707/0.711 0.717/0.706 0.776/0.833 0.752/0.791 0.860/0.772 ROUGE-L 0.729/0.760 0.735/0.807 0.703/0.710 0.792/0.855 0.722/0.717 0.719/0.721 0.763/0.800 0.792/0.835 0.729/0.311 BLEU 0.687/0.677 0.716/0.747 0.671/0.644 0.806/0.884 0.715/0.708 0.711/0.710 0.775/0.819 0.745/0.794 0.868/0.745 BERTScore 0.758/0.816 0.691/0.702 0.682/0.636 0.823/0.916 0.757/0.792 0.739/0.761 0.816/0.871 0.743/0.798 0.747/0.645 MiniLM 0.723/0.748 0.687/0.638 0.833/0.894 0.875/0.943 0.770/0.789 0.802/0.844 0.780/0.817 0.773/0.814 0.821/0.811 sent-T5 0.766/0.816 0.674/0.668 0.853/0.928 0.876/0.944 0.770/0.827 0.848/0.902 0.728/0.760 0.797/0.871 0.762/0.806 C400: SK-Codex 400; Comb.: SK-Combined; Find: SK-FindKG; Oreg.: SK-Oregano; W-Swap: Wikipedia Entity-Swap. Table 3. LLM evaluation results (mean over N =400 samples, α=0.5).
Table 2. TAU performance vs. KEA baseline.
Metric
KEA
Ours
Micro Precision Micro Recall Macro Precision Macro Recall Micro F1 Macro F1
0.853 0.738 0.953 0.641 0.836 0.622
0.847 0.949 0.803 0.946 0.895 0.782
4.3. LLM Contextual Understanding Evaluation
Dataset
Model
GoldSim CtxSim
CUS
MesaQA
Gemma-7B Llama-2-7B Mistral-7B Falcon-7B
0.6889 0.6570 0.6491 0.6036
0.7031 0.7236 0.7356 0.6458
0.6774 0.6752 0.6780 0.6063
PubMedQA
Gemma-7B Llama-2-7B Mistral-7B Falcon-7B
0.5235 0.5220 0.5138 0.4541
0.6401 0.6567 0.7331 0.5560
0.5587 0.5651 0.5923 0.4800
To evaluate LLMs, S3KG is applied across three KGs (i) (i) (i) per QA instance i: KG LLM , KG gold , and KG ctx . Gold(i)
(i)
Sim(i) = SS3KG (KG LLM , KG gold ) measures factual accu(i)
(i)
racy; CtxSim(i) = SS3KG (KG LLM , KG ctx ) measures contextual faithfulness. Since neither alone reflects true understanding, CUS penalises imbalance between the two as CUS(i) =
2 · GoldSim(i) · CtxSim(i) . GoldSim(i) + CtxSim(i)
and across all metrics. On the 5% lowest similarity pairs, Mistral-7B remains the most reliable, recovering more reference-aligned triplets than other models, while Gemma-7B is the weakest, particularly on MesaQA. PubMedQA hard cases show sharply lower aligned-triplet recovery across all models, confirming that biomedical terminology and entity variability dominate failure modes.
(3)
The dataset-level score is the mean of CUS(i) over all N samples. Table 3 reports mean GoldSim, CtxSim, and CUS over N = 400 samples per model–dataset combination at α = 0.5. Mistral-7B leads on CUS across both datasets (MesaQA: 0.678; PubMedQA: 0.592), driven by the highest CtxSim in each case (0.736 and 0.733), reflecting strong contextual faithfulness. CtxSim exceeds GoldSim in all eight model–dataset combinations, indicating that instruction-tuned models systematically elaborate on context rather than producing concise reference-style responses. A consistent ≈10 percentage-point CUS gap between MesaQA and PubMedQA across all models reveals domain complexity, biomedical vocabulary, and reasoning as key bottlenecks rather than model size or architecture. Falcon-7B shows the weakest performance on both datasets
5. Discussion S3KG moves beyond surface-level metrics via typeseparated, one-to-one label alignment and WL kernel multihop sensitivity. The GoldSim/CtxSim decomposition exposes model-specific trade-offs invisible to aggregate metrics; the consistent PubMedQA–MesaQA domain gap confirms Mistral-7B handles domain-specific health knowledge more reliably than general health QA. KG extraction quality remains the primary bottleneck, and future work will incorporate GPT-4 and attention-based directional alignment to better distinguish semantically inverse relations. 4
Evaluation of Contextual Understanding in Large Language Models
6. Conclusion
Izacard, G. and Grave, E. Leveraging passage retrieval with generative models for open domain question answering. In EMNLP, 2021.
We presented a KG-based evaluation pipeline combining S3KG and the TAU for reproducible, interpretable measurement of LLM contextual understanding in QA. S3KG achieves best-per-dataset F1 of 0.766–0.932 and AUROC up to 0.973, consistently outperforming lexical and neural baselines on KG-rich datasets while remaining competitive on short-text settings. The framework is dataset-agnostic and readily extensible, supporting broader efforts toward trustworthy and verifiable AI evaluation.
Jin, Q., Dhingra, B., Liu, Z., Cohen, W., and Lu, X. Pubmedqa: A dataset for biomedical research question answering. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 2567–2577, 2019. Li, X. V. and Sanna Passino, F. FinDKG: Dynamic knowledge graphs with large language models for detecting global trends in financial markets. In Proceedings of the 5th ACM International Conference on AI in Finance (ICAIF 2024), pp. 573–581, 2024. doi: 10.1145/3677052.3698603. URL https://arxiv. org/abs/2407.10909. Financial knowledge graph extracted from news articles using LLMs.
Acknowledgments The work of Kamal Premaratne (KP) was supported by the U.S. National Science Foundation (NSF) under Award No. 2530256. The authors also acknowledge the developers and open-source communities behind the Falcon, Mistral, Llama, and Gemma large language models, as well as the Hugging Face platform for providing access to open-source models and tools that supported this research.
Lin, C.-Y. ROUGE: A package for automatic evaluation of summaries. In Text Summarization Branches Out: Proceedings of the ACL-04 Workshop, pp. 74–81, 2004.
References Ni, J., Ábrego, G. H., Constant, N., Ma, J., Hall, K. B., Chang, M., and Yang, Y. Sentence-t5: Scalable sentence encoders from pre-trained text-to-text models. In Findings of the Association for Computational Linguistics: ACL 2022, pp. 272–279, 2022.
Agirre, E. et al. SemEval-2012 task 6: A pilot on semantic textual similarity. In Proceedings of the First Joint Conference on Lexical and Computational Semantics (*SEM), 2012. Bordes, A., Usunier, N., Garcı́a-Durán, A., Weston, J., and Yakhnenko, O. Translating embeddings for modeling multi-relational data. In Advances in Neural Information Processing Systems, volume 26, pp. 2787–2795. Curran Associates, Inc., 2013.
Papineni, K., Roukos, S., Ward, T., and Zhu, W.-J. Bleu: A method for automatic evaluation of machine translation. In Proceedings of the 40th Annual Meeting of the Association for Computational Linguistics, pp. 311–318, Philadelphia, PA, USA, 2002.
Boudin, M., Diallo, G., Drancé, M., and Mougin, F. The OREGANO knowledge graph for computational drug repurposing. Scientific Data, 10:871, 2023. doi: 10.1038/ s41597-023-02757-0. URL https://www.nature. com/articles/s41597-023-02757-0. Food ontology and natural compound knowledge graph for drug repurposing.
Poelen, J. H., Simons, J. D., and Mungall, C. J. GloBI: Global biotic interactions. [Online]. Available: https: //www.globalbioticinteractions.org, 2014. Accessed: Jan. 15, 2025. Reimers, N. and Gurevych, I. Sentence-BERT: Sentence embeddings using Siamese BERT-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP), pp. 3982–3992, Hong Kong, China, 2019. Association for Computational Linguistics. doi: 10.18653/v1/D19-1410.
Dolan, W. B. and Brockett, C. Automatically constructing a corpus of sentential paraphrases. In Proceedings of the Third International Workshop on Paraphrasing (IWP2005), 2005. Edge, D., Trinh, H., Cheng, N., Bradley, J., Chao, A., Mody, A., Truitt, S., Metropolitansky, D., Ness, R. O., and Larson, J. From local to global: A graph RAG approach to query-focused summarization. arXiv preprint arXiv:2404.16130, 2024.
Safavi, T. and Koutra, D. Codex: A comprehensive knowledge graph completion benchmark. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, 2021.
Haskins, R. and Adams, B. Kea explain: Explanations of hallucinations using graph kernel analysis. arXiv preprint arXiv:2507.03847, 2025.
Sansford, H., Richardson, N., Maretic, H. P., and Saada, J. N. Grapheval: A knowledge-graph based llm hallucination 5
Evaluation of Contextual Understanding in Large Language Models
evaluation framework. arXiv preprint arXiv:2407.10793, 2024.
Yan, J., Wang, C., Huang, J., and Zhang, W. Do large language models understand logic or just mimick context? arXiv preprint arXiv:2402.12091, 2024.
Sennrich, R., Haddow, B., and Birch, A. Neural machine translation of rare words with subword units. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (ACL), pp. 1715–1725. Association for Computational Linguistics, 2016. doi: 10.18653/v1/P16-1162. URL https: //aclanthology.org/P16-1162/. Subword segmentation method used for node replacement perturbations.
Zhang, T., Kishore, V., Wu, F., Weinberger, K. Q., and Artzi, Y. BERTScore: Evaluating text generation with BERT. In International Conference on Learning Representations (ICLR), 2020. Zhang, Y., Baldridge, J., and He, L. PAWS: Paraphrase adversaries from word scrambling. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pp. 1702–1711, 2019.
Shervashidze, N., Schweitzer, P., van Leeuwen, E. J., Mehlhorn, K., and Borgwardt, K. M. Weisfeiler-Lehman graph kernels. Journal of Machine Learning Research, 12:2539–2561, 2011.
Zhu, Y., Moniz, J. R. A., Bhargava, S., Lu, J., Piraviperumal, D., Li, S., Zhang, Y., Yu, H., and Tseng, B.-H. Can large language models understand context? In Findings of the Association for Computational Linguistics: EACL 2024, pp. 2004–2018, mar 2024.
Sun, Z., Deng, Z.-H., Nie, J.-Y., and Tang, J. RotatE: Knowledge graph embedding by relational rotation in complex space. In Proceedings of the 7th International Conference on Learning Representations (ICLR), 2019. Togninalli, M., Ghisu, E., Llinares-López, F., Rieck, B., and Borgwardt, K. Wasserstein Weisfeiler–Lehman graph kernels. In Advances in Neural Information Processing Systems, volume 32, pp. 6439–6449. Curran Associates, Inc., 2019. Wang, J.-I., Huang, H.-H., and Chen, H.-H. MESAQA: A dataset for multi-span contextual and evidencegrounded question answering. In Proceedings of the 31st International Conference on Computational Linguistics, pp. 10891–10901, Abu Dhabi, UAE, January 2025. Association for Computational Linguistics. URL https://aclanthology.org/2025. coling-main.724/. Wang, W., Wei, F., Dong, L., Bao, H., Yang, N., and Zhou, M. MiniLM: Deep self-attention distillation for taskagnostic compression of pre-trained transformers. In Advances in Neural Information Processing Systems, volume 33, pp. 5776–5788, 2020. Wei, J. and Zou, K. EDA: Easy data augmentation techniques for boosting performance on text classification tasks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. 6382–6388. Association for Computational Linguistics, 2019. doi: 10.18653/v1/D19-1670. URL https: //aclanthology.org/D19-1670/. NLP-based perturbation techniques including deletion operations. Wei, J., Wang, X., Schuurmans, D., Bosma, M., Ichter, B., Xia, F., Chi, E., Le, Q., and Zhou, D. Chain-of-thought prompting elicits reasoning in large language models. In Advances in Neural Information Processing Systems (NeurIPS), 2022. 6
Evaluation of Contextual Understanding in Large Language Models
A. S3KG Similarity Pipeline
Figure 2. S3KG similarity pipeline. Triple-level SBERT matching identifies semantically relevant triples in G2 ; soft label alignment resolves lexical mismatches between entity and relation labels; the normalised WL kernel and SBERT mean-pool scores are blended via mixing coefficient α to produce the final S3KG score.
B. Hyperparameter α selection This appendix reports the complete S3KG α sweep results across all evaluation datasets. Each table shows F1 and AUROC for α ∈ {0.0, 0.1, . . . , 1.0}, where α = 0.0 recovers a pure KG structural embedding and α = 1.0 recovers a pure dense sentence-transformer representation, as defined in Equation 2. The best-performing α per dataset (by F1) is shown in bold. B.1 Short-Text Datasets Table 4 presents the α sweep for MRPC, PAWS-Wiki, and STS12. The optimal α differs notably across datasets: MRPC peaks at α = 0.3, PAWS-Wiki at α = 0.5, and STS12 at α = 0.1. The consistently low optimal values indicate that retaining a meaningful KG structural component is beneficial for short-text paraphrase detection, and that a pure dense-embedding representation (α = 1.0) is suboptimal for all three tasks. 7
Evaluation of Contextual Understanding in Large Language Models Table 4. S3KG α sweep on short-text datasets (F1 / AUROC). The best α per dataset by F1 is shown in bold.
MRPC
PAWS-Wiki
STS12
α
F1
AUROC
F1
AUROC
F1
AUROC
0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0
0.676 0.681 0.683 0.692 0.680 0.680 0.681 0.688 0.684 0.683 0.681
0.654 0.664 0.669 0.673 0.674 0.675 0.674 0.674 0.672 0.671 0.671
0.694 0.745 0.760 0.764 0.764 0.766 0.764 0.764 0.764 0.764 0.764
0.739 0.781 0.790 0.793 0.795 0.795 0.795 0.795 0.795 0.795 0.731
0.783 0.786 0.780 0.780 0.780 0.775 0.768 0.764 0.762 0.756 0.756
0.828 0.834 0.834 0.831 0.827 0.824 0.820 0.815 0.810 0.805 0.790
B.2 KG-Perturbed Paragraph Datasets Table 5 presents results for the five KG-perturbed paragraph datasets. Four of the five datasets favour a balanced blend of structural and dense signal: SK-Codex 400 and SK-Combined peak at α = 0.5, SK-GloBI at α = 0.6, and SK-Oregano at α = 0.4. SK-FindKG is the single exception, where the pure KG embedding (α = 0.0) yields the highest F1 of 0.767, suggesting that the structural signal in FindKG is particularly discriminative and is diluted rather than enhanced by dense representations.
Table 5. S3KG α sweep on KG-perturbed paragraph datasets (F1 / AUROC). The best α per dataset by F1 is shown in bold.
SK-Codex 400
SK-Combined
SK-FindKG
SK-GloBI
SK-Oregano
α
F1
AUROC
F1
AUROC
F1
AUROC
F1
AUROC
F1
AUROC
0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0
0.871 0.922 0.927 0.927 0.929 0.932 0.932 0.932 0.932 0.932 0.932
0.969 0.973 0.973 0.973 0.973 0.973 0.973 0.973 0.973 0.973 0.932
0.776 0.791 0.819 0.828 0.833 0.834 0.832 0.833 0.832 0.828 0.828
0.801 0.825 0.829 0.830 0.829 0.829 0.828 0.828 0.828 0.827 0.799
0.767 0.760 0.762 0.748 0.744 0.741 0.736 0.734 0.731 0.729 0.728
0.796 0.798 0.793 0.788 0.783 0.779 0.776 0.772 0.769 0.765 0.756
0.811 0.883 0.888 0.889 0.891 0.890 0.892 0.892 0.892 0.892 0.892
0.898 0.934 0.937 0.936 0.936 0.935 0.935 0.935 0.935 0.934 0.917
0.777 0.803 0.811 0.812 0.812 0.810 0.812 0.812 0.810 0.810 0.810
0.884 0.892 0.892 0.892 0.892 0.892 0.892 0.891 0.891 0.891 0.770
B.3 Wikipedia Entity-Swap (Anti-Circularity Control) Table 6 presents the α sweep for the Wikipedia Entity-Swap dataset, which serves as an anti-circularity control. The best performance is achieved at α = 0.1 (F1 = 0.872, AUROC = 0.890, Precision = 0.780), confirming that even a small contribution from the KG structural embedding improves over the pure dense baseline, while heavier structural weighting (α ≥ 0.2) offers no further benefit. 8
Evaluation of Contextual Understanding in Large Language Models Table 6. S3KG α sweep on Wikipedia Entity-Swap (F1 / AUROC / Precision / Recall). The best α by F1 is shown in bold.
α
F1
AUROC
Precision
Recall
0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0
0.821 0.872 0.869 0.868 0.865 0.868 0.865 0.865 0.865 0.865 0.865
0.892 0.890 0.890 0.890 0.890 0.890 0.889 0.889 0.889 0.889 0.852
0.698 0.780 0.783 0.773 0.777 0.773 0.777 0.777 0.777 0.777 0.777
0.995 0.990 0.975 0.990 0.975 0.990 0.975 0.975 0.975 0.975 0.975
C. Temperature Sensitivity Results The sampling temperature is a key hyperparameter governing how deterministically a language model generates text. A temperature of zero corresponds to greedy decoding, where the model always selects the most probable next token, yielding highly consistent and context-adherent outputs. As temperature increases, the sampling distribution becomes broader, allowing the model to explore a wider range of responses. Low temperatures such as 0.3 preserve factual grounding while introducing modest lexical variation, whereas a mid-range value of 0.7 is commonly adopted in practice to balance fluency and diversity. At temperature 1.0, the model samples directly from its raw output distribution, producing the most varied responses but with a greater risk of factual drift away from the provided context. We selected T ∈ {0.0, 0.3, 0.7, 1.0} to span the full practical operating range of instruction-tuned models, from fully deterministic inference to high-entropy generation. This allows us to examine whether contextual faithfulness, as measured by GoldSim and CtxSim, is robust to generation stochasticity or degrades meaningfully as randomness increases. Tables 7 and 8 report these results. Most models remain stable within ±0.01–0.02 across all settings; the notable exception is Falcon-7B on MesaQA, where GoldSim falls from 0.6036 at T =0.0 to 0.4662 at T =1.0, indicating that higher sampling randomness substantially degrades factual alignment for this model. Table 7. Mean GoldSim per model across temperatures.
Dataset MesaQA
PubMedQA
Model
T=0.0
T=0.3
T=0.7
T=1.0
Llama-2-7B Gemma-7B Mistral-7B Falcon-7B
0.6570 0.6889 0.6491 0.6036
0.6588 0.7060 0.6567 0.6085
0.6580 0.7015 0.6523 0.5558
0.6406 0.6931 0.6444 0.4662
Llama-2-7B Gemma-7B Mistral-7B Falcon-7B
0.5220 0.5235 0.5138 0.4541
0.5185 0.5149 0.5121 0.4330
0.5143 0.5153 0.5085 0.4184
0.5087 0.5204 0.5023 0.3852
This appendix provides complete benchmark results referenced in the main paper, including short-text datasets (Section 4.1), KG-perturbed paragraph datasets, the Wikipedia Entity-Swap anti-circularity control, and a summary heatmap visualization. D.1 Short-Text Datasets Table 9 presents the full F1 and AUROC results for MRPC, PAWS-Wiki, and STS12. S3KG achieves competitive performance, with sentence-T5-base leading on STS12 due to the short-text nature of the dataset. 9
Evaluation of Contextual Understanding in Large Language Models Table 8. Mean CtxSim per model across temperatures.
Dataset MesaQA
PubMedQA
Model
T=0.0
T=0.3
T=0.7
T=1.0
Llama-2-7B Gemma-7B Mistral-7B Falcon-7B
0.7236 0.7031 0.7356 0.6458
0.7212 0.7217 0.7585 0.6571
0.7248 0.7045 0.7287 0.6122
0.7026 0.7013 0.7223 0.5195
Llama-2-7B Gemma-7B Mistral-7B Falcon-7B
0.6567 0.6401 0.7331 0.5560
0.6563 0.6540 0.7352 0.5165
0.6594 0.6434 0.7314 0.5120
0.6468 0.6319 0.7003 0.4571
Table 9. Performance on short-text datasets (F1 / AUROC).
Method
MRPC
PAWS-Wiki
STS12
S3KG
0.692 / 0.673 0.766 / 0.795 0.786 / 0.834
ROUGE-1 ROUGE-2 ROUGE-L BLEU BERTScore MiniLM sentence-T5-base
0.745 / 0.784 0.720 / 0.721 0.729 / 0.760 0.687 / 0.677 0.758 / 0.816 0.723 / 0.748 0.766 / 0.816
0.678 / 0.490 0.715 / 0.721 0.735 / 0.807 0.716 / 0.747 0.691 / 0.702 0.687 / 0.638 0.674 / 0.668
0.725 / 0.754 0.681 / 0.656 0.703 / 0.710 0.671 / 0.644 0.682 / 0.636 0.833 / 0.894 0.853 / 0.928
D.2 KG-Perturbed Paragraph Datasets Table 10 reports results for the five KG-perturbed paragraph datasets. S3KG achieves the highest F1 on four of five datasets, with gains up to +7.6 F1 points on SK-GloBI. Table 10. Performance on KG-perturbed paragraph datasets (F1 / AUROC).
Method S3KG
C400
Comb.
Find
GloBI
Oreg.
0.932/0.973 0.834/0.829 0.767/0.796 0.892/0.935 0.812/0.892
ROUGE-1 0.835/0.917 0.732/0.728 0.745/0.745 0.784/0.833 0.745/0.782 ROUGE-2 0.822/0.894 0.707/0.711 0.717/0.706 0.776/0.833 0.752/0.791 ROUGE-L 0.792/0.855 0.722/0.717 0.719/0.721 0.763/0.800 0.792/0.835 BLEU 0.806/0.884 0.715/0.708 0.711/0.710 0.775/0.819 0.745/0.794 BERTScore 0.823/0.916 0.757/0.792 0.739/0.761 0.816/0.871 0.743/0.798 MiniLM 0.875/0.943 0.770/0.789 0.802/0.844 0.780/0.817 0.773/0.814 sent-T5 0.876/0.944 0.770/0.827 0.848/0.902 0.728/0.760 0.797/0.871 C400: SK-Codex 400; Comb.: SK-Combined; Find: SK-FindKG; Oreg.: SK-Oregano.
D.3 Wikipedia Entity-Swap Results Table 11 presents results on the Wikipedia Entity-Swap anti-circularity control. S3KG achieves the highest F1 (0.872) and AUC (0.890), confirming that gains are not an artefact of the KG construction pipeline.1 1 ROUGE-1 is excluded due to a surface-form artefact: entity-swapped pairs share nearly all surrounding tokens, making unigram overlap trivially near-perfect.
10
Evaluation of Contextual Understanding in Large Language Models Table 11. Results on Wikipedia Entity-Swap (N = 400).
Method
F1
AUC
Prec.
Rec.
S3KG
0.872
0.890
0.780
0.990
ROUGE-2 ROUGE-L BLEU BERTScore MiniLM sentence-T5-base
0.860 0.729 0.868 0.747 0.821 0.762
0.772 0.311 0.745 0.645 0.811 0.806
0.773 0.573 0.776 0.615 0.729 0.621
0.970 1.000 0.985 0.950 0.940 0.985
D.4 Best S3KG Variant Summary Table 12 summarizes the best-performing α variant per dataset, selected by maximum F1 via grid search over α ∈ {0.0, 0.1, . . . , 1.0}. Table 12. Best S3KG variant per dataset.
Dataset
F1/AUROC
Verdict
MRPC (α = 0.3) 0.692/0.673 Behind (sparse KG) PAWS-Wiki (α = 0.5) 0.766/0.795 Strong (+3.1 F1) STS12 (α = 0.1) 0.786/0.834 Behind (short text) SK-Codex 400 (α = 0.5) 0.932/0.973 Strong (+5.7 F1) SK-Combined (α = 0.5) 0.834/0.829 Strong (+6.0 F1) SK-FindKG (α = 0.0) 0.767/0.796 Behind (domain noise) SK-GloBI (α = 0.6) 0.892/0.935 Strong (+7.6 F1) SK-Oregano (α = 0.4) 0.812/0.892 Comparable (+1.2) Wiki Swap (α = 0.1) 0.872/0.890 Best meaningful score
D.5 Performance Heatmap Figure 3 visualizes the F1 and AUROC scores of S3KG and all baseline methods across all benchmark datasets. Gold borders indicate the best-performing method per dataset. This appendix provides a detailed worked example from the PAWS-Wiki dataset, illustrating how S3KG correctly identifies semantic opposition where ROUGE-1 fails. Table 13 compares a positive (similar) pair and a negative (not similar) pair. In the positive example, word-order variation yields identical KG triplets and simα = 1.0, correctly predicting Similar. In the negative example, reversed subject–object roles expose semantic opposition despite identical surface tokens: ROUGE-1 incorrectly predicts Similar, while S3KG correctly predicts Not Similar.
11
Evaluation of Contextual Understanding in Large Language Models
Figure 3. F1 score (left) and AUROC (right) of S3KG and baseline methods across all benchmark datasets. Gold borders indicate the best-performing method per dataset. S3KG (navy border, leftmost column) represents the best-performing α variant selected per dataset from the full α sweep (Appendix B).
12
Evaluation of Contextual Understanding in Large Language Models
Table 13. S3KG pipeline scores for two PAWS-Wiki sentence pairs.
Positive Example
Negative Example
Text s1
His father returned as a finished violinist of the Renzo Furlan won 6–3, 6–4 against Thomas Russian School to Bombay. Johansson in the finals.
Text s2
His father returned to Bombay as a finished Thomas Johansson won 6–3, 6–4 against Renzo violinist of the Russian school. Furlan in the finals.
Triplet (s1 ) Triplet (s2 )
(father, returned to, Bombay)
(Renzo Furlan, won against, Thomas Johansson)
(father, returned to, Bombay)
(Thomas Johansson, won against, Renzo Furlan)
simTP (α=0) simST (α=1) simα (α=0.5) Threshold τ
1.0000 1.0000 1.0000 0.92
0.8310 0.1361 0.4835 0.92
S3KG predic- Similar ✓ tion ROUGE-1 1.0000 score ROUGE-1 pre- Similar ✓ diction True label
Not Similar ✓ 1.0000 Similar ×
Positive (1)
Negative (0)
13