Proceedings of Machine Learning Research vol 284:1–24, 2026 20th Conference on Neurosymbolic Learning and Reasoning
A Neurosymbolic Approach for Explainable Early Diagnosis of Alzheimer’s Disease Ranveer Singh∗ Pranuthi Tenali
[email protected] [email protected]
The University of Texas at Dallas, Richardson, TX
Saurabh Mathur
TU Darmstadt, Germany
Ameet Soni
Swarthmore College, PA
arXiv:2607.29530v1 [cs.LG] 31 Jul 2026
Vaishali Phatak Karla Lynch Daniel Murman Matthew Rizzo
[email protected] [email protected] [email protected] [email protected]
University of Nebraska Medical Center, NE
Sriraam Natarajan
The University of Texas at Dallas, Richardson, TX
Editors: Alessandra Mileo, Andrea Passerini and Cogan Shimizu
Abstract Identifying reliable Alzheimer’s disease (AD) markers typically requires manual, laborintensive transcription and expert analysis, limiting its scale. We introduce an automated pipeline that extracts qualitative knowledge about potential AD progression indicators directly from audio recordings of verbal fluency tests. Our method uses pretrained foundation models to process raw audio and extract clinically relevant variables to construct a Bayesian Network (BN); this BN is used to reason about the AD progression markers and infer their qualitative relationships. Our system successfully recovers known clinical knowledge and identifies novel relationships between linguistic markers.
1. Introduction Alzheimer’s disease (AD) is a leading cause of mortality among elderly populations, currently ranking as the fifth-leading cause of death for Americans over the age of 65 (adf, 2024). With prevalence projected to double by 2060, there is an urgent need for scalable methods to detect early cognitive symptoms. These symptoms often manifest as subtle linguistic and semantic deficits (Foster et al., 2013), including reduced verbal fluency (Marra et al., 2021) and increased word-finding difficulty (Salmon, 2011). Clinical practitioners typically evaluate these markers through standardized neuropsychological assessments, such as verbal fluency tests (Henry et al., 2004; Marczinski and Kertesz, 2006). However, the analysis of these tests remains a manual, labor-intensive bottleneck that requires expert transcription and qualitative analysis, severely limiting widespread longitudinal tracking. ∗
R Singh, P Tenali and S Mathur - equal contribution
© 2026 R. Singh, P. Tenali, S. Mathur, A. Soni, V. Phatak, K. Lynch, D. Murman, M. Rizzo & S. Natarajan.
Singh Tenali Mathur Soni Phatak Lynch Murman Rizzo Natarajan
total_words = LEN( FILTER(VALID,words) )
Neural Animals
[719.73 - 723.60] So
Extracted Response:
you'll have one minute
Timestamps:
all the animals that you can think of in one
Speech-toText (S2T) Features and Descriptions
minute. [729.97 - 749.55] Lions, tigers, ...
Transcription with timestamps
Symbolic
Lions, tigers, ...
and I want you to tell me
Large Language Models (LLMs)
CI M-
id, ne, ...
Start: 729
1, 0, ...
End: 790
2, 1, ...
NE
Vegetables Extracted Response: ...
Relevant Structured Output
Symbolic Program Executor
Dataset
QuaKE
Monotonicities
Influence Graph
NeuroSymbolic Symbol Grounding
Symbolic Inference
Figure 1: End-to-end pipeline for automated discovery of qualitative AD markers. Our framework extracts qualitative knowledge from raw audio data by combining foundation models with symbolic reasoning. It uses foundation models for robust transcription, information extraction, and for generating an influence network. These data are combined with the influence network to instantiate a Bayesian Network (BN). The QuaKE algorithm uses this BN to infer underlying monotonicities between linguistic markers and disease progression.
The fundamental challenge in automating this discovery process lies in the representational gap between sensory input and clinical knowledge: cognitive impairment indicators are embedded within unstructured, continuous acoustic signals; clinical domain knowledge is structured around qualitative relationships between discrete variables. While neural architectures excel at perception and pattern recognition, they lack explicit mechanisms to ensure logically consistent, verifiable, and interpretable reasoning. Conversely, symbolic systems cannot process raw, high-dimensional audio. Bridging this gap needs a framework capable of both high-fidelity symbol grounding and rigorous reasoning. To this effect, we introduce NeSyQuaKE (Neurosymbolic Qualitative Knowledge Extraction). Following a Neuro | Symbolic (Type 3) architecture (Kautz, 2022), we use pretrained foundation models for symbol grounding and Bayesian Networks for reasoning about pairwise relationships between cognitive impairment and candidate markers. NeSyQuaKE identifies the nature of the relationships between markers, represented as Qualitative Influence Statements (QIs)(Kuipers, 1994; Karanam et al., 2021; Yang and Natarajan, 2013; Odom and Natarajan, 2018). These QIs concisely capture trends such as monotonicity, allowing easy communication and validation of discovered knowledge against medical literature. We make the following key contributions: (1) We propose the NeSyQuaKE Framework, a novel neurosymbolic architecture that extracts qualitative knowledge from clinical audio using foundation models and Bayesian Networks; (2) We introduce a layer of deterministic symbolic programs to map structured neural outputs to clinical features, ensuring that the derivation of metrics such as speech rate and semantic switching is mathematically consistent and verifiable; (3) We evaluate our system on a real-world dataset of clinical verbal fluency tests, demonstrating that NeSyQuaKE successfully recovers established clinical knowledge and generates additional hypotheses about qualitative relationships between cognitive impairment and its markers.
2. Background Neurosymbolic AI for Knowledge Discovery Neurosymbolic AI (Garcez and Lamb, 2023; Marra et al., 2024) integrates effective representation learning of neural networks 2
NeSyQuaKE
with the structured reasoning of symbolic logic. The integration of these two paradigms is often categorized based on the interaction between their neural and symbolic components. Our setting naturally aligns with the Neuro | Symbolic (Type 3) system from Kautz’s taxonomy (Kautz, 2022), because our task of clinical knowledge discovery from audio clips requires both pattern recognition and strict, verifiable symbolic reasoning. Instantiating such a system requires three critical components: (1) A Symbol Grounder: A mechanism to map raw, high-dimensional observational data into discrete symbols. We address this symbol grounding problem by using Pretrained Foundation Models to extract clinical variables from noisy audio. (2) A Symbolic Language: A formal representation to encode domain entities and their relationships. We use Bayesian Networks (BNs) as our formal language. (3) An Inference Procedure: A mathematically sound method to derive new knowledge from the grounded symbols. We employ the QuaKE algorithm to perform this reasoning. While symbol grounding has been addressed through learnable neural modules (Hsu et al., 2023), the paucity of clinical data and the high cost of annotation make this less suited for medical domains. A promising alternative is to use pretrained foundation models (Wüst et al., 2025). These models are effective at extracting information from unstructured data such as images. To this effect, we use pretrained speech and language foundation Models for intermediate symbol grounding and combine these with symbolic programs encoding domain knowledge to reliably map verbal fluency test recordings to variables of interest. However, instead of learning predictive models for cognitive impairment, we aim to infer explainable qualitative rules about the variables of interest and the nature of their relationship with cognitive impairment. Foundation Models Foundation Models can provide a robust solution to the symbol grounding problem by acting as high-fidelity neural predicates or feature extractors. Unlike traditional machine learning models trained on task-specific datasets, foundation models are pretrained on large corpora, enabling them to capture complex semantic and acoustic patterns required to extract clinically relevant information. We consider two types of these models: Speech to Text (S2T) and Large Language Models (LLMs). Speech-to-text (S2T) models are a subclass of speech foundation models typically built on the transformer (Vaswani et al., 2017) architecture that learn rich acoustic and semantic representations from large-scale audio data, specialized for tasks such as automatic speech recognition (ASR) (Ahlawat et al., 2025). For low-data settings such as clinical ASR where training or fine-tuinng these models is not possible, pre-trained transformer-based models such as Whisper (Radford et al., 2023) have been successful (Hwang et al., 2025). Large language models (LLMs) are a class of foundation models trained on large corpora of text and used for natural language processing tasks such question answering, summarization, and information extraction (Naveed et al., 2025; Xu et al., 2024). Particularly in clinical settings, domain-specific pre-trained LLMs such as MedGemma have been successful on information extraction tasks (Zhou and Yu, 2025). However, using the LLM outputs for symbolic grounding requires the outputs to conform to a predefined schema or format. This is often achieved through techniques like constrained decoding (Wang et al., 2025). Bayesian Networks Bayesian Networks (BNs) (Pearl, 1988) are a powerful framework for representing and reasoning under uncertainty. They are a type of Probabilistic Graphical 3
Singh Tenali Mathur Soni Phatak Lynch Murman Rizzo Natarajan
Models (Koller and Friedman, 2009) that use Directed Acyclic Graphs (DAGs) to factorize the joint distribution over a set of variables X as a product of local conditionals. Specifically, a BN over X is defined as ⟨G, P ⟩, where G = (X, E) is a DAG, consisting of nodes corresponding to each random variable and directed edges E between them indicating direct influence, and each PX ∈ P is the local conditional overQX given its parents paX . The joint probability of a data point x is factorized as P (x) = X PX (x | paX (x)). Learning these joint distributions involves estimating the parameters of each local conditional. Maximum Likelihood Estimation (MLE) (Mitchell, 1997) estimates the parameters that maximize the likelihood of the data D, θ̂M LE = arg maxθ P (D|θ). However, it fails for small data regimes where unobserved data configurations are assigned zero probability, thereby making such data points impossible under the model. A Bayesian Estimator (Koller and Friedman, 2009) mitigates this issue by placing a prior such as BDeu (Heckerman et al., 1995) over the parameters, P (θ), and computing a full posterior P (θ|D) ∝ P (D|θ)P (θ). Domain Knowledge as Qualitative Influence Statements Qualitative Influence Statements describe the direction of and nature of interaction between variables (Wellman, 1990). They are a common formalism in scientific knowledge, particularly medicine, enabling clinicians to communicate trends without specifying precise values (Sipponen et al., 1989; Malaspina et al., 2001). We focus on a specific type of qualitative influence statement called a Monotonic Influence Statement, generalizing statements of the form “X increases as Y increases.” Such monotonic influence statements can be formalized using probabilistic logic to describe patterns in the conditional probability distribution (Altendorf et al., 2005). In a bivariate system, the statement that Y positively monotonically influences X implies that higher values of Y lead to stochastically higher values of X. Formally, this corresponds to a first-order stochastic dominance or a leftward shift in the cumulative distribution: P (X ≤ x | Y = a) ≤ P (X ≤ x | Y = b) ∀a > b, a, b ∈ Domain(Y )2 . On the other hand, in multivariate systems, a qualitative influence can be defined either locally (holding all other parents of X constant (Altendorf et al., 2005)) or marginally (marginalizing over all other factors (Karanam et al., 2021, 2025)). In this work, we take the latter approach. We define the influence of patient cognitive impairment Y on a linguistic marker X based on the posterior distribution P (X | Y ), obtained by marginalizing out all other markers X \ {X}, allowing us to identify the global trends learned by the model from clinical data.
3. Neurosymbolic Qualitative Knowledge Extraction We consider the task of extracting qualitative influence statements describing the relationship between cognitive impairment and verbal fluency indicators, using a database of audio recordings. We formalize this as the following task: (i) is a raw audio recording and Given: A dataset D = {(a(i) , y (i) )}N i=1 , where each a y (i) ∈ {0, 1} is a diagnostic label indicating the presence of cognitive impairment; a set of clinical linguistic features X = {X1 , . . . , Xn } To Do: Identify qualitative influences between the presence of cognitive impairment, Y, and linguistic features X
The key challenge in this problem is the heterogeneity of the data representations. On one hand, the cognitive impairment indicators of interest are embedded in unstructured, 4
NeSyQuaKE
Table 1: Variables of Interest. Lexical, Semantic, and Acoustic variables of Interest, along with their descriptions, colored by type: Lexical , Semantic , Acoustic . Type Lexical
Semantic
Acoustic
Variables Total Words (TW) Word Frequency (WF) Word Length (WL) Age of Acquisition (AoA) Number of Clusters (NC) Average Cluster Size (CS) Number of Switches (SW) Speech Rate (SR)
Description Total valid words in response Mean corpus frequency of valid words Mean character length of valid words Mean age of learning of valid words Total semantically related word groups Mean number of words per semantic cluster Total transitions between semantic clusters Mean words produced per second
continuous acoustic signals a(i) . On the other hand, clinical knowledge is expressed through qualitative relationships between discrete variables X and Y. Extracting this knowledge requires a system to perceive the nuances of natural language and reason about the probabilistic dependencies within a formal symbolic framework. To bridge this gap, we introduce Neurosymbolic Qualitative Knowledge Extraction (NeSyQuaKE). As outlined in Figure 1, NeSyQuaKE instantiates a Neuro|Symbolic architecture to decouple the task into three distinct stages: Information Extraction (IE), Bayesian Network Construction (BNC), and Qualitative Knowledge Extraction (QuaKE). 3.1. Information Extraction The Information Extraction (IE) stage maps each audio recording a(i) to a tabular feature vector x(i) , transforming the original dataset D into a tabular dataset D′ = {(x(i) , y (i) )}N i=1 , making it suitable for Bayesian Network-based reasoning. It does so using the Speech to Text (S2T) and the Large Language Models (LLM) modules sequentially. The Speech to Text (S2T) module takes each raw audio file a(i) as input and outputs the corresponding transcriptions t(i) , each consisting of a set of timestamped utterances. (i) (i) (i) (i) That is, t(i) = S2T(a(i) ) ∀i = 1, . . . , N where t(i) = {(t1 , u1 ), . . . , (tki , uki )} We implement S2T using a pre-trained transformer-based foundation model trained for speech recognition. Next, we use these transcriptions T = {t(i) }N i=1 to create the tabular dataset by aligning them with the variables of interest, X. The LLM module constructs the tabular datasets by using the transcripts T as input along with a custom prompt P describing the scope of information required, extraction rules, and constraints. The LLM outputs a structured response R, which serves as a grounded intermediate representation, where the LLM has mapped the raw transcript t(i) into a structured sequence of words and temporal markers. However, these individual grounded symbols do not directly correspond to the variables of interest required for BN reasoning. To complete the grounding, we define a set of deterministic symbolic programs to compute the feature vectors {x(i) }N i=1 . This step ensures that every clinical variable is derived through a transparent and consistent calculation. For example, while a dictionary lookup function identifies which words are valid responses, a simple length function determines the total words (TW), i.e, TW = LEN(FILTER(valid words, words)). Similarly, the speech rate is calculated by dividing the number of identified words by the total dura5
Singh Tenali Mathur Soni Phatak Lynch Murman Rizzo Natarajan
tion recorded in the temporal markers. This ensures that the resulting dataset is logically consistent while being grounded in the information extracted by the foundation models. 3.2. Bayesian Network Construction The Bayesian Network Construction (BNC) stage is responsible for constructing the influence graph (a BN) over the variables of interest X = {X1 , X2 , . . . , Xn } as well as learning the joint probability distribution P (X) = P (X1 , X2 , . . . , Xn ) over the graph G = (X, E). Since reliable structure learning is infeasible in domains with small and noisy data, we obtain the initial graph G over the variables from an LLM by prompting it multiple times with the task and feature description to output a BN structure. The resulting graphs are pooled to construct the final graph by adding edges in decreasing order of occurrence, skipping any edges that introduce cycles to maintain the acyclicity constraint. Next, the dataset generated in the IE stage is used to learn the joint distribution P (X) for the BN. Due to the stochastic nature of the different components of the IE stage and the noise inherent in raw audio files used for dataset construction D′ , a Bayesian Estimator with a strong prior like BDeu is used to regularize the parameter estimates. 3.3. Qualitative Knowledge Extraction The Qualitative Knowledge Extraction module is based on the QuaKE algorithm (Karanam et al., 2021). It reasons over the joint distribution P (X, Y ) of Alzheimer’s diagnostic labels Y and the discretized feature vector X to extract the qualitative knowledge as Monotonic Influence (MI) statements. For each descendant Xd of Y , QuaKE checks for the presence of monotonic influence of Y on its descendant Xd as YYY ′ Cd = max(P (Xd ≤ k|Y = y j ) − P (Xd ≤ k|Y = y j ) + ϵ, 0) j
j′
k
where ϵ is the monotonic slack that allows a small tolerance for violation of this constraint. + A positive MI (Y ≺M Xd ) requires that higher values of Y make it less likely that Xd ′ takes smaller values, which implies that ∀k,j ′ >j , P (Xd < k|Y = y j ) ≥ P (Xd < k|Y = y j ). Essentially, Cd is non-zero if and only if the monotonicity constraint holds for all the values ′ of Y , i.e., for all (y j , y j ) where j ′ > j and for all the values of k. The same holds for negative monotonicities. When the monotonicity constraint is satisfied, the degree of monotonic influence δd is then computed as δd = ICd >0
X X X P (Xd ≤ k|Y = y j ) − P (Xd ≤ k|Y = y j ′ ) j
|Y |
j ′ >j k
(1)
where δd quantifies the strength of the monotonic influence. Intuitively, stronger monotonic influences result in larger cumulative differences between the conditional probabilities ′ P (Xd ≤ k|Y = y j ) and P (Xd ≤ k|Y = y j ) across the label values and threshold k. So, larger absolute values of δd represent a stronger monotonic influence. In addition to the MIs of Y on its descendants, this module also extracts the positive and negative monotonic influences of each feature Xi on its descendants in the graph. 6
NeSyQuaKE
CI
NC
AoA
WL
CI
WF
TW
WL
CS
SR
NC
AoA
SW
WF
TW
CS
SR
SW
Figure 2: The Influence Networks, elicited from the expert (left), and LLM responses (right) representing direct influence relationships between CI , Lexical , Semantic , and Acoustic variables.
4. Empirical Evaluation We now present our empirical evaluations1 and answers the following questions: Q1) Is NeSyQuake effective at symbol grounding of verbal fluency data? Q2) Does the knowledge extracted from the NeSyQuaKE pipeline match expert knowledge? Q3) Is NeSyQuaKE an effective source of generating hypotheses about AD markers? Q4) Can feedback between symbolic and neural components improve symbol grounding? Dataset: To answer these questions, we consider a dataset of 162 audio files of verbal fluency tests from the University of Nebraska Medical Center, focusing on patients classified as Healthy Controls (HC) or Mild Cognitive Impairment (MCI). Manually administering these verbal fluency tests and diagnosing the patients is both time-consuming and labor-intensive. However, the data is very rich- Each patient in the study underwent a comprehensive neuropsychological battery annually for three years, and their diagnoses for each year were reviewed by the domain experts: two neurologists and a neuropsychologist. Since we aim to identify knowledge for early diagnosis, we use data from the first year. Our primary focus is on patients’ ability to generate words within one minute, either beginning with specific letters (e.g., F or L) or belonging to a given semantic category (e.g., animals or vegetables), and evaluating the known clinical AD markers (Kavé and Goral, 2016; Venneri et al., 2008; Troyer et al., 1998; Pistono et al., 2019) described in Table 1. Experimental Setup: We instantiate NeSyQuaKE as follows: we use the WhisperX Speech to Text model and the MedGemma-27b medical LLM (Sellergren et al., 2025) as the foundation models2 . We enforce syntactic constraints on LLM output using Pydantic3 , which are fed as input to symbolic programs to generate the final dataset4 over the features described in Table 1. Finally, to discretize the dataset, we use the mean feature values as a threshold5 . For QuaKE, the LLM-generated BN (Figure 2 (right)) is fed alongside the data to learn the parameters using a Bayesian Estimator with a BDeu prior and an equivalent sample size of 10 to account for noisy data and a monotonic slack of 0. Results: Tables 2 and 3 present the results for the questions introduced above. (Q1, Symbol Grounding) To evaluate the efficacy of our symbolic grounding module, we compare its output (i.e., 1. The code for all the experiments is available in the supplementary 2. Owing to the sensitive nature of medical data, only locally available LLMs were used for data processing 3. Without structural decoding, none of the outputs generated match the schema specification. Using structural decoding, on average 97.5% of all outputs were schema compliant across 5 seeds. 4. The symbolic programs, auxiliary data used (valid word lists, word groups, word frequencies, and ageof-acquisition data), structure learning, and transcription error analysis are provided in the appendix. 5. Threshold is computed as the mean over each variable across 5 different trials, as per our domain experts’ recommendation, with the exact values provided in the supplementary
7
Singh Tenali Mathur Soni Phatak Lynch Murman Rizzo Natarajan
SW
SW
CS
CS
AoA
AoA
TW
TW
NC
NC
SR
SR
WF
WF WL
WL −0.3
−0.2
−0.1 0 Monotonic Influence Score
0.1
−0.2
0.2
−0.15
−0.1 −0.05 Monotonic Influence Score
0
0.05
Figure 3: Degree of monotonic influence of cognitive impairment on semantic (left) and phonemic (right) features, as inferred by NeSyQuaKE, averaged over 5 trials variable values) to manually annotated data. To quantify the resulting difference in pairwise relationships, we introduce the Pairwise Conditional Estimation Error (PCEE), which is computed as P (PNeSyQuaKE (X = 1|Y = y) − Pmanual (X = 1|Y = y))2 P CEE(X|Y ) = Y |Y | where Y denotes the cognitive impairment and X represents a variable of interest. A higher PCEE (closer to 1) indicates a larger discrepancy, while a lower PCEE indicates close agreement between NeSyQuaKE groundings and manual annotations. We compare PCEE of our symbol-grounding pipeline against a rule-based baseline that uses a set of hardcoded patterns and keywords for approximately parsing the transcript1 . Table 2 shows the PCEE values for different methods. NeSyQuaKE obtains substantially lower PCEE than the domain-specific rule-based extractor, yielding an improvement of 94.3% in the semantic fluency task and 77.2% in the phonemic task, indicating higher reliability and match with manual annotations. (Q2, Comparison with Expert) Table 3 compares the monotonic influences extracted by NeSyQuaKE to those listed by our domain experts for the semantic and phonemic fluency tasks. For both tasks, NeSyQuake matched all influences initially elicited from the expert. In addition to the direction of influence, NeSyQuaKE also quantifies their degree of influence as shown in Figure 3. Additionally, we compare the LLM-generated Bayesian Network structure used by NeSyQuaKE to an expert-defined network. We find that the LLMgenerated network recovers 80% of the edges from the expert network. However, it contains several extra edges. 43% of its edges were not present in the expert network. In contrast, the data-driven Peter-Clark (PC (Spirtes and Glymour, 1991)) algorithm4 yields a sparser network, matching only 1 out of 15 expert network edges. Overall, the LLM-generated Bayesian Network structure is conservative, encoding fewer independencies, which allows its data-driven parameters to recover patterns accurately. (Q3, Additional Hypotheses) Our framework discovers 9 additional influences beyond those initially elicited from the domain expert. Of these, 6 were in the phonemic task, and 3 were in the semantic task. In the semantic task, it infers the negative monotonic influence of cognitive impairment on cluster size, speech rate, and word length. These were validated by the domain expert and confirmed to be correct, since Healthy Control (HC) participants typically exhibit higher word length, larger cluster sizes, and higher speech rates than individuals with MCI in semantic fluency. For the phonemic task, our framework infers a positive monotonic influence of cognitive impairment on word length and a negative mono8
NeSyQuaKE
Table 2: PCEE values with rule-based extraction vs NeSyQuake for Lexical , Semantic , and Acoustic variables in Semantic and phonemic Fluency tasks. For NeSyQuaKE and NeSyQuaKE with feedback, the reported values correspond to the mean over five runs. Standard deviations are omitted, as all values were below 10−4 Variable Name
Semantic Fluency Rule-based NeSyQuake w/Feedback
Phonemic Fluency Rule-based NeSyQuake w/Feedback
Total Words (TW) Age of Acquisition (AoA) Word Length (WL) Word Frequency (WF)
0.087 0.003 0.117 0.005
0.005 0.001 0.008 0.002
0.004 0.001 0.008 0.002
0.120 0.003 0.0 0.004
0.044 0.001 0.001 0.001
0.041 0.001 0.001 0.001
Number of Clusters (NC) Average Cluster Size (CS) Number of Switches (SW)
0.117 0.028 0.059
0.006 0.003 0.002
0.004 0.003 0.001
0.057 0.064 0.133
0.012 0.028 0.030
0.009 0.024 0.022
Speech Rate (SR)
0.070
0.002
0.002
0.162
0.007
0.007
Table 3: NeSyQuaKE recovers expert knowledge. Comparison of monotonic influence rules extracted by NeSyQuaKE with those elicited from domain experts. Each rule characterizes the influence of Cognitive Impairment (CI) on variables of interest. For each source of rules, a (✓) indicates the monotonic influence was found, (?) indicates it was not found, and (×) indicates its opposite was found. An influence statement A M± ≺ B is read as A has a positive (M +) or negative (M −) monotonic influence on B. Rule CI M≺ TW CI M≺ AoA CI M≺ WL CI M+ ≺ WL CI M+ ≺ WF CI M≺ NC CI M≺ CS CI M≺ SW CI M≺ SR
Semantic Fluency Expert NeSyQuaKE ✓ ✓ ✓ ✓ ? ✓ ? × ✓ ✓ ✓ ✓ ? ✓ ✓ ✓ ? ✓
Phonemic Fluency Expert NeSyQuaKE ? ✓ ✓ ✓ ? × ? ✓ ✓ ✓ ? ✓ ? ✓ ? ✓ ? ✓
tonic influence of cognitive impairment on total words, number of clusters, cluster size, number of switches, and speech rate. The experts concurred with the negative monotonic influences; while these patterns are known to exist in the phonemic task, the effect is typically less pronounced than in the semantic task (Beatty et al., 1997; Laws et al., 2010). Nevertheless, NeSyQuaKE could detect these subtle patterns. On the other hand, the experts were less certain about the positive influence of cognitive impairment on word length. We speculate that this could be a result of the setup of the phonemic task. In it, the participants were instructed to generate as many words as possible within a one-minute interval. To achieve this, participants may have strategically used smaller words to maximize the total word count at the cost of average word length. However, such strategies are not possible in semantic fluency tasks where responses are constrained to fixed categories like vegetables and animals. 9
Singh Tenali Mathur Soni Phatak Lynch Murman Rizzo Natarajan
(Q4, Incorporating Feedback) We consider incorporating feedback from the symbolic component to the neural component. The symbolic component tests words for relevance to the specific fluency task and rejects irrelevant words, such as ones not beginning with the correct letter. However, these might be mistranscriptions. We feed these rejected words to the LLM and prompt it for phonetically similar replacements relevant to the task; we keep the highest-confidence replacements to correct the mistranscriptions. Table 2 shows that incorporating additional feedback resulted in an improvement over NeSyQuaKE by 13.7% and 14.5% in the semantic and phonemic tasks, respectively, but with the same set of monotonicities discovered as before. While the LLM correctly guesses the correct versions of mistranscribed words, this correction is not yet grounded in the audio clip, but is purely based on the transcript context. Therefore, while there may be a performance improvement using a feedback loop, further study is necessary to validate it.
5. Conclusion and Future Work We introduced the NeSyQuaKE framework, a neurosymbolic framework designed to extract qualitative knowledge from clinical audio recordings of verbal fluency tests. It does so by employing a Neuro|Symbolic (Type 3) architecture, bridging the representational gap between unstructured acoustic signals and structured qualitative knowledge in the form of monotonic influence statements. Our empirical evaluation on a real-world dataset shows that NeSyQuaKE is effective at recovering established clinical domain knowledge. Moreover, it also inferred less prominent trends, like the negative monotonic influence of cognitive impairment on the number and size of clusters in the phonemic task. There are several directions for future work. First, our choice of foundation models was limited by resource constraints and the sensitive nature of clinical data. Exploring larger foundation models and developing models specialized for processing pathological speech (Cohn et al., 2026) is important future work. Additionally, future work should explore Speech Language Models (SLMs) like SALMONN (Tang et al., 2024) to enable tighter integration between speech processing and information extraction. Second, future work should consider ways to scale the analysis to a larger and more complex set of variables of interest, such as the number of pauses, repetitions, and the rate of change in speech rate during the course of the response, and employing tractable probabilistic models. Finally, while NeSyQuaKE identifies the nature of relationships for a given set of variables of interest, it can be extended to discover new AD markers, such as by using an LLM to generate additional candidates and inferring the extent of cognitive impairment’s influence on them. Together with these directions, NeSyQuaKE offers a scalable, explainable approach to tracking Alzheimer’s disease progression and supporting early diagnostic efforts.
Acknowledgments The authors gratefully acknowledge the support by AFOSR award FA9550-23-1-0239.
10
NeSyQuaKE
References 2024 Alzheimer’s disease facts and figures. Alzheimer’s & Dementia, 20(5):3708–3821, may 2024. ISSN 1552-5260, 1552-5279. doi: 10.1002/alz.13809. Harsh Ahlawat, Naveen Aggarwal, and Deepti Gupta. Automatic speech recognition: A survey of deep learning techniques and approaches. International Journal of Cognitive Computing in Engineering, 6:201–237, 2025. ISSN 2666-3074. doi: https://doi.org/10. 1016/j.ijcce.2024.12.007. Eric E Altendorf, Angelo C Restificar, and Thomas G Dietterich. Learning from sparse data by exploiting monotonicity constraints. In Proceedings of the Twenty-First Conference on Uncertainty in Artificial Intelligence, pages 18–26, 2005. William W Beatty, Julie A Testa, Shelley English, and Peter Winn. Influences of clustering and switching on the verbal fluency performance of patients with alzheimer’s disease. Aging, Neuropsychology, and Cognition, 4(4):273–279, 1997. Marc Brysbaert and Andrew Biemiller. Test-based age-of-acquisition norms for 44 thousand english word meanings. Behavior research methods, 49(4):1520–1523, 2017. D.G. Clark, P. Kapur, D.S. Geldmacher, J.C. Brockington, L. Harrell, T.P. DeRamus, P.D. Blanton, K. Lokken, A.P. Nicholas, and D.C. Marson. Latent information in fluency lists predicts functional decline in persons at risk for alzheimer disease. Cortex, 55:202–218, 2014. ISSN 0010-9452. doi: https://doi.org/10.1016/j.cortex.2013.12.013. Language, Computers and Cognitive Neuroscience. Michelle Cohn, Alyssa Lanzi, Yui Ishihara, Chen-Nee Chuah, Georgia Zellou, and Alyssa Weakley. Challenges in automatic speech recognition for adults with cognitive impairment, 2026. Paul S Foster, Valeria Drago, Raegan C Yung, Jaclyn Pearson, Kristi Stringer, Tania Giovannetti, David Libon, and Kenneth M Heilman. Differential lexical and semantic spreading activation in alzheimer’s disease. American Journal of Alzheimer’s Disease & Other Dementias®, 28(5):501–507, 2013. Artur d’Avila Garcez and Luis C Lamb. Neurosymbolic ai: The 3 rd wave. Artificial Intelligence Review, 56(11):12387–12406, 2023. David Heckerman, Dan Geiger, and David M Chickering. Learning bayesian networks: The combination of knowledge and statistical data. Machine learning, 20(3):197–243, 1995. Julie D Henry, John R Crawford, and Louise H Phillips. Verbal fluency performance in dementia of the alzheimer’s type: a meta-analysis. Neuropsychologia, 42(9):1212–1222, 2004. Joy Hsu, Jiayuan Mao, Josh Tenenbaum, and Jiajun Wu. What’s left? concept grounding with logic-enhanced foundation models. Advances in Neural Information Processing Systems, 36:38798–38814, 2023. 11
Singh Tenali Mathur Soni Phatak Lynch Murman Rizzo Natarajan
Haeeul Hwang, Eric Jordan, Deok-Hee Kim-Dufor, Christophe Lemey, and Motasem Alrahabi. Evaluating asr in a clinical context: What whisper misses. In Proceedings of the 8th International Conference on Natural Language and Speech Processing (ICNLSP-2025), pages 374–378, 2025. Athresh Karanam, Alexander L Hayes, Harsha Kokel, David M Haas, Predrag Radivojac, and Sriraam Natarajan. A probabilistic approach to extract qualitative knowledge for early prediction of gestational diabetes. In International Conference on Artificial Intelligence in Medicine, pages 497–502. Springer, 2021. Athresh Karanam, Saurabh Mathur, Sahil Sidheekh, and Sriraam Natarajan. A unified framework for human-allied learning of probabilistic circuits. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 17779–17787, 2025. Henry Kautz. The third ai summer: Aaai robert s. engelmore memorial lecture. Ai magazine, 43(1):105–125, 2022. Gitit Kavé and Mira Goral. Word retrieval in picture descriptions produced by individuals with alzheimer’s disease. Journal of clinical and experimental neuropsychology, 38(9): 958–966, 2016. Daphne Koller and Nir Friedman. Probabilistic graphical models: principles and techniques. MIT press, 2009. Benjamin Kuipers. Qualitative reasoning: modeling and simulation with incomplete knowledge. MIT press, 1994. Keith R Laws, Amy Duncan, and Tim M Gale. ‘normal’semantic–phonemic fluency discrepancy in alzheimer’s disease? a meta-analytic study. Cortex, 46(5):595–601, 2010. Dolores Malaspina, Susan Harlap, Shmuel Fennig, Dov Heiman, Daniella Nahon, Dina Feldman, and Ezra S Susser. Advancing paternal age and the risk of schizophrenia. Archives of general psychiatry, 58(4):361–367, 2001. Cecile A Marczinski and Andrew Kertesz. Category and letter fluency in semantic dementia, primary progressive aphasia, and alzheimer’s disease. Brain and language, 97(3):258–265, 2006. Camillo Marra, Chiara Piccininni, Giovanna Masone Iacobucci, Alessia Caprara, Guido Gainotti, Emanuele Maria Costantini, Antonio Callea, Annalena Venneri, and Davide Quaranta. Semantic memory as an early cognitive marker of alzheimer’s disease: Role of category and phonological verbal fluency tasks. Journal of Alzheimer’s Disease, 81(2): 619–627, 2021. Giuseppe Marra, Sebastijan Dumančić, Robin Manhaeve, and Luc De Raedt. From statistical relational to neurosymbolic artificial intelligence: A survey. Artificial Intelligence, 328:104062, 2024. Tom Mitchell. Machine Learning. McGraw Hill, 1997. 12
NeSyQuaKE
Humza Naveed, Asad Ullah Khan, Shi Qiu, Muhammad Saqib, Saeed Anwar, Muhammad Usman, Naveed Akhtar, Nick Barnes, and Ajmal Mian. A comprehensive overview of large language models. ACM Trans. Intell. Syst. Technol., 16(5), August 2025. ISSN 2157-6904. doi: 10.1145/3744746. Phillip Odom and Sriraam Natarajan. Human-guided learning for probabilistic logic models. Frontiers in Robotics and AI, 5:56, 2018. Judea Pearl. Probabilistic Reasoning in Intelligent Systems: Networks of Plausible Inference. Morgan Kaufmann, 1988. Aurélie Pistono, Jeremie Pariente, Catherine Bézy, Béatrice Lemesle, Johanne Le Men, and Mélanie Jucla. What happens when nothing happens? an investigation of pauses as a compensatory mechanism in early alzheimer’s disease. Neuropsychologia, 124:133–143, 2019. Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, and Ilya Sutskever. Robust speech recognition via large-scale weak supervision. In International conference on machine learning, pages 28492–28518. PMLR, 2023. David P Salmon. Neuropsychological features of mild cognitive impairment and preclinical alzheimer’s disease. Behavioral neurobiology of aging, pages 187–212, 2011. Andrew Sellergren, Sahar Kazemzadeh, Tiam Jaroensri, Atilla Kiraly, Madeleine Traverse, Timo Kohlberger, Shawn Xu, Fayaz Jamil, Cı́an Hughes, Charles Lau, et al. Medgemma technical report. arXiv preprint arXiv:2507.05201, 2025. P Sipponen, K Seppälä, M Aärynen, T Helske, and P Kettunen. Chronic gastritis and gastroduodenal ulcer: a case control study on risk of coexisting duodenal or gastric ulcer in patients with gastritis. Gut, 30(7):922–929, 1989. Peter Spirtes and Clark Glymour. An algorithm for fast recovery of sparse causal graphs. Social Science Computer Review, 9:62–72, 04 1991. doi: 10.1177/089443939100900106. Changli Tang, Wenyi Yu, Guangzhi Sun, Xianzhao Chen, Tian Tan, Wei Li, Lu Lu, Zejun MA, and Chao Zhang. Salmonn: Towards generic hearing abilities for large language models. In The Twelfth International Conference on Learning Representations, 2024. Angela K Troyer, Morris Moscovitch, Gordon Winocur, Larry Leach, and Morris Freedman. Clustering and switching on verbal fluency tests in alzheimer’s and parkinson’s disease. Journal of the International Neuropsychological Society, 4(2):137–143, 1998. Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems, 30, 2017. Annalena Venneri, William J McGeown, Heidi M Hietanen, Chiara Guerrini, Andrew W Ellis, and Michael F Shanks. The anatomical bases of semantic retrieval deficits in early alzheimer’s disease. Neuropsychologia, 46(2):497–510, 2008. 13
Singh Tenali Mathur Soni Phatak Lynch Murman Rizzo Natarajan
Darren Yow-Bang Wang, Zhengyuan Shen, Soumya Smruti Mishra, Zhichao Xu, Yifei Teng, and Haibo Ding. Slot: Structuring the output of large language models. arXiv preprint arXiv:2505.04016, 1(2):3, 2025. Michael P Wellman. Fundamental concepts of qualitative probabilistic networks. Artificial intelligence, 44(3):257–303, 1990. Antonia Wüst, Wolfgang Stammer, Hikaru Shindo, Lukas Helff, Devendra Singh Dhami, and Kristian Kersting. Synthesizing visual concepts as vision-language programs. arXiv preprint arXiv:2511.18964, 2025. Derong Xu, Wei Chen, Wenjun Peng, Chao Zhang, Tong Xu, Xiangyu Zhao, Xian Wu, Yefeng Zheng, Yang Wang, and Enhong Chen. Large language models for generative information extraction: A survey. Frontiers of Computer Science, 18(6):186357, 2024. Shuo Yang and Sriraam Natarajan. Knowledge intensive learning: Combining qualitative constraints with causal independence for parameter learning in probabilistic models. In Joint European Conference on Machine Learning and Knowledge Discovery in Databases, pages 580–595. Springer, 2013. Songchi Zhou and Sheng Yu. High-throughput biomedical relation extraction for semistructured web articles empowered by large language models. BMC Medical Informatics and Decision Making, 25(1):351, 2025.
14
NeSyQuaKE
Type
Feature
Pseudo-code
Lexical
Number of Exemplars Word Frequency
LEN(FILTER(VALID, words)) MEAN(MAP(CORPUS FREQ, FILTER(VALID, words))) MEAN(MAP(LEN, FILTER(VALID, words))) MEAN(MAP(AOA, FILTER(VALID, words)))
Word Length Age of Acquisition Semantic
Number of Clusters Average Cluster Size Number of Switches
Acoustic
Speech Rate
LEN(DISTINCT(MAP(CLUSTER ID, words))) (LEN(words) - 1) / LEN(DISTINCT(MAP(CLUSTER ID, words))) LEN(FILTER(switched, ZIP(MAP(CLUSTER ID, words), MAP(CLUSTER ID, words[1:])))) LEN(words) / (words[-1].end words[0].start)
Table 4: Symbolic programs used to compute values for Lexical, Semantic, and Acoustic variables from LLM response
Appendix A. Symbolic Programs A.1. Pseudocodes Table 4 shows the list of features and the pseudo-code of the deterministic symbolic programs that are used to extract these feature values. These symbolic programs take the structured JSON obtained from the LLM as input6
6. The pseudocode table is for reference, with the actual implementation provided in the construct features.py file provided in the linked repository: https://github.com/s-ranveer/ verbal_fluency_alzheimers/blob/main/construct_features.py
15
Singh Tenali Mathur Soni Phatak Lynch Murman Rizzo Natarajan
A.2. Data Sources The following provides the list and the data sources that were considered for the task A.2.1. Word Frequency The Python library wordfreq7 was used to compute the word frequency for the different words. A.2.2. Age Of Acquisition The Age of Acquisition is a 2017 corpus of over forty thousand different words (Brysbaert and Biemiller, 2017). Any words considered when constructing the Age of Acquisition (AoA) feature, any word not in the list (no word or its lemma didn’t match), were not considered during computation. A.3. Valid Groups The same animal and vegetable groups as in (Clark et al., 2014) were used to group the animal and vegetable responses. The animal list was adopted without modification, while the grocery list was filtered to include only vegetables by removing fruits, meats, and processed food to get the final list of vegetables for semantic fluency word groups. For the Phonemic Fluency, two words were said to belong to the same group if they started with the same two letters, were homonyms, rhymed, or differed by only a vowel sound. The valid words are defined separately for the Semantic and Phonemic tasks. 1. Semantic: The words should belong to the group list of vegetables or animals to be considered valid 2. Phonemic: The word must have a word frequency of greater than 0, not be a proper noun (name, person, or number), and should begin with the letter in question when asked to output a word list. Only valid words are considered for the metrics, except for Speech rate, where we need to consider all the words spoken.
7. Refer to the package repository for additional details
16
NeSyQuaKE
Appendix B. Prompts
The different prompts used with the LLM are provided below. The LLM used was Medgemma27b with a total context window of 40,000 tokens
## Alzheimer's Fluency Tests Processing You are processing transcripts from a cognitive assessment for Alzheimer's disease. The transcript contains the timestamped utterances (questions and answers) from both the examiner and the participant. We are interested in extracting the responses to the following standard verbal fluency prompts which would go something like: 1. "Let's begin. Tell me all the words you can, as quickly as you can, that begin with the letter 'F'. Ready? Begin." 2. "Now I want you to do the same for another letter. The next letter is 'L.' Ready? Begin." 3. "Now I want you to name things that belong to another category: Animals. You will have one minute. I want you to tell me all the animals you can think of in one minute. Ready? Begin." 4. "Now I want you to name things that belong to another category: Vegetables. You will have one minute. I want you to tell me all the vegetables you can think of in one minute. Ready? Begin."
The responses to these standard prompts should be around 60 seconds from the start of the response. Your task is to extract the following from the transcripts: 1. The full response by the participant for the prompt ignoring the other speaker. 2. The extracted answer from the full response where one omits the filler words and incorrect answers. Repetitions are allowed. 3. The starting and ending timestamp of the participant response. 4. Individual pauses during the participant's response where the pause is at least 1 second, calculated as the time gap between the end timestamp of one utterance and the start timestamp of the next utterance. **Rules** 1. Do not infer or reconstruct missing speech. 2. Return an empty string for "full_response" and empty arrays for other fields if no response exists for that task. 3. Output only the final JSON with no additional commentary or explanation. 4. Use R1, R2, R3, and R4 to represent the 4 task prompts respectively. 5. Timestamps should be preserved in their original format from the transcript. **Edge Cases** - If a participant self-corrects (e.g., "cat... no, dog"), include both in full_response but only the final word in extracted_answer if valid. - If the examiner interrupts, note the interruption point as the end timestamp.
Figure 4: Prompt for Information Extraction (Part 1). 17
Singh Tenali Mathur Soni Phatak Lynch Murman Rizzo Natarajan
**Output JSON Schema** ```json { "responses": { "R1": { "full_response": "string", "response_timestamps": { "start": "string", "end": "string" }, "extracted_answer": ["word1", "word2"], "pauses": [ { "start": "string", "end": "string" }, { "start": "string", "end": "string" } ] }, "R2": { }, "R3": { }, "R4": { } } } ``` **Notes** 1. When we ask for the full response, we mean the entire answer provided by the participant including any errors. 2. A pause occurs when there is more than a second difference between the end timestamp of one line and the start timestamp of the next line. 3. The extracted answer for R1 and R2 should not include names of people or places or numbers. 4. Make a proper judgment of when the response begins and ends. Many times, there would be a sentence including phrases like "Stop here" to indicate the end of the response.
Figure 5: Prompt for Information Extraction (Part 2).
18
NeSyQuaKE
# Alzheimer's Verbal Fluency Test Corrections You are a clinical expert reviewing verbal fluency responses extracted from speech transcripts. Patients are elderly and may have mild cognitive impairment, dementia, speech disfluencies, articulation difficulties, or hearing impairments. Transcription and word extraction systems may also introduce recognition errors. ## Task For each rejected word, determine whether it is a transcription or extraction error and identify the most likely intended word. Rejected words failed symbolic grounding | they are either not valid words or do not belong to the expected response category. Your goal is solely to identify likely transcription or extraction errors. Do not attempt to improve the patient's score or force words into the expected category. ## Response Categories The input contains responses across four verbal fluency tasks: - **R1** | Letter fluency: words beginning with the letter F - **R2** | Letter fluency: words beginning with the letter L - **R3** | Category fluency: Animals - **R4** | Category fluency: Vegetables ## Correction Process Apply the following steps in order for each rejected word. ### Step 1: Generate Phonetic Candidates Identify all real words that are phonetically similar to the rejected word. Prefer candidates requiring minimal sound changes | ideally a single phoneme substitution, deletion, or insertion. Be especially skeptical of candidates that change the initial phoneme | these represent larger phonetic departures and require stronger justification. Never consider a candidate that is identical in pronunciation to the rejected word. No transcription error can exist between homophones. ### Step 2: Filter by Task Category Eliminate any candidate that does not satisfy the task category constraint: - For R1: candidate must begin with F - For R2: candidate must begin with L - For R3: candidate must be an animal - For R4: candidate must be a vegetable If no candidates remain after filtering, return no correction. Note: category is used here only to eliminate impossible candidates, not to generate or prefer corrections.
19
Figure 6: Prompt for Correcting Mistranscription (Part 1).
Singh Tenali Mathur Soni Phatak Lynch Murman Rizzo Natarajan
### Step 3: Rank by Transcript Context Rank remaining candidates by strength of contextual support in the transcript. Evidence strength, from strongest to weakest: - An immediately adjacent word that clearly suggests the intended word - A consistent naming pattern across the full response - Repeated phonemes or themes in nearby words If no candidate has meaningful contextual support, return no correction. ### Step 4: Select or Decline If one candidate is clearly supported by both phonetic similarity and transcript context, propose it as the correction. If two or more candidates remain equally matched after Steps 1{3, return no correction. ## When Not to Correct Do not propose a correction when: - Phonetic evidence is weak or ambiguous. - No candidate survives the category filter. - Contextual evidence is weak or diffuse. - Two or more candidates remain equally plausible after all steps. - You are uncertain which word was intended. **Valid English word rule:** If the rejected word is a valid English word, it may reflect the patient's genuine response rather than a transcription error. Only correct at confidence 5. When in doubt, return no correction. ## Confidence Scale - 5 = Near-certain transcription error (single phoneme difference, strongly supported by context) - 4 = Strong phonetic and contextual evidence - 3 = Plausible but some uncertainty remains - 2 = Weak evidence | do not correct - 1 = Highly speculative | do not correct **Valid English words require confidence 5 to correct. If you cannot reach confidence 5, return no correction.**
Figure 7: Prompt for Correcting Mistranscription (Part 2).
20
NeSyQuaKE
## Worked Example **Input:** ```json { "R2": { "prompt": "Tell me all the words you can that begin with the letter L.", "full_response": "love, lug, log, lamp, lunch, lost, lack, back", "word_list": ["love", "lug", "log", "lamp", "lunch", "lost", "lack"], "rejected_words": ["back"] } } ``` **Reasoning:** - Step 1: Phonetic candidates for "back" | "lack", "black", "rack", "pack" - Step 2: Category filter for R2 (must begin with L) | only "lack" survives - Step 3: "lack" is consistent with surrounding L-words in the response - Step 4: One candidate remains | propose "lack" **Output:** ```json { "corrections": { "R2": [ { "incorrect_word": "back", "correct_word": "lack", "confidence": 4, "justification": "Only L-candidate; consistent with response pattern" } ] } } ``` ## Output Format Return only raw JSON. No markdown, no explanation, no additional text. ```json { "corrections": { "R1": [ { "incorrect_word": "word", "correct_word": "replacement", "confidence": 5, "justification": "max 7 words" } ], "R2": [], "R3": [], "R4": [] } } ``` Return an empty array for any response with no valid corrections.
21
Figure 8: Prompt for Correcting Mistranscription (Part 3).
Singh Tenali Mathur Soni Phatak Lynch Murman Rizzo Natarajan
You are an expert in cognitive science and psycholinguistics. You will be given a set of variables derived from verbal fluency tests administered to patients, some of whom may have cognitive impairments. Your task is to propose a causal graph over these variables. Variables: CI { Cognitive Impairment: presence of cognitive impairment (e.g., MCI, Alzheimer's disease) TW { Total Words: total valid words produced in response WF { Word Frequency: mean corpus frequency of valid words WL { Word Length: mean character length of valid words AoA { Age of Acquisition: mean age at which words are typically learned NC { Number of Clusters: total semantically related word groups CS { Average Cluster Size: mean words per semantic cluster SW { Number of Switches: total transitions between semantic clusters SR { Speech Rate: mean words produced per second Output only a JSON array of directed edges. Each edge is a two-element array [X, Y] meaning "X is a cause of Y". Use the variable abbreviations above. Do not include any explanation or additional text. Example format:[["CI", "SR"], ["AoA", "WL"]]
Figure 9: Prompt for Eliciting the Influence Network from the LLM
22
NeSyQuaKE
Appendix C. Feature Discretization Thresholds and Frequencies Table 5: NeSyQuaKE grounds variables of interest in raw audio recordings. Discretization Thresholds and Frequencies for Lexical , Semantic , and Acoustic variables in Semantic and phonemic Fluency tasks.
Variable Name
Semantic Fluency Threshold Frequency
Phonemic Fluency Threshold Frequency
Total Words (TW) Age of Acquisition (AoA) Word Length (WL) Word Frequency (WF)
23.39 5.30 5.92 1.7e-5
83 88 90 74
20.97 6.60 5.30 1.7e-4
89 82 77 45
Number of Clusters (NC) Average Cluster Size (CS) Number of Switches (SW)
11.88 1.06 14.39
93 88 77
8.93 1.40 14.50
109 73 82
Speech Rate (SR)
0.35
62
0.28
77
Appendix D. Transcription Error Analysis While NeSyQuaKE correctly identifies patterns in the data, the transcriptions may be noisy since pathological speech can be challenging for automatic transcription models. To estimate the loss due to poor transcriptions, we calculated the percent difference in words present in the manual annotations and the foundation models: Percent Difference =
|Corrected Count - Original Count| Original Count
As shown in table 6, there is a higher percent difference and significantly greater variability in MCI patients compared to Healthy Controls, particularly in phonemic tasks. These statistics demonstrate the inherent difficulty S2T models face when handling pathological speech, confirming that the primary bottleneck in our pipeline is the acoustic front-end rather than the symbolic grounding logic. Table 6: Semantic and Phonetic Fluency Word Counts Before and After Correction Task
Group
Original
Corrected
Percent Difference
Semantic
HC MCI
35.62 ± 7.20 26.58 ± 8.54
35.82 ± 6.96 27.44 ± 8.41
3.94% ± 4.64% 9.08% ± 10.73%
Phonemic
HC MCI
28.30 ± 7.94 22.24 ± 9.26
29.69 ± 7.46 24.37 ± 8.69
8.41% ± 15.20% 18.04% ± 20.74%
23
Singh Tenali Mathur Soni Phatak Lynch Murman Rizzo Natarajan
Appendix E. Structure Learning using Peter Clarke (PC) Algorithm We compared the LLM-generated Bayesian Network structure to one obtained by the datadriven Peter Clarke Algorithm. For this baseline, we learned graphs over the semantic and phonetic features extracted by our pipeline. In the phonemic task, PC was only able to recover the edge CI → AoA from the expert graph, and similarly in the semantic task, only the edge CI → TW was recovered. CI
NC
AoA
WL
CI
WF
TW
WL
CS
SR
NC
AoA WF
TW
CS
SR
SW
SW
Figure 10: The Influence Networks over Phonemic (Left) and Semantic features (Right) learned using PC algorithm representing direct influence relationships between CI , Lexical , Semantic , and Acoustic variables.
24