Conceptio › Archive › arXiv CS
arXiv CSopen access

It's All in the Way You Say It: The Role of Information Representation in LLM-Based Glycemic-Event Prediction

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

It’s All in the Way You Say It: The Role of Information Representation in LLM-Based Glycemic-Event Prediction Andrea Apicellaa , Pasquale Arpaiab , Matteo Oreficeb , Andrea Pollastrob , Roberto Preveteb a Department of Information Engineering, Electrical Engineering, and Applied Mathematics (DIEM), University of Salerno, Fisciano (SA), Italy b Department of Electrical Engineering and Information Technology, University of Naples Federico II, Naples, Italy

Abstract

arXiv:2609.08772v1 [cs.AI] 8 Sep 2026

Large Language Models (LLMs) are increasingly being investigated for physiological time-series prediction, yet their effectiveness may depend not only on the model itself, but also on how physiological information is represented and presented at inference time. This study investigates prompt-based general-purpose LLMs for postprandial hyperglycemia and hypoglycemia prediction in individuals with type 1 diabetes. Using the OhioT1DM dataset, we evaluate multiple open-weight LLMs under zero-shot and few-shot inference across prediction horizons of 30, 60, and 90 minutes. The analysis varies both the textual representation of the available physiological information and the amount of information exposed to the model, ranging from glucose observations alone to derived descriptors and additional contextual variables related to insulin, meals, carbohydrates, and physical activity. Performance is compared with conventional patient-specific supervised models and with Gluco-LLM, a language-model-based architecture explicitly adapted to glucose time-series forecasting. Results show a marked task-dependent behavior. Conventional supervised models achieve the strongest performance for hyperglycemia prediction, whereas the best observed prompt-based LLM configurations improve performance for hypoglycemia across all investigated horizons. The effectiveness of prompt-based inference is also strongly influenced by how physiological information is represented, while providing additional contextual information does not lead to a systematic improvement. Overall, these findings highlight physiological information representation as a central design factor in prompt-based LLM approaches to glycemic-event prediction. Keywords: LLM, T1DM, Diabete, Hypoglycemy, Hyperglycemy, Machine Learning, Large Language Models

1. Introduction Type 1 diabetes (T1D) is a chronic autoimmune condition characterized by the destruction of pancreatic β-cells, resulting in little or no endogenous insulin production [1]. Consequently, blood glucose regulation relies on exogenous insulin administration and continuous monitoring of glucose levels. Inadequate glucose control may lead to hyperglycemia, i.e., excessively high blood glucose levels, or hypoglycemia, i.e., abnormally low glucose levels, both of which can have clinically relevant consequences. Continuous Glucose Monitoring (CGM) systems provide frequent measurements of interstitial glucose and therefore represent a key source of information for anticipating such adverse glycemic events. Large Language Models (LLMs) have recently emerged as a promising alternative for addressing a wide range of medical prediction and decision-support tasks [2, 3, 4, 5]. Importantly, LLM-based and LLM-assisted approaches have attracted increasing attention for blood glucose forecasting in people with T1D, showing encouraging results in personalized prediction and in the integration of heterogeneous physiological information [6, 7]. This development is particularly relevant given the growing global burden of T1D [8] and the challenges that conventional Machine Learning (ML) approaches still face in capturing highly individual glucose dyThis work has been submitted to a journal for peer review.

namics and effectively exploiting heterogeneous physiological inputs [9, 10, 11, 12]. In this work, we specifically focus on postprandial glycemic-event prediction, where the objective is to anticipate hyperglycemic and hypoglycemic events following meal intake. The postprandial period represents a particularly challenging and clinically relevant setting, as glucose dynamics are strongly affected by the interplay between several factors, such as carbohydrate absorption, insulin administration, and individual physiological responses. Accurate prediction in this phase may therefore support earlier identification of potentially adverse glycemic excursions and enable more timely preventive interventions [13]. Furthermore, general purpose LLMs may offer a practical advantage in terms of accessibility. Unlike conventional supervised models, they can be applied to a new task without task-specific parameter training and can expose their predictive capabilities through a natural language interface. Thus, this interaction paradigm can mitigate the technical barrier for non-specialist users, who would not be required to configure, train, and execute a dedicated ML pipeline. In the context of postprandial glycemic-event prediction, however, accessibility also depends on how physiological information can be communicated to the model. If clinically relevant descriptors can be expressed in natural language without a substantial loss in predictive performance, LLM-based systems could support the development of more intuitive user tools. Moreover, LLM-

based inference also offers practical properties that are not captured by predictive performance alone. A general purpose LLM can address the task without patient-specific parameter training, potentially reducing the computational and technical effort required to train and maintain a dedicated predictive model. Unlike conventional ML algorithms, however, general purpose LLMs are not explicitly designed to process numerical physiological time series such as Continuous Glucose Monitoring (CGM) signals. Consequently, before they can be employed for glucose prediction, physiological observations must first be transformed into a textual representation that can be interpreted through natural language processing. This transformation is not merely an implementation detail. The effectiveness of prompt-based processing depends on two complementary factors. The first concerns how physiological observations are represented before being presented to the LLM. For example, the original CGM sequence may be provided directly, summarized through clinically meaningful descriptors, or verbalized as natural language statements. The second concerns which physiological and clinical information is made available to the model. Besides the glucose measurements themselves, additional descriptors or contextual variables related to insulin administration, carbohydrate intake, and physical activity may also be incorporated. These two dimensions jointly determine the evidence available to the LLM and can substantially influence its prediction capabilities. Although several recent studies have explored the use of LLMs for glucose prediction in T1DM and Type 2 Diabetes Mellitus (T2DM), existing approaches generally follow one of two directions. The first adapts pretrained LLM to time series forecasting through task-specific optimization [14, 15, 16, 6]. The second directly queries LLMs through proper prompts [17, 18] and evaluates their predictions against conventional supervised learning approaches. While both directions have demonstrated encouraging results, they provide only limited understanding of a more fundamental question: how should physiological information be represented before being presented to a language model, and how much physiological information is actually required to support reliable clinical prediction? This work addresses these issues from a different perspective. Rather than proposing a new prediction architecture or further adapting an existing LLM, this work provides an empirical evaluation of how physiological information and its different representations can affect prompt-based processing. To this aim, we exploit the multimodal nature of the OhioT1DM dataset [19], which provides CGM measurements together with insulin therapy data, meal and carbohydrate information, self-reported behavioral and contextual events, and wearablederived physiological signals. In the present study, a selected subset of these modalities is used to define alternative information settings, ranging from CGM-derived information to configurations augmented with additional clinical and behavioral variables. To provide a comprehensive assessment, prompt-based inference is evaluated together with two complementary prediction paradigms. The first consists of conventional supervised

ML models trained specifically for postprandial glycemic-event prediction. The second is represented by Gluco-LLM [14], a language model explicitly adapted to physiological T1DM time series forecasting through task-specific architectural modifications. Comparing these three paradigms allows us to assess the potential of prompt-based processing and whether task-specific adaptation remains necessary for accurate glucose prediction. The experimental evaluation is organized around three complementary research questions: • RQ1. Can current general purpose LLMs provide predictive performance comparable to conventional supervised learning models for postprandial glycemic-event prediction? • RQ2. How do the representation of physiological information and the type of available physiological information in LLMs influence the capabilities of current general purpose LLMs? • RQ3. Can a non-adapted general purpose LLM achieve predictive performance comparable to an LLM explicitly adapted for physiological time series forecasting, such as Gluco-LLM? To address RQ1, we evaluate general purpose LLMs with respect to conventional supervised learning, treating these approaches as alternative prediction paradigms rather than simply performance baselines. To address RQ2, we analyze three alternative representation strategies, namely Raw, Structured, and Narrative. They represent three different and complementary ways to build the LLM input (prompt) from raw physiological data. Raw preserves the original physiological data via textual serialization, while Structured and Narrative represent derived physiological descriptors in terms of key–value pairs and natural language, respectively. The latter two representations are provided in two different modalities corresponding to different sets of physiological information, namely CGMderived and context CGMderived : the former contains descriptors derived from the historical CGM signal, whereas the latter additionally incorporates selected contextual variables related to insulin administration, carbohydrate intake, and physical activity (see Sec. 3 for further details). Finally, to address RQ3 we investigate whether prompt-based processing with general purpose LLMs can approach the predictive capability of a language model specifically adapted to physiological time series forecasting, thereby quantifying the benefits provided by task-specific architectural adaptation. Our results show that the relative effectiveness of the considered approaches depends on the prediction task. Conventional supervised models achieve the strongest performance for postprandial hyperglycemia prediction, whereas general purpose LLMs are particularly competitive for hypoglycemia prediction. We further observe that LLM performance is sensitive to how physiological information is represented, with no representation proving uniformly optimal across models, tasks, and prompting conditions. When performance is summarized by selecting the best observed result across LLM families, Structured 2

inputs are most frequently associated with the strongest configurations. However, within matched model configurations, Raw inputs become more competitive under few-shot inference, where a small number of labeled examples are provided in the prompt as in-context demonstrations without updating the model parameters [20], while Narrative representations remain preferable in selected settings. Finally, Gluco-LLM generally outperforms prompt-based general purpose LLMs in the regression experiments, although general purpose prompt-based inference remains competitive in selected conditions.

al. [16] proposed DiabLLM, an LLM-based framework for blood glucose prediction in T1D that adapts recent languagemodel-inspired forecasting architectures, namely Time-LLM and Chronos, to CGM data. Unlike prompt-based approaches, DiabLLM exploits pretrained LLM backbones for numerical time series forecasting. Differently, Healey et al. [17] investigated the use of GPT4 for the analysis and summarization of continuous glucose monitoring (CGM) data rather than for glucose prediction, investigating if LLMs can effectively transform raw physiological time series into clinically meaningful narrative descriptions. Healey and Kohane [29] proposed a benchmark for conversational analysis of CGM data that evaluates GPT-4-based text, code generation, and agentic frameworks on a range of objective glucose-related queries. He et al. [30] developed a personalized diabetes treatment-support system by fine-tuning compact general purpose LLMs with LoRA [31] on electronic health records. Ding et al. [32] investigated how heterogeneous clinical information can be represented and integrated by language-model-based predictors, combining clinical notes with both numerical and textualized laboratory values for newonset T2D prediction. Taken together, these studies demonstrate the growing interest in LLM-based approaches to physiological time series analysis and glucose prediction, while adopting different strategies for representing and integrating physiological and contextual information. However, to the best of our knowledge, no prior work has investigated how the representation and type of physiological information affect prompt-based glucose prediction, particularly in the context of postprandial glucose forecasting in T1DM.

2. Related Works Glucose forecasting in T1D has been extensively investigated using ML approaches and evaluated across multiple prediction horizons [21, 22, 23, 24]. Cui et al. [25] addressed the joint prediction of postprandial hyperglycemia and hypoglycemia from CGM data, proposing a patient-specific LSTM model that uses historical glucose measurements to estimate the maximum and minimum glucose values within a given prediction horizon. Their formulation provides an explicit protocol for constructing postprandial prediction instances and defining hyperglycemic and hypoglycemic events under a 30-minute prediction horizon. More recently, several studies have demonstrated that general purpose LLMs can be applied to time series forecasting by representing numerical observations as textual sequences. For instance, Gruver et al. [26] introduced LLMTime, showing how general purpose LLMs can perform zero-shot time series forecasting by representing numerical sequences as text and formulating forecasting as a next-token prediction task. Their work established a paradigm for treating numerical time series as textual sequences for prompt-based forecasting, providing a methodological foundation that can also be applied to physiological time series tasks. Alredaini et al. [18] compared traditional ML, DL, and LLMs for multi-horizon glucose forecasting. The authors formulate the task as a natural language reasoning problem by encoding historical CGM measurements and patient-specific characteristics into textual prompts, allowing fine-tuned LLMs such as GPT-4.1 [27] and LLaMA [28] to directly predict future glucose values, showing the feasibility of prompt-based formulations for physiological time series prediction. Lara-Abelenda et al. [6] investigated the use of Time-LLM for personalized glucose forecasting in people with T1D reprogramming pretrained LLM backbones to process continuous glucose monitoring time series through patch reprogramming. Li et al. [14] introduced Gluco-LLM, a personalized glucose-forecasting framework based on TimeLLM, in which a frozen pretrained LLM backbone is coupled with trainable time series reprogramming and output-projection modules to integrate CGM and patient-specific contextual information. Gao et al. [15] proposed GlyLLM, an LLMpowered framework for personalized glycemic assessment in Type 2 Diabetes (T2D) that integrates continuous glucose monitoring data, wearable sensor signals, and patient-specific metadata to explore if incorporating personalized textual metadata together with wearable sensor information improves predictive performance over conventional ML baselines. Mahmoudi et

3. Method This work investigates postprandial glycemic-event prediction through the comparison of three distinct and complementary approaches. The first approach comprises conventional supervised ML models trained specifically for glucose prediction from patient-specific physiological observations. The second formulates glucose prediction as a prompt-based task performed by an unchanged general purpose LLM. The third approach is represented by Gluco-LLM [14]. Unlike supervised ML models, which learn a prediction function directly from labeled physiological data, and unlike Gluco-LLM, which adapts a pretrained LLM through dedicated trainable components, the chosen prompt-based framework intentionally leaves the underlying LLM unchanged throughout the entire experimental protocol. Consequently, any observed variation in predictive performance can be attributed to differences in the information provided to the model rather than to parameter optimization or architectural modifications. This design enables a principled investigation of prompt-based processing independently of task-specific model adaptation. Figure 1 summarizes the overall experimental framework. Starting from the same postprandial prediction setting, the three considered approaches are evaluated on aligned prediction tasks 3

and horizons. Standard supervised models directly process numerical physiological inputs, whereas Gluco-LLM operates on numerical time series data through its dedicated time series adaptation pipeline. In contrast, the prompt-based configuration used in our experiments converts the available physiological information into a textual representation, which is then incorporated into a prompt and submitted to a general purpose LLM. Within the prompt-based framework, two complementary aspects of the input are investigated. The first concerns physiological information representation, namely how the available patient observations are expressed before being presented to the language model. To investigate how the form in which physiological information is presented affects LLM-based prediction, we define three alternative representation strategies: Raw, Structured, and Narrative. Raw preserves the original CGM sequence through direct textual serialization, whereas Structured and Narrative transform derived physiological descriptors into, respectively, a key–value representation and a naturallanguage description. The second aspect concerns the physiological information made available to the model. In particular, for Structured and Narrative representations, two information settings are evaluated: i) CGMderived contains only descriptors derived exclusively from the historical CGM sequence, whereas context ii) CGMderived augments these descriptors with additional clinical and therapeutic variables related to insulin administration, carbohydrate intake, basal insulin, meal timing, and physical activity. All these variables are summarized in Table 1. Notice that the effect of linguistic representation can be examined by comparing Structured and Narrative under the same information setting, whereas the contribution of additional physiological information can be assessed by comparcontext ing CGMderived and CGMderived within the same representation. Comparisons involving Raw are complementary, since Raw preserves the original CGM sequence whereas Structured and Narrative operate on descriptors derived from that sequence; such comparisons therefore jointly reflect differences in representation and information transformation. Under this perspective, the prompt acts as the interface through which different representations of the available physiological information are communicated to the language model, while the underlying model parameters remain unchanged.

of anticipating adverse glycemic events following meal intake, where glucose dynamics are strongly influenced by the interaction between carbohydrate absorption and insulin action. As above discussed, the information made available to the prediction model may correspond either to the historical CGM sequence itself, to a set of descriptors derived from that sequence, or to the same descriptors augmented with additional contextual information. We denote by I(s, t) the information available to the model for sample t under information setting s. Specifically,  Xt , s = CGMonly ,    I(s, t) = ϕ(Xt ), s = CGMderived ,    context [ϕ(Xt ), Ct ], s = CGMderived , where Xt = {gt−T +1 , gt−T +2 , . . . , gt } denotes the historical CGM sequence available at prediction time t, T is the observation-window length, and gi is the glucose concentration measured at time i. The function ϕ(·) maps the historical CGM window Xt to a set of statistical and temporal descriptors computed exclusively from the CGM observations available up to time t. In contrast, Ct denotes the collection of contextual variables available at prediction time t that cannot be derived dicontext rectly from Xt . Accordingly, the CGMderived setting combines the CGM-derived descriptors ϕ(Xt ) with the additional contextual information Ct . The variables included in ϕ(Xt ) and Ct are summarized in Table 1, while the corresponding preprocessing procedures are detailed in Section 4.1. The three feature settings were designed to capture complementary ways of making patient information available to the prediction model. The CGMonly setting preserves the original temporal structure of the CGM signal and requires the model to infer relevant temporal patterns directly from the glucose trajectory. By contrast, as already above discussed and summarized in Table 1, the CGMderived setting instead adopts only set of CGM-derived characteristics describing the patient’s current glycemic state, recent temporal evolution, variability, and exposure to clinically relevant glucose ranges. Finally, the context CGMderived setting complements these CGM-derived descriptors with contextual factors that may affect postprandial glucose dynamics but are not directly observable from CGM alone, including information related to insulin administration, carbohydrate intake, meal timing, and physical activity. This design allows us to investigate whether LLM-based prediction benefits from direct access to the original CGM trajectory, from an explicit representation of characteristics derived from that trajectory, or from the availability of complementary contextual information beyond CGM

3.1. Problem formulation To address the research objectives of this work, we adopt a meal-centered formulation. Each meal defines a postprandial episode, and prediction samples are generated within the corresponding postprandial period. Following the protocol proposed by Cui et al. [25], each meal defines a four-hour monitoring interval. The first two hours are considered for hyperglycemia prediction, whereas the following two hours are considered for hypoglycemia prediction. This temporal separation is consistent with the different temporal profiles of postprandial glycemic excursions: hyperglycemic events are more likely to occur during the early postprandial phase, while hypoglycemic events typically emerge later as the effects of insulin action become more prominent. This formulation reflects the objective

3.1.1. Type of predictions The objectives of this study are addressed in the context of a postprandial glycemic event prediction task, formulated as a binary decision concerning the occurrence of hyperglycemia or hypoglycemia within a selected prediction horizon. Following Cui et al. [25], two alternative formulations are investigated. In the direct-classification formulation, the available physiological information is mapped directly to a binary event predic4

Figure 1: The framework compares three complementary approaches for postprandial glycemic-event prediction from continuous glucose monitoring (CGM) data: (i) conventional supervised ML models, (ii) prompt-based inference with a general purpose LLM using alternative physiological information representations, and (iii) Gluco-LLM [14], a recently proposed LLM-based architecture explicitly adapted to glucose time series forecasting. The chosen prompt-based framework provides the primary methodological basis for investigating the three research questions addressed in this work.

tion. In the regression-based formulation, the model first estimates a continuous future glucose outcome, from which the binary event prediction is subsequently obtained by applying the corresponding clinical threshold on the maximum for hyperglycemia or the minimum for hypoglycemia. Depending on the considered approach, the relevant future glucose extreme, i.e. the maximum for hyperglycemia or the minimum for hypoglycemia, is either predicted directly or extracted from a predicted future glucose trajectory.

mance to be assessed both at the continuous glucose level and, after thresholding, at the event-classification level. Given the physiological information I(s, t) available at prediction time t, the objective is to estimate whether an adverse glycemic event (namely, hyperglycemia and hypoglicemia) will occur within the selected prediction horizon h for the classification problem, and the future glucose extreme occurring within the selected prediction horizon h for the regression problem. For the classification problem, hyperglycemia and hypoglycemia classification target (i.e., the labels) are generated according to the postprandial prediction protocol proposed by Cui et al. [25], assigning a positive label when at least two consecutive CGM readings within the prediction horizon are above 180 mg/dL or below 70 mg/dL, respectively, and a negative label otherwise. Instead, for the regression problem, the regression target is defined as the maximum or the minimum CGM reading within the prediction horizon in the hyperglycemia and hypoglycemia regression problem, respectively.

We highlight that, as in Cui et al. [25], this threshold-based event decision obtained from the predicted glucose extreme is not strictly equivalent to the direct-classification event definition, which requires at least two consecutive CGM readings beyond the corresponding threshold. The regression-based formulation should therefore be interpreted as an approximate eventdetection formulation derived from the continuous prediction target rather than as an exact reformulation of the direct binaryclassification task. However, this formulation enables perfor5

Table 1: Information available at prediction time t under the considered information settings. The CGMderived setting contains descriptors ϕ(Xt ) derived from the context setting additionally includes contextual variables C available at prediction time t. historical CGM window Xt , whereas the CGMderived t Setting

Attribute

Notation

Description

CGMonly

Historical CGM sequence

Xt = {gt−T +1 , . . . , gt }

CGM measurements collected during the historical observation window ending at t.

CGMderived

Current glucose Mean glucose Standard deviation Minimum glucose Maximum glucose Recent minimum Coefficient of variation Glucose slope Most recent glucose change Time above range Time below range Time in range Number of readings Window duration

gt µ(Xt ) σ(Xt ) min(Xt ) max(Xt ) min(Xtrecent ) CV(Xt ) slope(Xt ) gt − gt−1 TAR(Xt ) TBR(Xt ) TIR(Xt ) |Xt | ∆(Xt )

Most recent CGM measurement available at prediction time t. Mean glucose concentration over the historical CGM window. Variability of glucose measurements within the historical window. Minimum glucose concentration observed within the historical window. Maximum glucose concentration observed within the historical window. Minimum glucose concentration over the predefined recent portion of the historical window. Glucose variability normalized with respect to the mean glucose level. Estimated temporal trend of glucose measurements within the historical window. Difference between the two most recent CGM measurements. Fraction of the historical window spent above the target glucose range. Fraction of the historical window spent below the target glucose range. Fraction of the historical window spent within the target glucose range. Number of valid CGM measurements available in the historical window. Temporal duration covered by the available CGM measurements.

context CGMderived

Insulin on board Carbohydrates on board Basal insulin rate Number of boluses Number of meals Time since last meal Physical-activity duration

IOBt COBt Basalt Nbolus,t Nmeal,t ∆tmeal,t At

Estimated amount of previously administered insulin still active at time t. Estimated amount of previously ingested carbohydrates still being absorbed at time t. Basal insulin delivery rate available at prediction time t. Number of bolus insulin administrations considered before prediction time t. Number of recorded meals considered before prediction time t. Elapsed time between the most recent recorded meal and prediction time t. Duration of recent physical activity considered at prediction time t.

context information setting s ∈ {CGMonly , CGMderived , CGMderived }, a prediction task τ ∈ {Classification, Regression}, a prediction horizon h, the prompt-generation function Pt (r, s, τ, h, c) constructs the textual input submitted to the LLM for the sample observed at time t. Details about the prompt function will be provided in the following of this section. The resulting prediction is expressed as ŷt,h = LLM (Pt (r, s, τ, h, c)) , where (LLM) denotes the investigated general purpose LLM.

To evaluate whether the continuous prediction also correctly identifies the corresponding clinical event, binary predictions are subsequently derived by applying the same glucose thresholds adopted in the direct classification setting, i.e. hyperglicemia if the predicted value is above 180 mg/dL and hypoglicemia if the predicted value is below 70 mg/dL. This formulation therefore permits the models to be evaluated both in terms of similarity to the predicted glucose extreme and in terms of their ability to identify relevant glycemic events after thresholding. For conventional supervised ML models, the information I(s, t) available for a patient sample is provided directly to the classifier. Differently, general purpose LLMs interact with the prediction task through a textual interface. Consequently, the available physiological information must be encoded into a compatible representation before being provided to the model. As above introduced, we investigate three alternative prompt representation modalities, referred to as Raw, Structured, and Narrative. The Raw representation preserves the historical CGM sequence through direct textual serialization, whereas the Structured and Narrative represent CGM-derived descriptors and, when available, additional contextual information through, respectively, explicit key-value pairs and natural language descriptions. Each prompt is generated deterministically from the physiological information available at the corresponding prediction time, using predefined templates and representation rules. Furthermore, two inference conditions are considered, zero-shot inference (ZS) and few-shot inference (FS). Under ZS, the prompt contains the task instructions and the target sample without labeled examples. Under FS, the target query is preceded by K labeled in-context demonstrations. Finally, the prompt reference answer is a binary label for classification and the corresponding future glucose extreme for regression. More formally, given an inference condition c ∈ {ZS, FS}, a representation modality r ∈ {Raw, Structured, Narrative}, an

3.1.2. Prompt construction function Pt For a patient sample observed at time t, the prompt Pt (r, s, τ, h, c) submitted to the LLM is defined as Pt (r, s, τ, h, c) = S ⊕ E(r, s, τ, h, c) ⊕ Rr (I(s, t)) ⊕ Q(τ, h), where ⊕ denotes textual concatenation, I(s, t) denotes the information available under information setting s as discussed in Section 3.1, Rr (·) denotes its textual encoding according to the representation modality r, Q(τ, h) specifies the prediction task τ and the prediction horizon h, and E(r, s, τ, h, c) accounts for the possible in-context examples included in the LLM prompt under the zero-shot/few-shot inference condition c ∈ {ZS , FS }. The system instruction S is kept fixed across all experimental configurations: You are an AI assistant in a clinic with the task of aiding doctors by predicting glycemic events in Type 1 Diabetes patients based on continuous glucose monitor data.

Rr , E, and Q are defined in the following. Raw Modality prompt component (RRaw ). With this representation modality, the LLM receives the original historical glucose observations. For the Raw representation, RRaw is obtained by directly serializing the historical CGM sequence, without derived descriptors: 6

RRaw (CGM) = serialize (gt−T +1 , . . . , gt ) . The serialize reports in a textual representation the CGM measurements in chronological order. The most recent measurement gt , corresponding to the current glucose level, is additionally reported explicitly in the textual template. No feature extraction or clinical interpretation is performed. The resulting textual block returned by the serialize function follows the deterministic template

60 minutes. The current glucose level is [value] mg/dL. [Automatically generated physiological description]

Therefore, keys and their values are inserted into predefined natural language expressions and combined into coherent sentences describing the patient’s recent glucose dynamics and, when available, the additional contextual information. Thus, for a fixed information setting s, Structured and Narrative representations operate on the same underlying information I(s, t) and differ only in their textual realization: the former uses explicit key–value pairs, whereas the latter expresses the same content through natural-language sentences.

CGM readings (last 60 minutes, every 5 minutes): [gt−T +1 , ..., gt ] mg/dL Current glucose: gt mg/dL

Structured modality prompt component (RStructured ). The Structured representation expresses the available patient information through explicit descriptors organized as key–value pairs. Unlike the Raw representation, the language model does not need to infer basic descriptive properties directly from the glucose sequence, since relevant statistical quantities and, when available, additional contextual variables are explicitly provided. The key–value organization also makes the semantic role of each variable explicit, allowing the model to operate directly on higher-level descriptors of the patient’s recent state. For the Structured representation, the textual encoding is obtained by applying the representation function RStructured to I(s, t): Lms RStructured (I(s, t)) = j=1 format(k j , v j ), where (k j , v j ) denotes the j-th key–value pair associated with the selected information setting, m s is the corresponding number of variables, and ⊕ denotes textual concatenation. The resulting textual representation given by the format(·) function has the general form

Zero-shot and few-shot prompt component (E). Prompt-based inference is investigated under both zero-shot (ZS) and fewshot (FS) conditions. In both cases, the parameters of the language model remain unchanged, and task-specific conditioning is provided exclusively through the prompt. Following the prompt formalization introduced above, the inference condition is denoted by c ∈ {ZS, FS} and determines whether labeled in-context examples are included through the term E(r, s, τ, h, c). In the ZS condition, no labeled examples are provided and therefore E(r, s, τ, h, ZS) returns the empty textual sequence. Consequently, the prompt reduces to Pt (r, s, τ, h, ZS) = S ⊕ Rr (I(s, t)) ⊕ Q(τ, h). The model is therefore required to perform the prediction without observing any labeled examples of the considered task. In the FS condition, E(r, s, τ, h, FS) contains K labeled incontext examples selected according to the procedure described in Section 4. Each example is constructed using the same representation strategy r, information setting s, prediction task τ, and prediction horizon h adopted for the target sample, and includes the corresponding input representation together with its expected output. LK Formally, E(r, s, τ, h, FS) = i=1 E i (r, s, τ, h), where Ei (r, s, τ, h) denotes the i-th labeled in-context example and ⊕ denotes textual concatenation. The corresponding few-shot prompt is Pt (r, s, τ, h, FS) = S ⊕ E(r, s, τ, h, FS) ⊕ Rr (I(s, t)) ⊕ Q(τ, h). For classification, the expected outputs associated with the in-context examples are Yes or No; whenever possible, positive and negative examples are balanced to avoid introducing an artificial preference toward either class. For regression, the expected output is the corresponding numerical glucose extreme expressed in mg/dL. The resulting text from the E(r, s, τ, h, FS) function when r = Narrative and τ = Classification has the form:

attribute name: value Therefore, the final RStructured returns a textual representation of the form Patient: Type 1 Diabetes Monitoring window: 60 minutes Current glucose: [value] mg/dL Mean glucose: [value] mg/dL ...

Narrative modality prompt component (RNarrative ). The Narrative representation expresses the information available under a given information setting through natural language sentences. context For a fixed information setting s ∈ {CGMderived , CGMderived }, it is generated from the same underlying information I(s, t) used by the Structured representation, differing only in the way this information is textually presented to the language model. Formally, the Narrative representation is obtained by applying a verbalization function to the available information: RNarrative (I(s, t)) = verbalize(I(s, t)). The function verbalize(·) encodes the available variables as follows:

Example 1: A patient with type 1 diabetes has been monitored during the last 60 minutes. The current glucose level is [value] mg/dL. [Automatically generated physiological description]

A patient with type 1 diabetes has been monitored during the last

7

prediction. The model is instructed to return a single integer expressed in mg/dL, which is subsequently parsed for automatic evaluation. For the classification formulation, the task-specific instruction is defined as

Will this patient experience [hyperglycemia/hypoglycemia] in the next [horizon] minutes? Answer: [Yes/No] Example 2: A patient with type 1 diabetes has been monitored during the last 60 minutes. The current glucose level is [value] mg/dL. [Automatically generated physiological description] Will this patient experience [hyperglycemia/hypoglycemia] in the next [horizon] minutes? Answer: [Yes/No]

Will this patient experience hyperglycemia in the next h minutes? Answer Yes or No.

for hyperglycemia prediction, and Will this patient experience hypoglycemia in the next h minutes? Answer Yes or No.

for hypoglycemia prediction. For the regression formulation, the task-specific instruction is defined as

... Example K: A patient with type 1 diabetes has been monitored during the last 60 minutes. The current glucose level is [value] mg/dL. [Automatically generated physiological description] Will this patient experience [hyperglycemia/hypoglycemia] in the next [horizon] minutes? Answer: [Yes/No]

Predict the maximum glucose value (mg/dL) the patient will reach in the next h minutes. Answer with a single integer.

for hyperglycemia prediction, and Predict the minimum glucose value (mg/dL) the patient will reach in the next h minutes. Answer with a single integer.

for hypoglycemia prediction. In Tab. 2 the prompt construction is summarized.

Now consider the following case:

3.1.3. Output parsing LLM responses are automatically converted into the output format required for quantitative evaluation. The parsing procedure depends on the considered prediction formulation and is applied consistently across all models and experimental conditions. For direct classification, the model is instructed to return exclusively Yes or No. The resulting response is mapped to the corresponding positive or negative class, respectively, enabling direct comparison with the predictions produced by supervised classifiers. For regression, the model is instructed to return a numerical estimate of the future glucose extreme, expressed in mg/dL. Since LLMs may occasionally include additional text in their response, the generated output is automatically parsed to recover the predicted glucose value. Specifically, the first valid numerical value occurring in the generated response is extracted and interpreted as the predicted glucose value. If the model returns a numerical interval, its midpoint is used as the prediction. Responses from which no valid numerical value can be recovered are treated as invalid predictions. If no valid glucose value can be identified, the response is considered inconclusive and excluded from the evaluation. The parsed regression output is evaluated directly as a continuous glucose prediction and is additionally converted into a binary event prediction by applying the corresponding clinical threshold. This enables the regression formulation to be assessed both in terms of numerical prediction accuracy and glycemic-event detection.

The same ZS and FS protocols are adopted for the regression task, with the task instruction and expected response adapted to the continuous prediction target. Specifically, whereas classification demonstrations associate each input with a binary Yes/No label, regression demonstrations associate it with the corresponding future glucose extreme expressed in mg/dL. In every configuration, the in-context demonstrations are constructed using the same physiological representation (r) and information setting (s) as the target query. Accordingly, demonstrations based on the Structured and Narrative representations context are generated under either the CGMderived or CGMderived information setting, consistently with the evaluated query. The number (K) of in-context demonstrations and their selection procedure are detailed in the experimental setup, as they constitute experimental design choices rather than intrinsic components of the proposed framework. Task-specific prompt component (Q). For the classification formulation, Q(τ, h) instructs the language model to determine whether a hyperglycemic or hypoglycemic event will occur within the specified prediction horizon and to answer exclusively with Yes or No. Restricting the expected output format enables deterministic mapping of the generated response to the corresponding binary prediction. For the regression formulation, Q(τ, h) asks the model to estimate the relevant future glucose extreme within the selected prediction horizon. Specifically, the requested output corresponds to the maximum glucose value for hyperglycemia prediction and to the minimum glucose value for hypoglycemia 8

Table 2: Representative prompt structures obtained by combining the common system instruction S , the in-context component E, the task-specific component Q, and the representation component Rr . The table shows hyperglycemia prediction at a generic horizon h. Raw prompts use the CGMonly information setting, whereas context setting, the same templates additionally include the available contextual Structured and Narrative prompts are illustrated under CGMderived . Under the CGMderived variables. In few-shot prompts, only the first of the K demonstrations is shown explicitly. Common component S E ZS

Q Classification

Regression

FS

Classification

Regression

You are an AI assistant in a clinic with the task of aiding doctors by predicting glycemic events in Type 1 Diabetes patients based on continuous glucose monitor data. Rr Raw Structured Narrative CGM readings (last 60 minutes, every Patient: Type 1 Diabetes A patient with type 1 diabetes has 5 minutes): Monitoring window: 60 minutes been monitored during [gt−T +1 , ..., gt ] mg/dL Current glucose: [value] mg/dL the last 60 minutes. The current Current glucose: gt mg/dL Mean glucose: [value] mg/dL glucose level is Will this patient experience [Additional CGM-derived descriptors] [value] mg/dL. [Physiological hyperglycemia in the next Will this patient experience description] h minutes? Answer Yes or No. hyperglycemia in the next Will this patient experience h minutes? Answer Yes or No. hyperglycemia in the next h minutes? Answer Yes or No. CGM readings (last 60 minutes, every Patient: Type 1 Diabetes A patient with type 1 diabetes has 5 minutes): Monitoring window: 60 minutes been monitored during [gt−T +1 , ..., gt ] mg/dL Current glucose: [value] mg/dL the last 60 minutes. The current Current glucose: gt mg/dL Mean glucose: [value] mg/dL glucose level is Predict the maximum glucose value [Additional CGM-derived descriptors] [value] mg/dL. [Physiological (mg/dL) the patient Predict the maximum glucose value description] will reach in the next h minutes. (mg/dL) the patient Predict the maximum glucose value Answer with a will reach in the next h minutes. (mg/dL) the patient single integer. Answer with a will reach in the next h minutes. single integer. Answer with a single integer. Example 1: Example 1: Example 1: CGM readings (last 60 minutes, every Patient: Type 1 Diabetes A patient with type 1 diabetes has 5 minutes): Monitoring window: 60 minutes been monitored during [gt−T +1 , ..., gt ] mg/dL Current glucose: [value] mg/dL the last 60 minutes. The current Current glucose: gt mg/dL Mean glucose: [value] mg/dL glucose level is Will this patient experience [Additional CGM-derived descriptors] [value] mg/dL. [Physiological hyperglycemia in the next Will this patient experience description] h minutes? Answer Yes or No. hyperglycemia in the next Will this patient experience Answer: [Yes/No] h minutes? Answer Yes or No. hyperglycemia in the next .. Answer: [Yes/No] h minutes? Answer Yes or No. . .. Answer: [Yes/No] Example K: [same format] . .. Now consider the following case: Example K: [same format] . CGM readings (last 60 minutes, every 5 minutes): [gt−T +1 , ..., gt ] mg/dL Current glucose: gt mg/dL Will this patient experience hyperglycemia in the next h minutes? Answer Yes or No.

Now consider the following case: Patient: Type 1 Diabetes Monitoring window: 60 minutes Current glucose: [value] mg/dL Mean glucose: [value] mg/dL [Additional CGM-derived descriptors] Will this patient experience hyperglycemia in the next h minutes? Answer Yes or No.

Example 1: CGM readings (last 60 minutes, every 5 minutes): [gt−T +1 , ..., gt ] mg/dL Current glucose: gt mg/dL Predict the maximum glucose value (mg/dL) the patient will reach in the next h minutes. Answer with a single integer. Answer: [integer] .. .

Example 1: Patient: Type 1 Diabetes Monitoring window: 60 minutes Current glucose: [value] mg/dL Mean glucose: [value] mg/dL [Additional CGM-derived descriptors] Predict the maximum glucose value (mg/dL) the patient will reach in the next h minutes. Answer with a single integer. Answer: [integer] .. .

Example K: [same format] Now consider the following case: CGM readings (last 60 minutes, every 5 minutes): [gt−T +1 , ..., gt ] mg/dL Current glucose: gt mg/dL Predict the maximum glucose value (mg/dL) the patient will reach in the next h minutes. Answer with a single integer.

Example K: [same format] Now consider the following case: Patient: Type 1 Diabetes Monitoring window: 60 minutes Current glucose: [value] mg/dL Mean glucose: [value] mg/dL [Additional CGM-derived descriptors] Predict the maximum glucose value (mg/dL) the patient will reach in the next h minutes. Answer with a single integer.

Example K: [same format] Now consider the following case: A patient with type 1 diabetes has been monitored during the last 60 minutes. The current glucose level is [value] mg/dL. [Physiological description] Will this patient experience hyperglycemia in the next h minutes? Answer Yes or No. Example 1: A patient with type 1 diabetes has been monitored during the last 60 minutes. The current glucose level is [value] mg/dL. [Physiological description] Predict the maximum glucose value (mg/dL) the patient will reach in the next h minutes. Answer with a single integer. Answer: [integer] .. . Example K: [same format] Now consider the following case: A patient with type 1 diabetes has been monitored during the last 60 minutes. The current glucose level is [value] mg/dL. [Physiological description] Predict the maximum glucose value (mg/dL) the patient will reach in the next h minutes. Answer with a single integer.

specifically designed to adapt pretrained language models to glucose time series forecasting. Gluco-LLM builds upon the Time-LLM framework [33] and represents a fundamentally dif-

3.2. Time series-adapted LLM: Gluco-LLM In addition to prompt-based inference with general purpose LLMs, this study considers Gluco-LLM [14], an approach 9

Table 3: Number of prediction instances from the OhioT1DM dataset, reported for the training and test splits for each task and prediction horizon.

ferent use of language models from the prompt-based approach considered in the previous sections. Rather than expressing physiological observations as textual inputs and querying an unchanged general purpose LLM, Gluco-LLM introduces trainable components that map numerical time series information into the representation space of a pretrained language-model backbone. The historical CGM sequence is first divided into patches and mapped into an embedding space through a trainable patchembedding module. A trainable mechanism subsequently aligns these time series representations with the embedding space of a pretrained GPT-2 backbone. The parameters of the language-model backbone remain frozen, whereas the adaptation components surrounding it are optimized for the target forecasting task. Finally, a trainable output projection maps the resulting hidden representation to the required numerical glucose prediction. Gluco-LLM therefore provides an alternative approach between conventional task-specific forecasting models and prompt-based general purpose LLMs: it exploits a pretrained language-model backbone while learning dedicated interfaces for physiological time series prediction. In the present study, Gluco-LLM is adapted to the postprandial prediction formulation introduced in Section 3.1. Rather than forecasting glucose over the complete CGM trajectory, the model estimates the future glucose extreme within the considered prediction horizon, i.e., the maximum value for hyperglycemia and the minimum value for hypoglycemia. The corresponding binary event prediction is subsequently obtained by applying the same clinical thresholds adopted throughout this study. Its specific training configuration and the protocol adopted for comparison with the supervised regression baseline are described in the experimental setup.

Task Hyperglycemia

Hypoglycemia

Horizon 30 min 60 min 90 min 30 min 60 min 90 min

Training samples 16 384 16 194 16 011 25 106 16 984 8 620

Test samples 3 211 3 198 3 154 5 560 3 739 1 896

patients. CGM measurements are recorded at five-minute intervals and are accompanied by additional patient information, including insulin administration, meal-related carbohydrate intake, fingerstick glucose measurements, and physical activity. Throughout this work, the historical observation window spans 60 minutes, resulting in T = 12 measurements. The data used in this study are summarized in Table 3. The prediction windows exhibit a marked class imbalance, reflecting the uneven occurrence of postprandial glycemic events. This imbalance is substantially more pronounced for hypoglycemia, whose positive instances remain rare across prediction horizons, whereas hyperglycemic events occur more frequently. Moreover, class prevalence varies with the prediction horizon, as longer horizons increase the probability of observing a threshold-crossing event. This characteristic is taken into account both during model training and in the selection of the evaluation metrics. Data are processed independently for each patient. CGM measurements are ordered chronologically and represented on their nominal five-minute sampling grid following the protocol used in [25]. Short missing intervals of up to 15 minutes are handled through forward filling, thereby relying exclusively on previously observed glucose values. Longer gaps are left missing, and any observation window containing such gaps is discarded. For input variables requiring normalization, scaling parameters are estimated independently from the training data of each patient and subsequently applied unchanged to validation and test observations. No statistics computed from the validation or test sets are therefore used during preprocessing of the training data [34]. Postprandial samples are subsequently extracted according to the meal-centered protocol described in Section 3.1. Each prediction instance contains the previous 60 minutes of CGM history, corresponding to 12 measurements. Cui et al. [25] originally considered a 30 minute prediction horizon. In the present study, we retain their meal-centered extraction principle while extending the prediction horizon to h ∈ {30, 60, 90} minutes. For each horizon, prediction instances are generated only when the complete future interval required to define the target is available within the corresponding task-specific postprandial region. Consequently, the latest admissible prediction time depends on h, resulting in a decreasing number of valid prediction instances as the prediction horizon increases.

4. Experimental assessment The experimental assessment is designed to evaluate the three research questions introduced in Section 1 under a common and controlled evaluation framework. All prediction approaches are evaluated on the same postprandial glucoseprediction setting, while the specific model configurations and comparison protocols are adapted to the objective of each research question. The following sections first describe the dataset, preprocessing pipeline, input configurations, and evaluation criteria shared across the experiments. The experimental protocols adopted to address each research question are then presented separately. 4.1. Dataset and preprocessing Experiments are conducted on the OhioT1DM dataset [19], a publicly available benchmark for glucose prediction in individuals with T1DM. The dataset contains physiological data from 12 patients monitored over several weeks. Following the postprandial protocol of Cui et al. [25], subject 567 is excluded because no meal information is available in the official test partition. The final experimental cohort therefore consists of 11

4.2. Compared models The experimental assessment involves models belonging to the three prediction approaches introduced in Section 3: con10

Table 4: Hyperparameter search spaces considered for the supervised classification models.

ventional supervised learning, prompt-based inference with general-purpose LLMs, and time series adaptation of a pretrained language model. Since the considered research questions address different prediction formulations, the role of the supervised reference models differs across the experimental protocols.

Model

CNN1D

Supervised classification baselines. For direct glycemic-event classification experiments, four supervised models are considered: Logistic Regression, XGBoost, a one-dimensional Convolutional Neural Network (CNN1D), and a Long Short-Term Memory network (LSTM) [35]. The selected models cover complementary inductive biases, including a linear classifier, a nonlinear tree-based ensemble, and two neural architectures capable of modeling temporal patterns in glucose observations. To account for class imbalance during the training of the neural models, we adopt a weighted binary cross-entropy loss [36]. The weight is estimated exclusively from the training partition for each patient, prediction task, and prediction horizon, thereby preventing information from the validation or test sets from influencing model fitting. This weighting increases the penalty assigned to misclassified positive instances and counteracts the tendency to favor the majority negative class, which is particularly relevant for hypoglycemia prediction. All four models are trained separately for each patient and evaluated on the same patient-specific test samples used for the corresponding prompt-based LLM experiments. Depending on the considered input configuration, they receive either the original CGM sequence or the engineered descriptors introduced in Section 3.

LSTM

XGBoost Logistic Regression

Hyperparameter Number of filters Kernel size Dropout Learning rate Batch size Maximum epochs Hidden size Number of layers Dropout Learning rate Batch size Maximum epochs Number of estimators Maximum depth Learning rate Maximum iterations

Search space {16, 32, 64, 128} {2, 3, 5, 6} {0.2, 0.5} {0.001, 0.005, 0.01, 0.05, 0.1} {64} {50} {32, 64, 128} {2, 4, 5, 7} {0.2, 0.5} {0.001, 0.005, 0.01, 0.05, 0.1} {64} {50} {100, 200} {4, 6} {0.005, 0.01, 0.05} {500, 1000}

therefore performed exclusively through the ZS and FS inference procedures described in Section 3.1.2. To ensure a consistent inference setting, the context window is limited to 4096 tokens for all considered LLMs. Time series-adapted LLM. Gluco-LLM represents the time series-adapted LLM paradigm. As described in Section 3.2, the model combines a frozen GPT-2 backbone with trainable adaptation components specifically designed to map physiological time series data into the language-model representation space. Gluco-LLM is evaluated exclusively within the regression formulation and compared directly with the Cui et al. LSTM under an aligned postprandial prediction protocol. This subsection describes the model selection criteria and implementation choices adopted for the evaluated predictive approaches. We first detail the conventional supervised models and then present the implementation of the general purpose and time series-adapted LLMs, including their training or inference configurations where applicable.

Regression reference model. For the regression-based experiments, the supervised reference is the LSTM model following the postprandial prediction setting proposed by Cui et al. [25]. Unlike the direct classification baselines, this model predicts a continuous glucose target, namely the maximum future glucose value for hyperglycemia and the minimum future glucose value for hypoglycemia. The corresponding binary event prediction is subsequently obtained by applying the clinical thresholds defined in Section 3.1. The Cui et al. [25] model is employed as the supervised reference model both for the comparison with prompt-based LLM regression and for the evaluation of Gluco-LLM. This choice provides a common reference across the two regression-based experiments and enables the different LLM paradigms to be compared against the same task-specific forecasting approach.

Supervised classification models. For the direct-classification experiment, model selection is performed independently of the test set. For each patient, the last 15% of the chronological training partition is used as a validation set, preserving temporal order. Hyperparameters are selected through grid search [37] by maximizing the Matthews Correlation Coefficient (MCC, [38]) on the validation data. The grid search hyperparameter space is summarized in Table 4 To account for variability due to stochastic optimization and initialization, the neural and stochastic supervised models are evaluated over multiple random seeds, and results are reported as mean and standard deviation where applicable.

General purpose LLMs. Prompt-based experiments are conducted using four open-weight instruction-tuned language models from different model families: Llama 3.1 8B, Mistral 7B, Qwen 2.5 7B, and Gemma 2 9B. Considering multiple model families reduces the risk of attributing model-specific behavior to the prompt-based paradigm as a whole. All models are executed locally through Ollama1 and are used without task-specific parameter updates. Prediction is

Prompt-based LLMs. Prompt-based LLMs require neither training nor hyperparameter tuning. All four models are executed locally through Ollama under a common inference setup. For all models, the generation temperature is set to 0, the maximum number of generated tokens is limited to 20, and the

1 https://ollama.com/

11

context-window size is fixed to 4096 tokens. The same promptgeneration, output-parsing, and ZS/FS procedures are applied consistently across model families. In the FS condition, For each target prediction, K = 6 examples (Whenever possible, three positive and three negative) are selected exclusively from the corresponding patient-specific training partition, thereby preventing test information from entering the prompt.

glycemic event prediction. The primary comparison is conducted under the direct binary classification formulation introduced in Section 3.1. The four supervised classification baselines, namely Logistic Regression, XGBoost, CNN1D, and LSTM, are compared with the four general purpose LLMs described in Section 4.2. The supervised baselines are trained separately for each patient. Their reported performance is obtained over 10 random seeds and summarized through mean and standard deviation. Prompt-based LLMs do not undergo taskspecific training and are reported as a single prediction run for each zero-shot or few-shot experimental configuration.

Regression-based models. The regression experiments follow a different model selection protocol in order to remain aligned with the reference postprandial forecasting setting. In particular, the Cui et al. model [25] and Gluco-LLM are trained according to the optimization protocols adopted in their respective original studies.

4.4.2. RQ2: Influence of physiological information representation and availability Since RQ2 investigates how the representation and the types of information made available to the model influence promptbased prediction with general purpose LLMs, this analysis focuses exclusively on the four prompt-based LLMs and examines the sensitivity of their predictions to alternative ways of presenting the patient information. The effect of representation is first investigated by comparing the Raw, Structured, and Narrative strategies introduced in Section 3. A second analysis investigates the effect of providing different kind of information, i.e., performance under the CGMderived configuration is compared with the corresponding context CGMderived configuration in Structured and Narrative representations. The former contains only information derived from the CGM observations, whereas the latter additionally incorporates the clinical and therapeutic context described in Section 3. Comparing these configurations therefore isolates the effect of enriching the model input with information beyond glucose measurements.

4.3. Evaluation metrics Given the class imbalance characterizing the prediction problems, the MCC is adopted as the primary metric for model comparison, as it accounts for all four entries of the confusion matrix and provides a balanced assessment even when the class distribution is highly uneven. MCC ranges from −1 to +1. Positive values, particularly those approaching +1, indicate increasingly good predictive performance, whereas values close to or below zero indicate poor performance. For the regression formulation, prediction performance is evaluated using Root Mean Squared Error (RMSE) [35] between the predicted and actual future glucose extreme. For hyperglycemia, the target corresponds to the maximum glucose value within the prediction horizon, whereas for hypoglycemia it corresponds to the minimum. Lower RMSE therefore indicates more accurate prediction of the future glucose extreme. Additionally, the predicted continuous value is converted into a binary glycemic-event prediction using the same thresholds adopted for direct classification. The resulting decisions are evaluated using the MCC. RMSE and MCC therefore capture two complementary aspects of performance: numerical accuracy in predicting the future glucose extreme and the ability to correctly identify clinically relevant events. For stochastic supervised and time series-adapted models evaluated over multiple runs, results are reported as mean and standard deviation across runs, consistently with several previous glucose-prediction studies [22, 25, 24]. Prompt-based general-purpose LLMs are instead evaluated using a single inference run for each experimental configuration and are therefore reported without an across-run standard deviation.

4.4.3. RQ3: Prompt-based inference vs. time series-adapted LLMs RQ3 investigates whether prompt-based inference with general purpose LLMs can approach the predictive performance of an LLM architecture explicitly adapted to physiological time series forecasting. To establish a meaningful comparison between these two paradigms, the analysis is conducted under the regression formulation introduced in Section 3.1, using the model proposed by Cui et al. [25] as a common supervised reference. The first stage compares the four general purpose LLMs with the Cui et al. model. We highlight that the primary comparison for RQ3 concerns the relative performance of the two LLM paradigms, while the Cui et al. model provides a common task-specific reference. Both approaches are required to estimate the future glucose extreme, namely the maximum glucose value for hyperglycemia and the minimum glucose value for hypoglycemia, at the different investigated prediction horizons. To preserve input comparability, prompt-based LLMs are evaluated exclusively using the Raw representation of the CGM sequence, without engineered descriptors or additional clinical information. Each LLM is evaluated under both XS and FS inference.

4.4. Experimental protocols In this section, the experimental protocols adopted to address each research question are described. 4.4.1. RQ1: General purpose LLMs vs. supervised learning RQ1 investigates whether prompt-based inference with general purpose LLMs can provide predictive performance comparable to conventional supervised ML models for postprandial 12

The Cui et al. model is trained separately for each patient on the complete patient-specific training partition following the training protocol adopted in the reference setting. Its results are reported as mean and standard deviation over 10 initialization seeds, whereas each prompt-based LLM configuration produces a single inference result. The second stage evaluates Gluco-LLM against the same Cui et al. model as reference. For this comparison, the Gluco-LLM data-loading procedure is explicitly aligned with that of the Cui et al. model so that both models operate on the same postprandial samples, observation windows, regression targets, and chronological data partitions. Performance is evaluated both on the continuous glucose prediction, through RMSE, and on the corresponding threshold-derived event prediction, through MCC.

The advantage of prompt-based inference is observed across all prediction horizons and is not restricted to the inclusion of incontext demonstrations. One possible contributing factor is the severe class imbalance characterizing the hypoglycemia task, for which the FS setting may partially mitigate the effect of the imbalance by explicitly exposing the model to labeled examples from both classes. This interpretation, however, cannot fully explain the observed advantage, since competitive performance is also obtained under ZS inference. More generally, these results suggest that prompt-based LLMs may be comparatively less affected by the scarcity of positive hypoglycemic examples than conventional supervised models, although further investigation would be required to isolate the mechanisms underlying this behavior. 5.2. RQ2: Influence of physiological information representation and availability

5. Results Tables 5-7 provide a comprehensive overview of the directclassification results. Table 5 reports the MCC achieved by the supervised models across prediction tasks, horizons, and information settings, whereas Tables 6 and 7 report the corresponding results for the general purpose LLMs under ZS and FS inference, respectively, considering the different physiological information representations and information settings. Based on these complete results, the following subsections provide targeted analyses for the three research questions. We first compare the predictive performance of general purpose LLMs with conventional supervised models (RQ1), then examine how physiological information representation and availability affect prompt-based prediction (RQ2), and finally investigate the regression-based comparison between prompt-based inference and the time series-adapted Gluco-LLM approach (RQ3).

RQ2 investigates how the representation and the amount of physiological information made available to the model affect prompt-based prediction. Unlike RQ1, which focuses on the relative predictive performance of the considered paradigms, this analysis examines the behavior of the general purpose LLMs across the different input configurations. The zero-shot results (Table 6) do not reveal a general preference for the original CGM trajectory. Indeed, when Raw is compared with Structured and Narrative under CGMderived (thereby keeping the underlying physiological source limited to CGM), Structured achieves the highest MCC in the greatest part of the cases in hyperglycemia, while in hypoglycemia the results are more heterogeneous across models and horizons. Under FS inference (Table 7), the relative advantage shifts toward the Raw representation, corresponding to the use of CGMonly information, which achieves the highest MCC in most model–horizon configurations. Nevertheless, representations context information settings based on the CGMderived and CGMderived yield the best overall few-shot performance for hyperglycemia at 30 minutes and for hypoglycemia at all three prediction horizons. This indicates that Structured or Narrative representations can still produce the strongest overall result when considering the best-performing model for each task and horizon. A second aspect of RQ2 concerns whether providing additional physiological information improves prediction. The context comparison between CGMderived and CGMderived does not reveal a systematic advantage for either setting. Under ZS incontext ference, the CGMderived setting is more frequently preferred for Structured inputs, whereas the CGMderived setting clearly performs better for Narrative inputs. Consequently, the inclusion of insulin, carbohydrate, meal, basal-rate, and exercise information does not consistently improve the predictive performance of the evaluated LLMs. Importantly, this result should not be interpreted as evidence that these variables are predictively irrelevant. Rather, it indicates that the models do not systematically benefit from their inclusion when the additional information is provided through the considered prompt representations.

5.1. RQ1: General purpose LLMs vs. supervised learning RQ1 investigates whether prompt-based general purpose LLMs can achieve predictive performance comparable to conventional supervised models for postprandial glycemic-event prediction. Building on the complete results reported in Tables 5-7, while Table 8 provides a descriptive best-case comparison by reporting, for each prediction task and horizon, the highest test MCC observed among the evaluated supervised, ZS LLM, and FS LLM configurations. For each approach, the best result is selected across all applicable models and information configurations. The comparison reveals a clear task-dependent behavior. For hyperglycemia, the best supervised configuration achieves the highest MCC at all three prediction horizons. Although the gap between the best supervised and FS configurations progressively narrows as the prediction horizon increases, the best supervised configuration retains the highest performance at every horizon. A substantially different pattern emerges for hypoglycemia. In this task, the best prompt-based LLM configuration achieves a higher observed MCC than the best supervised configuration at every prediction horizon, including under ZS inference. 13

Table 5: MCC obtained by the supervised models across prediction tasks, input configurations, and prediction horizons. Results for stochastic models are reported as mean (standard deviation) over 10 runs. Bold values indicate the best supervised result for each task, input configuration, and prediction horizon.

Task

Input CGMonly

Hyperglycemia

CGMderived

context CGMderived

CGMonly

Hypoglycemia

CGMderived

context CGMderived

Horizon 30 min 60 min 90 min 30 min 60 min 90 min 30 min 60 min 90 min 30 min 60 min 90 min 30 min 60 min 90 min 30 min 60 min 90 min

LogReg 0.558 0.459 0.389 0.570 0.456 0.366 0.560 0.523 0.403 0.331 0.328 0.312 0.375 0.345 0.351 0.296 0.280 0.291

XGBoost 0.511 (0.003) 0.382 (0.004) 0.301 (0.004) 0.540 (0.005) 0.406 (0.005) 0.318 (0.002) 0.565 (0.005) 0.497 (0.004) 0.455 (0.004) 0.350 (0.006) 0.305 (0.008) 0.290 (0.002) 0.383 (0.003) 0.330 (0.001) 0.366 (0.002) 0.311 (0.008) 0.308 (0.008) 0.321 (0.003)

CNN1D 0.454 (0.009) 0.296 (0.038) 0.236 (0.036) 0.475 (0.013) 0.377 (0.013) 0.271 (0.047) 0.431 (0.011) 0.309 (0.074) 0.274 (0.088) 0.231 (0.022) 0.202 (0.023) 0.179 (0.015) 0.270 (0.085) 0.316 (0.039) 0.283 (0.056) 0.225 (0.092) 0.302 (0.028) 0.257 (0.051)

LSTM 0.505 (0.028) 0.391 (0.018) 0.308 (0.024) 0.411 (0.074) 0.365 (0.028) 0.253 (0.060) 0.364 (0.020) 0.329 (0.076) 0.313 (0.066) 0.230 (0.050) 0.233 (0.016) 0.262 (0.050) 0.210 (0.031) 0.232 (0.033) 0.229 (0.031) 0.129 (0.023) 0.186 (0.043) 0.190 (0.022)

backbone than prompt-based inference without task-specific parameter adaptation. This advantage is not universal, however: prompt-based inference obtains the highest LLM MCC for hypoglycemia at 90 minutes and marginally lower RMSE in two of the six conditions. Overall, RQ3 shows that none of the investigated LLM approaches consistently achieves the performance of the taskspecific model. Indeed, the Cui et al. model currently remains the strongest and most consistent approach for the postprandial regression tasks considered in this study.

5.3. RQ3: Prompt-based inference vs. time-series-adapted LLMs RQ3 investigates whether prompt-based inference with general purpose LLMs can approach the performance of a LLM explicitly adapted to physiological time series forecasting. The comparison is conducted under the regression formulation, using the Cui et al. model as a common supervised reference. Tab. 9 summarizes the results in terms of MCC and RMSE, the two metrics that are directly available across all three approaches. The comparison with the prompt-based LLMs shows a clear advantage for the task-specific model: Cui et al. model achieves the highest MCC in 4 of the 6 task–horizon combinations and the lowest RMSE in all six. The gap is relatively small for shorthorizon hyperglycemia, but increases substantially at longer horizons. At 90 minutes, for example, the corresponding MCC values are 0.234 and 0.436, respectively. The only case in which a prompt-based LLM exceeds the Cui et al. baseline in MCC is hypoglycemia at 90 minutes; however, its continuous glucose prediction remains less accurate in terms of RMSE. Considerng instead Gluco-LLM, its MCC is closer to that of Cui et al. in most experimental conditions, particularly for hypoglycemia. However, this difference is small relative to the variability observed across different runs. A more consistent difference emerges for continuous glucose prediction. Cui et al. achieves a lower RMSE than Gluco-LLM in every taskhorizon combination. Thus, the competitive event-detection performance occasionally observed for Gluco-LLM does not translate into more accurate prediction of the underlying glucose extreme. Within the regression-based comparison reported in Tab. 9, Gluco-LLM remains closer to the Cui et al. model than the best observed prompt-based LLM configuration in most conditions, suggesting that explicit adaptation to physiological time series may enable a more effective use of a pretrained language-model

6. Discussion In this work, we have provided a detailed investigation of the capabilities and limitations of language model-based approaches for postprandial glycemic-event prediction. Across the three research questions, the results reported in Section 5 show that general purpose LLMs can produce non-trivial predictions without task-specific parameter optimization. However, their performance strongly depends on the model, physiological representation, prompting strategy, target event, and prediction horizon. General purpose LLMs can be competitive in some settings. The results for RQ1 reported in Tables 5- 8 show that promptbased LLMs cannot yet be regarded as a general replacement for conventional supervised models. For hyperglycemia, the best observed supervised configuration achieves the highest MCC at every prediction horizon. Conversely, for hypoglycemia, the best observed prompt-based configuration outperforms the best observed supervised result at all three horizons, including under ZS inference. Furthermore, the effectiveness of prompt-based inference is strongly configuration-dependent. Performance varies across 14

Table 6: MCC of the general-purpose LLMs under zero-shot inference. For Structured and Narrative representations, values are reported as CGMderived / context . Raw corresponds to the CGM configuration. Bold values indicate the best input configuration for each LLM, prediction task, and horizon. A CGMderived dash denotes an undefined MCC due to a constant prediction output.

Task

Horizon 30 min

Hyperglycemia

60 min

90 min

30 min

Hypoglycemia

60 min

90 min

Representation Raw Structured Narrative Raw Structured Narrative Raw Structured Narrative Raw Structured Narrative Raw Structured Narrative Raw Structured Narrative

Gemma 2 0.298 0.376 / 0.392 0.204 / 0.202 0.299 0.357 / 0.369 0.168 / 0.162 0.271 0.331 / 0.338 0.139 / 0.126 0.341 0.458 / 0.507 0.132 / 0.140 0.373 0.456 / 0.403 0.183 / 0.195 0.386 0.436 / 0.406 0.240 / 0.247

LLM families, prompting strategies, physiological representations, target events, and prediction horizons. The best observed FS configuration outperforms the best observed ZS configuration in the greatest part of task–horizon combinations. Similarly, no representation is uniformly optimal: Structured produces the best overall result in most task–prompting–horizon combinations, whereas Raw and Narrative remain preferable in specific settings. This task-dependent behavior indicates that general purpose LLMs can extract predictive information from short glucose histories and derived physiological descriptors, even though numerical time series modeling is not their native training objective. However, LLMs cannot yet be regarded as a consistently superior replacement for models trained specifically for glycemic-event prediction. Therefore, the comparison should not be framed solely in terms of which paradigm “wins”. Rather, it reveals a trade-off between the overall reliability and consistency of supervised models and the training-free adaptability, accessibility, and representation flexibility offered by prompt-based LLM inference.

Llama 3.1 0.254 0.357 / 0.346 0.227 / 0.200 0.244 0.329 / 0.313 0.205 / 0.182 0.206 0.290 / 0.290 0.172 / 0.140 0.088 –/– 0.270 / 0.271 0.092 –/– 0.336 / 0.318 0.073 –/– 0.369 / 0.357

Mistral 0.335 0.270 / 0.408 0.238 / 0.214 0.310 0.256 / 0.360 0.200 / 0.184 0.284 0.189 / 0.265 0.151 / 0.146 0.274 – / -0.002 0.198 / 0.137 0.270 – / -0.004 0.283 / 0.200 0.337 – / -0.005 0.315 / 0.264

Qwen 2.5 0.367 0.372 / 0.485 0.314 / 0.247 0.305 0.398 / 0.407 0.280 / 0.204 0.240 0.331 / 0.375 0.245 / 0.194 0.421 -0.002 / -0.004 0.377 / 0.314 0.429 -0.003 / -0.008 0.345 / 0.282 0.452 -0.005 / -0.011 0.376 / 0.322

ables are represented and integrated into the prompt rather than simply on their availability. FS prompting also exhibits a configuration-dependent behavior. Although it improves the best observed result in several cases, it increases MCC in only a limited number of cases of the matched comparisons. In-context examples should therefore not be regarded as uniformly beneficial; their effectiveness depends on the model, representation, information setting, task, and prediction horizon. Taken together, these findings show that representation design is not a secondary implementation choice when LLMs are applied to physiological time series. Neither preserving the original signal nor transforming it into structured or narrative descriptors provides a universal advantage. Instead, the most effective interface depends on the interaction among the input representation, the available physiological information, and the specific LLM. Explicit time-series adaptation narrows, but does not close, the gap. RQ3 examines whether the limitations observed for prompt-based inference can be mitigated by explicitly adapting a pretrained language-model backbone to physiological timeseries forecasting. Gluco-LLM generally provides a stronger competitor to the Cui et al. model than general purpose models, particularly for hypoglycemia, indicating that learned time series interfaces can exploit a language model backbone more effectively than textual prompting alone. Nevertheless, the Cui et al. model remains the most reliable model in the regression experiments. It achieves the lowest RMSE in every considered task–horizon combination and the strongest MCC in almost all comparisons with Gluco-LLM. The isolated MCC advantage obtained by Gluco-LLM for hypoglycemia at 60 minutes is small relative to the variability across runs and therefore does not indicate a systematic superi-

How physiological information is represented could matter at least as much as the kind of information provided. The results do not identify a representation of the information that is uniformly preferable across all experimental conditions. As reported in Section 5, descriptor-based representations can produce the strongest overall configurations, even though preserving the original CGM sequence remains advantageous for several individual models. context The comparison between CGMderived and CGMderived further shows that increasing the amount of physiologically relevant information is not sufficient to ensure better prompt-based inference. Consequently, the effect of insulin, carbohydrate, basalrate, meal, and exercise information depends on how these vari15

context . Table 7: MCC of the general-purpose LLMs under few-shot inference. For Structured and Narrative representations, values are reported as CGMderived / CGMderived Raw corresponds to the CGMonly configuration. Bold values indicate the best input configuration for each LLM, prediction task, and horizon. A dash denotes an undefined MCC due to a constant prediction output.

Task

Horizon 30 min

Hyperglycemia

60 min

90 min

30 min

Hypoglycemia

60 min

90 min

Representation Raw Structured Narrative Raw Structured Narrative Raw Structured Narrative Raw Structured Narrative Raw Structured Narrative Raw Structured Narrative

Gemma 2 0.399 0.370 / 0.345 0.265 / 0.273 0.415 0.193 / 0.235 0.190 / 0.185 0.388 0.285 / 0.302 0.145 / 0.144 0.360 0.165 / 0.176 0.149 / 0.154 0.319 0.253 / 0.240 0.201 / 0.199 0.283 0.230 / 0.189 0.272 / 0.271

Llama 3.1 0.291 0.077 / 0.099 0.324 / 0.328 0.244 0.111 / 0.109 0.258 / 0.245 0.235 0.091 / 0.090 0.190 / 0.198 0.419 0.289 / 0.089 0.111 / 0.118 0.437 0.061 / 0.212 0.222 / 0.150 0.113 0.219 / 0.087 0.219 / 0.154

Mistral 0.441 0.395 / 0.319 0.257 / 0.263 0.360 0.227 / 0.245 0.238 / 0.200 0.280 0.273 / 0.171 0.170 / 0.187 0.165 0.453 / 0.349 0.170 / 0.138 0.345 0.390 / 0.375 0.445 / 0.277 0.220 0.397 / 0.499 0.339 / 0.276

Qwen 2.5 – 0.145 / 0.487 0.361 / 0.312 0.139 0.045 / 0.263 0.116 / 0.221 0.222 0.324 / 0.353 0.178 / 0.244 0.296 0.305 / 0.568 0.338 / 0.340 0.312 0.302 / 0.417 0.314 / 0.347 0.359 0.113 / 0.286 0.403 / 0.402

Table 8: Best observed test MCC among the evaluated supervised, zero-shot LLM, and few-shot LLM configurations for each prediction task and horizon. The best supervised result is selected across models and information settings, whereas the best LLM results are selected across models, representations, and information settings. For stochastic supervised models, values are reported as mean (standard deviation) over 10 runs. Bold values indicate the best overall MCC for each task and prediction horizon.

Task

Horizon Best supervised MCC Best zero-shot LLM MCC Best few-shot LLM MCC context context 30 min LogReg (CGMderived ) 0.570 Qwen (Structured/CGMderived ) 0.485 Qwen (Structured/CGMderived ) 0.487 context context Hyperglycemia 60 min LogReg (CGMderived ) 0.523 Qwen (Structured/CGMderived ) 0.407 Gemma (Raw/CGMonly ) 0.415 context context 90 min XGBoost (CGMderived ) 0.455 (0.004) Qwen (Structured/CGMderived ) 0.375 Gemma (Raw/CGMonly ) 0.388 context context 30 min XGBoost (CGMderived ) 0.383 (0.003) Gemma (Structured/CGMderived ) 0.507 Qwen (Structured/CGMderived ) 0.568 Hypoglycemia 60 min LogReg (CGMderived ) 0.345 Gemma (Structured/CGMderived ) 0.456 Mistral (Narrative/CGMderived ) 0.445 context 90 min XGBoost (CGMderived ) 0.366 (0.002) Qwen (Raw/CGM) 0.452 Mistral (Structured/CGMderived ) 0.499

ority of the adapted LLM. This is consistent with the observation that event detection after thresholding does not necessarily imply an equally accurate estimate of the underlying glucose value. The result is also informative with respect to the applicability of language-model-based time series architectures. GlucoLLM was originally developed for continuous glucose forecasting, whereas the present study considers a distinct prediction setting based on short, meal-centered observation windows and future glucose-extreme estimation. This difference in task formulation may partly explain why its behavior in the present postprandial setting differs from that observed in conventional glucose forecasting [14].

evaluation under a controlled chronological protocol, conclusions drawn from a single dataset cannot establish the generality of the observed behavior. Differences in population characteristics, sensing devices, treatment strategies, and data-collection protocols may affect both supervised models and LLM-based approaches. Future work should therefore reproduce the proposed evaluation framework on additional CGM datasets and, where possible, assess cross-dataset robustness. A related limitation concerns the training data of the adopted LLMs. Because the complete pretraining corpora are not publicly available, it is not possible to exclude with certainty that the evaluated models were exposed to OhioT1DM data, derived material, or closely related distributions during pretraining [34]. Such contamination cannot be quantified within the present study. Future work should investigate strategies for evaluating pretraining contamination and memorization in biomedical timeseries benchmarks, particularly when general purpose foundation models are compared with models trained exclusively on the experimental training partition. Second, the prompt-based experiments consider open-weight language models in the 7-9B parameter range. This choice favors local execution, reproducibility, and experimental control, but it does not characterize the behavior of larger open-weight

7. Limitations and future work The findings of this study should be interpreted in light of several limitations, which also identify relevant directions for future research. First, the experiments rely exclusively on the OhioT1DM dataset. Although this dataset provides a well-established benchmark for glucose prediction and enables patient-specific 16

Table 9: Regression-based comparison among the Cui et al. model, prompt-based general-purpose LLMs under zero-shot and few-shot inference, and Gluco-LLM. For the prompt-based approaches, the best observed test result across the four evaluated LLMs is reported separately for zero-shot and few-shot inference. Cui and Gluco-LLM results are reported as mean (standard deviation) over 10 seeds. Bold values indicate the best overall result for each task, prediction horizon, and metric. Lower RMSE is better.

Task Hyperglycemia

Hypoglycemia

Horizon 30 min 60 min 90 min 30 min 60 min 90 min

Cui 0.618 (0.012) 0.492 (0.008) 0.436 (0.013) 0.470 (0.028) 0.380 (0.043) 0.299 (0.056)

MCC ZS LLM FS LLM 0.456 0.560 0.343 0.353 0.234 0.231 0.265 0.409 0.186 0.240 0.222 0.378

models or proprietary frontier models. The conclusions should therefore be restricted to the class of models investigated here. Evaluating a broader range of model scales and architectures would help determine whether the observed limitations are intrinsic to prompt-based physiological reasoning or partly attributable to model capacity. Third, the variability of prompt-based inference is less extensively characterized than that of the trained models. Supervised stochastic models, the Cui et al. model, and GlucoLLM are evaluated over multiple random runs, where applicable, to account for variability introduced by stochastic model training and initialization. In contrast, prompt-based LLM configurations do not involve task-specific training or parameter updates; therefore, each configuration is evaluated through a single inference run under a fixed prompting setup. Consequently, the reported differences do not quantify sensitivity to alternative demonstration sets or other sources of inferencetime variability. Future studies should repeat prompt-based experiments across multiple decoding configurations, including different temperature settings, and report the corresponding variability and uncertainty estimates. Finally, the investigated representations cover only a limited subset of the ways in which multimodal physiological information can be presented to an LLM. In particular, the lack of a systematic benefit from the additional contextual variables should not be interpreted as evidence that insulin, carbohydrate, meal, or physical-activity information is intrinsically uninformative. Rather, this finding is specific to the way such information is represented and incorporated in the present study. Future work could investigate representations that retain the temporal evolution of insulin, meal, carbohydrate, and physical-activity information, rather than summarizing these variables into aggregate descriptors.

Gluco 0.577 (0.007) 0.471 (0.014) 0.379 (0.015) 0.460 (0.044) 0.397 (0.030) 0.285 (0.040)

Cui 18.45 (0.21) 30.53 (0.22) 37.64 (0.53) 13.75 (0.23) 23.94 (0.29) 31.51 (0.46)

RMSE (mg/dL) ZS LLM FS LLM 23.64 21.53 40.77 37.67 52.72 46.76 17.75 15.45 27.44 27.62 35.50 34.11

Gluco 20.79 (0.26) 35.38 (0.22) 46.80 (0.33) 15.08 (0.45) 26.12 (0.31) 35.26 (0.26)

With respect to RQ1, the comparison did not establish a uniform ranking between supervised ML and LLM-based inference, indicating that general purpose LLMs are not intrinsically unsuitable for glucose prediction, although their effectiveness is strongly dependent on the target event and experimental configuration. From the observed results, few-shot prompting should be regarded as a configuration-dependent adaptation mechanism rather than as a uniformly beneficial strategy. RQ2 showed that the representation of physiological information has a substantial effect on prompt-based prediction, but did not identify a universally optimal representation. When the best result was selected across models, Structured inputs achieved the highest MCC in most task–prompting–horizon combinations. Conversely, Raw was more frequently preferred across matched model configurations under FS inference, while Narrative became competitive in specific hypoglycemia settings. The effectiveness of a representation therefore depends on the interaction among the LLM family, prompting strategy, target event, and prediction horizon. Similarly, expanding the input from CGMderived to context CGMderived did not produce a systematic improvement. The additional insulin, carbohydrate, meal, basal-rate, and exercise information was beneficial in some configurations but not in others, and its effect varied with the adopted representation. This does not imply that these variables are physiologically or predictively irrelevant. Rather, it suggests that making additional information available is not sufficient unless the model can effectively integrate it through the chosen input interface. This aspect is particularly relevant when considering the potential use of LLMs for hypoglycemic or hyperglycemic event prediction in patient-oriented applications, where their ability to process heterogeneous information through a natural language interface could help reduce some of the theoretical and practical barriers associated with standard ML approaches. In this context, however, the effective representation and integration of clinically relevant information remains a key requirement for translating such flexibility into reliable predictive performance.

8. Conclusion This study investigated the use of language models for postprandial glycemic-event prediction in individuals with T1DM, comparing three complementary paradigms: conventional supervised ML, prompt-based inference with general-purpose LLMs, and time series adaptation of a pretrained LLM. The analysis considered both direct event classification and regression of future glucose extrema under a common patient-specific and temporally causal evaluation framework.

Finally, RQ3 showed that explicitly adapting a language model backbone to physiological time series narrows the gap with task-specific forecasting models but does not eliminate it. In particular, Gluco-LLM was generally more competitive with the Cui et al. model than prompt-only inference, partic17

ularly for hypoglycemia. Nevertheless, the recurrent baseline achieved a lower RMSE in every evaluated condition and the highest MCC in most comparisons. Thus, explicit time series adaptation appears more effective than prompt-only inference for exploiting a pretrained language model backbone, but the task-specific recurrent model remains the strongest and most consistent approach for the considered regression task. Taken together, these results suggest that a central challenge for LLM-based glucose prediction lies in effectively representing and modeling patient-specific temporal dynamics. Specialized models trained directly on physiological data remain the most reliable choice in several experimental conditions, particularly for hyperglycemia classification and continuous glucose prediction. However, their advantage is not universal: appropriately selected general purpose LLM configurations achieve competitive or superior event-detection performance for hypoglycemia without task-specific parameter training. The resulting picture is therefore not one of uniform superiority, but of a trade-off between the consistency of task-specific models and the training-free adaptability and representational flexibility of prompt-based inference. Future work should extend the evaluation to additional glucose-monitoring datasets, larger and more diverse language models, repeated prompt-based inference protocols, and adaptation strategies capable of preserving the temporal relationships among heterogeneous physiological and therapeutic signals. Dedicated ablation studies should also examine the effect of the number, selection, and class distribution of in-context demonstrations, particularly for severely imbalanced prediction tasks.

[4] Y. Peng, S. Yan, Z. Lu, Transfer learning in biomedical natural language processing: an evaluation of bert and elmo on ten benchmarking datasets, in: Proceedings of the 18th BioNLP workshop and shared task, 2019, pp. 58–65. [5] J. Lee, W. Yoon, S. Kim, D. Kim, S. Kim, C. H. So, J. Kang, Biobert: a pre-trained biomedical language representation model for biomedical text mining, Bioinformatics 36 (4) (2020) 1234–1240. [6] F. J. Lara-Abelenda, D. Chushig-Muzo, P. PeiroCorbacho, A. M. Wägner, C. Granja, C. Soguero-Ruiz, Personalized glucose forecasting for people with type 1 diabetes using large language models, Computer Methods and Programs in Biomedicine 265 (2025) 108737. [7] J. C. Wolber, M. E. Samadi, J. Sellin, A. Schuppert, Multimodal large language models and mechanistic modeling for glucose forecasting in type 1 diabetes patients, Journal of Biomedical Informatics (2025) 104945. [8] G. D. Ogle, F. Wang, A. Haynes, G. A. Gregory, T. W. King, K. Deng, D. Dabelea, S. James, A. J. Jenkins, X. Li, et al., Global type 1 diabetes prevalence, incidence, and mortality estimates 2025: Results from the international diabetes federation atlas, and the t1d index version 3.0, Diabetes Research and Clinical Practice 225 (2025) 112277. [9] H. Nemat, H. Khadem, J. Elliott, M. Benaissa, Datadriven blood glucose level prediction in type 1 diabetes: a comprehensive comparative analysis, Scientific reports 14 (1) (2024) 21863.

Acknowledgement

[10] B. Cinar, L. van den Boom, M. Maleshkova, Review of machine learning models in short-and long-term glucose forecasting and hypoglycemia classification, Informatics in Medicine Unlocked (2025) 101723.

The authors would like to thank Prof. Giovanni Annuzzi for his valuable clinical insights and constructive discussions, which contributed to the framing and interpretation of this work. This work was partially funded by the PNRR MUR project PE0000013-FAIR (CUP: E63C25000630006).

[11] G. Annuzzi, A. Apicella, P. Arpaia, L. Bozzetto, S. Criscuolo, E. De Benedetto, M. Pesola, R. Prevete, Exploring nutritional influence on blood glucose forecasting for type 1 diabetes using explainable ai, IEEE journal of biomedical and health informatics 28 (5) (2023) 3123– 3133.

References

[12] G. Annuzzi, A. Apicella, P. Arpaia, L. Bozzetto, S. Criscuolo, E. De Benedetto, M. Pesola, R. Prevete, E. Vallefuoco, Impact of nutritional factors in blood glucose prediction in type 1 diabetes through machine learning, IEEE Access 11 (2023) 17104–17115.

[1] M. A. Atkinson, G. S. Eisenbarth, A. W. Michels, Type 1 diabetes, The lancet 383 (9911) (2014) 69–82. [2] Y. Gu, R. Tinn, H. Cheng, M. Lucas, N. Usuyama, X. Liu, T. Naumann, J. Gao, H. Poon, Domain-specific language model pretraining for biomedical natural language processing, ACM Transactions on Computing for Healthcare (HEALTH) 3 (1) (2021) 1–23.

[13] S. Oviedo, I. Contreras, C. Quiros, M. Gimenez, I. Conget, J. Vehi, Risk-based postprandial hypoglycemia forecasting using supervised learning, International journal of medical informatics 126 (2019) 1–8.

[3] E. Alsentzer, J. Murphy, W. Boag, W.-H. Weng, D. Jindi, T. Naumann, M. McDermott, Publicly available clinical bert embeddings, in: Proceedings of the 2nd clinical natural language processing workshop, 2019, pp. 72–78.

[14] Q. Li, K. R. Amat, J. Li, Llm-powered personalized glucose prediction in type 1 diabetes, Computational and Structural Biotechnology Reports (2025) 100068. 18

[15] Y. Gao, Y. Gong, Y. Shi, Y. Guo, Llm-powered personalized glycemic assessment in type 2 diabetes with wearable sensor data, arXiv preprint arXiv:2606.12699 (2026).

[27] OpenAI, Introducing gpt-4.1 in the api, https:// openai.com/index/gpt-4-1/, accessed: 2026-08-21 (Apr. 2025).

[16] A. Mahmoudi, G. Farahani, P. Domanski, B. Farahani, F. Firouzi, K. Chakrabarty, Diabllm: an llm-based framework for blood glucose prediction in type 1 diabetes, IEEE Journal of Biomedical and Health Informatics (2026).

[28] A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, et al., The llama 3 herd of models, arXiv preprint arXiv:2407.21783 (2024).

[17] E. Healey, A. L. M. Tan, K. L. Flint, J. L. Ruiz, I. Kohane, A case study on using a large language model to analyze continuous glucose monitoring data, Scientific reports 15 (1) (2025) 1143.

[29] E. Healey, I. Kohane, Llm-cgm: A benchmark for large language model-enabled querying of continuous glucose monitoring data for conversational diabetes management, in: Biocomputing 2025: Proceedings of the Pacific Symposium, World Scientific, 2024, pp. 82–93.

[18] R. Alredaini, M. Abulkhair, H. Almisbahi, Interpretable glucose forecasting for type 2 diabetes across traditional, deep, and large language models, Scientific Reports (2025).

[30] S. He, Y. Zhang, J. Li, Personalized diabetes treatment support using large language models fine-tuned on electronic health records: development and evaluation study, JMIR Formative Research 10 (2026) e71541.

[19] C. Marling, R. Bunescu, The ohiot1dm dataset for blood glucose level prediction: Update 2020, in: CEUR workshop proceedings, Vol. 2675, 2020, p. 71.

[31] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, Lora: Low-rank adaptation of large language models, arXiv preprint arXiv:2106.09685 (2021).

[20] T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, et al., Language models are few-shot learners, Advances in neural information processing systems 33 (2020) 1877–1901.

[32] J.-E. Ding, P. N. M. Thao, W.-C. Peng, J.-Z. Wang, C.C. Chug, M.-C. Hsieh, Y.-C. Tseng, L. Chen, D. Luo, C. Wu, et al., Large language multimodal models for newonset type 2 diabetes prediction using five-year cohort electronic health records, Scientific reports 14 (1) (2024) 20774.

[21] K. Plis, R. C. Bunescu, C. Marling, J. Shubrook, F. Schwartz, A machine learning approach to predicting blood glucose levels for diabetes management., in: AAAI Workshop: Modern Artificial Intelligence for Health Analytics, Vol. 31, 2014, pp. 35–39.

[33] M. Jin, S. Wang, L. Ma, Z. Chu, J. Zhang, X. Shi, P.-Y. Chen, Y. Liang, Y.-F. Li, S. Pan, et al., Time-llm: Time series forecasting by reprogramming large language models, in: International conference on learning representations, Vol. 2024, 2024, pp. 23857–23880.

[22] J. Martinsson, A. Schliep, B. Eliasson, O. Mogren, Blood glucose prediction with variance estimation using recurrent neural networks, Journal of Healthcare Informatics Research 4 (1) (2020) 1–18.

[34] A. Apicella, F. Isgrò, R. Prevete, Don’t push the button! exploring data leakage risks in machine learning and transfer learning, Artificial Intelligence Review 58 (11) (2025) 339.

[23] M. De Bois, M. A. E. Yacoubi, M. Ammi, Glyfe: review and benchmark of personalized glucose predictive models in type 1 diabetes, Medical & Biological Engineering & Computing 60 (1) (2022) 1–17.

[35] C. M. Bishop, H. Bishop, Deep learning: Foundations and concepts, Vol. 1, Springer Cham, Switzerland, 2024.

[24] M. Voegeli, S. Laguna, H. Leutheuser, M. Pfister, M.-A. Burckhardt, J. E. Vogt, Beyond glucose-only assessment: Advancing nocturnal hypoglycemia prediction in children with type 1 diabetes, arXiv preprint arXiv:2504.09299 (2025).

[36] T. H. Phan, K. Yamamoto, Resolving class imbalance in object detection with weighted cross entropy losses, arXiv preprint arXiv:2006.01413 (2020).

[25] R. Cui, C. J. Nolan, E. Daskalaki, H. Suominen, Jointly predicting postprandial hypoglycemia and hyperglycemia using continuous glucose monitoring data in type 1 diabetes, in: 2023 45th Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC), IEEE, 2023, pp. 1–7.

[37] P. Dangeti, Statistics for machine learning: techniques for exploring supervised, unsupervised, and reinforcement learning models with Python and R, Packt Publishing Ltd, 2017. [38] D. Chicco, G. Jurman, The advantages of the matthews correlation coefficient (mcc) over f1 score and accuracy in binary classification evaluation, BMC genomics 21 (1) (2020) 6.

[26] N. Gruver, M. Finzi, S. Qiu, A. G. Wilson, Large language models are zero-shot time series forecasters, Advances in neural information processing systems 36 (2023) 19622– 19635. 19

Record · ID 668094 · SHA-256 932d39d405cc6b1f
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.