Context-Aware Optimization of Follow-Up Intervals for Type 2 Diabetes Care Using Markov Decision Processes Parisa Lotfibaghaa , Kristen Millerb, c, d , William J. Gallagherb, d , Elizabeth B. Seldene , Muge Capana
arXiv:2606.19092v1 [stat.AP] 17 Jun 2026
a
University of Massachusetts Amherst, United States, b MedStar Health Center for Diagnostic Systems Safety, United States, c National Center for Human Factors in Healthcare, MedStar Health Research Institute, United States, d Georgetown University School of Medicine, United States, e Medstar Georgetown University Hospital, United States ARTICLE HISTORY Compiled June 18, 2026 ABSTRACT Chronic disease management relies on regular patient-provider interactions to followup on disease progression and control. For Type 2 Diabetes (T2D), current guidelines prescribe fixed time intervals between subsequent primary care visits for all patients, overlooking heterogeneity in clinical trajectories and patient characteristics. This study introduces a Contextual Markov Decision Process (CMDP) model to optimize subpopulation-specific follow-up interval decisions using Electronic Health Record (EHR) data from 22,154 T2D patients across 10 primary care clinics. Contexts are identified by: i) dimensionality reduction of variables representing the individual health trajectories utilizing Principal Component Analysis, and ii) assigning patients to contexts via principal components and additional patient-level features using clustering. Two distinct contexts emerged, representing a lower- and a higher-risk subpopulation. CMDP-derived policies recommend: (i) follow-up within 1 month if lab value at current visit is unmeasured; (ii) up to 3 months for elevated lab values or recent hospitalizations; and (iii) 6 to 12 months for sustained glycemic control, with shorter follow-up intervals for patients in high-risk context. The optimal policies achieved lower expected cumulative cost than benchmarks (e.g., in the higher-comorbidity context, the CMDP policy reduced cost by about 34.8%, and in the lower-comorbidity context by about 6.4%, relative to an American Diabetes Association-like fixed interval follow-up policy. These findings demonstrate how context-aware approaches can inform adaptive follow-up strategies, and have the potential to advance chronic care management in primary care by synthesizing machine learning and probabilistic decision models. KEYWORDS Contextual Markov Decision Process, follow-up, primary care, Type 2 diabetes
1. Introduction The increasing prevalence of chronic diseases is pressuring healthcare systems to design adaptive follow-up strategies that reflect patient heterogeneity and lifelong care requirements. Type 2 Diabetes (T2D), affecting over 589 million adults globally, exemplifies this challenge (Ceriello and Colagiuri, 2025). Regular post-diagnosis follow-up CONTACT M. C. Author. Email: [email protected]
is essential to maintain disease control and prevent further T2D-induced health complications (American Diabetes Association Professional Practice Committee, 2025c). In the United States, post-diagnosis follow-up of T2D commonly occurs during visits with primary care physicians (PCPs) to achieve and maintain glycemic control (Kushner, Cavender, and Mende, 2022). Glycemic control is measured by a biomarker – hemoglobin A1c (HbA1c) – which represents the average blood sugar level over the past three months. The American Diabetes Association (ADA) guideline defines glycemic control using the target value of HbA1c ≤7% (53 mmol/mol) (American Diabetes Association Professional Practice Committee, 2025b). Without regular follow-up, majority of adult T2D patients struggle to achieve and maintain this HbA1c target (Njonnou, Nguedoung, Balti, Demanou, Lekpa, Ouankou, Boli, Wafo, Tchounchui, and Choukem, 2024; Yahaya, Doya, Morgan, Ngaiza, and Bintabara, 2023). ADA guidelines recommend quarterly PCP follow-up for T2D patients whose therapy has changed or who have not achieved the HbA1c target, and at least twice yearly for those who have achieved glycemic control (ElSayed, Aleppo, Aroda, Bannuru, Brown, Bruemmer, Collins, Hilliard, Isaacs, Johnson, et al., 2023). Finding the optimal and personalized follow-up interval between subsequent PCP visits is a complex decision problem. This decision requires a delicate balance. If the time between subsequent PCP visits is too short, this can increase the frequency of visits and overwhelm patients as well as the healthcare system. On the other hand, long follow-up intervals increase the risk of failing to detect lack of glycemic control (Yahaya et al., 2023). Several studies explored the problem of identifying a risk-appropriate follow-up interval for individuals diagnosed with T2D. For example, a randomized trial in adults with T2D who exhibit HbA1c values consistently ≤7.5% found that monitoring every three and every six months yielded comparable health outcomes (Wermeling, Gorter, Stellato, De Wit, Beulens, and Rutten, 2014). Another study found that regular followup interval of two months maintained glycemic control at rates comparable to monthly follow-up (Ukai, Ichikawa, Sekimoto, Shikata, and Takemura, 2019). For patients who failed to achieve glycemic control, a cohort study demonstrated that having more than 4 PCP visits per year increased the likelihood of achieving glycemic target (HbA1c ≤ 9.5% [80 mmol/mol]) compared to lower frequency follow-up groups (Asao, McEwen, Crosson, Waitzfelder, and Herman, 2014). The positive impact of increased followup frequency on achieving glycemic control was confirmed by another cohort study, which found that more frequent encounters were associated with faster attainment of an HbA1c target of <7% (53 mmol/mol)–by about 35% in non-insulin-treated patients and 17% in insulin-treated patients (Morrison, Shubina, and Turchin, 2011). However, these studies focus on associations between fixed visit intervals (e.g., 3 vs. 6 months) and outcomes, typically analyzed using descriptive or regression-based methods that estimate population-average effects, and lack a prescriptive framework for optimizing visit intervals considering patients’ evolving clinical trajectories. In this study, we approach finding the optimal post-diagnosis follow-up interval for T2D patients as a dynamic decision problem under uncertainty. Specifically, we introduce a Contextual Markov Decision Process (CMDP) model that uses longitudinal EHR data, and integrates dimensionality reduction and clustering to learn contextspecific follow-up policies. Markov Decision Process (MDP) models have been widely used in chronic disease screening and treatment decisions (Ayer, Alagoz, and Stout, 2012; Hajjar and Alagoz, 2023). Specifically in T2D, studies compared MDP-derived medication treatment strategies to local clinical practice. Results showed that MDP policy led to up to 0.27 additional QALYs for high-risk patients (e.g., age at diagnosis 2
of 60, smokers, or with HbA1c levels of 9% and above), whereas the improvements for low-risk patients were negligible (Meng, Sun, Heng, and Leow, 2020). Focusing on the screening of undiagnosed T2D populations aged 45 and older, Wu and Suen (2022) formulated a partially observable MDP to determine the optimal frequency of HbA1c testing. Their model recommended less frequent screening (e.g., every 3-6 years) for lower-BMI individuals aged 45-65, but significantly more frequent screening (e.g., every 1-3 years) for obese adults (BMI ≥30) aged 65 and older. Despite these advances, existing MDP models for T2D have largely overlooked the problem of optimizing post-diagnosis follow-up intervals in PCP. Prior studies have focused on treatment selection or one-time screening frequency and estimated policies for broad risk groups. Thus, there is a need to capture the dynamic clinical trajectories inherent and provide context-specific follow-up recommendations for T2D patients. To address this gap, we formulate a CMDP to explicitly model trade-offs between time spent in uncontrolled glycemic health states and healthcare utilization, and inform context-aware follow-up for individuals diagnosed with T2D. Clinical contributions of this study include the conceptualization of trade-offs between time spent in uncontrolled glycemic states and hospitalizations between subsequent PCP visits using Electronic Health Records (EHR) data and operationalization of follow-up interval decisions in PCP. The methodological novelty relies on the application of dimensionality reduction on complex post-diagnosis clinical trajectories and mixed-data clustering methods, and their synthesis within a Markov decision modeling framework to inform context-aware follow-up decision making in primary care setting.
2. Materials and Methods 2.1. Study Population This is a retrospective observational cohort study of patients who received care at one of the ten primary care clinics within an integrated healthcare system in the mid-Atlantic region of the United States. The study period spanned January 1, 2017, through December 31, 2023. The inclusion criteria were: (i) age 18 years or older at the time of their initial primary care encounter; (ii) a documented diagnosis of T2D, as identified by relevant International Classification of Diseases, Tenth Revision (ICD10) codes (Appendix A) (World Health Organization, 2022); and (iii) at least three PCP visits within the study window. The sample size was 20,483 unique patients and 1,399,615 unique visits. The research protocol was approved by an Institutional Review Board (IRB), and the health system authorized the use of de-identified EHR data. 2.2. Methods Our analytical framework, illustrated in Figure 1, is a multi-stage approach to learn context-specific follow-up policies from EHR data. The framework consists of four main stages: (i) Data Preprocessing, where raw static and dynamic clinical data are transformed into standardized variable sets; (ii) Dimensionality Reduction, where highdimensional patient trajectories compressed into dense embeddings; (iii) Clustering, where patients are clustered into different contexts; and (iv) Policy Optimization, where a CMDP is solved to derive optimal follow-up intervals for each subpopulation.
3
Data Preprocessing The first stage involves preprocessing raw EHR observations into a standardized patient-level variable set comprising both static and dynamic variables. The static variables refer to data elements observed for each patient at their index PCP visit and do not change over time. These include: (i) demographics (e.g., sex, age at diagnosis); (ii) clinical characteristics (e.g., presence of comorbid conditions such as cardiovascular disease (CVD), chronic kidney disease (CKD), and T2D complication status indicating health complications induced by T2D, defined using appropriate ICD-10 codes as detailed in Appendix A); and (iii) healthcare utilization indices (e.g., total number of hospitalizations and total number of PCP visits over study period). The selection of these static variables is guided by evidence that demographic characteristics, comorbidities, and patterns of care utilization are key determinants of glycemic control, and adverse outcomes. Age at diagnosis and sex have been associated with differences in glycemic control and complication risk among individuals with T2D (Mahmood, Daud, and Ismail, 2016), while comorbid conditions such as CVD and CKD have been linked to poor glycemic control and high complication burden (Bitew, Alemu, Jember, Tadesse, Getaneh, Seid, and Weldeyonnes, 2023). The dynamic variables refer to time-dependent observations at each PCP visit (t) for each patient throughout the study period. These variables characterize current glycemic control, short-term changes in glycemic level, and recent acute events. • Glycemic status at time t: The patient’s HbA1c level, discretized into three categories based on the most recent measurement at or before visit t: In-Control (HbA1c <6.5% [48 mmol/mol]), Out-of-Control (HbA1c ≥6.5% [48 mmol/mol]), or Unmeasured if no prior HbA1c is on record. The threshold of 6.5% (48 mmol/mol) is consistent with American Association of Clinical Endocrinologists clinical guideline that recommends an HbA1c target of 6.5% (48 mmol/mol) for patients with T2D (Garber, Handelsman, Grunberger, Einhorn, Abrahamson, Barzilay, Blonde, Bush, DeFronzo, Garber, et al., 2020). • Hospitalizations between subsequent PCP visits: A binary indicator of any inpatient admission between two consecutive PCP visits (1 = had at least one or more hospitalizations between subsequent PCP visits, 0 = otherwise). Interval hospitalization serve as markers of acute clinical deterioration and multimorbidity burden and are associated with subsequent adverse outcomes and higher healthcare utilization (Khalid, Raluy-Callado, Curtis, Boye, Maguire, and Reaney, 2014; Schneider, Kalyani, Golden, Stearns, Wruck, Yeh, Coresh, and Selvin, 2016), making them relevant for decision about intensifying follow-up (American Diabetes Association Professional Practice Committee, 2025a). • HbA1c trend between subsequent PCP visits: A categorized variable capturing the change ∆HbA1c = HbA1ct − HbA1ct−1 from PCP visit at time t − 1 to PCP visit at time t, categorized into the following categories: ◦ Improving: ∆HbA1c < −0.5%, ◦ Stable: |∆HbA1c| ≤ 0.5%, ◦ Worsening: ∆HbA1c > 0.5%, and . ◦ Undefined: If glycemic status was unmeasured at PCP visit at either time t, time t − 1, or both. The reason for including short-term HbA1c trend is because visit-to-visit glucose fluctuations have been linked diabetes-related health complications (Suh and Kim, 2015; Zhou, Sun, Huang, Zhu, and Bian, 2020). Evidence shows that reducing short-term glycemic fluctuations can improve the severity of uncon4
trolled T2D (Monnier, Colette, and Owens, 2018). We use a cut-off of 0.5% to define improvement and worsening because HbA1c changes of this magnitude are commonly interpreted as clinically meaningful differences in studies of glycemic outcomes (Chen, Yi, Wang, Wang, Yu, Zhang, Hu, Xu, Wu, Hou, et al., 2022) Principal Component Analysis The second stage in the analytical approach is dimensionality reduction (Figure 1). In this study, the high dimensionality of the clinical trajectories arises from representing each patient’s sequence of clinical states across PCP visits over the study period. At each PCP visit, the clinical state is defined as the Cartesian combination of glycemic status, HbA1c trend, and interval hospitalization. With 3 glycemic categories, 4 trend categories, and 2 hospitalization categories, this results in 24 distinct clinical trajectory states (3 × 4 × 2 = 24) that a patient can be in at a given visit t ∈ {1, · · · , T }. All trajectories are normalized to a uniform length Tmax = 16, chosen to cover 95% of visit counts in the cohort; for longer sequences, we retain the most recent Tmax visits, and shorter sequences are right-padded with a Loss-to-Follow-Up (LTF) token. Let X = (xij ) denote the resulting feature matrix, with rows representing unique patients i = 1, . . . , n and columns representing variables preprocessed in previous stage j = 1, . . . , p. We use principal component analysis (PCA) to reduce the dimensionality of X while transforming correlated variables into a smaller set of uncorrelated components that retain the most of variation in the original dynamic variables (Boehmke and Greenwell, 2019). We determined the number of principal components (PCs) using the cumulative proportion of variance explained. After computing the eigenvalues of the variable covariance matrix, we examined the cumulative proportion of variance explained by successive PCs and retained the smallest number of components that together accounted for at least 95% of the total variance. This threshold balances dimensionality reduction with preservation of the information contained in the original trajectory features. Clustering The goal of the clustering stage is to assign patients to contexts. Using the 83 selected PCs together with 5 static variables (sex, age at diagnosis, presence of comorbid conditions such as CVD, CKD, and T2D complication status) as input, we identified K patient contexts with the K-Prototypes algorithm, a mixed-type clustering method designed to handle numeric and categorical variables (Ahmad and Dey, 2007; Huang, 1998). We treated the PCs as numeric variables and the static patient variables as categorical variables. Prior to clustering, the PCs for each patient were standardized to a zero mean and unit variance. Standardization is important for distance-based clustering methods such as K-prototype to ensure that all variables contribute comparably to the distance calculation and prevents those with larger ranges from dominating the results (Hastie, Tibshirani, Friedman, et al., 2009). The K-Prototype algorithm, which is used for mixed numeric and categorical data, constructs clusters by minimizing the total within-cluster distance by using the mean for numerical variables and the mode for categorical variables (Huang, 1998; Pasin and Gonenc, 2023; Preud’Homme, Duarte, Dalleau, Lacomblez, Bresso, Smaı̈l-Tabbone, Couceiro, Devignes, Kobayashi, Huttin, et al., 2021). We selected the number of clusters K that maximized the mean silhouette score. The final model assigned each patient a context label z ∈ {1, . . . , k} for downstream use in the CMDP transition and cost models. 5
Contextual Markov Decision Process (CMDP) In the last stage of the methodological approach, we model finding the optimal returnto-clinic interval as an infinite-horizon discounted contextual Markov decision process. Each decision epoch corresponds to a PCP visit at which the follow-up interval is selected. The CMDP is defined by the tuple (S, A, {Pz }, {Cz }, γ), where S denotes the clinical state space, A is the set of feasible follow-up intervals,Pz is the context-specific state transition, Cz is the immediate context-specific cost function that aggregates clinically interpretable components, and γ ∈ (0, 1) is the discount factor. Each patient is assigned a context z ∈ {1, . . . , K}, with context-specific state transition probabilities Pz (st+1 | st , a) and immediate cost function Cz (st , a, st+1 ). For each context z, our objective is to compute an optimal policy πz : S → A that minimizes expected discounted cumulative cost. The components of the CMDP are described in detail below. State Space. The state at each decision epoch is defined as the patient’s clinical status. Specifically, at decision epoch t (a completed primary care visit), the patient’s clinical status is
I ∆H st = sH , t , st , st
(1)
H = where the components are categorical variables defined in Section 2.2: sH t ∈ S {In-Control, Out-of-Control, Unmeasured}(glycemic status), sIt ∈ S I = {No, Yes} (interval hospitalization), and s∆H ∈ S ∆H = {Improving, Stable, Worsening, Undefined} t (short-horizon HbA1c trend). The full state space, S, is defined as the union of the Cartesian product of these non-absorbing component spaces and a terminal absorbing state:
S = S H × S I × S ∆H ∪ {LTF},
(2)
with S H ×S I ×S ∆H = 3×2×4 = 24. The absorbing state, LTF (Loss-to-Follow-Up), is entered when a patient has no subsequent PCP visit within the study observation window. Upon entering this state, the decision process terminates. Action Space. At decision epoch t, given state st , the action at ∈ A denotes the return-to-clinic interval for the next PCP visit. Because explicit provider follow-up prescriptions are not recorded in the EHR, we used the realized time to the next PCP visit as a proxy for the historical action during model estimation. This continuous interval is discretized into five month-based categories (M = month), defining A = {1M, 3M, 6M, 12M, >12M}.
(3)
For each patient, the final PCP visit observed during the study period (i.e., when no subsequent PCP visit is recorded) is represented by a terminal absorbing state and assigning the longest allowable follow-up action (at =>12M). Cost Structure. Costs are non-monetary and quantify the clinical burden accumulated in each health state and each action. Total immediate cost collected between two subsequent PCP visits is the sum of four components: (i) time spent in uncontrolled glycemic state represented as the integrated area above the HbA1c target level 6
(≥ 6.5%) multiplied with the follow-up time interval, (ii) measurement uncertainty due to delayed or unavailable HbA1c results, (iii) the burden from interval inpatient admissions, and (iv) visit utilization burden reflecting the frequency of primary care follow-up. Each component is described below. • Time and Severity of Uncontrolled Glycemic State (C G ): This component measures cumulative exposure to hyperglycemia above the clinical threshold τ = 6.5% (48 mmol/mol) over the inter-visit interval of length a (months). We integrate the area above the target HbA1c level over the follow-up interval to represent the glycemic burden accumulated between two subsequent PCP visits and scale this quantity by a state-transition adversity weight W (st , st+1 ) that differentially weights the penalty according to the type of transition (e.g., assigning higher weights to transitions that worsen or increase uncertainty about glycemic control, such as In-Control → Out-of-Control, and lower weights to transitions that improve control, such as Out-of-Control → In-Control). Assuming a linear evolution of HbA1c between visits, we approximate the weighted integral using the trapezoidal rule as C G (st , at , st+1 ) = W (st , st+1 ) ·
h
i
1 2 (Es + Est+1 ) · a ,
(4)
where Est = max{0, HbA1c(st ) − τ } and Est+1 = max{0, HbA1c(st+1 ) − τ } denote exceedance above target at the beginning (state s) and end (state st+1 ) of the interval. Here, HbA1c(st ) (and analogously HbA1c(st+1 )) is the numeric value associated with the state: when a measurement exists at or before the index visit, we use last-observation-carried-forward (LOCF); otherwise, we use a pre-specified belief value κ (and κLTF if the transition terminates in LTF). The weight W (st , st+1 ) is dimensionless and reflects the relative adversity of the transition, with larger values assigned to clinically undesirable transitions and smaller values to clinically favorable transitions. • Cost of Unobserved HbA1c (C M ): This component penalizes follow-up decisions made when the patient’s current HbA1c level is unobserved or based on outdated measurements. The penalty increases both with the degree of uncertainty about HbA1c and with the length of the selected follow-up interval. We define a nonnegative uncertainty level ut ≥ 0 that summarizes how informative the available HbA1c information is at decision epoch t: ◦ if sH t ∈ {In-Control, Out-of-Control}, ut is derived from the time since the last observed HbA1c measurement; ◦ if sH t = Unmeasured and this is not the first PCP visit, ut is derived from the length of the previously chosen follow-up interval; ◦ if sH t = Unmeasured at the first PCP visit, there is no prior measurement or action, and we instead apply a separate baseline penalty. Let dt denote the number of days since the last available HbA1c at epoch t. When a measurement exists, dt is defined as the time difference between the current visit and the last HbA1c result. When no prior HbA1c measurement is available but t > 1 and sH t = Unmeasured, dt is approximated using the length of the previous follow-up interval. We then map dt to an ordinal uncertainty
7
level via a nondecreasing step function, 0, 1, ut = λu (dt ) = 2, 4, 16,
dt ≤ 45, 45 < dt ≤ 90, 90 < dt ≤ 180, 180 < dt ≤ 366, dt > 366.
The resulting uncertainty cost at epoch t is C M (st , at , st+1 ) =
( η at ,
sH t = Unmeasured and t is the first PCP visit,
ut at ,
(5)
otherwise,
where η ≥ 0 is a scaling parameter and the impact of varying η is examined in sensitivity analyses (Section 3.6). In this formulation, longer follow-up intervals are penalized more heavily when HbA1c information is less recent or unavailable, reflecting the increased clinical risk associated with prolonged periods of glycemic uncertainty. • Cost of Hospitalization (C I ): Between consecutive PCP visits, a patient may have hospital admissions. Let = Yes } H(st+1 ) = 1{ st+1 I indicate that at least one admission occurred in the interval (t, t+1] (as recorded at visit t + 1); otherwise H(st+1 ) = 0. We penalize any such interval by C I (st , at , st+1 ) = λI H(st+1 ),
(6)
where λI > 0 is a scaling parameter that reflects the relative burden associated with hospitalization between PCP visits. • Premature Return to PCP Cost (C F ): Considering that HbA1c reflects average glycemic control over approximately 3 months (ElSayed et al., 2023), very short very short return intervals to PCP may not add any additional information value on disease trajectory. On the contrary, it may impose unnecessary burden on patients and the delivery system due to increased frequency of PCP visits. This cost element reflects the required balance between following up too soon (and not obtaining new information) vs. too late (and missing the opportunity to observe elevated HbA1c values and intervene) when deciding the length of follow-up time interval between subsequent PCP visits. To reflect this, we include a utilization cost that depends only on the action at in months. We choose a concave, decreasing function of at via a transformed logarithm, C F (st , at , st+1 ) = λF [log(
20 1.5 )] , at
(7)
where λF > 0 is a scaling parameter, and the constant 20 (months) sets a reference horizon beyond the longest action in A. This term strongly penalizes
8
extremely frequent follow-up (e.g., 1 month) relative to more routine intervals, while assigning negligible cost to the longest follow-up option.
3. Results 3.1. Descriptive Statistics of Study Population The final analytic cohort included 20,483 unique patients. The cohort was predominantly female (59.94%) and African American (66.55%), with a majority (91.90%) identifying as Non-Hispanic. The patient population was largely middle-aged and older, with 52.59% of patients (n=10,772) being 60 years or older. Nearly half of all patients (47.45%) had a documented T2D-related complication, 34.45% had CVD, and 25.77% had CKD. Table 1 summarizes the study population characteristics. 3.2. PCA Results Figure 2 plots the cumulative explained variance as a function of the number of principal components. Figure 2 illustrates a sharp initial increase followed by a plateau, indicating that each additional PC explains progressively less variance. As shown by the intersection with our 95% threshold, 83 PCs were retained from the original 409dimensional variable space (400-dimensional trajectory sequence and 9 aggregate variables). 3.3. Clustering Results The optimal number of clusters, K, for the K-Prototypes algorithm were selected based on the mean silhouette score computed using Gower dissimilarity matrix. Results showed that the silhouette score was maximized at K = 2 (0.252). The score dropped substantially for K = 3 (0.103), k = 4 (0.032), and remained low at K = 5 (0.104). We therefore proceeded with K = 2 representing two distinct contexts for developing the CMPD model. The comparison of these two contexts with regards to demographic and clinical characteristics is shown in Table 2. Based on clinical domain expert insights (W.G. and E.S.), these two contexts represent distinct patient subpopulations with different follow-up needs as observed in primary care practice. Specifically, context (cluster) 1 represents an older, higher-risk cohort. Patients in this context were predominantly 60 years or older (64.1%), with a substantially higher prevalence of comorbid CKD (34.45% vs. 23.96%), CVD (40.69% vs. 33.15%), and existing T2D-related complications (57.95% vs. 45.27%) compared to context (cluster) 2. Because older age, CKD, CVD, and diabetes complications are associated with increased risks of hospitalization, cardiovascular events, and mortality among individuals with T2D (Schneider et al., 2016; Zoungas, Woodward, Li, Cooper, Hamet, Harrap, Heller, Marre, Patel, Poulter, et al., 2014), we interpret context 1 as a higher-risk group and context 2 as a lower-risk group. 3.4. Optimal Context-aware Follow-up Policies The CMDP learned context-specific policies, πz (s), for the high-comorbidity (context 1) and lower-comorbidity (context 2) cohorts, as visualized in Figure 3. The policies 9
demonstrate a complex, data-driven stratification of care, with the optimal action depending on both the patient’s immediate clinical state and their underlying context. Three key results emerged from the policies. First, when HbA1c is unmeasured at a given PCP visit, the follow-up policy recommends a 1-month return in both contexts, with only minor deviations when the short-horizon trend is undefined. This result indicates that prolonged periods without measurement elevate the risk of unobserved hyperglycemia, thus the optimal policy prioritizes rapid re-measurement. Second, for ”In-Control” states, context 2 (the lower-risk group) frequently receives 12-month returns (e.g., for In-Control with Stable or Improving trends), whereas context 1 (older, more comorbid) is capped at 6 months. Conversely, for Out-of-Control states both contexts tighten to 1–3 months, with the shortest interval (often 1 month) triggered by adverse signals such as a worsening HbA1c trend (Worsening) or an interval hospitalization. Third, in a few cases, “In-Control with an Undefined HbA1c trend”, Context 2 optimal policy recommends a 1-month return while Context 1 optimal policy recommends 6 months. Thus, context-aware follow-up policies implement a parsimonious rule set that is easy to operationalize: (i) measure soon when HbA1c is Unmeasured; (ii) lengthen intervals for sustained control, especially in the lower-risk context; and (iii) shorten intervals to 1–3 months when control is poor, trending worse, or when a hospitalization occurred. Crucially, for nearly every clinical state the recommended interval in Context 1 is no longer than in Context 2, indicating that the EHR-derived contexts translate into systematically more intensive follow-up for the higher-risk subpopulation. 3.5. Comparative Policy Evaluation We benchmarked the learned CMDP policy against three heuristics: (i) fixed ADA guidelines, (ii) a fixed 3-month follow-up policy, and (iii) a fixed 6-month follow-up policy. Figure 4 shows the expected cumulative cost for each policy, where lower values indicate lower cost. For context 1, the CMDP policy achieved the lowest expected cost (108.7). This was substantially lower than all three baselines: the ADA rule (166.7), the always-3M policy (172.9), and the always-6M policy (184.3). The cost values shown in Figure 4 are unitless as they combine multiple cost components accumulated between subsequent PCP visits as outlined in Section 2.2. For Context 2, the CMDP policy achieved the lowest cost (185.5). This represented a smaller but consistent improvement over the baselines (198.2 for the ADA rule, 195.9 for always-3M, and 213.9 for always-6M). In this context, the fixed 3-month policy was marginally better than the ADA heuristic, but both were inferior to the adaptive CMDP policy. 3.6. Sensitivity Analysis of Cost Parameters We conducted sensitivity analyses to assess the robustness of the learned CMDP policy to alternative specifications of cost function. We varied four key parameters: (i) the belief HbA1c level κ used when glycemic status is unobserved or lost to follow-up in the glycemic burden term; (ii) the baseline uncertainty penalty η, which scales the cost of acting under outdated or unavailable HbA1c information; (iii) the weight on any inpatient admission between visits, λI ; and (iv) the state-transition adversity weight W (st , st+1 ) that amplifies glycemic burden for clinically plausible range while holding all other parameters fixed at their baseline values, re-solved the CMDP, and 10
recomputed the expected total cost the test cohort. Figure 5 compares the CMDPderived policy with four benchmark strategies (an ADA guideline policy and fixed 3-, 6-, and 12-month follow-up intervals) across these parameters ranges. In every scenario, the CMDP policy achieves the lowest expected total cost. The rank ordering of policies remains invariant: the context-aware follow-up policy consistently dominates the guideline and fixed-interval strategies. As η and λI increases, all policies become more costly, which reflects greater penalties for prolonged glycemic uncertainty and hospital admissions; however, the CMDP policy remains clearly separated from comparators, indicating that its advantages is preserved even when decision makers place substantially higher weight on these outcomes. Varying κ and W (st , st+1 ) induces only modest shifts in expected cost for all policies, with no evidence of policy cross over. Overall, these analyses show that our main conclusions are robust to reasonable reweighting of the cost components: across a wide range of preferences regarding glycemic burden, measurement uncertainty, and hospitalizations, the CMDP-based policy remains the most efficient strategy in terms of expected clinical burden.
4. Discussion In this study, we developed and evaluated a Contextual Markov Decision Process (CMDP) framework to personalize return-to-clinic intervals for patients with Type 2 Diabetes (T2D) using de-identified EHR data. By combining trajectory-based PCs with K-Prototypes clustering, we identified clinically meaningful patient contexts and derived context-specific follow-up policies. Across 22,154 patients, the learned policies were interpretable, and aligned with clinical intuition, suggesting that CMDP-based approaches are a feasible pathway toward adaptive, data-driven follow-up decisions in primary care. Clustering uncovered two distinct and clinically meaningful patient subpopulations. One context consisted of older patients with higher comorbidity burden and more intensive utilization patterns, while the other captured comparatively younger and healthier patients with fewer complications. These contexts emerged from highdimensional representations of longitudinal trajectories and enabled the CMDP to learn policies that allocate more frequent follow-up to the higher-risk context while avoiding unnecessary visits for lower-risk patients. Learned policies exhibited clear patterns when expressed over the state space st = (sH , s∆H , sI ). For patients in an unmeasured state, both context-aware follow-up policies converged on an intensive 1-month follow-up. This can be interpreted as the lack of recent glycemic data is, in itself, an important signal for follow-up interval decisions. This is aligned with the clinical practice where extended, unobserved periods can increase the risk of undetected hyperglycemia and its associated complications. The model’s recommendation to re-observe these patients underscores the high cost of uncertainty in managing T2D in primary care setting. For Out-of-Control states, the CMDP favored 1–3 month intervals, with slightly more frequent follow-up when short-term trends indicated worsening control or recent inpatient care. In contrast, for In-Control, stable patients without interval hospitalizations, the model was more willing to extend follow-up to longer intervals, especially in the healthier context. Together, these patterns mirror and refine prevailing recommendations, more frequent assessments in poorly controlled or unstable patients and more relaxed intervals in stable ones, while adding an explicit role for measurement uncertainty and hospitalization history. This asymmetry is plausibly explained by the interaction of (i) the 11
measurement-uncertainty cost (which is larger when trends cannot be computed ) and (ii) context-specific baselines (Context 2 patients are more likely to have limited recent testing despite otherwise good control). By contrast, in Context 1, the combination of adequate recent testing and higher utilization burden can make a 6-month visit optimal when glycemia is currently in control. Finally, the novel cost structure explicitly integrated three clinically relevant dimensions: (i) glycemic burden above a target threshold, (ii) uncertainty induced by infrequent or missing HbA1c measurements, and (iii) the disutility of frequent visits. This cost design forces the CMDP to resolve a genuine trade-off: visiting too often is penalized, but so is failing to measure or control HbA1c promptly, particularly when patients show signs of clinical instability. This paper has some limitations that can be addressed in future research. First, the data that used in this study was derived from ten PCPs affiliated with the same healthcare delivery system which may have compromised the generalizability of findings. To address this limitation future should study should include more diverse clinics and study population. Further, the study focuses on in-person primary care follow-up visits and does not model alternative follow-up modalities such as telehealth visits or online portal–based encounters, which are increasingly used in T2D management. Future work should explicitly incorporate these modalities and compare how different followup strategies affect visit frequency, laboratory testing, and clinical outcomes. Finally, although the current state representation captures key clinical variables (glycemic status, short-term HbA1c trend, and hospitalizations), it does not include additional information sources such as patient-reported outcomes, medication adherence measures, or continuous glucose monitoring. Incorporating these data sources would allow for more refined context definitions.
Acknowledgement(s) The authors would like to thank the clinicians at MedStar Georgetown University Hospital and data scientists at MedStar Health Research Institute, particularly Laura Schubel for project management support and Sonita S. Bennett for data acquisition and cleaning. The authors are grateful for the funding and support from the Agency of Healthcare Research and Quality (AHRQ).
Disclosure Statement The authors declare that they have no conflicts of interest.
Data Availability Statement The raw datasets used and/or analyzed during the current study are not publicly available due to data privacy restrictions. Derived and summarized data supporting the findings of this study are available from the corresponding author (MC) on request.
12
Funding This work was supported by the Agency of Healthcare Research and Quality (AHRQ) under Grant 1R01HS029792-01.
References Amir Ahmad and Lipika Dey. A k-mean clustering algorithm for mixed numeric and categorical data. Data & Knowledge Engineering, 63(2):503–527, 2007. American Diabetes Association Professional Practice Committee. 16. diabetes care in the hospital: Standards of care in diabetes–2025. Diabetes Care, 48:S321, 2025a. American Diabetes Association Professional Practice Committee. 6. glycemic goals and hypoglycemia: Standards of care in diabetes–2025. Diabetes Care, 48(Suppl 1):S128–S145, 2025b. . URL https://doi.org/10.2337/dc25-S006. American Diabetes Association Professional Practice Committee. Introduction and methodology: standards of care in diabetes—2025. Diabetes care, 48:S1–S5, 2025c. Keiko Asao, Laura N McEwen, Jesse C Crosson, Beth Waitzfelder, and William H Herman. Revisit frequency and its association with quality of care among diabetic patients: Translating research into action for diabetes (triad). Journal of diabetes and its complications, 28 (6):811–818, 2014. Turgay Ayer, Oguzhan Alagoz, and Natasha K Stout. Or forum—a pomdp approach to personalize mammography screening decisions. Operations Research, 60(5):1019–1034, 2012. Zebenay Workneh Bitew, Ayinalem Alemu, Desalegn Abebaw Jember, Erkihun Tadesse, Fekadeselassie Belege Getaneh, Awole Seid, and Misrak Weldeyonnes. Prevalence of glycemic control and factors associated with poor glycemic control: A systematic review and metaanalysis. INQUIRY: The Journal of Health Care Organization, Provision, and Financing, 60:00469580231155716, 2023. Brad Boehmke and Brandon M Greenwell. Hands-on machine learning with R. Chapman and Hall/CRC, 2019. Antonio Ceriello and Stephen Colagiuri. Idf global clinical practice recommendations for managing type 2 diabetes–2025. Diabetes Research and Clinical Practice, page 112152, 2025. Junxiang Chen, Qian Yi, Yuxiang Wang, Jingyi Wang, Hancheng Yu, Jijuan Zhang, Mengyan Hu, Jiajing Xu, Zixuan Wu, Leying Hou, et al. Long-term glycemic variability and risk of adverse health outcomes in patients with diabetes: A systematic review and meta-analysis of cohort studies. Diabetes research and clinical practice, 192:110085, 2022. Nuha A ElSayed, Grazia Aleppo, Vanita R Aroda, Raveendhara R Bannuru, Florence M Brown, Dennis Bruemmer, Billy S Collins, Marisa E Hilliard, Diana Isaacs, Eric L Johnson, et al. 6. glycemic targets: standards of care in diabetes—2023. Diabetes care, 46 (Supplement 1):S97–S110, 2023. Alan J Garber, Yehuda Handelsman, George Grunberger, Daniel Einhorn, Martin J Abrahamson, Joshua I Barzilay, Lawrence Blonde, Michael A Bush, Ralph A DeFronzo, Jeffrey R Garber, et al. Consensus statement by the american association of clinical endocrinologists and american college of endocrinology on the comprehensive type 2 diabetes management algorithm–2020 executive summary. Endocrine Practice, 26(1):107–139, 2020. Ali Hajjar and Oguzhan Alagoz. Personalized disease screening decisions considering a chronic condition. Management Science, 69(1):260–282, 2023. Trevor Hastie, Robert Tibshirani, Jerome Friedman, et al. The elements of statistical learning, 2009. Zhexue Huang. Extensions to the k-means algorithm for clustering large data sets with categorical values. Data mining and knowledge discovery, 2(3):283–304, 1998. JM Khalid, M Raluy-Callado, BH Curtis, KS Boye, A Maguire, and M Reaney. Rates and
13
risk of hospitalisation among patients with type 2 diabetes: retrospective cohort study using the uk general practice research database linked to english hospital episode statistics. International journal of clinical practice, 68(1):40–48, 2014. Pamela R Kushner, Matthew A Cavender, and Christian W Mende. Role of primary care clinicians in the management of patients with type 2 diabetes and cardiorenal diseases. Clinical Diabetes, 40(4):401–412, 2022. MI Mahmood, Faiz Daud, and Aniza Ismail. Glycaemic control and associated factors among patients with diabetes at public health clinics in johor, malaysia. Public health, 135:56–65, 2016. Fanwen Meng, Yan Sun, Bee Hoon Heng, and Melvin Khee Shing Leow. Analysis via markov decision process to evaluate glycemic control strategies of a large retrospective cohort with type 2 diabetes: the ameliorate study. Acta Diabetologica, 57(7):827–834, 2020. L Monnier, C Colette, and DR Owens. The application of simple metrics in the assessment of glycaemic variability. Diabetes & metabolism, 44(4):313–319, 2018. Fritha Morrison, Maria Shubina, and Alexander Turchin. Encounter frequency and serum glucose level, blood pressure, and cholesterol level control in patients with diabetes mellitus. Archives of internal medicine, 171(17):1542–1550, 2011. Sylvain Raoul Simeni Njonnou, Sybil Leslie Nguiloung Nguedoung, Eric Balti, Michelle Carolle Dongmo Demanou, Fernando Kemta Lekpa, Christian Ngongang Ouankou, Anne Mireille Ongmeb Boli, Ruth Michaella Mogoum Wafo, Herna Stella Chimy Tchounchui, and Siméon Pierre Choukem. Factors associated with poor glycemic control in patients with type 2 diabetes at the bafoussam regional hospital: a preliminary cross-sectional study at the bafoussam regional hospital (cameroon). Journal of Xiangya Medicine, 9, 2024. Ozge Pasin and Senem Gonenc. An investigation into epidemiological situations of covid-19 with fuzzy k-means and k-prototype clustering methods. Scientific Reports, 13(1):6255, 2023. Gregoire Preud’Homme, Kevin Duarte, Kevin Dalleau, Claire Lacomblez, Emmanuel Bresso, Malika Smaı̈l-Tabbone, Miguel Couceiro, Marie-Dominique Devignes, Masatake Kobayashi, Olivier Huttin, et al. Head-to-head comparison of clustering methods for heterogeneous data: a simulation-driven benchmark. Scientific reports, 11(1):4202, 2021. Andrea LC Schneider, Rita R Kalyani, Sherita Golden, Sally C Stearns, Lisa Wruck, Hsin Chieh Yeh, Josef Coresh, and Elizabeth Selvin. Diabetes and prediabetes and risk of hospitalization: the atherosclerosis risk in communities (aric) study. Diabetes care, 39(5):772–779, 2016. Sunghwan Suh and Jae Hyeon Kim. Glycemic variability: how do we measure it and why is it important? Diabetes & metabolism journal, 39(4):273, 2015. Tomohiko Ukai, Shuhei Ichikawa, Miho Sekimoto, Satoru Shikata, and Yousuke Takemura. Effectiveness of monthly and bimonthly follow-up of patients with well-controlled type 2 diabetes: a propensity score matched cohort study. BMC Endocrine Disorders, 19(1):43, 2019. PR Wermeling, KJ Gorter, RK Stellato, GA De Wit, JWJ Beulens, and GEHM Rutten. Effectiveness and cost-effectiveness of 3-monthly versus 6-monthly monitoring of well-controlled type 2 diabetes patients: a pragmatic randomised controlled patient-preference equivalence trial in primary care (effimodi study). Diabetes, Obesity and Metabolism, 16(9):841–849, 2014. World Health Organization. International classification of diseases, 11th revision (icd-11), 2022. URL https://www.who.int/standards/classifications/ classification-of-diseases. Accessed: 2023-01-01. Chou-Chun Wu and Sze-Chuan Suen. Optimizing diabetes screening frequencies for at-risk groups. Health Care Management Science, 25(1):1–23, 2022. James J Yahaya, Irene F Doya, Emmanuel D Morgan, Advera I Ngaiza, and Deogratius Bintabara. Poor glycemic control and associated factors among patients with type 2 diabetes mellitus: a cross-sectional study. Scientific Reports, 13(1):9673, 2023. Zheng Zhou, Bao Sun, Shiqiong Huang, Chunsheng Zhu, and Meng Bian. Glycemic variability:
14
adverse clinical outcomes and how to improve it? Cardiovascular diabetology, 19(1):102, 2020. Sophia Zoungas, Mark Woodward, Qiang Li, Mark E Cooper, Pavel Hamet, Stephen Harrap, Simon Heller, Michel Marre, Anushka Patel, Neil Poulter, et al. Impact of age, age at diagnosis and duration of diabetes on the risk of macrovascular and microvascular complications and death in type 2 diabetes. Diabetologia, 57(12):2465–2474, 2014.
Appendices Appendix A. Table A1.: List of ICD-10 Codes Representing the Presence of T2D Diagnosis Code
Diagnosis
E11.9 E11.40 E11.8 E11.65 E11.319
Type 2 diabetes mellitus without complications Type 2 diabetes mellitus with diabetic neuropathy, unspecified Type 2 diabetes mellitus with unspecified complications Type 2 diabetes mellitus with hyperglycemia Type 2 diabetes mellitus with unspecified diabetic retinopathy without macular edema Type 2 diabetes mellitus with diabetic cataract Type 2 diabetes mellitus with diabetic chronic kidney disease Type 2 diabetes mellitus with hypoglycemia without coma Type 2 diabetes mellitus with diabetic peripheral angiopathy without gangrene Type 2 diabetes mellitus with diabetic nephropathy Type 2 diabetes mellitus with diabetic polyneuropathy Type 2 diabetes mellitus with diabetic peripheral angiopathy with gangrene Type 2 diabetes mellitus with ketoacidosis without coma Type 2 diabetes mellitus with diabetic autonomic (poly)neuropathy Type 2 diabetes mellitus with other specified complication Type 2 diabetes mellitus with other diabetic arthropathy Type 2 diabetes mellitus with other diabetic ophthalmic complication Type 2 diabetes mellitus with other diabetic kidney complication Type 2 diabetes mellitus with diabetic neuropathic arthropathy Type 2 diabetes mellitus with other diabetic neurological complication Type 2 diabetes mellitus with foot ulcer Type 2 diabetes mellitus with proliferative diabetic retinopathy without macular edema, unspecified eye Type 2 diabetes mellitus with other skin complications Type 2 diabetes mellitus with hyperosmolarity without nonketotic hyperglycemic-hyperosmolar coma (NKHHC) Type 2 diabetes mellitus with other skin ulcer Type 2 diabetes mellitus with mild nonproliferative diabetic retinopathy without macular edema, unspecified eye
E11.36 E11.22 E11.649 E11.51 E11.21 E11.42 E11.52 E11.10 E11.43 E11.69 E11.618 E11.39 E11.29 E11.610 E11.49 E11.621 E11.3599 E11.628 E11.00 E11.622 E11.3299
Continued on next page
15
Table A1 – Continued from previous page Diagnosis Code
Diagnosis
E11.3313
Type 2 diabetes mellitus with moderate nonproliferative diabetic retinopathy with macular edema, bilateral Type 2 diabetes mellitus with other circulatory complications Type 2 diabetes mellitus with diabetic mononeuropathy Type 2 diabetes mellitus with mild nonproliferative diabetic retinopathy without macular edema, right eye Type 2 diabetes mellitus with mild nonproliferative diabetic retinopathy with macular edema, right eye Type 2 diabetes mellitus with moderate nonproliferative diabetic retinopathy without macular edema, bilateral Type 2 diabetes mellitus with mild nonproliferative diabetic retinopathy without macular edema, bilateral Type 2 diabetes mellitus with mild nonproliferative diabetic retinopathy with macular edema, bilateral Type 2 diabetes mellitus with proliferative diabetic retinopathy without macular edema, left eye Type 2 diabetes mellitus with stable proliferative diabetic retinopathy, left eye Type 2 diabetes mellitus with stable proliferative diabetic retinopathy, bilateral Type 2 diabetes mellitus with proliferative diabetic retinopathy without macular edema, bilateral Type 2 diabetes mellitus with proliferative diabetic retinopathy with traction retinal detachment involving the macula, right eye Type 2 diabetes mellitus with proliferative diabetic retinopathy without macular edema, right eye Type 2 diabetes mellitus with proliferative diabetic retinopathy with macular edema, bilateral Type 2 diabetes mellitus with proliferative diabetic retinopathy with traction retinal detachment not involving the macula, bilateral Type 2 diabetes mellitus with proliferative diabetic retinopathy with traction retinal detachment involving the macula, bilateral Type 2 diabetes mellitus with severe nonproliferative diabetic retinopathy without macular edema, left eye Type 2 diabetes mellitus with severe nonproliferative diabetic retinopathy with macular edema, right eye Type 2 diabetes mellitus with moderate nonproliferative diabetic retinopathy without macular edema, right eye Type 2 diabetes mellitus with mild nonproliferative diabetic retinopathy with macular edema, left eye Type 2 diabetes mellitus with severe nonproliferative diabetic retinopathy with macular edema, bilateral Type 2 diabetes mellitus with severe nonproliferative diabetic retinopathy with macular edema, unspecified eye Type 2 diabetes mellitus with proliferative diabetic retinopathy with traction retinal detachment not involving the macula, left eye Type 2 diabetes mellitus with proliferative diabetic retinopathy with combined traction retinal detachment and rhegmatogenous retinal detachment, left eye Type 2 diabetes mellitus with mild nonproliferative diabetic retinopathy without macular edema, left eye
E11.59 E11.41 E11.3291 E11.3211 E11.3393 E11.3293 E11.3213 E11.3592 E11.3552 E11.3553 E11.3593 E11.3521 E11.3591 E11.3513 E11.3533 E11.3523 E11.3492 E11.3411 E11.3391 E11.3212 E11.3413 E11.3419 E11.3532 E11.3542 E11.3292
Continued on next page
16
Table A1 – Continued from previous page Diagnosis Code
Diagnosis
E11.3219
Type 2 diabetes mellitus with mild nonproliferative diabetic retinopathy with macular edema, unspecified eye Type 2 diabetes mellitus with proliferative diabetic retinopathy with macular edema, right eye Type 2 diabetes mellitus with severe nonproliferative diabetic retinopathy without macular edema, bilateral Type 2 diabetes mellitus with severe nonproliferative diabetic retinopathy with macular edema, left eye Type 2 diabetes mellitus with stable proliferative diabetic retinopathy, right eye Type 2 diabetes mellitus with proliferative diabetic retinopathy with macular edema, unspecified eye Type 2 diabetes mellitus with unspecified diabetic retinopathy with macular edema Type 2 diabetes mellitus with moderate nonproliferative diabetic retinopathy without macular edema, unspecified eye Type 2 diabetes mellitus with moderate nonproliferative diabetic retinopathy without macular edema, left eye Type 2 diabetes mellitus with severe nonproliferative diabetic retinopathy without macular edema, unspecified eye Type 2 diabetes mellitus with hyperosmolarity with coma Type 2 diabetes mellitus with moderate nonproliferative diabetic retinopathy with macular edema, right eye
E11.3511 E11.3493 E11.3412 E11.3551 E11.3519 E11.311 E11.3399 E11.3392 E11.3499 E11.01 E11.3311
Table A2.: List of ICD-10 Codes Representing the Presence of Chronic Kidney Disease Diagnosis Code
Diagnosis
I12.0
Hypertensive chronic kidney disease with stage 5 chronic kidney disease or end stage renal disease Hypertensive chronic kidney disease with stage 1 through stage 4 chronic kidney disease, or unspecified chronic kidney disease Hypertensive heart and chronic kidney disease with heart failure and stage 1 through stage 4 chronic kidney disease, or unspecified chronic kidney disease Hypertensive heart and chronic kidney disease without heart failure, with stage 1 through stage 4 chronic kidney disease, or unspecified chronic kidney disease Hypertensive heart and chronic kidney disease without heart failure, with stage 5 chronic kidney disease, or end stage renal disease Hypertensive heart and chronic kidney disease with heart failure and with stage 5 chronic kidney disease, or end stage renal disease Chronic kidney disease, stage 1 Chronic kidney disease, stage 2 (mild) Chronic kidney disease, stage 3 (moderate) Chronic kidney disease, stage 3 unspecified Chronic kidney disease, stage 3a Chronic kidney disease, stage 3b Chronic kidney disease, stage 4 (severe) Chronic kidney disease, stage 5 Chronic kidney disease, unspecified
I12.9 I13.0 I13.10 I13.11 I13.2 N18.1 N18.2 N18.3 N18.30 N18.31 N18.32 N18.4 N18.5 N18.9
Continued on next page
17
Table A2 – Continued from previous page Diagnosis Code
Diagnosis
N19 Z94.0
Unspecified kidney failure Kidney transplant status
Table A3.: List of ICD-10 Codes Representing the Presence of Cardiovascular Disease Diagnosis Code
Diagnosis
I25.10
Atherosclerotic heart disease of native coronary artery without angina pectoris Heart failure, unspecified Cerebral infarction, unspecified Occlusion and stenosis of bilateral carotid arteries Occlusion and stenosis of unspecified carotid artery Cerebral atherosclerosis Cerebrovascular disease, unspecified Peripheral vascular disease, unspecified
I50.9 I63.9 I65.23 I65.29 I67.2 I67.9 I73.9
18
Tables Table 1.: Baseline characteristics of the study population. Characteristic
Count (%)
Sex Female Male Unknown
12278 (59.94%) 8202 (40.04%) 3 (0.01%)
Race African American White Asian Other Unknown
13632 (66.55%) 4913 (23.99%) 345 (1.68%) 1198 (5.85%) 395 (1.93%)
Ethnicity Non-Hispanic Hispanic Unknown
18824 (91.90%) 517 (2.52%) 1142 (5.58%)
Age 18–49 50–59 60–69 70+
4476 (21.85%) 5235 (25.56%) 5794 (28.29%) 4978 (24.30%)
Characteristic
Count (%)
CKD Yes No
5278 (25.77%) 15205 (74.23%)
CVD Yes No
7056 (34.45%) 13427 (65.55%)
T2D complication Yes No
9720 (47.45%) 10763 (52.55%)
Marital status Single Married Widowed Unknown
10171 (49.66%) 7823 (38.19%) 2420 (11.81%) 69 (0.34%)
Variables are grouped as demographic (sex, race, ethnicity, age, marital status) and clinical (chronic kidney disease [CKD], cardiovascular disease [CVD], Type 2 diabetes [T2D] complication). Values are counts, with percentages in parentheses.
19
Table 2.: Cluster characteristics (counts with percentages). Count (%) Cluster 1
Cluster 2
Sex Female Male Unknown
2278 (64.64%) 10000 (58.97%) 1246 (35.36%) 6956 (41.02%) 3 (0.02%)
Age 18–49 50–59 60–69 70+
441 (12.51%) 824 (23.38%) 1177 (33.4%) 1082 (30.7%)
CKD (ICD) CKD No CKD
1214 (34.45%) 4064 (23.96%) 2310 (65.55%) 12895 (76.04%)
CVD (ICD) CVD No CVD
1434 (40.69%) 5622 (33.15%) 2090 (59.31%) 11337 (66.85%)
T2D complication status With complication Without complication
2042 (57.95%) 7678 (45.27%) 1482 (42.05%) 9281 (54.73%)
4035 (23.79%) 4411 (26.01%) 4617 (27.22%) 3896 (22.97%)
Abbreviations: CKD = Chronic Kidney Disease; CVD = Cardiovascular Disease.
20
Figures
2. Dimensionality Reduction 4. Contextual Markov Decision Process …
1. Data Preprocessing Clinical Characteristics (Categorical Variables)
Trajectory Frequencies (Numerical) ! categorical variables) + ! visits
One-Hot Encoding (Categorical Variables) [0 0 1…1]
𝒂!! ⋮ 𝒂#!
Standardized Variables (Numerical Variables)
⋯ 𝒂!" ⋱ ⋮ ⋯ 𝒂#" #×%
𝑃𝐶!! ⋮ 𝑃𝐶#!
…
⋯ 𝑃𝐶!& ⋱ ⋮ ,𝑘 < 𝑝 ⋯ 𝑃𝐶#& #×'
3. Clustering 𝑃𝐶!! ⋮ 𝑃𝐶#!
⋯ 𝑃𝐶!& ⋱ ⋮ ⋯ 𝑃𝐶#& #×' Static Variables
Reduced Dynamic Variables
𝜋: 𝒮 × 𝒵 ⟶ 𝒜
K Prototype
Figure 1.: Overview of the methodological approach for deriving context-specific followup policies. PC stands for principal component. n represent the number of unique patients in study population. Dynamic variables are observed at each PCP visit at time t. After PCA, the dimension is reduced to k. st represents health state at PCP visit t, at the follow-up action, ct cost accumulated between two subsequent PCP visits at times t1 and t, and z stands for the context.
Figure 2.: Selection of principal components based on cumulative explained variance. The plot shows the cumulative proportion of variance (y-axis) explained by the principal components (x-axis). The dashed vertical line indicates that 83 components were required to capture the pre-specified 95% (red horizontal line) of the total variance in the standardized feature set.
21
Figure 3.: Each row denotes a clinical state s=(glycemic status, short-horizon HbA1c trend, interval hospitalization), with hospitalization coded as 0 (none) or 1+ (≥ 1 admission). Columns correspond to the two patient contexts. Cell values are the optimal follow-up intervals in months {1,3,6,12}.
Figure 4.: Comparative policy evaluation by patient context. Expected cumulative cost (lower is better) for the learned CMDP policy (Optimal Value) versus three baselines– ADA rule, fixed 3-month, and fixed 6-month intervals–shown separately for cluster 1 (higher comorbidity) and cluster 2 (lower comorbidity). Numeric labels above bars are absolute costs; the CMDP policy achieves the lowest cost in both clusters. 22
Figure 5.: Sensitivity analysis of cost parameters used in CMDP model.
Figure Captions • Figure 1. Overview of the methodological approach for deriving context-specific followup policies. PC stands for principal component. n represent the number of unique patients in study population. Dynamic variables are observed at each PCP visit at time t. After PCA, the dimension is reduced to k. st represents health state at PCP visit t, at the follow-up action, ct cost accumulated between two subsequent PCP visits at times t1 and t, and z stands for the context. • Figure 2. Selection of principal components based on cumulative explained variance. The plot shows the cumulative proportion of variance (y-axis) explained by the principal components (x-axis). The dashed vertical line indicates that 83 components were required to capture the pre-specified 95% (red horizontal line) of the total variance in the standardized feature set. • Figure 3. Each row denotes a clinical state s=(glycemic status, short-horizon HbA1c trend, interval hospitalization), with hospitalization coded as 0 (none) or 1+ (≥ 1 admission). Columns correspond to the two patient contexts. Cell values are the optimal follow-up intervals in months {1,3,6,12}. • Figure 4. Comparative policy evaluation by patient context. Expected cumulative cost (lower is better) for the learned CMDP policy (Optimal Value) versus three baselines– ADA rule, fixed 3-month, and fixed 6-month intervals–shown separately for cluster 1 (higher comorbidity) and cluster 2 (lower comorbidity). Numeric labels above bars are absolute costs; the CMDP policy achieves the lowest cost in both clusters. • Figure 5. Sensitivity analysis of cost parameters used in CMDP model.
23