Validation of algorithms for identifying people living with HIV in French medico-administrative databases: implications for HIV surveillance - PMC Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice BMC Med Res Methodol . 2026 Mar 12;26:90. doi: 10.1186/s12874-026-02820-5 Search in PMC Search in PubMed View in NLM Catalog Add to search Validation of algorithms for identifying people living with HIV in French medico-administrative databases: implications for HIV surveillance Marc-Florent Tassi Marc-Florent Tassi 1 Public Health and Prevention Division, CHRU de Tours, Tours, France 2 INSERM U1259 MAVIVHe, Université de Tours, Tours, France Find articles by Marc-Florent Tassi 1, 2, ✉ , Adrien Lemaignen Adrien Lemaignen 3 Infectious and Tropical Diseases Department, CHRU de Tours, Tours, France 4 UR 7505 - Education Ethique Santé, Université de Tours, Tours, France Find articles by Adrien Lemaignen 3, 4 , Karl Stéfic Karl Stéfic 2 INSERM U1259 MAVIVHe, Université de Tours, Tours, France 5 HIV National Reference Center, CHRU de Tours, Tours, France Find articles by Karl Stéfic 2, 5 , Leslie Grammatico-Guillon Leslie Grammatico-Guillon 1 Public Health and Prevention Division, CHRU de Tours, Tours, France 2 INSERM U1259 MAVIVHe, Université de Tours, Tours, France Find articles by Leslie Grammatico-Guillon 1, 2 Author information Article notes Copyright and License information 1 Public Health and Prevention Division, CHRU de Tours, Tours, France 2 INSERM U1259 MAVIVHe, Université de Tours, Tours, France 3 Infectious and Tropical Diseases Department, CHRU de Tours, Tours, France 4 UR 7505 - Education Ethique Santé, Université de Tours, Tours, France 5 HIV National Reference Center, CHRU de Tours, Tours, France ✉ Corresponding author. Received 2025 Jul 23; Accepted 2026 Mar 4; Collection date 2026. © The Author(s) 2026 Open Access This article is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License, which permits any non-commercial use, sharing, distribution and reproduction in any medium or format, as long as you give appropriate credit to the original author(s) and the source, provide a link to the Creative Commons licence, and indicate if you modified the licensed material. You do not have permission under this licence to share adapted material derived from this article or parts of it. The images or other third party material in this article are included in the article’s Creative Commons licence, unless indicated otherwise in a credit line to the material. If material is not included in the article’s Creative Commons licence and your intended use is not permitted by statutory regulation or exceeds the permitted use, you will need to obtain permission directly from the copyright holder. To view a copy of this licence, visit http://creativecommons.org/licenses/by-nc-nd/4.0/ . PMC Copyright notice PMCID: PMC13094274 PMID: 41820882 Abstract Background French medico-administrative databases (Système National des Données de Santé, SNDS) are increasingly used to study the epidemiology of human immunodeficiency virus (HIV) infections. However, algorithms used to target HIV-positive patient populations have not yet been evaluated. This study aimed to validate algorithms for identifying people living with HIV in French medico-administrative databases and thus estimate diagnosed HIV prevalence and incidence. Methods Two algorithms were evaluated to detect diagnosed HIV cases. A gold-standard cohort was created by matching clinical data from a hospital data warehouse to the SNDS. The performance of the algorithms was assessed using Bayesian models, then calculated adjusted national HIV prevalence and incidence were estimated on the entire SNDS based on algorithm performance. Results For the prevalently diagnosed HIV population, the algorithms showed high sensitivity and specificity (> 99%). For new diagnoses, the sensitivity was 82–95% with a specificity > 99.9%. The positive predictive values varied substantially among the algorithms. The adjusted 2020 prevalence estimates ranged from 135,448 − 145,704 individuals diagnosed with HIV. For 2017–2023, the algorithms estimated 27,283 − 45,346 new diagnoses, compared to 40,912 from mandatory notifications. Conclusion Medico-administrative data can provide useful automated HIV surveillance, but algorithm performance must be considered when interpreting the results. More sophisticated algorithms may improve the accuracy. These methods could complement existing HIV monitoring systems in France. Supplementary Information The online version contains supplementary material available at 10.1186/s12874-026-02820-5. Keywords: HIV, Medico-administrative data, Validation study, Surveillance Introduction The study of human immunodeficiency virus (HIV) epidemiology has traditionally relied on the implementation and surveillance of cohorts, encompassing the collection of clinical, biological, and pharmaceutical data pertaining to people living with HIV (PLHIV) on one hand, and the notification of new infections by healthcare professionals, on the other hand [ 1 ]. These strategies require the allocation of substantial human and financial resources to establish reliable databases aiming to achieve robust epidemiological indicators. The COVID-19 pandemic has highlighted the vulnerability of this epidemiological surveillance model, particularly when patient management and screening procedures are disrupted, or when the surge in the caregivers workload alters the priorities of those involved in data generation and collection [ 2 , 3 ]. In France, the COVID-19 crisis exacerbated the incompleteness of declarative mandatory notification. Consequently, the French public health agency ( Santé publique France ) had to postpone the publication of the 2019 estimate of newly diagnosed HIV cases by one year [ 4 ]. Owing to its universal health insurance system, France possesses a database encompassing all healthcare consumption of residents with social insurance. This medico-administrative database, called the National Health Data System ( Système National des Données de Santé , SNDS), allows the study of many aspects of HIV epidemiology using data routinely collected by the National Health Insurance Agency ( Caisse Nationale de l’Assurance Maladie , CNAM) [ 5 ]. Moreover, in the case of surveillance systems, these databases also provide a permanent source of comprehensive data that can be rapidly updated. The application of real-world data for epidemiological purposes is now widely accepted, and surveillance systems based on this data source benefit from a substantial reduction in human and financial resources [ 5 – 8 ]. The universal nature of the French health insurance system guarantees that diseases can be studied on the basis of virtually complete patient populations, including (1) all outpatient consumption of goods, services and medical cares reimbursed by the French health insurance, (2) hospitalisation data in public and private health facilities, and (3) death certificates [ 9 , 10 ]. Over the last 15 years, there has been a tremendous increase in research on the epidemiology of infectious diseases using SNDS data. In addition to academic research, many public agencies currently use medico-administrative data for public health purposes in the decision-making and development of health policies [ 6 ]. However, the main limitation of SNDS is the lack of direct information on patients’ clinical and biological data. This inherent limitation can be overcome by developing algorithms based on healthcare consumption data to identify specific medical conditions. Several studies have developed algorithms using SNDS databases to identify populations of interest for HIV epidemiological research [ 7 , 11 , 12 ]. Without validation, it is exceedingly challenging to quantify the bias resulting from deficiencies in SNDS algorithms [ 13 ]. Hence, numerous studies emphasised the need to develop reliable validated methods for detecting pathologies in the SNDS [ 6 , 14 , 15 ]. This study aimed to assess the performance of HIV SNDS algorithms in estimating HIV prevalence and detecting newly diagnosed cases in France, addressing the critical gap in the validation of these epidemiological tools using multi-source data from a clinical data warehouse matched to SNDS, then calculate adjusted national HIV prevalence and incidence estimates on the entire SNDS based on algorithm performance. Methods Study design This validation study of the medico-administrative data identification algorithms for HIV surveillance is performed and reported in accordance with the recommendations issued by Benchimol et al. (see Additional file 1) [ 16 ]. Definition of HIV algorithms used for assessment Two algorithms targeting the prevalent diagnosed PLHIV population were evaluated. The first algorithm is presented in version G11 (May 2024) of the medical methodology used by the Health Insurance Agency to produce an annual cartography of pathologies and medical expenditures [ 11 , 17 ]. This algorithm is most often used in HIV studies based on SNDS data. It will be referred to as the G11 algorithm in the remainder of this article. According to the G11 algorithm, a person is classified as HIV-positive in a given year (year n ) if at least one of the following conditions is recorded in the SNDS database. HIV-related long-term illness in year n (International Classification of Diseases 10th Revision: ICD-10 codes [roots B20, B21, B22, B23, B24, F024, and Z21]) Hospitalisation in years n-4 to n in an acute care setting or in a psychiatric ward with a standardised discharge summary that includes the concept of HIV-related illness (ICD-10 codes listed above). Hospitalisation in an acute care setting in year n for any other reason with the notion of HIV-related illness present in one of the medical unit summaries (ICD-10 codes listed above). Dispensation of at least one antiretroviral drug specific to HIV treatment on three different dates during year n (excluding tenofovir and emtricitabine). Receipt of a medical-biological procedure specific to the management of HIV during year n (plasma HIV RNA quantification, HIV genotypic resistance test, search for HLA B57:01 genotype, and antiretroviral concentration quantification). The second algorithm, derived from the G11 algorithm, was developed as a part of this study. This algorithm is referred to as ValORIS algorithm . Using the ValORIS algorithm, a person is considered HIV-positive if at least two of the following conditions are identified in the SNDS: HIV-related long-term illness, hospitalisation condition, dispensing of antiretroviral drugs, and each biological procedure independently. The criteria relating to drug dispensation and medical-biological procedures were modified as follows: 4. Dispensation of at least one antiretroviral drug specific to HIV treatment on five different dates over a rolling one-year period during years n-4 to n . 5. Receipt of a medical-biological procedure specific to the management of HIV during year s n-4 to n . To estimate the number of newly diagnosed cases within a specific year, two adjustments were implemented in the algorithms. All criteria previously applied to a five-year period (hospitalisations and antiretroviral dispensations) were restricted to a single year. All individuals identified as HIV-positive in the preceding years were presumed to have the condition at the start of the study year, and were consequently excluded. Gold standard: data source and constitution Launched in 2016, the Clinical Data Centre (CDC) of the Tours Teaching Hospital is a hospital data warehouse that centralises detailed clinical, biological, and pharmaceutical data for all patients admitted to the hospital for consultation, emergency care, or hospitalisation [ 18 – 20 ]. In 2019, the CDC obtained authorisation from the French Data Protection Authority ( Commission Nationale Informatique et Libertés , CNIL ) for research purposes (authorisation file 2212853, tacit agreement 12/11/2019). At the beginning of 2022, the CDC contained more than 130 million data elements for two million patients, providing a potential validation database for the HIV SNDS algorithms. In this study, the information available at the CDC before September 2022 was used to establish the gold standard. As the gold standard was created before matching with data from the SNDS, individuals were selected, and their HIV status was determined blindly to the classification results of the targeting SNDS algorithms evaluated in this study. All individuals > 18 years of age in 2020 (the year used to validate the prevalence algorithm) were eligible for inclusion in the gold standard. HIV-negative individuals were selected based on at least one of the following criteria. Data from the hospital virology laboratory: negative HIV serology. Text data from hospitalisation, consultation, and surgery reports containing the term HIV and in which a regular expression targeting HIV-negative individuals was found. HIV-negative participants were considered HIV-negative until the final date confirming this status. Subsequently, participants were censored. For PLHIV, the initial selection was performed based on at least one of the following criteria: Data from the hospital virology laboratory: positive or low positive immunoblot, plasma HIV RNA quantification, HIV serotyping, HIV genotypic resistance test, and HIV tropism test. Data from the hospital pharmacy: at least one antiretroviral drug was dispensed by the hospital pharmacy. Data from the standardised discharge summary for the hospital stay: ICD-10 diagnosis codes (roots B20, B21, B22, B23, B24, F024, O987, and Z21) and start and end dates of stay. Data from chart reports: textual data from hospitalisation, consultation, and surgery reports containing the term HIV or the name of an antiretroviral drug (original drug denomination or international non-proprietary name) and in which a regular expression targeting PLHIV was found. The seropositivity of all individuals identified as HIV-positive in the previous step was verified by reviewing their medical records available in the CDC. This manual verification was conducted by a pharmacist, and doctoral candidate in epidemiology from the INSERM MAVIVH unit. This manual review process was also used to document the seropositivity detection dates for all subjects. Subjects were excluded from the study if they could not be determined accurately (see Additional file 2). For all participants included in the gold standard, the following additional information was extracted from the CDC: date of birth, sex, country of birth, and dates of the elements used to classify them. Matching the gold standard with SNDS data Direct matching was performed between CDC extraction and SNDS based on the national health insurance number, derived from individual information such as sex, date of birth, place of birth, and birth certificate number. Algorithms evaluations Ability to identify the prevalent population The performance of the algorithms in identifying the prevalent population of diagnosed PLHIV was estimated for 2020 using data from the gold standard, selecting: (i) PLHIV diagnosed before January 1, 2021, and (ii) subjects for whom confirmation of seronegativity was available after December 31, 2020 (prevalent population specific inclusion criteria). For all subjects, the G11 and ValORIS algorithms were applied to their SNDS data (2016–2020). The sensitivity (Se) and specificity (Sp) parameters of the algorithms were estimated using a Bayesian generalised linear model with a binomial distribution and a logit link function [ 21 ]. To account for a potential spectrum effect linked to gender and country of origin (French born or born abroad, according to the data available in the CDC and in the SNDS), we considered a multilevel model followed by post-stratification [ 22 , 23 ]. For a given individual, the result of the classification by the algorithm ( ) follows a binomial distribution that depends on its status in the gold standard ( ), gender, and country of origin ( ) such that: The priors chosen for the coefficients of this model correspond to moderately informative priors, assuming that sensitivity and specificity are greater than 0.672 with a probability of 0.9 (see Additional file 3). For each of the four groups, defined by sex and origin, Se and Sp were defined as follows: Post-stratification was subsequently used to estimate Se and Sp in the general population. The relative proportions of each group for HIV-negative individuals were derived from the 2020 French population census, which was conducted by the National Institute of Statistics and Economic Studies ( Institut national de la statistique et des études économiques , INSEE) [ 24 ]. The ANRS-CO4 FHDH cohort report provided data for HIV-positive participants, encompassing the surveillance of approximately 110,000 PLHIV over the course of the same year (see Additional file 4). The Se and Sp in the general population were then defined as follows: Subsequently, the positive predictive value (PPV) and negative predictive value (NPV) were estimated using Bayes’ formula. This calculation employed the posterior predictive distributions of Se and Sp along with the estimated prevalence of diagnosed HIV infection in France (p). The prevalence of diagnosed infections was calculated using UNAIDS estimates of infection prevalence, adjusted by the estimated percentage of diagnosed individuals within the population [ 25 , 26 ]. We integrated this estimate into our model using a beta distribution with a mean of 2.55.10 − 3 and a variance of 3.5.10 − 8 . Ability to identify newly diagnosed infections The performance of the algorithms in identifying newly diagnosed infections was estimated for the years 2016 to 2020 using data from the gold standard, selecting (i) PLHIV diagnosed between January 1, 2016, and December 31, 2020, and (ii) subjects for whom confirmation of seronegativity was available after December 31, 2016 (newly diagnosed infections specific inclusion criteria). The G11 and ValORIS algorithms, limited to a single year, were employed annually on the SNDS data of all the subjects from 2016 to 2020. HIV-negative individuals were censored from the year of their most recently confirmed HIV-negative status. For those newly diagnosed with HIV, censoring was implemented in the year following the diagnosis. The Se and Sp parameters of the algorithms were calculated according to the previous methodology, with one key modification: a random intercept was added for each participant to account for multiple measurements from the same individual. The estimation of Sp for the general population used INSEE census data. Regarding Se, data from the French National Public Health Agency ( Santé publique France) from the specified period were employed to determine the proportions of newly diagnosed patients across different groups of sex and country of origin [ 27 ]. The PPV and NPV were estimated as previously described. Assuming an annual number of HIV-positive discoveries between 4,000 and 7,000, the probability of being diagnosed a given year was integrated as a beta distribution with a mean of 8.10 − 5 and a variance of 1.2.10 − 10 . Epidemiology indicators: estimation of the prevalent HIV diagnosed population and the annual number of new HIV diagnoses in the SNDS Eventually, the aforementioned performance parameters were used to adjust the estimates of the prevalent diagnosed population and annual number of seropositive discoveries at the national scale. For the prevalent population, these two algorithms were applied to the entire SNDS database to determine the number of PLHIV diagnosed before the end of 2021. For each algorithm, the adjusted estimator was then calculated by multiplying the number of cases identified by the PPV/Se ratio. For the number of newly identified HIV infections each year, the population of prevalent cases in 2016 was first subtracted by applying the two algorithms to all SNDS data for the period 2012–2016. The annual numbers of HIV-positive discoveries for the period 2017–2023 were then determined by applying the two restricted algorithms to each year and subtracting the individuals identified from one year to the next. For each year between 2017 and 2023, the adjusted number of seropositive discoveries was calculated using the same method as that used for the prevalent population. Sensitivity analysis This study included HIV-negative participants who had been tested for HIV. We hypothesised that HIV testing might be linked to other factors, such as the prescription of antiretroviral therapy for post-exposure prophylaxis or molecular testing for HIV. This may be the case in situations involving high-risk sexual contact or for donors of products of human origin. This theoretical association implies that people who opt for HIV screening might have a higher likelihood of being identified as seropositive by SNDS targeting algorithms. As the inclusion criterion for HIV-negative subjects in our cohort was a negative HIV serology test, but HIV screening is not frequent among the general population (7% of the French population underwent an HIV test in 2021 [ 28 ]), we anticipated that the false-positive rate in our cohort might be higher than that in the general population. To assess potential selection bias in our study population, we implemented a sensitivity analysis by reducing the weighting of individuals classified as false positives. Initially, all subjects in the regression models were assigned a relative weight of 1. The sensitivity analysis involved repeating these models while progressively applying weights of 0.75, 0.5 and 0.25 to false-positive individuals. Results Gold standard population The validation cohort enabled the assessment of the HIV status in 109,216 individuals (2,123 PLHIV and 107,093 HIV-negative individuals). Owing to the inability to match SNDS data for 55 PLHIV and 3,009 HIV-negative individuals, SNDS information was retrieved for 2,068 PLHIV and 104,084 HIV-negative subjects (see Additional file 5 for a comparison of characteristics between matched and unmatched individuals). The final study group included 1,861 PLHIV and 98,105 HIV-negative individuals after excluding deaths prior to 2016 and individuals without any healthcare reimbursement during the study period (Fig. 1 ). Fig. 1. Open in a new tab Flowchart of the study inclusion. CDC. Clinical data warehouse; ID, Individuals ; PMSI, Programme de Médicalisation des Systèmes d’Information, (Hospital discharge summary); SNDS, Système National des Données de Santé, (French national health insurance information system) The HIV-negative subjects included in the cohort were predominantly female (66%), with a median age of 40.1 years as of 2020, and were mostly born in France (85.7%) (Table 1 ). In contrast, PLHIV were mainly men (65.2%), with a median age of 50.2 years. The majority were French born (66.8%) and had received a diagnosis at a median of 14.7 years prior (see additional file 6 for a comparison of the characteristics of PLHIV included in this study and those provided by the Regional Coordination Committees for the Fight against Human Immunodeficiency Virus and STIs throughout France). Table 1. Cohort characteristics Women, N (%) PLHIV N = 1,861 HIV-negative N = 98,105 648 (34.8%) 64 729 (66%) Age in 2020, median [Q1-Q3] 50.3 [41-57.9] 40.1 [31.7–58.1] Deceased 2016–2020 93 4935 Born abroad N (%) 617 (33.2%) 14 062 (14.3%) Time since diagnosis (years), median [Q1-Q3] 14.7 [8-23.4] Diagnosis after 2020 29 HIV transmission Heterosexual 680 (51.6%) MSM 502 (38.1%) Injection drug use 69 (5.2%) Transfusion / AEB 42 (3.2%) Perinatal 25 (1.9%) unknown 543 Open in a new tab AEB accidental exposure to blood, PLHIV People living with HIV, MSM men who have sex with men Evaluation of diagnosed HIV prevalence algorithms After applying the prevalent population specific inclusion criteria, the validation sample for the prevalent population included 1,710 PLHIV and 21,384 seronegative participants. The algorithm classification results are presented in a contingency table (Table 2 ). Table 2. Classification results for the HIV prevalent population G11 algorithm ValORIS algorithm positive negative total positive negative total CDC gold standard PLHIV 1,701 9 1,710 1,696 14 1,710 Seronegative 32 21,352 21,384 2 21,382 21,384 Total 1,733 21,361 23,094 1,698 21,396 23,094 Open in a new tab CDC Clinical Data warehouse, PLHIV People living with HIV By estimating the probability of an individual being identified as HIV-positive by the algorithms combined with post-stratification according to sex and country of origin, we were able to estimate the Se and Sp values of these algorithms within the general population. For the G11 algorithm, Se was estimated as 0.9933 (95% credible interval [0.9880–0.9972]). For the ValORIS algorithm, Se was estimated to be 0.9907 [0.9855–0.9954]. The Sp estimate fluctuated depending on the weighting assigned to each participant in the analysis as it was influenced by the number of false positives. When no weight modification was applied, the G11 algorithm’s 1-Sp was calculated to be 1.57.10 − 3 [1.06.10 − 3 -2.15.10 − 3 ], whereas the ValORIS algorithm’s 1-Sp was determined to be 1.57.10 − 4 [1.97.10 − 5 -3.46.10 − 4 ]. Reducing the weight of false positives led to an increase in Sp, particularly for the G11 algorithm (see Additional file 7). Initially, the predictive values were calculated from models in which the weights of false-positive cases remained unaltered. For the G11 algorithm, the PPV was calculated to be 0.6169 [0.5301–0.7058], while the 1-NPV was 1.72.10 − 5 [6.99.10 − 6 -3.11.10 − 5 ]. Regarding the ValORIS algorithm, the PPV was estimated to be 0.9415 [0.8768–0.9904], with a 1-NPV of 2.36.10 − 5 [1.18.10 − 5 -3.82.10 − 5 ]. A notable enhancement in the PPV of the G11 algorithm was observed when the weighting for false-positive individuals was reduced. The algorithm’s PPV rose to 0.7593 [0.6678–0.8456] when these subjects were allocated relative weights of 0.5 (see Additional file 8). The number of people identified by the G11 and ValORIS algorithms in 2020 based on the full SNDS data were 176,702 and 150,375, respectively (see Additional file 9). Using PPV and Se estimates obtained from our cohort, the unweighted models yielded significantly different outcomes, with the G11 algorithm estimating a population of 110,061 individuals [94,609 − 125,963] and the ValORIS algorithm estimating 143,165 individuals [133,171 − 150,520] (Fig. 2 ). This disparity decreased considerably when lowering the weight assigned to false-positive subjects. Consequently, utilising a relative weight of 0.5, the estimate increased to 135,448 individuals [119,072–150,836] for the G11 algorithm and to 145,704 individuals [136,920 − 151,401] for the ValORIS algorithm. Fig. 2. Open in a new tab Estimated number of diagnosed PLHIV identifiable in the SNDS by each algorithm Evaluation of algorithms for the newly diagnosed infections The application of the newly diagnosed infections specific inclusion criteria yielded a sample of 278 PLHIV and 54,941 seronegative individuals. The G11 algorithm demonstrated a Se of 0.9526 [0.9206–0.9811], while the ValORIS algorithm showed a lower Se of 0.8240 [0.7584–0.8891]. For Sp, the G11 algorithm estimation of 1-Sp reached 5.60.10 − 4 [3.04.10 − 4 -7.02.10 − 4 ] and the ValORIS algorithm attained a value of 1-Sp equal to 1.45.10 − 5 [3.03.10 − 6 -4.10.10 − 5 ]. A slight enhancement in Sp was observed, particularly for the G11 algorithm, when the impact of false positives was diminished (see Additional file 10). Based on these estimators and the estimated incidence of HIV diagnosis in France, we determined the predictive value of the algorithms. For the G11 algorithm, the PPV was calculated to be 0.1191 [0.0787–0.1851], while the 1-NPV was equal to 3.7.10 − 6 [1.2.10 − 6 -6.5.10 − 6 ]. Regarding the ValORIS algorithm, the PPV was estimated to be 0.8221 [0.6357–0.9701], with a 1-NPV of 1.1.10 − 5 [6.1.10 − 6 -1.7.10 − 5 ]. Again, a significant improvement in the PPV of the G11 algorithm was noted when the weighting for false-positive individuals decreased. When these subjects were assigned relative weights of 0.5, the algorithm’s PPV increased to 0.1993 [0.1416–0.2642] (see Additional file 11). Eventually, these estimators were used to adjust the annual number of newly diagnosed PLHIV identified in the SNDS comprehensive database (see Additional file 12). These estimates were compared with those obtained by Santé Publique France from mandatory notifications for seropositive diagnoses (Fig. 3 ). From 2017 to 2023, the G11 algorithm estimated 16,335 new diagnoses [10,647 − 25,058], whereas the ValORIS algorithm projected 43,082 [32,8234-51,562]. This discrepancy decreased as the impact of false positives decreased. Upon adjusting the relative weight of these individuals to 0.5, the G11 algorithm’s prediction increased to 27,283 diagnoses [19,282 − 35,909], and the ValORIS algorithm’s estimate reached 45,346 [35,409 − 52,795]. For reference, Santé publique France reported 40,912 new diagnoses (95% confidence interval: [39,893 − 41,931]) during the same timeframe. Fig. 3. Open in a new tab Estimated annual number of newly diagnosed PLHIV identifiable in the SNDS by each algorithm Characteristics of individuals misclassified by the algorithms Analysing the outcomes of algorithmic classifications allowed to discern the potential factors leading to misclassification. Cohort members were categorised based on their serostatus and whether any algorithm misclassified them at least once (Table 3 ). This classification suggests that among the diagnosed PLHIV, women and persons born abroad may have a higher likelihood of being undetected by the algorithms. As one might expect, the probability of correct identification seemed to increase with the time elapsed between the diagnosis date and the year in which the algorithm was applied. Among HIV-negative individuals, men and those with foreign births appeared to be more prone to incorrect identification as HIV-positive. Table 3. Characteristic of individuals according to their algorithmic classification outcome PLHIV N = 1753 HIV-negative N = 54,943 True positives N = 1698 False negatives N = 55 True negatives N = 54, 830 False positives N = 113 Women, N (%) 593 (34.9%) 28 (50.9%) 35,570 (64.9%) 60 (53.1%) Age in 2020, median [Q1-Q3] 50.7 [41.4–58.2] 37.5 [29.7–44.5] 37.8 [29.3–60.0] 37.6 [31.0–55.0] Deceased 2016–2020 4 2 1878 3 Born abroad N (%) 551 (32.4%) 44 (80.0%) 8,035 (14.7%) 23 (20.4%) Time since diagnosis (years), median [Q1-Q3] 14.7 [8.0-23.2] 3.2 [1.6–4.5] HIV transmission Heterosexual 637 (51.2%) 26 (76.5%) MSM 477 (38.3%) 8 (23.5%) Injection drug use 64 (5.1%) Transfusion 41 (3.3%) Perinatal 25 (2%) unknown 454 21 Open in a new tab PLHIV People living with HIV, MSM men who have sex with men Of the 113 instances in which individuals were incorrectly identified as seropositive by any algorithm, 90 were attributed to the reimbursement of viral RNA quantification. The rationale for this biological test was ascertained for 65 individuals, with the primary reasons being evaluations prior to transplantation (36 instances), examinations related to reproductive medicine (15 instances), and exposure to high-risk sexual activities (six instances). Other instances of incorrect HIV-positive classification involved erroneous HIV disease diagnosis codes applied to hospitalisations (seven instances) and chronic illnesses (four instances), and multiple reimbursements for antiretroviral drugs (six cases related to consecutive post-exposure treatments). Among the 55 PLWH incorrectly identified as negative by one of the algorithms, 45 were diagnosed or came to France between 2016 and 2019. These individuals were mainly misclassified by the ValORIS algorithm targeting new HIV diagnoses due to its higher specificity. Within this group of recently diagnosed individuals, 32 had an algorithmic classification schift to positive in the year after the diagnosis. Among the remaining 10 individuals who received their diagnosis prior to 2016 and were misclassified as false negative, only two had no evidence of accessing HIV-related care in the SNDS. The other eight subjects appear to be people who have dropped out of care, since only one of them received antiretroviral drugs in 2020. Discussion The primary interest of this study lies in examining the prevalence and incidence of diagnosed HIV in France with a new method based on medico-administrative databases. This approach to HIV epidemiology in France is potentially valuable, as it can be readily automated, requires minimal human resources, and can generate estimates in accordance with the update frequency of SNDS databases. This validation study, which focused on algorithms for identifying HIV-diagnosed individuals, represents one of the initial endeavours to evaluate the ability of French medico-administrative data to recognise populations affected by chronic infections. It stands alongside the research conducted by Lam et al.. on hepatitis C virus as a pioneering effort in this field [ 29 ]. The findings of this investigation demonstrated that, despite their relative simplicity, the assessed algorithms exhibited noteworthy performance metrics. Notably, for HIV cohort or case-control research, it is crucial to employ algorithms with near-perfect Sp to ensure an acceptable PPV, thus yielding a population of HIV-positive individuals with a limited number of false positives. Given the low prevalence of HIV infection in the general population, the proportion of false positives among targeted individuals can quickly become significant, leading to a substantial decrease in the PPV of the algorithm. In the context of targeting the prevalent population of diagnosed PLHIV, this study revealed that a slight difference in Sp (approximately 1.10 − 3 ) between the two algorithms could result in a considerable PPV variation of approximately 0.3. Considering that SNDS targeting algorithms are predominantly used to define an individual’s status, it becomes evident that users are more concerned with the algorithms’ PPV and NPV rather than their Se and Sp. Consequently, we emphasise the importance of algorithm validation studies that report not only Se and Sp, but also provide estimates of post-test probabilities. In cases where the study design precludes the creation of a gold standard with a representative proportion of affected subjects, researchers should employ already published prevalence estimates to determine the predictive value of the algorithms under evaluation. However, selecting a more specific algorithm leads to a lower Se, as observed in this study, owing to the increased number of criteria applied. One can readily surmise that this reduction in Se will primarily affect HIV-positive individuals with limited healthcare utilisation, who are most likely to be incorrectly classified as seronegative. Presuming that these falsely classified seronegative individuals possess distinct characteristics and, more importantly, different outcomes than those included in the study, opting for a more specific algorithm will still introduce bias by selecting a non-representative sample of PLHIV. In this instance, a quantitative bias analysis based solely on the algorithm’s performance metrics would be insufficient to rectify the bias associated with the non-random misclassification of PLHIV. Consequently, it is essential to establish a priori the attributes and/or outcomes of PLHIV according to the classification results of the algorithm. However, this approach is not feasible without linking the SNDS to a gold standard. Any country with a consolidated national medico-administrative database may validate its use as an additional tool in the HIV surveillance system, as performed in Canada [ 30 – 33 ] The information contained within these Canadian medico administrative databases differs from SNDS, necessitating partially divergent algorithmic strategies. Indeed, Canadian databases allow the use of diagnostic codes in medical consultations, whereas the French SNDS enables the identification of outpatient medical biology tests. Despite these disparities, it is noteworthy that algorithms based on Canadian information, when identifying prevalent populations of diagnosed PLHIV, yielded Sp estimates similar to those observed in this study, consistently surpassing 0.99. Nevertheless, these Canadian-based algorithms exhibited lower Se compared to those utilised in the SNDS, ranging from 0.8 to 0.97 according to the various algorithms examined. This study also highlighted the difficulty of determining an HIV-negative reference group. Regarding both the existing number of diagnosed PLHIV and the number of new seropositive cases identified, we observed that the projections derived from each targeting algorithm, adjusted according to the algorithms’ performance metrics, showed significant differences. One initial hypothesis is that these differences could be due to a selection bias among the PLHIV included in this study. However, comparing their characteristics with those of PLHIV in the ANRS CO3 and ANRS CO4 cohorts across France shows that our sample is representative of all the characteristics we studied, except for the proportion of PLHIV born abroad (33% in our cohort compared to 42% nationally). Nevertheless, we hypothesise that this difference in proportion is less attributable to an underrepresentation of people born abroad in our cohort than to a classification bias in the SNDS. This is because the geographical origin of individuals is not a variable directly recorded in the SNDS, but must be deduced from other variables, such as the allocation of state medical aid, the encoding of social security numbers and the date on which healthcare consumption commenced. Therefore, it is highly likely that a significant proportion of individuals born abroad could not be identified as such in this study and were misclassified as individuals born in France. This hypothesis is supported by the proportion of women and the distribution of HIV transmission groups, which are very similar in our cohort and in the national data. Therefore, it seems much more likely that the discrepancy in estimates between the two algorithms is primarily attributed to the bias in measuring their Sp. Quantification of HIV RNA was the most prevalent element found in false-positive subjects. The G11 algorithm was noticeably more impacted because only one criterion was required for the HIV-positive classification. Our modified algorithm, ValORIS, used a combination of two criteria and therefore obtained better results. It was further noted that the differences in the estimates tended to diminish when the weight assigned to false positives was reduced, thereby increasing the Sp of the algorithms. This over-representation of false-positive individuals in our sample is most likely attributable to the selection of HIV-negative individuals based on negative HIV serological results. Consequently, it is hypothesised that, in HIV-negative subjects, the identification of negative serology increases the probability of observing care consumption related to the follow-up of PLHIV. Thus, if the constitution of a sample of HIV-negative subjects based on real-life serologies induces a bias due to the non-representativeness of these individuals among the population of HIV-negative people, the only alternative could have been to conduct this screening test within the framework of the study in a group of subjects identified as blind to their history of HIV screening practice. The use of SNDS for the epidemiological study of HIV also has certain limitations related to the nature of this medico-administrative database. First, it is important to note that SNDS essentially contains information on individuals’ healthcare consumption. Consequently, behavioural information, particularly that related to the modes of transmission, is unavailable. Similarly, HIV infection can only be dated based on biological tests, the results of which are not recorded in SNDS databases [ 34 , 35 ]. Because this information is crucial for characterising the evolution of the HIV epidemic and defining screening and prevention strategies, it would be unfeasible to restrict HIV epidemiological monitoring to SNDS data without profoundly enriching the information it contains. A more problematic aspect is the sustainability of the algorithmic performance indicators presented in this study. These targeting algorithms are based on data related to the care consumption of PLHIV, which evolve over time in accordance with the emergence of new therapeutic strategies and management recommendations. This study has shown that the ValORIS algorithm can be used to estimate the number of discovered HIV-positive cases, with a confidence interval that largely overlaps with estimates from the French National Public Health Agency. However, it should be noted that this convergence of estimates appears to be increasing with time, which suggests that the algorithm’s performance may be varying over time. While we cannot determine the cause of this change, we can hypothesise that subtle changes in medical practices or individual profiles, may have led to variations in the algorithm’s performance over time. Evidently, the introduction of long-acting injectable antiretrovirals, such as the cabotegravir-rilpivirne combination available in France in late 2021, affects the sensitivity of an algorithm that does not differentiate between the various treatments available, either in terms of type or frequency of administration. The inherent nature of the targeting algorithms used in SNDS necessitates their update in line with changes in treatment, and more importantly, requires regular re-estimation of their performance parameters. It is clear, therefore, that future developments in targeting algorithms within the SNDS should be based on sophisticated methods that allow for the detailed consideration of the longitudinal nature of medical and administrative data. This longitudinal dimension must be considered on both a global timescale, to account for the aforementioned changes, and an individual timescale, to capture the information contained within intervals of healthcare consumption. This information will undoubtedly be useful in distinguishing the underlying causes of consumption. To this end, supervised learning assembly methods such as random forests or XGBOOST, in which features are constructed by nesting healthcare consumption within time windows, are relatively easy to implement and have been extensively tested [ 36 ]. Although these methods still compete with deep learning methods on tabular data, the latter’s rapid evolution now makes it possible to consider the temporality of data more precisely than by simply defining time windows. Therefore, deep learning methods that use time-aware transformers to explicitly model time intervals between healthcare consumption could be particularly useful for medical-administrative databases [ 37 , 38 ]. In an era characterised by the proliferation and increasing enhancement of data sources, another significant challenge for contemporary epidemiology is the development of integrated surveillance systems that draw data from diverse information sources [ 39 , 40 ]. These monitoring systems will encounter numerous technical and methodological challenges, particularly those pertaining to data standardisation, integration, quality, and volume [ 19 , 20 ]. From an ethical perspective, the establishment of integrated surveillance systems inevitably raises questions regarding data security, anonymity, ownership, and citizen information. Despite the magnitude of these challenges, the potential benefits offered by the enhancement of our surveillance systems should motivate researchers to address these issues [ 39 , 40 ]. France has the advantage of possessing a national medico-administrative database that will undoubtedly enrich surveillance data, particularly in terms of monitoring PLHIV care pathways. However, SNDS is not the only source of information that could be valuable for monitoring HIV and other infectious diseases. Recent years have witnessed a proliferation of healthcare data warehouses in France, including hospital data warehouses [ 20 , 41 ], as well as those for medical biology, general medicine, and pharmacy laboratories, in addition to thematic data warehouses such as that of the DAT’AIDS cohort dedicated to PLHIV. Conclusion This study highlights the importance of validating SNDS-based algorithms and tacking into account their performance parameters for HIV surveillance. Although these algorithms may be useful for automated monitoring of HIV epidemiology in France, they have limitations in capturing detailed clinical and behavioural data. Future research should focus on defining more sophisticated algorithms to improve their performance and sustainability. Ultimately, SNDS-based surveillance may serve as a valuable complementary tool to the existing HIV monitoring systems in France. Supplementary Information 12874_2026_2820_MOESM1_ESM.docx (2.6MB, docx) Supplementary Material 1: (S1-S5 ; S7-S12). Supplementary Material 2: (S6). (17.8KB, xlsx) Acknowledgements The authors express their sincere gratitude to Florent Lot, Françoise Cazein, and Amber Kunkel of Santé publique France for their insightful discussions and invaluable guidance. Additionally, we extend our appreciation to Marjorie Boussac and the entire team at the Assurance Maladie DEMEX unit for their efforts to match and provide access to SNDS data. Abbreviations CDC Clinical Data Centre HIV Human Immunodeficiency Virus NPV Negative Predictive Value PLHIV People Living with HIV PPV Positive Predictive Value Se Sensitivity Sp Specificity SNDS Système National des Données de Santé Authors’ contributions MFT, LG, and KS designed this study. MFT extracted the data, performed statistical analysis, and wrote the first draft. All authors read and approved the final manuscript. Funding The authors declare that no funds, grants, or other support were received during the preparation of this manuscript. Data availability Data in this study were obtained from two sources. Data from the CDC of the CHRU of Tours can be made accessible after agreement with the structure producing the data. Data from the SNDS can be obtained after agreement with the CNIL and the Assurance Maladie. Declarations Ethics approval and consent to participate This study was approved by the French Data Protection Board (Commission Nationale de l’Informatique et des Libertés, CNIL authorisation DR-2022-066). This approval covered access to de-identified data in the SNDS database and its matching with data from the CHRU de Tours CDC. Consent for publication Not applicable. Competing interests The authors declare no competing interests. Footnotes Publisher’s note Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations. References 1. Nelson KE, Williams CM. Infectious disease epidemiology: theory and practice. 3rd ed. Jones & Bartlett. 2014. 2. Rebeiro PF, Duda SN, Wools-Kaloustian KK, Nash D, Althoff KN. Implications of COVID‐19 for HIV Research: data sources, indicators and longitudinal analyses. J Int AIDS Soc. 2020;23:e25627. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 3. European Centre for Disease Prevention and Contro. The impact of the COVID-19 pandemic on the HIV response in Europe and Central Asia - monitoring implementation of the Dublin declaration on partnership to fight HIV/AIDS in Europe and Central Asia. 2024. 10.2900/438904. 4. SPF. Bulletin de santé publique VIH-IST. Décembre 2020. https://www.santepubliquefrance.fr/maladies-et-traumatismes/infections-sexuellement-transmissibles/vih-sida/documents/bulletin-national/bulletin-de-sante-publique-vih-ist.-decembre-2020 . Accessed 13 Aug 2024. 5. Fonteneau L, Le Meur N, Cohen-Akenine A, Pessel C, Brouard C, Delon F, et al. [The use of administrative health databases in infectious disease epidemiology and public health]. Rev Epidemiol Sante Publique. 2017;65(Suppl 4):S174–82. [ DOI ] [ PubMed ] [ Google Scholar ] 6. Tassi M-F, le Meur N, Stéfic K, Grammatico-Guillon L. Performance of French medico-administrative databases in epidemiology of infectious diseases: a scoping review. Front Public Health. 2023;11:1161550. [ DOI ] [ PMC free article ] [ PubMed ] 7. Hassen-Khodja C, Gras G, Grammatico-Guillon L, Dupuy C, Gomez J-F, Freslon L, et al. Hospital and ambulatory management, and compliance to treatment in HIV infection: regional health insurance agency analysis. Med Mal Infect. 2014;44:423–8. [ DOI ] [ PubMed ] [ Google Scholar ] 8. Guillon A, Laurent E, Duclos A, Godillon L, Dequin P-F, Agrinier N, et al. Case fatality inequalities of critically ill COVID-19 patients according to patient-, hospital- and region-related factors: a French nationwide study. Ann Intensive Care. 2021;11:127. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 9. Bezin J, Duong M, Lassalle R, Droz C, Pariente A, Blin P, et al. The national healthcare system claims databases in France, SNIIRAM and EGB: Powerful tools for pharmacoepidemiology. Pharmacoepidemiol Drug Saf. 2017;26:954–62. [ DOI ] [ PubMed ] [ Google Scholar ] 10. Tuppin P, Rudant J, Constantinou P, Gastaldi-Ménager C, Rachas A, de Roquefeuil L et al. Value of a national administrative database to guide public decisions: From the Système National d’information Interrégimes de l’Assurance Maladie (SNIIRAM) to the Système National des Données de Santé (SNDS) in France. Revue d’Épidémiologie et de Santé Publique. 2017;65:S149–67. [ DOI ] [ PubMed ] 11. Rachas A, Gastaldi-Ménager C, Denis P, Barthélémy P, Constantinou P, Drouin J, et al. The Economic Burden of Disease in France From the National Health Insurance Perspective: The Healthcare Expenditures and Conditions Mapping Used to Prepare the French Social Security Funding Act and the Public Health Act. Med Care. 2022;60:655. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 12. Lorgis L, Cottenet J, Molins G, Benzenine E, Zeller M, Aube H, et al. Outcomes after acute myocardial infarction in HIV-infected patients: analysis of data from a French nationwide hospital medical information database. Circulation. 2013;127:1767–74. [ DOI ] [ PubMed ] [ Google Scholar ] 13. Funk MJ, Landi SN. Misclassification in administrative claims data: quantifying the impact on treatment effect estimates. Curr Epidemiol Rep. 2014;1:175–85. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 14. Fonteneau L, Le Meur N, Cohen-Akenine A, Pessel C, Brouard C, Delon F et al. Apport des bases médico-administratives en épidémiologie et santé publique des maladies infectieuses. Revue d’Épidémiologie et de Santé Publique. 2017;65:S174–82. [ DOI ] [ PubMed ] 15. Goldberg M, Carton M, Doussin A, Fagot-Campagna A, Heyndrickx E, Lemaitre M et al. Le réseau REDSIAM (Réseau données Sniiram) – Spécial REDSIAM. Revue d’Épidémiologie et de Santé Publique. 2017;65:S144–8. [ DOI ] [ PubMed ] 16. Benchimol EI, Manuel DG, To T, Griffiths AM, Rabeneck L, Guttmann A. Development and use of reporting guidelines for assessing the quality of validation studies of health administrative data. J Clin Epidemiol. 2011;64:821–9. [ DOI ] [ PubMed ] [ Google Scholar ] 17. Cartographie des pathologies et des dépenses de l’Assurance Maladie | L’Assurance Maladie. https://www.assurance-maladie.ameli.fr/etudes-et-donnees/par-theme/pathologies/cartographie-assurance-maladie . Accessed 22 Aug 2024. 18. Cuggia M, Combes S. Yearb Med Inf. 2019;28:195–202. The French Health Data Hub and the German Medical Informatics Initiatives: Two National Projects to Promote Data Sharing in Healthcare. [ DOI ] [ PMC free article ] [ PubMed ] 19. Cuggia M, Bayat S, Garcelon N, Sanders L, Rouget F, Coursin A, et al. A full-text information retrieval system for an epidemiological registry. Stud Health Technol Inf. 2010;160:491–5. [ PubMed ] [ Google Scholar ] 20. Cuggia M, Guillon-Grammatico L, Gourraud P-A, Chretien J-M, Bocquet F, Guiton V, et al. In: Mantas J, Hasman A, Demiris G, Saranto K, Marschollek M, Arvanitis TN, et al. editors. Studies in Health Technology and Informatics. IOS; 2024. The Ouest Data Hub: An Interregional Health Data Sharing Ecosystem for Research. [ DOI ] [ PubMed ] 21. Coughlin SS, Trock B, Criqui MH, Pickle LW, Browner D, Tefft MC. The logistic modeling of sensitivity, specificity, and predictive value of a diagnostic test. J Clin Epidemiol. 1992;45:1–7. [ DOI ] [ PubMed ] [ Google Scholar ] 22. Mulherin SA, Miller WC. Spectrum Bias or Spectrum Effect? Subgroup Variation in Diagnostic Test Evaluation. Ann Intern Med. 2002;137:598–602. [ DOI ] [ PubMed ] [ Google Scholar ] 23. Kennedy L, Gelman A. Know your population and know your model: Using model-based regression and post-stratification to generalize findings beyond the observed sample. Psychol Methods. 2021;26:547–58. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 24. IMG1A -. Population par sexe, âge et situation quant à l’immigration en 2020 – Recensement de la population. INSEE. https://www.insee.fr/fr/statistiques/7633125?sommaire=7633727&geo=FE-1#consulter-sommaire . Accessed 16 Jan 2025. 25. AIDSinfo | UNAIDS. https://aidsinfo.unaids.org/ . Accessed 12 Dec 2024. 26. Vourli G, Noori T, Pharris A, Porter K, Axelsson M, Begovac J, et al. Human Immunodeficiency Virus Continuum of Care in 11 European Union Countries at the End of 2016 Overall and by Key Population: Have We Made Progress? Clin Infect Dis. 2020;71:2905–16. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 27. VIH/sida. https://www.santepubliquefrance.fr/maladies-et-traumatismes/infections-sexuellement-transmissibles/vih-sida . Accessed 16 Jan 2025. 28. Surveillance du VIH et. des IST bactériennes en France en 2023. Saint-Maurice: Santé publique France; 2024. [ Google Scholar ] 29. Lam L, Fontaine H, Lapidus N, Bellet J, Lusivika-Nzinga C, Nicol J, et al. Performance of algorithms for identifying patients with chronic hepatitis B or C infection in the french health insurance claims databases using the ANRS CO22 HEPATHER cohort. J Viral Hepatitis. 2023;30:232–41. [ DOI ] [ PubMed ] [ Google Scholar ] 30. Anderson AJ, Lix L, Loeppky C, Van Caeseele P, Queenan JA, Mahar AL. Validation of algorithms to identify human immunodeficiency virus cases using administrative data in Manitoba. Can J Public Health. 2025;116:124–35. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 31. Nosyk B, Colley G, Yip B, Chan K, Heath K, Lima VD, et al. Application and Validation of Case-Finding Algorithms for Identifying Individuals with Human Immunodeficiency Virus from Administrative Data in British Columbia, Canada. PLoS ONE. 2013;8:e54416. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 32. Emerson SD, McLinden T, Sereda P, Lima VD, Hogg RS, Kooij KW, et al. Identification of people with low prevalence diseases in administrative healthcare records: A case study of HIV in British Columbia, Canada. PLoS ONE. 2023;18:e0290777. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 33. Antoniou T, Zagorski B, Loutfy MR, Strike C, Glazier RH. Validation of Case-Finding Algorithms Derived from Administrative Data for Identifying Adults Living with Human Immunodeficiency Virus Infection. PLoS ONE. 2011;6:e21748. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 34. Barin F, Meyer L, Lancar R, Deveau C, Gharib M, Laporte A, et al. Development and validation of an immunoassay for identification of recent human immunodeficiency virus type 1 infections and its use on dried serum spots. J Clin Microbiol. 2005;43:4441–7. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 35. Le Vu S, Meyer L, Cazein F, Pillonel J, Semaille C, Barin F, et al. Performance of an immunoassay at detecting recent infection among reported HIV diagnoses. AIDS. 2009;23:1773–9. [ DOI ] [ PubMed ] [ Google Scholar ] 36. Romero A, Deypalan RAY, Mehrotra MN, Jungao S, Sheils JT, Manduchi NE. Benchmarking AutoML frameworks for disease prediction using medical claims. BioData Min. 2022;15:15. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 37. An Y, Liu Y, Chen X, Sheng Y. TERTIAN: Clinical Endpoint Prediction in ICU via Time-Aware Transformer-Based Hierarchical Attention Network. Comput Intell Neurosci. 2022;2022:4207940. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 38. Yu J, Feng Z, Lu J, Cai T, Zhou D. Time-aware attention for enhanced electronic health records modeling. arXiv preprint arXiv:2507.14847, 2025 . https://arxiv.org/abs/2507.14847 39. Bansal S, Chowell G, Simonsen L, Vespignani A, Viboud C. Big Data for Infectious Disease Surveillance and Modeling. J Infect Dis. 2016;214 suppl4:S375–9. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 40. Simonsen L, Gog JR, Olson D, Viboud C. Infectious Disease Surveillance in the Big Data Era: Towards Faster and Locally Relevant Systems. J Infect Dis. 2016;214 suppl4:S380–5. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 41. Ansoborlo M, Salpétrier C, Nail L-RL, Herbet J, Cuggia M, Rosset P, et al. Feasibility of automated surveillance of implantable devices in orthopaedics via clinical data warehouse: the Studio study. BMC Med Inf Decis Mak. 2024;24:324. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials 12874_2026_2820_MOESM1_ESM.docx (2.6MB, docx) Supplementary Material 1: (S1-S5 ; S7-S12). Supplementary Material 2: (S6). (17.8KB, xlsx) Data Availability Statement Data in this study were obtained from two sources. Data from the CDC of the CHRU of Tours can be made accessible after agreement with the structure producing the data. Data from the SNDS can be obtained after agreement with the CNIL and the Assurance Maladie. Articles from BMC Medical Research Methodology are provided here courtesy of BMC ACTIONS View on publisher site PDF (2.4 MB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top