Skip to main content An official website of the United States government Here's how you know Here's how you know Official websites use .gov A .gov website belongs to an official government organization in the United States. Secure .gov websites use HTTPS A lock ( Lock Locked padlock icon ) or https:// means you've safely connected to the .gov website. Share sensitive information only on official, secure websites. Search Log in Dashboard Publications Account settings Log out Search… Search NCBI Primary site navigation Search Logged in as: Dashboard Publications Account settings Log in Search PMC Full-Text Archive Search in PMC Journal List User Guide PERMALINK Copy As a library, NLM provides access to scientific literature. Inclusion in an NLM database does not imply endorsement of, or agreement with, the contents by NLM or the National Institutes of Health. Learn more: PMC Disclaimer | PMC Copyright Notice J Am Med Inform Assoc . 2025 Oct 27;33(3):663–669. doi: 10.1093/jamia/ocaf172 Search in PMC Search in PubMed View in NLM Catalog Add to search A novel analysis methodology for assessment of re-identification risks for the National Cancer Institute cancer registry privacy preserving record linkage technique Murat Kantarcioglu Murat Kantarcioglu , PhD 1 Department of Computer Science, Virginia Tech, Blacksburg, VA 24061, United States Find articles by Murat Kantarcioglu 1, ✉ , Will Howe Will Howe , BS 2 Information Management Services, Inc., Calverton, MD 20705, United States Find articles by Will Howe 2 , Benmei Liu Benmei Liu , PhD 3 Surveillance Research Program, Division of Cancer Control & Population Sciences, National Cancer Institute, Bethesda, MD 20892, United States Find articles by Benmei Liu 3 , Valentina Petkov Valentina Petkov , MD, MPH 4 Surveillance Research Program, Division of Cancer Control & Population Sciences, National Cancer Institute, Bethesda, MD 20892, United States Find articles by Valentina Petkov 4 , Esmeralda Casas-Silva Esmeralda Casas-Silva , PhD 5 Center for Biomedical Informatics and Information Technology, National Cancer Institute, Bethesda, MD 20892, United States Find articles by Esmeralda Casas-Silva 5 , Diana Velasquez-Kolnik Diana Velasquez-Kolnik , MBA 6 Frederick National Laboratory for Cancer Research, Frederick, MD 21701, United States Find articles by Diana Velasquez-Kolnik 6 , Bradley A Malin Bradley A Malin , PhD 7 Department of Biomedical Informatics, Vanderbilt University Medical Center, Nashville, TN 37232, United States 8 Department of Computer Science, Vanderbilt University, Nashville, TN 37232, United States 9 Department of Biostatistics, Vanderbilt University Medical Center, Nashville, TN 37232, United States Find articles by Bradley A Malin 7, 8, 9 , Lynne Penberthy Lynne Penberthy , MD, PhD 10 Surveillance Research Program, Division of Cancer Control & Population Sciences, National Cancer Institute, Bethesda, MD 20892, United States Find articles by Lynne Penberthy 10 Author information Article notes Copyright and License information 1 Department of Computer Science, Virginia Tech, Blacksburg, VA 24061, United States 2 Information Management Services, Inc., Calverton, MD 20705, United States 3 Surveillance Research Program, Division of Cancer Control & Population Sciences, National Cancer Institute, Bethesda, MD 20892, United States 4 Surveillance Research Program, Division of Cancer Control & Population Sciences, National Cancer Institute, Bethesda, MD 20892, United States 5 Center for Biomedical Informatics and Information Technology, National Cancer Institute, Bethesda, MD 20892, United States 6 Frederick National Laboratory for Cancer Research, Frederick, MD 21701, United States 7 Department of Biomedical Informatics, Vanderbilt University Medical Center, Nashville, TN 37232, United States 8 Department of Computer Science, Vanderbilt University, Nashville, TN 37232, United States 9 Department of Biostatistics, Vanderbilt University Medical Center, Nashville, TN 37232, United States 10 Surveillance Research Program, Division of Cancer Control & Population Sciences, National Cancer Institute, Bethesda, MD 20892, United States ✉ Corresponding author: Murat Kantarcioglu, PhD, Department of Computer Science, Virginia Tech, 1160 Torgersen Hall, 620 Drillfield Drive, Blacksburg, VA 24061, United States ( [email protected] ) Received 2025 Mar 20; Revised 2025 Aug 6; Accepted 2025 Aug 10; Collection date 2026 Mar. Published by Oxford University Press on behalf of the American Medical Informatics Association 2025. This work is written by (a) US Government employee(s) and is in the public domain in the US. PMC Copyright notice PMCID: PMC12981660 PMID: 41144319 Abstract Objective The National Cancer Institute (NCI), part of the National Institutes of Health (NIH) supports efforts to address critical challenges in advancing cancer research. As part of this effort, NCI sponsored the development of a privacy-preserving record linkage (PPRL) software that transforms identifying patient information into multiple tokens through a set of cryptographically secure keyed hash functions. This project aims to evaluate the PPRL software in the perspective of re-identification risks and propose effective strategies to sufficiently mitigate these risks. Materials and Methods To achieve the goals, we developed a novel re-identification risk assessment framework, based on token frequency analysis, to estimate the privacy impact of hashed tokens shared for record linkage. We assessed privacy risk through empirical analysis on a state-level voter registration database, a public dataset commonly used for re-identification, under various scenarios. These scenarios are defined based on several factors, including the size of the dataset used for linkage and a group size parameter that determines when an adversary can claim that a record has been re-identified. Results We found that the re-identification risk based on frequency analysis attack is approximately 0.0002 (ie, 2 patients out of 10 000 are potentially identifiable) under reasonable adversarial settings, with a group size parameter of k = 12 and a dataset size of 400 000 patients. Additionally, our analysis reveals a negative correlation between dataset size and re-identification risk. Discussion Re-identification risk is deemed low for the new NCI PPRL software. Token frequency analysis provides a reliable estimate of the re-identification risk in token-based PPRL tools. Keywords: private record linkage, data integration, privacy Introduction The National Cancer Institute (NCI), part of the National Institutes of Health (NIH), supports efforts to tackle numerous challenges in discovering groundbreaking therapies for cancer by building better artificial intelligence (AI) and machine learning (ML) models. One factor complicating the development of AI/ML and analytic models for cancer research is the integration of patient data that is scattered across disparate healthcare organizations. This is particularly a problem for diseases such as cancer, where patients often receive care at different institutions within and across multiple geographic areas. As a result, studies based on a single healthcare organization will likely provide an inaccurate and incomplete picture. Moreover, this could occur in several ways. First, as has been shown without record linkage and deduplication, many conditions will be overrepresented. For instance, it was shown that deduplication across hospitals in the Chicago region led to a reduction in prevalence of 24.0%, 28.0%, and 10.9% of type 2 diabetes, asthma, and myocardial infarction, respectively. 1 Furthermore, there may be incompleteness in a patient’s record. For example, a cancer treatment may begin at one institution, while follow-up examinations and assessments of downstream outcomes take place at another, potentially in a different state. There have been various efforts to implement data integration across health care organizations and public health reporting systems such as central cancer registries. However, these efforts are limited by non-trivial security and privacy challenges. The Health Insurance Portability and Accountability Act of 1996 (HIPAA) outlines required procedures for protecting protected health information. And, in this respect, in certain situations, sharing unique IDs, such as their Social Security number (SSN), without a patient’s consent or a valid purpose permitted by the Privacy Rule, could be considered a violation of the HIPAA Privacy Rule. However, a patient’s information could be linked in a manner that meets the de-identification standard under the Privacy Rule. While various techniques have been developed, there remains a need for free-to-use, and privacy-preserving record linkage (PPRL) tools, particularly for use in non-treatment, payment, and healthcare operations (non-TPO) settings such as cancer registries. Over the past several decades, PPRL approaches have been proposed to enable data holders to integrate records without revealing the raw information (see Gkoulalas-Divanis et al 2 for a recent survey). One class is based upon secure multi-party computation (SMPC) (eg, Ben-David et al 3 ) which leverages cryptographically strong protocols. These protocols enable similarity comparisons to be performed without revealing any information, except for the output of the protocol. While secure in principle, computationally expensive encryption techniques, such as homomorphic encryption (eg, Bos et al 4 ) or secure circuit evaluation (eg, Ben-David et al 3 ) are required. These techniques do not scale well to very large databases, a problem for cancer registries that could contain large number of records in it. 5 Moreover, some SMPC solutions require multiple rounds of interaction, necessitating extensive data exchange among the involved parties. This increased computational and communication cost can make these approaches less practical for many real-world settings, where efficiency and scalability are critical constraints. As an alternative, a second class of PPRL, used in the software being evaluated under this work, utilizes a weaker form of security—one that is based on data transformation. These transformed values, referred to as encodings or tokens, are used in the PPRL process. For this class of methods, it is critical to strike a balance between the competing priorities of accuracy, security, and efficiency when identifying an appropriate data transformation method for PPRL. One such approach is based on keyed hash functions. 6 In this approach, each site participating in the record linkage protocol uses a shared secret key (also at times referred to as “salt”) to transform predetermined tokens through keyed hash functions. Two records are deemed to match if their hashed tokens are sufficiently similar according to a distance function and threshold. However, it has been shown that the naïve application of keyed hash functions can be easily attacked using frequency analysis. 6 PPRL techniques must also achieve sufficient accuracy to ensure that the quality of the linked data is adequate for downstream tasks, such as training improved machine learning models on integrated datasets. Consequently, evaluating the utility of a PPRL method is crucial. Interestingly, in some scenarios, PPRL techniques can achieve higher accuracy than traditional, non-privacy-preserving record linkage methods. For example, Durham et al 5 report that Bloom filter-based PPRL achieved higher true positive rates than a conventional Jaro-Winkler-based record linkage approach. This result is particularly notable because it demonstrates that effective privacy preservation does not necessarily require sacrificing linkage quality. However, it is equally important to recognize that if a PPRL method produces poor linkage results—even while protecting all sensitive information—the outcome is practically useless, as downstream tasks cannot benefit from inaccurate data integration. Hence, evaluating the accuracy of the PPRL methods is crucial. This effort was developed to specifically evaluate the potential privacy risks in the recently developed NCI-sponsored PPRL module for the Match*Pro software system, 7 a new tool that is now being deployed by cancer registries. This PPRL tool creates tokens for each patient record by (1) combining different personally identifying attributes (eg, patient’s first name, last name, date of birth, phone number, social security number and residential address) and (2) securely hashing these values. When 2 cancer registries use the same hash key, records belonging to the same patients should be likely to have common hashed tokens, thus facilitating the linkage of patient records. The Match*Pro’s PPRL developers conducted an extensive evaluation related to the utility of the proposed PPRL method. 8 The evaluation assessed the ability to accurately link records across datasets under 3 real-world scenarios: (1) highly accurate and complete personally identifiable information (PII), (2) data with missing Social Security Numbers (SSNs), and (3) data with intentionally introduced errors and missing fields to simulate degraded data quality. A probabilistic, non-privacy-preserving linkage was first conducted to create a manually verified gold standard of person-level matches, which served as the basis for evaluating the performance of multiple PPRL methods. The evaluation included 4 different PPRL methods and assessed true positives, false positives, false negatives, and related metrics, including precision, recall, F1 score, accuracy, specificity, and false discovery rate. It was noted that linkage performance varied substantially across scenarios. In the high-quality PII (Scenario 1), F1 scores ranged from 0.574 to 0.964. In the case of missing SSNs (Scenario 2), F1 scores ranged from 0.123 to 0.964, reflecting the impact of incomplete identifiers. And, in the case of inaccurate or missing PII (Scenario 3), performance declined further, with F1 scores between 0.052 and 0.786. These results demonstrate that while some PPRL methods can achieve near-optimal linkage under ideal conditions, their accuracy diminishes significantly as data quality degrades. The findings in NCI report 8 highlight the importance of selecting the appropriate PPRL method for the quality and completeness of available data to ensure reliable and privacy-preserving linkage. Using the insights gathered from NCI report, 8 the Match*Pro’s PPRL developers chose the current token fields for their proven linkage performance based on their evaluation of other PPRL methods’ performance reported in the NCI report. 8 Although the performance has only been documented in a technical report, stakeholders (eg, cancer registries) found it sufficiently convincing to proceed with the PPRL version of the tool for real-world deployment. Although the developed PPRL tool does not share any individually identifiable information in the clear, there is a potential for some information to be leaked since the hashed tokens are not guaranteed to have a uniform frequency distribution. At the same time, the NCI PPRL tool has several important variations that have not been analyzed before. Specifically, the number of hashed tokens generated per record may vary depending on the record’s values, which could potentially leak information and be exploited for re-identification. Although several mitigation strategies—such as padding each record to the maximum observed token length—have been considered, the current version employs variable token-size generation due to computational resource constraints. Building on prior NCI research focused on utility, 8 this paper focuses to privacy, introducing novel frequency analysis techniques to characterize the degree of information leakage. Furthermore, using a bootstrap-based approach, we quantify the privacy leakage under different settings. To better understand the privacy implications of using a PPRL tool that generates a varying number of hashed tokens per record, this work also examines different deployment settings based on 2 scenarios representing who has access to the hash keys used for generating the hashed tokens. For deployment scenarios where the hashing key is well protected, we consider the privacy risks based on an analysis of the hashed tokens frequency distribution. To achieve this goal, we developed a novel re-identification analysis framework based on hashed token counts, specifically tailored for evaluating the NCI’s PPRL tool. More specifically, using a publicly available resource (in this work, we focus on voter registration lists as they have been leveraged in the past for re-identification attacks 9 , 10 ), we investigated whether an attacker could identify any patients who are in the cancer registry datasets. Our results indicate that if the keys used for hashing are not accessible to an attacker, then even under very pessimistic adversarial assumptions, the re-identification risk for cancer patients will likely to be very limited. Materials and methods Overview of the NCI privacy-preserving record linkage tool Here, we provide an overview of the NCI PPRL tool. Before the linkage process starts, each data contributing site creates hashed token information for each record. More specifically, for the i th record, r i , that are stored at the j th site, S j , 4 token fields are generated as follows: t 1 ( r i ) = { r i . DOY | | r i . Standart ( Firstn ) . r i . Firstn ( 4 ) | | r i . SSN } t 2 ( r i ) = { r i . DoB | | r i . Standard ( Firstn ) | | r i . Standart ( Lastn ) | | r i . SSN ( 4 ) } t 3 ( r i ) = { r i . DoYM | | r i . Standart ( Firstn ) | | r i . Standart ( Lastn ) | | r i . PhoneN } t 4 ( r i ) = { r i . DoYM | | r i . Standart ( Firstn ) | | r i . Standart ( Lastn ) | | r i . StreetA } Here, r i . Standart ( Firstn ) represents the standardized version of the first name of a given record r i using the standardization process used by the tool 7 ; r i . Firstn ( 4 ) represents the first 4 characters of the first name; r i . Standart ( Lastn ) represents the standardized version of the last name 7 ; r i . DoB represents the full date of birth including year (4 digits), month (2 digits), and day (2 digits); r i . DoY represents the 4-digit birth year only; r i . DoYM represents the birth year (4 digits) and birth month (2 digits); r i . SSN represents the full 9-digit Social Security Number; r i . SSN ( 4 ) represents the last 4 digits of the Social Security Number; r i . PhoneN represents the full phone number (10 digits); r i . StreetA represents the house number and partial Street Name; and “‖” represents the string concatenation operation 11 (see online Appendix for more details on street information generation). If a record has multiple first names or last names, a separate token is generated for each. Therefore, each token field may contain multiple tokens. Table S1 presents a synthetic record, while Table S2 illustrates how the synthetic record is applied to create different token fields for the same record. As shown in Table S2 , Token field #4 has 8 distinct tokens for the given record. We assume that the sites are relying upon a semi-trusted third party (TSP) to facilitate the linkage for security reasons . In this setting, if a cancer registry A and a cancer registry B want to link their data, each organization generates one-time use, random key strings K a and K b , respectively, and share it securely with each other (ie, using public key cryptography) (Steps 1 and 2 in Figure 1 ). Later on, each organization sets a common one-time hashed token generation key K as K = K a. ⊕ K b, where ⊕ denotes the bitwise XOR operation (Steps 3 and 4 in Figure 1 ). In addition to token information, a hash of K (ie, H ( K )) is sent to the TSP to ensure that the same key is set by both parties (Steps 5 and 6 in Figure 1 ). Thus, the TSP does not have access to the K used for hashed token generation. Figure 1. Open in a new tab Third semi-trusted party based private record linkage protocol execution steps. Once the set of token sets ( t 1 ( r i ) ) , … , ( t 4 ( r i ) ) are created for each record r i at Site A, a Hash-based Message Authentication Code (HMAC) ( H k ) 12 of the token is computed using the hash key K . In addition, a random identifier rnd i is generated for the record r i to facilitate matching record generation at TSP. To enhance privacy protections during the PPRL execution, the random identifier rnd i can be freshly generated each time a record linkage is performed on the dataset. This ensures that random identifiers cannot be linked across different time periods, further safeguarding privacy. More formally, hashed token set (HS) for site A: H S A = { ∪ r i ∈ S A [ rn d i , H K ( t 1 ( r i ) ) , … , H K ( t 4 ( r i ) ) ] } Similarly, site B computes H S B . Table S3 provides an example of how plaintext tokens are converted into their HMAC encrypted counterparts. Next, the TSP performs record linkage using the hashed token data sent by sites A and B , that is, H S A and H S B , respectively. The TSP then sends back only the matching tuple pair random identifiers (ie, tuple (A, rnd i ) matched with tuple ( B , rnd j )). Hence, if needed, after receiving the matching tuple pairs, sites A and B can exchange further information about specific records of interest. In this setting, a potential attacker may attempt to hack into TSP. However, the attacker will not have access to key K . And, since the potential number of key values is extremely large (ie, 2 256 many different key values), it is not computationally plausible for the attacker to regenerate the hashed tokens. In addition, HMAC is assumed to be a pseudorandom function under certain assumptions. 12 Consequently, without knowing key K , the hashed tokens in HS A or HS B would be indistinguishable from random values—unless a hashed token is repeated. However, reusing key K across sites could induce a vulnerability. For example, if provisioned with access to HS I from different sites that use the same K , the attacker can correlate the tokens. For instance, an attacker may discover a patient visited multiple organizations and infer the number of visits. Thus, for each pair of organization A and B , a new K may be used, such that it will not be feasible to link hash values beyond the given linkage task. Potential attacker prior knowledge To perform a privacy risk analysis, we must formalize what the attacker may potentially know. This is because the attacker’s prior knowledge (eg, first name, last name, date of birth, etc, for a target individual the attacker seeks to identify) significantly influences the level of privacy risk. The following provides a list of the information that may be accessible to an attacker. Hashed token set (HS): Let us denote the hashed token set H S A for site S A that is generated based on the patient records stored as the site (ie, for all records r i ∈ S j ) using salt/key K . Hash salt/key K: In addition, an attacker may know the salt/key K used for HMAC to create the hashed token. This could happen if K is leaked during storage on one of the sites. Alternatively, an insider may see K and disclose it (either deliberately or inadvertently). Additionally, malware running on one of the sites may capture K . Publicly available datasets D : Various datasets (eg, voter registration lists or leaked credit report databases) 9 provide information about the individuals to whom the registry data corresponds, such as their address, first name, surname, phone number, and birth year. In some cases, the full birth date may be available. Given the non-trivial number of data records with SSN leaked due to recent cyberattacks, 13 in our analysis, we assume that the attacker knows the birth date, address, first name, surname, SSN and phone number for target individuals. Threat scenarios One of the most important factors for success in the privacy attack is whether the attacker has access to K . Thus, we consider 2 scenarios based on the accessibility of K . Scenario 1: K, D, HS are accessible to the attacker with complete information : In the worst case scenario, D has complete information about patients (eg, phone numbers are available for everyone in D ) and each person in HS is also included in D (ie, if patient is in HS , he/she is also in D ). In this case, the attacker can launch a brute force attack (ie, try all possible patient records in D ) to re-identify individuals. This means that, for each individual i in D , the attacker can generate the hashed tokens and check whether they are in HS . This implies that if D has 10 million people in it, then around 40 to 50 million HMAC operations will be needed to discover the patients’ information. According to benchmarks, 14 an older GPU can perform these many calculations in approximately a day. This implies that, if K is leaked, then the privacy protections of this private record linkage scheme would be minimal. Scenario 2: K is not accessible to the attacker, D and HS are accessible : As the above discussion highlights, if K i s disclosed, the entire process becomes insecure, making brute force attacks feasible. Thus, in the remainder of this paper, we focus on the scenario where the attacker cannot access key/salt K . Furthermore, we assume that the linkage is accomplished securely using a TSP that is not provided access to the key (see Figure 1 for further details). Frequency analysis One potential vulnerability of the hashed tokens is the potential for repetition of their value. For example, if we create a token with only a first and last name, even if a random HMAC key K is used, the token for “James Smith” (ie, H K ( James ∥ Smith ) is more likely to repeat in the dataset than other names. For example, in our experimental evaluation, there were 107 “James Smith” records. An attacker can easily identify the most common tokens and can find James Smiths from the token list. Therefore, it is critical that tokens belonging to different individuals should not repeat. For Token 2, while repetitions are possible, they are unlikely. For instance, there could be 2 individuals named “James Smith,” who are born on the exact same date, and who happen to share the last 4 digits of their SSN. Nevertheless, similar repetitions might occur in other rare cases across the broader population. To estimate this likelihood, we employ the collision probability estimation method, often referred to as the birthday paradox. 15 This method calculates the probability that among n individuals, at least 2 will share the same birthday, assuming each day of the year is equally probable for a birth. Given n random integers drawn from a discrete uniform distribution with range [1, d ], the probability p ( n ; d ) that at least 2 numbers are the same can be approximated as p ( n , d ) ≈ 1 - e - n . ( n − 1 ) 2 d (1) We report the frequency analysis results for these tokens in the experimental evaluation section. Token count fingerprint analysis Even if the individual tokens are unlikely to be repeated, we need to consider if the number of tokens generated for each individual could leak information. For example, we can check how many tokens will be created for our target individual and create a token count fingerprint for that individual. Consider the example in Table S3 —this record will exhibit a token count fingerprint of 2-2-2-8, indicating 2 hashed tokens each for tokens fields 1, 2, and 3, and 8 for token field 4. This illustrates how, for some records, there may exist a unique token count fingerprint that facilitates re-identification by an attacker—even if the attacker does not have access to key K . To estimate the re-identification risk based on the token count fingerprints, we randomly sampled n records from our dataset for different n values. For this random sample, we grouped the records based on their token count fingerprints. For the record given in Table S3 , we counted the number of records that have the token count fingerprint 2-2-2-8. For each unique token count fingerprint, if the number of records in the group is less than k (re-identification risk group size parameter), we mark the records with that fingerprint as re-identifiable. Materials The data for this evaluation are derived from the North Carolina Voter Register List (NCVRL), 16 which includes information on over 7 million registered voters in North Carolina. These data include a voter’s first name, surname, address, phone number, and birth year, but do not contain social security numbers, birth month, or birthday. Thus, for computational efficiency reasons, we generated random SSNs, birth months and birth days for randomly selected 500,000 NCVRL records, which we refer to as the 500K dataset. We generated the hashed tokens for these records using a random hash key K . Results Frequency analysis results In the analysis of the 500K NCVRL dataset, there were 2 439 892 generated tokens, of which there were only 6 collisions. Upon further inspection of the corresponding records, it was evident that they were likely the same person (eg, the same address, birth year, the same name but different voter registration numbers). In the 500K dataset, there were a total of 28 850 distinct dates of birth and 438 921 unique first and last name combinations. Assuming a uniform distribution for the birth, first name and last name combinations and a random last 4 digits of the SSN, this yields a total of d = 28 850 × 438 921 × 10 000 = 126 628 708 500 000. And, for a set of n = 100 000 individuals, based on eqn (1) , the probability that any 2 would have the same token 2 value is approximately 3.9 × 10 −5 . Thus, the collision probability would be extremely low for smaller datasets. To assess how sensitive the collision probabilities are to dataset size, we estimated these probabilities, when varying n from 100 000 to 1 000 000 (see Figure 2 for details). The implication of this analysis is that the collision probabilities for non-matching patients are estimated to be low (ie approximately around 0.004) for each token used. This further supports the expectation that the false-positive rate for the NCI PPRL tool is likely to be low in practice, given the nature of the tokens used. Figure 2. Open in a new tab Token collision probability vs dataset size. Token count fingerprint analysis results When we analyzed the 500K NCVRL dataset, as expected, for most records, only 1 or 2 tokens are generated for each token field. For example, as shown in Figure 3 , more than 90% of the records have at most 2 tokens generated for each field. At the same time, some records, especially for those who have multiple first or last names (eg, Mary Jane Smith Stone), may have different numbers of tokens generated. Figure 3. Open in a new tab The occurrence probabilities for the most prevalent token count fingerprints. Next, we adjusted the group size parameter k ( ie, assuming a record is re-identifiable if it belongs to a group of size k or smaller ) to estimate the proportion of the population that could potentially be considered re-identified. Figure 4 reports on the re-identification probability (ie, the proportion of records that are considered re-identified) for a dataset size of 250,000 as a function of k . To estimate the error bounds, we used a bootstrap sample of size 20. The results indicate that, when k is 20, only a small fraction of the dataset (ie, approximately 0.0004) is expected to have potentially re-identifiable token count fingerprints. Figure 4. Open in a new tab Re-identification group size k vs average re-identification risk. Finally, we investigated the impact of changing the sample size on re-identification risk. To do so, we varied the sample size from 100,000 to 500,000 while fixing k to 12. We use k = 12 because prior research on governmental data releases commonly adopts 11 as the minimum acceptable group size parameter. 17 For each sample size value, we again applied a bootstrap sample of size 20 to estimate the error bounds. The results, summarized in Figure 5 , indicate that, for lower sample sizes, some records appear to have token fingerprints that are more likely to be unique, but the fraction of such records remains very low (ie, approximately 0.0006). Figure 5. Open in a new tab Average re-identification risk vs sample size. Discussion In this section, we discuss some key considerations relevant to our findings and analysis. We would like to note that our analysis framework applies to any hashing‑based PPRL deployment that uses a trusted third party, relies on the person-specific attributes relied upon in this study, and creates multiple tokens per record. At the same time, our collision probability analysis is based on US naming conventions and may vary for populations with different naming customs. One issue that we need to consider is the coverage of the background information. In practice, not all patients will be in the public background dataset. For example, NCVRL contains data on more than 7.5 million individuals, but the state of NC has around 10.5 million people in total. Therefore, the coverage will not be perfect, and some potential matches identified using our token count fingerprint analysis may not be real matches. This suggests that, in practice, the real re-identification risk could be lower than what we have reported. In addition to public voter registration lists, other online resources can serve as background information. For example, platforms like GoFundMe enable cancer patients to raise funds for medical treatment and related expenses. These online campaigns often disclose personal information, such as diagnosis dates, patient and names (and, at times, their relatives). At the same time, the availability of such online information does not change our analysis since we already assume that the attacker has access to full plaintext values used for record linkage purposes. It is clear from our analysis that if the key/salt used for linkage is leaked, then all the patients’ records could be recovered from the hashed token sets. Thus, the security of K is paramount. In addition, the secure linkage methodology using a TSP could be used to potentially limit the leakage of K . In addition, our analysis implicitly assumes that K is generated randomly at each time and should not be reused, and that HS and K should be transmitted securely in an encrypted format. For enhanced security, K should be destroyed as soon as possible to prevent linkage with the past data by any site. While our analysis indicates the overall re-identification risk for the PPRL technique is expected to be quite low, it is evident that some people could be left in a potentially re-identifiable situation. There are, however, additional strategies that can be employed to mitigate re-identification risk. First, each record could be padded with random dummy tokens to standardize token count fingerprints across all records. While this strategy would enhance privacy, it will inevitably increase the number of tokens that need to be processed and compared, which could significantly impact computational efficiency. Second, each site could help reduce the risk of hashed token repetitions by deduplicating its information locally. Third, data sharing agreements that require sites to use enhance security mechanisms including the deletion of linkage data after usage could be used to reduce the privacy risks even further. Conclusion In this work, we developed a novel frequency-based analysis technique to understand the privacy leakage and re-identification risks of using the recently developed NCI PPRL tool that is a key component of the Match*PRO software developed primarily with a focus on secure linkages to support the caner surveillance system. Our analysis demonstrated that the potential re-identification risk associated with the tool remains very low when appropriate security measures, such as safeguarding hash keys, are implemented. Our analysis further supports the expectation that false-positive rate resulted from the NCI PPRL tool will be minimal in practice. In the future, we plan to extend our analysis to include other PPRL tools that are widely used in the healthcare domain. Supplementary Material ocaf172_Supplementary_Data ocaf172_supplementary_data.docx (56.3KB, docx) Acknowledgments This project has been funded in whole or in part with Federal funds from the National Cancer Institute, National Institutes of Health, under Contract No. HHSN261201500003I, Task Order HHSN26100038. The content of this publication does not necessarily reflect the views or policies of the Department of Health and Human Services, nor does mention of trade names, commercial products or organizations imply endorsement by the U.S. Government. Contributor Information Murat Kantarcioglu, Department of Computer Science, Virginia Tech, Blacksburg, VA 24061, United States. Will Howe, Information Management Services, Inc., Calverton, MD 20705, United States. Benmei Liu, Surveillance Research Program, Division of Cancer Control & Population Sciences, National Cancer Institute, Bethesda, MD 20892, United States. Valentina Petkov, Surveillance Research Program, Division of Cancer Control & Population Sciences, National Cancer Institute, Bethesda, MD 20892, United States. Esmeralda Casas-Silva, Center for Biomedical Informatics and Information Technology, National Cancer Institute, Bethesda, MD 20892, United States. Diana Velasquez-Kolnik, Frederick National Laboratory for Cancer Research, Frederick, MD 21701, United States. Bradley A Malin, Department of Biomedical Informatics, Vanderbilt University Medical Center, Nashville, TN 37232, United States; Department of Computer Science, Vanderbilt University, Nashville, TN 37232, United States; Department of Biostatistics, Vanderbilt University Medical Center, Nashville, TN 37232, United States. Lynne Penberthy, Surveillance Research Program, Division of Cancer Control & Population Sciences, National Cancer Institute, Bethesda, MD 20892, United States. Author contributions Murat Kantarcioglu conducted the privacy analysis and prepared the initial draft of the manuscript. Will Howe performed the token generation for the PPRL protocol. All authors contributed to the revision, editing, and finalization of the manuscript. Supplementary material Supplementary material is available at Journal of the American Medical Informatics Association online. Funding This work was supported by the National Institutes of Health, National Cancer Institute grant numbers HHSN26100038 and HHSN261201500003I. Conflicts of interest The authors have no competing interests. Data availability The data underlying this article will be shared on reasonable request to the corresponding author. References 1. Kho AN, Cashy JP, Jackson KL, et al. Design and implementation of a privacy preserving electronic health record linkage tool in Chicago. J Am Med Inform Assoc. 2015;22:1072-1080. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 2. Gkoulalas-Divanis A, Vatsalan D, Karapiperis D, Kantarcioglu M. Modern privacy-preserving record linkage techniques: an overview. IEEE Trans Inform Forensic Secur. 2021;16:4966-4987. [ Google Scholar ] 3. Ben-David A, Nisan N, Benny P. FairplayMP: a system for secure multi-party computation. In: Proceedings of the 15th ACM Conference on Computer and Communications Security (CCS '08). ACM; 2008: 257-266. [ Google Scholar ] 4. Bos JW, Lauter K, Loftus J, Naehrig M. Improved security for a ring-based fully homomorphic encryption scheme. In: Stam M, ed. Cryptography and Coding. IMACC 2013. Lecture Notes in Computer Science, Vol 8308. Springer; 2013. 10.1007/978-3-642-45239-0_4 [ DOI ] [ Google Scholar ] 5. Durham E, Xue Y, Kantarcioglu M, Malin B. Quantifying the correctness, computational complexity, and security of privacy-preserving string comparators for record linkage. Inf Fusion. 2012;13:245-259. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 6. Christen P. Data Matching: Concepts and Techniques for Record Linkage, Entity Resolution, and Data Deduplication. Springer; 2012. [ Google Scholar ] 7. National Cancer Institute. MatchPro*. National Cancer Institute; 2023. [ Google Scholar ] 8. National Cancer Institute. Evaluating the performance of privacy preserving record linkage systems (PPRLs). 2023. Accessed October 16, 2025. https://surveillance.cancer.gov/reports/TO-P2-PPRLS-Evaluation-Report.pdf 9. Benitez K, Malin B. Evaluating re-identification risks with respect to the HIPAA privacy rule. J Am Med Inform Assoc. 2010;17:169-177. [ DOI ] [ PMC free article ] [ PubMed ] [ Google Scholar ] 10. Sweeney L. K-anonymity: a model for protecting privacy. Int J Uncertainty Fuzziness Knowledge-Based Syst 2002;10:557-570. [ Google Scholar ] 11. String concatenation. 2024. Accessed October 16, 2025. https://en.wikipedia.org/wiki/Concatenation 12. Krawczyk H, Bellare M, Canetti R. HMAC: Keyed-Hashing for Message Authentication. Internet Engineering Task Force RFC: 2104; 1997. [ Google Scholar ] 13. Flajolet P, Gardy D, Thimonier L. Birthday paradox, coupon collectors, caching algorithms and self-organizing search. Discrete Appl Math (1979). 1992;39:207-229. [ Google Scholar ] 14. Gosney JM. 2024. Accessed October 16, 2025. https://gist.github.com/epixoip/a83d38f412b4737e99bbef804a270c40 15. Picchi A. Hackers may have stolen your social security number in a massive breach. Here’s what to know. Lee A. M., ed. 2024. Accessed October 16, 2025. https://www.cbsnews.com/news/social-security-number-leak-npd-breach-what-to-know/ [ Google Scholar ] 16. North Carolina Voter Registration Database. Accessed November 10, 2022. ftp://www.app.sboe.state.nc.us/data 17. Yao AC. Protocols for secure computations. In: IEEE Annual Symposium on Foundations of Computer Science. IEEE. 1982:160-164. [ Google Scholar ] Associated Data This section collects any data citations, data availability statements, or supplementary materials included in this article. Supplementary Materials ocaf172_Supplementary_Data ocaf172_supplementary_data.docx (56.3KB, docx) Data Availability Statement The data underlying this article will be shared on reasonable request to the corresponding author. Articles from Journal of the American Medical Informatics Association: JAMIA are provided here courtesy of Oxford University Press ACTIONS View on publisher site PDF (875.3 KB) Cite Collections Permalink PERMALINK Copy RESOURCES Similar articles Cited by other articles Links to NCBI Databases Cite Copy Download .nbib .nbib Format: AMA APA MLA NLM Add to Collections Create a new collection Add to an existing collection Name your collection * Choose a collection Unable to load your collection due to an error Please try again Add Cancel Follow NCBI NCBI on X (formerly known as Twitter) NCBI on Facebook NCBI on LinkedIn NCBI on GitHub NCBI RSS feed Connect with NLM NLM on X (formerly known as Twitter) NLM on Facebook NLM on YouTube National Library of Medicine 8600 Rockville Pike Bethesda, MD 20894 Web Policies FOIA HHS Vulnerability Disclosure Help Accessibility Careers NLM NIH HHS USA.gov Back to Top