ConceptioArchivearXiv CS
arXiv CSopen access

Analysis of Personal Data Exposure in Thailand

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

Analysis of Personal Data Exposure in Thailand∗ Suphannee Sivakorn1 , Sasawat Malaivongs2 , Nuttaya Rujiratanapat1

arXiv:2604.23538v1 [cs.CR] 26 Apr 2026

Received: date / Accepted: date

Abstract In the digital era, personal data, particularly sensitive identifiers such as the Social Security Number and National Identification Number, have become a highly valuable asset, raising significant concerns regarding privacy and security. This study examines the risks associated with the online exposure of the Thai National Identification Number, a key element of identity verification in both governmental and commercial transactions. Similar to the Social Security Number in the United States, this unique identifier is crucial for various legal, financial, and welfare-related activities. However, the increasing digitization of personal records has heightened its vulnerability to unauthorized access and misuse, particularly through search engines that inadvertently index sensitive information. This research identifies publicly exposed Thai National Identification Numbers across major search engines, assessing the potential threats to individual privacy and national security. The study reveals the exposure of over 1.2 million unique National Identification Numbers, along with other highly sensitive personal data, e.g., addresses, contact details, employment status, disability status, and health information. Notably, the analysis indicates that a significant majority of these exposures originate from the Thai government sector websites, highlighting critical vulnerabilities in public data management practices. This widespread exposure not only increases the risk of identity theft and financial fraud but also underscores the urgent need for enhanced cybersecurity measures, stricter regulatory enforcement, and improved data governance within government agencies to prevent future breaches. Addressing these issues is essential to safeguarding citizens’ personal information and ensuring compliance with Thailand’s data protection laws in an increasingly digitized world. ∗

This manuscript is under review.

Suphannee Sivakorn E-mail: suphannee [email protected] Sasawat Malaivongs E-mail: [email protected] Nuttaya Rujiratanapat E-mail: nuttaya [email protected] 1 Department of Computer Science, Faculty of Science and Technology

Rajamangala University of Technology Tawan-ok, Thailand 2 ACinfotec Co., Ltd., Thailand

Keywords Personal Data Leakage · Sensitive Data Exposure · Thai National Identification Number · Government Data Breach

1 Introduction In the digital age, personal data has become a commodity of immense value, facilitating transactions, services, and often serving as a gateway to individual identities. Thailand, like many other nations, faces significant challenges in safeguarding the privacy and security of its citizens’ personal information, particularly concerning the Thai National Identification Number. Similar in function to Social Security Number (SSN) in the United States, the Thai National Identification Number is issued by the Thai government to individuals who are Thai citizens and/or registered in the household registration system. The Identification Number serves as evidence of identity, proof of personal status, and confirmation of identity for government transactions, service requests, and welfare benefits from state agencies. It is also used for various commercial transactions, legal activities, and other purposes such as job applications, opening bank accounts, asset and property transfers. However, its misuse and exposure through online platforms pose substantial threats to both individual privacy and national security. The proliferation of Internet usage and the digitization of personal records have inadvertently exposed sensitive information, including Thai National Identification Numbers, to online platforms. Search engines, designed to index and retrieve vast amounts of information, inadvertently become repositories of personal data. This exposure creates fertile ground for malicious actors engaged in identity theft, financial fraud, and other forms of cybercrime [1, 2]. The ease with which such information can be harvested and exploited underscores the urgent need for enhanced cybersecurity measures and regulatory frameworks. Existing legal frameworks, such as Thailand’s Personal Data Protection Act (PDPA) B.E. 2562 (2019) [3], aim to regulate the collection and use of personal data. However, the widespread exposure of sensitive information online suggests gaps in enforcement and the implementation of cybersecurity best practices. Previous studies on personal data exposure and cybercrime have focused on similar issues

Suphannee Sivakorn1 , Sasawat Malaivongs2 , Nuttaya Rujiratanapat1

2

in other countries [4–8], but research specific to the exposure of Thai National Identification Numbers remains scarce. In an effort to understand the prevalence and implications of personal data exposure online, this study investigates the current state of personal data leaks as revealed through search engine results. Utilizing a systematic search methodology across major search engines, we identify and collect search results related to personal information, with a particular focus on the Thai National Identification Number. These results are then analyzed across multiple dimensions to identify key patterns, assess the primary root causes of data exposure in Thailand, and propose strategies to mitigate future leaks and prevent the malicious exploitation of exposed data. The key contributions of this study are as follows: – To the best of our knowledge, this is the first practical study to systematically analyze personal data exposure online in Thailand. This research evaluates the prevalence and underlying factors contributing to personal data exposure, offering a data-driven perspective on the issue. – We designed and implemented an automated personal data collection system capable of retrieving and extracting exposed personal information, particularly Thai National Identification Numbers from online sources. This system serves as a monitoring tool for identifying personal data leaks and safeguarding sensitive information. Through this approach, we successfully collected over 1.2 million instances of exposed personal data. – Our analysis identifies the primary sources of personal data leak, revealing that a significant proportion originates from government sector websites, particularly local government offices. Additionally, we uncover specific data patterns and instances of publicly disclosed sensitive information that could be exploited for malicious activities such as phishing, scams, and identity theft, posing a direct threat to individual privacy. The remainder of this paper is structured as follows: Section 2 provides background information and discusses related work, including personal data regulations, relevant laws, and the theoretical context surrounding the Thai National Identification Number, personal data exposure, and cybercrime trends in Thailand. Section 3 presents the technical background, covering search engine indexing, website structures, domain names, and Thailand’s domain name system. Section 4 details our methodology for developing the personal data collection system, including keyword selection, data retrieval processes, and personal data extraction techniques. Section 6 presents the analysis of the collected data, including statistical insights and key findings. We explore and discuss general countermeasures against the personal data exposure in Section 7. Section 8 discusses the ethical considerations of this research, including responsible disclosure practices. Finally, Section 10 concludes the paper and outlines future research directions.

2 Background 2.1 Personal Data and Thai Personal Data Protection Act B.E. 2562 (2019) 2.1.1 Overview The Thai Personal Data Protection Act (PDPA), B.E. 2562 (2019), Thailand’s first comprehensive data protection law, was published in the Royal Thai Government Gazette (Ratchakitchanubeksa) on May 27, 2019, and became fully effective on June 1, 2022. Modeled closely on the European Union’s General Data Protection Regulation (GDPR) [9], the PDPA seeks to align with global data privacy and protection standards. By adopting many of the GDPR’s principles, the law aims to strengthen data privacy in Thailand and ensure consistency with international best practices. Like the GDPR, this law grants individuals significant control over their personal data, ensuring that the data collection, processing and storage activities are transparent, lawful and secure. Both regulations emphasize the importance of obtaining explicit consent from data subjects and provide individuals with a range of rights over their data, such as the right to access, correct, and delete personal data, as well as the right to withdraw consent and transfer data across borders. In terms of enforcement, the PDPA also follows the GDPR’s approach by establishing strict penalties for non-compliance and for breaches of data protection regulations. However, when it comes to enforcement, there is a notable difference in the penalty structures. The GDPR allows for fines based on a percentage of an organization’s annual global revenue up to 4% of annual global turnover or 20 million EUR, whichever is higher. In contrast, the PDPA imposes fixed penalties, with fines not exceeding 5 million THB (approximately 141,000 USD) [10]. Both the PDPA and the GDPR emphasize the protection of sensitive personal data (such as health, religious beliefs, and biometric information), recognizing the increased risk of harm that can result from unauthorized disclosure or misuse of such information. Although the PDPA is tailored to the context of Thailand’s legal and cultural framework, it shares many similarities with the GDPR, positioning Thailand as part of the global movement toward stronger data protection practices. 2.1.2 Key Categorizations and Personal Data Owner Rights The PDPA defines “Personal Data” and “Sensitive Personal Data” as follows: Personal Data refers to any information related to an individual that can be used to identify that person, either directly or indirectly. Examples of personal data include names, various identification numbers (e.g.Thai National Identification Numbers, passport numbers, bank account numbers, credit

Analysis of Personal Data Exposure in Thailand

card numbers), contact information (e.g.addresses, email addresses, phone numbers), and asset and property information (e.g.vehicle registration numbers, house registration number), and other data that can be linked to personal information (e.g.date and place of birth, medical records, educational records, financial information, employment records). Sensitive Personal Data includes high-risk information such as race, ethnic origin, political opinions, cult, religious or philosophical beliefs, sexual behavior, criminal records, health information, disabilities, biometric data or of any data which may affect the data subject in the same manner. The PDPA imposes stricter safeguards and higher penalties for the unauthorized disclosure of these categories due to their potential for discriminatory misuse. Personal Data Owner Rights. Under the PDPA, data owners are entitled to access and obtain copies of their personal data, request corrections, and demand erasure or destruction when appropriate. They may withdraw consent at any time, request data portability for information obtained directly from them, and object to processing undertaken on legal grounds. Additionally, they may request temporary restrictions on data processing. Organizations must comply or provide documented justification for any refusal.

3 Table 1 Classification of Individuals by the First Digit of the National ID Number Number 1&2

3 4 5 6 7 8

Description of Individual Category Thai nationals born after Jan 1, 1984, with the number 1 indicating those who registered their birth within 15 days, and the number 2 indicating those who failed to register within that period. Thai nationals/foreigners in the household registry prior to May 31, 1984. Thai nationals/foreigners establishing residency before an initial ID assignment. Individuals added later due to previous omissions or special circumstances. Temporary residents, illegal entrants, or ethnic groups awaiting citizenship. Children born in Thailand to parents classified under Category 6. Legal foreign residents or those acquiring Thai nationality after May 31, 1984.

Sequential Identifiers (Digits 6–12): serve as internal classification groups or sequential birth certificate numbers, providing a unique serial identifier within the local registry. Checksum Validation (Digit 13): is the final checksum digit, calculated to verify the correctness of the preceding 12 digits of the National ID number. 2.2.2 National ID Number Checksum Validation

2.2 Thai National Identification Number The Thai National Identification (ID) Number or Thai ID Number or Citizen ID Number, is a 13-digit number displayed on the Thai National ID Card, a government-issued document for Thai citizens who are registered in the household registration system. This card serves as a primary form of identification for Thai nationals and is used to verify and authenticate an individual’s identity. The National ID number serves as the primary key for accessing services, financial systems, and legal transactions in Thailand, particularly for digital authentication and verification purposes. 2.2.1 The Meaning of National ID Number The Thai National ID Number, consisting of 13 digits, holds specific meanings for each digit, as outlined below. Digit 1 (Individual Classification): signifies the category of the individual. These categories, detailed in Table 1, are fundamental to the demographic analysis of data exposure patterns conducted in this study. Geographic Identifiers (Digits 2–5): represents the issuing registration office: Digit 2 denotes the administrative region (1–9), Digit 3 identifies the province, and Digits 4–5 specify the district or municipality.

Fig. 1 Format of National ID Number as Displayed on the Thai National ID Card

Here, we describe the method used to calculate the checksum (Digit 13). For malicious actors, this checksum facilitates “data scrubbing,” allowing any leaked numeric string to be programmatically verified as a valid National ID, thereby increasing the efficiency of mass data harvesting. Step 1: Multiply each digit of the National ID number by its corresponding position multiplier (from 13 down to 2), then sum all the results. Table 2 presents the example of calculating the sum of all multiplication results (351) Step 2: Divide the result obtained in Step 1 by 11 and calculate the remainder (modulus operation): 351 mod 11 = 10 Step 3: Subtract the remainder obtained in Step 2 from 11. In this example: 11 − 10 = 1 Step 4: If the result of the subtraction is a two-digit number, use the units digit. Therefore, the 13th digit of the ID number is 1.

2.3 Cybercrime Landscape in Thailand Cybercrime in Thailand has been on the rise in recent years. According to statistics from the Royal Thai Police, between March 2022 and May 2023, over 296,000 cybercrime-related reports were filed, an average of 525 cases per day, resulting in

Suphannee Sivakorn1 , Sasawat Malaivongs2 , Nuttaya Rujiratanapat1

4 Table 2 Example of Checksum Calculation for the 13th Digit Verification Position

1

2

3

4

5

6

7

8

9

10

11

12

13 (Checksum)

Position Multiplier Example Multiplication result Checksum Validation

13 1 13

12 2 24

11 3 33

10 4 40

9 5 45

8 7 6 7 48 49 351

6 8 48

5 9 45

4 1 4

3 0 0

2 1 2

1

financial damages exceeding 40 billion THB (approximately 1.2 billion USD), or an average of 74 million THB per day. The majority of these crimes involve various forms of online scams and fraud which include Fraudulent online transactions involving goods and services (37%), Money transfer scams (13%), Loan fraud (12%), Investment scams (8%) and Cyber extortion (7%) [11]. This trend is driven by the widespread adoption of real-time payment systems, which, while efficient, have positioned Thailand sixth globally in online scam frequency [12]. 2.3.1 Systemic Data Breaches and Identity Risk The efficacy of these scams is directly linked to the availability of leaked PII, which facilitates high-fidelity spear-phishing. Thailand has faced several landmark breaches that underscore this systemic vulnerability: Mass Population Exposure: In 2024, a hacker known as “9near” claimed the exposure of 55 million citizens’ data, affecting nearly 83% of the population [13]. Commercial and Institutional Leaks: Also in 2024, significant breaches involving 5 million records from loyalty programs [14] and the illicit sale of customer data by bank employees [15] highlight vulnerabilities across both private and financial sectors. Beyond financial loss, these breaches inflict long-term identity scarring, where victims face years of credit recovery and psychological distress [16, 17]. 2.3.2 Thai Cybercrime Law B.E. 2566 (2023) Recognizing the urgent need for a coordinated response to cybercrime, the Thai government enacted the Emergency Decree on Measures for the Prevention and Suppression of Technological Crimes, B.E. 2566 (2023) [18]. This legislation enhances collaboration between key stakeholders, including banks, telecommunications providers, internet service providers (ISPs), law enforcement, and financial institutions [19]. Key provisions of the decree include: – Strengthening information-sharing mechanisms among relevant authorities to detect and respond to cybercrime. – Granting account holders the right to temporarily freeze and report suspicious transactions to prevent fraud. – Imposing stricter legal penalties on individuals involved in cybercriminal activities, such as, money laundering,

11 − (351 mod 11) = 1

money muling [20], and fraudulent identity schemes e.g., opening bank accounts, registering phone numbers or creating fake social media profiles to impersonate government authorities for scam operations.

2.4 Criminal Exploitation of National ID Numbers The exploitation of personal data, particularly the Thai National ID number, has become a critical concern in the digital age. Criminals leverage stolen personal information to engage in fraudulent activities, financial crimes, and identity theft, posing serious risks to individuals and institutions. Below are key examples of how the Thai National ID number can be misused for illicit purposes: Identity Theft and Financial Fraud. Cybercriminals can exploit an individual’s National ID number to impersonate them and gain unauthorized access to financial services [21]. The risk is further heightened when National ID numbers are coupled with the Laser Codes, an alphanumeric identifier located on the back of the Thai National ID card, which allows criminals to bypass GDX-based authentication to open “mule” accounts, apply for fraudulent loans, or intercept social welfare benefits [22]. Online Phishing Scams. Cybercriminals frequently use stolen personal data, including National ID numbers, in phishing schemes designed to deceive victims into divulging additional sensitive information. The more information they possess about a target, the more convincing and effective their phishing attacks become [23]. Dark Web Monetization: Verified citizen registries including National ID numbers are trafficked on hacker forums, where they serve as foundational datasets for long-term identity theft and automated scam operations [13, 15].

2.5 Emerging Risks: AI Training and Persistent Privacy Risks The risk of exposure of personal information has evolved considerably with the rapid proliferation of Large Language Models (LLMs). Large-scale web scraping of publicly accessible internet content constitutes a primary data source for training AI models [24, 25]. This practice raises significant privacy concerns, as training pipelines frequently ingest vast amounts of public data without reliably filtering PII or obtaining explicit consent from data subjects [26–28]. Such

Analysis of Personal Data Exposure in Thailand

large-scale data ingestion introduces a persistent identity memorization risk, whereby sensitive information, including National ID Numbers, may become implicitly embedded within model parameters. Unlike traditional databases, where compromised records can be deleted or access revoked, removing specific data instances from a trained model (e.g., machine unlearning) remains technically challenging and often infeasible in practice [29, 30]. Consequently, the exposure of Thai citizens’ personal data may result in long-term and potentially irreversible security and privacy risks within the global AI ecosystem.

3 Technical Background This section presents the technical background necessary to understand the approach used in this study.

3.1 Search Engine A search engine is a program designed as a tool for retrieving information from various online sources. Most search engines enable users to search for information on websites by entering keywords or phrases. The engine then processes the search query and presents relevant results to the user. The term entered for a search is referred to as “search keyword”. 3.1.1 Search Engine Processes In general, search engines operate through several key stages to create and manage their databases. This section details the critical steps involved in search engine operations, which are essential for understanding the focus of this paper. 1. Crawling refers to the process of collecting data from websites across the internet. Search engines gather information by following links within the content of a website that direct users to other sites. 2. Indexing involves creating an index of the content collected during the crawling process. This index is generated and stored based on the content of each website. While search engine indexes can be generated automatically by the search engine itself, website creators can customize their indexes using metadata (e.g., meta tags) [31], robots.txt files [31], or other mechanisms. The primary objective of indexing is to organize the collected data in a way that allows it to be efficiently queried and used to deliver relevant search results in response to user 3. Ranking is the process by which search results are ordered before being presented to users. The ranking of search results is determined by several factors, including the relevance of the search term to the indexed content, the quality of the content, and the volume of web traffic,

5

among others. Typically, web pages that are most closely related to the search keyword are ranked higher, with results presented in descending order of relevance. This ensures that users can more easily access the information they are seeking. 3.1.2 Advanced Search In addition to basic keyword searches, modern search engines support advanced search features or special search operators, which enhance the accuracy and precision of search results [32–34]. Examples of Advanced Search Techniques – Quote (" "): This method involves using quotation marks around a search term to instruct the search engine to only return search results that contain the exact phrase. For example, entering "dog" in the search engine will return results that only contain the word “dog” in the content [32]. – Exclude (-): This method involves using a minus sign before a search term to exclude results that contain that term. For example, "-cat" will show search results that do not contain the word “cat”. – Filetype: This operator is used to search for specific file types. For example, using filetype:pdf will show results only from PDFs, and filetype:doc will return search results from DOC files [32, 33]. – Site: This search operator restricts results to a specific website. For instance, site:example.com will return only search results from pages within the domain example.com [32]. – Operators (AND, OR): This allows the user to refine their search based on specific conditions. For example: – AND: Refines the search to return search results containing all specified terms. For example, cat AND dog will show search results that contain both “cat” and “dog.” – OR: Expands the search to return search results containing any of the specified terms. For example, dog OR cat OR bird will return search results containing any of the three terms. – Parentheses: This operator is used to group terms and organize the logic of the search, similar to its use in mathematics. For example, dog AND (cat OR bird) will show search results that contain the word “dog” and either “cat” or “bird.” Users can combine these advanced search operators to refine their searches. For example: – site:example.com filetype:pdf "dog" AND ("cat" OR "bird") This query would return results with the following conditions:

Suphannee Sivakorn1 , Sasawat Malaivongs2 , Nuttaya Rujiratanapat1

6

– Search results are limited to web pages within example.com (e.g. www.example.com/pet). – The content must come from PDF files. – The page must contain the word “dog” and also either “cat” or “bird.”

3.2 URL and Domain Name URL (Uniform Resource Locator) is the address used to access a specific resource on the internet, such as a web page, image, or file. A URL serves as a standardized reference that allows users to locate and retrieve web content efficiently. The URL is composed of several key components, including the protocol, domain name, and path. For example, in the URL https://example.com/page1, https is the protocol, example.com is the domain name, and /page1 is the specific path to the resource. Domain name is a human-readable identifier used to represent an IP address in a more memorable and userfriendly format. Each domain name corresponds to an IP address, which is a numerical identifier used by computers to locate resources on the internet [35]. Domain names are used to identify websites, email servers, and other online services.

reflecting a specific category or purpose. While some TLDs, such as .com, stand alone without an SLD, others, like country code TLDs (.th in "example.co.th"), incorporate an SLD within their structure. This tiered system facilitates the organization and categorization of domain names based on their intended use or the entity they represent (e.g. government, education, businesses, non-profits). This hierarchical structure enhances user recognition of website ownership and purpose, contributing to a more streamlined and navigable domain name system. The subsequent section will delve into the specific categories of Thailand’s SLDs.

3.4 Thailand’s Domain Name System and the .th ccTLD Thailand began its connection to the internet in 1986 [37]. Subsequently, the use of the .th Top-Level Domain (TLD) was introduced to represent domain names associated with Thailand. Today, the registration of domain names under the .th TLD is organized into seven categories of SLDs outlined in the Table 3. Each SLD category is intended to serve a specific sector, ensuring that domain registration aligns with the intended use, for example, governmental, academic, commercial, military, non-profit, individual, or network-related.

3.3 Top-Level Domain Name (TLD) 4 Methodology Top-Level Domain (TLD) refers to the final segment of a domain name, which comes after the last dot. The TLD serves as a classification system for domain names and can denote various organizational, geographical or functional attributes. For example: – Generic Top-Level Domains (gTLDs): These are the most commonly used TLDs and consist of three or more alphabetical characters. Examples include .com (commercial), .info (information), and .org (organization). These TLDs are often used by a wide variety of organizations and businesses for general Internet purposes. – Country Code Top-Level Domains (ccTLDs): These TLDs are two-letter domain suffixes that are designated to represent specific countries or territories, according to the ISO 3166 international standard [36]. Examples include .us (United States), .th (Thailand), and .cn (China). ccTLDs are typically used to indicate a geographic or national connection and often serve as an indicator of the website’s primary audience or origin. 3.3.1 Second-Level Domain Name (SLD) The Second-Level Domain (SLD) resides immediately preceding the Top-Level Domain (TLD). It functions as a unique identifier within a broader domain namespace, potentially

In this section, we will outline the details and methodology employed in the development of a system designed for the collection of personal data.

4.1 Personal Data Searching The collection of personal data in this system relies on search results retrieved from search engines. Our system supports searches and data aggregation from Google and Bing, two of the most widely used search engines, based on a survey conducted in 2024 [38]. To facilitate this, we have developed a Search Engine Crawler, which is responsible for sending search queries to search engines and collecting the results. The collected data consists of the following components: (1) Search Keywords, (2) URLs resulting from the search, and (3) the timestamp indicating when the search was conducted and the results were retrieved. These data are then stored in our search result database. Our Search Engine Crawler can be configured with the following parameters: (1) Search Keyword(s), (2) Preferred Search Engine(s) (either Google and/or Bing), (3) Search Delay, which controls the interval between search queries and the collection of results from each search results page, (4) Download Timeout and Download Max Retry, which set

Analysis of Personal Data Exposure in Thailand

7

Table 3 Types and Purposes of SLD under the .th TLD .th TLD

Purpose of Domain Name

.go.th

Government agencies and organizations

.ac.th

Academic institutions and organizations

.co.th

Commercial entities and businesses

.mi.th

Military organizations

.or.th

Non-profit organizations

.in.th

Individuals and general organizations

.net.th Network-related organizations

Example Domain Names thaigov.go.th (Thai Government), dla.go.th (Department of Local Administration) rmutto.ac.th (Rajamangala University of Technology Tawan-Ok), mahidol.ac.th (Mahidol University) google.co.th (Google Thailand), thailandpost.co.th (Thailand Post) rta.mi.th (Royal Thai Army), navy.mi.th (Royal Thai Navy) glo.or.th (Government Lottery Office), set.or.th (Stock Exchange of Thailand) bnn.in.th (BaNANA IT products) uni.net.th (Office of Information Technology for Educational Development), cat.net.th (National Telecommunications Public Company Limited)

the maximum time and number of attempts to download the result URLs, (5) Maximum Search Result Pages, and (6) Desired File Type for the downloaded results. 4.1.1 Maximum Number of Search Results The maximum number of search results refers to the total number of relevant results returned for a given search query, which are ranked according to their relevance, as discussed in Section 3.1. These results are then distributed across multiple pages of search results. Our experiments indicate that most search engines typically display between 6-10 results per page, although this number may vary. The Search Engine Crawler we developed is configurable, allowing users to define the maximum number of pages to crawl, with a default setting of 10 pages. Furthermore, the crawler is capable of detecting the actual number of pages, even if this number is fewer than the user-defined or default settings.

page on which the item appears, the download URL, the download status, and the date and time of retrieval. Additionally, the system calculates a SHA-256 hash for each file to facilitate integrity checks and support subsequent comparisons. This hash is particularly useful for identifying duplicate files that may appear across different search results or websites. The downloader further verifies that the downloaded files match the expected file types, ensuring the accuracy and completeness of the retrieved content. This verification is accomplished by examining the file extension specified in the Content-Disposition header, a standard HTTP header that indicates the file should be downloaded along with its name and extension. Once verified, the file is renamed with the corresponding hash value of its content, and both the hash and file type are recorded in the database, along with the previously noted information.

4.2 Search Keyword Selection 4.1.2 Search Result Retriever and Collection Upon retrieving the search results, our system utilizes the Search Result Downloader to collect the URLs from each result page. The downloader operates according to two key parameters: Download Timeout and Download Max Retry, which govern the maximum duration and the number of retry attempts for downloading the result URLs. Once all results on a given page have been successfully downloaded or the timeout is reached, the downloader enters a brief idle state before proceeding to fetch the next page, as determined by the specified search delay. This delay is designed to simulate real user behavior, as users typically spend some time reviewing the search results before advancing to the next page. 4.1.3 Search Result Information and File Download After each download, the system records key information in the database, including the search keyword, the search result

In order to obtain search results that capture the most personal data of Thai citizens, the selection of search keywords is of paramount importance. The following principles guide the selection of search keywords:

4.2.1 Keywords Related to National ID Numbers Examples of such keywords (all translated from Thai) include ID card number, National ID number ID card code, number, ID card, etc. These search terms are aimed at locating information related to National ID numbers, typically used in conjunction with advanced search techniques such as quotes to ensure the search results must include these specific terms. Additionally, advanced search operators such as "OR" can be used to encompass multiple keywords within a single search query. For example: "ID card number" OR "National ID number".

8

4.2.2 File Type-Specific Keywords (filetype) Examples of such search terms include "filetype:xls", "filetype:pdf", etc. Personal data is often stored in file formats such as PDF, XLS, and DOC files. Specifying the file type in a search query helps to locate information contained within these files, which are commonly uploaded to various websites. Our crawler currently restricts searches to only three file types, aligning with its capability to read and extract data from these formats: (1) PDF documents (.pdf) (2) Spreadsheet documents (.xls, .xlsx) and (3) Word Processing documents (.doc, .docx). 4.2.3 Search Keywords Based on National ID Number Prefixes

Suphannee Sivakorn1 , Sasawat Malaivongs2 , Nuttaya Rujiratanapat1

keywords to further gather personal data. For example, some keywords targets the list of individuals receiving government allowances, and this term was found to appear frequently in search results, proving useful in collecting personal data. To maximize the potential for gathering personal data, a combination of these search keywords was used. The search results, along with the used search terms, were stored for future reference. Examples include: – site:ac.th filetype:xlsx OR filetype:xls "number" "citizen" "Mr." – filetype:xls "National ID number" "name" – filetype:pdf "1-3501-" "number" "citizen" – "certificate of tax withholding" filetype:pdf site:go.th ("Miss" AND "Mr.") – filetype:pdf ("ID card number" OR "National ID number" OR "number") "list" "1-1001-"

Given the specific structure of National ID numbers, especially the first five digits, as detailed in Section 2.2, these digits can be used to create targeted search keywords. Examples are provided in Table 4 to illustrate how searches can identify individuals in specific districts, sub-districts, or provinces. Furthermore, advanced search techniques using Quotes ensure that the results contain the specific number pattern and its sequence. To comprehensively search for personal data across all provinces, the first digit of the National ID number is used as ‘1’, while digits 2-5 correspond to the district code for every province in Thailand. These structured searches help gather extensive personal data.

The collection of personal data begins with the extraction of text from files, followed by a thorough review of the extracted content to identify relevant personal data for storage. To evaluate the potential risk of personal data leakage, we have specifically focused on extracting National ID numbers as a representative marker for such leakage. The development process for this procedure is outlined as follows:

4.2.4 Keywords Related to Personal Names

4.3.1 Text Extraction

Examples (all translated from Thai) include "Mr.", "Mrs.", "Miss", "Ms.", "Master"/"Mstr" (boy), "Miss" (girl), etc. These keywords are intended to refine search results to include data related to personal names by specifying common prefixes used in Thai culture.

Text extraction is the process of retrieving or extracting text from files for use or analysis. This is achieved through the creation of a parser tailored to the specific type of file collected. The parsers developed for this purpose are all based on open-source libraries and have been designed to handle the following file types: PDF Documents (.pdf): The PDF parser was developed using various PDF parser libraries, including (1) PyPDF2 [39], (2) pdfplumber [40], and (3) pytesseract [41]. These libraries selected were chosen for their effectiveness in extracting Thai text. PyPDF2 and pdfplumber employ methods to extract text and characters directly from the file, ensuring high speed. However, if the PDF file is in image format, such as one created by scanning a document, PyPDF2 and pdfplumber will be unable to extract the text, as they do not support imagebased text extraction. To address this, pytesseract is also used to extract text. Pytesseract is an OCR (Optical Character Recognition) tool developed from Google’s Tesseract OCR Engine, which uses machine learning techniques for text and word recognition within images. The disadvantage of this method is that it requires more time for extraction compared

4.2.5 Keywords Identifying Website Types Examples include "site:go.th" and "site:ac.th". These keywords are used to restrict search results to specific types of websites, such as government or academic sites. For example, "site:go.th" filters search results to only those originating from government websites, as outlined in Section 2. 4.2.6 Other Related Search Terms Examples (all translated from Thai) include “name list”, “list of senior citizens”, “list of those eligible for allowances”, and “certificate of tax withholding”. These search terms were derived from analysis of search results and personal data found, and frequently recurring results were used to create new search

4.3 Personal Data Collection

Analysis of Personal Data Exposure in Thailand

9

Table 4 Example of Search Keywords Based on National ID Number Prefixes Search Keyword "1-1001-"

"3-2007-"

Meaning of Search Keywords Using Number Prefixes This search term targets National ID numbers for individuals born on or after January 1, 1984 (digit 1) AND National ID numbers of individuals born in the Phra Nakhon District of Bangkok (digits 2-5). This search term targets National ID numbers for individuals born before January 1, 1984 (digit 1) AND National ID numbers of individuals born in the Si Racha district of Chonburi province (digits 2-5).

to the other two methods. Our developed PDF parser extracts text using all three methods to ensure the maximum amount of text is collected. Spreadsheet Documents (.xls, .xlsx): The Spreadsheet (or Excel) parser was developed using the Pandas library and its read excel function [42], which allows users to select an engine for reading spreadsheet files. To ensure the Spreadsheet parser can read both old (.xls) and new (.xlsx) formats, the system uses three different engines: openpyxl, odf, and pyxlsb. This combination allows for the best text extraction from spreadsheet files. Word Processing Documents (.doc, .docx): The Word Processing (or Doc) parser was developed using Document Parser libraries, including (1) docx2txt [43], (2) textract [44], and (3) antiword [45]. To ensure the Doc parser can handle both old (.doc) and new (.docx) file formats, our system also uses all the aforementioned libraries to extract the maximum amount of text from the documents.

4.3.2 National ID Number Validation To ensure that only valid National ID numbers are stored, our system performs a validation check on every number extracted from the files during the text extraction process. This check is divided into the following steps: Step 1: Format Search and Verification. Since the Thai National ID numbers consist of 13 digits, with a clearly defined format (as discussed in Section 2.2). Therefore, we can prepare Regular Expression (RegEx) patterns for searching and extracting numbers that match the expected formats (e.g.number of expected digits, position of spaces and hyphens between digits). To ensure that our system can handle all Thai numerals, we extract these specific numerals and convert them into Arabic numerals (0-9). This conversion guarantees that they can be checked against the defined patterns. Only numbers that match the specified regex patterns are processed in the subsequent steps. Step 2: Checksum Validation. Once a number matches the format from Step 1, the system performs a checksum validation using the checksum rule for Thai National ID numbers, as explained in Section 2.2.2. Numbers that fail the checksum validation are discarded and not stored.

Step 3: Prefix Validation. This final validation step verifies the first five digits of the number that passed the Step 1 and 2, as discussed in Sections 2.2 and 2.2.2, to verify the following: – Digit 1: Must be a number between 1 and 8. – Digit 2-5: Must match the district or sub-district code in Thailand. This validation guarantees that all collected numbers accurately represent a Thai National ID number. 4.3.3 Personal Data Storage After the National ID number has been validated, the valid number is stored in the prepared database. Each stored National ID number is linked to its source, including the search engine used, the search keywords, the date and time of discovery, the URL of the source, and the location of the downloaded document. This information allows for precise tracking of the source of each ID number, aiding in analysis and interpretation of the results. Additionally, this metadata may be useful in case the information needs to be disclosed for legal requests or other regulatory purposes. Data Privacy. We are committed to protecting the privacy of all individuals whose data was collected during this research. The dataset is securely stored in an offline environment and safeguarded using industry-standard authentication protocols. Access is strictly limited to authorized personnel only. Under no circumstances will any personally identifiable information be disclosed. Further details regarding our ethical practices and responsibilities are provided in Section 8. 5 Adversary Model and Economic Analysis This section establishes the adversary model used to assess the threat landscape. We categorize potential attackers by their technical capabilities and the financial costs associated with large-scale data acquisition. 5.1 Adversary Capabilities We categorize potential attackers based on their technical resources and the realistic harm they can inflict using the exposed data.

Suphannee Sivakorn1 , Sasawat Malaivongs2 , Nuttaya Rujiratanapat1

10

Opportunistic Adversaries: Individual actors or “script kiddies” with limited technical resources. Their capability is restricted to manual or semi-automated search engine queries (Dorking). While the barrier to entry is near-zero, the harm is typically limited to small-scale identity theft or individual harassment. Systematic Scrapers: Organized groups with the infrastructure to perform bulk scraping and data indexing. These actors can automate the extraction of large numbers of identifiers and cross-reference them with other leaked databases (e.g., historical bank breaches). This enables systemic harm, such as building comprehensive “Personas” for large-scale financial fraud or state-level profiling.

5.2 Economic Analysis The effort required for exploitation is categorized by the “cost of acquisition,” which in this context is effectively negligible. This is primarily due to the ease of retrieving indexed search results. Web Scraping Retrieval: An adversary can implement a search engine crawler using headless automated web engines such as Selenium1 , allowing results to be retrieved gradually to bypass bot-detection mechanisms. However, as search engines increasingly implement stricter protections to prevent unauthorized scraping for AI training [46], this method may face scalability challenges in the long term. API-Based Retrieval: Alternatively, official or thirdparty Search Engine Results Page (SERP) APIs significantly lower the technical barrier. As of 2026, the Google Custom Search API [47] provides 100 free daily queries, with additional requests costing $5 per 1,000 queries. Similarly, third-party services such as SerpAPI2 offer tiered pricing ranging from $25 to $275 per month for 1,000 to 30,000 searches, respectively. These price points demonstrate that an attacker could aggregate over a million sensitive records for a financial investment of less than a few hundred dollars.

6 Analysis of Personal Data Leak This section provides a systematic analysis of the collected personal data and evaluates its impact on the security landscape of Thai citizens. Our methodology—conducted over a three-month window from March to May 2024, utilized 619 targeted search queries to identify 6,097 unique URLs. From these, we successfully extracted 1,263,268 unique National ID numbers from 6,004 documents. To put this in perspective, this figure represents approximately 2% of Thailand’s total population (65.9 million) [48]. 1 Selenium: https://www.selenium.dev 2 SerpAPI: https://serpapi.com

This volume of exposure suggests a significant systemic failure in data protection, as these identifiers are the primary keys for accessing government and financial services in Thailand. The following analysis categorizes these exposures by (1) Download File Type, (2) Registered Domain, and (3) Top-Level Domain and Ownership to identify the structural drivers of this leakage.

6.1 Threats to Validity Before detailing the results, it is necessary to address the inherent limitations and potential biases of the data collection process: Search Engine and Algorithmic Bias: The reliance on Google and Bing introduces platform-specific biases. Search results are governed by proprietary ranking algorithms and search filters, which may exclude certain exposed records from our view. Temporal Snapshot: Data collection occurred over a three-month snapshot. Because search engine indices are dynamic, our results reflect a specific temporal window and do not account for data that may have been de-indexed or exposed outside this period. Lower Bound Estimation: Most importantly, the figures presented in this section should be interpreted as a lower bound of the total exposure. Our methodology only captures indexed “surface web” documents; it does not account for the Deep Web, un-indexed databases, or files protected by robots.txt that remain accessible via direct links.

6.2 Download File Types Analysis of the downloaded file types (Table 5) reveals a significant correlation between technical format and the volume of exposed PII. While PDF and Word documents are more numerous in terms of unique URLs, they typically contain “narrative” or localized records, such as individual forms or small individual lists. In contrast, spreadsheet formats are structurally optimized for bulk data aggregation. From a security perspective, this disparity represents a critical Aggregation Risk. Although spreadsheets account for only 42.5% of files containing personal data, they are responsible for over 81% of the total extracted National ID numbers. Such exposure often stems from Security Misconfigurations (OWASP A02:2025) [49], specifically where default server settings allow directory listing or fail to implement proper access control headers on administrative staging folders.

Analysis of Personal Data Exposure in Thailand

11

Table 5 Number of Downloaded Files and National ID Numbers Collected from the Internet Categorized by File Type

File Type

File Ext.

Spreadsheet

.xls, .xlsx .pdf .doc, .docx

PDF Word Total

Downloaded Files URLs Containing Personal Data 2,370 1,467

Unique National IDs 1,032,236

3,298 429

1,938 48

241,093 4,062

6,379

3,453

1,263,268

“Unique National IDs” refers to unique Thai National ID numbers extracted from the datasets. Finding 1: A significant number of Thai individuals have had their National ID numbers and personal data exposed online. – Systemic Scale: Over three-months period (March to May 2024), our system gathered personal data through search queries, resulting in the extraction of 1.2 million National ID numbers accounting for nearly 2% of the Thai population. – Aggregation Risk: While spreadsheet documents represent roughly 42% of the downloaded files containing personal data, they account for over 81% (1, 032, 236) of the total extracted National ID numbers. This disparity highlights that the exposure of tabular, bulk-aggregated data is the primary driver of large-scale PII leakage.

for Thai government agencies, hosts the highest density of exposure, with over 756k unique National IDs. However, a critical interpretive finding is the substantial volume of IDs discovered on commercial (.com), academic (.ac.th), and organization (.org) TLDs, which collectively expose over 380k unique National IDs. Analysis of these results reveals a structural correlation between decentralized administration and TLD selection. While central departments generally adhere to .go.th protocols, many local government entities (provincial, district, and subdistrict administrations) frequently register domains under commercial TLDs like .com or .in.th for perceived ease of deployment. This creates a fragmented security perimeter where official civil registries, often containing the same depth of PII as centralized databases, are hosted on third-party infrastructure. This suggests that the vulnerability is not a flaw in the .go.th infrastructure itself, but rather a systemic lack of centralized security oversight for decentralized administrative units. In the following sections, we provide a granular analysis of the specific TLDs and organizational sectors that serve as primary vectors for PII exposure. By scrutinizing these recurring patterns of leakage, we aim to identify structural vulnerabilities and provide actionable recommendations to mitigate future data exposure. Finding 2: Majority of exposed Thai National ID Numbers were collected from government websites.

6.3 Registered Domain This section examines the registered domains serving as the primary vectors for PII exposure, revealing a significant institutional concentration of risk. As shown in Table 6, the top 10 domains alone account for over 547k exposed National ID numbers; however, a critical interpretive finding is the Cross-Domain Correlation between high-sensitivity data and non-government TLDs. Specifically, the most severe leak (Rank 1, 112k National ID numbers) originated from a .com domain (pokkrongnakhon.com), while the fourth largest (60k National ID numbers) occurred on a .org domain (chpao.org), both of which are owned by government agencies. This indicates that sensitive civil registries are frequently migrated to commercial or third-party hosting environments that may lack the rigorous security oversight of the official .go.th infrastructure. Consequently, the following section provides a deeper analysis of how these institutional failures manifest across different top-Level domains to identify broader patterns of exposure.

6.4 Top-Level Domain Distribution and Ownership Table 7 categorizes the exposure by TLD, highlighting significant variations in how sensitive data is distributed across the Thai web ecosystem. The .go.th domain, the official TLD

– PII Density: The .go.th domain exhibits the highest data density, accounting for 58.8% of all exposed IDs. This concentration reflects the centralized nature of government databases, where institutional leaks have a disproportionately higher impact than those in other sectors. – Cross-Domain Persistence: A significant volume of leaks on non-government TLDs (.com, .org) are traced back to local government entities. This suggests that the vulnerability is rooted in public-sector data handling practices rather than the security of the TLD infrastructure itself.

6.4.1 TLD: go.th The .go.th domain represents the core of the Thai government’s digital presence. This section scrutinizes the distribution of PII exposure across various ministerial bodies to identify which institutional sectors serve as the primary sources of leakage. As shown in Table 8, we identified a significant concentration of risk within three specific ministries: Interior, Education, and Agriculture. The Ministry of Interior accounts for the largest share of exposure, with over 409k unique National ID numbers. This is interpreted as a result of the ministry’s role as the primary custodian of civil and local administrative data. The Ministry of Education follows with approximately 234k unique National ID numbers, while the Ministry of Agriculture and Cooperatives accounts for 86k.

Suphannee Sivakorn1 , Sasawat Malaivongs2 , Nuttaya Rujiratanapat1

12

Table 6 Top 10 Registered Domain with the Highest Number of National ID Numbers Collected Rank Domain Name

Registered Domain Owner

URLs Downloaded Files

1 2 3 4 5 6 7 8 9 10

Nakhon Si Thammarat Provincial Administration Department of Learning Encouragement Department of Community Development Chaiyaphum Provincial Administration Organization Maha Sarakham Primary Educational Service Area Office (Area 3) Department of Fisheries Royal Thai Army Chachoengsao Primary Educational Service Area Office (Area 2) Royal Thai Navy Educational Data System Development Center

12 49 81 1 1 74 9 3 10 2

pokkrongnakhon.com nfe.go.th cdd.go.th chpao.org mkarea3.go.th fisheries.go.th rta.mi.th ccs2.go.th navy.mi.th edudev.in.th

12 49 81 1 1 74 9 3 10 2

FQDNs Unique National IDs 1 112,048 15 92,433 21 70,075 1 60,626 1 47,187 2 44,208 4 32,783 1 29,754 5 29,062 1 29,039

Table 7 Number of Download URLs, Files, and National ID Numbers Collected Categorized by TLD TLD go.th com ac.th org mi.th in.th N/A* or.th net ac co.th Others

Downloaded Files 2,305 159 579 36 20 37 33 201 14 9 53 7

FQDNs 983 76 311 18 10 18 1 38 11 7 19 5

Registered Domains 776 66 160 16 3 18 19 34 10 1 17 5

Unique National IDs 756,278 167,236 155,139 66,796 61,842 48,586 15,411 6,057 5,763 1,706 510 73

Registered Domain Examples nfe.go.th, cdd.go.th,mkarea3.go.th pokkrongnakhon.com, tecs4.com, sktcoop.com tupr.ac.th, ubu.ac.th, veis1.ac.th chpao.org, saraburi2.org, cupsakol.org rta.mi.th, navy.mi.th, tdc.mi.th edudev.in.th, ssk.in.th, abt.in.th 122.154.253.83, 122.155.168.174, 203.157.184.6 baac.or.th, nfcrbr.or.th, mea.or.th utdone.net, kkict.net, phsc.net thai.ac pea.co.th, skybook.co.th, pwa.co.th 1stdirectory.co.uk, thaiconsulate.jp,

National ID Numbers by Ministry Division in Ministry of Interior

Collectively, these three ministries are responsible for over 729k unique National IDs, or roughly 96% of all leaks within the .go.th domain. This high concentration suggests that PII exposure is not a uniform problem across the government, but is specifically clustered in agencies that manage nationwide beneficiary programs, educational registries, and local community datasets. The following subsections provide a granular analysis of these top-tier ministerial exposures to identify the specific document types and administrative processes driving these leaks. Ministry of Interior. Figure 2 presents the number of National ID numbers associated with government agency domains under various divisions of the Ministry of Interior. Notably, the majority of these domains belong to local government offices, including the Subdistrict Administrative Organization, Subdistrict Municipality Office, Province, Town Municipality Office, Provincial Local Administration Office, Provincial Administrative Organization, and City Municipality Office, collectively accounting for over 76% (314k) of all exposed National ID numbers within the Ministry. However, central government agencies, e.g. central government offices, also contribute significantly to the total number of exposed ID numbers, over 98k. To gain a deeper understanding of the root causes of these leaks, a manual audit of the top 100 high-impact files was conducted (Figure 3). Our analysis reveals a recurring

Subdistrict Administrative Organization Central Government Agency (e.g., Ministry, Departments, Offices) Subdistrict Municipality Office Province Town Municipality Office Provincial Local Administration Local Department of Provincial Administration Provincial Administrative Organization City Municipality Office 0

50000

100000

150000

Fig. 2 Distribution of Exposed National ID Numbers by Ministry Division within the Ministry of Interior

Institutionalized Disclosure Pattern: the majority of leaks are not the result of malicious breaches, but rather the publication of official administrative registries. Notably, over 60% of these files comprise lists of vulnerable populations including elderly allowance recipients, persons with disabilities, and disaster victims. From a security perspective, these documents facilitate High-Fidelity Identity Reconstruction. By publishing full names, National IDs, bank accounts, and home addresses alongside sensitive status indicators (e.g.disability or pregnancy), these agencies provide malicious actors with a complete profile for targeted social engineering. This is particularly concerning given the growing cybercrime landscape in

Analysis of Personal Data Exposure in Thailand

13

Table 8 Top 10 Government Ministries Registered under TLD “go.th” with the Highest Number of Exposed National ID Numbers Unique Domain Owner Examples National IDs Ministry of Interior 409,012 Community Development Department, Department of Local Administration, Juab Subdistrict Administrative Organization Ministry of Education 234,400 Department of Learning Encouragement, Maha Sarakham Primary Educational Service Area Office (Area 3), Chachoengsao Primary Educational Service Area Office (Area 2) Department of Fisheries, Department of Agricultural Extension, Ministry of Agriculture and Cooperatives 86,218 Department of Royal Irrigation Ministry of Public Health 10,251 Ministry of Public Health, Kantharawichai Hospital, National Institute for Emergency Medicine Ministry of Transport 5,553 Department of Highways, Department of Rural Roads Ministry of Labour 5,533 Ministry of Labour, Department of Employment, Department of Skill Development Government Agencies not under the Prime 3406 National Office of Buddhism, Royal Thai Police, Anti-Money Laundering Minister’s Office, Ministries, or Departments Office Office of the Prime Minister 1,617 Office of the Official Information Commission, Open Government Data of Thailand, Office of the Permanent Secretary The Excise Department, The Revenue Department, Ministry of Finance 1,279 The Customs Department Ministry of Natural Resources 1,198 Ministry of Natural Resources and Environment, Department of Groundwater and Environment Resources, Department of Mineral Resources

Rank Domain Owner Ministry 1 2

3 4 5 6 7 8 9 10

Thailand; the exposure of phone numbers (found in 5% of the sampled documents) and income levels directly enables the “call center” scamming syndicates currently targeting Thai citizens. Ultimately, these findings suggest a fundamental conflict in local administrative processes: the pursuit of public transparency regarding government spending and subsidies currently lacks the technical safeguards (such as data masking or access-controlled portals) necessary to protect the constitutional privacy of the beneficiaries. Ministry of Education. The Ministry of Education ranks as the second-largest source of PII exposure, characterized by a distinct demographic risk, the large-scale leakage of data belonging to minors. As detailed in Table 9, over 133k unique National IDs were exposed through the Office of the Basic Education Commission, primarily via Primary and Secondary Educational Service Area Offices. For example, sensitive student registries from Prathom (grades 1–6) and Mathayom (grades 7–12) levels were found publicly indexed in regions such as Sukhothai and Maha Sarakham. This is particularly concerning as these data involve not only personal information but also the personal details of minors, which means the leakage affects children from a young age. Mirroring the patterns observed in the Ministry of Interior, the leakage is driven by decentralized administrative systems. Local educational commissions and vocational offices regularly upload student lists such as enrollments and registries, without adequate data-masking protocols. A particularly egregious discovery was the exposure of data from Special Education Centers, which links National IDs to sensitive indicators of special needs or disabilities.

Ministry of Agriculture and Cooperatives. The Ministry of Agriculture and Cooperatives represents the thirdlargest institutional source of PII exposure, characterized by the leakage of comprehensive agrarian registries. As shown in Table 10, the Department of Fisheries and the Department of Agricultural Extension are the primary contributors, collectively exposing over 62k unique National IDs. Our qualitative analysis reveals a unique and highly invasive data pattern, the coupling of PII with geospatial and asset Metadata. For example, registries published by the Department of Fisheries do not merely list National ID numbers and names; they include precise geographical coordinates (latitude and longitude), farm dimensions, and livestock types. From a security perspective, this creates a Physical-Digital Risk Linkage, where a malicious actor can not only identify an individual but also remotely assess their physical assets, land value, and precise location. Furthermore, the discovery of agricultural subsidy lists containing bank account information and the exposure of retired officer registries mirroring the patterns in the Ministry of Interior. The inclusion of professional metadata, such as occupation and farm registration dates, enables sophisticated spear-phishing campaigns, where scammers can pose as government officials to “verify” agricultural grants or land reform status, leveraging the victim’s data to build trust. Ministry of Public Health. The Ministry of Public Health exhibits a lower total volume of unique National ID numbers (approx. 10.5k) compared to other ministries, yet the qualitative sensitivity of the exposed data is significantly higher. As shown in Table 11, leaks are distributed across central agencies and provincial hospitals, representing a direct

Suphannee Sivakorn1 , Sasawat Malaivongs2 , Nuttaya Rujiratanapat1

14 List of Eligible Recipients for Elderly Allowance Payments

List of Elderly Individuals

List of Lands and Properties

List of Disabled and Elderly Individuals List of Members of the Role Development Fund List of Disaster Victims List of Recipients of Compensation for the Impact of the Coronavirus Pandemic Report on Community Household Data List of Eligible Candidates for Skills and Knowledge Examinations

Others

Gender

Payment Amount

100

75

50

25

0 National ID Full Name Number

Address

Phone Number

Birthday

Age

Bank Account

Occupation

Income

Fig. 3 Top 10 Most Frequently Occurring Document Title in the 100 Files Containing the Highest Number of Exposed National ID Numbers from Domains under the Ministry of Interior Table 9 Ministry Divisions under the Ministry of Education with the Highest Number of Exposed National ID Numbers Unique Domain Owner Examples National IDs Office of the Basic Education Commission 133,396 Sukhothai Secondary Educational Service Area Office, Chiang Mai Primary Educational Service Area Office (Area 4), Maha Sarakham Secondary Educational Service Area Office Central Government Agencies 94,540 Department of Learning Encouragement Office of the Private 3,677 Narathiwat Office of the Private Education, Songkhla Office of the Private Education, Education Commission Yala Office of the Private Education Office of the Vocational 2,284 Office of the Vocational Education Commission Education Commission Provincial Education Office 652 Rayong Provincial Educational Office, Trang Provincial Educational Office, Loei Provincial Educational Office Special Education Center 53 Chonburi Special Education Center (Area 12) District Learning Encouragement Center 1 Sung Men District Learning Encouragement Center Ministry Division

Table 10 Top 10 Ministry Divisions under the Ministry of Agriculture and Cooperatives with the Highest Number of Exposed National ID Numbers Unique Rank Ministry Division National IDs 1 Department of Fisheries 44,008 2 Department of Agriculture Extension 18,907 3 Department of Royal Irrigation 8,154 4 Office of the Permanent Secretary for Ministry 4,448 of Agriculture and Cooperatives 5 Department of Livestock Development 3,192 6 Department of Cooperative Auditing 2,292 7 Department of Sericulture 1,770 8 Department of Land Development 1,521 9 Office of Agricultural Land Reform 1,143 10 Department of Royal Forest 955

compromise of the intersection between civil identity and healthcare metadata. From the top 10 files that exposed the highest number of National ID numbers, we discovered several alarming instances of personal data leak. For instance, we identified a list of over 2,200 government officers, which included their full names, National ID numbers, workplace and salaries. Additionally, we found approximately 3,000 Thai withholding tax documents, which typically contained full names, National

Table 11 Ministry Divisions under the Ministry of Public Health with Highest Number of Exposed National ID Numbers Ministry Division Central Government Agencies Hospital Public Organization Regional Health Promotion Center Local Public Health Office Regional Health Provider Office

Unique National IDs 5,689 2,635 973 562 530 113

ID numbers, addresses, and wages. The most critical finding, however, is the exposure of pediatric registries containing the National ID numbers of approximately 500 children. Unlike previous educational leaks, these records include clinical metadata, such as medication information. The exposure of a minor’s medical history alongside their permanent National ID number creates a permanent, non-remediable privacy violation. Such data is highly sought after for insurance fraud or sophisticated social engineering. Other Government Ministries As present in the Table 8, other ministries, including the Ministry of Labour, Ministry of Transport, and various government agencies not directly under the Prime Minister’s Office, were found to have exposed approximately 20k National ID numbers. This constitutes

Analysis of Personal Data Exposure in Thailand

about 5% of the total National ID numbers disclosed across government websites with the .go.th domain. To further analyze the scope of this data breach, we identified the top 10 files containing the highest number of exposed National ID numbers. These files included sensitive information such as lists of retired government officials, individuals receiving government subsidies due to unemployment, monks and their affiliated temples, recent graduates, individuals participating in government project bidding, and individuals who had received government compensation for injuries sustained during political violence. Other personal data were also exposed. Some of these documents contained additional sensitive details, including phone numbers, home addresses, height, weight, health conditions, salaries, and ages. Particularly concerning were those documents that linked individuals’ personal information to their political affiliations, raising significant privacy and security concerns. Finding 3: The exposure of personal data on local government websites presents serious privacy and security concerns. – Decentralized Risk in Local Government: The majority of exposed data originates from the Ministry of Interior, Education, Agriculture, Public Health, and Labour. Notably, the leakage is driven by local government agencies (e.g.subdistrict offices and provincial service areas), suggesting that decentralized administrative units lack the robust security oversight present in central ministries. – Multidimensional Identity Exposure: In addition to National ID numbers, the datasets include phone numbers, bank details, and income levels. This “data clustering” creates a high-fidelity profile for each victim, significantly lowering the barrier for highly targeted financial fraud and social engineering (e.g., call center scams). – Sensitive Social Indicators: The exposure of data related to disabilities, pregnancy, and political affiliations represents a profound breach of privacy. Such sensitive indicators could lead to social discrimination or predatory targeting, moving the risk from mere identity theft to potential long-term human rights implications.

6.4.2 TLD: com The .com TLD represents the second-largest vector of exposure, accounting for 13% (167k) of all unique National IDs. While traditionally reserved for commercial enterprise, a substantial portion of personal data exposure, as highlighted in Finding 1, originates from government agency websites, as presented in Table 12. Many of these domains remain active or have not been properly decommissioned after their intended use, for example, Chachoengsao Primary Educational Service Area Office (Area 2) website. As a result, certain files containing sensitive data remain accessible online. Another notable pattern is the storage or hosting of personal information on external websites, such as filethaischool1.com and wordpress.com, or file-sharing services like filesesbuy.com.

15 Table 12 Top 10 Website Categories of Websites Registered under .com TLD with Highest Number of Exposed National ID Numbers Rank Website Category Unique National IDs 1 Government and Legal Organizations 128,955 2 Education 12,979 3 Finance and Banking 8,369 4 Dynamic DNS 6,195 5 Business 6,054 6 Health and Wellness 3,409 7 Web Hosting 695 8 Newsgroups and Message Boards 314 9 File Sharing and Storage 152 10 General Organizations 123

Business Sector Leakage and Data Spillover. To understand the extent of personal information exposure within the business sector, we focused our analysis on websites categorized under the Business website category. We identified a concerning trend of government data spillover. Business websites were found hosting documents that should theoretically be restricted to state registries, including property ownership and disability pension recipient lists. In contrast, a more justifiable instance of personal information exposure within the business sector is the publication of employee lists within an organization, as such records are typically maintained for internal administrative purposes. 6.4.3 TLD: ac.th The .ac.th Top-Level Domain, designated for academic institutions, represents 12% (155k) of the unique National ID numbers in our dataset. To gain a deeper understanding, we categorized the .ac.th domains based on institution types, following the classification by the Ministry of Higher Education, Science, Research, and Innovation (MHESI) [50]. As shown in Table 13, the highest volume of exposure (over 62k IDs) originates from Primary and Secondary Schools (Grades 1–12), followed by autonomous and public universities. This distribution highlights a critical early-stage identity compromise, where sensitive PII is leaked at the very beginning of an individual’s civic life. Furthermore, we identified the top 10 documents with the highest number of exposed National ID numbers and conducted a content analysis to determine additional leaked data. These documents typically contain students’ full names, along with their academic level, class, and grade. However, some documents also disclose highly sensitive personal details, including date of birth, nationality, race, religion, disability status, height, weight, blood type, phone number, and home address even for minors under the age of 18. Additionally, some records include parental information, further exacerbating privacy concerns. Additionally, we observed a concerning pattern where university students who have taken out student loans are identifiable within the leaked data. This is particularly sensitive

16

Suphannee Sivakorn1 , Sasawat Malaivongs2 , Nuttaya Rujiratanapat1

as it publicly reveals individuals who are in financial debt, further extending the scope of exposed personal information.

exposures originate from the government sector. Specifically, local government agencies that registered their domains outside the .go.th TLD contribute significantly to these leaks. Interestingly while some domains do not officially 6.4.4 TLD: org belong to government agencies such as abt.in.th, thai.ac, and thaischool.in.th, they are typically still connected to the govThe .org TLD is primarily designated for organizations, ernment sector, primarily serving as website hosting services particularly non-profit entities. Notable websites using the for local government agencies and school websites. .org extension include the Red Cross and Wikipedia. However, IP-address websites account for 15k unique National an interesting observation is that a significant portion of this ID numbers, representing approximately 1% of all exposed TLD in Thailand are still associated with government sectors. National ID numbers identified in our analysis. In total, there Examples include websites of the Chaiyaphum Provincial are 34 IP addresses that contain personal information. We Administrative Organization, Saraburi Primary Educational audited the top 10 source IPs with the highest number of Service Area Office (Area 2), and Local Public Health Offices. exposed National ID numbers and performed a reverse lookup Mirroring the vulnerabilities of the .go.th sector, these on their IP Whois registration information to gather details .org domains frequently host non-anonymized beneficiary about their geolocation and registered owners. Our findings registries for allowance payments, confirming a systemic revealed that all 10 IPs are located in Thailand. In some cases, failure in data privacy standards regardless of the TLD. the domains are linked to government sectors, such as the Corporate and Other Organizations. Beyond governmentMinistry of Public Health and the Ministry of Higher Educaaffiliated .org domains, we also identified personal data extion, Science, Research, and Innovation. We also identified posure in documents hosted by corporate entities and other that some file also contains full names, dates of birth, ages, organizations. These documents contain full names and Naand addresses. All these top-10 documents are associated with tional ID numbers, along with other sensitive information various government-related duties, such as lists of children such as passport numbers and marital status, further highin compulsory education and the survey list of individuals lighting the widespread risk of personal data leakage across with risk behaviors for non-communicable diseases. This furdifferent sectors. ther highlights the extensive scope of personal data exposure within government-sector documents. 6.4.5 TLD: mi.th The .mi.th TLD is designated for domains associated with the Thai military. Our analysis revealed that .mi.th-registered domains ranked fifth among all identified TLDs in terms of personal data exposure, accounting for 4.8% of all exposed National ID numbers. The majority of these leaks originated from two primary domain owners: the Royal Thai Army and the Royal Thai Navy, with over 32k and 29k records exposed online, respectively. Our qualitative analysis of high-impact military files reveals a pattern of institutionalized structural exposure. These files exclusively pertain to military personnel records, containing names, National ID numbers, and military ranks. Notably, some documents also include lists of personnel assigned to specific missions and sub-organizational units, information that could be considered sensitive in terms of military and national security confidentiality.

Finding 4: Most personal data exposures stem from the government sector even when hosted on non-government domains. – Public Sector Data Gravity: While non-government TLDs (e.g. .com, .org, .in.th) exhibit high exposure, the owners are almost exclusively local government agencies. This indicates the vulnerability is rooted in public-sector data practices rather than TLD infrastructure. – Activity-Driven Exposure: The highest-density leaks are strongly correlated with mass administrative activities, such as welfare distribution, student registration, and public health tracking. This confirms that bulk-data processing is the primary risk vector for large-scale exposure. – Lifecycle Management Failure: Manual review reveals that many exposed records belong to legacy websites. This indicates a “long-tail risk” where data remains indexed long after its operational necessity has passed, highlighting a systemic lack of data decommissioning protocols.

6.5 Search Keyword and Search Engine 6.4.6 Other TLDs and IP Addresses Other TLDs such as .in.th, .or.th, and .net account for approximately 62k unique National ID numbers, representing around 5% of all exposed National ID numbers identified in our analysis. Our findings indicate a recurring pattern observed in other TLDs, where the majority of data

In this section, we analyze the correlation between search query structures and the volume of exposed unique National ID numbers. Our results demonstrate that search engines, particularly Google, which accounted for over 99% of all retrieved data, act as highly efficient, unintentional aggregators of PII. The most effective queries utilized a combination

Analysis of Personal Data Exposure in Thailand

17

Table 13 Categorization of ac.th TLD Registered Domains by Institution Type and Their Respective Exposed National ID Numbers Institution Type Primary and Secondary School (Grade 1-12) Autonomous University Public University Rajamangala Universities of Technology Rajabhat University Vocational College Non-affliated Academic Institution under the MHESI Private University Private College Private Institution Open Public Universities

Unique National IDs 62,929 32,856 18,645 14,752 11,591 10,172 2,774 962 625 40 5

Fig. 5 Number of National ID Numbers Grouped by Individual Categories (National ID Number First Digit) Finding 5: Google search is the primary source of exposed National ID numbers driven by specific operations and keywords. – High-Value Query Combinations: Exposure is driven by the synergy between filetype-specific operators and Thai-specific PII keywords. This confirms that targeted Google dorking remains a low-effort, high-reward vector for mass data harvesting. – Search Rank Criticality: The highest concentration of PII appears on the initial search result pages. This indicates that the most significant vulnerabilities are not hidden in the “long-tail” of search results but are highly ranked, maximizing exposure to even unsophisticated actors. – Platform Dominance: While both engines were used, Google’s superior indexing of Thai government subdomains makes it the primary discovery source, underscoring the need for search-engine-specific remediation.

Fig. 4 Association between Result URLs containing Personal Information and Result Page Number

of filetype operators (e.g.spreadsheets or PDFs) and Thailanguage PII identifiers, suggesting that mass data harvesting requires minimal technical sophistication. Table 14 summarizes the top-performing keyword categories. To ensure ethical compliance and prevent the replication of these findings for malicious use, we have redacted specific numeric patterns and exact keyword strings. The data reveals a synergistic query effect, the combination of contextual Thai-language terms commonly found alongside National ID numbers, domain-specific constraints (e.g. site:go.th) and specific file extensions (e.g. .xlsx) yielded the highest density of unique National ID numbers. Additionally, we analyzed the ranking distribution of search results to understand where exposed personal data appears on result pages. As illustrated in Figure 4, we observed that PII-heavy URLs are not buried in the “long-tail” of search results; instead, they are frequently indexed within the first 10 pages. This emphasizes that the primary risk is not just the existence of the files, but their high search-rank accessibility, which maximizes exposure to even unsophisticated actors.

6.6 Demographic Profile: Individual Category (National ID Number First Digit) To move beyond a simple volume count, we analyzed the first digit of the unique National ID numbers to reconstruct the demographic profile of the exposed population. As detailed in Section 2.2, the first digit serves as a proxy for age and registration status. As illustrated in Figure 5, over 61% of the collected IDs begin with the digit 3. This cohort represents Thai nationals and long-term residents registered before May 1984, meaning the majority of these individuals are aged 41 or older as of 2025. This age skew is highly significant from a security perspective; this demographic often possesses higher financial assets while simultaneously exhibiting lower digital literacy, making them the primary targets for sophisticated social engineering and “call center” scams currently prevalent in Thailand. The second-largest cohort (33%) consists of individuals with the first digit 1, representing Thai nationals born after 1984. Combined, these two categories account for over 93% of the total exposure. This concentration suggests that the current state of PII leakage in Thailand creates a multi-generational security crisis, where the elder population faces immediate financial risk, while the youth population faces a compromised digital future.

Suphannee Sivakorn1 , Sasawat Malaivongs2 , Nuttaya Rujiratanapat1

18

Table 14 Top 10 Keyword Categories Associated with the Highest Number of Exposed URLs and National ID Numbers Search Unique URLs Engine National IDs 1 National ID number term filetype:xlsx google 40 147,345 2 filetype:xls OR filetype:xlsx site:go.th National ID number term Name prefix term google 141 115,783 3 filetype:xlsx OR filetype:xls Partial National ID number* google 10 107,620 4 filetype:pdf National ID number term Partial National ID number* google 150 70,085 5 filetype:xls OR filetype:xlsx site:ac.th National ID number term Name prefix term google 80 55,886 6 National ID number term filetype:pdf google 104 51,672 7 National ID number term filetype:pdf google 128 51,267 8 filetype:xls OR filetype:xlsx (National ID number terms) Partial National ID number* google 1 47,187 9 site:in.th filetype:xlsx OR filetype:xls National ID number term Name prefix term google 12 46,167 10 National ID number term filetype:xls google 34 41,913 *The exact numbers used have been omitted to prevent the search keywords from being replicated to obtain personal information Rank Search Keyword

Finding 6: Senior individuals and those over 41 represent the majority of exposed personal data.

Table 15 Top 10 Provinces with the Highest Number of Personal Data Exposure

– High-Risk Demographic Skew: Over 60% of exposed individuals are aged 41 or older. This demographic skew indicates that the leaks primarily affect a population segment that is often targeted by social engineering and financial fraud due to lower digital literacy. – Generational Impact: The inclusion of 33% of individuals under 41–including student and minor records– highlights a long-term identity theft risk that may persist for decades as these individuals enter the workforce and financial systems.

Rank Province 1 2 3 4 5 6 7

6.7 Geographic Distribution: Province and District To identify the regions most acutely impacted by PII exposure, we mapped the collected National ID numbers to their geographic origins by extracting digits 2–4, which correspond to the registration province and district. By benchmarking these findings against the 2024 population statistics from the Bureau of Registration Administration [48], we identified a significant geographic variance in data exposure across Thailand. As presented in Table 15, Nakhon Si Thammarat recorded the highest absolute volume of exposure (136k unique National ID numbers), while Bangkok followed with 70k. However, when normalized against population size (Table 16), a different risk profile emerges: Satun province exhibited the highest per-capita exposure, with over 14.8% of its total population affected. This suggests that while major urban centers like Bangkok generate high volumes of data, smaller provincial administrations may suffer from higher institutional vulnerability densities, where a single leak impacts a disproportionately large segment of the local citizenry. District-Level Saturation. To gain deeper insights, we conducted the same statistical analysis at the district level. Table 17 presents the top 10 districts with the highest number of personal data exposures identified in our study. Notably, five districts from Nakhon Si Thammarat exhibit exposure

8 9 10

Nakhon Si Thammarat Bangkok Chaiyaphum Satun Khon Kaen Nakhon Ratchasima Maha Sarakham Chachoengsao Chiang Mai Kalasin

Province Unique Population1 Code National IDs

%

80

136,375

1,531,727

8.90

10 36 91 40

70,302 65,334 48,057 43,602

5,352,831 1,105,008 324,390 1,768,366

1.31 5.91 14.81 2.47

30

42,135

2,615,039

1.61

44

40,776

929,056

4.39

24 50 46

34,699 34,445 29,749

729,218 1,635,983 961,369

4.76 2.11 3.09

Table 16 Top 10 Provinces with the Highest Number of Personal Data Exposure, Ranked by the Percentage of Exposed Data Relative to the Population1 Rank Province 1 2 3 4 5 6 7 8 9 10

Satun Nakhon Si Thammarat Chaiyaphum Chachoengsao Maha Sarakham Nan Narathiwat Phatthalung Kalasin Sukhothai

Province Unique Population1 % Code National IDs 91 48057 324390 14.81 80

136375

1531727

8.90

36 24

65334 34699

1105008 729218

5.91 4.76

44

40776

929056

4.39

55 96 93 46 64

17134 28032 17182 29749 17241

468670 822827 519103 961369 572575

3.66 3.41 3.31 3.09 3.01

rates ranging from 8% to 19% of their district populations, while two districts from Satun show the highest exposure rates, ranging from 19% to 21%. These findings highlight significant regional disparities in data exposure. Table 18 ranks the top 10 districts based on the percentage of exposed personal data relative to their current population. Notably, Sanam Chai Khet Sub-district Municipality in Cha-

Analysis of Personal Data Exposure in Thailand

choengsao province exhibits an exposure rate that exceeds its recorded population. Furthermore, Chiang Yuen Sub-district Municipality and Kosum Phisai Sub-district Municipality in Maha Sarakham province have exposure rates surpassing 50% of their respective populations, highlighting significant vulnerabilities in these areas. These saturation leaks are interpreted as a failure of local administrative registries, where entire community databases, including historical records and non-resident registrants, have been published en masse. Finding 7: Specific provinces and sub-districts exhibit localized “saturation leaks” with exposure rates exceeding 50% of their populations. – Regional Exposure Saturation: While Nakhon Si Thammarat has the highest volume (136k), Satun exhibits a higher per-capita risk, affecting 14% of its population. This indicates that smaller provinces may face higher systemic exposure relative to their size. – Sub-district Vulnerability Peaks: In specific municipalities (e.g. Sanam Chai Khet, Chiang Yuen), exposure rates exceed 50% of the local population. These “saturation leaks” suggest that a single administrative error can compromise an entire local community’s privacy.

19 Finding 8: Most personal information exposure comes from a single source but repeated leaks across multiple sources heighten the risk of misuse. – Single-Source Criticality: Over 90% of the exposed National ID numbers originate from a single source URL, reflecting a significant concentration of data exposure in a limited number of sources. – Cross-Source Correlation Risk: While most individuals appear in only one leak, a subset appears across multiple URLs. This multi-source exposure allows attackers to cross-reference fragmented data facilitating more sophisticated social engineering.

7 Countermeasures and Discussion Our work highlights the privacy implications of personal data exposure in public domains, demonstrating the severity and widespread nature of such incidents.

7.1 Strengthening Government Data Policies and Governance

6.8 Repeated Exposure: Fragmented Identity Aggregation In this section, we analyze the frequency of an individual’s personal data appearing across multiple data sources. The aim is to evaluate the extent of repeated exposure, which could heighten the risk of further data exposure, enabling potential attackers to gather more detailed information about individuals. By cross-referencing National ID numbers as identifiers across sources, we determine how often an individual’s data appears on different sources. Table 19 presents the distribution of National ID numbers exposed across 1 to 15 distinct source URLs, revealing that over 90% of the exposed data comes from a single source URL. We analyzed National ID numbers appearing across 15 distinct sources. We observed that, beyond the basic personal data such as full name, National ID number, and gender, additional data points such as place of work, job position, educational background, salary and contracts were also exposed. This additional information significantly increases the risks associated with data exposure, potentially allowing malicious actors to gain more comprehensive profiles of individuals.

1 Population Survey as of December 2024 [48].

As demonstrated in Section 6, personal data exposure is both prevalent and large-scale, with the majority of leaks originating from government websites especially those operated by local administrative offices. While publishing certain personal information may support transparency and operational efficiency (e.g. publishing lists of welfare recipients or school registrants), it must be handled with far greater caution. As shown, these disclosures can be exploited by malicious actors. Website administrators, especially at local government levels, may lack awareness of data subjects’ rights under the Personal Data Protection Act (PDPA). Consequently, they may publish sensitive information without consent. This underscores the urgent need for the government to enforce stronger privacy practices, where key measures include: – Educating local web administrators on PDPA compliance. – Providing clear standard guidelines for personal data disclosure. – Mandating the use of anonymization or pseudonymization techniques when personal data must be published for transparency [51]. Additionally, as discussed in Section 6.4, many exposed documents were found on outdated, decommissioned, or inactive websites. Since such websites may not be actively maintained, exposed personal data could remain accessible for years until the domains expire. This highlights the need for a comprehensive data governance framework to ensure responsible data lifecycle management from collection and storage to decommissioning and deletion.

Suphannee Sivakorn1 , Sasawat Malaivongs2 , Nuttaya Rujiratanapat1

20

Table 17 Top 10 District with the Highest Number of Personal Data Exposure Province Code District

Rank Province 1 2 3 4 5 6 7 8 9 10

Nakhon Si Thammarat Satun Nakhon Si Thammarat Satun Chaiyaphum Khon Kaen Maha Sarakham Nakhon Si Thammarat Nakhon Si Thammarat Nakhon Si Thammarat

80 91 80 91 36 40 44 80 80 80

District Code

Nakhon Si Thammarat City Satun City Tha Sala La-ngu Chaiyaphum City Khon Kaen City Kosum Phisai Ron Phibun Chawang Thung Song

8001 9101 8008 9105 3601 4099 4403 8013 8004 8009

Unique Population1 % National IDs 17,599 161,786 10.88 14,183 66,389 21.36 13,592 113,346 11.99 13,363 68,932 19.39 12,246 135,262 9.05 12,056 98,880 12.19 10,959 107,980 10.15 10,110 61,153 16.53 9,858 52,103 18.92 9,593 115,660 8.29

Table 18 Top 10 District with the Highest Number of Personal Data Exposure, Ranked by the Percentage of Exposed Data Relative to the Population1 Rank Province 1 2 3 4 5 6 7 8 9 10

Province Code District

Chachoengsao Maha Sarakham Maha Sarakham Chachoengsao Satun Surat Thani Satun Phatthalung Satun Nakhon Si Thammarat

24 44 44 24 91 84 91 93 91 80

District Code

Sanam Chai Khet Sub-district Chiang Yuen Sub-district Kosum Phisai Sub-district Thung Sadao Sub-district Khuan Don Surat Thani City Satun City Khuan Khanun Sub-district La-ngu Hua Sai

Table 19 Distribution of Exposed National ID Numbers Across Multiple Source URLs Source URLs Unique National IDs % 15 5 0.0004 13 6 0.0005 12 56 0.0044 11 7 0.0006 10 7 0.0006 9 61 0.0048 8 79 0.0063 7 774 0.0613 6 3795 0.3003 5 3332 0.2637 4 9584 0.7584 3 16229 1.2843 2 89890 7.1138 1 1139443 90.2008

2481 4494 4496 2480 9102 8401 9101 9395 9105 8016

Unique Population1 % National IDs 6983 4,231 165.04 3067 4,481 68.44 4964 8,834 56.19 1523 5,782 26.34 5019 22,228 22.58 4223 19,177 22.02 14183 66,441 21.35 394 1,962 20.08 13363 69,137 19.33 8178 43,277 18.90

7.3 Search Engine Indexing Controls As outlined in Section 3.1, search engines index websites based on content, making personal information retrievable via targeted queries. As a last-resort mitigation strategy, web administrators can use the noindex directive to prevent pages from being indexed by search engines [52]. While this does not block direct access to exposed content, it increases the difficulty for malicious actors to locate personal data via search engines, buying valuable time for data owners or agencies to remediate exposure.

7.2 Centralized and Secure Services 7.4 Legal Accountability and Transparency To minimize risks, personal data verification services should be offered via secure, centralized platforms rather than through loosely governed, distributed local websites. A governmentbacked centralized service with mandatory access controls can ensure that personal data is only accessible to authorized users. This approach would reduce the need for each local agency to publish personal information and would offer a safer, more reliable digital service model aligned with national e-government initiatives.

The government must uphold transparency and enforce legal accountability when personal data is leaked. Section 2.3 details several previous incidents that should have led to investigations and clear consequences under PDPA. Enforcing these laws uniformly across both government and private sectors will incentivize organizations to adopt better data protection practices, as seen following GDPR enforcement in the European Union [53].

Analysis of Personal Data Exposure in Thailand

7.5 Proactive Monitoring While strong policies and governance frameworks are essential for long-term change, proactive monitoring is crucial to detect and mitigate leaks in real time. Search Engine Monitoring. As demonstrated in this study, search engines play a central role in exposing personal data. Stakeholders should proactively monitor search engine indexes using specialized tools and keyword tracking. In cases of exposure, search engines can be contacted to request removal of indexed results or cached content [54, 55]. This serves as a critical first response measure, allowing agencies to initiate takedown procedures with website owners before further harm occurs. Threat Intelligence Monitoring. As noted in Section 2.3, exposed personal data often appears on underground forums and dark web marketplaces. Monitoring these platforms is vital for early detection. Government agencies and security teams should invest in threat intelligence services capable of scanning and analyzing these sources, enabling timely interventions before data is weaponized for fraud, phishing, or identity theft.

21

retrieved documents and extracted National ID numbers are stored in a secure, offline environment. This “security-first” storage protocol ensures that collected PII remains inaccessible to external networks, effectively eliminating the risk of a secondary breach during the research lifecycle. Respect for Autonomy through Anonymization: While the scale of the exposed data rendered individual informed consent infeasible, we uphold the privacy and dignity of the affected citizens through rigorous data hygiene. We are committed to ensuring that no PII is disclosed to third parties. Only fully anonymized metadata, such as registered domain names and TLD distributions, will be shared, and exclusively for the purposes of advancing security research and mitigating future leakages. No raw personal data will be disclosed under any circumstances. Systemic Evidence Base: The collection of 1.2 million records was conducted not for individual profiling, but to establish the scale of a systemic national vulnerability. The public interest in identifying and remediating a flaw that affects nearly 2% of the Thai population outweighs the transient risk of a controlled, academic analysis. Our findings serve as a critical evidence base for the development of automated monitoring systems and the enforcement of Thailand’s PDPA.

8 Ethics and Responsible Disclosure 8.3 Responsible Disclosure and Takedown Rationale Given the sensitive nature of the National ID numbers and personal sensitive data identified in this study, our methodology adheres to the ethical principles outlined in the Menlo Report [56], focusing on public interest and harm minimization.

8.1 Ethical Justification and Proportionality The collection of this data is justified by the significant public interest in identifying systemic vulnerabilities in Thailand’s digital infrastructure. The scale of the collection (over 1.2 million records) was necessary to demonstrate the proportionality of the threat; a smaller sample size would not have sufficiently illustrated the systemic risk posed by search engine indexing. This study did not involve interaction with human subjects; rather, it analyzed publicly available, albeit sensitive, metadata to identify security failures in government and public sectors.

We followed a responsible disclosure process rather than an immediate public takedown to ensure systemic remediation. Immediate notification to thousands of individual domain owners was deemed impractical and potentially harmful, as it might alert malicious actors to the vulnerability before a systemic fix was in place. Instead, we have prioritized: Institutional Notification: We are in the process of coordinating with the relevant Thai government authorities to provide them with the comprehensive list of vulnerable URLs for coordinated remediation. Strategic Mitigation: By providing these insights to the government sector, we facilitate the development of automated monitoring systems that can prevent future indexing of sensitive files and information, addressing the root cause.

9 Related Work 9.1 Public Sector and Government Data Risks

8.2 Harm Mitigation and Data Handling Our methodology prioritizes the principles of Beneficence (minimizing harm while maximizing public benefit) and Respect for Persons: Harm Minimization via Air-Gapped Storage: In accordance with Menlo Report guidance on data protection, all

Our study reveals that personal data exposure from Thai government websites is not an isolated issue, but part of a broader global challenge tied to digital transformation. As governments worldwide adopt e-Government (e-gov) initiatives to improve efficiency and service delivery [57], they also collect and process vast amounts of personal data. Without strong safeguards, these systems can inadvertently introduce new

22

risks, making them susceptible to data breaches and privacy violations. Several studies have examined the implications of personal data exposure in the public sector. For instance, Zulfiani (2023) reviewed data leakage incidents in Indonesia and emphasized the urgent need to enhance personal data protection [4]. Similarly, India’s Aadhaar identity system, which stores biometric and demographic data of over a billion residents, has faced repeated scrutiny over data leaks [5, 6]. These cases underscore the challenges of maintaining data security in centralized government systems, which parallels our findings on National ID number exposure in Thailand. At a regional level, Chaipipat et al. [58] examine the fragmented legal frameworks for data protection across ASEAN, highlighting the absence of a unified, enforceable privacy regime. The COVID-19 pandemic further amplified privacy concerns globally. Governments collected massive volumes of personal data for vaccination tracking, health monitoring, and contact tracing often under urgent timelines and minimal oversight. Reports from China [59], Indonesia [60, 61], India [62], and South Korea [63] highlight how emergency responses led to unintended data leaks, public exposure of health records and third-party data sharing. Thailand also faced similar incidents, including the exposure of personal data through vaccination registration application [13], and public releases of recipient lists for COVID-19 financial relief, as analyzed in Section 6. 9.2 Thailand and Personal Data Privacy Thailand’s national development plan, “Thailand 4.0”, envisions a digitally connected society that leverages big data and advanced technologies under the banner of “Digital for All” [64]. This vision necessarily involves the digitization of personal information across numerous sectors. However, challenges persist in balancing technological advancement with robust privacy protections. Recent research by Kulrujiphat and Wuttidittachotti (2024) found that public trust in Thailand’s e-Government services remains moderate, with concerns over personal data breaches cited as the top issue [65]. These findings reinforce our call in Section 7 for stronger infrastructure, centralized access control, and clearer policy enforcement. Ramasoota and Panichpapiboon highlight key barriers to privacy in Thailand, including national security narratives, weak information practices, and overuse of cybercrime laws for surveillance. Despite the PDPA, these issues continue to erode public trust and legal effectiveness [66]. Compliance with the PDPA also remains a significant hurdle. Chatsuwan et al. (2023) surveyed Thai SMEs and found widespread deficiencies in their privacy practices, including inadequate privacy policies and limited understanding of PDPA [67]. These findings reflect a broader gap between

Suphannee Sivakorn1 , Sasawat Malaivongs2 , Nuttaya Rujiratanapat1

regulatory intent and implementation, a gap that our study underscores through real-world evidence of data exposure.

9.3 Search Engine and Personal Data Exposure Search engines have long been recognized as powerful tools for uncovering inadvertently exposed documents containing sensitive personal data [68–71]. Our study confirms this phenomenon within the Thai web ecosystem. We show that National ID numbers and other personal details are retrievable via carefully crafted search queries for personal data fields. These findings align with previous research that explored search engine leakage across various platforms. For example, a case study involving Scribd, a public document-sharing platform, revealed the exposure of millions of personal records including passport numbers, birth certificates, and phone numbers by using targeted keyword searches refined with Boolean operators [72]. Similarly, data leaks have been identified through search queries within social networking platforms [71, 73].

10 Conclusion The findings of this study underscore the critical risks associated with the online exposure of personal data, particularly the Thai National ID Number, which serves as a fundamental component of identity verification in both governmental and commercial transactions. Over the course of a threemonth data collection period, our research identified 1.2 million exposed National ID numbers, affecting nearly 2% of Thailand’s population. This significant data exposure raises serious concerns regarding identity theft, financial fraud, and other forms of cybercrime, emphasizing the need for stronger data protection measures. A key discovery of this study is that the majority of these exposures originate from government sector websites, with domains under the .go.th TLD accounting for 58.8% of all leaked National ID numbers. Further analysis revealed that even non-government TLDs (.com, .org, .in.th) were often linked to government-affiliated entities, particularly local administrative bodies, and academic institutions. Additionally, sensitive personal details including phone numbers, addresses, financial data, and even health-related information, were found to be publicly accessible, increasing the potential for malicious exploitation. The implications of this study call for immediate action from the Thai government and relevant stakeholders to strengthen data governance frameworks, enhance cybersecurity protocols, and enforce stricter compliance with the PDPA. Government agencies, in particular, must implement robust data protection measures to prevent unauthorized disclosures. Furthermore, increased public awareness and education on

Analysis of Personal Data Exposure in Thailand

digital security best practices are essential to reducing the likelihood of identity fraud and other cyber threats. Ultimately, this research serves as a call to action for policymakers, regulatory bodies, and cybersecurity professionals to collaborate in addressing Thailand’s data privacy challenges. By implementing proactive measures such as automated monitoring systems, improved access controls, and stronger legal enforcement, the risk of personal data breaches can be significantly reduced, safeguarding the privacy and security of Thai citizens in an increasingly digital world. 11 Funding The authors received financial support for this study. References 1. Predescu, P.A., Bălan, D.: The implications and effects of data leaks. In: Proceedings of the International Conference on Cybersecurity and Cybercrime, vol. 10, pp. 170–177 (2023) 2. Clough, J.: Data theft? cybercrime and the increasing criminalization of access to data. In: Criminal Law Forum, vol. 22, pp. 145–170. Springer (2011) 3. Personal Data Protection Act, B.E. 2562 (2019) (2019). URL https://www.etda.or.th/th/Useful-Res ource/law/pdpa.aspx 4. Zulfiani, Y.N.: Prevention of personal data privacy leakage in e-government, as the government’s responsibility. Annals of Justice and Humanity 1(1), 29–37 (2021) 5. Vijay, P., Ramesh, P., Keshav I, S.: A survey of ethics in aadhaar, cybersecurity, and healthcare data. In: International Conference on Recent Trends in Machine Learning, IOT, Smart Cities & Applications, pp. 437–454. Springer (2024) 6. Sadhya, D., Sahu, T.: A critical survey of the security and privacy aspects of the aadhaar framework. Computers & Security 140, 103,782 (2024). DOI https://doi.org/10.1 016/j.cose.2024.103782. URL https://www.scienc edirect.com/science/article/pii/S016740482 400083X 7. Krishnamurthy, B., Wills, C.E.: On the leakage of personally identifiable information via online social networks. In: Proceedings of the 2nd ACM Workshop on Online Social Networks, WOSN ’09, p. 7–12. Association for Computing Machinery, New York, NY, USA (2009). DOI 10.1145/1592665.1592668. URL https://doi.org/10.1145/1592665.1592668 8. Holtfreter, R.E., Harrington, A.: Data breach trends in the united states. Journal of Financial Crime 22(2), 242–260 (2015) 9. European Commission: Data protection in the EU (n.d.). URL https://commission.europa.eu/law/law -topic/data-protection_en

23

10. OneTrust DataGuidance: Comparing privacy laws: GDPR v. Thai Personal Data Protection Act (2024). URL https://www.dataguidance.com/sites/def ault/files/gdpr_v_thailand_updated.pdf 11. Buabkhom, A.: The Technological Crimes : Law and Integrative Practical Prevention Strategies. Journal of Roi Kaensarn Academi 8(12), 726–743 (2023). URL https://so02.tci-thaijo.org/index.php/JRKS A/article/view/266556 12. Kompetch Krongkrachang: Bank of Thailand – Global Online Financial Scams Compilation (2023). URL ht tps://www.bot.or.th/th/research-and-publi cations/articles-and-publications/bot-mag azine/Phrasiam-66-3/globaltrend_financial fraud.html 13. Resecurity: Cybercriminals leaked massive volumes of stolen PII data from Thailand in Dark Web (2024). URL https://www.resecurity.com/blog/article/cy bercriminals-leaked-massive-volumes-of-s tolen-pii-data-from-thailand-in-dark-web 14. DataBreaches.Net: Thai loyalty membership card data of 5 million customers put up for sale on hacking forum (2024). URL https://databreaches.net/2024/ 11/20/thai-loyalty-membership-card-data-o f-5-million-customers-put-up-for-sale-o n-hacking-forum/ 15. Money & Banking Online: The Thai Bankers Association rushes to find the root cause. Bank employees sell customer information call for confidence (2024). URL https://en.moneyandbanking.co.th/2024/9191 1/ 16. Amarin TV: Call center scam victim weeps, wants to die after being deceived out of 2.5 million Baht. (2024). URL https://www.amarintv.com/news/social/2 27071 17. Kom Chad Luek: 34-year-old man takes his own life, leaves a farewell letter demanding the arrest of the call center scam gang (2024). URL https://www.komcha dluek.net/news/crime/582613 18. Emergency Decree on Measures for the Prevention and Suppression of Technological Crimes B.E. 2566 (2023) (2023). URL https://mdes.go.th/law/detail/ 7455-Emergency-Decree-on-Measures-for-the -Prevention-and-Suppression-of-Technologic al-Crimes--B-E--2566--202319. Kowit Somwaiya, U.U.a.: Thailand Issues Law to Combat Technology Crimes (2024). URL https://www.lawp lusltd.com/2023/04/thailand-issues-law-t o-combat-technology-crimes 20. FBI.gov: Money Mules (n.d.). URL https://www.fb i.gov/how-we-can-help-you/scams-and-safet y/common-frauds-and-scams/money-mules

24

21. Thai PBS: How to safely use your id card – what happens if your information gets leaked? (2025). URL https: //www.thaipbs.or.th/news/content/347899 22. Digital Government Development Agency (DGA): Government Data Exchange Center : GDX (n.d.). URL https://kb.dga.or.th/gdx/1about/ 23. Alkhalil, Z., Hewage, C., Nawaf, L., Khan, I.: Phishing attacks: A recent comprehensive study and a new anatomy. Frontiers in Computer Science 3, 563,060 (2021) 24. Amarikwa, M.: Internet openness at risk: Generative ai’s impact on data scraping. Rich. JL & Tech. 30, 533 (2023) 25. Weatherbed, J.: Google confirms it’s training Bard on scraped web data, too. The Verge (2023). URL https: //www.theverge.com/2023/7/5/23784257/googl e-ai-bard-privacy-policy-train-web-scrap ing 26. White, J.: How Strangers Got My Email Address From ChatGPT’s Model. The New York Times (2023). URL https://www.nytimes.com/interactive/2023/1 2/22/technology/openai-chatgpt-privacy-e xploit.html 27. Sirisuriya, S.D.S.: Importance of web scraping as a data source for machine learning algorithms - review. In: 2023 IEEE 17th International Conference on Industrial and Information Systems (ICIIS), pp. 134–139 (2023). DOI 10.1109/ICIIS58898.2023.10253502 28. Grynbaum, M.M., Mac, R.: The Times Sues OpenAI and Microsoft Over A.I. Use of Copyrighted Work. The New York Times (2023). URL https://www.nytime s.com/2023/12/27/business/media/newyork-t imes-open-ai-microsoft-lawsuit.html 29. Layne, R.: How to Make AI ’Forget’ All the Private Data It Shouldn’t Have. Harvard Business School (2024). URL https://www.library.hbs.edu/working-k nowledge/qa-seth-neel-on-machine-unlearn ing-and-the-right-to-be-forgotten 30. Snyder, A.: Machine forgetting: How difficult it is to get AI to forget. Axios. URL https://www.axios.com/ 2024/01/12/ai-forget-unlearn-data-privacy 31. Google Search Central: Overview of crawling and indexing topics. Google (2025). URL https://developers .google.com/search/docs/crawling-indexing 32. Miller, M.: Using Google Advanced Search. Que Publishing (2011) 33. Microsoft: Advanced search options. Microsoft (n.d.). URL https://support.microsoft.com/en-us/ topic/advanced-search-options-b92e25f1-0 085-4271-bdf9-14aaea720930 34. Microsoft: Advanced search keywords. Microsoft (n.d.). URL https://support.microsoft.com/en-us/ topic/advanced-search-keywords-ea595928-5 d63-4a0b-9c6b-0b769865e78a 35. IETF: RFC 1035 – DOMAIN NAMES - IMPLEMENTATION AND SPECIFICATION (1987). URL https:

Suphannee Sivakorn1 , Sasawat Malaivongs2 , Nuttaya Rujiratanapat1

//datatracker.ietf.org/doc/html/rfc1035 36. ISO: ISO 3166 Country Codes (1974). URL https: //www.iso.org/iso-3166-country-codes.html 37. THNiC: .th Internet country code top-level domain (ccTLD) for Thailand (2020). URL https://www. thnic.or.th/en/th-internet-country-code-t op-level-domain-cctld-for-thailand/ 38. Seitz, L.: The Most Popular Search Engine (What Is It?) (2024). URL https://www.broadbandsearch.net/ blog/most-popular-internet-search-engine 39. Fenniak, M.: PyPDF2 (n.d.). URL https://pypdf2.r eadthedocs.io/en/3.x/ 40. Singer-Vine, J.: Github – pdfplumber (2025). URL https://github.com/jsvine/pdfplumber 41. madmaze: Python-tesseract – pytesseract 0.3.13 (2024). URL https://pypi.org/project/pytesseract/ 42. pandas: pandas documentation (2024). URL https: //pandas.pydata.org/docs/reference/api/pan das.read_excel.html 43. Shah, A.: Github – python-docx2txt (2025). URL https: //github.com/ankushshah89/python-docx2txt 44. Malmgren, D.: textract (2024). URL https://textra ct.readthedocs.io/en/stable/ 45. gentoo linux: antiword (2023). URL https://wiki.g entoo.org/wiki/Antiword 46. Olivier de Segonzac: Search Engine Land: Inside SearchGuard: How Google detects bots and what the SerpAPI lawsuit reveals (2026). URL https://searchengine land.com/inside-google-searchguard-467676 47. Google: Google for Developers – Custom Search (2026). URL https://developers.google.com/custom -search/v1/overview 48. The Bureau of Registration Administration of Thailand (BORA): Official statistics registration systems (2024). URL https://stat.bora.dopa.go.th/stat/stat new/statyear/ 49. OWASP: OWASP Top 10:2025 (2025). URL https: //owasp.org/Top10/2025/ 50. System and Strategic Information Management Division, Ministry of Higher Education, Science, Research and Innovation: Higher Education Institution Data (2024). URL https://info.mhesi.go.th/homestat_acad emy.php 51. Lapwattanaworakul, J., Srisa-An, C., Angsirikul, S.: Guideline for data anonymization for data privacy in thailand. In: 2022 6th International Conference on Information Technology (InCIT), pp. 211–215 (2022). DOI 10.1109/InCIT56086.2022.10067859 52. Google: Google Search Central – Block Search indexing with noindex (2025). URL https://developers.g oogle.com/search/docs/crawling-indexing/bl ock-indexing

Analysis of Personal Data Exposure in Thailand

53. Data Privacy Manager: 20 biggest GDPR fines so far [2025] (2025). URL https://dataprivacymanager .net/5-biggest-gdpr-fines-so-far-2020/ 54. Google: Google Search Help – Remove web results from Google Search (n.d.). URL https://support.goog le.com/websearch/answer/11080680?hl=en 55. Safety Net Project: Removing Sensitive Content from the Internet (2022). URL https://www.techsafety .org/removing-sensitive-content 56. Bailey, M., Kenneally, E., Maughan, D., Dittrich, D.: The Menlo Report . IEEE Security & Privacy 10(02), 71–75 (2012). DOI 10.1109/MSP.2012.52. URL https://doi.ieeecomputersociety.org/10.110 9/MSP.2012.52 57. Adnan, M., Ghazali, M., Othman, N.Z.S.: E-participation within the context of e-government initiatives: A comprehensive systematic review. Telematics and Informatics Reports 8, 100,015 (2022). DOI https://do i.org/10.1016/j.teler.2022.100015. URL https: //www.sciencedirect.com/science/article/pi i/S2772503022000135 58. Chaipipat, S.: Asean governance on data privacy: challenges to regional protection of data privacy and personal data in cyberspace. Independent study, Chulalongkorn University (2019). DOI https://doi.or g/10.58837/CHULA.IS.2019.80. URL https: //digital.car.chula.ac.th/chulaetd/6946/ 59. Wang, Z., Hu, F., Su, J., Lin, Y.: Information source characteristics of personal data leakage during the covid19 pandemic in china: Observational study. JMIR Med Inform 12, e51,219 (2024). DOI 10.2196/51219 60. Andani, S.R.: Analysis of information security in data leaks in the pedulilindungi application. Int J Informatics Comput Sci 5(3), 246–9 (2021) 61. Ravizki, E.N.: Criminal liability in cases of personal data leakage amid the covid-19 pandemic. Nusantara Science and Technology Proceedings pp. 91–95 (2022) 62. Malhotra, S.: India investigates alleged leak of personal data from covid vaccination database. BMJ 381 (2023). DOI 10.1136/bmj.p1407. URL https://www.bmj.co m/content/381/bmj.p1407 63. Jung, G., Lee, H., Kim, A., Lee, U.: Too much information: Assessing privacy risks of contact trace data disclosure on people with covid-19 in south korea. Frontiers in Public Health 8 (2020). DOI 10.3389/fpubh.2020.00305. URL https://www.frontiersin.org/journals/p ublic-health/articles/10.3389/fpubh.2020. 00305 64. Ministry of Industry: THAILAND 4.0 The Next Revolution (2017). URL https://www.industry.go.th/w

25

eb-upload/1xff0d34e409a13ef56eea54c52a2911 26/m_magazine/12668/373/file_download/b29e 16008a87c72b354efebef853a428.pdf 65. Kulrujiphat, S., Wuttidittachotti, P.: A guideline of security and privacy improvement for promoting a thai government e-service digital trust. In: 2024 Research, Invention, and Innovation Congress: Innovative Electricals and Electronics (RI2C), pp. 5–8 (2024). DOI 10.1109/RI2C64012.2024.10784415 66. Ramasoota, P., Panichpapiboon, S.: Online privacy in thailand: Public and strategic awareness. Journal of Law, Information and Science 23(1), [97]–136 (2014). URL https://search.informit.org/doi/10.3316/ie lapa.347924655938976 67. Chatsuwan, P., Phromma, T., Surasvadi, N., Thajchayapong, S.: Personal data protection compliance assessment: A privacy policy scoring approach and empirical evidence from Thailand’s SMEs. Heliyon 9(10) (2023) 68. Xu, Y., Wang, K., Zhang, B., Chen, Z.: Privacy-enhancing personalized web search. In: Proceedings of the 16th International Conference on World Wide Web, WWW ’07, p. 591–600. Association for Computing Machinery, New York, NY, USA (2007). DOI 10.1145/1242572.12 42652. URL https://doi.org/10.1145/1242572. 1242652 69. Khan, R., Islam, M.A., Ullah, M., Aleem, M., Iqbal, M.A.: Privacy exposure measure: a privacy-preserving technique for health-related web search. Journal of Medical Imaging and Health Informatics 9(6), 1196– 1204 (2019) 70. Zimmer, M.: The externalities of search 2.0: The emerging privacy threats when the drive for the perfect search engine meets web 2.0. First Monday 13(3) (2008). DOI 10.5210/fm.v13i3.2136. URL https://firstmonday. org/ojs/index.php/fm/article/view/2136 71. Guo, G., Yang, T., Liu, Y.: Search engine based proper privacy protection scheme. IEEE Access 6, 78,551– 78,558 (2018). DOI 10.1109/ACCESS.2018.2885073 72. Adeeb Abdul Rahim, M.A., Mohamad, A.M., Kamaruddin, S., Wan Rosli, W.R.: Data leaks through public digital document libraries: A growing concern in relation to personal data protection and cyber security regulations. In: 2024 7th International Conference on Internet Applications, Protocols, and Services (NETAPPS), pp. 1–6 (2024). DOI 10.1109/NETAPPS63333.2024.10823567 73. Onete, C.B., Vargas, V.M., Chita, S.D.: Study on the implications of personal data exposure on the social media platforms. Transformations in Business & Economics 19(2) (2020)

Record · ID 138861 · SHA-256 e7261cbd5a9c1301
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.